Default path
Qwen3.8 27B under local, single-user, single-stream use with stable MTP/speculative draft treated as the normal path. SGLang and DFlash/DFlash2 stay outside the default median.
ModelDock research
Your workload
Reading is the model taking in your prompt, files, and chat history. Writing is the answer it generates. Long documents and coding sessions are usually read-heavy, so the default starts at 75% reading / 25% writing. Stable MTP is treated as normal, not a separate scenario.
Model Quantization
2× GPU build note Dual-card rigs require chassis clearance and slot spacing for wide cards, stable high-headroom power delivery and airflow, plus more involved installation and multi-GPU scheduling than a single-card node.
| Hardware | Capability | Economics for your workload | Evidence | Market | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Assumptions
75% reading / 25% writing · 24 active h/day
Whole-node power · Qwen3.8 27B · ordinary local use
Transparent calculation
total GPU rig = GPU price + $1,000 hostlocal $/MM = (energy + hardware allocation) / daily capacitybreak-even days = total rig / (API value/day − energy/day)Sequential single-user model: reading time + writing time; no concurrency or overlap assumed.
Evidence data
Open the source spread behind published hardware records.
Methodology
Performance uses a source-visible MTP-ready median. The economics below are adjustable because they describe your usage—not the benchmark run.
Qwen3.8 27B under local, single-user, single-stream use with stable MTP/speculative draft treated as the normal path. SGLang and DFlash/DFlash2 stay outside the default median.
Reading (prefill) and writing (decode) are medianed independently. Source conditions—quant, context, runtime, topology, and draft depth—remain visible after clicking Median.
Median links to the source spread. ~ Rate estimate is bandwidth-scaled from measured and reported reading and writing data; it remains separate from source medians. Pending means no ordinary-lane source metric has been collected yet.
Discrete GPU totals use the dated GPU price plus a $1,000 host/platform allowance. Break-even divides total rig price by equivalent API value per day less direct energy cost; it is not an investment-return guarantee.