Default path
Qwen3.8 27B under local, single-user, single-stream use with stable MTP/speculative draft treated as the normal path. Source-marked short-context probes, SGLang, and DFlash/DFlash2 stay outside the default median.
ModelDock research
Your workload
Reading is the model taking in your prompt, files, and chat history. Writing is the answer it generates. Long documents and coding sessions are usually read-heavy, so the default starts at 75% reading / 25% writing. Stable MTP is treated as normal, not a separate scenario.
Model Quantization
2× GPU build note Dual-card rigs require chassis clearance and slot spacing for wide cards, stable high-headroom power delivery and airflow, plus more involved installation and multi-GPU scheduling than a single-card node.
| Hardware | Capability | Economics for your workload | Evidence | Market | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| RTX 309024GB · single-card node | 24 GB | 705 | 55.9 | Good | 500W | $1,350 | $2,350 | 15.6 | $0.198 | $1.80 | $2.73 | 2,524 d | Median ↗ | |
| RTX 3090 Ti24GB · single-card node | 24 GB | 1,458.3 | 69.8 | Good | 600W | $1,400 | $2,400 | 21.1 | $0.165 | $2.16 | $3.69 | 1,570 d | Median ↗ | |
| AMD Radeon RX 7900 XTX24GB GDDR6 | 24 GB | 877.8 | 43.9 | Good | 475W | $800 | $1,800 | 13.2 | $0.204 | $1.71 | $2.31 | 3,007 d | Median ↗ | |
| 2× RTX 5060 Ti16GB each · dual-card node | 32 GB | 723 | 65.9 | Good | 480W | $1,760 | $2,760 | 17.9 | $0.181 | $1.73 | $3.13 | 1,969 d | Median ↗ | |
| 2× RTX 5070 Ti16GB each · dual-card node | 32 GB | 3,250 | 113.3 | Good | 620W | $2,400 | $3,400 | 35.4 | $0.116 | $2.23 | $6.20 | 857 d | Median ↗ | |
| 2× RTX 408016GB each · dual-card node | 32 GB | ~1,700 | ~58 | Good | 760W | $1,900 | $2,900 | 18.2 | $0.238 | $2.74 | $3.18 | 6,501 d | Estimate | |
| RTX 409024GB · single-card node | 24 GB | 2,692.3 | 68.8 | Good | 550W | $2,500 | $3,500 | 22.1 | $0.176 | $1.98 | $3.86 | 1,858 d | Median ↗ | |
| RTX 509032GB · single-card node | 32 GB | 3,571.4 | 132.5 | Good | 700W | $4,200 | $5,200 | 41.2 | $0.130 | $2.52 | $7.21 | 1,109 d | Median ↗ | |
| Intel Arc Pro B7032GB GDDR6 | 32 GB | 424.8 | 49.3 | Good | 350W | $1,299 | $2,299 | 12.6 | $0.199 | $1.26 | $2.21 | 2,412 d | Median ↗ | |
| RTX PRO 5000 Blackwell48GB GDDR7 ECC | 48 GB | 6,288 | 112.5 | Good | 420W | $9,085 | $10,085 | 36.9 | $0.191 | $1.51 | $6.46 | 2,040 d | Median ↗ | |
| RTX PRO 6000 Blackwell96GB GDDR7 ECC | 96 GB | 7,756 | 109.3 | Good | 720W | $13,250 | $14,250 | 36.2 | $0.287 | $2.59 | $6.34 | 3,803 d | Median ↗ | |
| Mac Studio M3 Ultra 512 GB80-core · 512GB unified memory | 512 GB | 462.5 | 44.3 | Good | 330W | — | $9,499 | 11.9 | $0.538 | $1.19 | $2.08 | 10,657 d | Median ↗ | |
| Mac Studio M4 Max 64 GB64GB unified memory · current retail reference | 64 GB | ~274 | ~53.3 | Good | 145W | — | $3,499 | 11.6 | $0.210 | $0.52 | $2.04 | 2,312 d | Estimate | |
| Mac Studio M4 Max 128 GBsource-backed oMLX profile | 128 GB | — | 56.8 | Good | 145W | — | $6,500 | — | pending | — | — | — | Median ↗ | |
| MacBook Pro M5 Max 128 GB128GB unified memory | 128 GB | 745 | 59.2 | Good | 120W | — | $6,699 | 16.5 | $0.248 | $0.43 | $2.89 | 2,725 d | Median ↗ | |
| Mac mini M5 Pro 64 GB20-core GPU · 64GB unified memory | 64 GB | ~373 | ~29.6 | Good | ~110W | — | $2,899 | 8.3 | $0.240 | $0.40 | $1.45 | 2,762 d | Estimate | |
| Mac Studio M5 Max 64 GB40-core GPU · 64GB unified memory | 64 GB | ~745 | ~59.2 | Good | ~180W | — | $3,799 | 16.5 | $0.165 | $0.65 | $2.89 | 1,694 d | Estimate | |
| Mac Studio M5 Max 128 GB40-core GPU · 128GB unified memory | 128 GB | ~745 | ~59.2 | Good | ~180W | — | $5,399 | 16.5 | $0.218 | $0.65 | $2.89 | 2,407 d | Estimate | |
| Mac Studio M5 Ultra 96 GB30-core CPU · 64-core GPU · 96GB unified memory | 96 GB | ~1,075 | ~91 | Good | ~360W | — | $5,499 | 25.1 | $0.172 | $1.30 | $4.39 | 1,778 d | Estimate | |
| Mac Studio M5 Ultra 256 GB36-core CPU · 80-core GPU · 256GB unified memory | 256 GB | ~1,075 | ~91 | Good | ~360W | — | $10,799 | 25.1 | $0.288 | $1.30 | $4.39 | 3,492 d | Estimate | |
| DGX Spark128GB unified memory | 128 GB | 1,225.5 | 44 | Good | 170W | — | $4,699 | 13.7 | $0.232 | $0.61 | $2.40 | 2,623 d | Median ↗ | |
Assumptions
75% reading / 25% writing · 24 active h/day
Whole-node power · Qwen3.8 27B · ordinary local use
Transparent calculation
total GPU rig = GPU price + $1,000 hostlocal $/MM = (energy + hardware allocation) / daily capacitybreak-even days = total rig / (API value/day − energy/day)Sequential single-user model: reading time + writing time; no concurrency or overlap assumed.
Evidence data
Open the source spread behind published hardware records.
Methodology
Performance uses a source-visible MTP-ready median. The economics below are adjustable because they describe your usage—not the benchmark run.
Qwen3.8 27B under local, single-user, single-stream use with stable MTP/speculative draft treated as the normal path. Source-marked short-context probes, SGLang, and DFlash/DFlash2 stay outside the default median.
Reading (prefill) and writing (decode) are medianed independently. Source conditions—quant, context, runtime, topology, and draft depth—remain visible after clicking Median.
Median links to the source spread. ~ Rate estimate uses a disclosed conservative hardware ratio derived from measured and reported reading and writing data; it remains separate from source medians. Pending means no ordinary-lane source metric has been collected yet.
Discrete GPU totals use the dated GPU price plus a $1,000 host/platform allowance. Break-even divides total rig price by equivalent API value per day less direct energy cost; it is not an investment-return guarantee.