ModelDock research

Benchmark Data

Ordinary-lane hardware · source medians · adjustable economics

Your workload

Reading is the model taking in your prompt, files, and chat history. Writing is the answer it generates. Long documents and coding sessions are usually read-heavy, so the default starts at 75% reading / 25% writing. Stable MTP is treated as normal, not a separate scenario.

Model Quantization

21 listed hardware profilesOriginal catalog order · click a numeric column to sort

2× GPU build note Dual-card rigs require chassis clearance and slot spacing for wide cards, stable high-headroom power delivery and airflow, plus more involved installation and multi-GPU scheduling than a single-card node.

HardwareCapabilityEconomics for your workloadEvidenceMarket
RTX 309024GB · single-card node24 GB70555.9Good500W$1,350$2,35015.6$0.198$1.80$2.732,524 dMedian
RTX 3090 Ti24GB · single-card node24 GB1,458.369.8Good600W$1,400$2,40021.1$0.165$2.16$3.691,570 dMedian
AMD Radeon RX 7900 XTX24GB GDDR624 GB877.843.9Good475W$800$1,80013.2$0.204$1.71$2.313,007 dMedian
2× RTX 5060 Ti16GB each · dual-card node32 GB72365.9Good480W$1,760$2,76017.9$0.181$1.73$3.131,969 dMedian
2× RTX 5070 Ti16GB each · dual-card node32 GB3,250113.3Good620W$2,400$3,40035.4$0.116$2.23$6.20857 dMedian
2× RTX 408016GB each · dual-card node32 GB~1,700~58Good760W$1,900$2,90018.2$0.238$2.74$3.186,501 dEstimate
RTX 409024GB · single-card node24 GB2,692.368.8Good550W$2,500$3,50022.1$0.176$1.98$3.861,858 dMedian
RTX 509032GB · single-card node32 GB3,571.4132.5Good700W$4,200$5,20041.2$0.130$2.52$7.211,109 dMedian
Intel Arc Pro B7032GB GDDR632 GB424.849.3Good350W$1,299$2,29912.6$0.199$1.26$2.212,412 dMedian
RTX PRO 5000 Blackwell48GB GDDR7 ECC48 GB6,288112.5Good420W$9,085$10,08536.9$0.191$1.51$6.462,040 dMedian
RTX PRO 6000 Blackwell96GB GDDR7 ECC96 GB7,756109.3Good720W$13,250$14,25036.2$0.287$2.59$6.343,803 dMedian
Mac Studio M3 Ultra 512 GB80-core · 512GB unified memory512 GB462.544.3Good330W$9,49911.9$0.538$1.19$2.0810,657 dMedian
Mac Studio M4 Max 64 GB64GB unified memory · current retail reference64 GB~274~53.3Good145W$3,49911.6$0.210$0.52$2.042,312 dEstimate
Mac Studio M4 Max 128 GBsource-backed oMLX profile128 GB56.8Good145W$6,500pendingMedian
MacBook Pro M5 Max 128 GB128GB unified memory128 GB74559.2Good120W$6,69916.5$0.248$0.43$2.892,725 dMedian
Mac mini M5 Pro 64 GB20-core GPU · 64GB unified memory64 GB~373~29.6Good~110W$2,8998.3$0.240$0.40$1.452,762 dEstimate
Mac Studio M5 Max 64 GB40-core GPU · 64GB unified memory64 GB~745~59.2Good~180W$3,79916.5$0.165$0.65$2.891,694 dEstimate
Mac Studio M5 Max 128 GB40-core GPU · 128GB unified memory128 GB~745~59.2Good~180W$5,39916.5$0.218$0.65$2.892,407 dEstimate
Mac Studio M5 Ultra 96 GB30-core CPU · 64-core GPU · 96GB unified memory96 GB~1,075~91Good~360W$5,49925.1$0.172$1.30$4.391,778 dEstimate
Mac Studio M5 Ultra 256 GB36-core CPU · 80-core GPU · 256GB unified memory256 GB~1,075~91Good~360W$10,79925.1$0.288$1.30$4.393,492 dEstimate
DGX Spark128GB unified memory128 GB1,225.544Good170W$4,69913.7$0.232$0.61$2.402,623 dMedian
capability efficiency cost ~ hardware-scaled rate estimate or incomplete evidence

Assumptions

75% reading / 25% writing · 24 active h/day
Whole-node power · Qwen3.8 27B · ordinary local use

Transparent calculation

total GPU rig = GPU price + $1,000 host
local $/MM = (energy + hardware allocation) / daily capacity
break-even days = total rig / (API value/day − energy/day)Sequential single-user model: reading time + writing time; no concurrency or overlap assumed.

Evidence data

Open the source spread behind published hardware records.

Methodology

The conditions behind the table.

Performance uses a source-visible MTP-ready median. The economics below are adjustable because they describe your usage—not the benchmark run.

Default path

Qwen3.8 27B under local, single-user, single-stream use with stable MTP/speculative draft treated as the normal path. Source-marked short-context probes, SGLang, and DFlash/DFlash2 stay outside the default median.

Per-metric median

Reading (prefill) and writing (decode) are medianed independently. Source conditions—quant, context, runtime, topology, and draft depth—remain visible after clicking Median.

Evidence labels

Median links to the source spread. ~ Rate estimate uses a disclosed conservative hardware ratio derived from measured and reported reading and writing data; it remains separate from source medians. Pending means no ordinary-lane source metric has been collected yet.

Pricing basis

Discrete GPU totals use the dated GPU price plus a $1,000 host/platform allowance. Break-even divides total rig price by equivalent API value per day less direct energy cost; it is not an investment-return guarantee.

Market search

Find equivalent offers

Search links open on the indicated marketplace. No affiliate links, live-price feed, or seller ranking is used.