22 evaluated open-weight models in the catalog · 8 frontier references. The default shows a selection; search explores the full catalog.
Choose models 9/36
Choose hardware · 3/6Select a machine below the axis to explore its capacity and price.
9 models in this view
Swipe the chart sideways, or use Table for a compact comparison.
Higher = stronger benchmark result. Left = less estimated memory. Click a point for its source. The Y-axis adapts to the selection.
* USD, US market. Select a machine for the exact configuration, price basis and dated official source. Hardware photos are illustrative. PNG exports the chart and active capacity overlay.
Outside the scatter: DeepSeek V4.1 Flash · max (81.1; memory estimate pending)
Parameter basis: language component. Dense. Context can be extended up to 1,000,000 tokens with additional configuration; memory requirements change. Technical details checked 2026-09-22.
Where the model performs best
LiveBench 2026-06-25 · same tasks, 0–100 scale
80.0
75.7
61.4
86.2
76.6
74.3
72.7
Q4 weight estimates exclude runtime and cache. Planning memory adds a heuristic allowance; neither is a measured runtime requirement. Category results describe the exact benchmark variant, not a local quantization.
Sources, assumptions & limits
01 / LiveBench
The overall score is the equally weighted mean of seven category means (0–100). All scores use one release. This is not the Artificial Analysis Intelligence Index; the two scales must not be mixed.
Planning formula: total parameters × 0.625 bytes, plus 20% working allowance and 8 GB reserve, rounded up in decimal GB. This is a heuristic, not a measured minimum or guarantee. Quantization, context, KV cache, vision components, runtime and offloading change requirements. MoE uses total parameters, not only active ones.
Qwen counts refer to language parameters; GLM-5.3 and DeepSeek V4 use repository tensor totals; GLM-5.3 Flash uses the manufacturer’s 320B specification. Architectures without a reviewed estimate remain outside the memory axis.
03 / Reference, not a local test
LiveBench evaluates the named model and reasoning effort. Its score is not a measurement of the local Q4 variant. Hardware links are official destinations without affiliate tracking. Stock, exact configuration and regional availability can change.
04 / Updates without paid APIs
Existing results are checked when the page is visited and cached for up to 24 hours where supported. No inference is run. If a model disappears, a score shifts over 10 points, or the schema/release changes, the verified saved snapshot stays visible. New releases and hardware mappings require review.