← Lab / Open models

Open models. Real hardware.

How much memory does intelligence take? Compare open weights with the models you know.

LiveBench ↗Sources & methodology
Benchmark release: 2026-06-25Data checked: 2026-09-22Verified snapshot · checking source

22 evaluated open-weight models in the catalog · 8 frontier references. The default shows a selection; search explores the full catalog.

Choose models 9/36
Open weights
Frontier references · proprietary
Choose hardware · 3/6
Select a machine below the axis to explore its capacity and price.

9 models in this view

Swipe the chart sideways, or use Table for a compact comparison.

Intelligence × local memoryPROPRIETARY REFERENCES506070809016326412825651210242048Estimated memory · GB · logarithmic scaleQwen3.8 27BDeepSeek V4 FlashGLM-5.3 FlashGLM-5.3Nemotron 3 Ultra 550B A55BGPT-6 Astra · max 82.2GPT-5.6 Sol · max 81.1Claude Opus 5 · max 80.1LiveBench 2026-06-25 · reference scores; Q4 memory estimatesChecked 2026-09-22 · livebench.airicardoguia.com

Higher = stronger benchmark result. Left = less estimated memory. Click a point for its source. The Y-axis adapts to the selection.

* USD, US market. Select a machine for the exact configuration, price basis and dated official source. Hardware photos are illustrative. PNG exports the chart and active capacity overlay.

Outside the scatter: DeepSeek V4.1 Flash · max (81.1; memory estimate pending)

Alibaba / Open weights

Qwen3.8 27B

Official model card
LiveBench
75.3
Total parameters
27B
Disk space · Q4 weights estimate
~16.9 GB
Planning memory
~29 GB
Active parameters
27B
Documented context
262,144 tokens
Input → output
Text, image, video → text
License
Apache 2.0
Local performance test
Not measured

Parameter basis: language component. Dense. Context can be extended up to 1,000,000 tokens with additional configuration; memory requirements change. Technical details checked 2026-09-22.

Where the model performs best

LiveBench 2026-06-25 · same tasks, 0–100 scale

80.0
75.7
61.4
86.2
76.6
74.3
72.7

Q4 weight estimates exclude runtime and cache. Planning memory adds a heuristic allowance; neither is a measured runtime requirement. Category results describe the exact benchmark variant, not a local quantization.

Sources, assumptions & limits

01 / LiveBench

The overall score is the equally weighted mean of seven category means (0–100). All scores use one release. This is not the Artificial Analysis Intelligence Index; the two scales must not be mixed.

CSV ↗ · Categories · LiveBench ↗ · Data reuse policy

02 / Memory estimate

Planning formula: total parameters × 0.625 bytes, plus 20% working allowance and 8 GB reserve, rounded up in decimal GB. This is a heuristic, not a measured minimum or guarantee. Quantization, context, KV cache, vision components, runtime and offloading change requirements. MoE uses total parameters, not only active ones.

Qwen counts refer to language parameters; GLM-5.3 and DeepSeek V4 use repository tensor totals; GLM-5.3 Flash uses the manufacturer’s 320B specification. Architectures without a reviewed estimate remain outside the memory axis.

03 / Reference, not a local test

LiveBench evaluates the named model and reasoning effort. Its score is not a measurement of the local Q4 variant. Hardware links are official destinations without affiliate tracking. Stock, exact configuration and regional availability can change.

04 / Updates without paid APIs

Existing results are checked when the page is visited and cached for up to 24 hours where supported. No inference is run. If a model disappears, a score shifts over 10 points, or the schema/release changes, the verified saved snapshot stays visible. New releases and hardware mappings require review.