Local AI Wizard

Text models

LiveBench · 2026-06-25 · Scores for 22 open models · Checked 2026-09-22

OPEN-WEIGHT LEADERS

Start with quality. Explore the hardware when you need it.

01
81.1Overall
Memory under reviewEstimated local memory
02
79.5Overall
~2,108 GBEstimated local memory
03
79.2Overall
~2,108 GBEstimated local memory
Choose models 8/36
Open weights
Frontier references · proprietary
Choose hardware · 4/6

8 models in this view

Swipe the chart sideways, or use Table for a compact comparison.

LiveBench × local memoryPROPRIETARY5060708090163264128256512102420484096Estimated memory · GBLogarithmic scaleDeepSeek V4 FlashGLM-5.3Smaug AgenticSmaug MiniNemotron 3 Ultra 550B A55BGPT-6 Astra 82.2GPT-5.6 Sol 81.1Claude Opus 5 80.1LiveBench 2026-06-25 · LiveBench · reference scores; Q4 memory estimatesChecked 2026-09-22 · livebench.airicardoguia.com

Hardware by memory capacity

32 GB

128 GB

256 GB

Memory is estimated · not a local performance test USD · US prices
SOURCE-BACKED GUIDES

Take a closer look

Ten checkpoint profiles, five hardware configurations and three task comparisons. Each page keeps its source, date and limits visible.

Sources, methodology & data status
Benchmark release: 2026-06-25Data checked: 2026-09-22Verified snapshot · checking source

01 / LiveBench

The overall score is the equally weighted mean of seven category means (0–100). All scores use one release. This is not the Artificial Analysis Intelligence Index; the two scales must not be mixed.

CSV ↗ · Categories ↗ · LiveBench ↗ · Data reuse policy ↗

02 / Memory estimate

Planning formula: total parameters × 0.625 bytes, plus 20% working allowance and 8 GB reserve, rounded up in decimal GB. This is a heuristic, not a measured minimum or guarantee. Quantization, context, KV cache, vision components, runtime and offloading change requirements. MoE uses total parameters, not only active ones.

Qwen counts refer to language parameters; GLM-5.3 and DeepSeek V4 use repository tensor totals; GLM-5.3 Flash uses the manufacturer’s 320B specification. Architectures without a reviewed estimate remain outside the memory axis.

03 / Reference, not a local test

LiveBench evaluates the named model and reasoning effort. Its score is not a measurement of the local Q4 variant. Hardware links are official destinations without affiliate tracking. Stock, exact configuration and regional availability can change.

04 / Automatic updates without paid APIs

The deployed product checks scores, approved official model organizations, US prices and stock daily, and hardware specifications weekly. Due visits can also refresh safely. D1 preserves the last valid snapshot. Incompatible methodology is quarantined; unavailable sources never erase valid data. New models remain pending benchmark or insufficient data until official facts exist.