Local AI Wizard
Image models
51 open-weight entries in the Artificial Analysis snapshot checked 2026-09-22, with deeper technical review where primary evidence is available.
i
The complete ranked set comes from one Artificial Analysis arena snapshot per modality, so ranks and Elo scores are comparable only inside that modality. API variants can share the same public checkpoint.
A reviewed badge means license, components or hardware were checked against primary sources. Benchmark-only entries remain visible for coverage, with unknown facts clearly marked instead of inferred.
Top image models in this snapshot
Search, filters & my hardware 51/51
51 models · 51 ranked · 9 technical profiles reviewed · benchmark snapshot ↗ 2026-09-22
Benchmark score vs documented GPU memory
19 of 51 visible models have a documented memory path. Open the rankings for every model.
Compare selected models
| Field |
|---|
| Task |
| Parameters |
| Repository size |
| Components |
| Hardware |
| License |
Choose models to compare.
Video models
7 open-weight entries in the Artificial Analysis snapshot checked 2026-09-22, with deeper technical review where primary evidence is available.
i
The complete ranked set comes from one Artificial Analysis arena snapshot per modality, so ranks and Elo scores are comparable only inside that modality. API variants can share the same public checkpoint.
A reviewed badge means license, components or hardware were checked against primary sources. Benchmark-only entries remain visible for coverage, with unknown facts clearly marked instead of inferred.
Top video models in this snapshot
Search, filters & my hardware 13/13
13 models · 7 ranked · 6 technical profiles reviewed · benchmark snapshot ↗ 2026-09-22
Benchmark score vs documented GPU memory
2 of 13 visible models have a documented memory path. Open the rankings for every model.
Compare selected models
| Field |
|---|
| Task |
| Parameters |
| Repository size |
| Components |
| Hardware |
| License |
Choose models to compare.
Audio models
16 open-weight entries in the Artificial Analysis snapshot checked 2026-09-22, with deeper technical review where primary evidence is available.
i
The complete ranked set comes from one Artificial Analysis arena snapshot per modality, so ranks and Elo scores are comparable only inside that modality. API variants can share the same public checkpoint.
A reviewed badge means license, components or hardware were checked against primary sources. Benchmark-only entries remain visible for coverage, with unknown facts clearly marked instead of inferred.
Top audio models in this snapshot
Search, filters & my hardware 19/19
19 models · 16 ranked · 3 technical profiles reviewed · benchmark snapshot ↗ 2026-09-22
Benchmark score vs documented GPU memory
8 of 19 visible models have a documented memory path. Open the rankings for every model.
Compare selected models
| Field |
|---|
| Task |
| Parameters |
| Repository size |
| Components |
| Hardware |
| License |
Choose models to compare.
Explore your hardware
Explore high-scoring models within your memory budget.
Configuration · Q4 · 8,192 tokens · llama.cpp
llama.cpp setup documentation · Needs a supported architecture and GGUF checkpoint; choose the correct hardware backend. Choosing a runtime does not verify model compatibility. System RAM does not evaluate loading peaks or offload.
Start with a capacity
Select a preset or enter a custom capacity to reveal the strongest scored open models that fit.
Check local speed · import a benchmark
Run an optional test on your own machine, then import the JSON file. This website does not run inference or upload the file. Results remain in this tab and are not automatically matched to recommendations.
llama-bench -m /path/to/model.gguf -p 512 -n 128 -d 8192 -r 5 -o json > benchmark.jsonReplace the example path with an existing GGUF checkpoint. Requires llama-bench installed locally. Tokenization and sampling are excluded from its timing.
Browse published local measurements
Published measurements
Reported, not independently reproduced
These measurements belong to the exact checkpoint, quantization, backend, context and hardware shown. They are reference results, not a speed guarantee for another model ID, preset or runtime.
| Model / checkpoint | Hardware | Reported decode | Protocol | Source |
|---|---|---|---|---|
gpt-oss 20B MXFP4 MoEggml-org/gpt-oss-20b-GGUF | DGX Spark · 1× NVIDIA GB10MXFP4 MoE · CUDA | 83.43 ± 0.59tok/s · Not reported by llama-bench | 0 in → 32 outcontext depth=0; n_prompt=0 (tg32 row)Protocol & environment
| llama.cpp bench ↗ |
Qwen3 Coder 30B A3B Q8_0ggml-org/Qwen3-Coder-30B-A3B-Instruct-Q8_0-GGUF | DGX Spark · 1× NVIDIA GB10Q8_0 · CUDA | 61.06 ± 0.23tok/s · Not reported by llama-bench | 0 in → 32 outcontext depth=0; n_prompt=0 (tg32 row)Protocol & environment
| llama.cpp bench ↗ |
Qwen3.8 27B NVFP4 + MTPInferact/Qwen3.8-27B-NVFP4 | DGX Spark · 1× GB10NVFP4 + MTP (3 speculative tokens) · vLLM official | 18.5tok/s · TTFT 341 ms · C1 | 128 in → 128 outmax_model_len=262,144; KV dtype=FP8Protocol & environment
| NVIDIA Developer Forum ↗ |
Qwen3.8 27B NVFP4 + DSparkRadixArk/Qwen3.8-27B-NVFP4 | DGX Spark · 1× GB10NVFP4 + DSpark · SGLang | 36.6tok/s · TTFT 246 ms · C1 | 128 in → 128 outmax_model_len not stated in this recipe excerpt; do not inferProtocol & environment
| NVIDIA Developer Forum ↗ |
Reported results are kept separate from planning estimates. A missing match means “not measured here”, not zero performance.
Text models
LiveBench · 2026-06-25 · Scores for 22 open models · Checked 2026-09-22
Start with quality. Explore the hardware when you need it.
Choose models 8/36
Choose hardware · 4/6
8 models in this view
Swipe the chart sideways, or use Table for a compact comparison.
Hardware by memory capacity
32 GB
128 GB
256 GB
Take a closer look
Ten checkpoint profiles, five hardware configurations and three task comparisons. Each page keeps its source, date and limits visible.
Models
Qwen3.5 9Btext + image + videoQwen3.8 27Btext + image + videoQwen3.5 35B A3Btext + image + videoDeepSeek V4 Flash · 0731textGLM-5.3 Flashtext + imageKimi K2.6 · thinkingtext + imageFLUX.2 [dev]imageQwen-ImageimageWan2.1 T2V 1.3BvideoStable Audio Open SmallaudioSources, methodology & data status
01 / LiveBench
The overall score is the equally weighted mean of seven category means (0–100). All scores use one release. This is not the Artificial Analysis Intelligence Index; the two scales must not be mixed.
CSV ↗ · Categories ↗ · LiveBench ↗ · Data reuse policy ↗02 / Memory estimate
Planning formula: total parameters × 0.625 bytes, plus 20% working allowance and 8 GB reserve, rounded up in decimal GB. This is a heuristic, not a measured minimum or guarantee. Quantization, context, KV cache, vision components, runtime and offloading change requirements. MoE uses total parameters, not only active ones.
Qwen counts refer to language parameters; GLM-5.3 and DeepSeek V4 use repository tensor totals; GLM-5.3 Flash uses the manufacturer’s 320B specification. Architectures without a reviewed estimate remain outside the memory axis.
03 / Reference, not a local test
LiveBench evaluates the named model and reasoning effort. Its score is not a measurement of the local Q4 variant. Hardware links are official destinations without affiliate tracking. Stock, exact configuration and regional availability can change.
04 / Automatic updates without paid APIs
The deployed product checks scores, approved official model organizations, US prices and stock daily, and hardware specifications weekly. Due visits can also refresh safely. D1 preserves the last valid snapshot. Incompatible methodology is quarantined; unavailable sources never erase valid data. New models remain pending benchmark or insufficient data until official facts exist.