Local AI Wizard
CONFIGURATION / CAPACITY

RTX 5090 · 32 GB VRAM

One CUDA GPU with 32 GB dedicated VRAM. A model below 32 GB in the generic Q4 estimate is a planning candidate, not a tested fit. CPU offload needs system RAM and may reduce speed. manufacturer specification ↗.

Text models within the planning estimate

These are capacity candidates under a generic Q4 formula, not measured executions. The total estimate does not separately prove KV cache, vision encoder, runtime support or multi-GPU placement.

ModelQ4 planningQualityCheckpoint
Qwen3.5 9B~15 GB—source ↗
Qwen3.8 27B~29 GB75.3 · LiveBench 2026-06-25source ↗

Documented media paths

These are publisher-documented CUDA memory paths; exact GPU architecture, ARM64, precision, offload, resolution and runtime may differ from this machine.

Specification reviewed: 2026-09-24. No price or stock claim is made here.