CONFIGURATION / CAPACITY
DGX H100 · 8 × 80 GB
Eight 80 GB GPUs provide 640 GB in aggregate, not one 640 GB address space. Sharding, interconnect and runtime configuration are required for models exceeding one GPU. manufacturer specification ↗.
Text models within the planning estimate
These are capacity candidates under a generic Q4 formula, not measured executions. The total estimate does not separately prove KV cache, vision encoder, runtime support or multi-GPU placement.
| Model | Q4 planning | Quality | Checkpoint |
|---|---|---|---|
| Qwen3.5 9B | ~15 GB | — | source ↗ |
| Qwen3.8 27B | ~29 GB | 75.3 · LiveBench 2026-06-25 | source ↗ |
| Qwen3.5 35B A3B | ~35 GB | — | source ↗ |
| GLM-5.3 Flash | ~248 GB | 71.6 · LiveBench 2026-06-25 | source ↗ |
Documented media paths
These are publisher-documented CUDA memory paths; exact GPU architecture, ARM64, precision, offload, resolution and runtime may differ from this machine.
- FLUX.2 [dev] · 24 GB · Quantized consumer-GPU path with a remote text encoder · evidence ↗
- Wan2.1 T2V 1.3B · 8.19 GB · Official T2V 1.3B configuration · evidence ↗
Specification reviewed: 2026-09-24. No price or stock claim is made here.