RUN 125B+ MoE ON CONSUMER GPUS
Powered by Niko1221's dynamic sparse router offloading. Run server-grade frontier models like Qwen3.8-Flash-Next 125B on a single 12GB NVIDIA graphics card and 64GB system RAM—with zero cloud dependencies.
git clone https://github.com/Niko1221/Strata.git
cd Strata
.\START-HERE.bat How Strata Solves the VRAM Wall
Traditional inference engines crash when a 125B model exceeds GPU memory. Strata utilizes the sparse nature of Mixture-of-Experts architectures to stream only active weights:
Predictive Sparse Routing
In MoE models like Qwen 125B, only 2-8 out of 64 experts fire per token. Strata predicts which expert layers will be called 1-2 tokens in advance, queuing them for instant DMA PCIe transfer.
Tiered VRAM-RAM Cache
Critical attention heads and frequently triggered experts reside permanently inside your 12GB GPU VRAM. The remaining 90GB of dormant weights stream through high-speed 64GB DDR5 memory.
Zero-Copy Pinned Memory
Weights are transferred over asynchronous CUDA streams via PCIe 4.0/5.0 directly into tensor cores, completely bypassing CPU compute cycles and preventing stutter.
System Requirements Matrix
- Qwen3.8-Flash-Next 125B (4-bit Q_K)
- DeepSeek-V2-Lite 16B
- Qwen3.8-Flash-Next 125B (5-bit Q_M)
- Mixtral 8x22B Instruct
- Full-context 125B-236B MoE Models
- Continuous Batching Serving
Engine Performance Comparison
Testing Qwen3.8-Flash-Next (125 Billion Parameters) on a standard single RTX 4070 (12GB VRAM) paired with 64GB DDR5:
| Inference Engine | Hardware Target | Tokens / Sec | RAM Allocation | Status & Stability |
|---|---|---|---|---|
| ⚡ Strata LLM (Niko1221) | RTX 4070 (12GB) + 64GB RAM | 21.4 tok/s | 52 GB | Native Support (Zero OOM Crash) |
| Ollama (Standard) | RTX 4070 (12GB) + 64GB RAM | Fail / OOM | Exceeds limit | Out of Memory / Requires Quantization Drop |
| vLLM (CPU-Offload) | RTX 4070 (12GB) + 64GB RAM | 3.8 tok/s | 58 GB | Severe PCIe Bus Bottleneck |
Frequently Asked Questions
Everything you need to know about setting up Strata and configuring MoE quantization.