OPEN SOURCE Strata Engine by Niko1221 • Run 125B+ MoE Models (Qwen3.8-Flash-Next) on 12GB VRAM GPUs!
⚡
STRATA LLM
Debugging & Performance Tuning

Troubleshooting Guide

Encountering errors while compiling sparse kernels or loading 125B weights? Consult these verified resolutions for hardware edge-cases:

1. CUDA Out of Memory (OOM) During Long Context Ingestion

Symptom: Engine crashes with "RuntimeError: CUDA out of memory. Tried to allocate 1.80 GiB"

Root Cause: The KV cache expands with context length and exceeded remaining VRAM headroom on a 12GB GPU.

Resolution Checklist:
  • Reduce context window from 128K to 32K or 64K using: --max-context 65536
  • Enable FP8 KV cache quantization: --kv-cache-dtype fp8
  • Lower VRAM allocation threshold to reserve 1.5GB for Windows DWM: --vram-utilization 0.85

2. Generation Stutter or Slow Speed (< 6 Tokens/sec)

Symptom: Tokens generate slowly with periodic 1-2 second freezes between sentences.

Root Cause: System RAM bandwidth bottleneck or weights stored on a mechanical hard drive/slow SATA SSD.

Resolution Checklist:
  • Ensure the model file resides on a high-speed PCIe 3.0/4.0 NVMe SSD (M.2 drive).
  • Verify in BIOS that RAM is operating in Dual Channel mode with XMP / EXPO memory profiles enabled (minimum DDR4-3200 or DDR5-5600). Single channel memory halves transfer bandwidth.
  • Close background browsers or RAM-heavy applications to prevent Windows from paging model weights to disk.

3. Windows VirtualAlloc / Pagefile Out of Memory Error

Symptom: Console outputs: "OSError: [WinError 1455] The paging file is too small for this operation to complete."

Root Cause: Windows default automatic pagefile size is too small to back 64GB+ of mapped tensor memory.

Resolution Checklist:
  • Open Windows System Properties -> Advanced -> Performance Settings -> Advanced -> Virtual Memory.
  • Uncheck "Automatically manage paging file size for all drives".
  • Select your fastest NVMe SSD and set Custom Size: Initial Size: 32768 MB (32GB), Maximum Size: 65536 MB (64GB).
  • Click "Set" and restart your computer.

4. WSL2 Fails to Detect NVIDIA GPU

Symptom: Running nvidia-smi inside Ubuntu WSL2 reports "No devices were found".

Root Cause: Installing native Linux display drivers inside WSL2 breaks the Microsoft D3D12/DirectX virtual passthrough layer.

Resolution Checklist:
  • Only install the latest Game Ready or Studio Driver on the Windows host machine.
  • Do NOT run apt-get install nvidia-driver inside WSL2.
  • Install only the CUDA toolkit user-mode libraries: apt-get install -y cuda-toolkit-12-x

Need Additional Developer Support?

Submit bug reports or pull requests directly to Niko1221's official repository on GitHub.