OPEN SOURCE Strata Engine by Niko1221 • Run 125B+ MoE Models (Qwen3.8-Flash-Next) on 12GB VRAM GPUs!
⚡
STRATA LLM
MoE Architecture Catalog

Supported Models

Strata is optimized for sparse Mixture-of-Experts architectures where only a fraction of weights activate per token. Below are verified production models and memory allocation presets:

Alibaba Cloud / Qwen Team

Qwen3.8-Flash-Next (125B MoE)

64 Experts (8 Active per Token)

The primary model Strata was designed for. Rivals Claude 3.5 Sonnet on programming benchmarks while generating at 20-35 tok/s on an RTX 4070.

Total / Active 125 Billion (14 Billion)
Min VRAM 11.5 GB
Min System RAM 52 GB
Context Window 128K Tokens
Terminal Launch Command:
strata-cli run --model qwen3.8-125b-q4 --vram 11.5 --ram 52
DeepSeek AI

DeepSeek-V2-Lite MoE (16B / 236B)

Multi-head Latent Attention (MLA) + MoE

Extreme efficiency with Multi-head Latent Attention, reducing KV cache consumption by 70%.

Total / Active 16B / 236B (2.4B / 21B)
Min VRAM 10.0 GB
Min System RAM 38 GB
Context Window 64K Tokens
Terminal Launch Command:
strata-cli run --model deepseek-v2-lite --vram 10.0
Mistral AI

Mixtral 8x22B Instruct

8 Experts (2 Active per Token)

Strong reasoning and European multilingual capabilities. Recommended for RTX 4080 (16GB) or higher.

Total / Active 141 Billion (39 Billion)
Min VRAM 15.5 GB
Min System RAM 70 GB
Context Window 64K Tokens
Terminal Launch Command:
strata-cli run --model mixtral-8x22b-instruct --vram 15.5 --ram 70