OPEN SOURCE DEEP DIVE
Strata: Run a 125B-Parameter MoE Model on Your Gaming PC
Strata is an open-source local inference engine built on llama.cpp/ggml that runs the 125-billion-parameter Qwen3.8-Flash-Next MoE model on consumer gaming PCs with as little as 12GB VRAM, using tiered expert caching (GPU→RAM→SSD) and guess-and-check speculative decoding to achieve 53-94 tokens/second.
Overview
Strata is an open-source local inference engine that runs the 125-billion-parameter Qwen3.8-Flash-Next model on ordinary gaming PCs. Built on llama.cpp / ggml, it supports NVIDIA (RTX 20-50 series) and AMD (RX 6000-9000 series) GPUs with as little as 12 GB of VRAM. All inference happens locally — no data leaves your machine.
The Problem
A 125-billion-parameter MoE model normally requires server-grade hardware with hundreds of gigabytes of graphics memory. Strata's core innovation is distributing inference work across the entire PC — GPU, RAM, CPU, and SSD work together to make consumer hardware capable of running a hundred-billion-parameter model.
Core Architecture
Strata's technical design revolves around two key ideas:
MoE Expert Tiered Caching
Qwen3.8-Flash-Next is a MoE model with 24,576 experts, but each token activates only 10 of them. Strata exploits this property with tiered caching:
- GPU VRAM: Caches the few thousand most frequently requested experts (hot experts)
- System RAM: Holds all 24,576 experts
- CPU: Processes GPU cache misses in parallel
- SSD: Stores a large lookup table, loaded on demand
Think of it like a kitchen layout: ingredients used all the time stay on the counter (GPU), the rest waits in the pantry (RAM/SSD), fetched as needed.
Guess-and-Check Speculative Decoding
A small helper model guesses the next few tokens, and the large model verifies all guesses in a single forward pass. Correct guesses are accepted immediately — effectively generating multiple tokens per forward pass, yielding a 1.6-1.8x overall speedup.
Chunked Long-Context Reading
Long texts (such as 32K-token documents or code) are read in large chunks of up to 8,192 tokens, achieving over 1,000 tokens/second input processing speed.
Performance Benchmarks
Measured on two ordinary gaming PCs:
| Config | Quant | Generation | Prompt Processing |
|---|---|---|---|
| RTX 5070 (12GB) + 64GB RAM | Q2_0 | 94 tok/s | 2,650 tok/s |
| IQ2_XS | 79 tok/s | 2,090 tok/s | |
| IQ3_XXS | 62 tok/s | 1,750 tok/s | |
| IQ3_S | 53 tok/s | 1,620 tok/s | |
| RX 9070 XT (16GB) + 47GB RAM | Q2_0 | 60 tok/s | 1,160 tok/s |
| IQ2_XS | 52 tok/s | 1,110 tok/s |
60 tokens/second is faster than human reading speed. Larger-VRAM cards (e.g., RTX 3090 24GB) are expected to reach 100-140 tokens/second.
Model Variants and Quantization Levels
Strata offers multiple model variants to fit different RAM capacities:
| System RAM | Recommended | Notes |
|---|---|---|
| 32 GB | Coder | Code-specialized with half the experts removed; 91% of full model's SWE-bench Verified score |
| 48 GB | IQ2_XS or Q2_0 | Larger levels don't fit |
| 64 GB | IQ2_XS (recommended) / IQ3_XXS / IQ3_S | All levels fit; IQ3_S is best quality, slowest |
| 96 GB+ | IQ3_S or Unsloth 4-bit | Room for the largest levels |
Additional variants: Swift 1.5 (thinks shorter, answers sooner at similar quality) and Unsloth UD-Q4_K_XL (experimental, closest to full model but mostly read from SSD at 7-8.5 tok/s).
Installation and Usage
Installation is one-click: Windows users double-click START-HERE.bat, Linux users run ./setup.sh. The installer auto-detects the GPU, selects the right engine and model, downloads ~70 GB of model files (with resume support), and opens the browser at http://127.0.0.1:8080.
Also supports automated installation via AI coding assistants (Claude Code, Cursor, Codex, GitHub Copilot) or management through its MCP server.
External interfaces:
- OpenAI-compatible API:
http://127.0.0.1:8080/v1 - Anthropic-compatible API:
http://127.0.0.1:8080/v1/messages(Claude Code:ANTHROPIC_BASE_URL=http://127.0.0.1:8080) - Supports image input, thinking level control (off/low/medium/high), multi-GPU distribution, phone/remote access
Technical Position
Strata belongs to the LLM inference optimization domain. Its core contribution is tiering MoE expert caches across the full storage hierarchy of consumer hardware (GPU → RAM → SSD), combined with guess-and-check speculative decoding to achieve near-server-class inference speeds. It is not a new model — it is an inference engine that makes existing large models usable on ordinary PCs, dramatically lowering the hardware barrier to local hundred-billion-parameter model deployment.
License
Strata is open-source under the MIT License. The underlying model Qwen3.8-Flash-Next is developed by the Qwen team, with quantized versions by ISTA-DASLab, UkisAI, and Unsloth. The engine is built on llama.cpp / ggml. Individual components and models may carry their own licenses.
SOURCE LINKS