Qwen3.8-Flash-Next is a 125B-parameter model that only activates 6B per token, and a free tool called Strata lets a gaming PC run it. It splits the model across your GPU, RAM, CPU and SSD. The real requirement is RAM, not VRAM: 32 GB runs the Coder version, 64 GB runs every size. On Strata's own measurements, a 12 GB RTX 5070 with 64 GB of RAM hits 94 tokens per second in short chat with the smallest size. Strata exposes an OpenAI-compatible endpoint at http://127.0.0.1:8080/v1, so any coding tool can point at it. Expect 2-3 bit quality, a long first load, and one request at a time. Fair warning, we build a local-first AI coding tool, so we are not neutral.
On October 4, a Hacker News thread titled “Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s” hit 922 points. People wanted to know if it was real. Strata's own measurements back most of it, with conditions. This is the hands-on version: what you need, which size to pick, what the speeds look like, how to install it, and where it falls short.
What is Qwen3.8-Flash-Next?
Qwen3.8-Flash-Next is an open-weights mixture-of-experts model from Qwen with 125B parameters and 6B active per token, plus a 51B n-gram embedding and a 4B draft layer. Qwen calls it an experimental preview of the architecture that will underpin Qwen4. It uses 512 experts, 10 routed plus 1 shared per token, across 48 layers, with a hybrid of Gated DeltaNet and Qwen Sparse Attention. It also takes image input and has a 262,144-token native context, extensible to 1,000,000. The license is the Qwen Community License 1.0, which we cover in the caveats.
The coding numbers are why people care. These are Qwen-reported, at full precision, compared with Qwen3.8-27B:
| Benchmark (Qwen-reported) | Flash-Next | Qwen3.8-27B |
|---|---|---|
| SWE-bench Pro | 62.5 | 61.7 |
| DeepSWE 1.1 | 58.7 | 42.2 |
| SWE-bench Multilingual | 81.0 | 73.8 |
| NL2Repo-Bench | 48.1 | 42.3 |
| LiveCodeBench v6 | 91.9 | 90.3 |
| Toolathlon Verified | 73.5 | 67.1 |
The gap on SWE-bench Pro is small; the bigger gaps are on DeepSWE and SWE-bench Multilingual. These are vendor numbers, and none of them were run at the 2-3 bit sizes you will actually use here.
What does Strata do?
Strata spreads the model across every part of your PC instead of forcing it all into VRAM, which is how a model this size fits outside a server. It is an MIT-licensed project by Niko1221 that runs one model, Qwen3.8-Flash-Next, in several sizes. The split works like this:
- GPU. Keeps the most-used few thousand of the model's 24,576 experts.
- RAM. Holds all of them.
- CPU. Computes the experts that are not on the GPU, in parallel.
- SSD. Stores a 29 GB lookup table, the 51B n-gram embedding.
On top of that it uses speculative decoding with the model's built-in MTP draft layer. Strata says that gives the same answer 1.6 to 1.8 times sooner. Long prompts are read in chunks of up to 8,192 tokens.
What hardware do you need?
You need 32 GB of RAM or more, a 12 GB+ GPU, and about 80 GB of free disk, and RAM is the number that decides what you can run. Strata lists NVIDIA RTX 20/30/40/50 cards or listed AMD Radeon cards with 12 GB or more, Windows 10/11 or Linux, and a current driver. An SSD is recommended. macOS is not listed.
This is why our GPU guide and the VRAM calculator only get you halfway here. For this model, check your RAM first. Strata's guidance by RAM:
| Your RAM | What to run |
|---|---|
| 32 GB | Coder version |
| 48 GB | IQ2_XS (or Q2_0) |
| 64 GB | IQ2_XS recommended, or IQ3_XXS / IQ3_S (IQ3_S with little else open) |
| 96 GB or more | IQ3_S with room to spare, or Unsloth's ~4-bit UD-IQ4_XS |
Which size should you pick?
Strata recommends IQ2_XS for 64 GB of RAM, and the Coder version for 32 GB. The installer picks a size for your machine, but you should know what you are choosing between. The RAM-plus-VRAM column is the combined memory the size needs.
| Size | RAM + VRAM | Strata's rating |
|---|---|---|
| Q2_0 | 37.6 GB | Good, fastest |
| IQ2_XS | 39.2 GB | Better (recommended) |
| IQ3_XXS | 47.0 GB | Great |
| IQ3_S | 54.8 GB | Best, slowest. Strata says it matches the full model on the published tests (original model only) |
The three smaller sizes are a 66 to 76 GB download, plus a roughly 6 GB MTP draft layer. Budget the disk.
The Coder version comes from ISTA-DASLab. It keeps 256 of the 512 experts, chosen on code data, which is how it fits in 32 GB of RAM. Its authors report 91% of the full model's SWE-bench Verified score and 99% of LiveCodeBench. The tradeoff is that it is weaker outside code and in non-English languages. A Strata GitHub issue (#438) describes CJK answers looping. If you only prompt in English, that may not matter. If you do not, it will.
How fast is it?
On Strata's own measurements, a 12 GB RTX 5070 with 64 GB of RAM runs 43 to 94 tokens per second depending on size and context. These are Strata's numbers on its own test rigs, not ours, in tokens per second.
| Size (RTX 5070 12 GB, 64 GB RAM) | Short chat | At 128K context |
|---|---|---|
| Q2_0 | 94 | 76 |
| IQ2_XS | 79 | 63 |
| IQ3_XXS | 62 | 49 |
| IQ3_S | 53 | 46 |
| Coder | 55 | 43 |
On an AMD RX 9070 XT 16 GB with 47 GB of RAM, Strata measured Q2_0 at 60 tokens per second.
Two numbers you will see quoted elsewhere. Strata estimates an RTX 3090 24 GB at about 100 to 140 tokens per second, and that is an estimate, not a measurement. The “RTX 4090 at 100T/s” in the HN title comes from community reports, not from Strata's own table.
Prompt length matters for agent work. Strata's README budgets about a minute per 30K tokens for the first message of a chat, which is read in full; its Q2_0 benchmark on the 5070 is faster than that. Follow-ups start in seconds.
How do you install Strata?
Download or clone the repo, run the start script for your OS, and let the installer pick a size. Check the repo for current instructions before you start, since the project is only weeks old.
- Make sure you have about 80 GB free and a current GPU driver.
- Get Strata from
github.com/Niko1221/Strata. - On Windows, run
START-HERE.bat. On Linux, run./setup.sh. - Let the installer choose a size for your RAM, or override it using the tables above.
- Wait for the download, then open the app at
http://127.0.0.1:8080. It has Chat, Monitor and About tabs.
Expect the first start to be rough. Strata's README says your PC can be slow or stop responding for 1 to 3 minutes while it loads 35 to 55 GB into RAM. That is normal. Do not kill it.
How do you connect a coding tool?
Point your tool at Strata's local endpoint. It speaks the OpenAI API, the Anthropic API, and the OpenAI Responses API. Any API key works, and so does any model name. Reasoning effort can be set to off, low, medium or high.
| Endpoint | Use it for |
|---|---|
http://127.0.0.1:8080/v1 | OpenAI-compatible clients, including custom providers |
/v1/messages | Claude Code (Anthropic API) |
/v1/responses | Codex CLI (OpenAI Responses API) |
For Claude Code, set the base URL and start it as usual:
# bash
export ANTHROPIC_BASE_URL=http://127.0.0.1:8080
claude
# PowerShell
$env:ANTHROPIC_BASE_URL="http://127.0.0.1:8080"
claudeFor Bodega One Code, add a custom OpenAI-compatible provider. Our app takes any OpenAI-compatible endpoint, on top of its 10+ LLM provider presets (see BYOLLM). Set the base URL to http://127.0.0.1:8080/v1, use any value for the API key, and type any model name. One thing to know: our app does not ship Flash-Next in its built-in model catalog, because the smallest file is 68 GB and runtime support is still settling. Strata is one third-party way to run it locally today.
We have not benchmarked Bodega One Code with Strata. It uses the same OpenAI-compatible endpoint any client uses, and that is all we are claiming. With a local model and air-gap mode, nothing leaves the machine. For a general setup walkthrough, see how to run a local LLM for coding.
If you want to reach Strata from another machine, rerun setup with the host and key flags, and always set a key. On Windows (the Linux steps are in Strata's install guide):
START-HERE.bat --setup --host 0.0.0.0 --api-key <secret>What are the honest caveats?
It makes a big model runnable on a gaming PC, with real limits. Know these before you spend an evening on it.
- Quality at 2-3 bit. Qwen's benchmark numbers are full precision. Strata's sizes are quantized to 2 or 3 bits. Strata says IQ3_S matches the full model on the published tests, but the smaller sizes and the Coder version give something up. Test on your own code.
- First-start freeze. 1 to 3 minutes of a sluggish PC while it loads into RAM.
- One request at a time. The default is a single request. There is a
"parallel": 2option, but it is slower per answer on a 12 GB card. Multi-agent workflows will queue. - No macOS. Windows and Linux only, as far as Strata lists.
- Third-party and young. Strata is not our project. The repo was created on September 24, and we can only vouch for what its public docs say.
- The license has conditions. Strata is MIT. The model is under the Qwen Community License 1.0, which is not Apache or MIT. It allows use, modification, distribution, sale, hosting and fine-tuning, with two conditions. Products with over 100M monthly active users or over US$20M monthly revenue must prominently display the model name. And a licensee running a “Model as a Service” or “AI Work Assistant” business, meaning an independent AI product mainly for AI-assisted coding or office productivity, needs a separate license from Qwen before commercial use. Internal use is exempt as long as the model, its outputs or its capabilities are not made available to third parties. In plain terms, using it for your own coding looks fine, and building a coding-assistant product on it needs Qwen's say-so. Read the license, not this paragraph. This is not legal advice.
When is the simpler choice better?
If you have a single 24 GB card and under 64 GB of RAM, run Qwen3.8-27B instead. At about 17 GB in Q4, it fits entirely in VRAM, loads quickly, runs in any standard local runner, and skips the first-start freeze, the one-request limit and the 80 GB download. On SWE-bench Pro, Qwen reports 61.7 for it against 62.5 for Flash-Next, both at full precision. Here you would be comparing a Q4 27B against a 2-3 bit Flash-Next, so that small gap may close or flip. For most day-to-day coding, the 27B is the better trade.
Flash-Next earns the setup when you already have the RAM and want Qwen's larger model, especially on DeepSWE and SWE-bench Multilingual (many programming languages), where the Qwen-reported gap is widest. For a ranked list of what runs on which hardware, see local LLMs for coding. And if you want to try the whole thing inside a local-first IDE, Bodega One Code is free for everyone during the open beta.
Sources
Common questions
- Can I run Qwen3.8-Flash-Next on a gaming PC?
- Yes, with Strata, if you have an NVIDIA RTX 20/30/40/50 or a listed AMD Radeon card with 12 GB or more of VRAM, at least 32 GB of RAM, and about 80 GB of free disk. RAM is the real requirement. Strata runs on Windows and Linux, and it does not list macOS.
- How much RAM do I need for Qwen3.8-Flash-Next with Strata?
- 32 GB of RAM runs the Coder version only. 48 GB fits IQ2_XS or Q2_0. 64 GB runs every size (IQ3_S with little else open). 96 GB or more gives IQ3_S room to spare or fits Unsloth's ~4-bit UD-IQ4_XS build. The model weights load into system RAM, so the GPU alone is not enough.
- How fast is Qwen3.8-Flash-Next on Strata?
- In Strata's own measurements on an RTX 5070 12 GB with a Ryzen 5 7600 and 64 GB RAM, the Q2_0 size ran 94 tokens per second in short chat and 76 at 128K context. IQ2_XS ran 79 and 63. Strata estimates an RTX 3090 at about 100 to 140 tokens per second, but that is an estimate, not a measurement.
- Is a 2-bit or 3-bit Qwen3.8-Flash-Next as good as the full model?
- Not guaranteed. Qwen's published benchmarks are at full precision, not at Strata's 2 to 3 bit sizes. Strata says its largest size, IQ3_S, matches the full model on its published tests, but only for the original model. Smaller sizes trade quality for speed and fit. Test on your own code before you rely on it.
- Can I use Strata with Claude Code, Codex, or Bodega One Code?
- Strata serves an OpenAI-compatible API at http://127.0.0.1:8080/v1, an Anthropic-style API for Claude Code, and an OpenAI Responses API for Codex CLI. Bodega One Code accepts any OpenAI-compatible endpoint as a custom provider. We have not benchmarked Bodega One Code with Strata, so treat it as a standard endpoint hookup.
- Should I run Qwen3.8-Flash-Next or Qwen3.8-27B?
- If you have a single 24 GB card and less than 64 GB of RAM, the simpler choice is Qwen3.8-27B at about 17 GB in Q4. It fits in VRAM, loads fast, and needs no special engine. Flash-Next is worth the setup when you have the RAM and want Qwen's larger model.
Written by the Bodega One team. We build Bodega One Code, the local-first AI IDE, and we write here about local models, AI costs, and what we learn shipping it. More about the team and why we build local-first on the about page.
Related posts
Keep learning
Free, vendor-neutral courses and guides in the Bodega One AI Academy, from what a model is to shipping an app.
Stay in the loop
Build-in-public updates, model picks, and Copilot/Cursor news as it breaks.
Follow @BodegaOneAI on X →