Big LLMs on Gaming Hardware with Strata

Strata runs Qwen3.8-Flash-Next, a 125-billion-parameter model that normally needs a server, on an ordinary gaming PC with a 12 GB graphics card. I went through its docs and community benchmarks to work out which budget hardware is actually worth it.

  • AI
  • LLM
  • Hardware

How it fits on a gaming PC

Qwen3.8-Flash-Next is a mixture of experts: 24,576 small specialist networks, of which each token uses only 10. Strata exploits that by spreading the model across the whole PC:

  • Graphics card: the few thousand experts that are used most often
  • RAM: all of the experts, with the processor computing the ones the card doesn’t hold
  • SSD: a 29 GB lookup table that stays on disk

On top of that, a small draft model guesses the next few tokens and the big model checks them in one go, which the project says makes answers 1.6–1.8 times faster with the same result. The model is heavily compressed (2- to 3-bit), and you pick how much.

This leads to the one rule that decides every hardware choice: RAM decides which model fits, VRAM decides how fast it runs. According to the docs, every extra gigabyte of VRAM holds about 700 more experts, and every expert on the card is one the processor doesn’t have to compute.

The minimum

  • Graphics card: NVIDIA RTX 20, 30, 40 or 50 series, or a recent AMD Radeon (RX 6800 and up, RX 7700 XT and up, RX 9060 XT and up), with 12 GB of VRAM or more
  • RAM: 32 GB, which is enough for the Coder version only
  • Processor: any x86-64 desktop CPU with AVX2, so roughly anything from the last eight years
  • Disk: about 80 GB free, ideally on an NVMe SSD
  • System: Windows 10/11 or Linux. There’s no Mac version

Budget options, from cheapest up

1. A 12 GB card you already have, with 32 GB of RAM

With 32 GB of RAM the installer picks Coder, a version with half of the experts removed. Its authors report 91% of the full model’s SWE-bench Verified score, so it’s good at programming but noticeably weaker at everything else, including languages other than English. On the project’s RTX 5070 test PC it wrote 55 tokens per second. That PC had 64 GB of RAM, though, so treat it as a best case for this tier.

2. Upgrade the RAM to 48 or 64 GB

Usually the cheapest upgrade with the biggest effect. With 48 GB the full model fits in its two smallest sizes (Q2_0 and IQ2_XS), and with 64 GB every size fits. Measured on the same RTX 5070 (12 GB) with 64 GB of RAM:

SizeWrites answersReads prompts
Q2_0 (fastest)94 tokens/s2,650 tokens/s
IQ2_XS (recommended)79 tokens/s2,090 tokens/s
IQ3_S (best quality)53 tokens/s1,620 tokens/s

For comparison, 60 tokens per second is already faster than you can read.

3. A 16 GB card

On AMD, an RX 9070 XT (16 GB) with 47 GB of RAM measured 60 tokens per second with Q2_0. On NVIDIA, the project estimates about 80 tokens per second for an RTX 5060 Ti 16 GB with 64 GB of RAM, give or take 20%.

4. A used 24 GB RTX 3090

This is the card the docs keep coming back to: an older generation, but with twice the VRAM of a 12 GB card, it takes most of the work away from the processor. The estimate is 100–140 tokens per second with 64 GB of RAM. A 24 GB card also unlocks a low-RAM mode: with only 32 GB of RAM, it can run the full model (Q2_0 and IQ2_XS) instead of just Coder.

5. Two cards instead of one

Strata can split the model’s layers across two or three cards, and each card caches experts for its own layers, so two cards hold about twice as many. It doesn’t need NVLink, and cards in slow x4 or x1 slots work. So a second-hand card next to the one you have is a real option. Cards must all be NVIDIA RTX 20 or newer, or all AMD; you can’t mix the two brands.

6. Old datacenter cards (experimental)

This is where it gets interesting for a homelab. Community members got Strata running on retired server cards, which can be very cheap second-hand:

  • Tesla P40 (24 GB): 30–33 tokens per second with the IQ3_S size, as reported by one user.
  • Two AMD Instinct MI50s (16 GB each) in a server with an old Xeon E5-2666 v3 (Haswell generation), 32 GB of RAM and a SATA SSD: Coder at 50 tokens per second, and still 46 tokens per second with a 128,000-token prompt. Between them the two cards held all of Coder’s experts, so the old CPU barely mattered.

The catch: these paths are community-written, the maintainers don’t own the cards, and they’re meant for people who build from source. The MI50 needs a community build of ROCm, because AMD no longer ships libraries for it. Also remember that many server cards, like the P40, have no fan of their own and expect a server’s airflow.

What doesn’t pay off

  • 8 GB cards run, but slowly. Most of the model then has to be computed by the processor.
  • Models bigger than your RAM. The 4-bit Unsloth version reads experts from the SSD while it answers and manages 7–8.5 tokens per second on a 64 GB PC. Buy RAM before you go there.
  • GTX 10 series and older Tesla cards work only through the experimental CUDA 12 engine. Fine for tinkering, not something to buy for this.
  • Intel Arc is experimental and Linux-only, built from source.

Trying it

Download or clone the repository, then run START-HERE.bat on Windows or ./setup.sh on Linux. The installer detects your card and RAM, recommends a model and size (pressing Enter accepts the recommendation), and downloads about 70 GB. The chat app then opens at http://127.0.0.1:8080.

The useful part is the API. Strata speaks both the OpenAI and the Anthropic formats, so most tools can use it as a local model:

# any OpenAI-compatible app: base URL, any API key, any model name
http://127.0.0.1:8080/v1

# Claude Code
ANTHROPIC_BASE_URL=http://127.0.0.1:8080 claude

Things to know first

  • The first start can freeze your PC for 1–3 minutes. It loads 35–55 GB into RAM and locks part of it for the graphics card. Wait, and don’t close the window.
  • It answers one request at a time by default. Running two in parallel is possible but makes each answer slower on a 12 GB card.
  • Opening it up to your network needs a key. The docs use --host 0.0.0.0 --api-key <secret>. Without the key, anyone on your network can use your GPU.
  • It’s very new. The project started in late September 2026, the engine is at version 0.1.x, and it changes daily. The speeds above are the project’s and its users’ own measurements, and the 16 GB NVIDIA and 3090 figures are estimates.

Wrap-up

If you already have a gaming PC with a 12 GB card, the best value is more RAM: 64 GB turns it into a machine that runs a 125B model at reading speed or better. If you’re buying a card for this, VRAM beats speed, which makes a used 24 GB RTX 3090 the sweet spot. And for tinkerers, two old MI50s in a retired server get surprisingly far, as long as Coder is enough and you don’t mind building from source.

← Back to Scratchpad