Anime, manga, and games, with a take · A Yukimedia publication

← all stories other 1 sources · 54m ago ·

Strata Runs A 125B Model On Gaming Pcs

The reported token speeds are real, but the looping pi output shows what the quantized tiers cost: Strata makes a 125B model runnable on a desktop, not fully intact.

Reporting from 1 source: GIGAZINE.

Strata Runs A 125B Model On Gaming Pcs

Strata is an app that runs a quantized version of the 125-billion-parameter Qwen3.8-Flash-Next on consumer gaming PCs. The repository lists requirements of an NVIDIA or AMD GPU with 12GB or more of VRAM, 32GB or more of RAM, 80GB or more of storage, and Windows 10/11 or Linux. Reported runs hit 94 tokens/second on an RTX 5070 with 64GB RAM, but the 2-bit quantized build produced looping output in one prompt.

The repo splits the model by memory. With 32GB of RAM only the GSQ-RCO Coder build runs, a lightweight version with non-coding expert models removed. At 64GB the 2-bit IQ2_XS and the 3-bit IQ3_XXS and IQ3_S builds fit, and higher-precision quantizations need more still.

Speed reports circulated on X. One post measured Q2_0 at 94 tokens/second on an RTX 5070, Ryzen 5 7600 and 64GB RAM, and IQ3_S at 53 tokens/second. Another measured Q2_0 at 60 tokens/second on a Radeon RX 9070 XT, Ryzen 9 3900X and 47GB RAM. A third reported 140 tokens/second for IQ2_XS on a 4090, then showed the same build looping on a pi prompt.

Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.

Sources