A GitHub project named Strata is drawing attention for a bold claim: running Qwen3.8 Flash Next, a 125-billion-parameter model, on consumer hardware at 100 tokens per second. The repository, published by developer Niko1221, advertises one-click installation for Windows and Linux and has already collected more than 11,000 stars.
The hook is simple. Most 100B-plus models normally demand data-center GPUs or multi-card rigs. Strata says a single RTX 4090 is enough, and wraps the whole setup in an OpenAI- and Anthropic-compatible API served on localhost, with optional image input.
What Strata actually offers
Strata is not just a model download. The repository bundles a custom inference engine, also called Strata, along with setup scripts for both Windows and Linux. Build files point to CMake, a Dockerfile, and a third_party/ggml folder, the same foundation used by many local inference tools.
The project ships SETUP.bat and START-HERE.bat files for Windows users, plus setup.sh and update.sh for Linux, suggesting the team wants a near-zero-configuration experience. A bench/results directory in the repo suggests the 100T/s figure comes from the developers' own benchmarks.
Qwen3.8 Flash Next on consumer hardware
The model at the center of the story is Qwen3.8 Flash Next, a 125B parameter model. Running it locally would normally require serious GPU memory, which is why the 100T/s claim on a single RTX 4090 is the number everyone is repeating.
Strata also exposes an OpenAI- and Anthropic-style API on localhost, meaning developers can point existing tools and SDKs at their own machine instead of a cloud endpoint. Optional image input is listed as a feature, hinting at multimodal use.
As with any GitHub benchmark, the numbers are self-reported and not independently verified. Treat the 100T/s figure as a developer claim until third parties reproduce it.


