TL;DR
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A GitHub project called Strata says it can run the 125-billion-parameter Qwen3.8-Flash-Next model on supported gaming PCs, with reported results from an RTX 5070 and an AMD RX 9070 XT. Its published table lists up to 94 tokens per second for generated text on the RTX 5070, not an RTX 4090 result or 100 trillion tokens per second. The measurements are project-reported and do not include independent verification in the supplied material.
The open-source Strata project says it can run the 125-billion-parameter Qwen3.8-Flash-Next AI model on supported consumer PCs, with its published tests reporting up to 94 tokens per second for generated text on an RTX 5070. The supplied report does not show a test on an RTX 4090 or support the stated “100T/s” figure; the project’s unit is tokens per second, and its results are its own measurements.
Strata’s benchmark table lists results from two systems: an NVIDIA RTX 5070 with 12 GB of VRAM, Ryzen 5 7600 and 64 GB of system memory, and an AMD RX 9070 XT with 16 GB, Ryzen 9 3900X and 47 GB of RAM. On the NVIDIA system, the listed generation rates range from 53 to 94 tokens per second across the reported model formats. On the AMD system, the listed rates range from 44 to 60 tokens per second.
The table measures two different tasks. “Writes answers” reports the rate at which the model generates text, while “Reads your prompt” measures processing of a 32,000-token prompt; the NVIDIA figures for that task range from 1,620 to 2,650 tokens per second, and the AMD results range from 1,110 to 1,420. These prompt-processing rates are not the same as answer-generation speed. Strata notes that its NVIDIA Q2_0 result used engine version 0.1.36, while other NVIDIA rows used 0.1.26, with 4,000-token answers and 32,000-token prompts.
The project says it supports specified NVIDIA and AMD cards with at least 12 GB of VRAM, Windows or Linux, and at least 32 GB of RAM. It says the model loads roughly 35–55 GB into system memory, and the download is about 70 GB for the described setup. Strata presents the operation as local, saying data stays on the user’s PC. The supplied material does not include outside testing of that privacy claim or the benchmark results.
Local Access to a 125B Model
If the project’s results hold on other supported machines, Strata could let users run a large language model on hardware they own rather than sending prompts to a hosted service. That may matter to developers, researchers and privacy-conscious users who want local access for chat, coding, image input or connections to other applications. Local operation can reduce reliance on an external provider, though it does not by itself establish the model’s quality, security or suitability for every task.
The reported speed is also useful for setting expectations: on the tested RTX 5070, the fastest listed configuration generated 94 tokens per second, while other formats were slower. The project describes 60 tokens per second as faster than a person can read, but actual experience can vary with the chosen model format, prompt length, available memory and other PC workloads. The headline claim’s reference to an RTX 4090 and “100T/s” is not demonstrated by the included results, so buyers should not treat it as a benchmark for that card.
As an affiliate, we earn on qualifying purchases.
What Strata’s Benchmarks Measure
Strata is a software project hosted on GitHub that packages model setup and operation for local computers. According to its instructions, its installer detects the graphics card and selects an engine, then lets users choose a model format and context size. Users can also run setup manually on Windows or Linux. The project describes itself as free and open source; the supplied information does not provide a separate license assessment or independent audit.
The model options trade memory use and speed against model capacity and, according to the project, some degree of quality. Strata recommends different formats based on system RAM: its guide suggests a coding-focused option for 32 GB, and several smaller formats for 64 GB. It says its “Coder” variant removes half the experts and that its authors report 91% of the full model’s SWE-bench Verified score. That score is an attributed author claim, not an independently established result in the supplied material.
Strata also says a graphics card with more VRAM can be faster and estimates that an RTX 3090 with 24 GB could generate about 100–140 tokens per second. That is a projection from the project, not a result shown in the two-system benchmark table. The article’s “100T/s” wording should not be confused with that estimate: the report uses tokens per second, not trillions of tokens per second.
““Nothing leaves your PC.””
— Strata project
RTX 4090 Claim Lacks a Test
The supplied report does not identify an RTX 4090 benchmark, explain the “100T/s” unit in the headline claim, or show how that figure was measured. Its published results concern an RTX 5070 and an RX 9070 XT, and list rates in tokens per second. It is unclear whether the RTX 4090 wording refers to a separate test, a projection, or a mistaken unit.
The material also does not give test dates, independent replication, full benchmark methodology or a direct comparison with hosted versions of the model. Although Strata links to fuller tables and community results, those records are not included here. The stated local-data behavior and model capability descriptions likewise come from the project rather than an external evaluation.
Independent Tests Would Clarify Speed
The next useful evidence would be a reproducible RTX 4090 test that names the exact model format, software and engine versions, prompt and output lengths, and memory configuration. Independent tests on additional cards could show how closely other systems match Strata’s reported figures. Readers considering installation can consult the project’s linked details and hardware requirements, but should check those materials for updated compatibility information before downloading a roughly 70 GB model.
Until such results are available in the supplied reporting, the confirmed scope is limited: Strata reports running the model locally and publishes measurements from two other gaming-PC configurations. Whether the advertised RTX 4090 performance reaches 100 tokens per second remains unconfirmed.
Key Questions
Does the supplied report confirm 100T/s on an RTX 4090?
No. The included benchmark table covers an RTX 5070 and an RX 9070 XT, not an RTX 4090. It reports tokens per second and does not explain or verify the “100T/s” wording.
What generation speed did Strata report on the RTX 5070?
Strata lists answer-generation rates of 53–94 tokens per second across the reported NVIDIA model formats. The 94-token result is for Q2_0 and used engine version 0.1.36, according to the project.
What hardware does Strata say is required?
The project lists supported NVIDIA and AMD graphics cards with at least 12 GB of VRAM, at least 32 GB of system RAM, and Windows 10 or 11 or Linux. It also recommends about 80 GB of free disk space.
Are the performance results independently verified?
Not in the supplied material. The figures are Strata’s own measurements; the report does not provide independent replication or a complete methodology.
Source: hn
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
