Every number

    Every number, and how we got it.

    Everything we publish, in one place: what we measured, what we compared it to, and the setup it ran on. We also share what our own checks threw out, and where we're still growing.

    The results

    33 results. Each is compared against the model's own standard setup on the same hardware.

    Every published Yantrion result, what it was compared to, and the setup
    AreaResultWhat we measuredCompared toSetup
    Speed1.60×Tokens per second with 4 users: 255 vs 159Standard SGLang, same GPUsKimi-K3, AMD Instinct MI350X, SGLang
    Speed1.43×Tokens per second with 32 users: 1,033 vs 722Standard SGLang, same GPUsKimi-K3, AMD Instinct MI350X, SGLang
    Speed≥2.1×Tokens per second with 128 users: 1,926 vs 905The standard stack could run only 54 of the 128 users. Yantrion's 1,926 is a conservative floor.Standard SGLang, same GPUsKimi-K3, AMD Instinct MI350X, SGLang
    Speed1.44–1.62×Speed with 28K tokens of context, 17 usersIn August this same test ran 1.48× slower. We found the cause and fixed it.Standard SGLang, same GPUsKimi-K3, AMD Instinct MI350X, SGLang
    Speed1.72×Speed with 128K tokens of context, 4 usersStandard SGLang, same GPUsKimi-K3, AMD Instinct MI350X, SGLang
    Speed2.32×Speed with 256K tokens of context, 4 usersStandard SGLang, same GPUsKimi-K3, AMD Instinct MI350X, SGLang
    Speedwithin ~5%Our standard-stack result (905 tokens/s per machine) vs independent published results for the same model on AMD MI355X (952)A fair baseline, not a weak one.Wafer, July 2026Kimi-K3
    Capacity2.54×Users at once in the same GPU memory: 193 vs 76Measured on a fixed set of real sessions, identical for both sides.The model's own standard memory formatQwen3.8
    Capacity2.00×Users at once in the same GPU memory: 32 vs 16The model's own standard memory formatKimi-K3
    Capacity+53%Tokens held in exactly the same GPU memory: 1,964,288 → 3,013,376The model's own standard memory formatDeepSeek-V4-Flash
    Capacity1.9×Memory made smallerThe model's own standard memory formatGLM-5.3
    Capacity4.49×Memory made smaller on the parts that can shrink; about 1.6× for all memory at 32K tokensSimulated on the real model, not yet run live.The model's own standard memory formatGPT-OSS-120B
    Capacity3.97×Memory for the model's recurrent layers made smaller3.74× across the whole memory pool.The model's own standard memory formatNemotron 3 Ultra, NVIDIA H200
    Reach1,048,576Tokens in a single prompt, with 128 of 128 identical answers on our strictest checkThe standard stack stops at 307,571 tokens on the same setupKimi-K3
    Reach4 GPUsA 1M-token context served on four NVIDIA H200sWith compact 4-bit model weights.Normally needs eightNemotron 3 Ultra, NVIDIA H200
    Agents150/150Agents resumed with memory intact after the whole fleet was cycledBest standard setup that would run: 0 of 150, same machine, same day150 agents, 36K tokens each
    Agents0 failuresFailed requests while the fleet grew from 150 to 300 agentsSlows down gracefully, never breaksAgent fleet, 36K tokens each
    Agents12.7 sTypical time to resume an agent with 150 running, vs 28.5 s starting over. With 96: 1.7 s vs 33.2 s.Starting overAgent fleet, 36K tokens each
    Agents580Saved agent sessions per machine116 without YantrionNemotron 3 Ultra, NVIDIA H200
    Quality128/128Multi-step agent conversations with word-for-word identical answersThe standard stack, no partial creditBoth test campaigns
    QualityPassNeedle-in-a-haystack long-document recallHidden fact foundKimi-K3, AMD Instinct MI350X, SGLang
    Quality15/16GSM8K math reasoningProblems solvedKimi-K3, AMD Instinct MI350X, SGLang
    Quality195Test runs saved and fingerprintedEvery run kept on recordBoth test campaigns
    Memory1.696×Memory set aside, still ready to reuseThe model's own standard memory formatNemotron 3 Ultra, NVIDIA H200
    Memory3.015×Memory set aside, still ready to reuseThe model's own standard memory formatDeepSeek-V4
    Memory1.89–2.14×Memory sealed and saved, portable between machinesThe model's own standard memory formatSaved memory
    Memory3.664×Smaller to send over the networkCounted on its own, never added to memory savings.Saved memory before sendingTransfer
    EcosystemSGLangHigh-speed open-source serving engineProven end to end with real modelsSupported and qualified
    EcosystemvLLMWidely used open-source serving engineProven end to end with real modelsSupported and qualified
    EcosystemTensorRT-LLMNVIDIA's high-performance serving engineProven end to end with real modelsSupported and qualified
    EcosystemNVIDIA DynamoAI serving at data-center scaleProven end to end with real modelsSupported and qualified
    EcosystemNIXLFast transfer of AI memory between GPUs and machinesProven end to end with real modelsSupported and qualified
    EcosystemLMCacheStores and reuses AI memoryProven end to end with real modelsSupported and qualified

    How we compare

    The rules every number follows.

    • Every comparison is against the model's own standard setup on the same hardware, never an easy baseline.
    • Nothing is compared unless we measured the standard stack on the exact same test.
    • We never mix different kinds of measurement in one comparison.
    • Savings on storage and transfer are reported on their own, never added to memory savings.
    • No speed tricks like speculative decoding are counted. Those gains would come on top.
    • We never claim to be number one. Other benchmarks measure things too differently to rank fairly.

    What our checks threw out

    Good-looking results that didn't hold up.

    • A compression setting that changed the model's answers. Thrown out.
    • A shortcut that saved more memory but hurt quality. Thrown out.
    • A test run where Yantrion never actually switched on. Discarded, because it can't count for us.
    • A faster 32-user setup that failed the quality check. We publish the slower result that passed: 30.045 ms per step.
    • Our own August result, 1.48× slower than the standard stack. We said so, found the cause, and fixed it. The same test now runs 1.44–1.62× faster.

    Where we're still growing

    Stated plainly.

    • If every agent is paused and resumed at the very same moment, 54 of 150 come back with memory intact today. Our next memory tier is designed to close that gap.
    • For a single user, speed isn't yet where our many-user results are. The biggest wins show up when many people use a model at once.
    • GPT-OSS-120B results are simulated on the real model, not yet run live.

    Quality comes first

    A result only counts if the model gives the same answers. No partial credit, and every run is kept on record.

    128/128

    Same answers on agent tasks

    Multi-step agent conversations matched word for word. No partial credit.

    Pass

    Long-document recall

    Needle-in-a-haystack tests: the hidden fact is found.

    15/16

    Math reasoning

    GSM8K grade-school math problems.

    195

    Runs on record

    Every test run saved and fingerprinted.

    Yantrion

    From algorithm to production performance.

    The universal AI memory platform: more AI, longer context and agents that keep their memory, on the GPUs you already own.

    Our mission

    Writing kernels, solving hard problems, and a love of the math underneath.

    Member of the NVIDIA Inception Program

    © 2026 Yantrion, Inc. All rights reserved.

    Measured on AMD Instinct MI350X and NVIDIA H200. Every number is real, never projected.