Performance
Inference speed, measured
Every number on this page comes from our production API in Munich, with the method and the date next to it. Run the same test yourself.
Output Throughput, p50, 10K input / 1K output, single request. Last measured: July 2026.
Tokens per Second, Measured
These numbers are from our production API with our servers in the EU. Not vendor marketing - real measurements you can reproduce.
gpt-oss-120b
Our fastest
Up to 772 tok/s on shorter prompts
MiniMax M2.7 Ultraspeed
Our reasoning model
Up to 444 tok/s on shorter prompts
Gemma 4 31B
Native Multimodal
Native multimodal understanding for images and text
Server-side metrics (p50). Measured using our open-source benchmark tool. Your client-side results will vary based on network location and conditions. Last measured: July 2026.
Fast LLM inference, by workload
Which number decides your speed depends on what you are building. Same measured runs, cut three ways.
Chat and voice
Time to First Token
388 ms
gpt-oss-120b - p50
Conversational and voice interfaces are judged on how fast the first word appears, not on how fast the rest arrives.
Agents and coding
Output Throughput
713 tok/s
gpt-oss-120b - p50
Agent loops and coding assistants wait for the whole response, so tokens per second sets the pace. For frontier reasoning across long runs, MiniMax M2.7 Ultraspeed measures 428 tok/s with a 192K context window.
Batch and RAG
End-to-End Latency
1.789 s
gpt-oss-120b - p50
Pipelines care about the total round trip for a complete response. Gemma 4 31B takes image and text in the same request for document and vision workloads.
Fast is half of it. Every number above was measured on hardware we own in Munich - see what EU sovereign actually means.
What happens under load
A request's output speed holds when more requests run in parallel: the chip serves a small batch at a time, and each answer keeps streaming at close to full speed. What grows is the time to first token, when requests have to wait for a place in the next batch. On the shared API that happens at peak times. If your workload needs planned capacity, Enterprise plans come with guaranteed rate limits, and a dedicated rack serves only your traffic.
Open Source
Run Your Own Benchmark
Our benchmark tool is fully open source. Run it against our API with your API key, or clone the repository and run it locally. Same code, same methodology, your results.
- 01
Synthetic Performance
Fixed input/output token counts for controlled comparisons across models
- 02
Real Workload Simulation
Variable request rates mimicking production traffic patterns
- 03
Custom Dataset
Upload your own prompts and measure performance on your actual workload
- 04
Interactive Chat
Per-response metrics in a live chat interface - see TTFT and throughput on every reply
How we measure
Three numbers, measured under the same conditions: server-side p50, 10K input and 1K output tokens, one request at a time. We re-measure every quarter.
Time to First Token (TTFT)
How quickly the model starts responding after your request. Critical for interactive applications and chat interfaces.
Output Throughput
Tokens generated per second after the first token. Determines how fast a complete response is delivered to the user.
End-to-End Latency
Total time from request to complete response. Includes TTFT plus full generation time. The number that matters for batch workloads.
Where these numbers come from
All measurements come from our production infrastructure in Munich.
- Location
- Munich, Germany
- Hardware
- SambaNova SN40L
- Ownership
- Hardware we own
- Certification
- ISO 27001 Certified
Your requests are processed on racks that belong to Infercom, not on rented cloud capacity.
About the technology, independently
These sources measured or reported on SambaNova's dataflow technology, mostly on SambaNova's own cloud in the US. They are not measurements of our service; ours are above.
- Artificial Analysis - SambaNova Benchmarks Artificial Analysis: Independent, continuously updated speed and latency measurements of SambaNova's cloud, alongside other providers.
- SambaNova Breaks 1,000 Tokens/Sec Barrier VentureBeat: 2024: SambaNova passes 1,000 tokens per second on Llama 3.
- DeepSeek R1 671B with 95% Fewer Chips TechRadar: 2025: SambaNova runs DeepSeek R1 671B on 16 chips, where GPU setups use far more.
- Speed Record on Llama 3.1 405B SambaNova: SambaNova's speed result on Llama 3.1 405B, measured by Artificial Analysis.
- Intelligence per Joule SambaNova: SambaNova on why energy per unit of intelligence matters at scale.
- Intelligence Per Watt Research Stanford Hazy Research: An independent method for measuring AI efficiency; the study includes the SN40L.
Speed and energy
Less data movement makes the chip fast and saves energy: a request that finishes sooner draws power for a shorter time.
10 kW
Per rack, typical
About 10 kW in typical use, 7-14.5 kW in inference.
Air-cooled
No liquid cooling
Fits data centers built for regular servers. The chips need no water for cooling.
Up to 5x
Energy efficiency
SambaNova: 2.5x-5.6x against an NVIDIA H200 with Llama 70B, depending on batch size (2025).
LLM Inference Speed: Frequently Asked Questions
Who has the fastest LLM inference in Europe?
We do not rank other providers; independent comparisons such as Artificial Analysis do that. What we publish is our own measurement: 713 tok/s on gpt-oss-120b, up to 772 tok/s on shorter prompts, measured on our production API in Munich (July 2026). Run the same test against any provider with our open-source benchmark tool.
Which model is fastest, and for what?
gpt-oss-120b leads output throughput at 713 tok/s and also has our lowest time to first token at 388 ms, which makes it the default for chat, voice and high-volume agent loops. MiniMax M2.7 Ultraspeed is our 229B frontier reasoning model at 428 tok/s with a 192K context window, for long agentic runs and hard coding tasks. Gemma 4 31B measures 199 tok/s and adds native image and text input. Chat and voice depend on time to first token; agents and coding depend on throughput.
How is this measured - can I reproduce it?
Yes. Every number is a server-side p50 at 10K input and 1K output tokens, single request, measured against our production API with our open-source benchmark tool. Run it online against our API with your own key, or clone the repository and run it locally against your own prompts. Your client-side results will differ depending on network location and conditions.
What happens to speed when many requests run at once?
Each request keeps streaming at close to full output speed, because the chip serves a small batch at a time. The time to first token grows when requests wait for the next batch, which happens on the shared API at peak times. For planned capacity: Enterprise rate limits or a dedicated rack.
Is the speed independently verified?
Artificial Analysis continuously measures SambaNova's cloud, and VentureBeat and TechRadar have reported on its speed; a Stanford study of energy efficiency includes the SN40L. Those sources cover the technology. The per-model numbers on this page are our own measurements on our own infrastructure in Munich, which is why we publish the benchmark tool so you can check them.
Why is dataflow faster than GPUs?
While writing an answer, a GPU runs a model layer as many separate steps and writes the results in between back to memory. A dataflow chip runs the whole layer as one step and keeps those results on the chip, so it spends its memory bandwidth on the weights and answers sooner.
Optimize Your Integration
Get the best performance from your Infercom integration with our developer resources.