Developer Resources
EU Sovereign AI Inference Glossary
Clear, sourced definitions across EU sovereign AI inference - from data residency, GDPR, and zero data retention to TTFT, throughput, and dataflow architecture. Every entry is backed by published sources and real benchmark data from our EU infrastructure.
Sovereignty & Compliance
- Data ResidencyWhere your data is physically stored and processed - a necessary part of sovereignty, but not the same as control over who can legally reach it.
- Data Processing Agreement (DPA)The GDPR Article 28 contract that governs how an inference provider may process the personal data in your prompts - and a basic test of whether a provider is enterprise-ready.
- GDPR for AI InferenceWhat Europe's data-protection law requires when your prompts contain personal data - a lawful basis, a processor agreement, and processing that stays within reach of EU law.
- Zero Data Retention (ZDR)When an inference provider doesn't store your prompts or outputs after serving a request, and never trains on them - shrinking your data exposure to the moment of processing.
Performance Metrics
- TTFT (Time to First Token)How long a user waits between sending a request and seeing the first token of the response.
- Inter-Token Latency (ITL)The average time gap between consecutive tokens during generation - also called TPOT.
- Tokens per SecondThe standard unit for LLM generation speed - and why the same number can mean two different things.
- Inference SpeedThe umbrella term: TTFT, inter-token latency, and throughput - and which one matters when.
Architecture
- RDU (Reconfigurable Dataflow Unit)SambaNova's AI processor - purpose-built AI chips designed for dataflow execution instead of instruction-by-instruction processing.
- Dataflow ArchitectureThe execution model where data streams through operations as a pipeline - eliminating the kernel-by-kernel round trips of GPU execution.
Models & Inference
- InferenceRunning a trained AI model to produce outputs - the production workload of AI, and the one whose cost and speed compound with usage.
- Throughput (LLM Serving)Tokens per second in two senses: per-request output throughput vs. system-wide capacity - and how batching trades one against the other.
- Prefill vs. DecodeThe two phases of LLM inference - parallel prompt processing vs. token-by-token generation.
- Latency vs. ThroughputThe fundamental serving trade-off: total system output vs. each user's speed.
- Open-Weight ModelA model whose trained parameters are published so anyone can run it themselves - the technical basis for sovereign inference.
- Context WindowThe maximum amount of text, in tokens, a model can consider at once - prompt plus output. Its length directly shapes inference speed and cost.
- ParametersA model's learned weights - the rough measure of its size and capacity, and the direct driver of its memory, speed, and cost.
- Temperature (Sampling)The parameter controlling randomness in token selection - where 1.0 is the baseline and 0 forces greedy decoding.
- Top-P (Nucleus Sampling)A sampling method that keeps only enough high-probability tokens to cover a cumulative probability p - adapting the candidate pool to the model's confidence.
- Top-K SamplingA sampling method that limits selection to the k highest-probability tokens - a hard cap on the candidate pool.
- Greedy DecodingThe decoding strategy that always picks the highest-probability token - deterministic, but prone to repetition loops.