Inference API

Simple, Transparent Pricing

Pay-as-you-go rates for the Inference API. No minimums, no commitment - pay only for the tokens you use.

  • EUR Native Pricing
  • No Minimums, No Commitment
  • EU Sovereign by Default

Inference API pricing

Pay-per-token rates for our shared, EU sovereign inference platform. Need reserved capacity or your own hardware? Those two are priced per engagement.

You're here

Inference API

Shared infrastructure, pay-per-token. Start in minutes, scale as you grow - no minimums, no commitment.

  • EU sovereign by default
  • Pay as you go
  • OpenAI-compatible API

Starting from

€0.20 / 1M tokens, excl. VAT

See all rates

Compare all three side by side

How API Pricing Works

The Inference API is priced per token. A token is roughly 4 characters in English (this varies by model and language). You pay for what you use - no minimums, no commitments.

What is a token?

Tokens are the basic units LLMs process. In English, 1 token ≈ 4 characters or ¾ of a word. 1,000 words ≈ 1,300 tokens. Other languages may use more tokens per character.

Input vs Output

You pay separately for input (your prompt) and output (the model's response). Output costs more than input because your prompt is read in a single parallel pass, while each output token needs its own full pass through the model - one token at a time.

Choosing a model

Each model excels at different tasks - there's no single "best" choice. EU Sovereign models (🇪🇺) guarantee your data stays in European jurisdiction.

EU

EU Sovereign Models

Full GDPR compliance, no US CLOUD Act exposure

EU Sovereign Models
ModelModalityInput/1MOutput/1MContext
MiniMax M2.7 UltraspeedText€0.60€2.40192K
Gemma 4 31BText + vision€0.20€0.35128K
gpt-oss-120bText€0.22€0.59128K
Whisper Large v3Speech-to-text€0.10--
E5-Mistral 7BEmbeddings€0.13-4K

Prices in EUR, excl. VAT. EU-hosted models include full data sovereignty. Speech-to-text and embeddings are billed on input tokens only (they return transcripts or vectors, not generated text).

Global

Global Model Catalog

Additional models via global infrastructure

Global Model Catalog
ModelModalityInput/1MOutput/1MContext
Llama 3.3 70BText€0.60€1.20128K
DeepSeek V3.1Text€3.00€4.50128K
DeepSeek V3.2Text€3.00€4.5032K

Prices in EUR, excl. VAT. Requests processed on global infrastructure outside the EU.

Estimate your monthly cost

Start from a typical workload, then adjust the numbers to match yours.

Estimated monthly cost (excl. VAT)

€324.00

Per request
€0.00108
Per day
€10.80
Rate (in / out per 1M)
€0.60 / €2.40

EU Sovereign - data stays in EU

Estimate only. Assumes ~30 days of steady usage. Actual token counts vary by language, content, and model tokenizer (~750 words ≈ 1,000 tokens for English). Speech-to-text and embeddings aren't shown here - they bill on input tokens only; see the rate table above. Check exact usage anytime at cloud.infercom.ai/plans/usage.

Access Tiers

Choose how you want to start - upgrade anytime as you grow.

Start here

Developer

Pay as you go

Sign up, create an API key, and pay only for the tokens you use. No minimums, no commitment.

View rate limits

 

Enterprise

List price, with volume discounts

The same per-token rates as your baseline, with committed-volume discounts, custom SLAs, priority support, and higher rate limits.

Contact sales

Frequently Asked Questions

How does billing work?

The Inference API is pay-as-you-go. You add a billing method - a credit card - and usage is drawn down per token against the published list prices, so you only ever pay for what you use, with no minimums and no monthly commitment. Cost is easy to model up front from the rates above.

Can I see my token usage and current spend?

Yes, in real time. Your live token consumption is at cloud.infercom.ai/plans/usage, and your current billing and invoices are at cloud.infercom.ai/plans/billing. You always know exactly what you've used and what it costs.

What are the rate limits?

Rate limits vary by plan and model. The Developer tier has standard rate limits documented in our rate limits documentation. Enterprise plans offer custom rate limits tailored to your workload. Rate limits control request frequency - they do not affect inference speed per request.

What payment methods do you accept?

For our inference service, we accept major credit cards through Stripe. Enterprise, dedicated capacity, and on-premises customers can also pay by invoice.

What's included in EU sovereignty?

For EU sovereign models, all data processing happens in our European data centers by default. Your data never leaves EU jurisdiction unless you explicitly opt to use Global Catalog models. Full GDPR compliance and AI Act readiness included.

Do you offer volume discounts?

Yes - both Enterprise and dedicated capacity contracts offer custom pricing based on your committed usage. Contact our sales team to discuss your requirements.

Ready to Build with Enterprise-Grade AI?

Start with a pilot, scale to production. Record-breaking performance with dedicated enterprise support.