On-Premises

On-premises AI inference. Full control, cloud speed.

The same models and the same speed as our cloud, on hardware you own, on your own network, with your data never leaving the building. Air-cooled, at about 10 kW per rack, so it fits the datacenter you already have.

Full data control, cloud speed

When your data can't leave, this is the answer

The same platform and the same speed as our cloud, entirely within your walls. On-premises is the right choice when data or source code physically cannot leave your building, when a regulator or client mandates it, or when you've decided to own your AI infrastructure outright. The racks are air-cooled, with no liquid cooling or special buildout, so they fit the facility you already have. And if a workload doesn't carry that kind of requirement, our cloud and dedicated options give you the same speed with none of the capital commitment - we'll tell you honestly which fits.

The Fit

When to choose On-Premises

Four situations where owning the platform is the right call - not a fallback.

Data that can't leave

Source code, IP, or sensitive data that legally or contractually must stay inside your perimeter.

A mandate to meet

A regulator, client contract, or sector rule that requires on-premises processing.

Full control

You hold every key and operate the platform yourself, on your own network.

Ownership

Inference capacity you own instead of rent, with no other tenants. Model bundles and software updates come through your contract with us.

What You Get

The same inference engine, in your building

Fast, low-latency responses on every request - the same speed that powers our cloud, now inside your own building. Because the chips are built for inference, the racks run air-cooled at about 10 kW in typical use: no liquid cooling, no special facility. They fit the datacenter you already have, and you add racks as you grow.

Hardware

  • Purpose-built SambaNova AI accelerators - not GPUs
  • Current generation: SambaRack SN40L
  • Air-cooled, standard rack footprint - no liquid cooling
  • A Core Services Rack for the control plane
  • Sized to your workload, expandable as you grow

Software & Support

  • SambaStack platform (Kubernetes-based)
  • The models we run in Munich + Bring Your Own Checkpoint
  • One contract with Infercom: sizing, installation, configuration and support, delivered together with SambaNova's engineers
  • Air-gapped deployment support
  • Model bundles, software updates and technical support

Your Side

What your team provides

SambaNova's site requirements, in short. We go through each point with you in the site survey.

  • Datacenter space, power (400/415 V three-phase), cooling and network
  • A Kubernetes cluster: your own, or installed on the Core Services Rack
  • PostgreSQL, logging and monitoring (SambaNova recommends Prometheus and Grafana)
  • DNS, NTP, two URLs with TLS certificates, and identity through OIDC, LDAP or Keycloak
  • Regular updates: you stay at most one release behind, so fixes reach you
  • Air-gapped only: an artifact staging server, a container registry and 15 TB of shared storage

One rack, in numbers

Chips
16 RDUs
Memory
1 TB HBM, 4-24 TB DDR
Power
about 10 kW typical, 16.4 kW max
Power feed
400/415 V three-phase, 2 PDUs
Cooling
Air, front to back
Control plane
One Core Services Rack

Source: SambaNova, SambaStack on-premises site requirements (March 2026). Exact figures for your site follow in the site survey.

SambaRack SN40L-16 rack
SambaRack SN40L-16. Image: SambaNova

The Process

From contract to production

  1. 01

    Consultation & site survey

    We assess your requirements, datacenter readiness, and model needs, and size the deployment.

  2. 02

    Delivery and installation

    The racks are delivered and installed in your datacenter, together with SambaNova's engineers.

  3. 03

    Configuration

    SambaStack setup, model bring-up and integration, including air-gapped setups.

  4. 04

    Production

    Your team operates the platform. Model bundles, updates and support come through your contract with us.

Typically ~90 days from contract to production

Timelines vary with hardware availability and datacenter readiness; a firm date follows the site survey.

Models

The full catalog, in your building

Every model we run in Munich runs on-premises too, plus further models SambaStack supports, and your own fine-tunes.

  • The models we run in Munich, plus models brought up on request
  • BYOC - fine-tunes of supported architectures load immediately
  • Models at the precision their makers released them in
  • Air-gapped deployment support

In air-gapped deployments, model and software updates are delivered and installed manually - which needs a staging environment on your side.

Powered by SambaNova

The same dataflow architecture

On-premises runs the same SambaNova dataflow architecture that powers our cloud - purpose-built AI chips delivering up to 10x faster inference than GPUs, air-cooled in a standard rack. You get the same performance; you own the hardware.

Learn about the technology

FAQ

Questions, answered straight

Is on-premises performance the same as your cloud?

Yes - it's the same hardware, so you get exactly the same speeds. What on-premises changes is location and ownership, not performance: the racks sit in your datacenter and you hold the keys.

Does owning hardware actually pay off versus consuming managed tokens?

On-premises is a capital investment with a lead time, versus zero-capex, instant managed tokens. Whether it pays off is case by case - at sustained, heavy utilization over a multi-year horizon, owning the hardware can work out economically, but that depends on how hard you run it, so we'd model it with you rather than promise it. Often the deciding factor isn't cost at all: source code or data that cannot leave your building, a client or regulator that mandates it, or a strategic decision to own the infrastructure. Performance is identical to our cloud either way. If your case has no hard requirement and the economics are marginal, the Inference API or Dedicated Capacity is usually the better call, and we'll tell you so.

Can we test before buying hardware?

Yes, and we recommend it. Prove the value on the shared Inference API first - it runs on the same hardware - then optionally on a dedicated rack in our data center in Munich, before any capital outlay. A small pilot belongs in the cloud, not on-premises.

How big does a deployment need to be?

We size it to your workload: how many users, how many at once, and what they run (short prompts, long context, agents). From that we propose the number of racks, and you can add racks later. If you are still proving the use case, start on the Inference API or Dedicated Capacity.

What do we need in our datacenter - space, power, cooling?

A standard datacenter with high-density power. The racks are air-cooled (no liquid cooling), draw about 10 kW in typical use and up to 16.4 kW, and take 400/415 V three-phase power through two PDUs, in a standard rack footprint. So they fit a datacenter you already have rather than needing a purpose-built facility. Beyond space, power, cooling and network, your team provides the services listed above. We run a site survey to confirm your power, cooling, network and storage before anything ships.

How does an air-gapped deployment work - and how do you update or diagnose it?

SambaStack runs fully offline. In an air-gap, software and model updates are shipped and installed manually: you provide an artifact staging server, a container registry and shared storage, and validate each update under your own security process. Diagnostics happen via file transfer and web conference, since nobody outside has network access to your cluster. It is a proven setup, but an air-gap trades some support speed for total isolation: every update is a deliberate, manual step rather than an automatic pull.

How long does deployment take?

Plan for around 90 days from contract to production, covering hardware delivery, installation, configuration, and model bring-up. Timelines vary with hardware order volume and availability, datacenter readiness on your side, and whether any new model architectures need bringing up - we commit to a firm date only after a site survey.

Which models can we run - and can we bring our own?

Every model we run in Munich runs on-premises, and SambaStack supports further models we can bring up for you. Bring Your Own Checkpoint works too: fine-tunes of supported architectures load immediately. Fine-tuning itself runs on GPUs - you train, then load the weights onto your racks. Support for entirely new architectures is a commercial conversation.

How do model updates work once it's in our building?

Each SambaStack release bundles new models. You choose which to deploy, test them, and apply the allocation yourself - for example shifting racks between two models. Reconfiguring a rack (say, long-context to short-context) takes about 30 minutes. You stay at most one release behind, so fixes reach you. In an air-gap, these updates are delivered and installed manually.

How does the platform handle contention when many users hit it at once?

The inference router lets you prioritize users and rate-limit them, so no single user can take over the cluster. When the cluster is fully loaded, requests queue: the time to first token grows, the output speed of a running request does not. You hold these controls, since you operate the platform.

Can we deploy in our own datacenter, or do you host it?

Both. We can sell the hardware into your datacenter, co-locate it for you, or offer colocation-as-a-service - whichever fits your control and operational requirements.

Can you contractually guarantee that nobody outside our organization can see our data?

On-premises and air-gapped, you hold all the keys, we have no runtime access, and nothing leaves your network. The KV cache lives in memory only and is never written to disk; it is not encrypted in memory.

How self-sufficient are we once it's installed?

Day to day, your team runs it. Model bundles, software updates, fixes and technical support are part of your contract with us, delivered together with SambaNova's engineers, who know the platform down to the chip. If you prefer, we operate it for you as a managed service.

What happens if a rack fails?

With several racks, the others keep serving and you lose only that rack's share of capacity. The control plane runs on three nodes in the Core Services Rack. Each rack has two power supplies.

Can we deploy in several countries?

Yes. Each site is its own deployment. Some countries need an export-control check before the hardware ships; we clarify it during the site survey.

Ready to explore on-premises?

Let's discuss your requirements and datacenter readiness.