Your inference runs in Europe. So does your provider.












Three products. One layer.
SERVERLESS INFERENCE
OpenAI- and Anthropic-compatible. Change the base URL and you are serving open models from Europe, per million tokens. Nothing to deploy, no commitment.
See models and pricesDEDICATED ENDPOINTS
Your own endpoint on 100% dedicated resources: the open model you choose, your own, or a fine-tune. Fixed monthly fee, unlimited tokens, and public-API spikes never touch you.
See the tiersAI CLOUD
Whole GPUs for you and virtual machines as sandboxes for code-running agents, on our own infrastructure in Spain. SSH in with your key and pay with the balance you already have.
See the catalogueThe whole catalogue of open models, behind one single endpoint.
See models and pricesPay as you go
No fixed fees. No minimums.
Change the base URL. Your code stays put.
OpenAI- and Anthropic-compatible: your agent keeps working exactly as it did.
Hardware, balancing and queues. You send an HTTP request.
- Load balancing · Spread across GPUs
- Autoscaling · Scales with your traffic
- Queues · Spikes without errors
- Fault tolerance · Retries and failover
- Observability · Usage and spend per key
- Compliance · Data inside the EU
Pick a card and it's running in minutes. No tickets, no sales calls.
GPU
Monthly fee
Monthly fee, not tokens.
Root over SSH. With your key.
Add cards or disk without reinstalling. See the fee first.
GPU × 1
See the exact fee on the pricing page.
You pick the machine
With a GPU or without, and how big: cores, memory and disk. The price moves with you, and it is the price you pay.
It builds itself
One to three minutes. We email you when it is ready — no ticket, no call.
You get in over SSH
With your key and as root. Whatever you install stays there until you terminate it.
- Hosted in Spain · EU
Why Nextbit: SovereigntyComplianceProven scaleOur own stackPerformanceKV-cache
Not another inference provider.
Sovereignty. Spanish company, hardware in Spain. The CLOUD Act does not reach us.
Compliance. GDPR and the EU AI Act with documents, not with badges.
Proven scale. Partner of OpenRouter, the largest inference AI Gateway in the world.
Our own stack. We build our own stack: router, scheduler and separated compute phases.
Performance. Better than comparable providers on the same model, measured by a third party.
KV-cache. When your agent repeats context, we serve it from cache.
Downloading a model is the easy part. The hard part comes after: managing GPU memory without fragmenting it, not recomputing the same context thirty times, keeping a long prompt from blocking the short requests behind it, and holding the 99th latency percentile when load spikes on a Tuesday at 11:00. That is inference at scale, and it is the only problem we work on.
We optimize every layer of inference. And deploy it wherever you say.
The same layer, the same team and the same optimizations run in three places, depending on what your case needs.
On our nodes
Our own infrastructure in Spain. You manage nothing: not hardware, not stack, not scaling.
In your cloud
We deploy and operate our stack inside your AWS, Azure or GCP account. You leverage your spend commitment and your network policies.






Final availability depends on the hyperscaler.
On your premises
Real on-premise: physical hardware in your datacenter, zero data egress. Operated end to end by us.
One of our engineers inside your project. Not a support ticket.
Inference at scale is not solved by reading docs. It is solved by looking at your real workload and tuning the stack to it. On dedicated and on-premise, an engineer is assigned to your project from sizing to production.
Technical session
We analyze your real workload and tell you what hardware and model you actually need. If the answer is "less than you thought", we say so.
PoC with success criteria
Latency, cost and quality, measurable and agreed in writing before starting. If they are not met, there is no contract.
Production
Deployment, engine calibration to your load profile, and integration with your observability (Grafana, Datadog, Prometheus).
Ongoing engineering
Every new kernel, quantization format or better-fitting model: we evaluate it and propose it. We do not wait for you to ask.
Response in 48 h · Deployment in 5 working days from signature
Where we are and how to buy us

NVIDIA Inception
Members of NVIDIA's program for companies building on its hardware: early access to architectures and engineering support.


Lanzadera + Angels Capital
Lanzadera, Juan Roig's accelerator, and Angels Capital, his investment firm. Nextbit is part of the Valencia program.

Buy with your AWS commitment
AWS Partner Network members, listed on AWS Marketplace. Buy Nextbit Serverless Inference with your AWS spend commitment or credits.
Available for Serverless Inference. Dedicated Inference is contracted directly with Nextbit on our own infrastructure in Spain.
Your first integration, in under 24 hours.
Change your OpenAI client's base URL, test with your real workload, and decide with your data, not ours.
Response in under 48 hours.