NEWEU-Only Inference Endpoint: eu.api.nextbit256.com/v1

Your inference runs in Europe. So does your provider.

OpenRouterAWS MarketplaceAWS Partner NetworkNVIDIA Inception Program1 Million BotNeuroblockOnorato AIRespanConcentrateOpperLanzaderaAngels Capital
Products

Three products. One layer.

LLM Inference

Catalogue

The whole catalogue of open models, behind one single endpoint.

See models and prices

Pay as you go

0tokens

No fixed fees. No minimums.

Pay per useNo commitment
Full compatibility

Change the base URL. Your code stays put.

OpenAI- and Anthropic-compatible: your agent keeps working exactly as it did.

What we run for you

Hardware, balancing and queues. You send an HTTP request.

  • Load balancing · Spread across GPUs
  • Autoscaling · Scales with your traffic
  • Queues · Spikes without errors
  • Fault tolerance · Retries and failover
  • Observability · Usage and spend per key
  • Compliance · Data inside the EU
AI Cloud

Deploy

Pick a card and it's running in minutes. No tickets, no sales calls.

🖥️
Machine ordered·0:00

GPU

Price

Monthly fee

Monthly fee, not tokens.

Configure yours
Access

Root over SSH. With your key.

It grows with you

Add cards or disk without reinstalling. See the fee first.

GPU × 1

Commitment, in months
Monthly fee

See the exact fee on the pricing page.

Pricing
01

You pick the machine

With a GPU or without, and how big: cores, memory and disk. The price moves with you, and it is the price you pay.

02

It builds itself

One to three minutes. We email you when it is ready — no ticket, no call.

03

You get in over SSH

With your key and as root. Whatever you install stays there until you terminate it.

  • Hosted in Spain · EU
Why Nextbit

Why Nextbit: SovereigntyComplianceProven scaleOur own stackPerformanceKV-cache

Not another inference provider.

SPAIN · EU

Sovereignty. Spanish company, hardware in Spain. The CLOUD Act does not reach us.

DPA · MODEL CARDS

Compliance. GDPR and the EU AI Act with documents, not with badges.

TOP 5 · DAILY VOLUME · EU

Proven scale. Partner of OpenRouter, the largest inference AI Gateway in the world.

CACHE-AWARE · SLA-AWARE

Our own stack. We build our own stack: router, scheduler and separated compute phases.

SAME MODEL · BETTER SERVED

Performance. Better than comparable providers on the same model, measured by a third party.

−95% REDUNDANT COMPUTE

KV-cache. When your agent repeats context, we serve it from cache.

The thesis

Downloading a model is the easy part. The hard part comes after: managing GPU memory without fragmenting it, not recomputing the same context thirty times, keeping a long prompt from blocking the short requests behind it, and holding the 99th latency percentile when load spikes on a Tuesday at 11:00. That is inference at scale, and it is the only problem we work on.

The model0%
Running it0%
The stack

We optimize every layer of inference. And deploy it wherever you say.

The same layer, the same team and the same optimizations run in three places, depending on what your case needs.

CLIENTSOVEREIGN AI● SPAIN · EUYOUR CLOUDON-PREMISENON-EU

On our nodes

Our own infrastructure in Spain. You manage nothing: not hardware, not stack, not scaling.

api.nextbit256.com · Running
Regioneu-south · EU
Uptime 90 d99,9 %
Node local time--:--:--
Serverless · Dedicated

In your cloud

We deploy and operate our stack inside your AWS, Azure or GCP account. You leverage your spend commitment and your network policies.

AWSAzureGCP
Runs inYour account onAWSAzureGCP
RegionThe one you choose
eu-west-1us-east-1me-central-1

Final availability depends on the hyperscaler.

Managed Dedicated Inference

On your premises

Real on-premise: physical hardware in your datacenter, zero data egress. Operated end to end by us.

Sovereign deployment
The same API in all three. Starting in one and moving to another is a deployment decision, not a migration.
How we work

One of our engineers inside your project. Not a support ticket.

Inference at scale is not solved by reading docs. It is solved by looking at your real workload and tuning the stack to it. On dedicated and on-premise, an engineer is assigned to your project from sizing to production.

01

Technical session

We analyze your real workload and tell you what hardware and model you actually need. If the answer is "less than you thought", we say so.

02

PoC with success criteria

Latency, cost and quality, measurable and agreed in writing before starting. If they are not met, there is no contract.

03

Production

Deployment, engine calibration to your load profile, and integration with your observability (Grafana, Datadog, Prometheus).

04

Ongoing engineering

Every new kernel, quantization format or better-fitting model: we evaluate it and propose it. We do not wait for you to ask.

Response in 48 h · Deployment in 5 working days from signature

Programs & partners

Where we are and how to buy us

NVIDIA Inception

Members of NVIDIA's program for companies building on its hardware: early access to architectures and engineering support.

Lanzadera + Angels Capital

Lanzadera, Juan Roig's accelerator, and Angels Capital, his investment firm. Nextbit is part of the Valencia program.

Buy with your AWS commitment

AWS Partner Network members, listed on AWS Marketplace. Buy Nextbit Serverless Inference with your AWS spend commitment or credits.

Available for Serverless Inference. Dedicated Inference is contracted directly with Nextbit on our own infrastructure in Spain.

Your first integration, in under 24 hours.

Change your OpenAI client's base URL, test with your real workload, and decide with your data, not ours.

Response in under 48 hours.