$25/mo free credits for every developer · Need it built for you? XAI1 builds complete AI systems end to end →
XAI1 — GPU INFERENCE CLOUD

GPU inference that optimizes itself.

Deploy in minutes. Scale to billions of tokens.

One command turns any open-source Hugging Face model into a production API. XAI1 selects the hardware, quantizes the weights, and tunes the CUDA kernels — you get an endpoint, billed per token.

$25/mo free credits · no credit card required · scale to zero, $0.00 idle
xai1 — deploy
TOKENS / SEC
TTFT
$0.00
IDLE COST
2.4B+ tokens served daily across five modalities
ACME·AIVOXLABNORTHBEAMPIXELFOLDQUERYONHELIOSTAT
// HOW IT WORKS

From model_id to production in three steps.

No GPUs to provision. No CUDA to tune. No servers to babysit.

01

Pick a model

Any open-source model on Hugging Face — LLMs, Whisper-class ASR, TTS, vision, embeddings. Paste the model_id, or bring your own fine-tune.

02

XAI1 optimizes

The engine profiles the model, selects the cheapest hardware that meets your latency target, quantizes where it's lossless-in-practice, and compiles tuned CUDA kernels.

03

Call the API

An OpenAI-compatible endpoint, live in minutes. Serverless autoscaling from zero to thousands of GPUs — billed per token, never for idle.

// MEASURED, NOT MARKETED

The optimization premium, in numbers.

Same model, same silicon. The difference is what XAI1 does to it before your first request lands.

Throughput — tokens/sec per GPU (batch 32, 1×H100)
vanilla transformers
XAI1-optimized ▮
Illustrative pre-launch figures from internal benchmark harness · methodology and reproduction scripts published with GA. Precise numbers per model in docs/benchmarks.
// FIVE MODALITIES, ONE API

Every workload, priced natively.

Per-token pricing breaks down outside LLMs, so XAI1 meters each modality in its own unit. All prices public. Batch −50%. Cached input −80%.

>_

LLMs

Llama, Qwen, DeepSeek, Mistral — chat, completion, JSON mode
from $0.03 /1M input tokens
((●))

Speech-to-text

Whisper-class transcription & translation, word timestamps
$0.05 /audio-hour
◠◡◠

Text-to-speech

Natural voices, cloning-ready, streaming synthesis
$12 /1M characters
[◫]

Vision

Image generation, captioning, detection, OCR
from $0.015 /image
⋮⋮⋮

Embeddings

Retrieval, clustering, rerankers — bge, gte, e5 family
$0.008 /1M tokens
// KNOW THE BILL BEFORE YOU DEPLOY

Estimate your cost in ten seconds.

Auto-selected hardware never means surprise metering — the dashboard shows a pre-deploy estimate and a per-request cost breakdown. Here's the simple version.

120M tokens / month
20% batched
ESTIMATED MONTHLY COST
$—
vs closed-model API: $—  
// PRICING

Free to start. Fair at scale.

Every price is public except enterprise contracts. Usage is metered per modality, and idle time is never billed — on any tier.

HOBBY

$0/mo
For individuals proving an idea. Card only needed past your credits.
  • $25/mo recurring free credits
  • All serverless models, all modalities
  • 3 seats · 5 concurrent GPU workers
  • 24-hour logs
  • Community Discord support
Start free
MOST POPULAR

PRO

$249/mo + usage
For teams shipping to production.
  • $100/mo included credits + volume discounts
  • Unlimited seats · 50 concurrent GPU workers
  • Priority fast-lane inference class
  • 30-day logs + observability dashboard
  • Deployment rollbacks · custom domains
  • SOC 2 report access included
  • Email + Slack support
Start Pro

ENTERPRISE

Custom
For corporations running inference as core infrastructure.
  • Committed-use discounts (10–40%)
  • 99.9%+ uptime SLA with credits
  • HIPAA BAA · data residency · region pinning
  • SSO/SAML + SCIM · RBAC · audit logs
  • VPC, self-host, or BYOC deployment
  • Reserved dedicated capacity
  • Named support engineer
Talk to an engineer
// ENTERPRISE

Production trust, from day one.

Run XAI1 in our cloud, your VPC, or your own metal. Compliance artifacts on request — no twelve-touch sales cycle required.

Talk to an engineer Response from a deploy engineer — not a sales deck — within one business day.
SOC 2 II
audited controls
HIPAA
BAA available
99.9% SLA
with credits
VPC / BYOC
your cloud, our engine
SSO + SCIM
SAML / OIDC
Residency
region pinning
// FAQ

The questions engineers actually ask.

What about cold starts?

Optimized snapshots keep warm-pool cold starts sub-second for most models under 20B parameters. Larger models use predictive pre-warming based on your traffic pattern; Pro's fast-lane class guarantees a warm floor.

Do I pay for idle GPUs?

Never, on any tier. Serverless endpoints scale to zero and bill per token (or audio-hour, character, image). Dedicated endpoints bill per GPU-second of active work only.

How do you choose my hardware — and can I trust the bill?

The engine profiles your model against a live fleet-price table and picks the cheapest configuration meeting your latency target. Before anything deploys, you see a cost estimate; after, every request carries a cost breakdown in the dashboard and API response headers.

Does quantization hurt quality?

Quantization is applied only when it passes automated quality evals against the full-precision baseline (perplexity + task-specific suites). You can pin full precision per endpoint with --precision fp16.

What happens to my data?

Prompts and outputs are never used for training, never shared, and retained only as long as your log window (24h Hobby, 30d Pro, custom Enterprise). VPC and self-host deployments keep data entirely inside your network.

Which model licenses can I serve?

Any model whose license permits your use case — the deploy flow surfaces the model's license (Apache-2.0, MIT, Llama Community, etc.) and flags restrictions before the endpoint goes live.

Deploy in minutes.
Scale to billions of tokens.

$25/mo free credits · no credit card required