All systems operational
MModelCloud
Pricing

Pay for the GPUs you run. Nothing else.

No per-token markup on your own models, no platform fee, and no charge for idle. Bring your weights; you're billed for the compute they use.

Starter

For evaluation and side projects.

$0 / month
no credit card required
  • 40 GPU-hours / month
  • 2 concurrent models
  • Shared capacity
  • Autoscaling to zero
  • Community support
Start free

Scale

For committed, steady-state capacity.

Reserved
committed-use discount
  • Reserved GPU capacity
  • Discounted GPU-hour rate
  • Warm pools always on
  • Multi-region serving
  • 99.95% SLA
Talk to sales

Enterprise

For regulated, high-throughput workloads.

Custom
annual agreement
  • Dedicated GPU fleet
  • VPC peering
  • Compliance: SOC 2, HIPAA, GDPR
  • SSO and audit logs
  • Dedicated support
Talk to us
Compute rates

Representative per-GPU-hour pricing

On-demand managed inference, billed to the second. Scale-to-zero means you're never billed for an idle replica.

AcceleratorMemoryBest forOn-demandReserved
L424 GBembeddings, small models$0.70 / hr$0.42 / hr
A10G24 GB7–8B models, high concurrency$1.10 / hr$0.66 / hr
L40S48 GB13–30B models, best price-per-token$1.90 / hr$1.14 / hr
A100 · 40 GB40 GB13–30B models$2.20 / hr$1.32 / hr
A100 · 80 GB80 GB30–70B models, tensor-parallel$3.50 / hr$2.10 / hr
H100 · 80 GB80 GB70B+ models, lowest latency$5.10 / hr$3.06 / hr

Rates are representative for on-demand managed inference. Reserved pricing locks a fixed rate against a minimum commit. Spot capacity is available at a further discount for batch, offline, and evaluation workloads that can tolerate interruption.

FAQ

Common questions

Do you charge a per-token fee?

No. You're running your own model, so there's no model provider taking a cut. You pay for GPU compute and minimal egress — nothing else.

What counts as billable compute?

Time a replica is loaded and serving. When an endpoint scales to zero, billing stops. Warm resumes are part of the replica's run time, not billed separately.

Do I need to move my weights to you?

No. Weights can be pulled from Hugging Face, a private registry, or your own object storage at deploy time. You keep full ownership and can remove them at any time.

How does quantization affect my bill?

Directly — a smaller memory footprint can fit the same model on a cheaper GPU. An 8B model that needs an A100 at BF16 might fit an A10G with INT4 at roughly half the hourly rate.

How is multi-GPU billed?

Per accelerator, by the second. A tensor-parallel deployment split across two A100s runs at twice the single-GPU rate, with no parallelism overhead fee on top.

Do you offer spot capacity?

Yes, at a discount for workloads that can be interrupted — nightly evals, bulk inference, and offline fine-tuning runs.

Start free. Scale when you're ready.

Deploy your first model on the starter tier and move to reserved or dedicated capacity as traffic grows.

Get started