modelcloud.cloud Model hosting & inference cloud
MModelCloud
Model hosting & inference cloud

You bring the model.
We run it.

Run your own open-weights and fine-tuned models on a managed GPU fleet — inference, training, and observability — without operating a single server. No per-token markup, no walled garden.

Free tier OpenAI-compatible API Autoscale to zero
https://console.modelcloud.cloud/deployments
The ModelCloud console showing the deployment matrix and GPU fleet overview ModelCloud console — deployments & fleet

Trusted by platform teams in fintech, healthcare, robotics, legal, media, and e-commerce

Products

Three products, one GPU fleet

Each product stands alone or composes with the others — from a single serving endpoint to a compliance-ready dedicated fleet.

Plus, Fleet — dedicated and reserved capacity with VPC peering, quotas, and compliance posture (SOC 2, HIPAA, GDPR). View fleet →
How it works

From weights to a running endpoint

Bring your model, pick a GPU, and ModelCloud stands up the endpoint, autoscaling, and telemetry — so you deploy in minutes instead of operating servers.

https://console.modelcloud.cloud/deployments
The ModelCloud deployment console for creating and managing a new model deployment Deployment console — deploy and scale a model
Pricing

Pay for the GPUs you run

Start free, scale per GPU-hour, or commit to reserved and dedicated capacity at a discount. No per-token markup on your own models.

Starter

For evaluation and side projects.

$0 / mo
no credit card required
  • 40 GPU-hours / month
  • 2 concurrent models
  • Shared capacity
  • Autoscaling to zero
  • Community support
Start free

Scale

For committed, steady-state capacity.

Reserved
committed-use discount
  • Reserved GPU capacity
  • Discounted GPU-hour rate
  • Warm pools always on
  • Multi-region serving
  • 99.95% SLA
Talk to sales

Enterprise

For regulated, high-throughput workloads.

Custom
annual agreement
  • Dedicated GPU fleet
  • VPC peering
  • Compliance: SOC 2, HIPAA, GDPR
  • SSO and audit logs
  • Dedicated support
Talk to us
Representative per-GPU-hour pricingbilled to the second
AcceleratorMemoryBest forOn-demandReserved
L424 GBembeddings, small models$0.70 / hr$0.42 / hr
A10G24 GB7–8B models, high concurrency$1.10 / hr$0.66 / hr
L40S48 GB13–30B models, best price-per-token$1.90 / hr$1.14 / hr
A100 · 40 GB40 GB13–30B models$2.20 / hr$1.32 / hr
A100 · 80 GB80 GB30–70B models, tensor-parallel$3.50 / hr$2.10 / hr
H100 · 80 GB80 GB70B+ models, lowest latency$5.10 / hr$3.06 / hr

Rates are representative for on-demand managed inference. Reserved pricing locks a fixed rate against a minimum commit. Spot capacity is available at a further discount for batch, offline, and evaluation workloads that can tolerate interruption.

Solutions

Run models the way your industry needs them

The same fleet, configured for the constraint that matters to you — latency, isolation, compliance, or cost.

Company

Built by people who ran production inference

Dana Whitfield

Dana Whitfield

Co-founder & CEO
Marcus Bell

Marcus Bell

Co-founder & CTO
Priya Nair

Priya Nair

Head of Engineering
Kenji Watanabe

Kenji Watanabe

Head of ML
Amara Diallo

Amara Diallo

Head of Product
Lucas Meyer

Lucas Meyer

Head of Infrastructure

Bring your model. We'll run it.

Start on the free tier, move to reserved or dedicated capacity as your traffic grows — and keep your weights yours the whole way.

Get started Talk to sales