Run your own open-weights and fine-tuned models on a managed GPU fleet — inference, training, and observability — without operating a single server. No per-token markup, no walled garden.
ModelCloud console — deployments & fleet
Each product stands alone or composes with the others — from a single serving endpoint to a compliance-ready dedicated fleet.
Serve any open-weights or fine-tuned model behind one OpenAI-compatible endpoint.
Fine-tune and train across many GPUs with checkpointing and deterministic resume.
Latency, throughput, token accounting, and cost attribution per model and deployment.
Bring your model, pick a GPU, and ModelCloud stands up the endpoint, autoscaling, and telemetry — so you deploy in minutes instead of operating servers.
Deployment console — deploy and scale a model
Start free, scale per GPU-hour, or commit to reserved and dedicated capacity at a discount. No per-token markup on your own models.
For evaluation and side projects.
For serving models in production.
For committed, steady-state capacity.
For regulated, high-throughput workloads.
| Accelerator | Memory | Best for | On-demand | Reserved |
|---|---|---|---|---|
| L4 | 24 GB | embeddings, small models | $0.70 / hr | $0.42 / hr |
| A10G | 24 GB | 7–8B models, high concurrency | $1.10 / hr | $0.66 / hr |
| L40S | 48 GB | 13–30B models, best price-per-token | $1.90 / hr | $1.14 / hr |
| A100 · 40 GB | 40 GB | 13–30B models | $2.20 / hr | $1.32 / hr |
| A100 · 80 GB | 80 GB | 30–70B models, tensor-parallel | $3.50 / hr | $2.10 / hr |
| H100 · 80 GB | 80 GB | 70B+ models, lowest latency | $5.10 / hr | $3.06 / hr |
Rates are representative for on-demand managed inference. Reserved pricing locks a fixed rate against a minimum commit. Spot capacity is available at a further discount for batch, offline, and evaluation workloads that can tolerate interruption.
The same fleet, configured for the constraint that matters to you — latency, isolation, compliance, or cost.
Sub-second fraud scoring with VPC isolation and multi-region failover.
HIPAA posture, private connectivity, and traceable model versioning.
Shift-aware autoscaling and per-line cost attribution for vision models.
Bulk document review on spot capacity, kept inside a VPC.
Burst-ready autoscaling and warm pools for launch-day traffic.
High-concurrency search and recommendations with per-request tracing.






Start on the free tier, move to reserved or dedicated capacity as your traffic grows — and keep your weights yours the whole way.