All systems operational
MModelCloud
Changelog

Release notes.

What shipped, in what version, and when — since the private beta in 2023.

v3.2.0
2026-07-21
autoscaling

Shift-aware autoscaling schedules

Autoscaling can now follow a weekly schedule, so fleets spin up for known peaks — first shift, settlement windows, a launch — and drop to zero outside them.

  • Recurring schedule rules in addition to metric-based scaling
  • Cost attribution API for per-deployment and per-tag spend
  • Scale-to-zero cooldown configurable down to 30 seconds
v3.1.0
2026-03-09
compliance

Compliance posture: SOC 2 & HIPAA

  • SOC 2 Type I report available under NDA
  • HIPAA posture for healthcare workloads on dedicated fleets
  • Audit logs and role-based access controls
v3.0.0
2025-11-14
major

Fleet product & reserved capacity

Dedicated and reserved GPU fleets graduate to their own product, with VPC peering and committed-use pricing.

  • Reserved capacity at a discounted GPU-hour rate
  • VPC peering and private connectivity
  • Enterprise tiers with dedicated support and SSO
v2.2.0
2025-06-17
runtime

Quantization: FP8, INT8, INT4

  • FP8, INT8, and INT4 quantization applied at deploy time
  • Automatic quantization matched to GPU memory and model size
  • KV-cache FP8 dtype to roughly double cache capacity
v2.0.0
2025-01-22
major

Training product enters beta

  • Distributed fine-tuning and training jobs
  • Spot capacity with automatic rescheduling on interruption
  • Deterministic resume from checkpoint
v1.4.0
2024-10-08
observability

Observability dashboard

  • Per-model latency, throughput, and GPU utilization
  • Token and cost accounting per deployment
  • OTLP export to Datadog, Grafana, and Prometheus
v1.3.0
2024-08-02
network

VPC peering & role-based access

  • VPC peering for private connectivity
  • Workspace roles and quotas
  • Multi-region serving
v1.0.0
2024-02-16
major

General availability

Managed inference for open-weights and fine-tuned models is generally available, with the first version of the pricing model.

  • Autoscaling to zero
  • Warm pools for cold-start reduction
  • OpenAI-compatible endpoint
v0.4.0
2023-09-28
autoscaling

Private beta: autoscaling & cold starts

  • Autoscaling to zero on request volume
  • Warm-container resumes
  • CLI deploy and logs
v0.1.0
2023-04-11
major

Private beta

Managed inference for open-weights models, offered to a small set of early teams.

  • Single-GPU serving
  • vLLM runtime
  • Bearer-token authentication