What shipped, in what version, and when — since the private beta in 2023.
Autoscaling can now follow a weekly schedule, so fleets spin up for known peaks — first shift, settlement windows, a launch — and drop to zero outside them.
Dedicated and reserved GPU fleets graduate to their own product, with VPC peering and committed-use pricing.
Managed inference for open-weights and fine-tuned models is generally available, with the first version of the pricing model.
Managed inference for open-weights models, offered to a small set of early teams.