The apparent choices are to reserve capacity in advance and risk underutilizing it, or wait and accept price and availability risk in the on-demand market. I’m curious how inference providers, enterprises running fine-tunes or evals, and teams with batch or seasonal workloads handle this in practice.
In particular: - How far forward do you reserve capacity? - How frequently do you end up underusing reservations? - Can providers resize, defer, or release commitments?
I know that larger cloud providers like AWS and CoreWeave have flexixbility/credits, but if you're largely getting bare metal capacity from Neoclouds, how do you handle this?