Off-peak tokens now cost half, and the peak windows are published
Peak and off-peak API rates took effect on 16 August, with off-peak at half of peak. Peak runs 01:00–04:00 and 06:00–10:00 UTC. Inference has started pricing like electricity.
From 16 August the same tokens cost different amounts depending on when you ask for them. Off-peak rates are half of peak. Peak windows are published: 01:00–04:00 and 06:00–10:00 UTC.
This is time-of-use pricing, and the industry it comes from is not software. Electricity is priced this way because generation capacity is fixed in the short run and demand is not, so the price does the work of moving load. Inference now has the same shape — GPUs are the plant, and they cannot be built between breakfast and lunch.
The published windows are the part worth reading closely. Two peaks rather than one, and both inside the European and Asian working day, tells you where the load is coming from. A provider that discloses its peaks is disclosing something about its customer base.
For anyone running batch work the arbitrage is immediate and large. Evaluation runs, backfills, document processing, synthetic data generation — none of that cares what hour it happens in, and all of it just halved. The teams that have a queue will save money. The teams that call the API synchronously from a user request will not.
It also raises a question about the headline price comparisons the whole market runs on. A published rate card with two rates is no longer a single number, and the price war being fought over coding agents gets harder to score.
Digital Applied — AI Model Releases: August 2026 Tracker and Dated Ledger → · AI Release Tracker — Latest AI Model Releases — August 2026 →