Weigh an AI subscription against pay-as-you-go API at your usage level. Subscriptions bill per time window rather than per token, so no exact break-even exists: results are given as ranges, with the cache sensitivity and the date each plan was checked.
Chat is about 1. An agent plans, searches, calls tools and summarises, so one prompt often triggers 5 to 30 calls. That is how a model with a low per-token price can still be expensive per task.
First-party list price, the like-for-like alternative to a first-party subscription.
≈ 4,800 model calls / month at this usage.
The quota costs about 100% less than buying the same volume through the API.
Based on GPT-5.6 Terra and the profile on the left. Cache-write charges are excluded, so real API cost is slightly higher.
The quota costs about 99% less than buying the same volume through the API.
The official quota includes further caps with no published numbers, so both this capacity and the discount above are overestimates.
Based on Claude Sonnet 5 and the profile on the left. Cache-write charges are excluded, so real API cost is slightly higher.
The quota costs about 99% less than buying the same volume through the API.
Based on Claude Fable 5.1 and the profile on the left. Cache-write charges are excluded, so real API cost is slightly higher.
Cost ranges come from cache-hit uncertainty (±15pp). Capacity ranges come from reading each window quota conservatively vs optimistically. The break-even is a region, not a point, because subscription quotas are time windows that cannot be reduced to a fixed token count.
Rolling windows shorter than a day are normalised against an assumed 176 active hours per month (8h × 22 days); day, week and month caps are calendar windows and are not discounted.
These publish no absolute quota, or bill in a unit that cannot be mapped to model calls. They are listed here rather than given an invented number.