AIAcademy · AIAcademy · 2026-05-16
Every frontier model is now priced in dollars per million tokens, split into input and output. The headline numbers look small. The bill is rarely small. Here is how to read one.
Output dominates. Across the major labs, output tokens cost 3–8× input tokens — usually closer to 4–5×. Claude Opus 4.7 is $5 input / $25 output per million. GPT-5.5 is $5 / $30. A long agent run that ingests a 50K-token document and emits a 5K-token answer pays roughly the same for those 5K output tokens as for all 50K input. If your usage is interactive chat with short replies, input dominates. If your usage is anything agentic — multi-step reasoning, tool calls, code generation — output dominates the bill. Most enterprise spend in 2026 is the second category.
Big price cuts are real and recent. Anthropic dropped Opus pricing 67% in February 2026 — from $15 input / $75 output per million to $5 / $25 — for the same flagship class. The driver was AWS Trainium economics and the Project Rainier capacity coming online; the practical effect was that production workloads built against Sonnet pricing six months earlier suddenly had a flagship-quality option in the same envelope.
The hidden discounts move the answer. Prompt caching cuts repeated-input cost by up to 90%. Batch APIs cut both sides by 50% for non-realtime workloads. Context-window discounts apply on cached document grounding. A naive sticker-price comparison between two models can be inverted by which one has working cache primitives in your stack. Read the pricing page and the caching docs.