Local & private AI · 14 min read

The AI price rises that have already happened.

By James Durkin, JDCS Updated 24 July 2026

This isn't a piece about an AI crash. There's no forecast in here at all. What follows is the published record of what AI has cost over the past two years, taken from vendor pricing pages, regulator filings and dated announcements, so you can look at it and draw your own conclusion. Nobody has to predict a price rise to justify planning for one.

The short version: flagship API prices have risen sharply while cheaper tiers have fallen, seat prices have been raised on the back of bundled AI features in Australia and elsewhere, and several guarantees that used to be purchasable have been withdrawn. None of that is speculation. All of it is on the record, and it tells you which parts of your setup deserve a fallback plan.

A note on currency: every API price below is in US dollars, because that's how the vendors publish them. Australian buyers carry exchange-rate movement on top of every change described here.

The record, with dates

Start with the flagship API tier, which is what most serious builds run on.

OpenAI's flagship went from US$1.25 per million input tokens and US$10 per million output when GPT-5 launched in August 2025, to US$5 and US$30 with GPT-5.5 in April 2026. Four times the input price and three times the output price in under a year. Price cuts announced at the end of July 2026 were real, and they applied to the cheaper tiers; the flagship stayed where it was.

Anthropic has done something unusually transparent, which is to publish the date its mid-tier model gets more expensive. Claude Sonnet 5's introductory rate of US$2 / US$10 runs through 31 August 2026, after which standard pricing of US$3 / US$15 applies. That's a 50% increase on the workhorse tier, announced in advance and printed on the pricing page. Most price changes are not that visible.

Now the other direction, because the record includes falls and an account that only lists rises isn't worth reading. Anthropic's Opus line came down from US$15 / US$75 to US$5 / US$25 with Opus 4.5, a 67% cut on the input price. GPT-4 launched in March 2023 at US$30 / US$60 and its successors arrived at a fraction of that. Cheap tiers have generally got cheaper; flagship tiers have generally got dearer. Which trend you experienced depended entirely on which model you happened to build on.

The rises that never appear on a price list

The published rates are the visible half. The rest moves without a headline.

Anthropic's own pricing page carries the clearest example. It notes that Claude 4.7 and later models use a newer tokeniser, and that it “produces approximately 30% more tokens for the same text”. The price per token did not change. Feed in the same document and the bill is bigger. There's nothing underhand about it, and it's exactly the kind of change that never reaches the person approving the invoice.

Then there's the surcharge layer, which has grown steadily and now runs to a decent list:

  • Long-context multipliers. Google charges double for input and 1.5 times for output beyond 200,000 tokens on its Pro tier; OpenAI applies a 2x multiplier. Anthropic, to be fair, charges no long-context premium on 4.6 and later, which include the full million-token window at standard pricing.
  • Priority and fast tiers. OpenAI's is a 2x multiplier, renamed Fast mode at the end of July 2026. Google's runs around 1.8x.
  • Data-residency uplifts. OpenAI adds about 10% for models released from 5 March 2026. Anthropic applies roughly 1.1x for pinned inference geography, and about 10% through Bedrock and Google Cloud for Sonnet 4.5, Haiku 4.5, Opus 4.5 and future models. Keeping data in a particular country has a price attached to it.
  • Cache write premiums. Anthropic charges 1.25 times the base rate for a five-minute cache write and double for a one-hour write, which pays for itself after two reads and costs you if it doesn't get them.
  • Storage by the hour. Google's context cache storage runs at US$1.00 per million tokens per hour, which is a billing dimension that didn't exist a couple of years ago. Neither did per-session-hour charging for managed agents.

Individually, each of these is defensible. Collectively they mean the headline rate has become a starting point rather than a price.

Why the bill can rise while every published price falls

This is the mechanism most finance conversations miss, and it explains more variance than everything above put together.

Reasoning models think before they answer, and that thinking is billed as output. One published analysis of the same workload run across different models found it cost US$10.31 on one model and US$522.48 on another, at broadly comparable headline rates, because the second generated roughly twelve times as many completion tokens. Same job, same nominal price per token, a fifty-fold difference in the invoice.

The clean way to hold both facts: the price of a unit of intelligence is falling fast, and the amount of intelligence consumed per job is rising faster. Which of those dominates your bill depends entirely on how the work is architected. A budgeting rule of thumb that survives contact with reality: when moving to a reasoning model, expect three to five times the output tokens you'd have planned for, and check the actual usage after a fortnight instead of trusting the estimate.

Seat prices, and the Australian case that went to court

Per-seat software is where most small businesses actually meet AI pricing, and the pattern there is more legible than the API market.

In January 2025 Microsoft bundled Copilot into consumer Microsoft 365 and raised Australian prices at the same time. Personal went from A$109 to A$159 a year, a 45% rise. Family went from A$139 to A$179, a 30% rise. Cheaper “Classic” plans existed which kept the existing features without Copilot. The ACCC subsequently sued Microsoft, alleging it misled subscribers by not disclosing that option. Microsoft acknowledged that its communication had been poor and offered affected subscribers an eight-week refund window. The allegations have not been determined by a court, and the sequence itself is a matter of public record.

The same movement shows up elsewhere without a regulator involved. Microsoft raised commercial list prices by between about 5% and 33% on 1 July 2026. Google raised Workspace prices by 17% to 22% in January 2025, when it bundled Gemini in.

Name the pattern and it becomes easy to spot: the add-on you can say no to becomes the bundle you cannot. An AI feature arrives as an optional extra at a separate price, adoption is measured, and at the next renewal it's part of the base product and the base price has moved. That has now happened across the two largest productivity suites in the world, in the same eighteen months, in this country.

What can be withdrawn, and how much notice you get

Price is only one of the terms. The others move too.

  • Model retirement. Anthropic's published commitment is at least 60 days' notice. Its last four retirements ran at 60, 61, 62 and 62 days, so the floor has become the norm. Claude Opus 4.1 shipped in August 2025 and was retired on 5 August 2026: about twelve months from flagship to switched off.
  • The stated reason. Anthropic says it out loud, which is more than most: models are retired “to ensure capacity for new model releases”. Retirement is a compute decision rather than a product one, which tells you it will keep happening while compute stays scarce.
  • Guarantees you used to be able to buy. Anthropic's Priority Tier, the paid capacity commitment, is now described as “no longer available for purchase”.
  • What your rate limits actually mean. Anthropic's documentation puts it plainly: limits “represent maximum allowed usage, not guaranteed minimums”. Read that before you build a customer-facing product on a standard tier.
  • Fine-tuning. From 6 January 2027, OpenAI customers cannot create new fine-tuning jobs at all. A business whose edge was a fine-tuned model keeps the model it has and loses the ability to retrain it, which means the asset stops being maintainable.

None of this is misconduct. It's a young industry rationing scarce compute and tidying its product line, and every vendor listed here has done things that made customers better off as well. The point is narrower: these are the terms, they change, and a business that has assumed otherwise has an unpriced dependency sitting in the middle of its operations.

A short note on the wider market

Two institutions worth listening to have commented, and their words are worth quoting exactly, not paraphrased into something more dramatic.

The Bank of England's Financial Stability Report of December 2025 said that equity valuations in the United States are “close to the most stretched they have been since the dot-com bubble”. The Bank for International Settlements, in its annual report of June 2026, noted that the terms of chip and compute deals between suppliers, labs and cloud providers are “typically poorly disclosed, with risks of the same asset being pledged multiple times”.

That's the whole of it. Neither institution forecast a collapse, and neither will I. It's context for the record above, not a prediction to plan around.

Capping the exposure

The useful response to all of this is architectural and fairly boring, which is generally a good sign.

  1. Keep an abstraction layer. Your application should talk to one internal interface, with the model behind it as a configuration setting. Switching provider then costs an afternoon instead of a project.
  2. Keep prompts and evaluations in your own repository. Prompts stored inside a vendor's platform are a dependency you'll have to reimplement when that platform changes. Evaluations are what tell you whether the replacement model is actually as good, and without them a migration is guesswork.
  3. Keep a second provider genuinely tested. Configured is not tested. Run your evaluation set against the alternative once a quarter so that switching is a decision rather than an emergency.
  4. Own your data. Your documents, your extracted fields, your embeddings and your fine-tuning datasets should live somewhere you control. Rebuilding those is the expensive part of any migration.
  5. Never let one vendor's roadmap become your product roadmap. If a feature you sell depends on a platform capability with a retirement schedule, you've inherited someone else's deprecation calendar.

There's one more option worth naming, since it removes the exposure instead of managing it. Where the work is a single-pass transformation, summarising, extracting fields, classifying, transcribing, the same job usually runs perfectly well on a model you host yourself, at a cost that cannot be repriced by anyone. That's the argument in what open models are actually good enough for, and the numbers behind it are in what local AI actually costs. It doesn't suit every workload, and where it fits, the pricing question stops applying.

Bottom line: the flagship tier quadrupled in price in under a year, a dated 50% rise on a mid-tier model is printed on a vendor's own pricing page, a tokeniser change added about 30% to the same document, and two productivity suites bundled AI in and raised Australian prices. That's the record, not a forecast. Build so that any single vendor's next decision is an inconvenience rather than a crisis.

Worried about what your AI bill does next?

The first conversation is free. You'll get an honest read on where you're exposed to one vendor, what a fallback would take, and which parts of the work could run on hardware you own. See AI consulting or pricing for how it works.

Start a conversation

Pricing questions, answered.

Have AI API prices gone up or down?
Both, in different places. OpenAI's flagship API went from US$1.25 per million input tokens in August 2025 to US$5 in April 2026, four times the price in under a year. Over the same period Anthropic's Opus line fell from US$15 to US$5 per million input tokens, a 67% cut. Cheaper tiers have generally got cheaper and flagship tiers have generally got dearer, so which direction you experienced depends on which model you were using.
Why did my AI bill go up when the published prices went down?
Two common reasons. Newer models emit far more billable tokens: one published study found the same workload cost US$10.31 on one model and US$522.48 on a reasoning model at broadly comparable headline rates. And tokenisation changed. Anthropic's pricing page notes the newer tokeniser used by Claude 4.7 and later produces approximately 30% more tokens for the same text, so an unchanged document costs more to process at an unchanged price per token.
Is Microsoft 365 more expensive in Australia because of Copilot?
In January 2025 Microsoft bundled Copilot into consumer Microsoft 365 and raised Australian prices, Personal from A$109 to A$159 a year and Family from A$139 to A$179. The ACCC has since sued Microsoft, alleging it misled subscribers by not disclosing cheaper Classic plans that kept the existing features without Copilot. Microsoft acknowledged its communication was poor and offered an eight-week refund window. Those allegations have not been determined.
How much notice do you get before an AI model is retired?
Anthropic's published commitment is at least 60 days, and its last four retirements ran at 60, 61, 62 and 62 days, so the floor has become the norm. Claude Opus 4.1 shipped in August 2025 and was retired on 5 August 2026, about twelve months from flagship to gone. Anthropic states the reason plainly: models are retired to ensure capacity for new model releases.
How do I protect my business from AI price rises?
Keep an abstraction layer so the model is a setting rather than a rewrite. Keep your prompts and evaluations in your own repository. Keep a second provider actually tested, not just configured. Own your data. And where the work is a single-pass job like summarising or extracting, the same task often runs fine on a model you host yourself, which caps the exposure entirely.