Skip to content

Operations

The upgrade is the risk: what happens to a live AI system when the frontier moves

Most writing about frontier models is written for the moment before you build. This is about the eighteen months after. A system that is live has three exposures to the frontier moving: the model can change underneath it, the economics can change without the model changing at all, and the model can be switched off. In one week in September all three were on display. None of them appears in a benchmark, and none of them is visible to anyone who stopped paying attention at go-live.

Author
Luka Kokot, Founder and Chief Executive, Applicat AI
Published
Reading time
9 min read

Benchmarks, context windows and a price per million tokens are the right things to look at in the week before you build. They also have a short shelf life, because the interesting part of a production system is the eighteen months afterwards, when it is running, somebody depends on it, and the ground underneath it keeps moving.

One week in September put all three of those exposures on the table at once, which makes it a useful worked example.

The model changes underneath you

On 22 September Anthropic released Claude Opus 5.5 at $4 per million input tokens and $20 on output, down from $5 and $25 for Opus 5. Cheaper, with a one million token context window and adaptive thinking always on. On paper this is the easiest upgrade decision anyone will make this quarter.

The same release notes list three call shapes that worked on Opus 5 and now return a 400. Thinking can no longer be disabled, so a system that switched it off to keep a step fast and deterministic gets an error. A tool_choice of any or tool is refused, so an agent that forces a tool call to guarantee a structured reply gets an error. And computer use on the Claude API and on Google Cloud requires a new toolset, while the older one is rejected, except on Amazon Bedrock where it keeps working.

Read that last clause again, because it is the one that catches people. The same model, released on the same day, behaves differently depending on which cloud you reached it through. A team on Bedrock and a team on the Claude API are not doing the same migration, and if both exist inside one organisation they will report different results from the same test.

What makes this class of change awkward is that it fails loudly and late. A model that gets slightly worse at your task shows up as a drift in a score, if someone is scoring. A model that returns a 400 shows up as a broken workflow at nine in the morning, and only in the paths that use the affected call shape, which may not be the paths anyone demonstrated.

The economics change while the model stays the same

On the same day, OpenAI put GPT-6 Sol and GPT-6 Luna underneath GPT-6 Astra. All three carry the same 1.05 million token context window and the same 128,000 token maximum output. Astra is $10 per million input tokens and $50 on output. Sol is $2 and $10. Luna is $0.10 and $0.50.

A hundredfold spread on input, at identical context. That turns a question which used to be architectural into a purely commercial one, and it is not a question you can answer from a leaderboard. Whether your extraction task survives on the cheapest tier depends on your documents, your tolerance for a wrong field, and what a human review step costs you. The only way to find out is to run your own cases on all three and look at what breaks.

Then there is the cliff. Above 272,000 input tokens the entire request is billed at doubled input and cache rates and one and a half times output. Not the tokens past the threshold. The whole request.

This is the one that arrives without anybody doing anything. A retrieval step that pulled 250,000 tokens at launch pulls 280,000 a year later, because the corpus grew, which is what corpora do. Nothing was deployed. No decision was taken. The bill for that request roughly doubles, and the only signal is a line on an invoice that somebody may or may not be reading. A system can get quietly more expensive purely because the business it serves got bigger.

The model goes away

The third exposure is the one nobody plans for, because at the point you choose a model it feels permanent. It is not. Both major labs publish retirement schedules, and the notice periods are worth knowing before you need them.

Anthropic commits to notifying customers with active deployments at least 60 days before retiring a publicly released model. OpenAI commits to at least six months for generally available models, at least three for specialised variants, and as little as two weeks for previews. Those are both reasonable policies. They are also different by a factor of three, which matters the moment you are running on more than one lab, because your real migration window is the shortest commitment among the models you depend on. Teams discover this in the wrong order.

The schedules make it concrete. On 23 October 2026 OpenAI shuts down gpt-4, gpt-3.5-turbo, o1, o1-pro and o3-mini, announced six months earlier. On 11 December 2026, gpt-5-2025-08-07 stops answering. That model was released in August 2025. A system built on it in its first month gets about sixteen months before the thing it was built on stops existing, with six months of warning in the middle. On the Anthropic side, claude-opus-4-20250514 was deprecated on 14 April 2026 and retired on 15 June, sixty-two days later.

None of this is a complaint about the labs. Retiring old models is how the price of the new ones stays sane. It is a statement about what owning a production system actually involves, which is that a dependency you did not write and cannot control has a published expiry date, and that date is closer than the depreciation schedule of every other piece of software in the building.

Why none of this shows up in a benchmark

Benchmarks measure a model against a public test set at a moment in time. All three exposures above are properties of the relationship between a model and a running system: an interface contract, a billing rule, a lifecycle date. A model can top every leaderboard and still be the wrong thing to have underneath you, because the version you integrated is deprecated, or the call shape you rely on has been removed, or the way you use it crosses a pricing threshold nobody mentioned in the announcement.

This is why we keep a public log of what moves at the frontier and what each change means for something already live. Not because the releases are interesting, though some are, but because the release note is usually where the operational news is hiding, several paragraphs below the benchmark chart.

What a team should actually do

  1. Pin the model string. Never run a production path on a floating alias. You want the version you tested, and you want a deliberate act to change it.
  2. Keep a set of your own cases, scored, with the current model's results recorded. This is the only instrument that tells you whether a new model is better at your work, and it takes an afternoon to start.
  3. Re-run that set on every model change, every prompt change and every retrieval change, before a user meets the result. Treat a price cut as a reason for more testing.
  4. Read the release notes for breaking changes before the benchmarks. On a multi-cloud estate, test on each platform you actually use, because the same model can behave differently on each.
  5. Put a ceiling on the context you send and alert when a request approaches a pricing threshold. Watch cost per transaction as closely as accuracy.
  6. Keep a register of every model string in production, with its published retirement date, and review it monthly. The shortest date on that list is your real planning horizon.

None of that is clever. It is the ordinary discipline of running something, applied to a dependency that changes faster than the rest of the estate. The reason it often does not happen is not difficulty, it is ownership: the project that built the system ended, the team moved on, and nobody was ever made responsible for the months in which the ground moved.

Sources

  1. Anthropic: Claude Platform release notes, 22 September 2026
  2. Anthropic: Model deprecations and retirement dates
  3. OpenAI: API deprecations and shutdown dates
  4. OpenAI API pricing for GPT-6 Astra, Sol and Luna
  5. GPT-6 Sol and Luna: pricing, release date and API access

Questions on this topic

Bring the frontier into production.

Tell us about the process you want to change. We reply within one business day.