Claude Sonnet 4.5 → 5.5 Migration Guide: Deadline, Prices, and 5 Breaking API Changes
Anthropic shipped Claude Sonnet 5.5 on September 28, 2026, and two days later deprecated the model most production workloads are still pinned to: Claude Sonnet 4.5. If your code calls claude-sonnet-4-5-20250929, you have a hard deadline — November 30, 2026 — after which that model ID stops serving requests on the Claude API.
This guide is the practical checklist for that move. It covers what actually breaks, what gets cheaper, and the exact request edits you need, with before-and-after Python for every change.
The deadline that matters
Sonnet 4.5 is not just “older” — it is on a published removal schedule:
| Event | Date |
|---|---|
| Sonnet 4.5 released | September 29, 2025 |
| Sonnet 5.5 released | September 28, 2026 |
| Sonnet 4.5 deprecated | September 30, 2026 |
| Sonnet 4.5 retires (Claude API) | November 30, 2026 |
“Deprecated” means Anthropic recommends you leave; “retires” means the endpoint stops answering. Two months is enough to test, but not enough to discover breakage in production during a busy week. If any workload still points at 4.5, that is the migration with a real clock on it.
Sonnet 4.5 vs 5.5 at a glance
| Spec | Sonnet 4.5 | Sonnet 5.5 |
|---|---|---|
| Model ID (Claude API) | claude-sonnet-4-5-20250929 | claude-sonnet-5-5 |
| Context window | 200K tokens | 1M tokens |
| Max output | 64K tokens | 128K tokens (300K via Batch beta) |
| Input price | $3 / MTok | $2 / MTok |
| Output price | $15 / MTok | $10 / MTok |
| Cache read | $0.30 / MTok | $0.20 / MTok |
| Thinking | Extended (no default effort) | Adaptive (effort defaults to high) |
| Knowledge cutoff | Jan 2025 | Jun 2026 |
| Status | Deprecated → retiring Nov 30, 2026 | Default Sonnet, GA |
The two changes that quietly reshape your architecture are the 5× context jump (200K → 1M) and the thinking-mode default flip (Extended → Adaptive). Both show up in the breaking changes below.
What the price change means for you
Sonnet 5.5 is cheaper per token on every axis. For a representative job of 1M input + 200K output tokens:
- Sonnet 4.5: $3.00 (input) + 0.2 × $15.00 (output) = $6.00
- Sonnet 5.5: $2.00 (input) + 0.2 × $10.00 (output) = $4.00
That is about a 33% lower bill for the same shape of request, before you count the fact that 5.5 often finishes agentic jobs in fewer steps. Cache reads also fall from $0.30 to $0.20 per million tokens, which matters most for long-agentic loops that replay a big system prompt on every turn.
Five breaking API changes
Anthropic documents five ways pre-5.5 request shapes break on Sonnet 5.5. If your 4.5 integration used the older syntax, these hit you too. Each one below has the before/after.
1. Turning thinking off: disabled → between_tools
Sonnet 5.5 ships with adaptive thinking on by default. The old thinking: {type: "disabled"} now returns a 400 telling you to use between_tools.
Before (Sonnet 4.5):
client.messages.create(
model="claude-sonnet-4-5-20250929",
max_tokens=16000,
thinking={"type": "disabled"},
messages=[{"role": "user", "content": "Summarize this thread."}],
)
After (Sonnet 5.5):
client.messages.create(
model="claude-sonnet-5-5",
max_tokens=16000,
thinking={"type": "between_tools"},
messages=[{"role": "user", "content": "Summarize this thread."}],
)
Two caveats: between_tools is only valid at low, medium, and high effort. At xhigh or max you must drop the thinking field (or send {"type": "adaptive"}) and accept up-front thinking. And between_tools takes no extra fields — adding display, budget_tokens, or block_binding to it returns a 400.
2. Forced tool use is gone
tool_choice values "any" and "tool" now return a 400. The replacement is auto plus strict tool use, which forces every call to match your input schema.
Before (Sonnet 4.5):
client.messages.create(
model="claude-sonnet-4-5-20250929",
max_tokens=1024,
tools=tools,
tool_choice={"type": "tool", "name": "get_weather"},
messages=[{"role": "user", "content": "What's the weather in Paris?"}],
)
After (Sonnet 5.5):
client.messages.create(
model="claude-sonnet-5-5",
max_tokens=1024,
# strict tool use: every call matches the tool's input_schema
tools=[{**tool, "strict": True} for tool in tools],
tool_choice={"type": "auto"},
messages=[{
"role": "user",
"content": "What's the weather in Paris? Use the get_weather tool.",
}],
)
Strict tool use requires additionalProperties: false on every object and allows at most 20 strict tools per request. On Amazon Bedrock, structured outputs (which include strict tool use) are not available for 5.5 — send auto without strict and validate tool input in your own code.
3. Thinking blocks are model- and account-bound
Sonnet 5.5 thinking blocks record which model and which conversation produced them. The failure mode that bites production is replay after an edit: if you change the system prompt, the tools, or an earlier message and then resend a cached thinking block, accounts created on or after August 31, 2026 get a 400.
The fix is to keep conversations append-only — use mid-conversation system messages for instruction changes rather than editing the original system field. If you must tolerate mismatches, send the thinking-binding-controls-2026-08-01 beta header and set thinking.block_binding.prefix_mismatch_behavior to "drop_block" (adaptive thinking only). Transitioning a session from Sonnet 5 onto 5.5 is clean — 5.5 reads 5’s blocks — but moving to Opus 5 / 5.5 / Fable / Mythos drops them.
4. The old computer-use tool is rejected
On the Claude API and Google Cloud, computer use now requires computer_toolset_20260801. The older computer_20251124 returns a 400.
Before:
tools=[{"type": "computer_20251124", "name": "computer", ...}]
After:
tools=[{"type": "computer_toolset_20260801", "name": "computer", ...}]
You also need to update your agent loop for member tool_use blocks, batched actions, and toolset_name on results. Amazon Bedrock still accepts computer_20251124 on 5.5, so Bedrock integrations do not need an immediate change.
5. The advisor tool narrowed its accepted advisors
The advisor tool no longer accepts Opus 4.8, Opus 4.7, or Sonnet 5 as advisors. If your 4.5-era config pointed an advisor at one of those, the call fails. Move advisor references to a current 5.x model.
Two more gotchas worth a test
These are not in the “five” list but break real integrations:
- Non-default
temperature,top_p, ortop_kreturn HTTP 400. Many shared clients settemperaturefor every model. Changing only the model string will not remove that field — make sampling config depend on the selected model’s capabilities, then verify with a real request. - Text between tool calls can go silent. On 5.5, longer explanatory text between tool calls is returned as thinking blocks. With the default
display: "omitted"the interface shows nothing and no error is raised. Setthinking.display(adaptive) or switch tobetween_toolsto get the text back. A UI that only listens for ordinary text blocks will look frozen even though the model is working.
Step-by-step migration checklist
- Find every call site. Grep your codebase and any gateway for
claude-sonnet-4-5,sonnet-4-5, and the aliasclaude-sonnet-4-5-20250929. - Swap the model string to
claude-sonnet-5-5and confirmmax_tokenscan go up to 128K if you need it. - Replace
thinking: {type: "disabled"}withbetween_tools(or drop it for adaptive atxhigh/max). - Remove forced
tool_choice. Useauto+strict: true, or restructure the prompt to state when the tool applies. - Audit computer use for
computer_20251124→computer_toolset_20260801(Claude API / Google Cloud). - Make sampling params conditional so
temperature/top_p/top_kare not sent to 5.5 unless intended. - Make conversations append-only if you cache and replay thinking blocks; add the binding beta header if you cannot.
- Update advisor references off Opus 4.8 / 4.7 / Sonnet 5.
- Run an eval on a copied dataset with test tools — not production write tools — and compare correctness, latency, and cost per completed task.
- Keep the old config behind a flag for fast rollback.
Rollback and risk plan
Do not cut over blindly. Keep the 4.5 model string and its request shape in a feature flag, run 5.5 against a staging clone of real traffic, and only shift production once the eval passes. Note that reverting the model does not undo a payment or a sent message, so test against copied data and test tools. For long agentic loops, watch the step where the model re-reads its own context — that is where a silent thinking-block drop or a forced-tool 400 shows up first.
Should you migrate now?
Yes, if the workload touches the Claude API and you cannot re-run evals this week, start the cutover anyway — the November 30 date does not move. The move is cheaper, gives you 5× the context, and the breaking changes are mechanical once you know them. If you have a tightly tuned 4.5 agent that is still meeting its SLA and nobody can validate a 5.5 run this sprint, schedule the eval before the deadline rather than after it.
The good news: tokenization is the same family, so your prompt tokens barely move, and the effort tiers are the main thing to re-tune — start at medium for agentic coding and high for harder multi-step work, then dial in.