Dateline: October 2, 2026 | Next update: October 9, 2026
The busiest week of the quarter, and the three vendors managed to launch on consecutive days without colliding. Anthropic shipped Claude Sonnet 5.5 on Monday the 28th. OpenAI held DevDay on Tuesday the 29th and announced more than twenty things at once. Google published Gemini 4 Argon on Wednesday the 30th. GitHub Copilot, which has to absorb whatever the labs ship, had two of the three models in its picker within a day of each release — and on Thursday it gained the ability to drive desktop applications.
If you read only one section this week, read the first: Sonnet 5.5 carries five changes that break code written for Sonnet 5, and Sonnet 4.5 now has a retirement date.
Claude / Anthropic
★ Claude Sonnet 5.5 ships, with five breaking changes
Sonnet 5.5 keeps the shape of its predecessor — one million tokens of context, 128K maximum output, $2 input and $10 output per million — and changes how it behaves underneath. Prompt cache reads are $0.20 per MTok, five-minute cache writes $2.50, one-hour writes $4. Adaptive thinking is on by default at high effort.
Five things that work on Sonnet 5 fail on Sonnet 5.5. To turn off up-front thinking you now send thinking: {"type": "between_tools"} rather than "disabled", and only at high effort or below. Forced tool use — tool_choice of type any or tool — returns a 400 error outright. Thinking blocks are tied to both the model and the conversation. The older computer_20251124 computer use tool is rejected on the Claude API and Google Cloud. And the advisor tool refuses Opus 4.8, Opus 4.7 and Sonnet 5 as advisors.
Model ID: `claude-sonnet-5-5` on the Claude API, Google Cloud, Microsoft Foundry and Claude Platform on AWS; `anthropic.claude-sonnet-5-5` on Bedrock | Context: 1M tokens | Max output: 128K synchronous, 300K on the Batches API with `output-300k-2026-03-24` | Pricing: $2 input / $10 output per MTok; cache read $0.20; 5m cache write $2.50; 1h cache write $4; Batch API 50% off | Thinking: adaptive, default effort `high`; lowest setting is `between_tools` | Also note: `temperature`, `top_p` and `top_k` return 400 if set to a non-default value; minimum cacheable prompt is 512 tokens | Knowledge cutoff: June 2026 | Retirement: not sooner than September 28, 2027
Best for: Most production traffic, once you have worked through the five changes. Do not treat this as a drop-in version bump — read the migration guide before changing the model string.
⚠ A sixth change breaks nothing and will still surprise you
Separate from the five errors, one change alters the response shape without failing any request: text between tool calls now comes back inside thinking blocks rather than as plain text.
An application that streams that commentary to its users will simply go quiet between tool calls. Nothing errors, nothing logs, and the agent looks like it has hung while it is working normally. The fix is either to set a `display` value that returns the text, or to turn off up-front thinking with `between_tools`. This is the kind of change that reaches you through a support ticket rather than a stack trace, which is why it is worth knowing before you migrate.
Behaviour: text between tool calls is returned in `thinking` blocks | Failure mode: none — requests succeed, but streamed interim text disappears from the user's view | Remedies: set a `display` value that returns the text, or send `thinking: {"type": "between_tools"}` to turn off up-front thinking | Related: thinking blocks produced by Sonnet 5.5 only work in the account that produced them, or a linked account; another account's blocks are silently dropped before the model sees them and the request still succeeds
Best for: Anyone whose product shows the model's intermediate reasoning to end users. Test the streaming path specifically, not just the final output.
⚠ Claude Sonnet 4.5 retires on November 30
Anthropic put a date on Sonnet 4.5. Two months' notice, with Sonnet 5.5 named as the migration target.
The awkward part is the overlap. The replacement for a model being retired on November 30 is a model that shipped two days before the deprecation notice and carries five breaking changes of its own. If you are still on 4.5, you are not doing a version bump — you are doing a migration, and the window is eight weeks.
Deprecated: `claude-sonnet-4-5-20250929` | Announced: September 30, 2026 | Retirement: November 30, 2026 | Migration target: Claude Sonnet 5.5 | Guidance: see the model deprecations page
Best for: Put this in the calendar now. Eight weeks is comfortable if you start this month and tight if you start in November.
★ Claude Mods — plugins that change behaviour, not just add commands
Claude Code plugins could add commands, agents and skills. Mods let them modify deeper behaviour. Anthropic shipped one built in, called "You should know" — a side agent that watches the session and flags things you or Claude might have missed, enabled with /plugin enable cc-plugin-you-should-know.
The same release changed something quieter that matters more to cost: Opus 4.7 and later, plus Fable, now use a one-million-token context window by default on Bedrock, Vertex, Foundry and the Claude apps gateway — with no `[1m]` suffix required. If your long-context behaviour was gated behind that suffix, it is now the default, and so is the pricing that goes with a long prompt.
Version: 2.1.287 (published October 1, 2026) | Claude Mods: plugins may modify deeper behaviour than commands, agents and skills | Built-in mod: "You should know" — a side agent that flags what you or Claude might miss; enable with `/plugin enable cc-plugin-you-should-know` | Context default: Opus 4.7+ and Fable use a 1M context window by default on Bedrock, Vertex, Foundry and the Claude apps gateway, no `[1m]` suffix | Opt-out available via `CLAUDE_CODE_DISABLE_*` environment variable | Also 2.1.288 (October 2): `--max-findings
Best for: Teams who had hit the ceiling of what a plugin could change. Check the context-window default against your Bedrock or Vertex bill first.
★ Administrators can restrict which API providers a machine may use
A managed setting named allowedProviders now limits which API providers a given machine is permitted to reach — the Anthropic API, a custom endpoint, Bedrock, Mantle, Vertex AI, Foundry or the Claude apps gateway.
Two other controls landed alongside it. `CLAUDE_CODE_DISABLE_WEB_FETCH` turns the WebFetch tool off entirely, which is the straightforward answer for an environment where outbound fetching is not acceptable. And `claude --desktop` opens the Claude desktop app on the current directory, or on a specific session with `--continue` or `--resume
Version: 2.1.285 (published September 29, 2026) | `allowedProviders`: managed setting limiting which API providers a machine may use — Anthropic API, custom endpoint, Bedrock, Mantle, Vertex AI, Foundry, Claude apps gateway | `CLAUDE_CODE_DISABLE_WEB_FETCH`: turns off the WebFetch tool | `claude --desktop`: opens the desktop app on the current directory, or with `--continue` / `--resume
Best for: Regulated environments where which endpoint a developer's machine can reach is an audit question rather than a preference.
Plans and pricing
Sonnet 5.5 enters at the same price as Sonnet 5 — $2/$10 per MTok — so the lineup now reads Fable 5.1 at $10/$50, Opus 5.5 at $4/$20, Sonnet 5.5 at $2/$10 and Haiku 4.5 at $1/$5. In Claude Code 2.1.284, Sonnet 5.5 became the default Sonnet model on the Anthropic API.
Sonnet 5.5: $2/$10 per MTok, cache read $0.20, 5m cache write $2.50, 1h cache write $4 | Opus 5.5: $4/$20 per MTok, cache read $0.20 | Fable 5.1: $10/$50 per MTok | Haiku 4.5: $1/$5 per MTok | Batch API: 50% off input and output on all models | Default Sonnet on the Anthropic API: Sonnet 5.5 from Claude Code 2.1.284 (September 28) | Retirement announced this period: Claude Sonnet 4.5 on November 30, 2026
Best for: No price rise to absorb. The work this week is migration, not budgeting.
ChatGPT / OpenAI
★ DevDay 2026 — more than twenty announcements in one day
OpenAI's largest DevDay to date, framed around opening ChatGPT as a surface that developers can build inside rather than only call from. The company put its weekly user base at 1.2 billion.
Rather than list everything, here is what changes a working team's options this quarter. The rest — Pages, collaborative slides, shareable profiles, the Meetings plugin, Bedrock Managed Agents — is in OpenAI's own recap.
Also announced: Private Intelligence (Zero Data Retention with Private Safety Processing; Private Inference preview coming this fall) | ChatGPT Space and Pages for Pro, Business and Enterprise | Collaborative slides in the coming weeks | Teams and shared tasks, and @ChatGPT in Slack and Microsoft Teams, for Business and Enterprise | Meetings plugin in beta on macOS for Pro and Business, audio deleted once notes are ready | Plugin extensions, Plugin Creator, MCP events for plugin automations | Sites can host plugins on Business, Enterprise, Healthcare and Edu | OpenAI Marketplace: enterprise commitment applied to 32 partner products | Bedrock Managed Agents, built with Amazon
★ GPT-6.1 Sol — near-Astra intelligence at a fifth of the price
OpenAI describes Sol 6.1 as delivering near-Astra intelligence at a fifth of Astra's standard input and output token prices, with the gains concentrated in agentic coding, computer use and professional work.
The pricing is the argument. $2 per million input and $10 per million output, with cached input at $0.10 — a twentieth of the uncached rate, which is an unusually steep cache discount and makes repeated-context workloads markedly cheaper to run. Tool calling requires the Responses API; multi-agent support is in beta there too.
Model: `gpt-6.1-sol` | APIs: `v1/responses`, `v1/chat/completions` | Pricing per 1M tokens up to 272K input: $2 input, $0.10 cached input, $2.50 cache write, $10 output | Tool calling: requires the Responses API | Multi-agent: beta, via the Responses API | Availability: all API, Plus, Pro, Business, Enterprise and Edu users | Also in Copilot from September 29
Best for: Agentic coding workloads currently on Astra. Re-run the cost model — a fifth of the token price with a twentieth-rate cache read changes which jobs are worth running.
★ The Agents API gains computer use, and OpenAI runs the infrastructure
Agents built on the Agents API can now operate software through its interface, in an OpenAI-hosted browser. The same release brings Codex's multi-agent capability, tool search, tool calling and context compaction into the API.
The operational point is who runs the browser. OpenAI hosts the execution environment, so the team using it does not stand up and secure a fleet of browsers — but website access approvals and sign-in are handled by the calling application, which is where the authorisation design work moves to. That is a sensible split, and it is also the part that needs a policy before anything touches a customer account.
Capability: agents complete tasks in an OpenAI-hosted browser | Application's responsibility: website access approvals and sign-in | Also included: Codex multi-agent, tool search, tool calling, context compaction | Availability: through the API, and in Codex and ChatGPT Work on Pro 500 and Enterprise plans | Related: Bedrock Managed Agents brings the same core capabilities natively into AWS
Best for: Workflows locked inside GUI-only software with no API. Write the approval policy before the first agent signs into anything.
★ Ultrafast, and a Pro tier at 25× the Plus allowance
Ultrafast is a premium speed tier: OpenAI quotes up to 8× faster token generation in Codex, at 300 tokens per second, and up to 6× in the API. GPT-6 Astra Ultrafast is available now; GPT-6.1 Sol Ultrafast is listed as coming soon. The new Pro 500 plan carries 25 times the ChatGPT Plus usage allowance and includes Ultrafast access.
Parameter: `service_tier: "ultrafast"` on `v1/responses` and `v1/chat/completions` with `gpt-6-astra` | Speed: up to 8× token generation in Codex (300 tokens/sec), up to 6× in the API | Availability: API customers with rate limits; US data residency only | In ChatGPT Work and Codex: Pro 500 and Enterprise plans | Pro 500: 25× the Plus usage allowance, includes Ultrafast | GPT-6.1 Sol Ultrafast: announced as coming soon
Best for: Interactive products where the wait is the product problem. Note the US data residency limit before planning a European rollout.
★ Codex moves to the cloud, and gains a security product
Codex now runs on a computer, from a phone, or in the cloud from any device, with reusable development environments that carry approved settings and permissions across a team. Alongside it, Codex Security Cloud scans GitHub repositories on demand or on a schedule, investigates findings, removes duplicates and prepares fixes — including while the laptop is closed.
Deduplication is the detail that decides whether a scanning product gets used. Any scanner can produce findings; the reason security backlogs rot is that the same issue arrives twenty times under different names. A tool that investigates and collapses duplicates before a human sees them is solving the actual bottleneck. Codex Security Cloud also includes access to models offered through Daybreak Blue without a separate application.
Codex in the cloud: available on Plus, Pro, Business, Healthcare, Education and Enterprise; reusable development environments with shared approved settings and permissions | Codex CLI refresh: voice control, new `/agents` view, worktrees, prompt editing, session resume — all plans | Code review in the ChatGPT desktop app: summaries, diffs, GitHub pull requests and GitLab merge requests, automatic cloud reviews — all plans | Codex Security Cloud: scan entire GitHub repositories on demand or scheduled, ongoing checks of new commits, investigates and deduplicates findings, prepares fixes; includes Daybreak Blue models without a separate application; Pro, Business, Enterprise and Edu on desktop and web
Best for: Security teams drowning in duplicate findings, and any team that wants a code review pass to run while nobody is at a desk.
Plans and pricing
GPT-6.1 Sol enters at $2/$10 per million with cached input at $0.10. Pro 500 is the new top consumer tier at 25× the Plus allowance with Ultrafast included. The OpenAI Marketplace lets eligible enterprise customers spend part of an existing OpenAI commitment on approved partner software.
GPT-6.1 Sol: $2 input / $0.10 cached input / $2.50 cache write / $10 output per 1M tokens, up to 272K input | Pro 500: 25× the ChatGPT Plus allowance, includes Ultrafast | Sign in with ChatGPT: plan allowance usable across 16 partners including Cognition's Devin, Notion, Vercel, T3, OpenClaw and Dactyl, with per-partner usage controls; identity available globally, plan usage for Plus and Pro in participating tools | OpenAI Marketplace: part of an existing enterprise commitment applied to 32 launch partners including Figma, Adobe, Salesforce, ServiceNow, Harvey, Palo Alto Networks and CrowdStrike | Reported platform gains: 45% lower API time to first token, over 30% faster tool calls and workflows
Best for: Enterprises with an unspent OpenAI commitment — the Marketplace turns it into procurement flexibility rather than a use-it-or-lose-it number.
Gemini / Google
★ Gemini 4 Argon — a million tokens of output, not input
Google's new frontier model moves the number everyone has been watching from the input side to the output side: Argon's output limit is one million tokens, up from 64K. Google positions it for real-world software engineering, enterprise knowledge work such as legal and finance, and cybersecurity vulnerability detection and patching.
A million-token output ceiling is a different kind of capability from a million-token context. Long context changes how much you can show a model; long output changes what it can produce in a single pass — a complete codebase migration, an entire document set, a full remediation patch series rather than a plan for one. Access is deliberately narrow for now: trusted cyber defenders through the Fairwind Program, and Google notes it is taking part in the US government's voluntary pre-release model access process.
Output limit: 1M tokens, up from 64K | Introductory pricing: $2 per million input, $10 per million output | Standard pricing after the introductory period: $4 per million input, $20 per million output | Cached input: 95% off the standard input rate | Access: limited rollout to trusted cyber defenders via the Fairwind Program, broader availability planned | Focus areas: software engineering, enterprise knowledge work (legal, finance), cybersecurity vulnerability detection and patching | Governance: Google states it is engaged in the US government's voluntary process for pre-release model access
Best for: Watch rather than plan around it — general access is not open. If you do get in through Fairwind, note that the introductory rate doubles later.
★ Skills arrive in Gemini and Workspace, and will replace Gems
Skills — reusable prompts that steer Gemini for a specific task — are rolling out across the Gemini app and Workspace, and Google says they will eventually replace Gems.
"Eventually replace Gems" is the operative phrase for anyone who has invested in building them. No retirement date has been published, but a replacement that is announced as a replacement is a migration on a timer you cannot yet see. If your team has standardised on Gems, start by inventorying which ones are actually in use — that list is the migration scope, and it is usually much shorter than the list of Gems that exist.
Skills: reusable prompts that customise Gemini for specific tasks | Surfaces: Gemini app and Google Workspace | Editions: all Workspace users, rolling out over coming weeks | Relationship to Gems: Google states Skills will eventually replace Gems; no retirement date announced
Best for: Workspace teams with a library of Gems. Inventory now, migrate on your own schedule rather than on the one you will eventually be given.
★ Workspace — Vids moves to 3.8 Flash Lite, and a Gemini hub for students
Google Vids switched its text-to-speech engine to the Gemini 3.8 Flash Lite model for AI voiceovers and avatar narration, which is the first visible consumer use of the TTS models that reached general availability the week before. Vids also gained caption customisation — fonts, typography, positioning and animation — which had not been adjustable before.
Separately, a Gemini Student Hub launched for Workspace for Education, collecting study notebooks, flashcards and practice quizzes in one place. It is available to Education users of all ages outside the EEA, and to personal account users globally — a geographic split worth noting if you are evaluating it from inside Europe.
Google Vids TTS: now uses `gemini-3.8-flash-lite-tts` for AI voiceovers and avatar narration; all Workspace customers | Vids captions: customisable fonts, typography, positioning and animations; all Vids users | Gemini Student Hub: study notebooks, flashcards, practice quizzes; Workspace for Education outside the EEA, and personal accounts globally
Best for: Teams producing internal video. The voiceover quality step comes free with the engine change — but check the EEA exclusion before promising the Student Hub to a European institution.
Plans and pricing
No changes were published to the Gemini API price list in this window. Gemini 4 Argon carries introductory pricing that is explicitly temporary.
No published Gemini API price changes between September 25 and October 2, 2026 | Gemini 4 Argon introductory: $2 / $10 per MTok, cached input 95% off standard input | Gemini 4 Argon standard, after the introductory period: $4 / $20 per MTok | Workspace: no price changes announced for the Vids, Skills or Student Hub updates
Microsoft Copilot
★ Copilot can now drive desktop applications
Computer use arrived in GitHub Copilot. It can read accessible app content and visual context, click controls, enter and edit text, press keys, scroll, drag, and move between applications.
The target is explicit and unglamorous: legacy and GUI-only software with no API, no command line and no MCP integration. That is a large share of what mid-market operations actually run on, and it has been the hardest category to automate because there was nothing to call. Copilot asks for approval before controlling an application, and apps you have set to always allow can be reviewed or reset — so the standing permission list is a thing somebody should own.
Status: public preview | Surfaces: GitHub Copilot CLI and the GitHub Copilot app, on macOS and Windows | Actions: read accessible content and visual context, click controls, enter and edit text, press keys, scroll, drag, navigate across applications | Enable in the CLI: `/computer on`, with `/computer show` for status and `/computer off` to disable | Enable in the app: Settings → Computer Use → Enable Computer Use | Control: approval requested before controlling an app; always-allow list can be reviewed or reset
Best for: Processes trapped in software with no integration surface. Treat the always-allow list as a permission register, not a convenience setting.
★ Sonnet 5.5 and GPT-6.1 Sol reach Copilot within a day of launching
Claude Sonnet 5.5 appeared in Copilot on September 28, the day Anthropic released it. GPT-6.1 Sol followed on September 29, the day OpenAI announced it.
Claude Sonnet 5.5 in GitHub Copilot: September 28, 2026 | GPT-6.1 Sol in GitHub Copilot: September 29, 2026 | Also shipped: HydraFusion in VS Code and the Copilot app (September 30); GitHub Copilot in VS Code September 2026 releases (October 1); dynamic workflows in the Copilot CLI and app (October 1); Copilot code review gains API support and a new default effort level (October 2)
Best for: Teams who standardise on Copilot rather than on a single lab. Same-day availability is becoming the expectation rather than the exception.
⚠ Four more models deprecated, two weeks after the last list
A second deprecation list in two weeks, each entry with a named replacement.
Gemini 3.5 Flash → Gemini 3.8 Flash. Gemini 3.6 Flash → Gemini 3.8 Flash. Kimi K2.7 Code → Kimi K3. Claude Opus 4.7 → Claude Opus 5.5. This comes on top of the six models announced on September 18 for retirement on October 19, so an organisation that pinned model names in workflow files now has ten migrations to do rather than six. Enterprise administrators may also need to enable the replacement models in their model policies before anything can switch.
Announced October 2, 2026: Gemini 3.5 Flash → Gemini 3.8 Flash; Gemini 3.6 Flash → Gemini 3.8 Flash; Kimi K2.7 Code → Kimi K3; Claude Opus 4.7 → Claude Opus 5.5 | Scope: all Copilot experiences | Administrator action: replacement models may need enabling through model policies | Previously announced, retiring October 19: Gemini 3.7 Flash, GPT-5.5, GPT-5.4, GPT-5.4 mini, GPT-5 mini, Grok 4.5
Best for: Whoever owns the workflow files. Do both lists in one pass rather than twice in three weeks.
Plans and pricing
No price changes were announced. The commercial news is availability and attrition: two new frontier models added within a day of launch, four older ones deprecated.
No Copilot price changes announced between September 25 and October 2, 2026 | Added: Claude Sonnet 5.5 (September 28), GPT-6.1 Sol (September 29) | Deprecated October 2: Gemini 3.5 Flash, Gemini 3.6 Flash, Kimi K2.7 Code, Claude Opus 4.7 | Retiring October 19, announced earlier: six further models | New capability in public preview: computer use against desktop applications
