← Back to Insights
11 min read
By Ben Gould · Published 9 September 2026
Most businesses are comparing the wrong things. Claude and ChatGPT are frontier assistants you choose deliberately; Copilot and Gemini are productivity assistants your software suite chose for you. The first pair is better at hard work. The second pair is better at knowing where your files are. Most teams need one of each, and neither should own your processes.
As of July 2026 this is no longer a purchasing decision anyway. Both suites now bundle the assistant into the licence and charge you for it whether anyone uses it or not. The only question still open is what you buy on top, and for whom.
The AI stopped being an add-on and became part of the rent.
Microsoft raised O365 business plan prices from 1 July 2026 - roughly 17% on Business Basic and 12% on Business Standard - and folded Copilot Chat into Word, Excel, PowerPoint, Outlook and OneNote as part of the deal. Existing customers meet the new pricing at their first renewal after that date (SysGroup). The standalone Microsoft 365 Copilot licence went up at the same time.
That is precisely the move Google made eighteen months earlier, when it dissolved the separate Gemini add-on into every Workspace tier and raised the base price for everyone (9to5Google). Two vendors, same playbook, and the same result for you: a per-seat AI bill for your entire headcount, decided by your suite vendor, arriving on renewal.
So "should we buy Copilot or Gemini" is a dead question. You have one. The real question is whether it is any good at the work you actually need doing, and the evidence on that is not flattering.
Not as much as the licence count suggests, and the reason is trust rather than training.
Recon Analytics has been tracking Copilot's accuracy Net Promoter Score, and it has been negative for over a year: -3.5 in July 2025, -24.1 by September 2025, and still -19.8 in January 2026. Among users who had stopped using Copilot, 44.2% gave distrust of the answers as their main reason. Most telling for this comparison: when workers are given a free choice between Copilot, ChatGPT and Gemini, only 8% pick Copilot (The Next Web, April 2026).
A negative accuracy NPS is a specific and damning signal. It doesn't mean people dislike the product. It means the average person who has tried it is more likely to warn a colleague off trusting its answers than to recommend it. You cannot build a process on a tool your team has quietly decided not to believe.
Governance, and it is getting worse rather than better as agents arrive.
Gartner surveyed IT leaders at 186 organisations between March and May 2026 for its Microsoft 365 Copilot and Agents research. The findings are consistent: 51% named oversharing and data loss as the top barrier to a successful Copilot deployment, 80% said they need additional governance controls before deploying Copilot agents widely, 68% are worried about agent sprawl, and 86% are actively limiting Copilot Studio deployments. Nearly half now pay for third-party tooling just to govern Microsoft 365, up from 40% a year earlier (Gartner 2026 research, as reported).
This is the structural problem with an assistant wired into everything you can already see. It applies no judgement about whether your permissions were sensible; it just answers. Every overshared SharePoint site and never-expiring link becomes searchable in plain English on day one.
The security researchers keep making the same point in a more alarming way. Twice in 2026, Microsoft has had to patch flaws where an assistant could be talked into handing over things it had access to - once through text typed into a public web form that an agent later read, and once, in the consumer Copilot, through a single link a user clicked, which was enough to reach their email, calendar and stored notes (The Hacker News). Both were fixed, and neither was exploited in the wild.
The detail matters less than the shape of it. An assistant that can read everything you can read is one convincing message away from repeating it to the wrong person. That is not a bug that gets fixed once; it is the trade-off you accept when you connect a language model to your whole company, and it is why the permissions clean-up is the real project rather than the rollout.
Gemini's version of the same trade-off is commercial rather than architectural. Google moved the Gemini app to compute-based usage limits in May 2026 after complaints that people were hitting caps too quickly, then softened them again following the backlash (9to5Google). And at Cloud Next in April 2026 it renamed Vertex AI to the Gemini Enterprise Agent Platform and absorbed Agentspace into it. Neither is fatal. Both tell you that what you're buying is a moving target set by someone else's roadmap.
No, and the honest reading of the evidence cuts both ways.
The suite assistants do save real time on admin. The UK's Department for Work and Pensions measured 19 minutes per day across 1,716 users against a 2,535-person control group (The Register, February 2026). The earlier cross-government trial of roughly 20,000 civil servants found 26 minutes a day, with 82% saying they would not want to go back (GDS report). Microsoft's own 2026 Work Trend Index, analysing over 100,000 Copilot conversations, found 49% of them supporting genuinely cognitive work rather than formatting and admin (Microsoft WorkLab) - Microsoft's own data on Microsoft's own product, so weight it accordingly.
The catch has been visible since the government trials of 2024 and 2025 and nothing since has overturned it. The Department for Business and Trade's evaluation of 1,000 Copilot licences concluded flatly: "We did not find robust evidence to suggest that time savings are leading to improved productivity." In Excel, Copilot users were slower and less accurate than colleagues without it; PowerPoint output arrived faster but needed correcting (DBT evaluation, via The Register).
Put those together and the shape is stable across two years of evidence. The suite assistants reliably shave minutes off admin: summarising a meeting you missed, finding a document, drafting a reply, unpicking a forty-message thread. They don't reliably improve work that needs judgement, and on data-heavy tasks they can make it worse. Nineteen minutes a day across a team is real money. It is not a transformation, and it compounds into nothing unless you spend the time on something.
The hard middle of the work: sustained reasoning over a messy brief, analysis you would otherwise hand a competent analyst, code, structured extraction from ugly documents, and anything where you want to argue with the output for twenty minutes rather than accept a first draft.
The spending pattern backs this up. Menlo Ventures' enterprise research put Anthropic at 40% of enterprise LLM API spend against OpenAI's 27% and Google's 21%, with Anthropic holding an estimated 54% of the enterprise coding market by mid-2026, up from about 42% six months earlier (Menlo Ventures). When a business builds something that has to work, it overwhelmingly reaches past whatever its suite shipped.
The frontier assistants are also walking into the suite's own territory. Claude for Excel became generally available to Pro subscribers in January 2026, Claude runs as a PowerPoint add-in, and Anthropic made enterprise-managed authorisation for its MCP connectors generally available in August 2026, wiring Claude into Slack, Notion, Atlassian, Linear and the rest under IT control. The capability gap is closing from the frontier side faster than it is from the suite side.
Increasingly, yes - and that is the most useful fact in this whole comparison.
Microsoft began offering Anthropic's models inside Microsoft 365 Copilot in September 2025 (Microsoft 365 blog), and by August 2026 Claude was selectable across Researcher, Copilot Studio, Excel, PowerPoint, Word, Copilot Chat and Cowork. The company with the strongest commercial incentive on earth to standardise on one model decided that model choice was a feature worth shipping.
Its customers reached the same conclusion independently. In that same Gartner 2026 survey, 66% of organisations deploying Microsoft 365 Copilot are also deploying at least two other enterprise AI assistants. Two thirds of Copilot customers are running Copilot and something else, on purpose. The multi-vendor argument I've been making forever is no longer a contrarian position - it is the majority behaviour of the market, and it is the practical version of what vendor-agnostic AI actually means.
If you're still planning to standardise your business on one assistant in 2026, you're taking a position that Microsoft, Gartner's respondents and the enterprise spending data have all abandoned. That is the case I set out in full in the dangers of a single-platform AI assistant.
| Claude / ChatGPT | Copilot / Gemini | |
|---|---|---|
| Why you have it | You chose it | It came with the licence |
| Best at | Reasoning, analysis, code, long documents | Summarising, retrieval, first drafts in the app |
| Knows your data | Only what you connect or paste | Everything you already have permission to see |
| Model choice | Yours, per task | Whatever the vendor routes you to |
| Sits where you work | A separate window, plus add-ins | Inside Word, Excel, Gmail, Docs, Teams |
| Cost shape | Per seat, for the people who need it | Bundled into everyone's plan |
| Proven ROI | Task-level, easy to demonstrate | Minutes per day, hard to attribute |
| Main risk | Data has to leave the suite deliberately | Inherits every permission mistake in your tenant |
Neither column wins. They answer different questions, and the mistake is buying one and expecting it to do the other's job.
For most UK SMEs of 20 to 200 people, a small amount of both, split deliberately:
On data protection, don't let the consumer headlines decide this. OpenAI and Anthropic both commit to not training on business or API data by default on commercial tiers, with zero-retention options available - roughly parity with what Microsoft and Google offer inside the suite. Read the terms for the specific tier you're buying.
You should never standardise on one. Picking one to start with is fine and usually sensible.
The test is the one I keep coming back to: could you move a critical workflow to a different provider in a week, without a rebuild? If your AI lives entirely inside a suite assistant, the answer is no - not because moving is hard, but because there is nothing to move. The workflow only ever existed as a habit in someone's sidebar. If it lives in an automation layer, or behind an AI gateway, swapping providers is a configuration change.
That distinction is worth more each year. Frontier models leapfrog each other in months. Suite assistants get repriced, rebundled and renamed on a quarter's notice, as both Google and Microsoft have now demonstrated in consecutive years. Any architecture that assumes today's winner is permanent is going to age badly, and quickly.
Claude and ChatGPT are capability purchases. Copilot and Gemini are distribution - in front of you because you already pay for them whether you asked or not. The evidence says the bundled assistants save real minutes on admin and lose the room the moment a task needs judgement or a spreadsheet; the frontier assistants do the harder work but know nothing about your business until you connect them. Two thirds of Copilot customers have already concluded they need both. Buy both narrowly, for the people who will use them, and keep your actual processes in a layer that doesn't care which one you picked.
That layer is what I build. If you're looking at a renewal quote with Copilot baked into it and aren't sure what else you need, that is exactly the kind of thing a thirty-minute discovery call sorts out - I'll tell you which seats are worth buying, which processes should never be with an assistant at all, and what to automate instead. See AI consulting and process automation for how I approach it.
The dangers of vendor lock-in with single-platform AI assistants
Read more →n8n vs Zapier vs Make: which automation tool for UK SMEs (2026)
Read more →What does "vendor-agnostic AI" actually mean in practice?
Read more →Or see how I put this into practice: services, case studies.