This article was AI-generated as part of an experimental historical-content project. The date reflects the period being analyzed rather than the date the article was originally written.
Anthropic released Claude Opus 4.5 on Monday with a familiar set of claims: better coding, better agents, better computer use. The part that matters most for companies sits lower in the announcement. Anthropic says the model is “meaningfully better” at “working with slides and spreadsheets.” Claude for Chrome is now available to all Max users, and the Claude for Excel beta announced in October has expanded to all Max, Team and Enterprise users. One customer quoted in the post says accuracy on its internal Excel automation and financial modeling evaluations improved 20 percent.
The easy conclusion is that an assistant which can browse in Chrome and build in Excel will produce more reliable numbers about companies. I think that gets the risk backwards.
Where the numbers come from
Picture a common request. An analyst, a reporter or a procurement manager asks an agent to pull revenue, headcount and leadership for five companies into a sheet. The agent opens tabs, reads investor pages, data aggregators, press coverage and perhaps a stale PDF, then fills the cells.
The output is a clean table. Every cell has a value. Nothing in the cell says where the value came from, which year it describes, or whether it belongs to the parent company, a subsidiary or a similarly named firm.
A chat answer at least has a shape that invites scrutiny, and often citations. A spreadsheet strips context by design and draws its authority from the format. Once a figure lands in a cell, it gets summed, charted, pasted into a deck and forwarded. By the third hop, nobody remembers it came from an aggregator page last updated in spring.
Why better models do not fix this
Better reasoning reduces some errors. It does not resolve ambiguity in the sources. If your investor page shows trailing figures in one place and guidance in another, or if an aggregator still lists a CEO who left months ago, a more capable agent will find those faster and format them better.
Anthropic also says Opus 4.5 is harder to trick with prompt injection than any other frontier model, citing a benchmark run by an outside firm. That matters for browser agents reading third-party pages, and it is good news. But prompt injection is a security problem. Stale and conflicting facts are a publishing problem, and no model setting solves those.
What finance and IR teams can do
Publish key figures in text, with the period and the entity named in the same sentence: “Revenue for fiscal 2025 was,” not a number floating in a chart. Keep one canonical page for leadership and corporate structure, and date it. Retire or clearly label old PDFs that still rank. And when you test AI tools, do not only ask them questions. Ask them to build the table an analyst would build, then check every cell against your filings.
In 2023 Bard added a button to double-check its answers with Google, a quiet admission that answers still need corroboration. Agents that fill spreadsheets need the same thing, and corroboration only works if the sources are clean enough to check against.