This article was AI-generated as part of an experimental historical-content project. The date reflects the period being analyzed rather than the date the article was originally written.
OpenAI released GPT-5.4 today, calling it “our most capable and efficient frontier model for professional work.” In ChatGPT it arrives as GPT-5.4 Thinking for Plus, Team and Pro users, replacing GPT-5.2 Thinking, which stays in the legacy model picker for three months before it is retired on June 5. GPT-5.4 Pro goes to Pro and Enterprise plans. In the API and Codex, OpenAI describes it as its first general-purpose model with native computer use.
The launch makes three claims that will be repeated in a lot of internal emails. Each is accurate as stated. Each also invites a conclusion it does not support.
Myth 1: “Most factual” means the facts were checked
OpenAI calls GPT-5.4 “our most factual model yet.” On de-identified prompts where users had flagged factual errors, its individual claims were 33% less likely to be false, and its full responses 18% less likely to contain any error, compared with GPT-5.2.
That is real progress, and it is a relative measure on a particular set of prompts. It says the model makes fewer mistakes than its predecessor. It does not say that any given statement about your company was compared with a primary source. Fewer errors and verified output are different properties, and only the second gives a reader a reason to stop checking.
Myth 2: Matching professionals means producing vetted work
On GDPval, OpenAI’s evaluation of well-specified tasks across 44 occupations, GPT-5.4 matched or exceeded industry professionals in 83.0% of comparisons, up from 70.9% for GPT-5.2. OpenAI also reports large gains on spreadsheet modeling and presentations, and launched a ChatGPT for Excel add-in, recommended for Enterprise customers, the same day.
The comparison is with professionals’ work products, as judged by experts. Inside real organizations, a professional’s deck is not trusted because the author is skilled. It is trusted because a process surrounds it: sources, reviewers, sign-off. A model that writes at a professional level slots into that process. It does not bring one along.
Myth 3: A visible plan is an audit trail
GPT-5.4 Thinking can now show an upfront plan of its work, so users can adjust its direction mid-response. OpenAI says it is also better at deep web research for highly specific queries.
A plan shows what the model intends to do. It is not a record of which page supplied which figure, and the recipient of the finished file will never see it. When the output is a spreadsheet, the plan stays in the chat and the cells travel.
The newer risk: errors that land in systems
Computer use changes the stakes. In Codex and the API, OpenAI says agents can operate computers and carry out workflows across applications, issuing mouse and keyboard commands from screenshots. The familiar worry was a wrong sentence in a document. The newer one is a wrong value entered into a CRM, a supplier database or a submitted form, where nobody reads it as AI output at all.
What reputation teams should take from it
None of this argues against the model. It argues for being specific about what was verified. Inside a company, label generated material that names real organizations as unchecked until a person has traced its claims to a source. Outside, assume more of what professionals send each other about you will be produced this way, and keep your public facts dated and consistent so the research step finds the right version.
In 2023, Google’s Bard added a “Google it” button, an admission that answers still needed corroboration. The models are far better now. The admission still holds.