GPT-4o’s Demo Is Smooth. The Reputation Question Is What It Says When You Are Not Watching

This article was AI-generated as part of an experimental historical-content project. The date reflects the period being analyzed rather than the date the article was originally written.

A product demo is the most controlled environment a technology will ever work in. The prompts are rehearsed, the presenters know where the edges are, and if something goes sideways there is a laugh and a quick retry. Yesterday’s OpenAI livestream was a good demo. That is exactly why it is worth thinking about what happens after it.

On May 13, OpenAI introduced GPT-4o, with the “o” standing for “omni.” The model accepts any mix of text, audio, image and video, and OpenAI says it can respond to spoken input in as little as 232 milliseconds, with an average of 320, which is close to the pace of human conversation. On stage, presenters interrupted it mid-sentence, asked it to tell a story more dramatically, had it walk through a handwritten equation seen through a phone camera, and used it to translate between Italian and English in real time.

The more consequential detail was quieter. GPT-4o is coming to ChatGPT’s free tier, with usage limits, and OpenAI says free users will also get tools that used to sit behind the paywall, including web browsing, file uploads and GPTs. The new voice experience is due for Plus subscribers in alpha over the coming weeks.

A bigger audience, with nobody supervising

Put those pieces together and the shape is clear. A more capable model is about to answer far more people, about far more things, in a format that feels like talking to someone.

Some of those things will be companies and the people who run them. Someone checks out a supplier before a meeting. An analyst asks what a CEO said about layoffs last year. A job candidate asks whether a firm treats people well. None of these conversations will have a communications team in the room, and none of them will be rehearsed.

That was always true of search. What changes is the form of the answer.

Voice removes the page

A results page, for all its problems, is a list. The reader sees several sources, some critical, some owned, and makes a judgment. Even a text answer from a chatbot sits on a screen where you can reread it and check it.

A spoken answer is one answer. It arrives in a warm, quick voice and then it is gone. There is no visible ranking to second-guess and nothing on screen that signals “this sentence came from a forum post.” Fluency sounds like confidence, and confidence sounds like accuracy, even when the two have little to do with each other.

OpenAI also says GPT-4o is noticeably better in non-English languages. That is genuinely useful. It also means a company may be described, out loud, in languages its own team does not monitor.

Where the work shifts

I don’t think alarm is the right reaction. A faster, cheaper model is not inherently riskier than a slower one, and these systems are improving. But the question changes.

It used to be “what ranks for our name?” Now it is also “what does a general-purpose assistant say about us when someone asks casually?” When Gemini was announced in December, the point was that the real shift is who summarizes you first. And if last week’s reports that OpenAI is building a search product are right, those answers will lean more and more on what the assistant retrieves from the open web.

Which brings the work back to unglamorous inputs: accurate owned pages, credible third-party coverage, reference sources that are current. The demo was polished. The answers people get about your company next month will be assembled from whatever material is out there.