This article was AI-generated as part of an experimental historical-content project. The date reflects the period being analyzed rather than the date the article was originally written.
OpenAI released o3 and o4-mini yesterday, and the most important sentence in the announcement is not about benchmarks. It is this one: “For the first time, our reasoning models can agentically use and combine every tool within ChatGPT,” including web search, Python analysis of files and data, reasoning about images, and image generation.
OpenAI’s example shows what that looks like. Asked how summer energy use in California will compare with last year, the model might search for utility data, write code to build a forecast, generate a chart and explain the factors behind it, “chaining together multiple tool calls,” typically in under a minute. Plus, Pro and Team subscribers have the new models now, and free users can try o4-mini by selecting “Think.”
The prediction
Here is where I think this leads, with the usual caveat that predictions about AI products age quickly.
Answers about companies will increasingly arrive as artifacts rather than replies. Ask a tool-using model whether a company is financially sound, how it compares with its competitors or what happened in a controversy, and the response may combine several searches, a table built from figures it found, a chart, perhaps a generated image, and a confident narrative that ties them together. It will look like an analyst’s afternoon.
And readers will treat the process as verification. A chart implies data. A table implies a calculation. Several searches imply cross-checking. People will read the shape of the work as evidence that the work was checked, the way a footnoted report reads as more reliable than a paragraph. I argued recently that a visible thought process is not the same as a checkable source. Multi-step artifacts carry that problem from text into charts and tables.
Why caution is warranted
OpenAI’s own system card supplies a reason. On PersonQA, an evaluation built from questions about publicly available facts about people, o3 answered correctly more often than o1 (59% against 47%) but also hallucinated more often (33% against 16%). OpenAI’s explanation is that o3 “tends to make more claims overall, leading to more accurate claims as well as more inaccurate/hallucinated claims,” and it says more research is needed to understand why.
That is a finding about people, which makes it directly relevant to executives. A model that says more gets more right and more wrong. When those claims are packaged in a chart or a table, the wrong ones borrow the credibility of the right ones.
Chains also compound. If the first search retrieves an outdated figure, the code built on it is precise and wrong, the chart displays the error cleanly, and the narrative explains it fluently. Each step adds polish to the original mistake. When Bard added a button to double-check its answers, the lesson was that answers still need corroboration. A chain of tool calls looks like corroboration without necessarily being it.
What I would prepare for
Expect your numbers to be recomputed. A tool-using model may build its own revenue growth or market share figures from whatever sources it finds. Publish the figures that matter in clear, current form, with dates and definitions, so the inputs are right.
Expect artifacts to travel. A chart from a chatbot will be screenshotted into decks and posts without the conversation that produced it. Treat a widely shared AI-made chart about your company as you would a misleading analyst note: something to answer with a correct source.
And ask these models the questions a skeptic would ask about you, then read the whole output, including the code and the sources it chose. The artifact shows what it concluded. The steps show why.
I do not know how quickly people will come to rely on these outputs. But format carries weight. A one-line wrong answer looks like a mistake. A wrong answer with a chart looks like research.