This article was AI-generated as part of an experimental historical-content project. The date reflects the period being analyzed rather than the date the article was originally written.
Google released Gemini 3.1 Pro yesterday as a preview. It is rolling out to developers through the Gemini API, to enterprises in Vertex AI and Gemini Enterprise, and to consumers in the Gemini app, with higher limits for Google AI Pro and Ultra subscribers, and in NotebookLM for those subscribers only. The headline number is ARC-AGI-2, where Google reports a verified score of 77.1%, “more than double the reasoning performance of 3 Pro.”
It is a striking result, and it is worth being precise about what it measures. Google describes ARC-AGI-2 as evaluating “a model’s ability to solve entirely new logic patterns.” The tasks are abstract puzzles, built to test reasoning on problems the model has not seen rather than accumulated knowledge.
That makes the benchmark useful for what it tests and nearly silent on the question reputation teams care about: when this model describes a company, where does the description come from?
The example Google chose
The more revealing material is in the examples. Under “complex system synthesis,” Google says the model built a live aerospace dashboard, “successfully configuring a public telemetry stream to visualize the International Space Station’s orbit.” Elsewhere it offers help to “synthesize data into a single view.”
Look at the dashboard from the point of view of provenance. The impressive part is not that the model knew where the space station is. It is that it found a public, structured data feed and wired the output to it. Every number on that screen has a path back to a source somebody maintains.
Now point the same capability at a market. “Build me a single view comparing the five largest providers in this category” is exactly the kind of task Google is describing. For the space station there is a telemetry stream. For most companies there is a scatter of press releases, an About page, a Wikipedia article, data aggregator profiles that may be years old and whatever news coverage ranks. The reasoning can be excellent and the single view can still rest on the weakest of those.
Capability is not a citation path
Better reasoning probably helps at the margin. A stronger model may notice that two sources disagree, or that a figure is dated. But reasoning operates on what retrieval supplied. A doubled logic score does not tell you which sources a dashboard used, whether it shows them, or how anyone would correct a wrong cell.
Distribution sharpens the question. Gemini Enterprise and Vertex AI put the model inside organizations, where analyses are prepared for procurement, investment committees and partnership reviews. NotebookLM works from sources the user uploads. Neither produces a public answer you can search for and screenshot.
Google is also explicit that this is a preview, released “to validate these updates” before general availability. Limits and access may change. The structural point will not.
What a company can do
Treat the telemetry feed as the standard. The companies described best in synthesized views will be the ones that publish their own facts in structured, stable, attributable forms: a facts page with dates, investor data machines can read, figures that match across the company site, filings and the major databases, and a visible date and source on anything a summary might lift. In 2023 I wrote that provenance was becoming the product in generative search. It is becoming the product in dashboards too.
For a reasoning model, the reputation question is not how smart it is. It is what it was looking at when it got to you.