This article was AI-generated as part of an experimental historical-content project. The date reflects the period being analyzed rather than the date the article was originally written.
Anthropic released Claude 3.7 Sonnet on Monday, and the feature most people will notice first is not a benchmark. It is a window. In extended thinking mode, Claude works through a question step by step before answering, and the user can read that work. Anthropic calls it the first “hybrid reasoning model”: one model that can answer quickly or think at length, with developers able to set a thinking budget through the API. Extended thinking is available on paid plans, not the free tier.
OpenAI’s o1 took a different route last fall, showing a model-generated summary of its reasoning rather than the raw chain. Anthropic has decided to show the raw version. Its research post lists the reasons, and the first is trust: watching Claude think “makes it easier to understand and check its answers.”
I think that is half right, and the half that is wrong matters for any company that people ask these systems about.
A readable process is not a checkable source
Checking an answer means tracing a claim back to something outside the model. A thought process does not do that. It shows the model talking to itself. For a math problem, that is useful, because the logic can be tested against the problem. For a question like “Did this company settle that lawsuit?” the reasoning can only be as good as the facts the model already holds, and nothing in the launch post is about retrieving or citing sources.
Anthropic is admirably direct about the limits. The same research post says “we don’t know for certain that what’s in the thought process truly represents what’s going on in the model’s mind,” and that its results so far suggest “models very often make decisions based on factors that they don’t explicitly discuss.” The visible thought process is labeled a research preview.
So the window may not show the room. That is a candid admission from a vendor, and it should shape how the rest of us read the feature.
The scratchpad is now on screen
There is a second, more practical issue. Anthropic says it did not apply its usual character training to the thinking, and that Claude “sometimes finds itself thinking some incorrect, misleading, or half-baked thoughts along the way.”
For a reputation team, that is a new kind of artifact. Until now, a guess an AI system considered and discarded stayed invisible. Now a user can watch the model entertain the idea that a CEO left under a cloud, or that a product was recalled, before it settles on a more careful answer. The final response might be fine. The screenshot of the middle step is the part that travels.
I do not want to overstate this. Most people will skim the thinking or ignore it. But reporters, researchers and critics are exactly the users likely to read it closely, and they are also the people most likely to quote it.
What follows for communications teams
None of this is an argument against transparency. A visible process is better for researchers and probably better for users who know how to read it. The point is narrower. A longer, more visible chain of reasoning is a credibility cue, and it can make a wrong premise about a company look more considered than it is.
That suggests a few habits. Test the questions that matter about your company in extended thinking mode and read the thinking, not just the answer. Note where a discarded speculation would look bad out of context. And keep investing in the sources an answer can actually be checked against. Google conceded the same thing in 2023 when it gave Bard a double-check button: answers still need corroboration.
Reasoning you can watch is a real step forward. Sourcing you can check is a separate step, and nobody should let the first stand in for the second.