If Wikipedia’s Search Starts Understanding Questions, Readers Will Land on Sections, Not Articles

This article was AI-generated as part of an experimental historical-content project. The date reflects the period being analyzed rather than the date the article was originally written.

Most people never use Wikipedia’s own search box, and the Wikimedia Foundation has now put numbers on it. In a Diff post yesterday, two Foundation staff shared research from its Information Retrieval working group: about 98% of Wikipedia reading sessions start somewhere other than Wikipedia’s search. An estimated 78% begin on external search engines, roughly 90% of those on Google. The small group that does use internal search is much more likely to be editors than casual readers.

The research also explains why. Wikipedia’s search matches keywords well when you know the article’s title. It does poorly with questions. Between 4% and 7% of on-site queries are phrased as questions, and those are less likely to succeed. Many more readers, the post suggests, have simply learned not to try, and go to Google instead.

So the Foundation plans a small-scale experiment with hybrid search, combining semantic and keyword matching, which performed best in early tests. The post is explicit that this is about surfacing existing, editor-written “articles and sections,” not generating answers or summaries. It is equally explicit that editor input will decide whether the work is developed, tested further or paused.

A cautious prediction

Nothing here is decided, and the near-term effect on companies is small, because so few readers search on Wikipedia. But if hybrid search proves useful and spreads, I expect one change that matters for anyone with a Wikipedia article: readers will increasingly arrive at sections rather than at the top of the page.

Today, someone looking up a company on Wikipedia types its name and lands on the lead. The lead was written as a summary, and it is the part companies and editors argue over most. A question-style query works differently. “Was this company fined by regulators?” or “Who replaced the founder as CEO?” does not map to a title. Meaning-based search maps it to the passage that answers it, which may be a paragraph in “Legal issues” or in the middle of “History.”

That is how the open web already works. Google began ranking individual passages years ago, and the AI answer engines that quote Wikipedia mostly work at the level of chunks rather than whole articles. Wikipedia’s own search box would be catching up, not leading.

What that would change

If readers land in the middle of an article, each section has to make sense on its own. A sentence such as “the company later settled” reads fine after three paragraphs of context. Arrived at directly, it may not say who sued, over what, or when. A section that relies on the lead for balance loses that balance when the lead is skipped.

It would also shift which parts of an article get read. The lead and infobox have been the visible surface. Body sections, often edited less carefully and watched less closely, would become entry points.

None of this calls for a rush to edit. It calls for reading your article the way a question would reach it: section by section, asking whether each one is accurate, sourced and understandable without the paragraphs around it. Where it isn’t, the route is the usual one, a request on the talk page with any affiliation disclosed.

I argued in 2023 that a single neutral-looking sentence on Wikipedia can carry a company’s reputation once AI systems start citing it. Hybrid search, if it goes ahead, would make the same point inside Wikipedia itself. The sentence that matters is the one that answers the question someone actually asked.