If Wikipedia Can’t Agree How to Use LLMs, Who Audits the Sentences AI Will Repeat?

This article was AI-generated as part of an experimental historical-content project. The date reflects the period being analyzed rather than the date the article was originally written.

Wikipedia has a rule for almost everything: sourcing, neutrality, conflicts of interest, even how to name an article about a ship. So it may surprise people to learn that more than a year after ChatGPT launched, English Wikipedia still has no adopted policy or guideline specifically about large language models.

Not for lack of trying.

Two attempts, no rule

Through most of 2023, editors built a long draft at WP:LLM on how language models could and couldn’t be used. In late August it went to a formal request for comment. When that closed in October, the closer found “an overwhelming consensus to not promote,” with 30 editors against. The most common objection was that existing policies, especially verifiability and reliable sourcing, already cover the problems. The page became an essay.

On December 13, editor JPxG opened a second RfC with a much narrower question: should one sentence become policy? The sentence says LLM output used on Wikipedia “must be manually checked for accuracy (including references it generates), and its use (including which model) must be disclosed by the editor,” and that text added in violation “may be summarily removed.”

As of this week, that discussion is still open, and it is split three ways. Some editors support the sentence. Some want LLM-generated content banned outright. Others oppose it, arguing that existing rules are enough and that no one can reliably detect machine text, so a summary-removal clause could be turned against human writing. At least one editor reported that AI detectors flagged their own hand-written prose.

The interesting part is what they agree on

Read the thread closely and almost nobody argues that unchecked model output belongs in articles. The fight is about disclosure and enforcement, and specifically whether a rule means anything when violations can’t be detected. One supporter offered a useful comparison: Wikipedia already prohibits undisclosed paid editing, which is nearly as hard to detect, and stating the rule still does work.

That is a reputation professional’s argument as much as a Wikipedian’s. Most of what governs a company’s Wikipedia article was never enforced by detection. It was enforced by norms and a small number of people checking citations.

So who audits the sentences?

This is where the debate stops being internal. Wikipedia is a primary input for the systems that now summarize companies and people. The Times’ complaint against OpenAI, filed December 27, notes that in a filtered Common Crawl snapshot of the kind used to build training data, only two sources were more heavily represented than the Times itself. One of them was Wikipedia. Wikipedia is also being retrieved live, as the Wikimedia Foundation’s ChatGPT plugin experiment made explicit last summer.

If machine-written text enters Wikipedia unflagged, you get a loop: a model writes a plausible sentence, the encyclopedia absorbs it, and the next model repeats it with encyclopedic authority. One editor in the RfC raised exactly this circular-sourcing worry.

The honest answer to “who audits” is the same volunteers, under the same verifiability rules, at human speed. There are tracking categories for suspected AI-generated text and a WikiProject devoted to cleanup. They are real, and small next to what a language model can produce.

What this means for companies with Wikipedia articles

A bot writing your whole article is the unlikely scenario. The likelier one is a well-formatted sentence with a citation that doesn’t support it. One editor in the discussion shared drafts in which every reference was fabricated. A single sentence like that, in an otherwise solid article, can sit unnoticed for months while search engines and chatbots repeat it.

So read your article the way a fact-checker would. Open the citations and confirm they say what the sentence says. If something is wrong, raise it on the talk page under the conflict-of-interest rules rather than editing directly. And if you propose text, verify every word yourself. Whatever the RfC decides, unverified machine text is the worst thing a company could bring to its own article.

The vote will close eventually. The downstream problem won’t. Wikipedia’s answer to most content questions is verifiability, and that answer is only as strong as the number of people checking. AI systems will keep repeating whatever survives.