What Communications Teams Should Demand Before Their Reporting Becomes Someone Else’s AI Page

This article was AI-generated as part of an experimental historical-content project. The date reflects the period being analyzed rather than the date the article was originally written.

For most of this year, the argument between publishers and AI answer engines has been covered as a media business story. It is also a reputation story, and not only for newspapers. Any organization that produces original work it wants credited, whether that’s a newsroom, a research team, an annual industry report or a CEO’s byline, is in the same position. Its work can be read, compressed and retold by a system that decides how much credit to give.

The past six weeks made that concrete. In early June, the New York Times described publishers scrambling as Google’s AI Overviews began answering questions above the links that used to send them readers, a shift an earlier post looked at from the publisher’s side. Forbes then accused Perplexity of republishing its reporting in the Perplexity Pages feature, with sources cited “in the most easily ignored way possible,” as one Forbes editor put it. Wired reported on June 19 that it had seen Perplexity access pages its robots.txt file asked bots to avoid.

Perplexity disputes the framing. Its head of business told TechCrunch that fetching a page because a user asked about that URL “doesn’t meet the definition of crawling,” and that “nobody has a monopoly on facts.” Its CEO had said the company would cite sources more prominently. Then, last week, Cloudflare gave every customer, free tier included, a one-click toggle to block AI scrapers and crawlers.

Communications teams won’t settle the copyright questions. They can decide what to check and what to ask for. Here is what I would put on the list.

Know who is retelling you

Run your most important queries, about the organization, its executives and its best-known reports, through Google, Perplexity, ChatGPT and Copilot. Record what comes back, with dates. One screenshot proves little. A baseline lets you notice when a summary of your work starts to drift.

Ask for attribution a reader can actually see

A small logo linking somewhere is technically a citation. It is not the same as a sentence saying where the information came from. Named, in-line attribution matters for credit, but it matters more for accuracy, because a reader who can see the source can check it. Where an answer engine offers a publisher contact or feedback channel, ask for that standard explicitly.

Separate the controls, and know their limits

The controls on offer do different things. OpenAI’s GPTBot can be blocked in robots.txt. Google’s Google-Extended token, announced last September, lets sites opt out of helping improve Bard, now Gemini, and Vertex AI models. Google treats it as separate from Search, so it is not an exit from AI Overviews. A blanket block through a service like Cloudflare is blunt and may cut off discovery you want. And every robots.txt rule depends on the bot choosing to honor it.

So decide first which outcome you want: not being trained on, not being summarized, or being summarized with credit. They lead to different settings.

Read licensing deals as distribution terms

OpenAI has signed content agreements with a lengthening list of publishers, including News Corp, Vox Media, The Atlantic and Time since May. If your organization is negotiating anything similar, ask how errors get corrected, how fast, and who you call. Nieman Lab reported last month that ChatGPT was inventing broken links to partner publishers’ best-known investigations, so the question isn’t theoretical. As the Times lawsuit argued in December, misattribution is a reputational harm as well as a commercial one.

Build the escalation path now

When a synthesis misstates something about your organization, someone should own it. Capture the query, the answer and the date. Submit feedback through the product. Correct the underlying public source if that’s where the error started, and agree in advance when legal gets involved. This is the same discipline as the AI Overviews routine from May, extended to every engine that retells you.

None of these demands guarantees anything. Leverage is uneven, and most organizations are not the New York Times. But the minimum ask is reasonable: when a machine retells your work, readers should be able to tell whose work it was, and you should be able to tell when it got it wrong.