How AI engines read creator sites
Schema is not the AI citation lever it is sold as
Short answer
Not on the evidence we have. Pages cited by AI engines carry schema markup more often than other pages, but that is a correlation. When Ahrefs tracked pages that added JSON-LD and compared them with matched pages that did not, citations in Google AI Overviews, AI Mode and ChatGPT showed no major uplift. Schema describes a page; it does not make the page worth quoting.
Why is schema sold as an AI lever?
Schema markup is structured data added to a page, usually as a block of JSON-LD, that tells machines what the page is: an article, a product, a person, a podcast episode. It has been part of search engine work for years, mostly as a way to earn rich results such as review stars or event dates.
With answer engines, it picked up a new pitch. The claim goes like this: AI systems read structured data, so pages with schema get cited more, so adding schema is the fastest way into AI answers. The pitch is attractive because schema is cheap. A plugin can add it to every page of a site in an afternoon, and an audit can count it in seconds.
There is a real observation behind the pitch. Pages that AI engines cite do carry schema more often than pages they ignore. The question is what that observation proves.
What happens when a page adds schema?
The best test we have found comes from Ahrefs. Instead of comparing cited and uncited pages, which mixes schema with every other quality signal, they looked at pages that changed. In the Ahrefs schema study, the team tracked 1,885 web pages that added JSON-LD schema between August 2025 and March 2026, matched them against 4,000 control pages, and measured citation changes across Google AI Overviews, AI Mode, and ChatGPT.
The result was flat. On AI Mode and ChatGPT, pages that added schema did about as well as the matched pages that did not, within the range of random noise. On AI Overviews, pages that added schema did slightly worse than their controls, but both groups were already falling together before the change, and the authors do not read the gap as proof that schema hurts. Their overall verdict is that adding schema produced no major uplift in citations on any of the platforms they measured.
The study explains the earlier correlation too. Schema tends to live on sites that are better maintained and more technically careful. Those same sites tend to publish stronger content, earn more links and do the other things that get pages cited. Schema rides along with quality. It does not create it.
So is schema useless?
No, and it would be wrong to read the study that way. Schema still does what it was built to do: it describes a page in a form machines can parse without guessing. That helps search engines show the right title, date, author or episode details, and it costs very little to add.
We use it ourselves. Every page on this site carries JSON-LD for the organisation, and each note marks up its author and dates. We ship it because it is accurate and cheap. We do not sell it as a way to get cited, because the evidence does not support that claim. The same goes for the llms.txt file we publish, for the reasons in our note on what llms.txt measurably does.
The trouble starts when schema is treated as the main lever. A creator who pays for a schema audit while their site still has no written content has fixed the label on an empty box. The engine can now read precisely what the page is. There is still nothing on it worth quoting.
What does move citations, then?
The honest answer is that nobody has a complete map yet. The studies that exist cover particular engines, particular time windows and particular kinds of sites. But the evidence we have points away from markup and toward content.
The clearest signal for creators comes from format. In a controlled experiment by OtterlyAI, the same stories were published as videos and as written articles built from their transcripts, and the written versions earned roughly 4 in 5 format citations (79.5%). That is one publisher’s test, not a law. Still, it suggests that what an engine quotes is a clear, readable answer, and that the format of the content matters far more than the tags around it. Which engines lean toward video and which toward text is the subject of our note on which AI engines cite video.
For a creator, that turns the usual checklist upside down. The order that the evidence supports looks like this:
- First, have text. Pages that answer real questions, written from your own work, on a domain you own.
- Second, make each page quotable. One question per page or section, a direct answer near the top, a heading that says what the section covers.
- Third, let the engines in. Check that your robots.txt does not shut out the search crawlers you want to appear in. Our note on what blocking AI crawlers costs covers the difference between training and search bots.
- Last, add schema. It is worth doing, and worth doing correctly. It is the final step, not the first.
What does this mean for creator sites?
In The Dark Archive, our study of long-form YouTube channels, most creators have no written archive at all. Our count of channels that keep one found it was a very small minority. For those creators, schema is not the missing piece. The missing piece is the writing.
That is also why we are wary of tools that promise AI visibility through markup alone. They solve the easiest part of the problem, report a score that goes up, and leave the hard part untouched. A creator is better served by a smaller site with real, structured answers than by a perfectly marked-up page that says nothing.
The full method behind our own scoring of creator sites, including which technical signals we measured and how much weight each one carries, is in The Dark Archive report.
To see what your own channel’s site is missing, run the free Dark Matter Check.
Sources
- Ahrefs: we tracked pages adding schema, AI citations barely moved, retrieved 2026-09-30
- OtterlyAI: video-to-blog citation experiment, retrieved 2026-09-30