# Does an llms.txt file help a creator's site get cited?

By Tai Nguyen, ExDarkMatter. Published 2026-10-07; updated 2026-10-07.
Canonical: https://exdarkmatter.com/notes/llms-txt-measured/

**Short answer:** Not on the evidence available today. While an llms.txt file is often promoted as a shortcut to get indexed or cited by artificial intelligence, empirical data shows that answer engines do not seek out these files to discover content. Furthermore, most deployed files receive virtually no traffic or turn out to be unauthored server fallbacks. Having readable, well-structured text on your own site remains the primary requirement for earning citations.

## Why is llms.txt pitched as an AI shortcut?

An llms.txt file is a markdown file placed at the root of a domain that summarizes what the site offers and points toward its most important content. Defined in a community proposal (see our [glossary](https://exdarkmatter.com/glossary/#llms-txt)), it was created to give automated systems a clean map of a website without forcing them to parse complex navigation trees.

The idea has since acquired a commercial pitch. The promise is that publishing an llms.txt file puts a website on the radar of [answer engines](https://exdarkmatter.com/glossary/#answer-engine) and large language models. Under this view, a simple text file is a signal that attracts citations, so that answer engines notice your work and cite your pages.

The pitch appeals to creators for the exact reason structured data audits appeal to them: it sounds fast, cheap, and technical. Adding a file takes minutes, requires no redesign, and creates the comforting sense that an automated gap has been closed. Much like the claims we examined in our note on [schema is not the AI citation lever](https://exdarkmatter.com/notes/schema-is-not-the-ai-citation-lever/) it is sold as, this promise mistakes an optional index file for a discovery engine.

## Do AI engines actually read or cite llms.txt files?

When researchers look at how automated systems behave in practice, the reality looks very different from the promotional pitch. In [the Ahrefs study of llms.txt adoption](https://ahrefs.com/blog/llmstxt-study/), 28% of the 137K domains using Ahrefs Web Analytics publish an llms.txt file. Plenty of site owners have tried the format.

However, adoption has not translated into crawler activity or citation volume. In that same [Ahrefs survey](https://ahrefs.com/blog/llmstxt-study/), 97% of those files received zero traffic in May 2026. Publishing the file and having it used are different things.

The explanation is straightforward. As [the Ahrefs investigation](https://ahrefs.com/blog/llmstxt-study/) points out, answer tools do not go hunting for an llms.txt file, so publishing one does not put a site on their radar. In our reading, assistants and search crawlers find content by following existing links, querying search indices, or evaluating documents retrieved during live search retrieval. They do not scour the web hunting for root text files on domains they have no prior reason to inspect.

Publishing the file does not create discovery where discovery did not already exist. If an engine has no reason to examine your domain in the first place, adding an index file changes nothing.

## What did our audit of creator websites find?

In [The Dark Archive](https://exdarkmatter.com/glossary/#dark-archive), our research into long-form YouTube creators, we scored 528 creator websites to measure how accessible and readable their material was to modern machines. As part of that investigation, our crawler flagged every domain that returned content at the standard file location.

When we re-fetched the files our crawler had flagged on a subset of creator domains, the results were telling. Most of the addresses flagged were not genuine, creator-authored files at all.

Our inspection revealed three distinct patterns across those responses:

- **Missing-page fallbacks.** Many server configurations serve a generic error template or a homepage redirect instead of a proper not-found status when a nonexistent file is requested.
- **Automated platform outputs.** The file was generated by the site's platform or a plugin, not written by the creator.
- **Authored summaries.** Only a small handful of the examined files showed signs of being deliberately drafted by the creator to summarize their own original work.

The finding underscores a common pitfall in web audits. Merely observing a server response at a specific path tells you nothing about whether the site owner made an intentional choice. In the subset we re-checked, a hand-written file was the rare exception; the rest were by-products of hosting software or fallback pages.

The full llms.txt and crawler audit, including our complete methodology and technical breakdown, is detailed in [The Dark Archive report](https://exdarkmatter.com/report/).

## Why do we publish an llms.txt file if it does not drive citations?

Given that the empirical evidence shows no citation lift, one might wonder why ExDarkMatter maintains an llms.txt file on this domain.

We publish one for the same reason we implement clean semantic markup and descriptive metadata: it is accurate, orderly, and inexpensive to provide. An llms.txt file offers an unambiguous summary of what our archive contains for any software agent explicitly designed to look for it. It costs very little to write and keep current, and it represents sound technical housekeeping.

What we do not do is pretend that the file acts as a shortcut to search citations. We do not tell creators that adding a text file will rescue a channel from obscurity. An index file can organize existing knowledge, but it cannot invent authority or replace written substance.

Creators frequently face competing advice about technical optimizations. Some prioritize crawler policies, as discussed in our note on [what blocking AI crawlers costs](https://exdarkmatter.com/notes/what-blocking-ai-crawlers-costs/) a creator's domain. Others spend days fine-tuning markup or tweaking server headers. In our reading, technical hygiene is useful only after a creator has solved the much harder problem: having substantial, original written answers on a domain they control.

## What should a creator build before worrying about llms.txt?

If an llms.txt file does not drive citations, where should a video creator focus their effort? The answer comes back to the fundamental mechanics of how answer engines operate.

Answer engines do cite YouTube directly. As we explore in our analysis of [which AI engines cite video](https://exdarkmatter.com/notes/which-ai-engines-cite-video/), several major assistants frequently link directly to video URLs when responding to user questions. However, citation behavior varies dramatically depending on the platform and query. While certain platforms incorporate video links, others lean toward written pages.

More importantly, a single, nuanced point buried midway through an hour-long video is harder to cite than the same point stated in writing. When an assistant quotes a source, it looks for clear statements that directly resolve a specific query. A written article with a descriptive heading and an explicit answer near the top provides the exact shape that automated systems extract and credit.

For creators seeking broader visibility across both human readers and search assistants, the hierarchy of priorities is clear:

First, build written articles derived from your video material. Every comprehensive video essay, documentary, or interview contains valuable knowledge that remains trapped in audiovisual form. Turning those insights into readable articles on an owned website creates the raw material that machines can quote.

Second, structure each article around a discrete question. State the answer clearly in the opening paragraph, provide supporting context beneath it, and link back to the exact timestamp in the original video.

Third, ensure your server does not block search crawlers. Allow standard discovery bots to access your pages so they can index your written work.

Only after these foundations are in place does technical formatting matter. Adding semantic markup, structured data, or an llms.txt file is a reasonable finishing touch. But treating the finishing touch as the starting point leaves a site with a well-labeled table of contents pointing to empty shelves.

## Where is the full audit?

This article highlights our findings regarding llms.txt adoption and server fallbacks. The complete audit of machine readability, crawler permissions, and platform tiers across long-form video channels is available in [The Dark Archive report](https://exdarkmatter.com/report/).

To evaluate how your own channel's website performs across machine readability standards, [run the free Dark Matter Check](https://exdarkmatter.com/check/).

## Sources

- [Ahrefs: llms.txt study](https://ahrefs.com/blog/llmstxt-study/), retrieved 2026-09-30
