The Writing King Your Ethical Ghostwriter. Your Story, Done Right.

Training Crawls and Live Queries Are Not the Same Thing

This entry is part 9 of 10 in the series Your Website Is Your Beacon
TL;DR: People call two completely different things AI traffic. A training crawl is a machine reading your site to teach a model, and it may never send you a human being. A live query is a person asking a question right now and a system fetching your page to answer it. Only the second one is traffic. Optimizing for the first is how people end up celebrating numbers that never turn into anybody.

Noah Landow said something on my podcast that changed how I read my own analytics. He has run an IT firm since 1996 and watched several waves of this. His point was that the useful distinction in AI search is not between the acronyms. AEO, GEO, and AI search optimization are the same thing wearing different labels. The distinction that matters is between a crawl that trains a model and a live query from a person.

Separate those two and most of the current advice sorts itself into useful and pointless.

What is the difference between a training crawl and a live query?

A training crawl is bulk collection. A company is building or updating a model and needs enormous quantities of text, so its crawler works through the web taking pages. Your article goes into a dataset. Whatever the model learns from it diffuses into weights, alongside millions of other documents, and nothing about that process involves a person wanting to know something today.

A live query is the opposite. Somebody types a question into an assistant. The system decides it needs current information, fetches a handful of pages right then, reads them, and composes an answer citing what it used. That request exists because a human being asked for something a few seconds ago.

Both show up in your server logs as a bot. That is the trap. They look similar and they mean different things for your business.

Why the difference matters

One of them can produce a client. The other mostly cannot.

A live query can put your name and your page in front of somebody at the exact moment they are trying to solve a problem you solve. Some of those people click through. More of them do not, and they still leave with your name attached to a good answer. That is a recommendation.

A training crawl gives you no attribution, no click, and no moment. Your material contributes to a model’s general competence and the model does not say where anything came from. There are arguments about whether that is fair, and separately there are arguments about whether being in the training data helps you get mentioned later. Neither argument is settled, and neither pays this quarter.

So when somebody tells you their AI traffic tripled, ask which kind. Tripling training crawls means your bandwidth bill went up.

Why the acronyms do not matter

AEO, GEO, LLMO, AI search optimization. A small industry is inventing terms for this, and the terms are interchangeable. Noah put it plainly: they are the same thing under different acronyms.

Watch what happens whenever a field produces four names for one practice inside two years. It means nobody has a durable method yet, and the naming is standing in for results. The same thing happened with content marketing, with growth hacking, and with a dozen others. Each time the useful core turned out to be small and unglamorous. The vocabulary was the product.

Do not ignore the field. Buy on evidence instead of terminology, and to ask anyone selling you a service which of the two crawl types they intend to move, and how they will show you it happened.

How do you tell them apart in your own logs?

The separation is imperfect, and do it anyway.

The major AI companies publish which of their user agents do what, and they generally run separate crawlers for training collection and for live retrieval. Those are documented, they change, and checking the current documentation once a quarter is all the maintenance there is.

The behavioral tells are more durable than the names. A training crawl moves in bulk: many pages, systematically, often deep into your archive, in patterns that look like somebody working through a sitemap. A live-query fetch is a small number of pages, often one. It targets something specific, usually recently published or unusually relevant to a narrow question. It arrives in ones and twos instead of waves.

Then there is referral traffic, which settles it. When a person follows a citation, that arrives as an ordinary visit with a referrer from the assistant. That is the real measure of whether any of this is working, and almost nobody looks at it. The number is usually small, and the crawl numbers are large and flattering.

What should you optimize for?

Live queries. The work is the same work you should be doing anyway.

Answer the question in the first sentence of the section, then explain. A retrieval system takes passages, and the system can lift a passage that opens with the answer whole. I go through the mechanics of this in writing to get found by AI instead of just ranked.

Write sections that stand alone. If a paragraph depends on three above it to make sense, nobody can use it out of context. Out of context is the only way a machine reads it.

Be current, visibly. Live retrieval exists because a model’s own knowledge is stale. A page with a real update date and current figures is more useful to that system than an excellent page from 2019.

Be corroborated somewhere other than your own site. Systems weigh whether a source appears to be a real entity with a track record, and anything on your own domain is easy to assert about yourself.

Should you block the training crawlers?

People ask me this constantly and I do not think there is one answer.

The case for blocking is that you are supplying free raw material to a commercial product that will not attribute you. The case against is twofold. Nobody can prove that being in the training data does not help you get named later. And blocking reverses nothing, because what has been taken has been taken.

What I would not do is block the retrieval crawlers. Those are the ones that fetch a page because a person asked a question, and blocking them removes you from exactly the moment you want to be in. Getting this wrong is easy, since the user agents look similar and a blunt rule catches both.

If you block anything, block narrowly, document what you blocked and when, and check your referral numbers before and after. That is how you find out whether the decision cost you anything.

The number that tells you the truth

One measure survives all of this: are people arriving from AI assistants, and are they the right people.

Everything else is a proxy. Crawl volume is a proxy for interest that may not exist. Citations you spot by asking the assistants yourself are a sample of one. Referral traffic with a decent time on page pays.

My own measurement produced an unwelcome result. My long guide carries close to eighty percent of the commercial citations my site earns. My service pages carry none. The guide answers questions and the service pages announce offerings, and answer engines quote the first kind. I could not see that gap until I separated what the crawlers took from what the answers cited.

The unglamorous conclusion

Most of what people sell as AI search optimization is one of two things. The ordinary discipline of writing clearly, or optimizing for a crawl that will never send you a person.

The real work is narrow. Answer questions directly. Write in passages that stand alone. Keep the material current. Be a real entity with corroboration off your own site. And measure referrals instead of crawls. There is no trick underneath it. That is why so much of the advice in this space needs dressing up.

A book does all of this better than a website can, because a book is the corroboration, and I set that out in why a book is the strongest AEO asset you can build. Noah’s full conversation is in his episode of Leaders and Their Stories, and if you want the skeptical case against all of it, I made it in the cracks in AI search.

The Guides That Get Your Book Written, Published, and Sold

Four short, practical guides on writing, publishing, and selling your book, plus the occasional note when there's something worth your time. No fluff, no daily inbox clutter. Drop your email and they're yours.

We use MailerLite to manage our list and send these emails. Your address is used only to send you what you signed up for. We will not sell it, share it, or use it for anything else, and you can unsubscribe anytime.

Frequently Asked Questions

What is the difference between an AI training crawl and a live query?
A training crawl collects text in bulk to build or update a model, with no attribution and no person involved at the time. A live query happens because somebody asked a question seconds earlier and the system fetches pages to compose an answer, usually citing what it used.
Does being in AI training data send you traffic?
Generally no. Training contributes your material to a model’s diffuse competence without attribution or clicks. Whether it improves your chances of being mentioned later is unresolved, so do not count it as traffic in any reporting.
How can you tell training crawls from retrieval fetches in server logs?
Check the published user agent documentation from the major AI companies, which generally separates the two, and watch behavior: training crawls move through many pages systematically, while retrieval fetches hit one or a few specific pages in response to a question.
What metric shows whether AI search is working for you?
Referral traffic from AI assistants, with reasonable time on page, is the only measure that reflects actual people. Crawl volume measures machine interest, and self-run citation checks are a sample of one.
Should you block AI crawlers?
If you block anything, block training crawlers narrowly and leave retrieval crawlers alone, since those are the ones fetching your page because a person asked a question. Document what you blocked and compare referral numbers before and after.
How do you write pages that AI answers use?
Answer the question in the first sentence of each section, write sections that make sense read alone, keep figures and dates current with visible updates, and build corroboration of who you are on sites other than your own.

📁︎ Book Marketing📁︎ Technology

🏷︎ AEO🏷︎ AI Search🏷︎ Answer Engine Optimization🏷︎ SEO

📝 Disclaimer

The views and opinions expressed in this blog post are solely those of Richard Lowe and are based on personal experience and research. This content is for informational purposes only and should not be construed as professional legal, financial, accounting, or business advice. Always consult with qualified professionals before making important business or legal decisions. Richard Lowe is not a lawyer, accountant, or licensed professional advisor, and this content does not establish any professional relationship.

Leave a Reply

Your email address will not be published. Required fields are marked *