Latest
It’s Not X, It’s Y: The AI Tell Shakespeare Wrote FirstWho Gets Rich From AI Data Centers? The Local, National and Global EconomyAre AI Data Centers Bad for the Environment? What’s True and What Isn’tAI Data Centers and Geopolitics: Chips, China, Spies and PowerAI Derangement Syndrome: Who Cares If the Cover Was Made With AI?How to Write a Plot Twist Readers Never See ComingBook Ads Getting Clicks but No Sales? The Problem Is the PageCan Writing a Memoir Make You Sick? The Toll Nobody Warns You AboutWhen Caregiving Ends and the Words Won’t ComeCan You Publish a Clean and a Spicy Version of the Same Book?Do You Need a Writing Buddy? What Works and What Doesn’tCan You Put Someone Who Wronged You in Your Novel?Kindle Unlimited or Wide? Where to Publish Your Debut NovelWriting Vampires: How to Build Your Own Vampire RulesWhat Software Do Novelists Use to Write a Book?What If Your Family Doesn’t Support Your Writing?ARC Reviews: Real Follow-Through Numbers, and Do Reviews Sell Books?The AI Singularity: Would a Conscious AI Even Care About Us?“Forbidden” AI Prompts: Why Viral Prompt Lists Are Mostly JunkAuthor Scam Emails: Fake Agents, Flattery and the $500 PitchWill AI Steal My Book? Turn Off Training Before You UploadThe Fear of Being Judged for Your Memoir: My Father Told Everyone to Burn MineShould You Only Use “Said” in Dialogue Tags? My Rules for Tags and AdverbsWhy a Good Book Isn’t Enough: Agents, Quiet Novels and What SellsShe Asked How to Format Her Comic for KDP. The Thread Went to War Over AI.Why Are Writing Groups So Hostile?My Ghostwriter Stopped Responding. What Do I Do?My Ghostwriter Missed the Deadline. What Are My Options?I Hate My Ghostwriter’s First Draft. Now What?Can I Get a Refund From a Ghostwriter?How Many Revisions Should a Ghostwriter Include?Can My Ghostwriter List My Book in Their Portfolio?My Family Doesn’t Want Me to Publish My Ghostwritten Memoir. Now What?Should a Ghostwriter Write a Free Sample Chapter?How Much Does a Ghostwriter Take Up Front?Can a Ghostwriter Get Me a Book Deal or a Literary Agent?Can I Work With a Ghostwriter Over Zoom or From Another Country?Ghostwriting a Tribute Book, Eulogy or Obituary for Someone You LostCan a Ghostwriter Write My Book in Spanish?Ghostwriting a Book for a Dental or Chiropractic PracticeWhy Keynote Speakers Need a Ghostwritten BookGhostwriting a Science or Research Book for General ReadersGhostwriting a Book for Teachers and EducatorsHiring a Ghostwriter for a Pilot’s or Aviation MemoirHiring a Ghostwriter for a Musician’s MemoirCan a Ghostwriter Help Me Write a Whistleblower Memoir?Hiring a Ghostwriter for an Immigrant’s StoryCan a Ghostwriter Help Me Write About Losing Someone to Suicide?Can a Ghostwriter Help Me Write a Depression or Mental Illness Memoir?Hiring a Ghostwriter for a Special-Needs Parent’s Memoir
The Writing King Your Ethical Ghostwriter. Your Story, Done Right.

Training Crawls and Live Queries Are Not the Same Thing

TL;DR: People lump two things together as AI traffic. A training crawl is a machine reading your site to teach a model, and it may never send you a human being. A live query is a person asking a question right now and a system fetching your page to answer it, and only that counts as traffic. Optimizing for crawls is how people end up celebrating numbers that never turn into anybody.

Noah Landow said something on my podcast that changed how I read my own analytics. He’s run an IT firm since 1996 and watched several waves of this. His point was that the useful distinction in AI search isn’t between the acronyms. AEO, GEO, and AI search optimization are the same thing wearing different labels. The distinction that matters is between a crawl that trains a model and a live query from a person.

Separate those two and most of the current advice sorts itself into useful and pointless.

Most of what gets sold in this field makes me impatient. Writers and small business owners are paying for reports full of crawl counts that will never put a single reader in front of them, and plenty of the people selling those reports know it. Noah’s distinction is the first thing I’d teach anyone before they spend money on AI search, because it exposes most of the waste in one step.

What is the difference between a training crawl and a live query?

A training crawl is bulk collection. A company is building or updating a model and needs enormous quantities of text, so its crawler works through the web taking pages.

Your article goes into a dataset. Whatever the model learns from it diffuses into weights, alongside millions of other documents, and nothing about that process involves a person wanting to know something today.

A live query is the opposite. Somebody types a question into an assistant. The system decides it needs current information, fetches a handful of pages right then, reads them, and composes an answer citing what it used. That request exists because a human being asked for something a few seconds ago.

Both show up in your server logs as a bot. That’s the trap. They look similar and they mean different things for your business.

Why does the difference between training crawls and live queries matter?

One of them can produce a client. The other mostly cannot.

A live query can put your name and your page in front of somebody at the exact moment they’re trying to solve a problem you solve. Some of those people click through. More of them don’t, and they still leave with your name attached to a good answer. That’s a recommendation.

A training crawl gives you no attribution, no click, and no moment. Your material contributes to a model’s general competence and the model doesn’t say where anything came from. There are arguments about whether that’s fair, and separately there are arguments about whether being in the training data helps you get mentioned later. Neither argument is settled, and neither pays this quarter.

So when somebody tells you their AI traffic tripled, ask which kind. Tripling training crawls means your bandwidth bill went up.

Why the acronyms do not matter

AEO, GEO, LLMO, AI search optimization. A small industry is inventing terms for this, and the terms are interchangeable. Noah put it plainly: they’re the same thing under different acronyms.

Watch what happens whenever a field produces four names for one practice inside two years. It means nobody has a durable method yet, and the naming is standing in for results. The same thing happened with content marketing, with growth hacking, and with a dozen others. Each time the useful core turned out to be small and unglamorous. The vocabulary was the product.

Don’t ignore the field. Buy on evidence instead of terminology. Ask anyone selling you a service which of the two crawl types they intend to move and how they’ll show you it happened. If they can’t answer, I’d walk away, because a vendor who can’t name what they’re moving is selling you the vocabulary.

How do you tell training crawls from live queries in your logs?

The separation is imperfect, and do it anyway.

The major AI companies publish which of their user agents do what, and they generally run separate crawlers for training collection and for live retrieval. Those are documented, they change, and checking the current documentation once a quarter is all the maintenance there is.

The behavioral tells are more durable than the names. A training crawl moves in bulk: many pages, systematically, often deep into your archive, in patterns that look like somebody working through a sitemap.

A live-query fetch is a small number of pages, often one. It targets something specific, usually recently published or unusually relevant to a narrow question. It arrives in ones and twos instead of waves.

Then there’s referral traffic. That settles it. When a person follows a citation, that arrives as an ordinary visit with a referrer from the assistant.

That’s the real measure of whether any of this is working, and almost nobody looks at it. The number is usually small, and the crawl numbers are large and flattering.

I think the flattering numbers are the most dangerous part of all this.

A business owner sees crawl volume climb, decides the strategy is working, and keeps paying for it for a year while the referrals stay flat. Nobody told them an outright lie. They were handed the wrong number and nobody bothered to correct them.

What should you optimize for?

Live queries. The work is the same work you should be doing anyway.

Answer the question in the first sentence of the section, then explain. A retrieval system takes passages, and the system can lift a passage that opens with the answer whole. I go through the mechanics of this in writing to get found by AI instead of just ranked.

Write sections that stand alone. If a paragraph depends on three above it to make sense, nobody can use it out of context. Out of context is the only way a machine reads it. Be current, visibly. Live retrieval exists because a model’s own knowledge is stale. A page with a real update date and current figures is more useful to that system than an excellent page from 2019. Be corroborated somewhere other than your own site. Systems weigh whether a source appears to be a real entity with a track record, and anything on your own domain is easy to assert about yourself.

Should you block the training crawlers?

People ask me this constantly and I don’t think there’s one answer.

The case for blocking is that you’re supplying free raw material to a commercial product that won’t attribute you. The case against is twofold. Nobody can prove that being in the training data doesn’t help you get named later. And blocking reverses nothing, because what has been taken has been taken.

This question gets ugly fast. The vitriol about AI inside writers groups shocked me, and I’d call it among the most divisive subjects I’ve ever seen. Most of that anger lands on the training crawl, and that crawl matters least to whether a reader finds you next week. I understand the anger. I’d still like writers to make this call with their referral numbers in front of them.

What I wouldn’t do is block the retrieval crawlers. Those are the ones that fetch a page because a person asked a question, and blocking them removes you from exactly the moment you want to be in. Getting this wrong is easy, since the user agents look similar and a blunt rule catches both.

If you block anything, block narrowly, document what you blocked and when, and check your referral numbers before and after. That’s how you find out whether the decision cost you anything.

The number that tells you the truth

One measure survives all of this: are people arriving from AI assistants, and are they the right people.

Everything else is a proxy. Crawl volume is a proxy for interest that may not exist. Citations you spot by asking the assistants yourself are a sample of one. Referral traffic with a decent time on page pays.

My own measurement produced an unwelcome result. My long guide carries close to eighty percent of the commercial citations my site earns.

My service pages carry none. The guide answers questions and the service pages announce offerings, and answer engines quote the first kind. I couldn’t see that gap until I separated what the crawlers took from what the answers cited.

The unglamorous conclusion

Most of what people sell as AI search optimization is one of two things. The ordinary discipline of writing clearly, or optimizing for a crawl that will never send you a person.

The work is narrow. Answer questions directly. Write in passages that stand alone. Keep the material current. Be a real entity with corroboration off your own site. And measure referrals instead of crawls. There’s no trick underneath it. That’s why so much of the advice in this space needs dressing up.

A book does all of this better than a website can, because a book is the corroboration, and I set that out in why a book is the strongest AEO asset you can build. Noah’s full conversation is in his episode of Leaders and Their Stories, and if you want the skeptical case against all of it, I made it in the cracks in AI search.

It makes me angry how much money writers have already spent on this. Most of it went to optimizing crawls that were never going to send anyone, sold under four acronyms for one idea. Write clearly, answer the question, stay current, and watch the referrals. Anyone charging you for more than that owes you proof, and I haven’t seen many of them produce it.

Frequently Asked Questions

What is the difference between an AI training crawl and a live query?
A training crawl collects text in bulk to build or update a model, with no attribution and no person involved at the time. A live query happens because somebody asked a question seconds earlier and the system fetches pages to compose an answer, usually citing what it used.
Does being in AI training data send you traffic?
Generally no. Training contributes your material to a model’s diffuse competence without attribution or clicks. Whether it improves your chances of being mentioned later is unresolved, so don’t count it as traffic in any reporting.
How can you tell training crawls from retrieval fetches in server logs?
Check the published user agent documentation from the major AI companies, which generally separates the two, and watch behavior: training crawls move through many pages systematically, while retrieval fetches hit one or a few specific pages in response to a question.
What metric shows whether AI search is working for you?
Referral traffic from AI assistants, with reasonable time on page, is the only measure that reflects actual people. Crawl volume measures machine interest, and self-run citation checks are a sample of one.
Should you block AI crawlers?
If you block anything, block training crawlers narrowly and leave retrieval crawlers alone, since those are the ones fetching your page because a person asked a question. Document what you blocked and compare referral numbers before and after.
How do you write pages that AI answers use?
Answer the question in the first sentence of each section, write sections that make sense read alone, keep figures and dates current with visible updates, and build corroboration of who you’re on sites other than your own.

About the Author
Richard Lowe, professional ghostwriter

Richard Lowe is a professional ghostwriter and author with 113+ books authored and 54+ ghostwritten. Before writing full time he spent 33 years in enterprise technology, including 20 years as Director of Computer Operations and Technical Services at Trader Joe's. He writes nonfiction, fiction and memoir, and works with executives and experts on books that build authority.

More about Richard Lowe →

Disclaimer

The views and opinions expressed in this blog post are solely those of Richard Lowe and are based on personal experience and research. This content is for informational purposes only and should not be construed as professional legal, financial, accounting, or business advice. Always consult with qualified professionals before making important business or legal decisions. Richard Lowe is not a lawyer, accountant, or licensed professional advisor, and this content does not establish any professional relationship.

0 comments

No comments yet. Yours can be the first.

Was this useful?

Leave a comment