Noah Landow said something on my podcast that changed how I read my own analytics. He’s run an IT firm since 1996 and watched several waves of this. His point was that the useful distinction in AI search isn’t between the acronyms. AEO, GEO, and AI search optimization are the same thing wearing different labels. The distinction that matters is between a crawl that trains a model and a live query from a person.
Separate those two and most of the current advice sorts itself into useful and pointless.
Most of what gets sold in this field makes me impatient. Writers and small business owners are paying for reports full of crawl counts that will never put a single reader in front of them, and plenty of the people selling those reports know it. Noah’s distinction is the first thing I’d teach anyone before they spend money on AI search, because it exposes most of the waste in one step.
What is the difference between a training crawl and a live query?
A training crawl is bulk collection. A company is building or updating a model and needs enormous quantities of text, so its crawler works through the web taking pages.
Your article goes into a dataset. Whatever the model learns from it diffuses into weights, alongside millions of other documents, and nothing about that process involves a person wanting to know something today.
A live query is the opposite. Somebody types a question into an assistant. The system decides it needs current information, fetches a handful of pages right then, reads them, and composes an answer citing what it used. That request exists because a human being asked for something a few seconds ago.
Both show up in your server logs as a bot. That’s the trap. They look similar and they mean different things for your business.
Why does the difference between training crawls and live queries matter?
One of them can produce a client. The other mostly cannot.
A live query can put your name and your page in front of somebody at the exact moment they’re trying to solve a problem you solve. Some of those people click through. More of them don’t, and they still leave with your name attached to a good answer. That’s a recommendation.
A training crawl gives you no attribution, no click, and no moment. Your material contributes to a model’s general competence and the model doesn’t say where anything came from. There are arguments about whether that’s fair, and separately there are arguments about whether being in the training data helps you get mentioned later. Neither argument is settled, and neither pays this quarter.
So when somebody tells you their AI traffic tripled, ask which kind. Tripling training crawls means your bandwidth bill went up.
Why the acronyms do not matter
AEO, GEO, LLMO, AI search optimization. A small industry is inventing terms for this, and the terms are interchangeable. Noah put it plainly: they’re the same thing under different acronyms.
Watch what happens whenever a field produces four names for one practice inside two years. It means nobody has a durable method yet, and the naming is standing in for results. The same thing happened with content marketing, with growth hacking, and with a dozen others. Each time the useful core turned out to be small and unglamorous. The vocabulary was the product.
Don’t ignore the field. Buy on evidence instead of terminology. Ask anyone selling you a service which of the two crawl types they intend to move and how they’ll show you it happened. If they can’t answer, I’d walk away, because a vendor who can’t name what they’re moving is selling you the vocabulary.
How do you tell training crawls from live queries in your logs?
The separation is imperfect, and do it anyway.
The major AI companies publish which of their user agents do what, and they generally run separate crawlers for training collection and for live retrieval. Those are documented, they change, and checking the current documentation once a quarter is all the maintenance there is.
The behavioral tells are more durable than the names. A training crawl moves in bulk: many pages, systematically, often deep into your archive, in patterns that look like somebody working through a sitemap.
A live-query fetch is a small number of pages, often one. It targets something specific, usually recently published or unusually relevant to a narrow question. It arrives in ones and twos instead of waves.
Then there’s referral traffic. That settles it. When a person follows a citation, that arrives as an ordinary visit with a referrer from the assistant.
That’s the real measure of whether any of this is working, and almost nobody looks at it. The number is usually small, and the crawl numbers are large and flattering.
I think the flattering numbers are the most dangerous part of all this.
A business owner sees crawl volume climb, decides the strategy is working, and keeps paying for it for a year while the referrals stay flat. Nobody told them an outright lie. They were handed the wrong number and nobody bothered to correct them.
What should you optimize for?
Live queries. The work is the same work you should be doing anyway.
Answer the question in the first sentence of the section, then explain. A retrieval system takes passages, and the system can lift a passage that opens with the answer whole. I go through the mechanics of this in writing to get found by AI instead of just ranked.
Write sections that stand alone. If a paragraph depends on three above it to make sense, nobody can use it out of context. Out of context is the only way a machine reads it. Be current, visibly. Live retrieval exists because a model’s own knowledge is stale. A page with a real update date and current figures is more useful to that system than an excellent page from 2019. Be corroborated somewhere other than your own site. Systems weigh whether a source appears to be a real entity with a track record, and anything on your own domain is easy to assert about yourself.
Should you block the training crawlers?
People ask me this constantly and I don’t think there’s one answer.
The case for blocking is that you’re supplying free raw material to a commercial product that won’t attribute you. The case against is twofold. Nobody can prove that being in the training data doesn’t help you get named later. And blocking reverses nothing, because what has been taken has been taken.
This question gets ugly fast. The vitriol about AI inside writers groups shocked me, and I’d call it among the most divisive subjects I’ve ever seen. Most of that anger lands on the training crawl, and that crawl matters least to whether a reader finds you next week. I understand the anger. I’d still like writers to make this call with their referral numbers in front of them.
What I wouldn’t do is block the retrieval crawlers. Those are the ones that fetch a page because a person asked a question, and blocking them removes you from exactly the moment you want to be in. Getting this wrong is easy, since the user agents look similar and a blunt rule catches both.
If you block anything, block narrowly, document what you blocked and when, and check your referral numbers before and after. That’s how you find out whether the decision cost you anything.
The number that tells you the truth
One measure survives all of this: are people arriving from AI assistants, and are they the right people.
Everything else is a proxy. Crawl volume is a proxy for interest that may not exist. Citations you spot by asking the assistants yourself are a sample of one. Referral traffic with a decent time on page pays.
My own measurement produced an unwelcome result. My long guide carries close to eighty percent of the commercial citations my site earns.
My service pages carry none. The guide answers questions and the service pages announce offerings, and answer engines quote the first kind. I couldn’t see that gap until I separated what the crawlers took from what the answers cited.
The unglamorous conclusion
Most of what people sell as AI search optimization is one of two things. The ordinary discipline of writing clearly, or optimizing for a crawl that will never send you a person.
The work is narrow. Answer questions directly. Write in passages that stand alone. Keep the material current. Be a real entity with corroboration off your own site. And measure referrals instead of crawls. There’s no trick underneath it. That’s why so much of the advice in this space needs dressing up.
A book does all of this better than a website can, because a book is the corroboration, and I set that out in why a book is the strongest AEO asset you can build. Noah’s full conversation is in his episode of Leaders and Their Stories, and if you want the skeptical case against all of it, I made it in the cracks in AI search.
It makes me angry how much money writers have already spent on this. Most of it went to optimizing crawls that were never going to send anyone, sold under four acronyms for one idea. Write clearly, answer the question, stay current, and watch the referrals. Anyone charging you for more than that owes you proof, and I haven’t seen many of them produce it.
