How AI Search Actually Works & How I’d Build for It From Scratch
Share
AIO and GEO become easier to understand once you stop starting with the acronyms.
Someone asks a question.
The system has to figure out what the person means, find information that might help and produce some kind of response.
That is the basic problem.
Google, OpenAI and Microsoft are obviously not running identical systems, and nobody outside those companies has the complete decision process showing exactly why one source is selected and another one is not.
I get suspicious pretty quickly when somebody claims to have that formula.
But enough pieces are public now that we can build a useful working model without pretending it is a disclosed universal architecture.
For the way I think about it:
Intent → discovery → understanding → retrieval → usefulness → corroboration → answer and citation
That is our practical model. The real systems are more complicated, differ by platform and contain proprietary mechanisms we do not get to see.
Google openly describes retrieval-augmented generation and query fan-out in its generative Search guidance. OpenAI documents that ChatGPT Search can rewrite a user’s question into one or more targeted searches and issue additional searches after reviewing initial results. Bing now gives publishers a limited window into citation activity through cited URLs and sample grounding-query phrases in Webmaster Tools. (Google for Developers)
That is enough to stop guessing at everything.
Start with what the customer actually means
Search marketing spent years training businesses to look at short phrases.
Outdoor kitchen contractor.
Inventory management software.
Best travel backpack.
People do still search that way. But conversational search lets them explain the whole problem.
Someone might now ask:
I want to rebuild an outdoor entertaining area, but the property gets heavy afternoon sun and strong seasonal wind. What should I plan before I start talking to contractors?
That question is about more than finding a contractor.
There are probably issues involving shade, materials, utilities, wind exposure, cooking equipment, drainage, layout, budget and maintenance.
Google describes query fan-out as a set of concurrent related queries generated to gather more information around the user’s original request. ChatGPT Search can also rewrite a question into one or more targeted searches and issue additional, more specific searches after seeing initial results. (Google for Developers)
This changes how I think about keyword research.
I still want to know what people search.
But I also want to know what they are trying to work out.
Those are not always the same thing.
A spreadsheet may tell me that outdoor kitchen cost gets searched frequently. Talking to the people who actually sell and build them may reveal that customers are much more worried about how utilities get routed, which materials survive exposure and what they need to decide before concrete work starts.
That second layer often produces the better content.
And then the information has to be available.
This sounds almost too obvious to mention, but there is a surprising amount of AI-search advice being sold to companies whose important pages are poorly linked, difficult to crawl or not indexable.
For Google, standard Search eligibility still matters: pages must meet the technical requirements for Search, be indexed and be eligible to appear with a snippet before they can qualify for generative Search visibility. Google has also begun rolling out a Search Console control governing whether a site’s content can appear as links and help ground responses in its generative Search features. Inclusion is the default, although the control is still available only to a subset of website owners while Google tests the rollout. (Google for Developers)
OpenAI uses OAI-SearchBot for automatic search crawling. OpenAI says any public website can potentially appear in ChatGPT Search and recommends allowing OAI-SearchBot to help ensure content can be discovered, surfaced and clearly cited and linked. SearchBot access is also required for site content to be included in ChatGPT summaries and snippets. Opting out limits search-answer use, but OpenAI notes that navigational links can still appear in some circumstances. (OpenAI Help Center)
Before I worry about being cited by an AI answer, I want to know that the important information can actually be found.
Finding a page is not the same as understanding it
A good site should make the relationships fairly obvious.
This company offers these services.
These are the products.
These people work here.
This is the type of customer the business serves.
These are the areas where it has actual experience.
These pages explain those areas in more detail.
A surprising amount of business copy works against that.
Take a consulting firm whose homepage says:
Creating transformative strategies for tomorrow’s leaders.
That could be an accounting company, management consultant, leadership coach, software vendor or somebody running corporate retreats.
Now compare it with a firm that says it advises privately held construction and manufacturing companies on ownership transitions, restructuring and financial planning.
Less clever.
Much more useful.
Structured data can help Google understand standardized information and make pages eligible for certain Search features, but Google is explicit that there is no special structured-data format required specifically for generative Search. (Google for Developers)
Schema is helpful plumbing.
It does not turn vague information into good information.
Retrieval is where things get more interesting
A search-backed AI system does not need to place the entire public web into the context of every answer.
Relevant information can be retrieved for the particular request.
Google describes this directly through retrieval-augmented generation, or RAG. In its generative Search systems, Google says its core Search ranking systems retrieve relevant, current pages from the Search index and that information can then be used to ground a response. (Google for Developers)
Bing’s terminology gives site owners another useful clue. Its AI Performance report includes grounding queries—key phrases used when retrieving content that was later referenced in supported AI-generated answers. Microsoft only shows a sample of overall citation activity, so this should not be mistaken for a complete view of retrieval. (Bing Blogs)
That gives marketers a better objective than “write content for AI.”
I do not know what writing “for AI” is supposed to sound like anyway.
The practical goal is to publish information useful enough to be available when the right question is being answered.
There is an important limit here, though.
Google tells us some of how its own retrieval and grounding work. Microsoft exposes some citation and grounding data. OpenAI publishes parts of how Search queries are formed and how publishers can make content available.
None of them publishes the complete source-selection and evaluation formula.
So when someone produces a chart claiming a certain type of mention is worth 1.7 times another one in “the LLM algorithm,” I would want to see where that number came from.
Usually that is where the conversation gets quieter.
Average information has become very cheap
Consider two companies selling outdoor equipment.
One publishes:
10 Tips for Your Next Camping Trip
The other publishes:
We Tested Four Tent Fabrics Through a Full Wet Season: What Stretched, What Leaked and What We Would Use Again
The second company has the opportunity to show actual testing, photographs, measurements, failure points and methodology.
Maybe the test is not perfect. Maybe one result surprises them.
Even better. At least there is something there to evaluate.
The first article could have been produced by almost anybody.
Google’s current generative Search guidance specifically encourages unique, expert-led, non-commodity content and warns against simply recycling information that already exists or could easily be produced by a generative model. (Google for Developers)
This is probably one of the bigger practical opportunities AI has created for good businesses.
It raised the floor on generic writing.
Everybody can now produce average information quickly.
The businesses that actually know something still have an advantage, provided they bother to publish it.
I would also pay attention to how information is written inside the page.
A generative Search system may surface a relevant portion of a page rather than treating every page as one indivisible answer. Google explicitly says its systems can understand multiple topics within a page and surface relevant material without requiring publishers to artificially break everything into tiny sections. (Google for Developers)
So a case study should actually explain what happened.
A comparison needs to make clear what is being compared.
A technical explanation should not hide the useful answer in three paragraphs of throat-clearing.
That does not mean writing like a machine.
Clear writing works fine.
Authority is not the number of times you call yourself an authority
Imagine a wellness brand claiming it is one of the leaders in its category.
The claim appears on its homepage.
Then in a blog post.
Then in a founder biography.
Then a press release the company wrote about itself.
Then ten social posts.
There are plenty of mentions, but it is still basically one source talking.
Now imagine another company whose products have been independently reviewed, whose founders have been interviewed in credible publications, whose customers have left substantial feedback and whose expertise gets referenced by other organizations.
Different situation.
That outside evidence is valuable because the business does not control all of it.
I would be careful, though, about turning that idea into another fake signal-building exercise. Google specifically warns against seeking inauthentic mentions across the web in the hope of influencing generative Search. (Google for Developers)
Do not build fake evidence.
Build a company that creates real evidence.
Sometimes this takes much longer than marketers would like. That does not make it less useful.
And I would use the word corroboration carefully here.
It is sensible to create an information environment where claims about a business are supported outside the company’s own website. But the platforms do not publish a universal rule telling us how many reviews, mentions, citations or independent sources are required before an AI system “trusts” a company.
That kind of precision is not public.
The answer can now sit between the search and the click
Traditional search generally made the user perform more of the synthesis.
Open one site.
Back out.
Open another.
Compare.
Search again.
Generative search can perform part of that work inside the answer itself.
A product may appear in a comparison.
A technical article may support a factual statement.
A business might be part of a shortlist.
A page may be cited even if the user does not immediately click it.
This is why visibility is getting broader than rankings and organic sessions alone.
But measurement needs the same restraint as everything else.
Google began rolling out dedicated Generative AI performance reporting in Search Console in June 2026, initially to a subset of websites. The reports can show how often URLs appeared in Google’s generative Search and Discover features, which pages appeared, country and device information, and changes over time. (Google for Developers)
That is useful.
It does not tell us every query behind every appearance or disclose why Google selected a particular page.
Bing goes a little farther in a different direction. Its AI Performance report—currently in public preview—shows citation counts, cited pages, visibility trends and a sample of grounding-query phrases across supported Microsoft AI experiences. Microsoft specifically says those citation figures do not indicate ranking, authority, page importance or where a citation appeared within an answer. (Bing Blogs)
So these tools are becoming good for questions like:
Are we appearing at all?
Which pages are involved?
What kinds of queries seem connected?
Is visibility increasing or falling over time?
They are not giving us the complete retrieval, evaluation or source-selection formula.
That difference is worth keeping in mind when somebody turns a dashboard into certainty.
If I were starting with an empty domain
I would not begin with GEO.
I would begin with facts.
What is this company actually called? Who works there? What does it sell? What can it legitimately claim expertise in? Which products, services and customer problems actually generate revenue?
That becomes the truth layer.
Then I would build the website around how customers understand the business, not around an internal org chart and not around a giant list of keywords.
A design-build company might need strong pages around its major services, project types, process, portfolio and expertise.
A product company is different. Categories, product pages, comparisons, support information and strong product data probably matter more.
Software is different again. Product, use cases, integrations, implementation, security and documentation may carry a huge portion of the weight.
I would build those commercial pages before producing fifty supporting articles.
Once the core is there, start collecting actual customer questions.
A specialty outdoor contractor might hear:
What should be designed before concrete is poured?
Which countertop materials hold up outside?
How should utilities be planned?
Which features turn into maintenance problems?
There are probably dozens more, and I would not turn every one into a 1,000-word page. Some deserve a paragraph. Some deserve video. A few may be worth a serious article.
Then I would look for material competitors cannot reproduce just by opening the same AI tool.
Project lessons.
Testing.
Original numbers.
Case studies.
Photos.
Comparisons.
Failed approaches.
Customer situations.
Internal expertise.
Sometimes the best authority piece is not particularly glamorous. It is just useful.
From there, build the outside evidence. Reviews, legitimate publications, professional references, partnerships, product reviews, expert commentary. Whatever makes sense for that category.
And then distribute.
A real field test could become a long article, a video, three shorter clips, a customer email and a useful addition to the product page.
A strong case study can support sales conversations, LinkedIn content and an industry presentation.
I do not think every platform needs its own invention every morning.
Use the knowledge you already paid to create.
What I would not spend much time on first
I would not begin by publishing an llms.txt file and calling the project complete.
I would not generate hundreds of near-duplicate question pages.
I would not rewrite normal sentences into awkward fragments because somebody decided language models require everything in “answer blocks.”
I would not manufacture brand mentions.
And I would not bury a simple business underneath layers of structured data that say more than the visible site actually does.
For Google Search specifically, its current guidance directly rejects the need for special AI markup, llms.txt optimization and artificial content chunking for its generative Search features. (Google for Developers)
Most companies have better things to fix first.
Make the business clear.
Make the important information accessible.
Publish something that comes from actual knowledge.
Give independent people reasons to validate it.
Then watch what search systems actually retrieve, cite and surface using the limited but improving data the platforms make available.
That is less exciting than promising to “rank a brand in ChatGPT.”
It is also a much more serious way to build one.