Skip to content
2pizza.teamBlog

Generative Engine Optimization: What the Data Actually Supports in 2026

Ivan Bolonikhin
Founder, 2pizza.team

TL;DR: The order that matters is access, entity, shelves, placements. Placements are the main lever and the slowest. llms.txt is placebo. General schema markup does not move citations. Nobody can guarantee you a hit rate, and the lag from publication to appearing in answers runs four to eight weeks. Everything below is sourced.

Generative Engine Optimization is eighteen months old as a category and almost entirely unpoliced. That combination produces a lot of confident decks. This post is the version of the playbook we actually run, with the evidence attached and the weak parts labelled as weak.

What GEO is, in one paragraph

Classic search optimization tries to rank a page. Generative engine optimization tries to get your brand named and cited inside an answer that an engine composes on the fly. The buyer never sees ten blue links - they see three names and a paragraph of reasoning. If you are not in the three, the ranking of your page is irrelevant to that particular buyer.

Where brands actually drop out

An engine assembles an answer in stages: retrieve candidate sources, resolve the entities involved, decide which names belong in the set. You can fall out at any of the three, and the fixes are completely different at each stage. Diagnosing the wrong stage is the most common way a GEO budget gets wasted.

  • Retrieval - the crawler never got your content. Blocked at the firewall, or served an empty shell because your app renders in the browser.
  • Entity - it retrieved you but is unsure who you are. Inconsistent naming, two brands under one legal entity, prices that disagree between pages.
  • Set membership - it retrieved you, it understands you, and it still left you out, because the sources it trusts for your category never mention you.

Layer one: access, and the Cloudflare problem

Since July 2025 Cloudflare has blocked AI crawlers by default on new zones. This is the single most common blocker we find, and most teams have no idea it is switched on. If OAI-SearchBot cannot fetch your pages, you are absent from ChatGPT Search regardless of the quality of anything else you have built. It costs nothing to check and usually one afternoon to fix, which makes it the highest return per hour in the entire discipline.

The second access failure is client-side rendering. Crawlers do not execute your JavaScript. A single-page app can serve a two-kilobyte shell to every engine while looking perfect in a browser. Check what the bot receives, not what you see. If the answer is an empty shell, prerendering is your first project and nothing else in this list matters until it ships.

Layer two: entity hygiene, which is hygiene and not growth

One Organization object with a stable identifier. Consistent name, founding date and address everywhere. sameAs pointing only at profiles that exist. No contradictory facts across pages - pricing is the usual offender, where an old landing page still shows a number the current pricing page abandoned. This work changes how an engine answers a question like 'is this company legitimate'. It does not get you into 'best X'.

A specific trap worth naming: if two brands sit under one legal entity, engines often merge them into a single confused entity. We found exactly this on our own properties - one company name appearing as the parent of two unrelated products, which produced hedged answers about both. Splitting the attribution is unglamorous work with a real payoff.

Layer three: shelves, or pages worth quoting

Engines quote passages. A page that never states a clean, extractable claim gives an engine nothing to lift. The pattern that works is a category hub, honest comparison pages, short answer paragraphs that stand alone without their surrounding context, FAQs, and a named author with a real profile behind the name.

The discipline here is negative rather than positive: only build the shelf for categories where you are genuinely strong. A page that overclaims does get retrieved, and then you are cited badly, which is worse than not being cited. We have turned down shelf-building work for exactly this reason more than once.

Layer four: placements, the main lever

This is where the budget should go, and it is the part most agencies skip because it is slow and cannot be shipped from a code editor. Ahrefs looked at roughly 75,000 brands and found brand mentions correlate with AI visibility at r = 0.664, while backlinks correlate at 0.218. Brands with independent media coverage are cited far more often than brands relying on their own pages.

Read that as correlation, because that is what it is. Well-known brands get both more mentions and more AI visibility, and the study cannot separate the two. But the direction is consistent enough, and the mechanism is plausible enough, that spending on third-party presence rather than on link buying is the better bet with the evidence available.

The practical method is to map, not guess. Run your buying prompts, record which domains the engine actually cites for your category, and treat that list as your target list. The sources an engine already trusts are not always the ones you would have picked.

One caveat on listicles

Analysis by Peec across roughly 200,000 answers found that being first in a listicle associated with a large visibility difference. That result is observational and everyone quoting it as a tactic is overreaching. More importantly, more recent model versions have started cutting listicle citations, so a placement strategy built entirely on roundups is fragile. Mix media coverage, review sites and directories deliberately.

What does not work, with sources

Two things dominate competitor decks and neither survives contact with the data.

  • llms.txt - Ahrefs examined 137,000 domains and found 97% of these files received zero requests. Retrieval bots made up about one percent of the small remainder. Google has said it will not use the file. Ship one in ten minutes if it makes someone happy; do not put it on an invoice.
  • General schema markup as a citation lever - across 1,885 pages Ahrefs found no citation effect. Specific Product and Review markup with real prices does matter. Everything else is entity hygiene, which is worth doing for a different reason at a different price.

How to measure without fooling yourself

Model answers are not deterministic. A single run of a single prompt tells you almost nothing, and any before-and-after built on single runs is noise dressed as a result. The workable method is a fixed panel of buying prompts, three runs per prompt, the same engine, recorded identically before and after.

  • Recall prompts - 'best X for Y in 2026'. Four or more, because this is what you are actually buying.
  • One brand prompt - 'is X legit, what do you know about them'. This measures entity health.
  • Citation prompts - 'how do I do Z'. These measure whether your content gets quoted at all.
  • Record three things per run: were you named, in what position, and which sources the answer cited.

Timeline expectations, honestly

The lag from a placement going live to it showing up in answers runs roughly four to eight weeks. That means a thirty-day engagement ends before the first real signal arrives. It also means that if someone shows you a dramatic before-and-after after three weeks, they measured noise. Plan in quarters or do not start.

What we would tell you not to buy

If your crawler is blocked, buy one afternoon of firewall configuration and re-measure in six weeks before spending anything else. If your site is a client-rendered app serving empty shells, buy prerendering and nothing else. If you are not genuinely strong in the category you want to be named in, buy product work rather than visibility work, because being cited for something you do badly is a liability.

We ran this whole pipeline on our own properties before we sold it to anyone, including the parts of the report that were uncomfortable to read. If you want the same diagnostic run on yours, the audit is at /geo and it stands alone - the findings are yours whether or not you continue with us. Ivan / 2pizza.team

Want us to look at your setup?

Free 30-min audit. We tell you what to automate first and what it would cost.

Book a free audit
GEO & AI Search
Does llms.txt Do Anything? The Data Says No
9 min read
GEO & AI Search
2,591 Bytes: The Shop That Did Not Exist for AI Crawlers
9 min read
Australia
The Privacy Act and AI Automation: What an Australian Business Actually Has To Do
13 min read