← All posts

What the research actually says about GEO (and what it doesn't)

· Strategy

A magnifying glass focuses on three connected icons (a document, a target and a rising chart) while scattered checklist, star and tag icons drift out of focus

The GEO industry has produced a great many checklists. The research community has produced rather fewer certainties. It's worth knowing the difference before you spend next quarter's content budget.

Generative engine optimisation grew up fast, and much of its received wisdom arrived before the evidence did. Add statistics. Add quotations. Add schema. Write longer. Rewrite for "citability". Some of that advice holds up. Some of it doesn't. This summer brought the most rigorous look yet at which is which.

The state of the evidence: messier than the checklists suggest

In July, researchers published Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026), reviewing 45 studies published between November 2023 and July 2026 (arXiv). Its headline conclusions are sobering:

  • The field doesn't yet agree on terms, metrics or standards of evidence. Different studies measure "visibility" in incompatible ways, which makes many results hard to compare.
  • No technique reviewed shows a stable, long-term effect across platforms. Tactics that work on one engine, one month, in one experimental set-up often fail to transfer.
  • Generic heuristics transfer poorly, and gains can be eroded by competition: when everyone applies the same trick, nobody benefits.
  • Citation-oriented rewrites can actually harm retrieval. Editing a page to look more "citable" can make it less likely to be found in the first place.

The one robust finding? Content that has already been retrieved can causally influence whether it gets cited. That points to a two-stage process, and it's the key to most of what follows.

Two gates, not one

Think of an AI answer as having two gates. First, your page has to be retrieved: found and pulled into the pool of candidate sources. Then it has to be cited: chosen from that pool to support the answer.

One large factorial experiment highlighted in the survey, 252,000 controlled trials across six LLMs and 18 content factors, tested what decides the second gate. Its conclusion: relevance and position are the primary determinants of which source gets cited first. Citation behaves like a bottleneck after retrieval, and it rewards the source that most directly answers the question, most prominently (arXiv survey).

That's consistent with the survey: topical relevance and the position of the answer within the context are the most reproducible levers. It's also a fairly unglamorous result. The most reliable way to be cited is to answer the question clearly, early and precisely.

What the large-scale data adds

Lab experiments are one thing; the live web is another. Two large studies from Ahrefs this summer fill in the picture:

  • Freshness matters. Across 17 million citations on seven AI platforms, AI assistants showed a preference for fresher content (Ahrefs).
  • Adding schema alone didn't move citations. In a study of 1,885 pages that added JSON-LD, Ahrefs found no meaningful uplift in AI citations (Ahrefs).

That second finding deserves an honest word from us. We've argued that JSON-LD is the hidden layer behind AI visibility, and we still think structured data matters. It tells machines unambiguously what a page is, who wrote it and which entities it's about, which helps with accuracy, entity recognition and eligibility for rich results. But the evidence says it isn't a citation lever in its own right. Schema helps AI systems get your facts right; it doesn't persuade them to choose you.

Meanwhile, on the content side, Search Engine Land's analysis found that the theme of a blog post predicts LLM traffic more reliably than almost any other variable, and that shorter posts built around unique data outperform long "comprehensive" guides (Search Engine Land). Length, it turns out, is not a proxy for citability.

What this means in practice

Pulling the threads together, here's what the evidence currently supports, from strongest to weakest:

  1. Be the most relevant answer. Relevance is the one lever every study agrees on. Pick narrower questions you can answer better than anyone, rather than broad topics you can only answer as well as everyone.
  2. Put the answer first. Position matters. Lead with the direct answer, the key figure or the definition, then elaborate.
  3. Get retrieved before you worry about being cited. Technical access, rendering and ordinary search visibility come first. You can't win the second gate without passing the first. (See Is your site blocking AI search without knowing it? and GPTBot can't see your JavaScript.)
  4. Keep important pages fresh. Update and re-date content when facts genuinely change.
  5. Publish something only you have. Original data, first-hand experience and specific examples. This is the thread running through the Search Engine Land findings and through E-E-A-T.
  6. Treat schema as hygiene, not a hack. Implement it properly for accuracy and entity clarity; don't expect it to lift citations on its own.
  7. Be sceptical of "citation rewrites". Rewriting pages to game AI citation risks harming retrieval, and Google's June 2026 spam update explicitly brought AI manipulation tactics into scope. We looked at the darker side of this in GEO as a threat surface.

And one organisational finding worth noting: in Semrush's 2026 AI Visibility Index, 81% of organisations that run SEO and AI visibility as one workflow reported more traffic or leads from AI platforms, against 36% of those managing the two separately (Semrush). That's a survey correlation, not proof of cause, but it matches Google's own position that GEO is, in large part, good SEO applied to a new surface.

Frequently asked questions

What actually gets content cited by AI search engines? The most consistent research finding is that relevance and position matter most: the source that answers the question most directly, and most prominently, tends to be cited first. Content must first be retrieved, so technical accessibility and search visibility come before citation.

Does adding schema markup increase AI citations? A 2026 Ahrefs study of 1,885 pages that added JSON-LD found no meaningful citation uplift. Structured data still helps AI systems interpret pages and entities accurately, but it isn't a citation lever on its own.

Do AI assistants prefer fresh content? Yes. Ahrefs' analysis of 17 million citations across seven AI platforms found a preference for fresher content.

Are long-form guides better for GEO? Not necessarily. Search Engine Land's analysis found shorter posts built around unique data outperformed long "comprehensive" guides, and that topic predicted LLM traffic better than length.

Is there proof that GEO techniques work? A 2026 survey of 45 GEO studies found no technique with a stable, long-term effect across platforms. The most reliable levers are relevance and the position of the answer, not platform-specific tricks.

The kicker

After two and a half years and 45 studies, the most reproducible finding in generative engine optimisation is that you should answer the question, clearly, near the top of the page. It's not the sexy answer the industry was hoping for, but it has the considerable advantage of being true.

Want to know how your pages score on the fundamentals that the evidence supports? Run a free CiteSite audit.

Sources: Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026), arXiv 2607.14035; Ahrefs: do AI assistants prefer to cite fresh content?; Ahrefs: does schema markup increase AI citations?; Search Engine Land: the SEO–GEO gap; Semrush: 2026 AI Visibility Index.