Generative Engine Optimization: What the Research Actually Tested
Sep 3, 2026 · 13 min read
What the original GEO paper measured, what it found, what it explicitly did not test, and which of its methods anyone has checked since. Written for the marketing lead being pitched a GEO retainer.

Generative engine optimization is the practice of editing content so a generative engine uses and cites more of it inside a written answer. The term comes from a November 2023 paper later published at KDD 2024, and its famous “40%” is a gain in a visibility metric the authors invented, not traffic, clicks, or rankings.
The paper’s own results table puts it precisely: the best methods beat the unoptimized baseline by 41% on position-adjusted word count and 28% on subjective impression, according to Aggarwal and colleagues in GEO: Generative Engine Optimization. One independent group has since reported a 30.96% gain on the same position-adjusted metric after fine-tuning a model to rewrite travel pages.
Neither number is a traffic figure. The paper says outright that it did not evaluate how its methods affect search rankings. We checked the ten results Google returned in the United States for this term on 31 August 2026: one repeated the paper’s headline figure, none named the metric it belongs to, and none mentioned the ranking limitation.
Key Takeaways
- The 40% is a metric, not money: it measures how much of a generated answer your source occupies, weighted by citation position. It is not clicks, sessions, revenue, or rank.
- Evidence density is the finding: adding citations, quotations, and statistics moved the metric most. Adding an authoritative tone produced no significant improvement.
- Contrarian: keyword stuffing scored worse than making no changes at all. The habit most lazily carried over from SEO is the one the research rules out.
- The gain is conditional on rank: the same edit that lifted a fifth-placed source by 115.1% cut a first-placed source by 30.3%.
- The paper tested one half of the job: every candidate source in the experiment was already a top-five Google result, so nothing in it explains how you get retrieved in the first place.
What Generative Engine Optimization Is
Generative engine optimization is the practice of changing a page so that a generative engine, the kind that writes an answer rather than listing links, includes more of that page in what it writes. The unit of success is space inside an answer, not a position on a results page. Written out in full, the phrase also appears as generative search engine optimization, and the two are the same discipline.
The definition matters because the measurement changes. Ranking asks where you sit in an ordered list. A generated answer has no list, so the research replaced position with two other questions: how many words in the answer were attributed to your source, and how prominently those citations were placed. That substitution is the intellectual core of the field, and it is why GEO numbers cannot be read like SEO numbers.
A generative engine does two separable jobs, and the distinction runs through everything below. First it retrieves a handful of candidate sources. Then it writes an answer from them. Retrieval decides whether you are in the room. Generation decides how much of the answer is yours. Most of what is sold under this label addresses the second job while being priced on the promise of the first.
Where the Term Came From
The term was coined in a single paper: GEO: Generative Engine Optimization, by Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, and Ameet Deshpande. It went up on arXiv on 16 November 2023, reached its third version on 28 June 2024, and was published on 24 August 2024 in the proceedings of KDD 2024, the ACM's peer-reviewed knowledge-discovery conference. The work was supported by the US National Science Foundation under grant 2107048.
The paper formalized “generative engines” as a category and named the optimization problem now called GEO. The paper supplied three things the field still uses: a name, a benchmark, and a pair of metrics designed for answers rather than lists. What it did not supply was a business case, and the gap between those two is where most vendor claims live.
What the Original Research Tested
The setup was deliberately narrow. The authors built GEO-bench, a benchmark of 10,000 queries split 8,000 for training and 1,000 each for validation and testing, preserving a real-world distribution of 80% informational queries and 10% each transactional and navigational. For every query they took the top five Google results, cleaned the text, and fed those five sources to a language model that wrote the answer. They then rewrote one source at a time and measured what changed.
They measured two things. Position-adjusted word count counts how many words in the answer are attributed to a source, discounting each citation by an exponentially decaying function of where it appears. Subjective impression is a set of seven scores covering relevance, influence, uniqueness, diversity, subjective position, subjective count, and the likelihood of a reader clicking the citation. Those seven were scored by G-Eval, an LLM-based evaluator, so they are a model's judgment of a model's output. Nine content edits were tested against both metrics.
The Nine Methods and What Each One Scored
Scores below are the paper's own absolute figures from its main results table. The unoptimized baseline sits at 19.3 on both metrics, so anything under that made the source less visible.
| Method | What it changes | Position-adjusted word count | Subjective impression |
|---|---|---|---|
| No optimization (baseline) | Nothing | 19.3 | 19.3 |
| Keyword Stuffing | Adds query keywords to the text | 17.7 | 20.2 |
| Unique Words | Adds uncommon vocabulary | 20.5 | 20.4 |
| Authoritative | Rewrites in a persuasive, authoritative tone | 21.3 | 22.9 |
| Easy-to-Understand | Simplifies the language | 22.0 | 20.5 |
| Technical Terms | Adds domain-specific terminology | 22.7 | 21.4 |
| Cite Sources | Adds citations to credible sources | 24.6 | 21.9 |
| Fluency Optimization | Improves the fluency of the prose | 24.7 | 21.9 |
| Statistics Addition | Replaces qualitative claims with quantitative ones | 25.2 | 23.7 |
| Quotation Addition | Adds quotations from credible sources | 27.2 | 24.7 |
Four results carry the paper. The evidence-adding trio of Cite Sources, Quotation Addition, and Statistics Addition delivered a relative improvement of 30 to 40% on position-adjusted word count and 15 to 30% on subjective impression. Stylistic edits, meaning Fluency Optimization and Easy-to-Understand, produced what the authors call a significant visibility boost of 15 to 30%.
Keyword stuffing lost ground against doing nothing, on one metric. It scored 17.7 against a 19.3 baseline on position-adjusted word count, and the paper files it under non-performing methods. On subjective impression it edged slightly above baseline at 20.2. The direction that matters is the word-count one, because that is the metric measuring how much of the answer you occupy.
The authoritative-tone result needs the same care. The paper's text reports no significant improvement and concludes that generative engines are "already somewhat robust to such changes". Its own table still places Authoritative among the high-performing methods, at 22.9 on subjective impression. Read together: a persuasive tone can make an answer's framing look better without winning your source more of the answer.
Combining methods helped, modestly. The best pair, Fluency Optimization with Statistics Addition, beat any single strategy by more than 5.5%, on a 200-example subset the authors ran because of cost. That footnote is worth carrying: the combination result rests on a smaller sample than the headline table.
Which of Those Methods Hold Up Outside the Paper
The authors tested their own methods once outside the lab, on Perplexity.ai. Quotation Addition again led on position-adjusted word count with a 22% improvement over baseline. Cite Sources and Statistics Addition showed improvements of up to 9% and 37% across the two metrics. Keyword stuffing scored 10% worse than baseline on position-adjusted word count, repeating the in-lab result on a deployed engine.
We found no published straight replication of GEO-bench. What exists instead is two peer-reviewed 2026 papers that rebuilt the experiment on harder ground, and they complicate the original rather than confirming it.
What the 2026 Follow-Ups Found
SAGEO Arena, published at KDD 2026, makes the sharpest criticism. Sunghwan Kim and colleagues at Yonsei and Konkuk universities argue, in an evaluation environment built for this gap, that existing benchmarks operate "on pre-determined candidate documents", abstracting away the retrieval and reranking stages of a real search pipeline. They built an environment that keeps those stages in. Running the optimization end to end, they report that existing approaches remain largely impractical, often degrading visibility in retrieval and reranking. Structural information helped, and effective optimization had to be tailored to each stage separately.
What Gets Cited, published at SIGIR 2026, went at citation choice directly. A team from Sprinklr ran 252,000 paired trials across six models in a controlled two-document experiment, changing one of eighteen content factors at a time and recording which of two sources the model cited first. Topical relevance and list position were the biggest drivers. Explicit price information and a recent timestamp helped consistently, trust cues added less, and formatting-only edits had little impact. The authors work for a vendor in this market, which is worth knowing, and the venue is peer-reviewed.
The third strand is supportive rather than critical. Lüttgenau, Colic, and Ramirez trained a model on 1,905 travel content pairs rewritten with citations, statistical evidence, and improved fluency, then measured the output with Llama-3.3-70B. They report 15.63% on absolute word count and 30.96% on position-adjusted word count, using the original paper's own metric. It is a preprint, covers one vertical, uses synthetic training pairs, and tests a single open-weight model.
So the 2024 findings are contested rather than confirmed or overturned. Under the original's conditions, with candidate sources already retrieved, evidence-dense edits win more of the answer. Put the same edits through retrieval and reranking and the advantage can reverse.
The Obvious Objection, and Why It Is Fair
A 2024 paper cannot describe engines that ship changes monthly. The authors say so themselves in their limitations section: methods may need to adapt over time as generative engines evolve, and query behavior changes too. The 2026 work suggests they were right to hedge. Any list of generative engine optimization best practices drawn from this paper alone is a hypothesis with two years on it, and one that a peer-reviewed follow-up has already pushed back on.
Two further limits are worth carrying into any vendor conversation. The authors’ own engine ran on GPT-3.5 and fetched exactly five sources per query. That is a fixed lab setup, not a description of how today’s assistants build every answer. And the subjective impression scores came from G-Eval, meaning a model grading a model, which is reasonable at benchmark scale and is not a human deciding a brand looked credible.
GEO as On-Page Work Against GEO as Source Building
Here is the detail almost nobody repeats. Every source in the experiment was already a top-five Google result for its query. The engine fetched five candidates, and the researchers then edited one of them. Nothing in the study describes how a page becomes one of those five. The whole apparatus measures what happens after retrieval has already happened.
The paper's own rank breakdown makes the point sharper. Relative visibility change by the source's position in the search results:
| Method | Rank 1 | Rank 2 | Rank 3 | Rank 4 | Rank 5 |
|---|---|---|---|---|---|
| Cite Sources | -30.3% | 2.5% | 20.4% | 15.5% | 115.1% |
| Quotation Addition | -22.9% | -7.0% | 3.5% | 25.1% | 99.7% |
| Statistics Addition | -20.6% | -3.9% | 8.1% | 10.0% | 97.9% |
| Authoritative | -6.0% | 4.1% | -0.6% | 12.6% | 6.1% |
| Fluency Optimization | -2.0% | 5.2% | 3.6% | -4.4% | 2.2% |
The same edit helps or hurts depending on where you already stand. Cite Sources lifted a fifth-placed source by 115.1% and cut a first-placed source by 30.3%. So GEO, as the paper defines it, is a redistribution mechanism among sources that already made the shortlist. It is not an entry ticket.
This is no longer only our reading. The KDD 2026 team built their benchmark because the earlier work ran "on pre-determined candidate documents", and once retrieval and reranking were put back in, the same optimization often degraded visibility. The SIGIR 2026 trials point the same way, finding list position among the two strongest drivers of which source gets cited first.
That is why the honest answer to how to optimize for AI search has two halves, and the paper covers only one. On-page edits change how much of the answer you occupy once you are in the evidence pool. Getting into that pool is a separate fight involving crawlability, relevance, retrieval, reranking, and the off-site signals that move an answer, none of which the study tested. Why on-site work hits a ceiling explains where the wider source footprint enters. Klarivo’s mature programs run roughly one part on-site to four parts off-site. That is Klarivo’s operating model. The research is silent on the four.
Common GEO Advice That Does Not Survive Scrutiny
Google published its own position, and it is unusually direct. The mythbusting section of Google's guide to optimizing for generative ai features, last updated 10 July 2026, lists five things site owners can ignore:
- LLMs.txt and other special markup: you do not need machine-readable files, ai text files, markup, or Markdown, because Google Search does not use them.
- Chunking content: there is no requirement to break pages into small pieces, and no ideal page length.
- Rewriting content just for AI systems: the systems understand synonyms and intent, so capturing every phrasing variant is unnecessary.
- Seeking inauthentic mentions: manufactured coverage is less useful than it looks, and spam systems are part of how the ai features are grounded.
- Overfocusing on structured data: it is not required for generative AI search, and no schema type buys you visibility.
Read the scope before you read the list. Google is describing Google Search, including AI Overviews and AI Mode. It is not describing ChatGPT, Claude, Perplexity, or Grok, which retrieve from different indexes under different rules. Treating the mythbusting list as a universal law about every engine overcorrects in the opposite direction.
Two further claims fail on the research rather than on Google's guidance. Keyword density is the first. Much of what is sold as SEO for AI search is keyword work in a new wrapper, and keyword stuffing was the only method to score below baseline on position-adjusted word count, in the lab and on Perplexity alike.
The unnamed metric is the second. When a page tells you to optimize for AI search on the strength of a 40% lift, that figure is a position-adjusted word count from a benchmark. On the top-ten check described earlier, the single page carrying the number did link the paper and did say "visibility" rather than traffic, and it still did not name the metric. Ask any vendor which metric their number belongs to. The answer separates the people who read the paper from the people who read a summary of it.
What a Realistic GEO Program Contains
Useful generative engine optimization strategies split cleanly into what the evidence supports, what it stays quiet about, and what it rules out.
What the evidence supports. Cite credible sources. Replace qualitative assertions with numbers where you have them. Quote credible third parties instead of paraphrasing them. Edit for fluency and readability. These were the four highest-scoring interventions, and they cost an editor rather than a platform.
What the evidence does not cover. Everything off your domain. That is where how a thread becomes citable matters, and where the budget question of what changes and what does not gets decided. If you are scoping an external program, what a GEO retainer covers is a separate question from what the research demonstrates.
What the evidence rules out. Keyword stuffing, an authoritative tone in place of actual evidence, and any promise of a ranking or traffic gain sourced to this paper. The authors did not measure rankings and said so.
Four Questions That Separate a Real Program From a Repackaged One
- "Which metric is your percentage?" A number with no metric attached is a number from a summary. Position-adjusted word count, mention rate, and citation rate measure different things.
- "What do you do off my domain?" The methods tested here are all on-page and editorial. A retainer priced above editorial rates has to earn the difference on the half the research never tested.
- "Which engines do you check, and how often?" Findings from a retrieval-augmented engine do not transfer cleanly to answers written from trained recall.
- "What will you show me that is not a score?" A score says a gap exists. A list of the domains an engine actually read says where to work next.
How Klarivo Approaches GEO
Klarivo works on the half the paper left out. The study showed that, within its setup, a source already in the retrieval set can win more of an answer through better evidence and cleaner prose. It never asked how a brand enters that set, and that entry problem is what a monitoring and publishing program exists to solve.
The bridge is a source list. Klarivo Monitor tracks five answer engines, ChatGPT, Claude, Perplexity, Gemini, and Grok, each independently, and reports the domains those engines drew on when they answered. A mention count tells you there is a gap. A domain list tells you which pages to go and get named on. The on-page edits the paper tested are cheap and you should make them. They cannot fix the absence of domains describing you in the first place.
The wrong fit is a team that already publishes off-site every week and only wants the numbers. Buy the tracking, keep your capacity, and skip the managed half.
Book a Klarivo discovery call if the question is which sources the engines read before they wrote you out. Fifteen minutes, a slot you choose, and instant confirmation, with no form to fill in first. The tracking side runs in Klarivo Monitor, and the program shape is set out in how the managed side runs.
Frequently Asked Questions About Generative Engine Optimization
Is GEO a real discipline or a rebranding of SEO?
Both descriptions have a claim. The metric is genuinely new, because a generated answer has no positions to occupy, and the paper's position-adjusted word count exists because no ranking measure fit. The tactics are less novel. Cite your sources, use numbers, write clearly. Any editor from 2010 would recognize all four winning methods. The honest reading is that GEO is a new measurement layer over largely familiar craft, with one genuine reversal: keyword stuffing cost visibility on the paper’s primary metric.
How much should GEO work cost, and can I do it in-house?
The four highest-scoring on-page methods are editorial work, so an in-house editor can do them at the cost of their time. Nothing in the research needs a platform or a license. What in-house teams struggle to fund is the off-site half, which is publishing and outreach across other people's domains, run weekly. Price a vendor on that half, and ask what they would do if the answer were simply “add citations to eleven pages.”
Does GEO work for a brand nobody writes about yet?
Less than the headline numbers imply. Every result in the paper was measured on sources already retrieved as top-five results, so the methods redistribute attention among candidates rather than creating candidates. A brand with no third-party footprint has a retrieval problem first. Klarivo’s sequence is coverage, then retrieval, then the on-page edits. The paper tested only the last step.
Did the paper test whether GEO changes Google rankings?
No, and it says so in its limitations. The authors state that owing to the black-box nature of search engine algorithms they did not evaluate how GEO methods affect search rankings. They add a reasoned guess that the changes are unlikely to affect rankings, since they alter text rather than domain, backlinks, or metadata. A guess in a limitations section is not a result, and it should never be quoted as one.
Which engine does the research actually generalize to?
The 2024 paper tested two directly: the authors' own two-step engine built on GPT-3.5, and Perplexity.ai. The SIGIR 2026 follow-up widened that to six models in a controlled two-document setup. All of them are retrieval-augmented, reading fetched sources at answer time. None of this work measures what a model writes from trained recall without retrieving anything, and that gap is rarely acknowledged by the pages citing it.
Start accelerating your Reddit presence
See how Klarivo can shift your visibility across AI, search and buyer communities.
Book a Demo


