Skip to content
7.3Advanced9 min

The Princeton GEO Study: Methodology, Results and Critique

Lucas Blochberger··Updated 11 June 2026
Definition

The Princeton GEO study (Aggarwal et al., ACM SIGKDD 2024) is the founding document of Generative Engine Optimization. The GEO-bench framework tested approximately 10,000 queries across nine datasets and proved that targeted content optimization can boost AI visibility by 22 to 41 percent.

Key Takeaways

  • GEO-bench tested 10,000 queries across 9 datasets and 7 domains
  • Statistics: +41% on Position-Adjusted Word Count, +37% on Subjective Impression
  • Attribution: +115.1% visibility for pages at position 5 (Equalizer Effect)
  • Citations: +28% on Subjective Impression
  • Keyword-Stuffing: -10% versus unoptimized baseline on Perplexity
  • Optimal Strategy: Fluency + Statistics (outperforms individual methods by 5.5%+)
  • Criticism: Zero-Sum setup with only 5 sources amplifies relative gains

The Princeton GEO study is the intellectual foundation of the entire GEO discipline. Whoever understands the study understands the mechanics behind AI visibility.

Methodology in detail

The GEO-bench framework by Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, and Deshpande simulated a two-stage generative engine pipeline. Google Search retrieved the top-5 sources per query, then GPT-3.5-turbo synthesized answers with citations. Nine different optimization methods were tested across approximately 10,000 queries over nine datasets and seven domains.

The nine tested methods were: citations, statistics, expert quotes, technical terminology, fluency optimization, keyword stuffing, authority claims, unique content, and simplicity.

The results hierarchy

Statisticsdelivered the strongest single improvement: 41 percent on the Position-Adjusted Word Count metric and 37 percent on Subjective Impression, validated on Perplexity.ai.

Citationsshowed the most dramatic equalizer effect. Pages at position 5 — normally with low visibility in AI answers — achieved 115.1 percent higher visibility. This confirms GEO as a leveling mechanism for pages outside top positions.

Quotesachieved 28 percent improvement on Subjective Impression.

Theoptimal strategycombined fluency optimization with statistics and outperformed every single method by more than 5.5 percent.

Keyword stuffing— the most iconic traditional SEO tactic — performed 10 percent worse than the unoptimized baseline on Perplexity.

Independent criticism

The study has three methodological limitations. The zero-sum setup with only five competing sources per query artificially amplifies relative gains. All optimizations were LLM-generated rather than human-applied. The 30-40 percent improvements are relative within a controlled environment — real effects in competitive niches are likely smaller.

Follow-up studies

Venkit et al. (FAccT 2025) found that 50 to 90 percent of citations in AI answers do not fully support the accompanying claims. Wu et al. (Nature Communications, April 2025) confirmed: even RAG-enabled GPT-4o leaves approximately 30 percent of individual statements without source support. Allouah et al. (Columbia/MIT, 2025) showed in an e-commerce sandbox that AI shopping agents exhibit choice homogeneity and small content changes can dramatically shift market share.

A large-scale study with 55,936 queries across six LLM search engines found that 37 percent of domains cited by LLMs are unique to AI and do not appear in traditional search results.

Data & Statistics

Bis zu 40 Prozent höhere Sichtbarkeit in generativen Antworten; 30 bis 40 Prozent relative Verbesserung auf Position-Adjusted Word Count für die Methoden Cite Sources, Quotation Addition und Statistics Addition

Aggarwal et al., GEO: Generative Engine Optimization, arXiv 2311.09735 (KDD 2024) (2024)

GEO-bench: 10.000 Anfragen (8.000 Training / 1.000 Validierung / 1.000 Test), 9 Datenquellen, 25 Domains, 80 Prozent informationale / je 10 Prozent transaktionale und navigationale Anfragen

Aggarwal et al., GEO: Generative Engine Optimization, arXiv 2311.09735 (KDD 2024) (2024)

KI-Sichtbarkeit korreliert mit Domain-Autorität (Pearson 0,65 / Spearman 0,57); Nofollow-Links zeigen nahezu denselben Zusammenhang wie Follow-Links; Basis 1.000 Domains

Semrush Blog, Do Backlinks Still Matter in AI Search? Insights from 1,000 Domains (2025)

Die Präsenz einer AI Overview korreliert mit einer um 58 Prozent niedrigeren durchschnittlichen Klickrate (vs. 34,5 Prozent in der Untersuchung von April 2025); Basis 300.000 Keywords

Ahrefs Blog (2026)

Ein KI-Suchbesucher ist auf Basis der Conversion Rate 4,4-mal so wertvoll wie ein klassischer organischer Suchbesucher

Semrush Blog, We Studied the Impact of AI Search on SEO Traffic (2025)

KI-Nutzung österreichischer Unternehmen: 20,3 Prozent (2024) vs. 10,8 Prozent (2023); 41 Prozent der KI-nutzenden Unternehmen setzen KI zur Sprachgenerierung ein

Statistik Austria, IKT-Einsatz in Unternehmen 2024 (2024)

FAQ

What is the Princeton GEO study?
The Princeton GEO study is the paper "GEO: Generative Engine Optimization" by Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, and Deshpande, published in 2024 at the ACM SIGKDD conference and as preprint arXiv 2311.09735. It coined the term Generative Engine Optimization and was the first to demonstrate in a controlled experiment that content can be deliberately optimized for higher visibility in AI-generated answers.
What does the result of 40 percent visibility increase mean?
The figure is a maximum value, not an average. According to the study, the three strongest methods achieved a relative improvement of 30 to 40 percent on the Position-Adjusted Word Count metric compared to the unoptimized baseline. It represents a relative improvement under favorable conditions, with particularly low-ranked sources benefiting disproportionately. The figure is not a guarantee for every individual piece of content.
Which GEO methods work best according to the study?
The three authority methods had the strongest effect: Statistics Addition (incorporating relevant statistics), Quotation Addition (adding relevant quotes), and Cite Sources (adding source citations). The combination of linguistic optimization and statistics outperformed any single method. Keyword stuffing, on the other hand, was among the weakest approaches and could even reduce visibility.
What are the weaknesses and limitations of the Princeton GEO study?
The main limitations are the synthetic, partly opaque black-box evaluation, the model state from 2023/2024, and a zero-sum setup with only a few competing sources that amplifies relative gains. Therefore, the exact percentage values cannot be directly transferred to today's systems like Google AI Overviews, Google AI Mode, Perplexity, or Claude. However, the fundamental mechanism is considered well-established.
Do the results of the Princeton study also apply to German-speaking regions?
The mechanism is language-independent: the fact that generative systems prefer well-documented, well-structured, and authoritative content is not a property of the English language. In this respect, the qualitative findings are transferable to the DACH B2B market. The exact percentage values are not, as they are tied to the tested systems and English-language datasets. Market relevance is growing: according to Statistics Austria, AI usage by Austrian companies reached 20.3 percent in 2024.
How do I implement the study's findings in practice?
The findings can be translated into four levers: front-loading (placing the most important answer at the beginning of a section), answer islands (self-contained, complete passages per question), high density of facts and statistics with documented, preferably proprietary data, and structured data with clean heading hierarchies. Since effectiveness varies by subject area, methods should be tested per topic field rather than applying a blanket strategy.
What does the study say about llms.txt?
Nothing. The idea of llms.txt is not part of the Princeton study and is not covered by it. Independently, no major AI provider has confirmed using llms.txt for answer generation, and industry analyses find no reliable correlation with higher citation rates. llms.txt should therefore be classified as experimental and low-priority.

Related Articles

How does your website perform?

Get a free, AI-powered SEO report of your website by email – technical SEO, on-page, keywords & competitors. No obligation.

Get a free SEO audit