The Princeton GEO Study: Methodology, Results and Critique
The Princeton GEO study (Aggarwal et al., ACM SIGKDD 2024) is the founding document of Generative Engine Optimization. The GEO-bench framework tested approximately 10,000 queries across nine datasets and proved that targeted content optimization can boost AI visibility by 22 to 41 percent.
Key Takeaways
- ✓GEO-bench tested 10,000 queries across 9 datasets and 7 domains
- ✓Statistics: +41% on Position-Adjusted Word Count, +37% on Subjective Impression
- ✓Attribution: +115.1% visibility for pages at position 5 (Equalizer Effect)
- ✓Citations: +28% on Subjective Impression
- ✓Keyword-Stuffing: -10% versus unoptimized baseline on Perplexity
- ✓Optimal Strategy: Fluency + Statistics (outperforms individual methods by 5.5%+)
- ✓Criticism: Zero-Sum setup with only 5 sources amplifies relative gains
The Princeton GEO study is the intellectual foundation of the entire GEO discipline. Whoever understands the study understands the mechanics behind AI visibility.
Methodology in detail
The GEO-bench framework by Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, and Deshpande simulated a two-stage generative engine pipeline. Google Search retrieved the top-5 sources per query, then GPT-3.5-turbo synthesized answers with citations. Nine different optimization methods were tested across approximately 10,000 queries over nine datasets and seven domains.
The nine tested methods were: citations, statistics, expert quotes, technical terminology, fluency optimization, keyword stuffing, authority claims, unique content, and simplicity.
The results hierarchy
Statisticsdelivered the strongest single improvement: 41 percent on the Position-Adjusted Word Count metric and 37 percent on Subjective Impression, validated on Perplexity.ai.
Citationsshowed the most dramatic equalizer effect. Pages at position 5 — normally with low visibility in AI answers — achieved 115.1 percent higher visibility. This confirms GEO as a leveling mechanism for pages outside top positions.
Quotesachieved 28 percent improvement on Subjective Impression.
Theoptimal strategycombined fluency optimization with statistics and outperformed every single method by more than 5.5 percent.
Keyword stuffing— the most iconic traditional SEO tactic — performed 10 percent worse than the unoptimized baseline on Perplexity.
Independent criticism
The study has three methodological limitations. The zero-sum setup with only five competing sources per query artificially amplifies relative gains. All optimizations were LLM-generated rather than human-applied. The 30-40 percent improvements are relative within a controlled environment — real effects in competitive niches are likely smaller.
Follow-up studies
Venkit et al. (FAccT 2025) found that 50 to 90 percent of citations in AI answers do not fully support the accompanying claims. Wu et al. (Nature Communications, April 2025) confirmed: even RAG-enabled GPT-4o leaves approximately 30 percent of individual statements without source support. Allouah et al. (Columbia/MIT, 2025) showed in an e-commerce sandbox that AI shopping agents exhibit choice homogeneity and small content changes can dramatically shift market share.
A large-scale study with 55,936 queries across six LLM search engines found that 37 percent of domains cited by LLMs are unique to AI and do not appear in traditional search results.
Data & Statistics
Bis zu 40 Prozent höhere Sichtbarkeit in generativen Antworten; 30 bis 40 Prozent relative Verbesserung auf Position-Adjusted Word Count für die Methoden Cite Sources, Quotation Addition und Statistics Addition
Aggarwal et al., GEO: Generative Engine Optimization, arXiv 2311.09735 (KDD 2024) (2024)GEO-bench: 10.000 Anfragen (8.000 Training / 1.000 Validierung / 1.000 Test), 9 Datenquellen, 25 Domains, 80 Prozent informationale / je 10 Prozent transaktionale und navigationale Anfragen
Aggarwal et al., GEO: Generative Engine Optimization, arXiv 2311.09735 (KDD 2024) (2024)KI-Sichtbarkeit korreliert mit Domain-Autorität (Pearson 0,65 / Spearman 0,57); Nofollow-Links zeigen nahezu denselben Zusammenhang wie Follow-Links; Basis 1.000 Domains
Semrush Blog, Do Backlinks Still Matter in AI Search? Insights from 1,000 Domains (2025)Die Präsenz einer AI Overview korreliert mit einer um 58 Prozent niedrigeren durchschnittlichen Klickrate (vs. 34,5 Prozent in der Untersuchung von April 2025); Basis 300.000 Keywords
Ahrefs Blog (2026)Ein KI-Suchbesucher ist auf Basis der Conversion Rate 4,4-mal so wertvoll wie ein klassischer organischer Suchbesucher
Semrush Blog, We Studied the Impact of AI Search on SEO Traffic (2025)KI-Nutzung österreichischer Unternehmen: 20,3 Prozent (2024) vs. 10,8 Prozent (2023); 41 Prozent der KI-nutzenden Unternehmen setzen KI zur Sprachgenerierung ein
Statistik Austria, IKT-Einsatz in Unternehmen 2024 (2024)FAQ
What is the Princeton GEO study?
What does the result of 40 percent visibility increase mean?
Which GEO methods work best according to the study?
What are the weaknesses and limitations of the Princeton GEO study?
Do the results of the Princeton study also apply to German-speaking regions?
How do I implement the study's findings in practice?
What does the study say about llms.txt?
Related Articles
How does your website perform?
Get a free, AI-powered SEO report of your website by email – technical SEO, on-page, keywords & competitors. No obligation.
Get a free SEO audit →