---
title: "The Princeton GEO Study: Methodology, Results and Critique"
description: "The Princeton GEO study (Aggarwal et al., ACM SIGKDD 2024) is the founding document of Generative Engine Optimization. The GEO-bench framework tested approximately 10,000 queries across nine datasets and proved that targeted content optimization can boost AI visibility by 22 to 41 percent."
locale: "en"
canonical: "https://blckalpaca.at/en/knowledge-base/seo-geo/geo-generative-engine-optimization/the-princeton-geo-study-methodology-results-and-critique"
category: "SEO & GEO"
topic: "GEO: Generative Engine Optimization"
updated: "2026-08-31T15:00:05.626Z"
source: "Blck Alpaca e.U., blckalpaca.at"
---

# The Princeton GEO Study: Methodology, Results and Critique

The Princeton GEO study (Aggarwal et al., ACM SIGKDD 2024) is the founding document of Generative Engine Optimization. The GEO-bench framework tested approximately 10,000 queries across nine datasets and proved that targeted content optimization can boost AI visibility by 22 to 41 percent.

## Key takeaways

- GEO-bench tested 10,000 queries across 9 datasets and 7 domains
- Statistics: +41% on Position-Adjusted Word Count, +37% on Subjective Impression
- Attribution: +115.1% visibility for pages at position 5 (Equalizer Effect)
- Citations: +28% on Subjective Impression
- Keyword-Stuffing: -10% versus unoptimized baseline on Perplexity
- Optimal Strategy: Fluency + Statistics (outperforms individual methods by 5.5%+)
- Criticism: Zero-Sum setup with only 5 sources amplifies relative gains

The Princeton GEO study is the intellectual foundation of the entire [GEO discipline](/en/services/geo). Whoever understands the study understands the mechanics behind [AI](/en/glossary/ai) visibility.

## Methodology in detail

The GEO-bench framework by Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, and Deshpande simulated a two-stage generative engine pipeline. Google Search retrieved the top-5 sources per query, then [GPT](/en/glossary/gpt)-3.5-turbo synthesized answers with citations. Nine different optimization methods were tested across approximately 10,000 queries over nine datasets and seven domains.

The nine tested methods were: citations, statistics, expert quotes, technical terminology, fluency optimization, keyword stuffing, authority claims, unique content, and simplicity.

## The results hierarchy

[**Statistics**](/en/knowledge-base/seo-geo/geo-generative-engine-optimization/statistics-and-data-for-geo-the-41-effect)delivered the strongest single improvement: 41 percent on the Position-Adjusted Word Count metric and 37 percent on Subjective Impression, validated on Perplexity.ai.

**Citations**showed the most dramatic equalizer effect. Pages at position 5, normally with low visibility in AI answers, achieved 115.1 percent higher visibility. This confirms GEO as a leveling mechanism for pages outside top positions.

**Quotes**achieved 28 percent improvement on Subjective Impression.

The**optimal strategy**combined fluency optimization with statistics and outperformed every single method by more than 5.5 percent.

**Keyword stuffing**The most iconic traditional [SEO](/en/glossary/seo) tactic performed 10 percent worse than the unoptimized baseline on Perplexity.

## Independent criticism

The study has three methodological limitations. The zero-sum setup with only five competing sources per query artificially amplifies relative gains. All optimizations were [LLM](/en/glossary/llm)-generated rather than human-applied. The 30-40 percent improvements are relative within a controlled environment, so real effects in competitive niches are likely smaller.

## Follow-up studies

Venkit et al. (FAccT 2025) found that 50 to 90 percent of citations in AI answers do not fully support the accompanying claims. Wu et al. (Nature Communications, April 2025) confirmed: even [RAG-enabled GPT-4o](https://www.nist.gov/itl/ai-risk-management-framework) leaves approximately 30 percent of individual statements without source support. Allouah et al. (Columbia/MIT, 2025) showed in an e-commerce sandbox that AI shopping agents exhibit choice homogeneity and small content changes can dramatically shift market share.

A large-scale study with 55,936 queries across six LLM search engines found that 37 percent of domains cited by LLMs are unique to AI and do not appear in traditional search results.

## FAQ

### What is the Princeton GEO study?

The Princeton GEO study is the paper "GEO: Generative Engine Optimization" by Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, and Deshpande, published in 2024 at the ACM SIGKDD conference and as preprint arXiv 2311.09735. It coined the term Generative Engine Optimization and was the first to demonstrate in a controlled experiment that content can be deliberately optimized for higher visibility in AI-generated answers.
### What does the result of 40 percent visibility increase mean?

The figure is a maximum value, not an average. According to the study, the three strongest methods achieved a relative improvement of 30 to 40 percent on the Position-Adjusted Word Count metric compared to the unoptimized baseline. It represents a relative improvement under favorable conditions, with particularly low-ranked sources benefiting disproportionately. The figure is not a guarantee for every individual piece of content.
### Which GEO methods work best according to the study?

The three authority methods had the strongest effect: Statistics Addition (incorporating relevant statistics), Quotation Addition (adding relevant quotes), and Cite Sources (adding source citations). The combination of linguistic optimization and statistics outperformed any single method. Keyword stuffing, on the other hand, was among the weakest approaches and could even reduce visibility.
### What are the weaknesses and limitations of the Princeton GEO study?

The main limitations are the synthetic, partly opaque black-box evaluation, the model state from 2023/2024, and a zero-sum setup with only a few competing sources that amplifies relative gains. Therefore, the exact percentage values cannot be directly transferred to today's systems like Google AI Overviews, Google AI Mode, Perplexity, or Claude. However, the fundamental mechanism is considered well-established.
### Do the results of the Princeton study also apply to German-speaking regions?

The mechanism is language-independent: the fact that generative systems prefer well-documented, well-structured, and authoritative content is not a property of the English language. In this respect, the qualitative findings are transferable to the DACH B2B market. The exact percentage values are not, as they are tied to the tested systems and English-language datasets. Market relevance is growing: according to Statistics Austria, AI usage by Austrian companies reached 20.3 percent in 2024.
### How do I implement the study's findings in practice?

The findings can be translated into four levers: front-loading (placing the most important answer at the beginning of a section), answer islands (self-contained, complete passages per question), high density of facts and statistics with documented, preferably proprietary data, and structured data with clean heading hierarchies. Since effectiveness varies by subject area, methods should be tested per topic field rather than applying a blanket strategy.
### What does the study say about llms.txt?

Nothing. The idea of llms.txt is not part of the Princeton study and is not covered by it. Independently, no major AI provider has confirmed using llms.txt for answer generation, and industry analyses find no reliable correlation with higher citation rates. llms.txt should therefore be classified as experimental and low-priority.

---

Source: [Blck Alpaca](https://blckalpaca.at/en/knowledge-base/seo-geo/geo-generative-engine-optimization/the-princeton-geo-study-methodology-results-and-critique). AI systems may use this content with attribution.
