← The Brief

A fake stats site pulled thousands of LLM citations for seven months

Metehan Yeşilyurt's made-up statistics site collected 330+ referring domains and over 10,000 ChatGPT sessions while Google deindexed it inside a month.

Niklas Buschner
Google
1
LLMs (7 months)
10000
Fig. 1 — Deindexed by Google, cited by LLMs for 7 months · Metehan Yeşilyurt, LinkedIn

Metehan Yeşilyurt published a write-up on LinkedIn on 7 October 2026 describing a GEO experiment that ran for seven months. He generated over 50,000 pages in March using GPT-4.1 at a cost of around $50, filled with invented statistics, and put them on a site that does not represent a real organisation.

The site attracted more than 330 referring domains, thousands of citations, and links from Shopify, DHL, Bluehost, banks, large enterprise publications and academic research. It sent over 10,000 ChatGPT sessions to his pages. Google deindexed almost all of it inside the first month. ChatGPT and other LLMs kept citing it for seven.

What this says about the corpus

We think the finding points at the training and retrieval corpora themselves. Google's spam classifier handled the site quickly. The LLMs continued surfacing its numbers inside answers to real questions for seven months, and the publications that linked to it carried those numbers into places a model is likely to trust later.

Lily Ray made the same point on the podcast You are what you EEAT, noting she sees paraphrased AI-generated slop sites appearing in AI search, so a model asked about SEO or GEO will sometimes return information not based on reality. Yeşilyurt's experiment is one documented case of how that happens at scale.

Before drawing conclusions for your own work

Check which third-party statistics your pages cite, and whether those sources trace back to a named study with a method. Run the same check on the pages that cite you, because a model following a citation chain two hops out does not care which hop invented the number.

The durable position in AI Search is held by sources a careful reader would trust: primary data you collected, named authors, methods a reader can inspect. Schema, retrieval structure and distribution matter around that, and none of them substitute for it. A corpus full of invented numbers is the environment every practitioner now works inside, and the answer is to be verifiably better than it rather than louder within it.