Peec AI publishes ChatGPT's internal retrieval engine list from Yeşilyurt leak
A GEO research team has released the full inventory of search engines ChatGPT calls behind the scenes, including dozens of Labrador variants for Reddit, news, finance, legal and medical queries.
David Konitzny of Peec AI's GEO Research Team published the full list on 17 September, drawn from a retrieval leak first surfaced by Metehan Yeşilyurt. The team has been analysing what happens between a prompt and the sources ChatGPT cites.
The data covers internal engines, external providers, and dozens of Labrador variants. Named domains include Reddit, news, finance, legal, medical, local search and images. Konitzny credits Tomek Rudzki, Yeşilyurt, Jan Ehrlinspiel and Malte Landwehr on the work.
Some engines in the list have no confirmed purpose yet. The team is still testing which engines trigger for which query types and how often each one runs. They have asked other practitioners to help identify the unknowns and reproduce triggers.
What the list is, and what it is not
This is an inventory, not a ranking model. It tells you which retrievers exist inside ChatGPT's stack. It does not tell you how any one of them scores or selects a source once it fires.
We think the useful move for a practitioner is to treat the list as a map of surfaces to test against. Pick the engines that match your category, Reddit for community-led queries, the finance or legal variants where relevant, and run prompts that should trigger them.
Log which engine returns your domain and which does not. That gives you a per-engine read on where you are reachable, which is more useful than a single citation count across all query types.
A retrieval engine still picks the best source it can reach for the query it was built for. The page work is the same. The list tells you which doors are in the building.