What AI Reads Before It Recommends: Inside 7,534 Perplexity Citations
~6 min readAITwo Perplexity model variants were asked to recommend the five best tools in each of 380 B2B software categories. Their number-one picks matched in 79% of categories, and their top-five lists overlapped by 78% on average. That consistency is the interesting part: AI software recommendations are not noisy coin flips. They are a stable funnel β and a stable funnel can be engineered.
A dataset released today by Trellner Research (Manufactured sources behind AI recommendations, CC BY 4.0, collected 2 September 2026) makes the funnel measurable. This article walks through the raw numbers, adds an independent classification of the 2,055 cited domains, and points out a gap that should worry European software vendors in particular.
The experiment
The methodology is pleasantly verifiable. Trellner put 380 buyer-intent software categories β "project management software", "password managers", "CRM software" β to two web-grounded Perplexity models (perplexity/sonar and perplexity/sonar-pro) via OpenRouter. One prompt per category per model, 760 calls, zero errors. Each answer returned a ranked top five of vendors plus the grounding URLs the model reported retrieving. Every cited domain was then matched against the Tranco top-1M list and the Wayback Machine.
That leaves 3,800 recommendation slots, 7,534 citation rows, and roughly 9.9 grounding URLs read per answer. All numbers below come from that corpus; the domain-level classification is my own analysis of the same CSV.
AI reads the long tail β and sometimes nothing at all
If AI answer engines only trusted the established web, every citation would land on a top-10,000 domain. They do not:
| Metric | Value |
|---|---|
| Citations to domains not in the Tranco top 1M | 23.4% |
| Cited domains not in the top 1M at all | 751 of 2,055 (36.5%) |
| Median Tranco rank of ranked citations | ~71,600 |
| Share of citations going to top-10 domains | 17.3% |
Read that fourth row again. Over 82% of everything Perplexity reads before recommending software comes from outside the web's ten most-visited sites. The grounding layer is a long tail β and the long tail is where manufacturing costs are lowest. Of the unranked domains, 98.1% do exist in the Wayback Machine, so these are not dead links; they are live, purpose-built content properties.
Three kinds of source, one dominant pattern
Classifying all 2,055 cited domains by hand-tuned heuristics (multi-category footprint plus listicle URL patterns) yields three groups: