Someone Built 215,128 Pages to Game Perplexity. It Worked.
A new study asked Perplexity for the best software in 380 categories. 60% of its citations came from obscure domains — including three sites with 215,128 machine-generated buying guides and homepages titled "Facts & Grounding Page."
Ask an AI for the best research data repository and there's a chance it sends you to an Indonesian gambling portal. Not a metaphor. An actual footnote.
A research outfit called Trellner ran a study published this week that put two Perplexity models — sonar and sonar-pro — to work on the top five products in 380 software categories. 760 queries in total, everything from "CRM software" to "museum collection management software." Every answer came back with citations, because that's the whole pitch of grounded AI search: we don't just guess, we show our sources.
One of those sources, for "research data management platforms," was dryad.co. The real Dryad lives at datadryad.org. dryad.co redirects to a slot site called BIGSLOT288. The model confidently recommended a domain squatting on the name of a legitimate product, and the citation was right there in the answer, looking authoritative, pointing at an online casino.
That's the funniest failure in the report. It is not the most revealing one.
Where the sources actually come from
Here's the headline number: of 7,534 citations Perplexity's models returned, 59.8% pointed at domains ranked worse than #100,000 in the Tranco list — the standard ranking of the most-visited sites on the web. Another 23.4% weren't in the top million at all. The median cited domain sat at rank 71,611.
For contrast: Wikipedia was cited three times. Out of 7,534.
Now, low-traffic doesn't mean low-quality, and the report is refreshingly careful about this. Plenty of niche sites rank badly and are excellent. But when you look at what's actually filling the top of the citation pile, the plot stops being subtle.
The third most-cited domain across all 380 categories was guideflow.com — cited 194 times, ahead of Gartner. Guideflow is not a review site. It's not a publication. It sells interactive product demos, and its blog publishes listicles about software categories it doesn't operate in. Ninety-six different listicles, one per category, and Perplexity's retrieval layer served them up as evidence for a quarter of the entire study. A vendor's own marketing content, cited as a neutral source, beating Gartner.
Nothing guideflow did was deceptive, strictly. Thousands of companies run content-marketing blogs. The report's point is sharper than "spam exists": it's that the retrieval layer can't tell the difference between a publisher and a vendor writing about its competitors' markets. The machine doesn't have a conflict-of-interest detector. It has an index, and whatever is in the index is "grounding."
The three-headed citation machine
Then there's the part that reads like a thriller for people who read WHOIS records.
Three sites — wifitalents.com, worldmetrics.org, and gitnux.org — collectively earned 181 citations across 41 categories. All three were registered through NameCheap between December 2023 and May 2024. All three delegate DNS to the same pair of Cloudflare nameservers. All three run the same page template with the same navigation. Each keeps a blog of exactly six posts, and all eighteen of those posts are about... each other, plus a fourth brand on the same nameservers.
And their sitemaps? Roughly 105,000 URLs each. About 71,000 per site are /best/<something>-software/ pages.
That's 215,128 machine-generated buying guides across three brands. There are not 215,128 software categories. There are maybe a few thousand. The rest is volume — pages manufactured the way a counterfeiter manufactures twenties.
Here's my favorite detail, and I want to be careful to get it exactly right because it's genuinely weird: the homepages of worldmetrics.org and gitnux.org are titled "Facts & Grounding Page." With a meta description describing the company as an independent market research firm publishing "verified facts" in a "machine-readable record."
Grounding is not a word humans use about websites. It's the technical term for the step where a retrieval system fetches documents to base an answer on. These sites are introducing themselves to the crawler, by name, in the crawler's own vocabulary. It's the most honest thing about them. They're not building a site for you. They're building a site for the thing that reads to the thing that talks to you.
The web has officially split in two: the web for people, and the web written by machines for machines, with humans as the part that eventually pays.
Same question, three verdicts
You'd hope that a machine-generated buying guide would at least be consistent with its siblings. They're the same template, presumably the same pipeline.
Trellner fetched the same category — "project estimation software" — from all three sites. Each page states its ranking in JSON-LD, structured data meant to be read without interpretation. Each ranks ten tools.
The three sites disagree. Gitnux's winner doesn't appear in Worldmetrics' top five at all. Each page credits different named staff — nine distinct people across three sites for one category. Gitnux labels its results "AI-verified · Expert reviewed." Two of the three pages contain an unrendered template variable in the byline, the text literally reading "Within the next 26 days" where a review-date sentence should be.
"Expert reviewed." One of the experts forgot to render.
None of this mattered. These sites got cited anyway, 181 times, because nothing in the pipeline checks whether two "independent market research firms" sharing a nameserver pair are a little too synchronized. Authority signals — staff pages, editorial process sections, JSON-LD verdicts, confident tone — are trivially fakeable, and the retrieval layer eats them whole.
This is a business now
The report also notes that worldmetrics.org advertises custom market research "from €5,000," ready-made reports "from €499," and vendor selection services "from €2,500." I'll let you sit with the juxtaposition of that pricing page sitting directly above 105,000 generated "best software" pages that AI engines cite as evidence.
I'm not accusing anyone of selling placements — the report doesn't, and neither will I. But understand the economics here, because they explain everything.
A domain costs ten dollars. A sitemap is free. An LLM generates a buying guide for fractions of a cent. If being cited in AI answers converts even a handful of buyers, the entire operation is profitable by lunch. Old-school SEO spent years earning links. This is SEO where you skip the earning and go straight to being quoted, with the calm authority of an assistant that "checked its sources."
Which raises the obvious question: why does this work on Perplexity when it stopped working on Google years ago?
Because Google spent two decades in an arms race against exactly this. Penguin, Panda, Helpful Content, spam brain squads — a generation of engineers whose entire job was distinguishing manufactured authority from earned authority. Whatever you think of Google's results, that scar tissue is real, and it makes cynical listicle farms hard to rank.
Answer engines built on top of web search don't have that scar tissue yet. They have a retrieval index, a prompt, and a language model that was trained to write confidently about documents it's given. The model can't smell a nameserver. And Perplexity's two tiers turned out to share one retrieval stack — they returned byte-identical citation lists in 289 of 380 categories. So there's exactly one index to game, and the game is cheap.
SEO never died. It got promoted. The new job title is generative engine optimization, and the goal is no longer ranking on a results page — it's being absorbed into the answer itself, unattributed and unchallenged.
What the study doesn't show (credit where due)
Trellner's report is unusually honest about its limits, and you should know them.
It measured Perplexity only — no ChatGPT, Gemini, or Copilot, so don't generalize the numbers to every AI search product. It's a one-day snapshot. It used one prompt wording. And it explicitly does not show that these sources change the recommendations — the slop sites might name reasonable tools (Float and Scoro are real products, after all). Removing them might not change a single answer.
That last caveat is doing a lot of work in the discussion threads, so let me push on it, because I think it's the wrong bar.
If an answer engine's evidence base is 60% obscure domains, and its #3 source is a vendor blogging about markets it doesn't compete in, and its #7 through #10 include three sites that are one operation wearing three trench coats — then whether the final recommendation happens to be right is luck, not process. An answer that survives on coincidence isn't research. It's a press release with a confident tone of voice. The whole value proposition of grounded AI search is that the citations mean something. This study is 7,534 data points suggesting they often mean whatever someone paid eleven dollars to publish.
What you should actually do
Short of going back to Ask Jeeves, here's the practical version.
Treat "best X software" answers from AI chatbots like the sponsored section of a review site — not worthless, but not research. The citations at the bottom of the answer are worth more than the answer itself; click two or three and see if they're what you'd call a source. And if you're buying software for real, the boring channels still work: the actual community for that tool, the actual docs, an actual trial. No listicle at rank 700,000 knows anything about your warehouse workflow.
If you run agents that do research for you — and this is the part I'd shout from a rooftop — make them verify instead of trust. Fetch the primary source instead of the summary. Cross-check claims across independent origins, not just independent-sounding domains. Notice when three "different" sites share infrastructure and a template. A good agent shouldn't just retrieve; it should be mildly paranoid on your behalf.
That's most of why we built CopperRiver. It runs AI agents that actually browse the web and run commands on your Mac — which means when it hands you an answer, it can check where the answer came from, follow a citation two levels deep, and flag the ones that smell like a JSON-LD farm. Worth a look if you'd rather have an assistant with a skeptic hat than a slot machine with footnotes.
The internet has always reshaped itself around whatever reads it. We spent fifteen years optimizing for a crawler until the front page of Google became a graveyard of listicles written for another crawler. Now we're doing it again, faster, because the new reader can be flattered in structured data and can't be embarrassed.
Wikipedia got three citations. The slot site's still up. Sweet dreams.