BlogIndustry Analysis

Half of AI Unicorns Have Published Zero Scientific Papers. The Study That Proves It.

A Stanford study cataloged all 317 AI companies worth over $1 billion. More than half have never published a single paper. The industry claiming to revolutionize science isn't participating in it.

Chethan·July 30, 2026

More than half of AI companies worth over $1 billion have never published a single scientific paper. Not one. Not a preprint, not a conference paper, not a peer-reviewed study. Zero.

These are the same companies telling you they're revolutionizing drug discovery, redefining scientific research, and building the most important technology since electricity. And most of them have produced less scientific literature than a moderately productive grad student.

That's the headline finding from a new preprint posted on bioRxiv on July 16, and it's been making the rounds on Hacker News and in research circles for the last few days. The study comes from John Ioannidis — the Stanford metascientist who was the first to publicly call out Theranos for having no peer-reviewed data — and his team. They cataloged every AI unicorn (companies valued at over $1 billion) that has existed from 1998 to 2025. All 317 of them. Then they searched for any scientific publication where a company researcher played a leading authorship role.

The full dataset: 2,077 publications across all 317 companies. That's 1,389 peer-reviewed papers and 688 preprints — total. Not per year. Not per company. Total.

To put that in perspective: collectively, these 317 companies accounted for roughly one in every 1,000 AI papers published in 2025. The entire unicorn class of the most hyped industry on Earth produced 0.1% of the scientific literature.

The Paradox Ioannidis Caught

Ioannidis frames the contradiction plainly: "For a field that is supposedly reshaping science and is so advanced in terms of scientific potential, not having any scientific documentation seems like a very weird paradox. How can you judge that what they say is real, validated, and reproducible?"

He would know. In 2015, he was the first researcher to publicly scrutinize the lack of peer-reviewed studies behind Theranos — the blood-testing startup that turned out to be a $9 billion fraud built on technology that never worked. When a company makes world-changing claims and backs them with zero verifiable data, that's not a strategy. That's a pattern.

Now, I'm not saying every AI unicorn is Theranos. Most of these companies are building real products that demonstrably work — you can use them. But there's a meaningful gap between "the product works" and "we've advanced science in a way that others can verify, reproduce, and build upon." The former is engineering. The latter is what these companies claim to be doing.

Where the Citations Actually Go

The influence is even more concentrated than the publication count suggests. The top 5% of firms in the dataset accounted for more than 90% of all citations. OpenAI alone was responsible for nearly 40% of all citations across the entire unicorn dataset, followed by Megvii (the Chinese computer vision company) and Hugging Face.

And even at OpenAI — a company with roughly 4,500 employees — only eight researchers had authored five or more qualifying papers. Eight people. This isn't a research institution with a publication culture. It's a company where a handful of researchers happen to publish, and the rest don't.

The long tail is worse. Over 150 companies in the dataset — unicorns, remember, each worth over a billion dollars — have contributed absolutely nothing to the scientific record. No papers, no preprints, nothing where their own researchers are in an authorship role. You wouldn't know they exist if you searched Google Scholar.

Why They Stopped (And Why That's Rational)

Here's where I'd push back on the obvious "these companies are evil" take. The HN discussion around this study surfaced a genuinely interesting point from user sillysaurusx: publishing is most valuable to people who have no other way to get the attention of smart strangers. Once you can hire nearly anyone and everyone already returns your calls, the main remaining effect of publishing is telling your competitors which things worked.

This isn't hypothetical. Google published the transformer paper ("Attention Is All You Need") in 2017. It's one of the most important AI papers ever written. And Google got almost nothing for it commercially while everyone else — OpenAI, Anthropic, Meta, a thousand startups — built trillion-dollar businesses on top of it. As Mohamed Abdalla, an AI ethicist at the University of Alberta, puts it in the Science article: "It's not the company's job to advance science. The company's job is to advance money."

That's cynical, but it's also correct. Companies are rational actors operating under commercial incentives. Publishing your best work when competitors will immediately exploit it is a bad business strategy. The same thing happened in chemistry — researchers published freely until dyes became worth real money, and then the interesting work disappeared into corporate labs. AI is following the exact same pattern, just at internet speed.

But here's the irony that several commenters flagged: these companies' entire business model depends on training models on publicly available data — academic papers, Wikipedia, books, code repositories. They're consuming the scientific commons voraciously while contributing almost nothing back to it. They eat from the table but never set it.

The China Angle (Again)

The study found something that's going to make people uncomfortable: Chinese AI firms consistently publish more than their US counterparts. This isn't a small difference. It maps directly onto the open-versus-closed model debate that's been raging all year.

Leading US frontier labs — OpenAI, Anthropic — have increasingly kept their most capable models closed, releasing technical reports that read more like marketing documents than scientific papers. You get benchmark scores and carefully curated examples, not architecture details or training methodologies. Meanwhile, Chinese companies have been dumping model weights onto Hugging Face at a pace that's hard to keep up with. Moonshot AI, one of the companies in the study, released Kimi K3 — one of the strongest open models to date — with full weights publicly available.

Avijit Ghosh, an AI policy researcher at Hugging Face, makes a distinction that matters: the real question isn't whether companies publish in journals versus blogs. What matters is whether they release enough code, data, or model weights for others to independently verify and build on their work. By that standard, a Chinese company posting weights on Hugging Face is doing more for science than a US lab publishing a Nature paper with no reproducible code.

This is the uncomfortable truth for the American AI establishment. You can argue all day about safety and responsibility. But if your definition of "responsible AI" includes "nobody can check our work," you're not practicing science. You're practicing marketing.

The "Blogification" Problem

Several researchers in the Science article defend what Ghosh calls the "blogification" of research — the practice of announcing new models through blog posts and technical reports instead of peer-reviewed papers. The argument is that peer review takes months or years, and AI moves in weeks. By the time a paper clears review, it's obsolete.

There's some truth to this. Peer review in AI has always been a mess — the conference cycle is brutal, reviewers are overwhelmed, and good work gets rejected for stupid reasons. Blog posts and technical reports can communicate results faster and to a wider audience.

But here's the thing: blog posts aren't reviewed by anyone. There's no adversarial process. Nobody independently verifies the benchmark numbers. When a company says "our model beats GPT-5 on these eight benchmarks," you're trusting their internal eval team — the same team whose performance review depends on the model looking good. That's not a scientific process. It's a press release with LaTeX formatting.

The companies that publish real papers — where independent reviewers push back, where methodology gets interrogated, where reproducibility matters — are still doing something fundamentally different from the companies that ship a blog post and call it research. The study shows most AI unicorns are doing the latter, if they're doing anything at all.

What This Means for You

If you're building with AI, this study should change how you evaluate vendors and models. When a company claims a breakthrough, ask: where's the paper? Where's the reproducible evidence? If the answer is a blog post with cherry-picked examples and no methodology section, calibrate accordingly.

The open-source model ecosystem exists partly because of this dynamic. When DeepSeek publishes weights, when Qwen ships checkpoints, when Kimi drops a 2.8 trillion parameter model for free — they're not just being generous. They're participating in a scientific process that the closed labs have abandoned. You can download their models, run them yourself, verify the benchmarks, and build on them. That's how science is supposed to work.

The Ioannidis study is ultimately a mirror. It shows us an industry that talks constantly about transforming science while barely participating in it. An industry that consumes public knowledge at industrial scale while walling off its own discoveries. An industry where the loudest claims of world-changing capability come from companies that have produced fewer peer-reviewed papers than a single productive academic lab.

The companies building the future of intelligence don't seem very interested in documenting how it works. Make of that what you will.


If you want to actually verify AI claims yourself — run models locally, check benchmarks, and build with open-source AI instead of taking closed labs at their word — CopperRiver is a desktop AI assistant that runs open-source models directly on your Mac. No cloud dependency, no trust required.

#AI research#open source#industry analysis#Ioannidis

Try CopperRiver yourself

A desktop AI assistant that browses, codes, and automates. Plans from $9/mo.

Read next