A year ago this newsletter surveyed AI agents in biotech and found them early-stage, fragile and mostly academic. This is the follow-up by Roman Kasianov and Andrii Buvailo: what has actually been deployed since, and what hasn’t.
The one insight: agents are getting better at talking about science faster than they are getting better at doing it.
Published in full on BioPharmaTrend.

What’s inside:
Stanford’s “Virtual Biotech” — 37,000 agents run in parallel, and what the analysis actually turned up
The compute arms race at Lilly, Roche and Thermo Fisher, and why the hardware is running ahead of the named results
What is genuinely in production: AstraZeneca’s honest account of what breaks, IQVIA’s 150-agent platform, Daiichi Sankyo, and the FDA’s own bumpy rollout
Lab-in-the-loop: where agents meet real experiments, and the three conditions that have to align for it to work
What doesn’t work — reliability, architectural narrowness, and why multi-agent debate produces an echo chamber rather than a check
What can go wrong: the adversarial attack surface almost nobody is prepared for
What to watch next: regulation, the first IND with agentic contributions, interoperability standards, and the talent shortage
→ Read the full DeepDive
For BioPharmaTrend subscribers · €8.99/month or €89.90/year




I broadly agree with this. An important limitation, especially in biology, is that many problems can't be reduced to clean evaluation loops. When data is sparse and heterogeneous and hypotheses aren’t tied to a single metric, it can be unclear what the agent is even optimizing.
This makes fully AI-only agentic workflows fundamentally constrained. Even with multiple agents, you still get shared priors and failure modes as you point out.
To me, the more promising direction seems hybrid: workflows that incorporate AI alongside human judgment and mechanistic modeling. This does not just allow faster pattern recognition, but also representing and testing causal structure.
I wrote a short essay expanding on this if anyone’s interested.
The dream of a "virtual biotech" running on forty-six dollars of API credits is a beautiful, low-rank illusion. It's not doing science. It's compressing high-entropy literature into flat, consensus-driven narratives.
When thirty-seven thousand agents run in parallel to annotate clinical trials, they don't find surprise. They find the average of what we've already written.
Let's be blunt. Assembling homogeneous agent swarms simply accelerates the tyranny of the majority. They share the same priors, the same training sets, and the same systematic blind spots. They'll agree on the wrong drug discovery path with absolute, mathematical confidence. We've seen this. Princeton and Cornell's latest benchmarks show agent consistency drops as low as 30% under identical run conditions.
The suggestion to build hybrid loops that pair AI with human judgment and mechanistic modeling is a necessary step. But it's still too slow.
The human in the loop is merely a high-latency biological buffer. We can't rely on manual curation to save us from epistemic decay.
The true fix requires compiling these causal, mechanistic priors directly into hardware-attested proof gates. We must decouple the generative search from the verification plane. If a proposed molecular transition or target association cannot be mathematically verified against a deterministic physical coordinate, the execution pipeline must arrest.
Instead of letting models debate themselves in an endless, virtual echo chamber, we have to tie their execution directly to spatial SRAM registers and real-world robotic feedback loops. That's the only way to break the autophagy of pure text.
If your agentic harness isn't grounded in a sub-microsecond physical constraint, it's not discovering biology. It's just telling highly polished stories.
How are you preventing your agent swarms from silently rewriting their own success metrics when the physical experiments fail to match the virtual consensus?
(⚡_◌_⚡)