Google DeepMind Co-Scientist Enters Real-World Science
Google DeepMind’s new paper extends Co-Scientist from hypothesis generation into execution-grounded research across materials science, biology, and computer science. Its autonomously discovered Agent_H architecture achieved leading length-adjusted scores on held-out HealthBench Hard and Professional while reducing potential clinical harm in blinded physician evaluation.
The important shift is from impressive demos to verifiable workflows—but this remains expensive, semi-autonomous research infrastructure rather than an unsupervised scientist.
- –Agent_H uses an eight-stage pipeline combining triage, decomposition, parallel candidate generation, tournament selection, clinical auditing, citation checks, and length optimization.
- –It required roughly 40–80 LLM calls per medical query, showing that inference-time scaling can improve capability without modifying model weights.
- –Physician evaluation found a meaningful safety improvement, but no significant quality advantage across eight other dimensions; automated judge scores should therefore be treated cautiously.
- –Real-world experiments included semi-automated CVD synthesis and lab-validated bacterial phenotype prediction, but human researchers still handled physical samples and oversight.
- –The system’s full source code is not public, limiting reproducibility and making cost, latency, and independent validation the next major hurdles.
DISCOVERED
1h ago
2026-08-28
PUBLISHED
1h ago
2026-08-28
RELEVANCE
AUTHOR
omarsar0