Samuel Schmidgall1,*, Xiaokai Zhu2,*, Marian Shaw3, Lin Yang1, Valentin Liévin1, Jingyun Yang2, YuchenZhuang1, Tim Strother1, Alex Bijamov1, Min Woo Sun1, Anil Palepu4, Justin Chen1, David Steiner1, JacquelineShreibati1, Wei-Hung Weng1, Yilin Zhao2, Xingjian Hu2, Nicholas Zahn2, Sadhya Garg3, Julia Kirby3, YuxiangGan5, Jiaoli Li5, Divy Thakkar1, Shekoofeh Azizi1, David Racz1, Juraj Gottweis1, Vivek Natarajan1, ChenglinWu5, Tal Danino3, Keran Rong1, Haozhe Wang2, Benoit Schillings1, Yong Cheng1, Quoc V. Le1and Tao Tu11Google DeepMind,2Duke University,3Columbia University,4Google Research,5Texas A&M University,*Equal contribution We present an extension and comprehensive real-world validation of Co-Scientist, a Gemini-basedmulti-agent system designed to accelerate end-to-end scientific research across hypothesis generation,experimentation, and manuscript generation. While previous iterations of the system largely focusedonin silicohypothesis generation, this new specialized configuration transitions Co-Scientist into anexecution-grounded research partner capable of advancing closed-loop scientific workflows. We validatethese extended capabilities across materials science, biology, and computer science, spanning a spectrumof autonomy and producing novel scientific results with real-world significance. In materials science,Co-Scientist interfaced with a semi-automated chemical vapor deposition (CVD) reactor to design anovel non-hazardous precursor route for MXenes; microscopic and diffraction analyses indicate that theas-synthesized lamellar two-dimensional (2D) material shares key structural similarities with theTi3C2T𝑥MXene lattice, while further experiments are needed to confirm the atomic structure. Furthermore,for 2D transition metal dichalcogenides (TMDs), by leveraging Gemini 3 Deep Think for rapid, lab-in-the-loop execution, Co-Scientist tailored synthesis recipes to laboratory constraints in minutes,enabling successful, single-attempt growth of monolayerMoS2,MoSe2, andWS2semiconductors. Inbiology, Co-Scientist built a system to predict emergent swarming phenotypes of engineeredE. coliacross inducer (IPTG) concentration gradients from sparse imaging data, largely matching unpublishedwet-lab morphological measurements, suggesting a potential for reducing experimental screeningcycles. In computer science, the extended Co-Scientist autonomously designed an inference-time scalingarchitecture that outperformed six frontier models on HealthBench (Hard and Professional) whileachieving a significant reduction in potential clinical harm under blinded physician evaluation. Finally,a double-blind study of end-to-end generated papers with 30 domain experts across 450 independentreviews provides empirical evidence that Co-Scientist’s reliability modules reduce hallucination andplagiarism while improving research safety.Together, these results demonstrate further progresstoward closed-loop multi-agent scientific AI systems capable of iterative self-improvement to acceleratereal-world scientific discovery.arXiv:2608.26701v1 [cs.AI] 27 Aug 2026 1. Introduction Artificial intelligence (AI)-assisted scientific discoveries are increasingly transitioning fromin silicoexperiments to physical reality. Recent agentic AI systems have autonomously solved open mathe-matical conjectures (Feng et al., 2025a,c; OpenAI, 2026a,d), proposed wet-lab validated biomedicalhypotheses (Ghareeb et al., 2026; Gottweis et al., 2026), identified clinically actionable biomark-ers (Kim et al., 2026), and produced complete AI research manuscripts end-to-end (Lu et al., 2026a;Schmidgall et al., 2025). As one of the first demonstrations of a multi-agent system for scientific discovery, Co-Scientist (Got- tweis et al., 2026), built on Gemini, has acted as a collaborative research partner in prior works (Ali-abadi et al., 2026; Guan et al., 2025; Penadés et al., 2025; Toghani et al., 2026; Wang et al., 2026a),demonstrating the ability to assist human experts in formulating hypotheses and interpreting complexbiological data. These advances point toward an emerging paradigm where agentic AI systems operatenot just as passive tools, but as active research partners capable of formulating hypotheses, designingexperiments, interpreting research outcomes, and self-refining through experimental feedback. Yet,systems that have produced validated discoveries typically require substantial human oversight, withresearchers decomposing problems, verifying intermediate steps, and executing physical experiments. Closing the gap between what autonomous systems can ideate computationally and what they canvalidate physically remains the central barrier to scalable, AI-accelerated scientific discovery in the realworld. On one end of the spectrum, purelyin silicoresearch agents incorporate ideation, coding, andmanuscript writing into unified pipelines (Jansen et al., 2025; Lu et al., 2026a; Schmidgall and Moor,2025; Schmidgall et al., 2025). However, because these