Google Research, Google DeepMind, and the Massachusetts Institute of Technology have introduced CoDaS (AI Co-Data-Scientist), a multi-agent AI system built to extract scientific findings from raw wearable-device data. The system can handle hypothesis generation, statistical analysis, adversarial validation, literature-grounded reasoning, and full paper drafting on its own. According to the study, a research and writing task that would normally take human experts 37 person-days can be completed by CoDaS in 6 to 8 hours.
Wearable data from nearly 10,000 participants produced mental and metabolic signals
In testing, the researchers supplied CoDaS with a large wearable dataset covering nearly 10,000 participants. The data included sleep, activity, heart rate, and smartphone-use patterns. Without human prompting, the system identified several meaningful health-related features, with one of the most notable findings tied to mental health.
CoDaS found that excessive late-night browsing of social media or negative news was significantly positively correlated with depression severity, with reported statistics of ρ = 0.177, p < 0.001, and n = 7,497. The system also coined its own label for the behavior: “late-night doomscrolling.” Outside mental health, it also identified a negative correlation between the ratio of daily step count to resting heart rate and insulin resistance, a marker linked to metabolic disease.
Adversarial validation was built in to reject weak scientific claims
CoDaS includes an adversarial validation mechanism designed to limit scientific hallucinations and empty statistical claims. The paper describes one example in which the system proposed using “glucose squared” to predict insulin resistance. The relationship looked strong in purely statistical terms, but the validation layer determined that the feature was a scientifically meaningless tautology and rejected it.
That safeguard is central to the system’s design. It is also one of the clearest distinctions between CoDaS and a standard text-generation model.
Blind review gave CoDaS papers an 86% non-rejection rate
The researchers also tested output quality through blind expert review. Papers generated by CoDaS achieved an 86% non-rejection rate, meaning they were scored as accepted, minor revision, or major revision at that level. By comparison, papers produced by other baseline AI science agents saw rejection rates ranging from 85% to 100%.
The study presents a case for multi-agent AI turning consumer-grade wearable data into clinically relevant research insights. The main takeaway from the discussion around the release is straightforward: the question is no longer whether AI can write a paper, but whether it can carry out the research workflow that comes before it.

