Microsoft has unveiled two new features for its Copilot Researcher — Critique and Council — seamlessly integrating OpenAI's GPT with Anthropic's Claude to elevate research capabilities. This marks a major leap forward in AI collaborative frameworks.
Critique: Sequential Collaboration for Factual Accuracy
The Critique mode employs a sequential collaboration model: GPT handles research planning and draft creation, while Claude ensures factual accuracy and citation quality. This division of labor makes the research process more rigorous, effectively reducing errors and misinformation in AI-generated content.
Council: Parallel Generation and Synthetic Output
The Council feature allows GPT and Claude to independently generate research reports, after which a third model synthesizes and optimizes their outputs. This parallel architecture captures the strengths of each model, producing more comprehensive analysis results.
Impressive DRACO Benchmark Results
On the authoritative DRACO benchmark, the Copilot Researcher with Critique achieved an impressive score of 57.4, outperforming the second-place model by nearly 14% and significantly surpassing the previous industry leader Claude Opus 4.6 (score 42.7). This result fully demonstrates the enormous potential of multi-model collaboration in complex research tasks.
Microsoft's update not only enhances Copilot's product competitiveness but also points to a new direction for the future development of AI research tools. As models like GPT and Claude continue to evolve, multi-agent collaborative systems are poised to become the core architecture of next-generation AI applications.

