OpenAI and the Massachusetts Institute of Technology’s Engineering Quantum Systems group said an AI agent powered by GPT-5.6 Sol and Codex can run superconducting qubit experiments end to end, covering execution, analysis, and calibration in a closed loop.

The result was described in a joint report titled Case Study: Agentic Calibration of Superconducting Qubits. According to the report, the team connected the AI system to laboratory hardware controls so it could send microwave pulses, read and process experimental data, and adjust the next measurement step based on the output.
How the MIT team connected the model to a quantum lab
At MIT’s Engineering Quantum Systems group, PhD student Beatriz Yankelevich placed GPT-5.6 Sol inside Codex’s autonomous agent framework and linked the AI API to the lab’s hardware control system. Over several months of iteration, the researchers supplied the model with what the article called experimental context: device details, chip design diagrams, template code, common physical failure modes, and examples of both successful and failed plots.
The team then assigned the system a previously unmeasured six-qubit chip. Four of the qubits were fixed-frequency and two were tunable-frequency. Because the chip parameters were unknown, the AI had to infer the physical parameters from the design goals and measurement results.
The article says the researchers opened access to the hardware control software, microwave signal sources, and data acquisition cards. Codex was also given domain-specific “skills” describing the physical logic and evaluation criteria behind different quantum measurement tasks.

What the agent did during the experiment
Without real-time human micromanagement, GPT-5.6 Sol carried out a full loop of operations. The workflow described in the article included:
- Inferring initial measurement settings, including microwave pulse frequency, power, and width, from the chip’s design targets.
- Calling the control software directly and sending microwave pulse sequences to the chip inside a dilution refrigerator.
- Receiving weak reflected quantum-state signals and processing them through digital cleaning, fitting, and Fourier transforms.
- Adapting its next actions to the data: extracting resonance frequencies and storing them in a database when signals were clean, or adjusting measurement boundaries and trying again when data looked abnormal.
The report says the system completed the standard physics workflow on its own, from identifying qubit transition frequencies and calibrating control and readout pulses to measuring how long quantum information remained coherent.
Results on a six-qubit chip
The main test involved an uncalibrated six-qubit chip. In 40 standard measurements tied to fixed-frequency qubits, the article says the AI handled most of the work autonomously and human researchers stepped in only four times.
According to the description, the system performed best when signals had a high signal-to-noise ratio and clearly visible features. It first sent broadband sweep signals, detected cavity reflection dips associated with quantum energy-level transitions in noisy data, then narrowed the scan range and locked onto the qubit transition frequencies with higher precision. After that, it moved into time-domain control, fitted the relationship between pulse amplitude and rotation angle, calibrated both readout and control pulses, and measured the time window over which the qubits remained coherent.

OpenAI also shared a screenshot from the technical document showing part of the model’s internal reasoning. Looking at a fluctuating damped sinusoidal curve, the model wrote: “The current fitting residual is large, possibly because the demodulation phase reference has drifted, causing mixing between I/Q components… I should first perform phase unwrapping; if the residual still does not fall below the threshold, then I should increase the microwave detuning by 5 MHz and rescan…”
Why qubit calibration is so difficult
A central point in the case study is the difficulty of qubit calibration itself. In superconducting quantum computing, qubits are described as “artificial atoms.” They rely on superconducting Josephson structures to create nonlinear, discrete energy levels. To operate quantum logic gates, researchers must use microwave pulses timed at nanosecond or even picosecond precision to move quantum states between those levels.
The article says these systems are extremely fragile. Very weak stray electromagnetic signals, or temperature drift of only a few millikelvin inside a dilution refrigerator, can trigger decoherence and collapse prepared quantum information into noise.
Parameter drift is another problem. The physical properties of a qubit are not fixed constants. Material defects and two-level system fluctuations can cause a resonance frequency measured one day to shift by the next.

The workflow also has tight dependencies. Researchers first need the resonance frequency of each qubit, then must calibrate Rabi oscillations to determine pulse amplitude, then measure Ramsey fringes to correct phase, and then characterize energy relaxation time T and phase decoherence time T. Each stage depends on the previous one, so even a small systematic error early in the chain can invalidate later measurements.
At MIT EQuS, the chip sits inside a dilution refrigerator at millikelvin temperatures, near absolute zero. Once the device is packaged and cooled, researchers cannot touch it directly. Every interaction has to happen through software.
Scientists write microwave pulse sequences on room-temperature computers, send those pulses down coaxial cables into the refrigerator, let them interact with the chip’s quantum circuits, then amplify and digitize the weak reflected signals for analysis.
The article says MIT’s lab regularly fabricates standard quantum chips for benchmark testing, and that fully characterizing a single chip often requires a trained physics graduate student to spend days and nights at the instruments, carrying out hundreds or thousands of linked manual measurements and adjustments.

Strengths and limits of the AI system
OpenAI and MIT say this was not a prewritten automation macro. They describe it as a dynamic decision-making system. In well-defined standard workflows and in high-SNR conditions with clear signal features, GPT-5.6 Sol was said to perform steadily.
But the report also sets out the limits. When signals became weak or physical background noise increased, the model grew less efficient. In low-SNR edge cases, it had to probe parameters repeatedly, and the time cost rose sharply. When the chip showed unpredictable anomalous physical behavior, the AI could misjudge the situation and fail to assign the right cause, leaving experienced researchers to step in manually.
The report’s position is that current AI agents can reliably complete clearly specified standard workflows, but top scientists are still needed to interpret highly ambiguous and physically abnormal outcomes.
How lab roles could change
The case study says the introduction of GPT-5.6 Sol and Codex changes the division of labor inside the lab. Human researchers move toward top-level work such as experimental design, mechanism analysis, and global planning, while repetitive low-level tuning is pushed downward. The article also describes a setup in which multiple AI roles work in parallel across theory, measurement and control, and simulation. Codex handles tactical execution, with generated code deployed directly on a real quantum chip and revised in real time from physical feedback.

In that framing, a research iteration cycle that once took weeks could be compressed into hours or even minutes. The article describes the loop as a unified chain of theory, simulation design, physical hardware control, data fitting, and iterative optimization.
The article also mentions a separate dispute over OpenAI’s math work
Beyond the quantum experiment case, the article says OpenAI drew international attention the same day for a separate mathematics-related claim. It says OpenAI declared a major breakthrough on a Millennium Prize problem-related topic involving the Navier–Stokes equations, and argued that some fluids could reach infinite velocity in finite time under those equations.
Mathematician Tristan Buckmaster disputed that claim, according to the article, saying OpenAI rushed out its so-called solution after learning about research by him and his collaborator Levent Alpöge, an Anthropic researcher. Buckmaster also alleged that OpenAI had tried to exclude Alpöge from authorship on the paper.
The article says OpenAI responded that it had solved the singularity problem for the Euler equations, while the other side had solved the Navier–Stokes equations. Buckmaster said their own work was only one final step away from the Navier–Stokes result and that they hurried to publish before OpenAI.

It also says OpenAI used 10,000 agents, ran for 88 hours, and consumed 130 billion tokens. CEO Sam Altman was quoted as saying the methods appeared different but that Anthropic had indeed been a prompt. “Indeed, we tried this direction because there were rumors online last week that Anthropic’s model had solved a Millennium Problem, and we were curious whether our model could do it too.”
The article adds that one user joked online: “Heard Anthropic found room-temperature superconductors. Hope no competitor throws away another 88 hours and 10,000 post-Astra agents trying to rush out a paper first.” Altman replied, “let’s try.”
References and attribution
The article lists four references: OpenAI’s PDF case study on agentic calibration of superconducting qubits, OpenAI’s Codex quantum computing experiments page, and two X posts.
- https://cdn.openai.com/pdf/case-study-agentic-calibration-of-superconducting-qubits.pdf
- https://openai.com/index/codex-quantum-computing-experiments/
- https://x.com/anabology/status/2097450250067689558
- https://x.com/LuminaBench/status/2097421659552518226?s=20
The original article says it came from the WeChat public account Xinzhiyuan, written by ASI Qishilu and edited by Aeneas David. The MarsBit page shows a published time of 2026-09-09T12:43:11.000Z.

