OpenAI Codex Agent Calibrates Six-Qubit Chip at MIT
OpenAI reports that an AI agent, utilizing its Codex tool, successfully performed standard calibration measurements on a six-qubit superconducting chip at MIT. This development reduced the time researchers spent on routine experiments. The agent, powered by an OpenAI model called GPT-5.6 Sol, managed the sequence with minimal oversight when data was clear. However, human intervention was required when data became noisy or ambiguous, according to an OpenAI blog post. This represents a single case study on a familiar chip, not a fully deployed tool capable of independently diagnosing new hardware.
Agent’s Specific Actions
According to OpenAI, Beatriz Yankelevich, a graduate student in MIT’s Engineering Quantum Systems Group, connected Codex to the lab software controlling the chip. The agent then selected measurement settings, operated the hardware, analyzed the resulting data, and determined subsequent steps. It successfully identified the qubits’ transition frequencies, calibrated the microwave pulses for control and readout, and measured the quantum information’s coherence time.
OpenAI noted that characterizing one of the group’s standard chips can occupy a researcher for several days due to the hundreds or thousands of interdependent preliminary measurements required for each chip. The company stated that the MIT group now regularly employs agents for these routine measurements. It also mentioned that researchers might still identify optimal settings faster than current models, implying that the benefit lies in reduced supervision and not in increased raw speed.
Yankelevich elaborated on her setup in the post.
“I’ve built infrastructure to guide agents through several parts of my work – measurement, theory, and chip design – and now it’s really starting to pay off. (…) I can have multiple agents working on different problems at once, and I spend most of my time on higher-level work – interpreting results, devising experiments, planning next steps for the agents, reading, and writing.”
Limitations of the Agent
OpenAI acknowledged that the agent struggled with weak signals or data buried in noise. Quantum devices are prone to drift over time, and hardware defects or environmental disturbances can generate misleading readings that experienced researchers can identify visually. For unusual data, Yankelevich assigns agents narrower objectives while retaining control over scientific decisions.
This demonstration involved a six-qubit chip and workflows already understood by the MIT group. OpenAI admitted that it did not prove the agent’s ability to independently diagnose novel hardware issues or interpret an unexplored experiment. A six-qubit system is a small demonstrator, and its calibration is a laboratory step, not a useful computation. Questions remain regarding reliability, oversight, and the transferability of this approach to larger systems.