Agentic AI for Scientific Reasoning in Autonomous Quantum Sensing Experiments
Takuya Isogawa, Ryotaro Okabe, Nutdech Phadetsuwannukun, Mingda Li, Paola Cappellaro
Why It Matters
What makes this one worth your time
This research is relevant for AI engineers and researchers interested in the intersection of AI and quantum experiments, showcasing how autonomous agents can enhance scientific reasoning and experiment control.
The paper introduces an AI-driven workflow for autonomous quantum sensing experiments with NV centers.
Summary
The paper presents an agentic AI workflow using a large language model for autonomous quantum sensing experiments with nitrogen-vacancy centers in diamond. It demonstrates an autonomous experiment workflow and introduces offline benchmarks to evaluate the agent's reasoning capabilities.
Key contributions
- Demonstration of an autonomous NV experiment workflow integrating project records, data analysis, and experiment control.
- Development of offline benchmarks for evaluating AI reasoning in quantum sensing experiments.
Notable insights
- The use of a large language model agent to autonomously select and calibrate NV centers for quantum sensing.
- The introduction of offline benchmarks to separately evaluate the reasoning capabilities of the AI agent.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2607.25145v1 Announce Type: cross Abstract: We implement an agentic AI workflow built around a large language model (LLM) agent for autonomous experiments with nitrogen-vacancy (NV) centers in diamond. NV centers are a widely used platform for quantum sensing, and the ability to control many measurements from a computer makes NV experiments a natural setting for autonomous workflows. We make two main contributions. First, we demonstrate an autonomous NV experiment workflow that combines persistent project records, quantitative calculation and data analysis tools, and deterministic experiment control. In one autonomous experiment, the agent selected a single NV center, calibrated its resonant frequency, measured \(T_2^\ast\) with Ramsey measurements, and added a Carr--Purcell--Meiboom--Gill (CPMG) measurement to check a weak feature that could be related to nearby \(^{13}\mathrm{C}\). Second, we introduce two offline benchmarks that evaluate the agent's reasoning separately from laboratory execution. We evaluated both benchmarks with GPT-5.4, GPT-5.5, and GPT-5.6 Sol. In the Ramsey checkpoint benchmark, greater reasoning effort generally improved recognition of a residual resonance calibration offset. By contrast, in the pulsed optically detected magnetic resonance (pODMR) data evaluation benchmark, pulse sequence information alone produced more false positive resonance judgments at higher reasoning effort. Requiring an expected signal calculation kept false positive rates low across all three models and reasoning settings. The results suggest a clear division of labor for autonomous experiments. The agent forms scientific hypotheses and uses quantitative tools to evaluate data, while deterministic code controls the hardware and enforces safety constraints.