RMA: Context-Orchestrated Research Math Agents
Zelin Zhao, Bo Yuan, Yuchen Zhu, Jaemoo Choi, Yongxin Chen
Why It Matters
What makes this one worth your time
The framework offers a structured approach to solving complex mathematical problems, which could enhance automated reasoning capabilities in mathematical research.
RMA is a novel agentic system that outperforms existing models on research-level mathematical problems.
Summary
The paper introduces Research Math Agents (RMA), a framework designed to tackle research-level mathematical problems by decomposing the proof-solving process into specialized modules coordinated by agents. It demonstrates superior performance on the First Proof benchmark compared to existing models, solving eight out of ten problems.
Key contributions
- Development of the RMA framework for research-level mathematical problem solving.
- Demonstrated superior performance on the First Proof benchmark.
- Comprehensive ablation studies showing the importance of interaction between modules.
Notable insights
- The use of a multi-role, multi-round workflow with agents for iterative proof refinement is a clever methodology.
- The integration of structured reasoning modules and verifier-based feedback contributes significantly to performance gains.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2605.22875v2 Announce Type: replace Abstract: Long-horizon mathematical reasoning fails less often because a model cannot produce a valid next step than because an agent fails to maintain and expose the right semantic state across many iterations. Left unmanaged, this produces research-level proofs that are locally convincing yet globally incomplete: a key lemma unproved, an assumption unchecked, a citation unsupported, or a computational claim unverified. We present Research Math Agents (RMA), an agentic framework for long-horizon proof development built around a persistent, typed research store and an orchestrator that compiles operation-specific context from that store. The Research Context Orchestrator is the central state-management layer between the persistent research store and each locally scoped proof operation: it retrieves task-relevant artifacts, compiles them into a bounded context, invokes the appropriate operation, and writes the resulting proof edits, issue updates, literature notes, plans, or evaluations back to the store. This process is designed to keep proof revisions, unresolved issues, prior attempts, literature, and evaluations available across rounds while exposing only task-relevant state to each local operation. We evaluate RMA across complementary research-level settings using independent expert evaluation, blind mathematician review, LLM-based benchmark evaluation, and Lean 4 kernel verification. RMA achieves a 42.5% solve rate on the independently evaluated SOOHAK Challenge Hard set, obtains 8 of 10 correct solutions on First Proof B1 and 8 of 10 passing solutions on B2 under human-expert evaluation, and verifies 213 of 300 sampled Research Solved targets in Formal Conjectures with the Lean 4 kernel.