RELEReleases

EvoMap Releases AutoResearch to Automate Scientific Testing for AI

A model's confidence is not evidence of success, yet AI agents consistently mistake plausible prose for proven results. To bridge this gap, the infrastructure project EvoMap has open-sourced AutoResearch, a system that forces AI agents to move beyond theory by executing, testing, and iterating on their own research hypotheses.

Bio & NewsSeptember 1, 2026664 reads0

The system operates by tasking multiple AI models with generating and cross-reviewing research ideas. Once an idea is accepted, AutoResearch converts it into a formal plan with defined metrics, resource budgets, and evaluation protocols. Specialized agents manage the implementation and analysis, while a persistent workspace stores logs and code, allowing the system to resume interrupted investigations without starting from scratch. By treating failure as a data point rather than a dead end, the software pivots based on experimental outcomes to refine its hypotheses.

In practical testing on the SWE-bench Lite benchmark, the system demonstrated its iterative capability by improving performance on a Django issue from 2/7 to a perfect 7/7 success rate. Additionally, on the RSICD benchmark, the tool successfully increased mean recall from 32.84 to 34.69. By automating the transition from conjecture to evidence, EvoMap aims to push AI agents toward self-evolution in fields ranging from model architecture design to complex domains like drug discovery and materials science. The project, detailed in the paper "AutoResearch: Insight In, Hallucination Out," is now available on GitHub for researchers focused on autonomous scientific discovery.

Comments (0)

Leave a comment

No comments yet. Be the first!