A Chinese artificial intelligence (AI) system, Qiushi Engine, has achieved the top ranking for autonomous scientific research, surpassing Anthropic’s Claude Code and other leading agents.
As of Tuesday, Qiushi Engine, developed by a team from Zhejiang University, leads the ResearchClawBench leaderboard, followed by Open Science Desktop in second place and Claude Code in third.
ResearchClawBench evaluates AI agents’ abilities to independently conduct research, comparing their outcomes to reference papers authored by humans to determine if they reach similar conclusions or exceed the original findings.
The benchmark was created by a team at the Shanghai Artificial Intelligence Laboratory to test whether AI agents can effectively perform the tasks they are designed for.
Qiushi Engine, launched last week, is a large language model-based agent intended for scientific research in real-world settings.
Developers of Qiushi Engine stated that, unlike some existing systems limited to specific tasks, it is capable of “end-to-end autonomous scientific discovery.”
On July 15, the development team reported that the system utilized multiple agents to adapt its strategies for complex research assignments.
A preprint paper submitted to arXiv in April outlined that Qiushi Engine achieved “the first demonstration of an AI agentic system autonomously identifying and experimentally validating a non-trivial, previously unreported physical mechanism.”
Yang Yihao, a professor at Zhejiang University and author of the study, stated that “Qiushi Engine is no longer just a research support aid, but is closer to a new type of independent research tool,” according to the People’s Daily.
AI systems based on large language models are becoming increasingly integrated into scientific instruments and used as research assistants.
The ultimate aim is for AI to generate ideas and carry out the full spectrum of research, analysis, and manuscript preparation.
The team behind Qiushi Engine notes that scientific discovery is one of the most challenging intellectual endeavors to automate comprehensively.
ResearchClawBench evaluated 40 tasks extracted from papers across ten disciplines, scoring agents on a 100-point scale, where a score of 50 indicates matching the conclusions of the original paper.
Scores below 50 signify that rediscovery was not achieved, while scores above 50 suggest the potential for new discoveries, as detailed in another preprint paper by ResearchClawBench developers in May.
The benchmark assesses agents in fields including astronomy, chemistry, earth science, energy science, information science, life science, materials science, mathematics, neuroscience, and physics.
As of Tuesday, all but one of the agents had overall scores in the 20s or lower, with Qiushi Engine, utilizing OpenAI’s GPT-5.5, achieving an overall score just above 30.
ResearchClawBench’s developers noted that even the most advanced agents are still far from achieving reliable rediscovery.




