Running Model Comparison · Autolab Research
Autolab (autolab.ai) publishes a continuously updated public benchmark that runs frontier models inside Claude Code on fixed autoresearch tasks, reporting mean, variance and API cost, with all raw trajectories made public. The standings are refreshed with every notable model release.
Aug 2, 2026
Autolab
·
Company Website