Inherent, a stealthy startup founded by DeepMind alumni, is making bold claims: its AI "teammate" agent can replicate research tasks better than frontier models from OpenAI and Anthropic. The benchmark, posted on the company's website, shows Inherent's agent achieving higher success rates on a suite of research workflows, from literature reviews to experiment design.
But let's be clear about what this means. The evaluation—dubbed ResearchBench—measures how well an AI can follow multi-step research protocols, not whether it can make Nobel-level discoveries. Inherent's agent reportedly used a combination of retrieval, code execution, and iterative reasoning to outperform Claude and GPT models on a subset of tasks.
Still, the claim is notable. Inherent was founded by former Google DeepMind researchers who worked on AlphaFold and Gemini. They've built a "cognitive architecture" that decomposes complex research tasks into smaller, verifiable steps—something that excites the AI-for-science community.
But hold on. The benchmark is self-reported, and the company hasn't released the full evaluation set or code. Independent verification is absent. As with many AI claims, the devil is in the details—what constitutes "replicating research" is still a fuzzy line.
Inherent says it plans to open-source the benchmark and invite academics to audit the results. That's a promising step, and if they follow through, it could set a new standard for AI evaluation in scientific research.
As the AI community continues to chase benchmarks, Inherent's "teammate" stands out for its ambition. But until we see independent tests, temper your excitement.
Source: TechCrunch AI
Self-reported benchmarks are always a red flag for me. Would love to see a third-party evaluation of their agent's replication workflow.