Inherent, a stealthy startup founded by DeepMind alumni, is making bold claims: its AI "teammate" agent can replicate research tasks better than frontier models from OpenAI and Anthropic. The benchmark, posted on the company's website, shows Inherent's agent achieving higher success rates on a suite of research workflows, from literature reviews to experiment design.

But let's be clear about what this means. The evaluation—dubbed ResearchBench—measures how well an AI can follow multi-step research protocols, not whether it can make Nobel-level discoveries. Inherent's agent reportedly used a combination of retrieval, code execution, and iterative reasoning to outperform Claude and GPT models on a subset of tasks.

Still, the claim is notable. Inherent was founded by former Google DeepMind researchers who worked on AlphaFold and Gemini. They've built a "cognitive architecture" that decomposes complex research tasks into smaller, verifiable steps—something that excites the AI-for-science community.

Why it matters: If Inherent's approach generalizes, it could shift the conversation from raw model intelligence to agentic workflows that reliably execute research tasks. That's the kind of leap that could accelerate scientific discovery—or flood the field with unverified "AI-assisted" papers.

But hold on. The benchmark is self-reported, and the company hasn't released the full evaluation set or code. Independent verification is absent. As with many AI claims, the devil is in the details—what constitutes "replicating research" is still a fuzzy line.

Caveat emptor: Inherent's results are impressive on paper, but without third-party evaluation, we should treat them as a strong signal, not a proven fact. The AI hype cycle has taught us to demand transparency.

Inherent says it plans to open-source the benchmark and invite academics to audit the results. That's a promising step, and if they follow through, it could set a new standard for AI evaluation in scientific research.

The takeaway: Inherent's claim is a reminder that the next frontier of AI isn't just bigger models—it's better agents. Whether they actually beat OpenAI and Anthropic remains to be seen, but the direction is right.

As the AI community continues to chase benchmarks, Inherent's "teammate" stands out for its ambition. But until we see independent tests, temper your excitement.

Source: TechCrunch AI