The Failure of Public Benchmarks

As AI models advance, traditional public benchmarks have become saturated and unreliable. Rayan Krishnan, CEO of Vals, argues that these benchmarks are often 'gamed' by labs to show high performance on open-source rubrics while failing to deliver on private, real-world tasks. A notable example cited was the release of Llama 4, which showed impressive public scores but underperformed on Vals’ held-out, private benchmarks. This disconnect highlights a systemic issue: when benchmarks are public and static, they become targets for optimization rather than accurate measures of intelligence.

The Need for Independent Evaluation

Krishnan and the a16z team compare the emerging field of AI evaluation to established industries like credit rating agencies or auditing firms. Just as Enron demonstrated the dangers of conflicts of interest in auditing, the AI industry faces a similar risk if the same entities building models are also the ones defining the tests. Vals maintains a strict policy against selling training data to labs, ensuring their evaluations remain objective. The goal is to provide a 'shared language' for model performance that enterprises can use to justify ROI and labs can use to prove genuine advancement.

Evaluating Agentic Systems

Moving from simple text-in/text-out models to complex, agentic systems requires a shift in infrastructure. Modern evaluations must account for tasks that unfold over hours, days, or weeks. Unlike static datasets like ImageNet, these evaluations involve fewer samples but significantly more complex rubrics. Krishnan notes that the biggest bottleneck is making 'fuzzy' human workflows—such as the difference between a junior associate and a partner at a law firm—explicit and legible enough to be tested by an AI.

The Economics of Intelligence

As enterprises integrate AI, they face a 'misvaluing' of intelligence. Krishnan shares an anecdote about a Fortune 10 company that arbitrarily set a $100-per-day token budget for engineers, leading to bizarre productivity patterns where work only happens when rate limits reset. As token spend begins to rival employee salaries, businesses need rigorous evaluations to determine which models provide the highest ROI for specific tasks, rather than relying on arbitrary caps or broad, unverified claims of capability.

The 'Always a Higher Peak' Philosophy

Vals operates on the motto 'always a higher peak.' Because foundation models are constantly 'hill climbing' to reach new levels of performance, benchmarks must be perpetually reconstructed. This involves retiring saturated benchmarks and ensuring that tests reflect the current state of the world—such as updating legal research benchmarks to account for new case law. This continuous, iterative process is the only way to keep pace with the rapid evolution of frontier intelligence.