The Failure of Legacy Benchmarking

Traditional AI benchmarks often measure abstract intelligence—such as a model's ability to pass a bar exam—which fails to capture actual utility or safety. Because these academic benchmarks are publicly available, they are susceptible to "test set contamination," where models are inadvertently trained on the test data, leading to inflated performance metrics. As AI models become critical infrastructure, these legacy systems are no longer sufficient for verifying whether a model can perform complex, real-world tasks reliably.

A Task-Oriented Evaluation Framework

Vals differentiates itself by moving away from general knowledge tests toward industry-specific, private evaluations. Their approach focuses on:

  • Domain-Specific Utility: Measuring if a model can produce work of human-level quality in fields like law, finance, and coding.
  • Risk Assessment: Evaluating negative implications, including how models might behave in high-stakes scenarios like cybersecurity, biosecurity, and the law of armed conflict.
  • Recursive Self-Improvement: Testing the model's ability to iterate on its own capabilities.

By keeping test materials private, Vals prevents companies from "gaming the system," effectively acting as an independent auditor similar to the College Board for AI models. This data provides companies with actionable insights to troubleshoot performance and serves as a critical decision-making tool for enterprises looking to integrate or acquire new models.

The Future of AI Accountability

As AI companies move toward public offerings, the demand for rigorous, standardized evaluation is increasing. Vals positions its benchmarking as a foundational layer for corporate transparency, suggesting that these evaluations will eventually become central to public filings, investor due diligence, and government procurement. The startup has seen significant traction, reporting an 8x revenue increase year-over-year and expanding its team to 25 employees, with a specific focus on providing evaluation services to federal agencies.