Startup Vals is challenging traditional AI benchmarking by developing a more rigorous, task-focused approach that tests models on practical capabilities across industries. Backed by a $40 million Series A funding round led by Andreessen Horowitz, Vals is expanding rapidly to meet growing demand.
- Vals raised $40M Series A led by Andreessen Horowitz
- Focuses on industry-tailored AI benchmarking beyond academic tests
- Revenue has grown eightfold in the past year with expanding team
What happened
Vals, an AI benchmarking startup founded in 2024, recently announced a $40 million Series A funding round led by Andreessen Horowitz. This investment follows an earlier seed round from 8VC and Bloomberg Beta, underpinning Vals’ growth trajectory and increasing industry presence. The startup offers a novel approach to benchmarking by keeping specific testing materials confidential and emphasizing practical, domain-specific tasks over traditional academic-style evaluations.
The company, co-founded by Stanford graduate Rayan Krishnan, evaluates AI models not just on data recall but on how well they perform complex tasks in fields like law, finance, cybersecurity, and even biosecurity. Vals has rapidly expanded its workforce, tripling its team size to 25 employees within a year, and is planning to move into a larger office to accommodate continued growth.
Why it matters
Benchmarking AI models effectively is critical as artificial intelligence integrates more deeply into various sectors. Legacy benchmarks often fail to reflect the evolving capabilities of modern AI or the real impact these models have in practical, professional environments. Vals’ approach helps companies assess whether AI outputs meet industry standards and produce results comparable to human expertise, which is essential for trust and effective deployment.
Additionally, Vals incorporates evaluations of potential negative outcomes resulting from widespread AI use, addressing ethical and safety considerations that are increasingly important to regulators, businesses, and users. This shift towards robust, nuanced benchmarking could become influential in guiding AI development and adoption.
What to watch next
Vals plans continued expansion, both in staffing and product offerings. Future benchmarks will cover even more specialized areas, including recursive self-improvement of AI, mental health applications, and understanding complex ethical frameworks like the Geneva Convention in automated systems. Observers will be watching how these developments influence industry standards and whether Vals becomes the preferred resource for AI reliability assessments.
Furthermore, the startup’s growing revenue and increased engagement with AI developers signal that its benchmarking service might play a pivotal role in acquisition and investment decisions within the AI ecosystem. How competitors and traditional benchmarks respond to this challenge will also be a key factor shaping the AI evaluation landscape.