
A16z Leads $40M Vals AI Series A After Finance-Agent Benchmark Caps at 52% Accuracy
Vals says its Finance Agent v2 test shows frontier models still fail about half of real analyst workflows, and it wants to be the market’s independent evaluator.
Vals AI raised a $40 million Series A led by Andreessen Horowitz at a $400 million valuation after its Finance Agent v2 benchmark put the best frontier model at 52% accuracy on multi-step finance analyst work. The pitch is that “AI is ready” narratives are now gated less by model releases and more by whether anyone can measure real workflow correctness fast enough to keep up.
A16z Backs Vals as Finance-Agent Benchmarks Put Frontier Models at 52% Accuracy
Vals AI closed a $40 million Series A at a $400 million valuation on Aug. 13, 2026, led by Andreessen Horowitz (a16z). Returning investors 8VC, Pear VC, and Bloomberg Beta participated again, with HRT Ventures and Next Ladder Ventures joining as new backers.
The round is being sold as infrastructure, not a model bet. Vals’ Finance Agent v2 benchmark recorded a top score of 52% accuracy in May 2026, meaning the best available system failed to correctly complete roughly half of the multi-step tasks a professional finance analyst would be expected to handle.
Mechanically, Finance Agent v2 is not a multiple-choice exam. It tests whether an agent can run a workflow end-to-end, including retrieving supporting data, synthesizing it without hallucinating figures, and chaining reasoning across steps. The tasks Vals lists include building relative-value models across peer companies, synthesizing sector catalysts, and generating investment recommendations grounded in actual document data.
Vals paired the funding with traction metrics: eightfold revenue growth over 2025, customer count doubling, and the team tripling over the past six months. It also said its evaluations have been cited in model cards, the official evaluation disclosures labs publish, from OpenAI, Anthropic, Google, Meta, and xAI.
From Leaderboards to Workflow Correctness: Why Vals Thinks Public Benchmarks Mislead
Vals’ core claim is that public leaderboards are increasingly measuring the wrong thing, or measuring it after the market has already learned how to game it. The failure modes it points to are benchmark saturation, where top models become statistically indistinguishable on a test, and benchmark contamination, where test questions leak into training data and inflate scores.
The company’s argument leans on a February 2026 ETH Zurich and Stanford preprint that analyzed 60 widely used large language model benchmarks and found 29 of 60 were saturated, with top-vs-second-best gaps within measurement noise. The same preprint is cited for a less comfortable conclusion for evaluators: private test sets do not systematically prevent saturation once a benchmark’s distributional characteristics are widely known.
Vals’ product response is lifecycle management. It treats benchmarks as “perishable” and retires them when they stop differentiating models. In May 2026, it replaced its CorpFin benchmark with an Excel test after CorpFin stopped providing enough separation between frontier models.
This is also where Vals draws a line against preference-vote leaderboards. Arena (formerly LMArena) collects millions of human preference votes to answer, “Which model do users prefer in chat?” Vals is explicitly trying to answer a different question: “which model completes investment research correctly at professional standards?” Those are different architectures, and they fail differently when you try to operationalize them inside a finance workflow.
Alongside the raise, Vals launched three products that widen its scope beyond finance-and-coding point tests: Vals Smith (custom coding benchmarks built from enterprise GitHub repositories), a frontier-risk benchmark suite (cybersecurity, mental health, AI safety), and an expanded Vals Index 2.0 with broader economic coverage. Vals Smith’s pitch is the gap between testing “can this model write Python?” and “can this model write our Python?” for firms with proprietary codebases.
Signals Traders Can Track: Benchmark Turnover, Product Adoption, and Who the 52% Model Was
The most market-relevant missing detail is basic: Vals does not name which specific frontier model scored 52% on Finance Agent v2 in May 2026. Until that mapping is public, the datapoint is more useful as a ceiling on the category than as a clean read-through to any one lab’s release cadence.
The next signal is whether that ceiling moves with subsequent model launches, or whether it stays sticky because the bottleneck is workflow reliability rather than raw capability. If Vals can keep turning around results “within hours” of getting model access, the benchmark becomes a near-real-time check on release hype.
On the product side, Vals Smith adoption is the cleanest proxy for whether enterprises are budgeting for measurement as a line item. Custom benchmarks built from private GitHub repos imply a buyer is past demo-stage and into procurement discipline.
Finally, watch Vals Index 2.0 coverage changes and benchmark retirements or replacements similar to the CorpFin-to-Excel swap. Fast turnover is a tell that tests are saturating quickly, and that “scorekeeper” credibility depends on continuously rebuilding the exam.
My Read: The Trade Isn’t ‘AI Is Smart’—It’s ‘Measurement Becomes the Bottleneck’
The threshold that matters here is not whether a model can top a public leaderboard, it is whether it can clear professional workflows with error rates low enough to survive production. A 52% accuracy ceiling on Finance Agent v2 is a blunt constraint on near-term “autonomous finance agent” narratives, because it implies the best system is still wrong often enough to require a human control loop.
If Vals keeps getting cited in model cards while also selling paid evaluations to the same ecosystem, the real test is whether its incentives stay aligned with buyers rather than vendors. If that independence holds and benchmark turnover stays high, measurement becomes the gating layer that decides which models get deployed, and that is the practical point where the category stops being narrative-driven and starts being procurement-driven.