Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking
AssessmentData ScienceComputer Science
THE AI ANGLE
Evaluating frontier models on complex domain-specific tasks and risksStartup Vals raised $40 million in Series A funding to address flaws in legacy AI benchmarking, arguing that older public tests allow companies to game metrics rather than reflect modern model capabilities. By keeping its testing materials private and focusing on complex, domain-specific tasks in fields like coding, law, and cybersecurity, Vals offers paid evaluations to verify real-world performance and risks. For faculty in computer science, data science, and assessment, this shift underscores how evaluation methodologies must evolve beyond abstract standardized tests to robust, task-oriented measures.
THE TEACHING ANGLE
Instructors can explore the tension between open-access benchmarks, which promote academic transparency but risk models 'teaching to the test', and proprietary, black-box evaluations that guard against cheating but limit external reproducibility.Read the original at techcrunch.com Generate teaching or study materials
More in Assessment
- Alibaba Open Sources OpenCodeReview for AI-Assisted Code ReviewInfoQ · September 20, 2026
- Anthropic picks Accenture for in-house AI safety evaluationPhys.org — Technology · September 20, 2026
- Google Agent Development Kit for Kotlin Reaches Feature Parity with Python, Supports On-Device AIInfoQ · September 20, 2026
- Will AI models achieve the ability to improve autonomously? Leading labs say the scenario is nearPhys.org — Technology · September 20, 2026
- Your AI agent failed. The model might not be the problem.The New Stack · September 20, 2026