Victor Barres
Research Scientist at Mercor.
I build and study conversational AI agents — systems that have to get real work done while sustaining long, coherent interactions with the people they work with. Doing both at once is where most of what’s hard about deploying them lives.
I’m a founding member of the APEX research team at Mercor, where I’m taking that beyond the single conversation: agents inside real organizations, working with people, and transforming the work itself — and how we measure that well enough to understand it, and help shape it. Before that, I led the τ-Bench family of agent benchmarks at Sierra (live leaderboard at taubench.com).
My background is in computational cognitive science and cognitive linguistics, and I’ve spent years building real-world conversational systems across several startups. More on how I think about the work →
the τ-Bench family
At Sierra I led the τ-Bench family of agent benchmarks (originally introduced there in 2024) — code, repo, public leaderboard, and a sequence of extensions, a line I still contribute to:
- τ²-Bench — extends τ-Bench to a dual-control setting where both the agent and the user can act on the world.
- τ-Knowledge — knowledge-retrieval domain.
- τ-Voice — first benchmark to measure full-duplex voice agents on realistic, grounded customer-service tasks.
- τ³-Bench — combines τ-Knowledge and τ-Voice with community-contributed task fixes and code improvements.
- τ-Multilingual — extends voice-agent evaluation across languages.
- Hyper-τ-Bench — flips τ-Bench around: the agent has to build the agent, from scattered requirements and a client who holds the rest, and we grade what it ships.
User simulation runs through all of it — to build and evaluate agents that talk with people, you need simulated people you can trust. I’m co-organizing the NeurIPS 2026 workshop on Grounded User Simulation for Model Evaluation and Training (Paris, December 12).
news
| Sep 04, 2026 | Hyper-τ-Bench released — do agents have what it takes to build agents? The best developer agent ships 23.9%; an engineer’s hand-built reference reaches 82.2%. Open source, with a leaderboard that takes submissions by pull request. |
|---|---|
| Jul 13, 2026 | Co-organizing the NeurIPS 2026 workshop on Grounded User Simulation for Model Evaluation and Training — Paris, December 12. |
| Jun 01, 2026 | Joined Mercor as a founding member of the APEX research team — conversational agents in real organizations, and how they’re transforming work. |
| May 13, 2026 | New blog post — τ-Knowledge: benchmarking agents on realistic knowledge. Frontier has moved from 25.5% → 37.4% Pass^1 since the March release, with ~63 pp of headroom still left. Includes a behavioral analysis of what separates the strong agents from the rest. |
| May 11, 2026 | Three τ-Bench family papers accepted to ICML 2026 — including τ²-Bench as an oral (slides): τ²-Bench (dual-control evaluation), τ-Knowledge (knowledge retrieval), and τ-Voice (full-duplex voice agents). See you in July! |
selected publications
- NAACLFrom Generating Answers to Building Explanations: Integrating Multi-Round RAG and Causal Modeling for Scientific QAIn Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Industry Track), 2025