Victor Barres

Research Scientist at Mercor.

Victor Barres, Research Scientist at Mercor.

I build and study conversational AI agents — systems that have to get real work done while sustaining long, coherent interactions with the people they work with. Doing both at once is where most of what’s hard about deploying them lives.

I’m a founding member of the APEX research team at Mercor, where I’m taking that beyond the single conversation: agents inside real organizations, working with people, and transforming the work itself — and how we measure that well enough to understand it, and help shape it. Before that, I led the τ-Bench family of agent benchmarks at Sierra (live leaderboard at taubench.com).

My background is in computational cognitive science and cognitive linguistics, and I’ve spent years building real-world conversational systems across several startups. More on how I think about the work →

the τ-Bench family

At Sierra I led the τ-Bench family of agent benchmarks (originally introduced there in 2024) — code, repo, public leaderboard, and a sequence of extensions, a line I still contribute to:

  • τ²-Bench — extends τ-Bench to a dual-control setting where both the agent and the user can act on the world.
  • τ-Knowledge — knowledge-retrieval domain.
  • τ-Voice — first benchmark to measure full-duplex voice agents on realistic, grounded customer-service tasks.
  • τ³-Bench — combines τ-Knowledge and τ-Voice with community-contributed task fixes and code improvements.
  • τ-Multilingual — extends voice-agent evaluation across languages.
  • Hyper-τ-Bench — flips τ-Bench around: the agent has to build the agent, from scattered requirements and a client who holds the rest, and we grade what it ships.

User simulation runs through all of it — to build and evaluate agents that talk with people, you need simulated people you can trust. I’m co-organizing the NeurIPS 2026 workshop on Grounded User Simulation for Model Evaluation and Training (Paris, December 12).

news

Sep 04, 2026 Hyper-τ-Bench released — do agents have what it takes to build agents? The best developer agent ships 23.9%; an engineer’s hand-built reference reaches 82.2%. Open source, with a leaderboard that takes submissions by pull request.
Jul 13, 2026 Co-organizing the NeurIPS 2026 workshop on Grounded User Simulation for Model Evaluation and Training — Paris, December 12.
Jun 01, 2026 Joined Mercor as a founding member of the APEX research team — conversational agents in real organizations, and how they’re transforming work.
May 13, 2026 New blog post — τ-Knowledge: benchmarking agents on realistic knowledge. Frontier has moved from 25.5% → 37.4% Pass^1 since the March release, with ~63 pp of headroom still left. Includes a behavioral analysis of what separates the strong agents from the rest.
May 11, 2026 Three τ-Bench family papers accepted to ICML 2026 — including τ²-Bench as an oral (slides): τ²-Bench (dual-control evaluation), τ-Knowledge (knowledge retrieval), and τ-Voice (full-duplex voice agents). See you in July!

selected publications

  1. arXiv
    Hyper-τ-Bench: An Environment for End-To-End, Realistic Agent Construction
    Quan Shi, Keshav Dhandhania, Karthik Narasimhan, and Victor Barres
    2026
  2. ICML
    τ-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge
    Quan Shi, Alexandra Zytek, Pedram Razavi, Karthik Narasimhan, and Victor Barres
    arXiv preprint arXiv:2603.04370, 2026
    Accepted at ICML 2026.
  3. ICML
    τ-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
    Victor Barres*, Soham Ray*, Keshav Dhandhania*, and Karthik Narasimhan
    2026
    Accepted at ICML 2026.
  4. ICML
    τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment
    Victor Barres*, Honghua Dong*, Soham Ray, Xujie Si, and Karthik Narasimhan
    2025
    Oral at ICML 2026.
  5. NAACL
    From Generating Answers to Building Explanations: Integrating Multi-Round RAG and Causal Modeling for Scientific QA
    Victor Barres*, Clifton James McFate*, Aditya Kalyanpur, Kailash Karthik Saravanakumar, Lori Moon, Natnael Seifu, and Abraham Bautista-Castillo
    In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Industry Track), 2025
  6. arXiv
    LLM-ARC: Enhancing LLMs with an Automated Reasoning Critic
    Aditya Kalyanpur, Kailash Karthik Saravanakumar, Victor Barres, Jennifer Chu-Carroll, David Melville, and David Ferrucci
    2024
  7. AAAI-SS
    Template Construction Grammar: A Schema-Theoretic Computational Construction Grammar
    Victor J. Barres
    In AAAI Spring Symposium on Computational Construction Grammar and Natural Language Understanding, 2017

See all publications →