xMAD-2 is now available to research partners — read the model card

xMAD.ai

Frontier AI agent research

Agents that reason
over long horizons.

xMAD.ai is a research company advancing large language model technology. We study how models plan, use tools, and stay reliable across thousands of steps — and we publish what we learn.

Founded
2023
Researchers
140+
Publications
62
Open weights
4 models
Compute
8 regions

Research collaborations & compute partners

  • Meridian Labs
  • Northgate Institute
  • Helio Compute
  • Cambridge ALIGN
  • Vector Foundry
  • Ostro University

Research agenda

Four problems we work on.

Our agenda is narrow on purpose. Every project either extends how far an agent can act without supervision, or makes its behaviour easier to measure and trust.

  1. 01

    Long-horizon reasoning

    Keeping an agent coherent across tens of thousands of tokens of state: hierarchical planning, memory that survives context rollover, and recovery from its own mistakes.

    14 papers

  2. 02

    Tool use & embodied control

    Grounding a model in real interfaces — browsers, terminals, APIs, robots — and training it to know when a tool is the wrong one.

    11 papers

  3. 03

    Evaluation & interpretability

    Benchmarks that resist saturation, sparse autoencoders over agent trajectories, and methods for auditing a decision after the fact.

    19 papers

  4. 04

    Safe deployment

    Capability thresholds, sandboxed execution, and monitoring that catches a drifting agent before it causes harm rather than after.

    18 papers

Publications

Recent work.

All 62 publications
  • New Preprint

    Sierra: an agent that recovers from its own failures

    A. Nováková, R. Okafor, L. Tanaka, J. Meyer, and 6 others

    We train a recovery policy on twenty million annotated failure trajectories. On a suite of unattended data-engineering tasks, Sierra completes 71.4% end-to-end, against 43.2% for a comparable agent without recovery training.

  • NeurIPS 2026

    Tessera: measuring what an evaluation set no longer tests

    M. Ferreira, D. Okonkwo, S. Bhatt, Y. Lindqvist

    Contamination and saturation are usually reported after the fact. We give a cheap online estimator for the marginal information a benchmark still carries about a model family, and show it flags seven widely used suites as exhausted.

  • ICML 2026

    Harbour: sparse features of agent trajectories

    K. Adeyemi, P. Rasmussen, N. Iyer, T. Volkov

    Applying dictionary learning to full agent trajectories rather than single activations yields features that transfer between tasks. We use them to predict, with 0.81 AUC, whether a rollout will end in a specification violation.

Models

Weights we release.

Four models are available under a permissive licence. Frontier checkpoints are shared with research partners under a capability agreement.

Released models and their evaluation results
Model Parameters Context AgentBench-2 Licence
xMAD-2 Mini 3B 128k 48.1 Apache 2.0
xMAD-2 Base 9B 256k 61.7 Apache 2.0
xMAD-2 Pro 52B 512k 72.9 xMAD Research
xMAD-2 Pro Long 52B 2M 74.3 xMAD Research

AgentBench-2 scores are pass@1 on the held-out split, averaged over three seeds. Full methodology in the model card.

Safety

Written down, not implied.

Every model we train passes a documented review before release. The review covers capability thresholds, evaluation on autonomy-relevant tasks, and a red-team pass with external reviewers. Findings are summarised in the model card, including the ones that went badly.

  • Capability thresholds published before training begins
  • Independent red-teaming for every frontier checkpoint
  • Incident reports published within 30 days
  • Sandboxed execution by default in all hosted agents

Read our safety policy

“The interesting question is no longer what a model knows. It is what a model does on the two-hundredth step, when nobody is watching.”

Dr. Imogen Hartley Chief Scientist, xMAD.ai