Company
A research company, not a product company.
xMAD.ai was founded in 2023 to work on one question: what happens when language models stop answering and start acting. Everything else follows from that.
About
We are 140 researchers, engineers and operators across three offices. About two thirds of the company works on research; the rest builds the infrastructure that makes the research reproducible — evaluation harnesses, training pipelines, and the sandboxes our agents run in.
The company is funded by research partnerships rather than by a consumer product. That choice is deliberate: it keeps our incentives aligned with publishing results instead of shipping features on a quarterly cadence.
How we work
- Research results are published, including negative results.
- Evaluation harnesses ship with the paper, not after it.
- No capability claim without a released evaluation protocol.
- Every frontier training run has a named safety reviewer outside the team.
Safety policy
Our policy has three parts: thresholds defined before training begins, an independent review before any release, and public incident reports within thirty days of a confirmed issue.
Thresholds cover autonomy-relevant capabilities — long-horizon self-direction, irreversible actions, and deception under evaluation. Where a threshold is approached but not crossed, we say so in the model card and publish the measurements.
Questions about the policy are welcome from researchers and regulators alike.
Team
The leadership team, with research staff listed in full on request.
Dr. Imogen Hartley
Chief Scientist
Previously led evaluation research at a national AI institute. Works on capability thresholds and measurement.
Rafael Okafor
Head of Research
Long-horizon planning and recovery. Author of the Sierra line of work.
Dr. Lena Tanaka
Head of Interpretability
Sparse methods for sequential state. Previously at a university ML group in Zürich.
Priya Bhatt
Head of Engineering
Training and inference infrastructure. Built the sandbox our hosted agents run in.
Dr. Tomas Volkov
Safety Lead
Independent reviewer for frontier releases. Background in formal methods.
Marguerite Ferreira
Head of Evaluation
Benchmark design and decay estimation. Maintains AgentBench-2.
Dr. Yusuf Adeyemi
Research Scientist
Trajectory interpretability and violation prediction.
Sofia Lindqvist
Chief Operating Officer
Research partnerships, compute procurement and publishing.
Press kit
Logos, the wordmark, team photography and our boilerplate are available for editorial use. For interviews, contact the communications desk and we will respond within two business days.