xMAD.ai

Careers

Six open roles.

We hire slowly and give people large problems. Every researcher publishes; every engineer owns a system in production.

Research Scientist, Long-Horizon Reasoning

Reasoning

San FranciscoFull-time

Design and run experiments on planning and recovery over very long trajectories. You will own a research direction end to end, from hypothesis to published paper.

Apply

Research Engineer, Evaluation Infrastructure

Evaluation

LondonFull-time

Build and maintain the harnesses behind AgentBench-2 and Terminal-2, including the tooling other labs use to reproduce our numbers.

Apply

Interpretability Researcher

Interpretability

San FranciscoFull-time

Extend the Harbour line of work: dictionary learning over agent trajectories, and methods for predicting failure before it happens.

Apply

Member of Technical Staff, Inference

Infrastructure

ZürichFull-time

Own the serving path for the xMAD-2 family, including 2M-token context and the sandboxed execution environment.

Apply

Research Engineer, Safety Tooling

Safety

LondonFull-time

Build the monitoring and drift-detection systems used in the pre-release review, and the internal tooling that makes red-teaming repeatable.

Apply

Research Intern, Agent Evaluation

Evaluation

San FranciscoInternship

A twelve-week project on benchmark decay estimation, with a paper target. Open to PhD students and exceptional undergraduates.

Apply

What we offer

  • Publication is expected, not negotiated. Travel to conferences is funded.
  • Compute allocation per researcher, decided by the research team rather than by committee.
  • Four weeks of paid leave, plus a company-wide shutdown in December.
  • Visa sponsorship and relocation support for all three offices.

We are an equal opportunity employer. We do not ask for a current salary, and we publish our bands internally.