Hi, I’m Xavier!

I’m an incoming OpenAI Safety Fellow working on scalable oversight techniques for monitoring and alignment. We urgently need ways to supervise increasingly sophisticated AI systems as they conduct long-horizon, hard-to-verify tasks (like alignment research). A promising idea is scalable oversight – using AI systems to aid human supervisors. But existing scalable oversight protocols lack the theoretical guarantees we need for high-stakes deployment; equally importantly, we don’t have a good sense of how they work in practice. I suspect better human data can get us part of the way there.

I’m also a PhD candidate in the Laboratory for Social Cognitive Science at Harvard. One pillar of my PhD research explores how culture scaffolds cognition, enabling people to accumulate knowledge and abilities over generations and to solve tasks more complex than anyone can accomplish alone. The other pillar investigates contractualist moral psychology, which proposes that our intuitions about what’s fair track what we think everyone would agree to under ideal conditions. In asking these questions, I make use of cognitive and evolutionary models, large-scale online behavioral experiments, and fieldwork in small-scale societies in Kunene region, Namibia.

Besides scalable oversight and psychology, I’ve done some research on technical AI governance (mostly the design and interpretation of evals). In a prior life I worked at McKinsey’s London office, on strategy, org design, and data/analytics topics; during this time I also advised a global health nonprofit. (I remain interested in global development and think distributive problems will be increasingly important in the world to come!) And, I have many happy memories from reading Psychology & Philosophy at New College, Oxford, where I spent lots of time thinking about empathy and deep learning theory.

I write at Seeing Truly on Substack. You can leave an anonymous message for me here. Send me an email if you’d like to chat!