On the front lines of AI safety

Kevin Chen MPP '27 reflects on several days at the AI safety hub Lighthaven, where he and peers participated in an intensive residential program aimed at contributing to the AI safety and governance landscape.

Published
Author
Kevin Chen

In September, I spent four days at Lighthaven, a campus in Berkeley, California, that has become one of the premier hubs for AI safety researchers, as one of 24 participants in BlueDot Impact's first Context Week. Those were some of the most important days I have spent honing my policy path since arriving at Jackson, and genuinely one of the most important incubators I have been a part of since becoming a student.

Kevin Chen at Lighthaven
Kevin Chen

BlueDot Impact is the organization behind the AI Safety Fundamentals courses that many people now working in the field started with. Context Week, a new experiment for them, was borne from the idea that most people trying to move into AI safety are held back by a lack of context — meaning the knowledge and judgment to understand where the field stands, which threat models matter, which strategies are being pursued and by whom, and where their own skills fit. BlueDot wanted to see whether that could be taught in an intensive residential setting to a small, hand-picked group.

It was a bit of a challenge to be accepted, especially in this “decades in weeks” moment for this summer in AI safety. The program was announced with about a week and a half of lead time, hundreds of applications came in from all over, and everyone who made it past the written round did an intensive interview that really explored your motivations in the field before the final cohort was chosen. Twenty-four of us were selected; I was one of very few people coming from a policy school with prior public sector experience, which turned out to be an asset.

The facilitators and guests who cycled through included engineers from the frontier labs, researchers from the evaluation and auditing firms that test those labs' models, and policy people who have spent time inside government thinking about how any of this gets regulated. Essentially, the kind of folks I usually only read at a distance from Substacks or Twitter threads. Late into the night, they were happy to argue about timelines, explain why a particular safety agenda had stalled, or tell you frankly which organizations were worth your time. That individualized candor is hard to get from a paper or a panel, especially in this burgeoning field.

The sessions ranged from the concrete to the speculative. One of the first had us work through the primary reports on the OpenAI and Hugging Face incident from this summer, and then debate what it implied for research and for governance. Another had each of us sketch an AI-enabled future we would endorse and we really delved into how desirable and how stable we thought those futures were. Running through all of it was a habit the facilitators drilled relentlessly. “Why do you believe that?” “What is the mechanism by which that intervention helps?” “What would change your mind?” I left with my own beliefs in AI safety exposed as shakier than I thought.

My research here at Jackson focuses on what the AI safety field calls epistemic risk — more specifically, the way governments and national security institutions are coming to rely on AI systems for drafting and analysis, and the gap between the AI safety community and the international relations community in taking that seriously. Context Week was four days of living and really scrutinizing this topic. 

I came back with a far better read on what the technical side actually worries about and where the policy world is talking past it. 

If I have one takeaway for other Jackson students it is that the AI safety world is far more open to policy people than it looks from the outside — and far more hungry for them. Jackson, through the Schmidt Program, is one of the few places training people who can move between the two. 

Thank you to Yale Jackson and the Schmidt Program for treating a week like this as part of the education, and to Professor Ted Wittenstein for his support and for pushing me toward this community. At BlueDot, thank you to Jack Douglass, who ran operations for two dozen strangers without visible strain, and to Harry Waterman and Joshua Landes, who designed the program and wrote up their own lessons from it. One last major thanks to James Coates for being an amazing research mentor who nurtured my interest in AI safety. 
 

Participants at Lighthaven having a discussion

Learn more