Bilgehan Sel

AI safety research

I work on AI safety: safe decision-making in foundation models. How language models reason and plan over long horizons, whether alignment holds up under adversarial pressure, and how to build systems that recognize the limits of their own certainty.

I completed my PhD in Electrical & Computer Engineering at Virginia Tech, advised by Ming Jin. Most recently I was a Member of Technical Staff at Anthropic, working on the adversarial robustness of safety classifiers; before that I was a Student Researcher at Google.


Selected Papers

All 22 publications →


Background

Anthropic Fellow (2025) · Amazon Fellow (2024)