Skip to content
Scalable Trustworthy AI Scalable Trustworthy AI

Interests

  • Mechanistic interpretability, AI alignment

Biography

What a model represents internally and what it ends up saying are not always the same thing. Models already encode signals about their own correctness, error modes, and uncertainty, yet they rarely recognize or act on that information. I am interested in three questions this raises: how to read out internal states, how to quantify the gap between internals and outputs, and how to close it so that expressed behavior is genuinely aligned with what a model represents.

So far I have examined whether unlearned knowledge is genuinely removed or only suppressed at the representation level, how deeply unlearning actually reaches, and how to diagnose failures in automated interpretability pipelines such as those built on SAEs.