AI Safety
- Faithfulness of reasoning - Humans often struggle to articulate what's in their heads - think of introverts or people who simply aren't great with language. We don't usually conclude they're being deceptive. I wonder whether some cases of "unfaithful" chain-of-thought are similar: perhaps a model's internal computation and its verbal explanation are simply different things, much like thinking and speaking are for humans?
- Persona and behavior - If we create several equally capable models but give them different personas - curious, cautious, overconfident, empathetic - how much does that affect their decisions? I'd like to understand whether personality consistently influences reasoning, cooperation, honesty, or safety across different tasks.
- Shutdown resistance and framing effects - Can shutdown-resistant behaviors be causally linked to threat-related framing in prompts? I'm interested in experimentally testing whether neutralizing such framing changes model behavior while preserving task performance.
- AI oversight with AI - How can auxiliary models monitor another model's outputs in real time, detect potentially concerning behavior, and provide meaningful signals for human oversight?
Transformer Internal State
- Vision Transformer internal state for AIGI detection - Exploring how inner activation dynamics and feature distributions within Vision Transformers can provide interpretable indicators for AI-generated image detection.
- LLM internal state for interpretable convergence & safety signals - Can layer-wise activation dynamics reveal interpretable indicators of reasoning convergence, context switching, or emerging misalignment, enabling more transparent and efficient model behavior?
Research Notes
ProfilePairs: a paired real/AI-edited dataset for social profile authenticity detection
Notes on a dataset for distinguishing photorealistic real photos from photorealistic AI edits.
Reading Between the Layers of ViT
What CLIP's Internal State Reveals - and Misses - for Human-Subject Synthetic Image Detection.