Sohan Venkatesh

preview.png

Hey there! I’m Sohan. I’m currently a research fellow at LASR Labs working with Bartosz Cywinski (Google DeepMind).

I work on AI safety and interpretability with the goal of reducing catastrophic risks caused by AI systems. To that end, I have previously worked on various projects including my recent papers on steering vectors and on mechanistic interpretability. Feel free to email me if you find my work interesting!

I’m motivated by effective altruism and I see AI safety as the most altruistic use of my technical skills. Outside of work/research, you will often find me reading, ranting about AI risks or going down Wikipedia rabbit holes at 2am (try it — follow the first link on any Wikipedia article 20 times and you’ll always end up on ‘Philosophy’).

news

02 Oct 2026 My paper in collaboration with MILA got accepted to NeurIPS Agents in the Wild Workshop 2026.
14 Sep 2026 Mentoring a project on CoT faithfulness in SPAR Fall’26.
20 Jul 2026 Started as a research fellow at LASR Labs, London.
11 Jun 2026 My papers on circuit localization and emotional valence in LLMs have been accepted at ICML Mechanistic Interpretability Workshop 2026!
21 Apr 2026 I will be heading to Rio to present my paper at ICLR Re-Align and CAO workshops!

publications

Re-Align@ICLR

On the Non-Identifiability of Steering Vectors in Large Language Models

Sohan Venkatesh, Ashish Mahendran Kurapath

Representational Alignment Workshop at ICLR 2026

PDF
MechInterp

Architecture, Not Scale: Circuit Localization in Large Language Models

Sohan Venkatesh

Mechanistic Interpretability Workshop at ICML 2026

PDF
MechInterp

Negative Before Positive: Asymmetric Valence Processing in Large Language Models

Sohan Venkatesh

Mechanistic Interpretability Workshop at ICML 2026

PDF
AIWILD

Where Empathy Goes Wrong: Emotional Appeals Compromise AI-to-AI Oversight

Sohan Venkatesh, Thomas Jiralerspong, Flemming Kondrup

Agents in the Wild Workshop at NeurIPS 2026

PDF
AIMS@COLM

Algorithmic Blindness in Large Language Models: A Calibration Study of Performance Prediction

Sohan Venkatesh, Ashish Mahendran Kurapath, Tejas Melkote

AI Measurement Science Workshop at COLM 2026

PDF
Preprint

Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs

Sohan Venkatesh

arXiv preprint

PDF

latest posts