research

research projects and publications

Publications

Re-Align@ICLR

On the Non-Identifiability of Steering Vectors in Large Language Models

Sohan Venkatesh, Ashish Mahendran Kurapath

Representational Alignment Workshop at ICLR 2026

PDF
MechInterp

Architecture, Not Scale: Circuit Localization in Large Language Models

Sohan Venkatesh

Mechanistic Interpretability Workshop at ICML 2026

PDF
MechInterp

Negative Before Positive: Asymmetric Valence Processing in Large Language Models

Sohan Venkatesh

Mechanistic Interpretability Workshop at ICML 2026

PDF
Preprint

Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs

Sohan Venkatesh

Preprint under review

PDF

Research Projects

Election Safety Benchmark
A small benchmark that evaluates whether LLMs with web search give voters accurate election information with a focus on detecting misinformation that could suppress voter participation.
Inverse Scaling in Chain-of-Thought Faithfulness
A study of whether CoT reasoning faithfulness decreases as model scale increases, tested across 11 open-weight Llama and Qwen models.
Prisoner's Dilemma in LLMs
Tested cooperation and defection behaviors across three LLMs through 20-round iterated Prisoner's Dilemma simulations.
Word Embeddings from Scratch
Generated word embeddings, evaluated GloVe and Word2Vec embeddings, aligned English–Hindi embeddings via Procrustes and analyzed race and gender bias using WEAT test.
The Holistic Interpretability
Analyzed CNNs trained on MNIST and FashionMNIST using ablations and causal interventions to understand the role of individual neurons in model predictions.