[Submitted on 12 May 2026 (v1), last revised 17 Jul 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:While deep transformer-based models have advanced rapidly, their internal mechanisms remain largely a mystery. Recent work has prioritized understanding text-based transformer models, leaving ASR systems largely unexplored. In order to address this gap, we examine the internal representations of Whisper's encoder using a sparse autoencoder. We find diverse monosemantic features across linguistic and non-linguistic boundaries, spanning a hierarchy from phonetic to semantic representations, and conduct a causal feature-steering campaign across this hierarchy, including cross-lingual steering. We further find that steering is more reliable for higher-level features than lower-level ones, an asymmetry that may reflect redundant encoding of lower-level information. Altogether, this work demonstrates that Whisper's encoder represents a surprisingly rich hierarchy of linguistic information that extends well beyond what is strictly necessary for transcription.

Submission history

From: Zachary Houghton [view email]
[v1] Tue, 12 May 2026 15:02:20 UTC (444 KB)
[v2] Fri, 17 Jul 2026 21:21:18 UTC (430 KB)