Exploring Phonological Subspaces in Self-Supervised Speech Representations
Kulkarni, Atharva Abhijit
Kulkarni, Atharva Abhijit
Citations
Altmetric:
Author
Supervisor
Department
Natural Language Processing
Embargo End Date
Type
Thesis
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Understanding how linguistic structure is represented in neural models of speech remains an open challenge. While self-supervised speech models such as HuBERT achieve strong performance on downstream tasks, the nature of the representations they learn is not fully understood. In particular, it is unclear whether these models encode interpretable phonological structure or rely on distributed statistical patterns without explicit linguistic grounding.
In this thesis, we investigate the extent to which phonological features are encoded in HuBERT representations. We adopt a feature-based framework grounded in phonological theory and extract articulatory feature labels using the PanPhon toolkit. Using sparse logistic regression probes, we identify low-dimensional subspaces corresponding to individual phonological features and analyze their structure across layers. Our analysis shows that phonological structure emerges hierarchically: early layers capture coarse acoustic distinctions, while deeper layers refine representations into more linguistically meaningful feature groupings. Single-unit analysis reveals that only coarse features such as sonorant and strident are reliably captured by individual units, whereas more fine-grained articulatory features are encoded in a distributed manner. We show that these subspaces are functionally meaningful by using them for phone recognition, achieving competitive performance while reducing dimensionality from 768 to approximately 300 dimensions. Overall, our findings suggest that phonological structure emerges naturally in selfsupervised speech models, but is encoded in distributed, low-dimensional subspaces rather than monosemantic feature units.
Citation
Kulkarni, Atharva Abhijit, "Exploring Phonological Subspaces in Self-Supervised Speech Representations," M.S. Thesis, Natural Language Processing, MBZUAI, 2026.
Source
Conference
Keywords
Phonology, Speech Processing, Mechanistic Interpretability
