Towards Reliable AI Fairness: Challenges in Steering Features within Bias-Implicated Neurons

Ismael Garrido-Munoz, Arturo Montejo-Raez, Fernando Martínez-Santiago


Abstract
LLMs perpetuate societal biases, such as gender stereotypes, reinforcing harmful norms and posing significant fairness risks in real-world applications. We investigate a fine-grained mitigation technique that moves beyond surface-level fixes. Our approach uses attribution graphs to identify and directly steer bias-implicated features within a Sparse Autoencoder’s (SAE) latent space. This method, known as feature steering, offers a theoretically precise, surgical intervention aimed at correcting bias at its neural source without costly retraining. We critically examine its practical reliability across various contexts. We find that steering effectiveness is highly sensitive to parameter tuning, often requiring unpredictable, context-specific adjustments. The intervention’s success exists in narrow “sweet spots,” outside of which performance can degrade catastrophically. This demonstrates that while direct intervention on learned features is a powerful analytical tool, significant challenges of brittleness and instability hinder its application as a consistent, broad-scale debiasing solution, necessitating research into more robust control mechanisms.
Anthology ID:
2026.lrec-1.306
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
3851–3860
Language:
External URL:
https://lrec.elra.info/lrec2026-main-306
DOI:
10.63317/2iexsnkqn3j6
Bibkey:
Cite (ACL):
Ismael Garrido-Munoz, Arturo Montejo-Raez, and Fernando Martínez-Santiago. 2026. Towards Reliable AI Fairness: Challenges in Steering Features within Bias-Implicated Neurons. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 3851–3860, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Towards Reliable AI Fairness: Challenges in Steering Features within Bias-Implicated Neurons (Garrido-Munoz et al., LREC 2026)
Copy Citation: