Back to papers
March 19, 2026cs.CLcs.AIAdvanced
UGID: Unified Graph Isomorphism for Debiasing Large Language Models
AI-Generated Summary
This paper addresses social biases embedded in large language models by proposing UGID, a method that treats the model's internal structure as a graph and forces it to process similar inputs (that differ only in sensitive attributes like gender or race) in the same way. The approach works by constraining both the attention mechanisms and hidden representations in bias-sensitive areas of the model, while using special techniques to prevent the model from losing its core knowledge and safety features.
Difficulty
Advanced
Categories
cs.CL, cs.AI
AI Tags
bias mitigationlarge language modelsinterpretabilityfairnessinternal representationsTransformers