Back to papers
March 19, 2026cs.CLcs.AIAdvanced

UGID: Unified Graph Isomorphism for Debiasing Large Language Models

AI-Generated Summary

This paper addresses social biases embedded in large language models by proposing UGID, a method that treats the model's internal structure as a graph and forces it to process similar inputs (that differ only in sensitive attributes like gender or race) in the same way. The approach works by constraining both the attention mechanisms and hidden representations in bias-sensitive areas of the model, while using special techniques to prevent the model from losing its core knowledge and safety features.

Difficulty
Advanced
Categories

cs.CL, cs.AI

AI Tags
bias mitigationlarge language modelsinterpretabilityfairnessinternal representationsTransformers