How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation
Ke-Han Lu, Szu-Wei Fu, Chao-Han Huck Yang, et al.
This paper investigates how much knowledge about sounds and audio Large Language Models (LLMs) naturally learn from text-only training, and whether this affects their performance when adapted to handle audio. The researchers test different LLMs in three ways: directly questioning them about audio concepts, having them reason about audio descriptions, and fine-tuning them with audio data. They find that the amount of audio knowledge varies significantly between different LLM families, and importantly, LLMs that show better audio understanding in text-only tests also perform better when actually processing audio.
Large Language ModelsAudio ProcessingMultimodal LearningModel Evaluation