Research author
Hung-yi Lee
8 AI research papers in the One9Founders library, with summaries and links to original sources.
Papers by Hung-yi Lee
Joint Optimization of Tool Creation and Use for Large Language Model Agents
Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen, et al.
Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models
Yu-Han Huang, Chih-Kai Yang, Ke-Han Lu, et al.
BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech
Ho Lam Chung, Bo-Xuan Zheng, Cheng-Chieh Huang, et al.
REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing
Cheng-Kang Chou, Ming-To Chuang, Ke-Han Lu, et al.
Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models
Chun-Yi Kuan, Wei-Ping Huang, Hung-yi Lee
All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation
Leonardo Haw-Yang Foo, Chih-Kai Yang, Chen-An Li, et al.
TiCo: Time-Controllable Training for Spoken Dialogue Models
Kai-Wei Chang, Wei-Chih Chen, En-Pei Hu, et al.
How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation
Ke-Han Lu, Szu-Wei Fu, Chao-Han Huck Yang, et al.
This paper investigates how much knowledge about sounds and audio Large Language Models (LLMs) naturally learn from text-only training, and whether this affects their performance when adapted to handle audio. The researchers test different LLMs in three ways: directly questioning them about audio concepts, having them reason about audio descriptions, and fine-tuning them with audio data. They find that the amount of audio knowledge varies significantly between different LLM families, and importantly, LLMs that show better audio understanding in text-only tests also perform better when actually processing audio.