Can AI Read the Room? USC Study Finds AI Is Better at Reading Than Listening
Today, users can do more than exchange text with artificial intelligence (AI). From holding conversations and asking questions aloud to sending voice messages, AI can now process audios and respond in real time. A new study led by USC researchers suggests the answer is no.
Today, users can do more than exchange text with artificial intelligence (AI). From holding conversations and asking questions aloud to sending voice messages, AI can now process audios and respond in real time.
A new study led by USC researchers suggests the answer is no. The team found that even today’s most advanced audio large language models (LLMs) struggle to interpret information beyond the spoken words, often taking language too literally and missing the nonverbal cues that.
What Happened
In everyday conversations, meaning extends far beyond words. Tone of voice, emotion, emphasis, pitch and other paralinguistic information provide important social and emotional context that helps people interpret what someone truly means.
This allows the model to access both the early layers of processing—which are better at capturing basic acoustic information such as pitch, volume and speaker characteristics—and the deeper layers, which specialize in understanding vocabulary.
Instead of relying only on the encoder’s final layer, PCLM draws information from multiple layers simultaneously before passing it to the language model.
As audio passes through an encoder, acoustic details such as pitch and tone are often suppressed in the deeper layers as the model becomes increasingly focused on words.
Key Details
USC professor Mohammad Soleymani led the research project to uncover why audio LLMs struggle to interpret these listening cues and developed new techniques that significantly improve their ability to understand the “how” of speech—not just the spoken words. Soleymani is a research.
First, the researchers developed a module called the Prompt-Conditioned Layer Mixer (PCLM), which helps the model determine exactly which part of the audio it should “listen” to based on the question it is being.
The team proposed a two-part solution designed to help Audio LLMs both retain and value what they hear.
The team also identified what they call a “utilization gap.” Although the correct acoustic information often still exists within the model’s internal representations, the model’s decision-making process is so heavily biased toward language that.
Why It Matters
He also leads the USC Intelligent Human Perception (IHP) Lab. The research team also included Soleymani’s PhD students Ashutosh Chaubey, and Jiacheng Pang, a former master’s student from his lab.
By examining the models’ internal computations in real time, the researchers found that acoustic details, such as pitch and tone, gradually become “stripped away” as audio data moves through the model’s layers toward its.
Soleymani’s team then set out to understand why Audio LLMs fail at interpreting paralinguistic information by probing the models’ internal layers—almost like performing “brain surgery” on AI.
What Reports Say
Coverage of the story so far points to:
Continued reporting by USC Viterbi School of Engineering as more details emerge