A new study reveals how the brain processes speaking and listening in real time, showing shared neural codes during short verbal exchanges and distinct timing networks when dealing with longer, overarching ideas. The findings, published in Nature Human Behaviour, map how we navigate the back-and-forth demands of everyday conversation.
Conversation is a dynamic task requiring quick context understanding, anticipation, and narrative coherence. To manage this, the brain must integrate language across multiple timescales—from immediate word processing to grasping whole paragraphs and concepts.
For decades, researchers have explored language processing with isolated sentences or long narratives. Such work established a hierarchical organisation of language in the brain, but how the brain handles spontaneous, real-time dialogue remained less clear.
Masahiro Yamashita and Shinji Nishimoto, scientists at Osaka University and the National Institute of Information and Communications Technology in Japan, led the team behind the study. They set out to determine whether speaking and listening recruit the same linguistic representations or rely on separate processes.
Eight native Japanese speakers participated in the experiment, lying in a functional MRI scanner while engaging in unscripted conversations with an experimenter through a microphone and earphones. Each session lasted about three hours and covered casual topics such as favourite classes and personal introductions.
To translate spontaneous speech into analysable data, the researchers used a large language model based on the GPT architecture. Transcripts of the conversations were fed into the model to extract mathematical representations of meaning, known as contextual embeddings, across context windows spanning one second to thirty-two seconds.
The researchers then constructed computer models to predict brain activity from the AI-derived embeddings. Initially, they tested a unified model that treated speaking and listening as a single pool of meaning. On short timescales of one to four seconds, the brain’s linguistic representations overlapped for both producing speech and listening, with shared neural codes appearing in the prefrontal, temporal and parietal cortices.
Conversations organised across multiple timescales in the brain
Yet the picture changed when examining longer stretches of conversation. Context windows of sixteen to thirty-two seconds revealed shared representations that were widely scattered across the brain and varied considerably from person to person. In some instances, activity extended into regions associated with inferring others’ thoughts and recalling memories, suggesting that while a universal mechanism supports immediate processing, individuals employ personalised strategies to integrate longer social context and conversational history.
The team then isolated brain activity specifically tied to speaking versus listening. They found an opposing timescale preference for each action: regions involved in producing speech were most active with short-term context, whereas regions tied to understanding the other person’s speech were most engaged by longer contexts spanning multiple sentences.
This division aligns with the demands of the two tasks—speaking requiring rapid responses to what was just said, and listening requiring the maintenance of information over time to form a stable mental model of the dialogue.
Additionally, the researchers identified bimodal brain zones that responded robustly to both speaking and listening, but in independent ways. These regions peaked in activity when context lengths reached eight seconds or more, encoding meaning for both tasks without sharing identical neural patterns. The authors interpret these areas as specialised to help people keep track of perspectives within a dialogue and to manage social juggling during conversation.
Using a statistical approach, the study also teased apart which word types drove the strongest brain responses at short timescales. They found that conversational fillers and short confirmations, such as “yeah” or “uh,” evoked distinct neural patterns. These tiny, low-effort words act as social glue, helping maintain the flow of dialogue and separating them from more logical or factual statements.
However, the study’s small size—eight participants—and the absence of functional localiser tasks mean the findings cannot yet be mapped onto precisely defined cognitive networks. The researchers say that future studies with larger groups and additional mapping techniques will be needed to confirm exactly which brain networks manage these conversational timelines.
The study, titled Conversational content is organized across multiple timescales in the brain, is authored by Masahiro Yamashita, Rieko Kubo and Shinji Nishimoto.
