NEWS
NEWS

A brain implant enables paralyzed individuals to not only speak but also gesture through an avatar

Updated

A study published in the journal 'Nature Neuroscience' paves the way for much more real and human social interactions for patients trapped in their own bodies after suffering a stroke or neurodegenerative diseases

A patient facing their avatar, communicating words and gestures.
A patient facing their avatar, communicating words and gestures.UCSF

Human communication is a complex system in which not only words are involved, but also gestures and looks. When we converse, we do not limit ourselves to emitting sounds. We also nod to agree, shrug our shoulders (expressing doubt), squint our eyes... However, individuals who suffer severe paralysis due to strokes or neurodegenerative diseases completely lose the ability to speak and gesture, greatly reducing their ability to interact with others. Over the past decade, science has achieved extraordinary milestones by reading the mind to translate thoughts into text or move robotic arms, but always analyzing speech or movement in isolation. Until now.

A clinical study coordinated by neurosurgeon Edward Chang at the University of California, San Francisco (UCSF) has demonstrated for the first time that a single brain device assisted by artificial intelligence (AI) can decode speech and physical gestures simultaneously. The breakthrough, which is published in the journal Nature Neuroscience, not only translates silent electrical impulses into words but also uses them to bring to life in real-time a full-body virtual avatar that replicates the patient's gestures while speaking.

Behind the engineering components of this trial, named BRAVO (Brain-Computer Interface Restoration of Arm and Voice), are the human stories of the participants. Among them stands out Bravo-1r, a man who suffered a massive brainstem stroke at the age of 20, leaving him with severe tetraplegia and the complete inability to articulate words. His native language is Spanish, and today, with his intellect intact, his method of communicating with the implant consists of trying to speak completely silently and sketching subtle and minimal gestures, an invisible effort that was previously trapped inside him.

The other main pillar of the study is Bravo-6, a patient diagnosed with amyotrophic lateral sclerosis (ALS) at the age of 56. In just nine months, the neurodegenerative disease completely robbed him of intelligible speech and the dexterity of his hands. Unlike Bravo-1r, Bravo-6's strategy involves trying to speak by emitting slight unintelligible vocalizations and covertly imagining the movements of his body without physically executing them. Another patient, Bravo-3, initially participated in the study, assisting in the mapping of isolated movements before withdrawing from the clinical protocol.

To achieve the miracle of multimodal communication, the team of surgeons implanted a thin flexible metal mesh equipped with 253 small high-definition electrodes in the left hemisphere of the patients' brains. This implant sits directly on the pia, the surface of the sensorimotor cortex, a key brain region that acts as the neural center where the brain orchestrates both the commands to move the body's muscles and the delicate movements of the lips, tongue, and vocal cords.

The electrical chaos in the cerebral cortex

The real scientific challenge lay not only in capturing the brain signals but in unraveling the chaos that occurs when a person tries to speak and gesture at the same time. Traditionally, neuroscience assumed a rigid and segregated organization of the brain's motor map: hand movements were controlled in an upper area and speech apparatus in a lower one. However, researchers at UCSF discovered that reality is much more intertwined and complex.

When the electrodes recorded the patients' brain activity in real-time, they observed that the neurons responsible for activating the arms and those planning speech overlap in a common space in the cortex. If the patient tried to say "hello" while waving their hand to greet, the electrical currents mixed in such a way that completely confused the computers. A decoder trained exclusively to recognize words failed miserably when the user gestured at the same time, interpreting the hand movement as an incorrect sound or causing system failures.

The solution came through a radical redesign in AI training: instead of processing the data independently, engineers developed state-of-the-art neural networks and trained them using combined databases. In other words, they fed the system with examples where the patient only spoke, examples where they only gestured, and crucially, mixed sessions where they performed both actions simultaneously.

This inclusive approach allowed the AI to learn the deep mathematical structure of coordinated thinking. By understanding how signals interact when they occur simultaneously, the algorithm autonomously learned to downplay the electrodes generating cross-interference and rely on the cleaner and more specific channels for each task. The combined training resolved the overlap and reduced system false alarms when the patient rested or performed only one of the actions.

An avatar with 100% accuracy

Real-time results show an unprecedented advancement. The patients' brain electrical signals were transmitted to software that processed the data in a few milliseconds (7.5). This information simultaneously fed two decoders that instantly sent commands to Unreal Engine, a three-dimensional graphic animation engine. On the screen, a full-body digital avatar, aesthetically designed according to each participant's personal preferences and pre-selected attire, immediately came to life.

During real conversation tests, where researchers asked informal questions, the system's performance demonstrated tremendous stability. Bravo-6, the ALS patient, achieved 85% accuracy in recognizing his gestures and 75% in transcribing his phrases, controlling a repertoire of 10 physical expressions and 10 verbal statements combinable with each other.

On the other hand, patient Bravo-1r recorded a 100% accuracy in both speech decoding and gestures throughout multiple blocks of active conversation. While the desired words appeared instantly written as subtitles at the bottom of the screen, his avatar waved, nodded, expressed doubt by shrugging, or celebrated a response by energetically closing his fist.

The research led by Chang has opened a very promising path by demonstrating a fundamental principle: artificial intelligence can combine complex actions that have never been seen together before. In advanced mathematical simulations, the algorithm showed that it does not need to individually memorize each possible pairing of gestures and words to interpret them. If the system already knows the isolated signal for the word "thank you" and the signal for the clapping gesture, it can successfully decode both actions the first time the patient decides to spontaneously combine them. This property is crucial for the future of technology, as it allows expanding the patients' vocabulary to thousands of words and dozens of everyday movements without the need for extensive calibration sessions in the laboratory.

Despite the widespread optimism, the study's authors themselves maintain caution and point out the logical limitations of an early-stage trial. At the moment, the tested vocabulary remains restricted and limited to controlled laboratory environments. Additionally, the processing speed and long-term stability of the implant -successfully evaluated over several months- will require additional tests in much larger patient groups with diverse clinical profiles before becoming a common medical assistance product.

Nevertheless, the milestone is there. By unifying speech and body through a single digital channel, medicine has not only restored the patients' ability to convey information but also their expressiveness, identity, and the naturalness of face-to-face conversation. It is a step forward for brain-computer interfaces to move away from being cold text processors and more accurately reflect human personality.