Before the First Note: Where Sound Is Born

Author: Inna Horoshkina One

Before the First Note: Where Sound Is Born-1
Silence enters the body as a golden vertical, unfolds in the heart and becomes breathing, a wave, a sound.

First — silence.

But this silence is not empty. It already contains that which strives to manifest: a barely discernible inner sounding that has not yet become a voice, a melody, or a movement.

The musician listens to it not only with hearing. He receives it with attention, allows it to enter his breath, to pass through his muscles, gestures, and hands. And only then does that which existed within, without physical sound, become a vibration of the air — a note that others are able to hear.

Silence gives birth to intention.
Intention enters the body.
The body translates the invisible into sound.

Now neuroscience has drawn closer to this boundary of manifestation. Researchers are learning to recognize brain activity that arises even before a person is able to utter a word, and to turn it into audible speech, facial expressions, and the movement of a digital body.

Science does not yet explain the mystery of the birth of music. But it is already capable of seeing: before a voice sounds in the outer world, an entire symphony takes place within a person.

To Restore Not Only Words

On 14 September 2026, in the peer-reviewed journal Nature Neuroscience a paper was published: Simultaneous speech and gesture decoding for multimodal communication in paralysis — “Simultaneous decoding of speech and gestures for multimodal communication in paralysis.”

The study was conducted by a team from the University of California, San Francisco, under the leadership of neurosurgeon and neuroscientist Edward Chang. The lead author was the researcher Samantha Brosler.

The experiment involved three people with varying degrees of paralysis of the speech apparatus and body. Electrocorticographic arrays — thin plates with electrodes that register neural activity — were placed on the surface of the motor cortex of their brains.

The participants attempted to utter given phrases and perform familiar gestures: waving a hand, giving a thumbs-up, shaking their head. Sometimes they did this separately, sometimes simultaneously.

Machine learning algorithms searched the resulting signals for patterns corresponding to an attempt to say certain words or make a movement. The recognized intention was then translated into the actions of a digital avatar.

As a result, two participants were able to control their speech and gestures simultaneously.

This is a fundamentally important step. Human communication is never limited to a sequence of words. We speak with our whole body: we change our facial expression, turn our head, move our hands, emphasize meaning with a glance and a pause.

The new technology strives to restore to a person not only the ability to convey information, but also a part of his individual presence.

The body does not accompany the voice — it takes part in it

One of the most interesting results of the study concerns how the brain connects speech and movement.

One might have assumed that uttering a phrase and gesturing simultaneously is simply the addition of two independent signals. However, the neural activity during combined expression differed from that which arose when words and movements were performed separately.

The algorithms recognized simultaneous speech and gestures more accurately if they had first been trained specifically on such combined actions.

In other words, human expression turned out not to be a mechanical sum of separate elements.

Voice, face, and movement are born as a single event.

We can observe this in music as well. A singer does not simply produce sound with the vocal cords. Breathing, the position of the spine, the chest, the muscles of the face, and an inner sense of space all take part in the sounding. A pianist touches the key with his whole body, even if the visible movement is made by a single finger. A dancer hears the rhythm not only with his ears — he translates it into weight, direction, and trajectory of movement.

The body is not packaging for sound. It becomes part of the musical utterance.

The sound that does not yet exist

On 23 September 2026, a major scientific review appeared on the arXiv platform Brain-to-Language Decoding: Tasks, Signals, Methods, Evaluation, Practical Use and Beyond. Its authors compared studies in which brain activity is being translated into text, synthesized speech, acoustic sound, and visual expression.

The authors examine three directions:

  • speech that a person is attempting to utter;
  • inner speech that sounds without external utterance;
  • the perception of heard language.

The review is currently a preprint and has not undergone peer review. This is not an announcement of the creation of a new system, and even less a proof of the possibility of freely reading human thoughts. Its value lies elsewhere: it shows how close various scientific fields have come to the space between intention and manifestation.

Researchers are already working not only with individual commands. Some experimental systems create streaming synthesized speech, bring its sound closer to a person's individual voice, and connect words with the movements of a digital face.

But the more expressive such systems become, the more important the question becomes: what exactly do they recognize?

A brain signal does not contain a ready-made sentence that one merely needs to "listen to." The algorithm interprets the complex activity of the nervous system, compares it with training data, and selects the most probable result.

Therefore, a coherently sounding phrase does not yet mean that the machine has fully and flawlessly restored the inner experience of a person. Part of the result may be completed by a language model.

This is not soul-reading. It is a complex translation — as yet imperfect, individually tuned, and requiring constant feedback with the person himself.

Who is the author of the restored voice

Until now, the main goal of speech neuroprostheses has been accuracy: how many words the system recognized correctly and how quickly it was able to reproduce them.

Now a deeper question arises — that of authorship.

If an algorithm helps a person express an intention, it becomes an instrument, similar to a musical instrument. But if the system begins to complete too much on its own, the boundary between assistance and substitution becomes less obvious.

Therefore, the development of such technologies cannot be assessed only by speed and recognition percentages. User control, the ability to correct the result, transparency of the algorithm's operation, and the preservation of a person's right to decide what exactly should sound on his behalf are necessary.

A true voice is not only the correct set of words.

It is timbre.
Rhythm.
Pause.
Gesture.
Emotional coloring.
And the freedom not to speak.

To return a voice means to return to a person the ability to recognize himself in his own expression.

What this opens up for music

Musicians have long known what neuroscience is only now finding language for: sounding begins before the physical note.

A composer can hear an inner musical line before touching the instrument. A singer is able to sense the direction of a phrase in advance and only then fill it with breath. An improviser does not always know which note will come next, but feels the movement that is already preparing to pass through the body.

Science describes these processes through motor planning, auditory imagery, memory, prediction, and the interaction of different areas of the brain. The musician experiences them as a moment of inner recognition:

here is the sound that wants to manifest.

This does not mean that a device is capable of recording finished music directly from consciousness or of revealing the source of inspiration. Between the inner sensation and the work there remains an enormous path — choice, training, technique, culture, memory, and the inimitable history of the person.

But new research allows us to see the very moment of creation differently.

Music is not born when the air begins to vibrate. Physical sound is already one of the last stages of the process. Before it there exists directed attention, a foreboding of form, and the body's readiness to embody what cannot yet be heard from outside.

Before the first note

We are accustomed to considering the beginning of music to be a strike on a drum, a touch on a key, or the first movement of the vocal cords.

But perhaps this is no longer the beginning, but the manifestation.

Before the first note there exists an inner space in which sound, feeling, and movement are not yet separated. Then intention arises. It gathers attention, changes the breath, sets neural and muscular processes in motion — and gradually enters into matter.

Neuroscience is learning to register individual stages of this transition. Technology is trying to restore to people with paralysis the ability to express themselves. And music reminds us that the voice was never merely an acoustic signal.

It is born of the whole person.

And perhaps the very first note does not sound in the air. It arises at that instant when the Silence within us chooses the form in which it wants to become audible.


17 Views

Comments

Did you find an error or inaccuracy?We will consider your comments as soon as possible.