Headphones recommended.It also works on a phone; the settings panels sit side by side on a wide screen and stack on a phone. Start with the volume low.
SIMULATOR · COCHLEAR IMPLANT

Vocoder Simulator

Pass a speech or music sample through a channel vocoder that works much like cochlear implant processing: change the number of channels, the carrier, the insertion depth, the current spread and the envelope cutoff, and listen to the result straight away. You can also try electric-acoustic stimulation, where acoustic hearing is kept in the low frequencies, channel selection (n-of-m) and switched-off electrodes, and export every parameter as a report.

For educational use · not an exact copy of what an implant user hears

  • 3 carriersnoise, sine and PSHC
  • Acoustic + electricEAS and the frequency gap
  • Processed in your browseraudio is never sent anywhere
0BEFORE YOU START

A short introduction to the vocoder

If you have never come across a vocoder, start here. Each box opens with a one-sentence summary; click it to expand and see the detail. The Try buttons set things up for you, play the sound and take you to the listening area.

The name comes from voice coder. Homer Dudley built the first one at Bell Laboratories in the 1930s to send speech over telephone lines with less information (Dudley, 1939). Today it is also known for the "robot voice" effect in music.

A vocoder works in four steps:

  1. Split. Filters divide the sound into a few frequency bands.
  2. Take the envelope. In each band the slow change in loudness over time is extracted; the fast vibrations are discarded.
  3. Rebuild. This envelope controls the loudness of an artificial sound, such as noise or a tone (the carrier).
  4. Add up. The bands are combined; the result is the sound you hear.

The surprising part: although most of the fine detail is thrown away, sentences are largely understood in quiet with only 3-4 bands (Shannon et al., 1995).

Try:

The speech processor of a cochlear implant divides the sound reaching the microphone into frequency channels that correspond to between 12 and 22 electrodes, depending on the manufacturer. Each channel's envelope sets the strength of the electrical pulses delivered by that electrode. Electrodes near the base of the cochlea carry high-pitched sounds, those near the top (the apex) low-pitched ones.

A vocoder builds the same logic acoustically: a carrier sound takes the place of the electrode and a normal-hearing ear the place of the auditory nerve. So you can hear how much spectral and temporal information the implant passes on.

Important: a vocoder does not show exactly what an implant user hears. The brain adapts to the implant's sound over months, and some properties of electrical stimulation cannot be imitated acoustically. The details are in Limits and sources.

Try:
Research

The effect of a single variable, such as the number of channels, insertion depth or current spread, on understanding speech is tested under controlled conditions in normal-hearing groups (Shannon et al., 1995; Friesen et al., 2001). New sound-processing ideas go through such trials before they reach implant users.

Teaching

For audiology students it makes the concepts met in programming (channels, frequency allocation, disabled electrodes, electric-acoustic stimulation) audible. Reading about the effect of a setting and hearing it are very different things.

Counselling

For families, teachers and partners it helps to show why noisy places are hard and why music may be perceived differently. While doing so, it is worth explaining that adaptation takes months and that the experience varies from person to person.

Try:
  1. Number of channels. Listen to 2, then 4, 8 and 16 channels. At what point does the sentence become intelligible?
  2. Noise. Add babble to the same 8 channels and compare with the original. Noise that hardly matters in the original makes the speech clearly harder to follow once it has gone through the vocoder.
  3. Music. Listen to the melody with 8 channels. The rhythm survives, but telling which note is higher becomes hard.
  4. Insertion depth. Set the shift to 6 mm. The voice gets thinner; the talker seems to be someone else.

Tip: you can switch with the Original and Processed buttons in the listening area while the sound is playing.

Key concepts to know

In an implant each channel corresponds to an electrode. More channels mean more frequency detail. Yet for implant users vowel and consonant recognition in noise improves little beyond about 7-8 electrodes, whereas normal-hearing listeners keep improving with a vocoder up to 20 channels (Friesen et al., 2001). In other words, the number of electrodes is not the same as the number of independent channels.

Try:

The envelope carries "when and how strong" a sound is; syllable rhythm and stress live here. The envelope cutoff sets how fast the changes that are kept can be: at 50 Hz syllable timing remains but voice pitch is lost; at 300-400 Hz the periodicity of the talker's voice pitch is partly kept (Cychosz et al., 2024).

Try:

Noise sounds hissy and represents current spread well, but it adds random fluctuations of its own. Sine carries the envelope most faithfully and sounds tonal, like whistling. PSHC is a pulse-train-like tone designed specifically to carry the envelope flat; some implant users found it closer to the sound of their own implant (Karoui et al., 2019).

Try:

If the array does not reach the apex, low frequencies are delivered to higher-pitched regions of the cochlea. This is called frequency-place mismatch, and the voice sounds thinner. This tool applies the shift in millimetres on the Greenwood map. Some implant users adapt to this shift over time.

Try:

Current from one electrode stimulates not only the nerve fibres opposite it but also their neighbours. As a result the channels blend into each other (channel interaction), and even with the same number of channels the sound gets blurred as if there were fewer. In this tool it is set in dB/octave: the smaller the number, the more mixing.

Try:

In people with residual low-frequency hearing the implant can take over only the high frequencies; acoustic and electric hearing are used together in the same ear. Natural low-frequency hearing makes voice pitch and telling talkers apart easier. Intelligibility drops as the gap between the acoustic and electric regions widens (Dorman et al., 2005).

Try:

Some processing strategies stimulate only the n strongest of m channels in each short time frame, for example 8 of 22. This keeps the important information while allowing a higher stimulation rate. The channels not selected are silent at that moment.

Try:

The horizontal axis is time and the vertical axis frequency; the brighter the colour, the more energy at that frequency at that moment. In the spectrogram of the original sound, thin horizontal lines (harmonics) are visible. In the vocoded sound they disappear and are replaced by as many broad stripes as there are channels.

1LISTENING AREA

Choose a sound, change the settings, listen

The sound is processed again after every change. You can switch between Original and Processed while it plays; playback continues from the same point. Headphones are recommended; keep the volume low at first.

Sound

Settings ready
Preset
Number of channels8
Carrier

Insertion depth0 mm
full depth · no shiftshallow · 8 mm basal

Current spreadnone

dB/octave · the smaller the value, the more neighbouring channels blend (channel interaction)

Envelope cutoff160 Hz

Hz · at 50 Hz syllable timing remains and voice pitch (F0) information goes; at 300-400 Hz pitch periodicity is partly kept

Expert settings

In each frame the n channels with the strongest pre-emphasised envelope are active and the rest are silent. Pre-emphasis is used only for the selection: it attenuates by 6 dB/octave below the corner so that the high frequencies are not pushed into the background.

Disabled channels spectral holes

Click a channel to switch it off. Drop: that band's information is lost. Split to neighbours: the band's power is shared equally between the nearest active channels (like reallocating a disabled electrode's frequencies to its neighbours).

Analysis range and band spacing
Synthesis (carrier) range

When entered manually, the analysis range is compressed or expanded into the synthesis range; the insertion depth slider is disabled.

Envelope extraction and filter slope

The filters are applied zero-phase (forward-backward); the slope is this combined effect. Current spread is applied separately, in the envelope domain; the filter slope does not change it.

Listen
Background noise

The noise is added before processing, so it enters the implant microphone together with the speech. SNR is calculated from the active (non-silent) RMS of the two signals.

With these settings
    2CHANNEL MAP

    Which frequency goes where in the cochlea?

    The top row shows the analysis bands: which range of the sound reaching the microphone falls into which channel. The bottom row shows the synthesis bands: the place that channel's carrier, that is, its electrode, stimulates. With a shift, the lines lean to the right, towards the base; this is frequency-place mismatch.

    3PROCESSING CHAIN

    What happens inside a single channel?

    The vocoder applies the same four steps in every channel. Choose a channel; a short window from where the sound is loudest is drawn step by step. Below, the spectrograms of the original and processed sound sit side by side.

    1. Analysis filterThe sound is filtered into the channel's frequency band.
    2. EnvelopeThe slow amplitude change of the band is extracted; the fine structure is discarded.
    3. CarrierThe envelope modulates a noise, sine or PSHC carrier.
    4. SummationAll channels are added together; the result is the sound you hear.
    Inside one channel
    Original with the noise, if added
    Processed dashed lines: carrier centres

    The spectrograms use a logarithmic frequency axis (100 Hz-10 kHz); the colour scale extends to 70 dB below the strongest point of each plot.

    4A/B COMPARISON

    Compare two settings on the same sentence

    Save the current settings from the listening area as A or B. You can switch between A and B while the sound plays; the parameters that differ are listed below.

    Setting A

    Not saved yet.

    Setting B

    Not saved yet.

    Once you save A and B, the differences appear here.

    5BLIND LISTENING

    What can your ear tell apart?

    Choose a question type and press "New question". The selected sound is processed with a hidden setting; you listen and guess. Apart from the question type, all settings stay the same as in the listening area.

    0 / 0

    Choose a question type to begin.

    6PARAMETER REPORT

    The full recipe for this setting

    Cychosz, Winn and Goupell (2024) recommend that a vocoder study report at least the following: carrier, number of channels, frequency range, the corner frequencies of each band, analysis and synthesis filter slopes, and envelope cutoff. The table below gives all of these plus this tool's additional assumptions.

    General
    Channels
    7LIMITS AND SOURCES

    What a vocoder shows, and what it does not

    How does a vocoder work?

    A filter bank divides the sound into frequency bands. The envelope of each band, its slow amplitude change, is extracted and the band's temporal fine structure is discarded. The envelope modulates a carrier and the channels are summed. Because cochlear implant processors also split sound into bands and put the envelope on a pulse train at each electrode, a vocoder is a simulation of the information an implant passes on that a normal-hearing ear can listen to (Shannon et al., 1995).

    Why does the number of channels matter so much?

    Shannon et al. (1995) showed that in quiet, sentences are largely understood even with three bands. In noise the picture changes: Friesen et al. (2001) found that normal-hearing listeners kept improving with a vocoder up to at least 20 channels, whereas implant users stopped improving in vowel and consonant recognition beyond about 7-8 electrodes. For words and sentences the limit was 7-10 electrodes. So 22 electrodes do not mean 22 independent channels.

    Choosing a carrier

    Sine carries the envelope faithfully but cannot imitate spread and sounds tonal. Noise represents spread better but adds random fluctuations that an implant does not have. PSHC is a harmonic complex whose pulse rate in each band is chosen so that the envelope stays as flat as possible at the output of an auditory filter (Hilkhuysen and Macherey, 2014). Mesnildrey et al. (2016) found speech reception thresholds in noise of 1.5 dB SNR for sine, 3.8 dB for PSHC and 4.9 dB for noise.

    How close is it to real implant sound?

    People who use an implant because of profound hearing loss in one ear can answer this, since they can compare their two ears. Karoui et al. (2019) reported that nine users rated PSHC as more similar than sine and noise, with no difference between sine and noise. Kopsch, Plontke and Rahne (2025) found sine more similar than noise in fifteen users (medians 5.9 and 3.6); in the first screening, simple low-pass, high-pass and band-pass filters were also rated as similar as the vocoders. With individual adjustment similarity rose to 9.7 out of 10, but the adjustment needed differed widely between people. In short, a vocoder simulates the information an implant passes on, not the sound of the implant.

    Insertion depth and adaptation

    The electrode array usually enters 22-30 mm of the roughly 35 mm cochlea. With shallow insertion, low frequencies are moved basally, to places of higher frequency; for example, peaks at 500 and 1200 Hz shift to about 900 and 2000 Hz (Cychosz et al., 2024). Users can get used to this shift over time; of Karoui et al.'s nine users, only one showed full and two partial adaptation. A few minutes of listening by a normal-hearing listener cannot imitate this adaptation.

    Electric-acoustic stimulation (EAS)

    Dorman et al. (2005) left the region below 500 Hz acoustic and delivered the region above it with a five-channel sine vocoder. With a gap of 0.5 kHz between them, sentence recognition was 90%, with 1 kHz 80%, with 1.5 kHz 58% and with 2.2 kHz 45%; the acoustic part alone stayed at 17%. In noise the combined score was higher than the sum of the two parts' separate scores.

    What does this tool not imitate?

    Cychosz et al. (2024) stress that vocoders should not be regarded as "real implant simulations". What cannot be imitated: the very narrow dynamic range of electrical stimulation and the high synchrony in the nerve fibres, spread being broader at the apex than at the base, the electrode-neuron interface varying from person to person, and the adaptation and learning implant users show, especially in the first 6-12 months. The Greenwood map is based on places along the organ of Corti; the frequency map of the spiral ganglion cells that electrodes stimulate departs from it, especially towards the apex.

    How is it calculated in this tool?

    The bands are placed at equal cochlear distances on the Greenwood map (A = 165.4, a = 2.1, k = 0.88, 35 mm). The filters are zero-phase Butterworth. Current spread follows the method of Oxenham and Kreft (2014): intensity envelopes are summed with a dB/octave weight based on the octave distance between carrier centres. PSHC phases are generated with the equations of Hilkhuysen and Macherey, and pulse rates with the formula of Mesnildrey et al. (37 + 151x + 0.17x², x in kHz). The whole calculation is verified by a test suite of 72 measurements: filter slopes, the Greenwood transform, envelope accuracy, PSHC pulse rate and crest factor, spread weights, SNR and level matching.

    References
    1. Cychosz, M., Winn, M. B., & Goupell, M. J. (2024). How to vocode: Using channel vocoders for cochlear-implant research. The Journal of the Acoustical Society of America, 155(4), 2407-2437. doi:10.1121/10.0025274
    2. Dorman, M. F., Spahr, A. J., Loizou, P. C., Dana, C. J., & Schmidt, J. S. (2005). Acoustic simulations of combined electric and acoustic hearing (EAS). Ear and Hearing, 26(4), 371-380.
    3. Dudley, H. (1939). The automatic synthesis of speech. Proceedings of the National Academy of Sciences of the United States of America, 25(7), 377-383. doi:10.1073/pnas.25.7.377
    4. Friesen, L. M., Shannon, R. V., Başkent, D., & Wang, X. (2001). Speech recognition in noise as a function of the number of spectral channels: Comparison of acoustic hearing and cochlear implants. The Journal of the Acoustical Society of America, 110(2), 1150-1163. doi:10.1121/1.1381538
    5. Greenwood, D. D. (1990). A cochlear frequency-position function for several species: 29 years later. The Journal of the Acoustical Society of America, 87(6), 2592-2605.
    6. Hilkhuysen, G., & Macherey, O. (2014). Optimizing pulse-spreading harmonic complexes to minimize intrinsic modulations after auditory filtering. The Journal of the Acoustical Society of America, 136(3), 1281-1294. doi:10.1121/1.4890642
    7. Karoui, C., James, C., Barone, P., Bakhos, D., Marx, M., & Macherey, O. (2019). Searching for the sound of a cochlear implant: Evaluation of different vocoder parameters by cochlear implant users with single-sided deafness. Trends in Hearing, 23, 1-15. doi:10.1177/2331216519866029
    8. Kopsch, A. C., Plontke, S. K., & Rahne, T. (2025). Acoustic simulation of cochlear implant sound to approximate the perceptual experience of electric hearing. Scientific Reports, 15, 38997. doi:10.1038/s41598-025-25711-z
    9. Mesnildrey, Q., Hilkhuysen, G., & Macherey, O. (2016). Pulse-spreading harmonic complex as an alternative carrier for vocoder simulations of cochlear implants. The Journal of the Acoustical Society of America, 139(2), 986-991. doi:10.1121/1.4941451
    10. Oxenham, A. J., & Kreft, H. A. (2014). Speech perception in tones and noise via cochlear implants reveals influence of spectral resolution on temporal processing. Trends in Hearing, 18, 2331216514553783. doi:10.1177/2331216514553783
    11. Shannon, R. V., Zeng, F.-G., Kamath, V., Wygonski, J., & Ekelid, M. (1995). Speech recognition with primarily temporal cues. Science, 270(5234), 303-304. doi:10.1126/science.270.5234.303

    Speech samples: three talkers from the Mozilla Common Voice English dataset (version 17.0, CC0 public domain dedication). The babble was made by overlaying six talkers from the Common Voice Turkish dataset (version 13.0, CC0). The recordings were trimmed and brought to the same active level. The melody (Ah vous dirai-je, maman) and the music (Minuet in G major, BWV Anh. 114, a simple two-voice arrangement) are public-domain tunes synthesised in this page.

    Try it in the simulator

    © 2026 Ahmet Alperen Akbulut, Auditory Scene. All rights reserved. It may not be copied, reproduced or distributed without permission.This covers the simulator's software, sound processing model and images. For permission requests: info@isitmeatolyesi.com