How does a vocoder work?
A filter bank divides the sound into frequency bands. The envelope of each band, its slow amplitude change,
is extracted and the band's temporal fine structure is discarded. The envelope modulates a carrier and the
channels are summed. Because cochlear implant processors also split sound into bands and put the envelope on a pulse
train at each electrode, a vocoder is a simulation of the information an implant passes on that a normal-hearing
ear can listen to (Shannon et al., 1995).
Why does the number of channels matter so much?
Shannon et al. (1995) showed that in quiet, sentences are largely understood even with three bands. In noise
the picture changes: Friesen et al. (2001) found that normal-hearing listeners kept improving with a vocoder up to at
least 20 channels, whereas implant users stopped improving in vowel and consonant recognition beyond about 7-8
electrodes. For words and sentences the limit was 7-10 electrodes. So 22 electrodes do not mean 22 independent channels.
Choosing a carrier
Sine carries the envelope faithfully but cannot imitate spread and sounds tonal. Noise represents
spread better but adds random fluctuations that an implant does not have. PSHC is a harmonic complex whose pulse
rate in each band is chosen so that the envelope stays as flat as possible at the output of an auditory filter
(Hilkhuysen and Macherey, 2014). Mesnildrey et al. (2016) found speech reception thresholds in noise of 1.5 dB SNR for
sine, 3.8 dB for PSHC and 4.9 dB for noise.
How close is it to real implant sound?
People who use an implant because of profound hearing loss in one ear can answer this, since they can compare their
two ears. Karoui et al. (2019) reported that nine users rated PSHC as more similar than sine and noise, with no
difference between sine and noise. Kopsch, Plontke and Rahne (2025) found sine more similar than noise in fifteen users
(medians 5.9 and 3.6); in the first screening, simple low-pass, high-pass and band-pass filters were also rated as
similar as the vocoders. With individual adjustment similarity rose to 9.7 out of 10, but the adjustment needed differed
widely between people. In short, a vocoder simulates the information an implant passes on, not the sound of the implant.
Insertion depth and adaptation
The electrode array usually enters 22-30 mm of the roughly 35 mm cochlea. With shallow insertion, low frequencies are
moved basally, to places of higher frequency; for example, peaks at 500 and 1200 Hz shift to about 900 and 2000 Hz
(Cychosz et al., 2024). Users can get used to this shift over time; of Karoui et al.'s nine users, only one showed full
and two partial adaptation. A few minutes of listening by a normal-hearing listener cannot imitate this adaptation.
Electric-acoustic stimulation (EAS)
Dorman et al. (2005) left the region below 500 Hz acoustic and delivered the region above it with a five-channel sine
vocoder. With a gap of 0.5 kHz between them, sentence recognition was 90%, with 1 kHz 80%, with 1.5 kHz 58% and with
2.2 kHz 45%; the acoustic part alone stayed at 17%. In noise the combined score was higher than the sum of the two
parts' separate scores.
What does this tool not imitate?
Cychosz et al. (2024) stress that vocoders should not be regarded as "real implant simulations". What cannot be
imitated: the very narrow dynamic range of electrical stimulation and the high synchrony in the nerve fibres,
spread being broader at the apex than at the base, the electrode-neuron interface varying from person to person, and
the adaptation and learning implant users show, especially in the first 6-12 months. The Greenwood map is based
on places along the organ of Corti; the frequency map of the spiral ganglion cells that electrodes stimulate departs
from it, especially towards the apex.
How is it calculated in this tool?
The bands are placed at equal cochlear distances on the Greenwood map (A = 165.4, a = 2.1, k = 0.88, 35 mm).
The filters are zero-phase Butterworth. Current spread follows the method of Oxenham and Kreft (2014): intensity
envelopes are summed with a dB/octave weight based on the octave distance between carrier centres. PSHC phases are
generated with the equations of Hilkhuysen and Macherey, and pulse rates with the formula of Mesnildrey et al.
(37 + 151x + 0.17x², x in kHz). The whole calculation is verified by a test suite of 72 measurements: filter slopes,
the Greenwood transform, envelope accuracy, PSHC pulse rate and crest factor, spread weights, SNR and level matching.
References
- Cychosz, M., Winn, M. B., & Goupell, M. J. (2024). How to vocode: Using channel vocoders for cochlear-implant research. The Journal of the Acoustical Society of America, 155(4), 2407-2437. doi:10.1121/10.0025274
- Dorman, M. F., Spahr, A. J., Loizou, P. C., Dana, C. J., & Schmidt, J. S. (2005). Acoustic simulations of combined electric and acoustic hearing (EAS). Ear and Hearing, 26(4), 371-380.
- Dudley, H. (1939). The automatic synthesis of speech. Proceedings of the National Academy of Sciences of the United States of America, 25(7), 377-383. doi:10.1073/pnas.25.7.377
- Friesen, L. M., Shannon, R. V., Başkent, D., & Wang, X. (2001). Speech recognition in noise as a function of the number of spectral channels: Comparison of acoustic hearing and cochlear implants. The Journal of the Acoustical Society of America, 110(2), 1150-1163. doi:10.1121/1.1381538
- Greenwood, D. D. (1990). A cochlear frequency-position function for several species: 29 years later. The Journal of the Acoustical Society of America, 87(6), 2592-2605.
- Hilkhuysen, G., & Macherey, O. (2014). Optimizing pulse-spreading harmonic complexes to minimize intrinsic modulations after auditory filtering. The Journal of the Acoustical Society of America, 136(3), 1281-1294. doi:10.1121/1.4890642
- Karoui, C., James, C., Barone, P., Bakhos, D., Marx, M., & Macherey, O. (2019). Searching for the sound of a cochlear implant: Evaluation of different vocoder parameters by cochlear implant users with single-sided deafness. Trends in Hearing, 23, 1-15. doi:10.1177/2331216519866029
- Kopsch, A. C., Plontke, S. K., & Rahne, T. (2025). Acoustic simulation of cochlear implant sound to approximate the perceptual experience of electric hearing. Scientific Reports, 15, 38997. doi:10.1038/s41598-025-25711-z
- Mesnildrey, Q., Hilkhuysen, G., & Macherey, O. (2016). Pulse-spreading harmonic complex as an alternative carrier for vocoder simulations of cochlear implants. The Journal of the Acoustical Society of America, 139(2), 986-991. doi:10.1121/1.4941451
- Oxenham, A. J., & Kreft, H. A. (2014). Speech perception in tones and noise via cochlear implants reveals influence of spectral resolution on temporal processing. Trends in Hearing, 18, 2331216514553783. doi:10.1177/2331216514553783
- Shannon, R. V., Zeng, F.-G., Kamath, V., Wygonski, J., & Ekelid, M. (1995). Speech recognition with primarily temporal cues. Science, 270(5234), 303-304. doi:10.1126/science.270.5234.303
Speech samples: three talkers from the Mozilla Common Voice English dataset
(version 17.0, CC0 public domain dedication). The babble was made by overlaying six talkers from the Common Voice
Turkish dataset (version 13.0, CC0). The recordings were trimmed and brought to the same active level. The melody
(Ah vous dirai-je, maman) and the music (Minuet in G major, BWV Anh. 114, a simple two-voice arrangement)
are public-domain tunes synthesised in this page.