Why Autotune Makes Your Esses Worse and How to Fix It
Share
Vocal tuning · 4 min read
A pitch corrector decides what note a sound is many times a second. An S has no note — it is broadband noise — so the tuner shifts and smears it anyway, and you get a longer, brighter, more artificial ess. Fix it by de-essing 2–4 dB before the tuner, not after it.
Fix it here
BEFORE the tuner
2–4 dB of de-essing first · 2 dB after
The 60-second answer
Move a de-esser to the top of the chain, ahead of the tuner, and pull 2–4 dB at the sibilance frequency. Raise retune speed from 0 to 20–40 wherever the part allows it. Then add a second small de-esser after the tuner, aimed 1–2 kHz higher, doing no more than 2 dB.
What a tuner does to an unvoiced sound
Pitch correction detects a fundamental and shifts the signal to move it. Voiced sounds have a fundamental. S, Sh, T, F and Ch do not — they are filtered noise. So the detector either holds the pitch it last saw or jumps around, and the shifter applies that decision to the noise regardless. What comes out is noise that has been stretched, resampled and phase-modulated, which the ear reads as a longer, thinner, spikier S than the one you recorded.
Why retune speed makes it worse
At retune speed 0 the corrector snaps to the target the moment it thinks the note changed, and that moment is often the consonant at the front of the word. So the transition into the note gets rewritten together with the S in front of it. At 20–40 the correction ramps in over a few tens of milliseconds, long enough for most consonants to pass through mostly unprocessed. That is why the hard-tuned sound and crunchy esses tend to arrive together.
Fixing tuner sibilance, step by step
| # | Do this |
|---|---|
| 01 |
Prove it is the tuner Bypass the corrector and loop one sibilant line. If the S gets shorter and softer, the tuner is the source and no de-esser after it fully fixes that. A/B: tuner bypassed |
| 02 |
Find the frequency on the raw take Sweep a narrow +12 dB bell through 4–12 kHz with the tuner off, so you are aiming at the real voice and not the artefact. Bell +12 dB, Q 8 |
| 03 |
Put the de-esser first, before the tuner Pull 2–4 dB. Less sibilant energy going in means fewer wrong pitch decisions and less noise for the shifter to smear. de-esser → tuner |
| 04 |
Raise retune speed where the song allows Use 0 only when the hard effect is the point. 20–40 is still obviously tuned and much kinder to consonants. Retune 20–40 |
| 05 |
Tune exposed lines offline instead Graphical tuners work note by note and leave the unvoiced noise between notes largely alone, so an intro or a bridge survives much better done by hand. Offline for exposed parts |
| 06 |
Second de-esser after the tuner Aim it 1–2 kHz above the first, because artefact energy sits higher than the natural S. If a whistle remains, notch it at Q 8. 2 dB max, 8–11 kHz |
The trick: de-ess into the tuner, not out of it
Once the corrector has processed an S, the damage is inside the noise: the length changed, the phase changed, and the peak moved up in frequency. A de-esser after that can only turn the artefact down, and turning it down far enough to hide it takes 6–8 dB, which is a lisp. Put 3 dB of de-essing in front of the tuner instead and you change what the algorithm sees. The pitch tracker has less broadband noise confusing it, so it makes fewer wrong decisions on consonants, and there is less sibilant energy for the shifter to stretch. The order that works on nearly every rap and pop vocal is trim, de-esser, tuner, compressor, saturation, then a second 2 dB de-esser at the end. Same total reduction, split across two places, and the esses stop sounding like plastic.

Stop guessing the chain.
The Vocal Chain Bible is 88 real vocal chains from Grammy-winning engineers — every plugin, every setting, in order, with free alternatives for each one. 5+ years, 100+ interviews, one PDF.
Get the Vocal Chain Bible — $48 →Mistakes that kill it
- De-essing only at the end of the chain. You are treating the symptom with a plugin that cannot reach the cause.
- Retune speed 0 on every line. Fine on a hook that wants the effect, rough on a verse full of S and T sounds.
- Using a high shelf cut instead. It is down for the whole song and takes the air with it, for a problem that lasts 100 ms per word.
- Tuning a take recorded 5 cm from a bright condenser. Nothing downstream fixes that. Step back to 15–20 cm and angle the mic off the mouth.
Keep going
FAQ
Should I de-ess before or after autotune?
Before, mostly. 2–4 dB in front of the tuner does more than 6 dB after it. Add a small second stage later in the chain for what the saturation creates.
Does Melodyne cause the same problem?
Less of it. Graphical tuning works note by note and leaves unvoiced noise between notes largely alone, which is why exposed lines are usually better tuned offline.
Why does my S whistle after tuning?
The shifter narrowed the noise and left a tonal peak, often between 8 and 11 kHz. A narrow dynamic band at Q 8 removes it without touching the air.

Who's behind this
Ivan Meshcheriakov · IvanFromRnD
5+ years ago I moved to LA with zero connections and started building from nothing — 14-hour days, meeting people and interning at studios. It took me 5+ years to meet and interview 100+ Grammy-winning engineers — then I turned it all into the Vocal Chain Bible.
Read my full story →