Vocal–Beat Balance Check

← All tools
Runs in your browserNo uploadLevel + masking in dBFree

Vocal–Beat Balance Check: How Loud Should Vocals Be in a Mix?

Is your vocal too quiet, too loud, or just masked? Drop the vocal and the beat, hear them together, get a number — not a feeling. The tool measures the level gap where the vocal is actually singing, how hard the beat fights it in the 1.5–5 kHz presence band, whether the low end clashes, and how much the phrases jump around. Everything runs locally with the Web Audio API; nothing is uploaded.

Compare vocal and beat
Vocal Drop or choose the vocal Dry or processed lead vocal · WAV, MP3, M4A, OGG, FLAC · first 60 s used
Beat / instrumental Drop or choose the beat The instrumental without the vocal · same formats · first 60 s used

Both files start at 0:00. Positive = the vocal starts later. Use it if your exports do not line up.

0.0 dB

Applied to playback and to the numbers below, so you can find the gain that lands in the window.

Nothing runs until you choose two files or tap the demo.

Nothing from your files leaves the browser. Decoding, analysis and playback use the Web Audio API on your device; there is no server, no microphone access and no tracking. The only network request this page makes is fetching the two demo clips when you tap "Load demo pair".

Reading the four numbers

Metric What it measures Starting window If it is outside
Level gap Vocal RMS minus beat RMS, only in the 400 ms frames where the vocal is active (within 35 dB of its loudest frame). 0 to +4 dB. Rap and pop often +2 to +5, rock and EDM 0 to +2. Below 0: raise the vocal or lower the beat. Above +4: the vocal floats on top; pull it in or glue it with reverb and compression.
Masking Beat energy minus vocal energy in 1.5–5 kHz during vocal frames (2048-point FFT, Hann window, averaged). Below −3 dB (vocal clearly wins the presence band). At −3 dB or above: cut the beat 2–3 dB around the reported frequency with a wide bell, or sidechain-duck it.
Low-end clash Vocal energy 60–200 Hz relative to its own 200–1000 Hz body. Below −10 dB clean; −10 to −4 dB watch it. At −4 dB or above: high-pass the vocal higher (100–150 Hz) so kick and bass get their room back.
Dynamics Spread of phrase RMS levels (P90 − P10; max − min when there are fewer than five phrases). Phrases are active runs separated by gaps over 250 ms. 6 dB or less. Over 6 dB: ride clip gain phrase by phrase or compress before touching the fader; a single fader position cannot be right for all of them.

All levels are RMS in dBFS over 400 ms windows with a 100 ms hop. Nothing here is LUFS or a loudness-normalized measurement; the point is the difference between the two files, not an absolute target.

How loud should vocals be in a mix?

The honest answer to "how loud should vocals be in a mix" is a range, not a number: in most finished records the lead vocal sits roughly 0 to +4 dB above the instrumental when you compare RMS level only in the moments the singer is actually singing. Rap and modern pop lean toward the top of that window and sometimes beyond it; rock, indie and EDM vocals sit lower and let the band or the drop carry. Nobody can give you one figure because the fader is only one of three things that decide whether a vocal sounds "in" the mix; the other two, spectral masking and phrase-to-phrase dynamics, have nothing to do with the fader at all. This tool measures all three so you know which one you are actually fighting.

Vocal not sitting in the mix: level, masking or dynamics?

When a vocal is not sitting in the mix, most people reach for the fader. Sometimes that is right: if the level gap reads −3 dB, the vocal is simply quieter than the beat and no EQ will fix it. But the same complaint has two other causes that feel identical in the moment. The first is masking: the beat has strong energy in the 1.5–5 kHz band where consonants and vowel formants live, so the words are covered even when the vocal is louder overall. You push the fader, the vocal gets loud and still sounds buried, so you push again and now it floats on top, too loud and still unclear. The second is dynamics: the chorus phrase is 8 dB hotter than the verse phrase, so any fader position is wrong for half the song. All three produce the same symptom at the speakers, which is why guessing fails and measuring helps.

Rule of thumb: if raising the vocal 2 dB makes it louder but not clearer, you have a masking problem. If it is clear in the hook and lost in the verse, you have a dynamics problem. Only if it gets clearer as it gets louder is it really level.

Masking vs level: why the presence band decides intelligibility

Speech intelligibility depends mostly on the region between about 1.5 and 5 kHz: the second and third vowel formants, the start of T, K and S, and the range where the ear is most sensitive. A vocal that owns this band can be a couple of dB under the beat and still read every word; one that loses it can be 4 dB on top and still sound behind a curtain. Beats fight for this band more than producers expect — snare crack, hi-hats, guitar and synth upper harmonics, and especially pads with a lot of 2–4 kHz content. The masking number this tool reports is the beat's energy minus the vocal's energy in that band, averaged only over frames where the vocal is active. Below about −3 dB the vocal clearly wins; at −3 dB or above they are fighting, and the fix is a small wide cut on the beat around the reported frequency, typically 2 to 3 dB, or a sidechain or dynamic EQ that ducks the beat only while the vocal is present. That is a far smaller move than the 4 dB fader push people make instead, and it leaves the beat at full energy between the lines.

Why the fader comes last

A fader sets one level for the whole track. If the phrases differ by more than about 6 dB, no single position works, and a compressor asked to close that gap alone will pump, dull the consonants and bring up breaths. The order that works: ride clip gain first so each phrase lands within roughly ±1.5 dB of the others, then compress a modest 3–6 dB for tone and density, then set the fader once. When the dynamics number here is above 6 dB, the fix is upstream of the fader and the Vocal Clip Gain Planner will write the move list for you.

Low end: the clash you hear as "muddy"

A vocal recorded close to a condenser carries a lot of energy below 200 Hz from proximity effect, room rumble and the fundamental itself. Alone it sounds warm; against a kick and an 808 it sounds muddy, because two sources are spending headroom on the same 60–200 Hz region. The tool compares the vocal's 60–200 Hz energy with its 200–1000 Hz body; when that ratio is high, a high-pass at 100–150 Hz with a 12 dB/octave slope usually fixes it and lets the vocal sit louder without the mix getting thicker.

How to use the result

  • Load the real pair. Export the vocal with its processing and the instrumental without it, both from 0:00. If they do not line up, set the offset until the overlay looks right.
  • Use the slider, not your DAW. Drag the vocal level until the gap reads about +2 dB and listen with "Play mix". The slider value is the fader move you need.
  • Hear the fix before you make it. "Play mix with suggested fix" applies the gain that lands in the window and, if the presence band is contested, a −3 dB wide bell on the beat at the reported frequency.
  • Check at three volumes. A balance that holds at a whisper, at normal level and loud is done; one that only works loud is masking you have not fixed yet.

FAQ

How loud should vocals be compared to the beat?

As a starting window, 0 to +4 dB above the beat's RMS level, measured only while the vocal is active. Rap and pop typically sit at the hot end or slightly above; rock and EDM at the quiet end. Genre, arrangement density and how much reverb is on the vocal all move the target, so use the number to compare against a reference you like, not as a rule.

Does this tool upload my vocal or beat?

No. Both files are decoded and analyzed in your browser with the Web Audio API and never leave your device. There is no server-side step, no account and no microphone access. The only network request is fetching the two demo clips if you tap "Load demo pair".

Why does my vocal still sound buried when it is louder than the beat?

That is masking. The beat has as much or more energy than the vocal in the 1.5–5 kHz band where words are understood, so turning the vocal up makes it louder without making it clearer. Cut the beat 2–3 dB with a wide bell at the frequency the tool reports, or duck that band from the vocal with a sidechain or dynamic EQ.

What is a good masking number?

Below −3 dB means the vocal owns the presence band. Between −3 and 0 dB the two are roughly equal and the vocal will sound a little veiled; above 0 dB the beat wins and words will drop out, especially on small speakers. The exact threshold depends on how dense the arrangement is, which is why the tool lets you hear the fix instead of only reporting it.

Is the level gap the same as LUFS?

No. The tool compares RMS level in dBFS over 400 ms windows, not LUFS or any loudness-normalized measurement. RMS is close enough to compare two files that share the same frames, but it does not model the ear's frequency weighting, so a bright beat reads a little quieter than it feels and a bass-heavy one a little hotter.

My vocal and beat exports do not start at the same time. What do I do?

Type the difference into the offset field: a positive number delays the vocal, a negative one pulls it earlier. The overlay on the waveform shows the alignment, and the analysis only uses the parts where both files overlap, so a rough offset within 50 ms or so is fine for level and masking.

What if the dynamics number is high but the level gap is fine?

The average is fine but individual phrases are not: a 6 dB or larger spread means the quiet lines will disappear and the loud ones will jump out, whatever the fader does. Ride clip gain phrase by phrase first, then compress lightly, then set the fader. The Vocal Clip Gain Planner on this site turns the same phrases into a move list.

Can I use a mastered 2-track beat?

Yes, but read the numbers with that in mind. A heavily limited beat has a high RMS level for its perceived loudness, so the level gap will read lower than it sounds and the tool may say "too quiet" when a dB less would do. Masking and dynamics are unaffected. If you cannot EQ the beat, a dynamic EQ or sidechain on the beat bus keyed from the vocal does the same job.

Streams follow sound

Know the problem. Now copy the chain that solves it.

The Vocal Chain Bible is a 187-page PDF with 88 real vocal chains from hit records, as reported by the engineers in public interviews — mic, preamp, plugins and settings, each chain also in a free-plugin version, plus the full Vocal Chain Maker app and Rap and Melodic Rap DAW templates. Once you know whether it is level, masking or dynamics, copy a chain whose compression and EQ were built to make a vocal sit, instead of guessing at the fader.

88 chains · every chain in a free-plugin version · Try one free chain first
Copied

Free download

Get the free Grammy Sauce PDF

The plugins and settings 100+ Grammy-winning engineers actually use — collected from 5 years of interviews, in one free PDF.

Want the full pro setup? Get the Vocal Chain Bible — 88 real chains →

Contact form