Audio Pitch & Bass Engineering Guide: Pitch Shifting, Transposing & Bass Boost Online

Data engineer who loves building high-performance data and web-related tools. Creator of SlowedReverbMaker.net, implementing browser-side digital signal processing (DSP) to democratize audio editing.
Pitch and bass are the two controls people reach for after speed, and they are the two most likely to produce a result that is worse than where you started.
Not because they are difficult. Because each one has a constraint that is invisible from the interface, and every guide on the internet hands out numbers as though the constraint did not exist. Pitch is not independent of tempo on most browser tools, which means transposing a track also moves its speed. And how much bass a file can absorb is a property of that specific file, not of your headphones — which makes the device-by-device dB charts everyone publishes, including the earlier version of this page, confidently wrong.
This guide covers what each control actually does, what limits it, and how to work out the right number for your track rather than copying one from a table.
Try this while you read: Bass Booster
Boost low-end and sub-bass online. Free, in your browser, no signup.
1. Pitch Shifting: What It Really Does Here
Pitch shifting moves a recording up or down in semitones. A semitone is one half-step — the distance from C to C sharp, or from one fret to the next on a guitar. Twelve semitones is an octave, which doubles or halves the frequency.
The critical detail is how the shift is achieved, because there are two methods with completely different consequences, and tools rarely say which they use.
The pitch shifter on this site uses varispeed — the tape approach. It reads the samples faster or slower, which raises or lowers every frequency together. Pitch and tempo move as one, because they are the same operation. Raising the pitch speeds the track up; lowering it slows the track down.
The alternative is time-stretching, which cuts audio into short overlapping windows and re-spaces them so the running time changes without the pitch following. Combine that with a varispeed shift of the same ratio and the tempo change cancels, leaving pitch alone.
Neither is better in general. Varispeed is mathematically exact — no sample is invented, nothing is spliced, so it cannot produce artefacts. Time-stretching preserves tempo but has to make decisions about where to overlap, and on transient-heavy material those decisions are audible as smearing or flamming.
1a. What Each Shift Costs in Running Time
Because this is varispeed, every semitone has a tempo consequence, and it is exactly calculable: n semitones sets the playback rate to 2 raised to the power n/12. Applied to a 3:30 track:
- −12 st — 0.500x, running time 7:00. Twice as long.
- −5 st — 0.749x, running time 4:40.
- −3 st — 0.841x, running time 4:09.
- −1 st — 0.944x, running time 3:42.
- +1 st — 1.059x, running time 3:18.
- +3 st — 1.189x, running time 2:56.
- +5 st — 1.335x, running time 2:37.
- +12 st — 2.000x, running time 1:45. Half as long.
1b. The Practical Consequence
At ±1 to ±3 semitones the timing change is small enough that most listeners will not register it as a speed edit. Past about ±5 they will, immediately.
This matters most for the use case people most often want pitch shifting for: transposing a backing track to fit a singer's range. That job specifically requires the tempo to stay put, and varispeed cannot do it — a track dropped three semitones to suit a lower voice also arrives 19% slower, which no singer will thank you for. If tempo must not move, you need a time-stretching tool. The 432 Hz converter on this site does run that engine, and shows the tempo cost of each approach on your file, but it is organised around reference-pitch targets rather than arbitrary semitones.
For creative edits — deeper slowed versions, brighter nightcore-style ones — varispeed is not a compromise at all. The tempo change is part of the effect you were after.
2. Why Big Shifts Sound Artificial
There is a hard limit on how far a voice can be shifted before it stops sounding like a person, and it is worth knowing why, because it tells you the limit is real rather than a quality problem you could solve with a better tool.
A voice carries two kinds of pitch information. The fundamental is the note being sung. The formants are resonant peaks created by the fixed size of the singer's throat, mouth, and nasal cavity — and they are what make a voice recognisable as that individual, and as an adult rather than a child. Formants do not move when a singer changes note; a bass and a soprano on the same pitch still sound like different people.
Varispeed scales every frequency by the same factor, so the formants move with the fundamental. Shift up and you imply a singer whose head shrank. Shift down and you imply a vocal tract the size of a doorway — which is exactly the cavernous quality deep slowed edits are chasing, so the same mechanism that ruins one effect creates another.
Instruments tolerate far more than voices, because a synth or a guitar has no formant structure your ear is auditing against a mental model of a human body. A pad can take ±12 and still sound intentional. A lead vocal usually gives the game away past about ±3. If you are shifting a full mix, the vocal is what breaks first — judge by it, not by the drums.
3. Bass: Why Device Charts Are the Wrong Model
Now the second control, and the more commonly mishandled one.
Every bass guide, including the previous version of this page, publishes a table like: earbuds +3 to +4 dB, over-ear +4 to +6 dB, car with subwoofer +6 to +9 dB. It looks authoritative and it is built on a mistaken premise — that the safe amount of boost is a property of what you are listening on.
It is not. Boosting is addition, not redistribution. Whatever you add has to fit under the digital ceiling at 0 dBFS, and how much room is available is a property of the file. The gap between a track's current peak and that ceiling is its headroom, and it is the only quantity that determines whether a boost will survive.
A quiet, dynamic recording peaking at −12 dBFS can absorb a great deal. A modern master pushed to −0.3 dBFS can absorb almost nothing, and the boosted low end gets flattened against the ceiling the moment the kick lands. Same setting, same headphones, opposite outcomes — because the variable that mattered was never the headphones.
Measured on a signal peaking at −0.45 dBFS, a +6 dB low shelf produced a peak of +5.39 dBFS. There is no such thing as +5.39 dBFS in a finished file; everything above zero is removed on export. That is not the tool failing. That is a boost the file never had room for.
3a. What Your Playback Device Does Determine
Devices are not irrelevant — they just answer a different question. They determine what you can hear, not what the file can take.
A phone speaker is a driver a few millimetres across and physically cannot move enough air to reproduce the range a bass shelf operates in. The boost is present in the file; the speaker will not render it. This creates a specific and common failure: you push the boost further and further trying to hear an effect the hardware was never going to produce, and the result is enormous everywhere else.
So the two rules are separate and both matter. How much you can boost is set by the file's headroom. Whether you can judge it is set by your playback. Get the first from a measurement and the second from headphones or a speaker with a real low-frequency driver.
4. Working Out the Right Number
The replacement for the device chart is a two-step procedure that takes about a minute and is right for your specific track.
- Measure first. Run the file through the volume booster, which reports sample peak, true peak, integrated loudness, and remaining headroom before it changes anything.
- Treat the headroom as a budget. If the file has 3 dB left, a +3 dB shelf is roughly where the loudest bass moments start hitting the ceiling. Do not spend more than you have.
- If there is no headroom, reduce first and boost after. Boosting into a wall and then turning the clipped result down does not undo the flattening — the damage is already in the samples.
- Leave extra margin for MP3. The waveform reconstructed between samples can exceed the sample values, so a file that measures safe can still overshoot on playback. A decibel of caution costs nothing audible.
- Judge on headphones, then verify where it will be heard. Dial it in somewhere honest, then check the car or the phone — but do not set it there.
5. What the Shelf Actually Touches
One more piece of information that most bass tools withhold and that changes how you use them: where the boost is centred.
The bass booster here applies a low shelf with its corner at 150 Hz. A shelf lifts everything below the corner by roughly the amount you set, with a transition band around it, and leaves the region well above it alone. That is different from a bell or peak filter, which picks out one band and leaves the frequencies either side untouched.
150 Hz sits above kick and bass fundamentals, so they get lifted fully. But it is also close to the bottom of a low male vocal, whose fundamental typically runs somewhere around 85 to 180 Hz. That overlap is why large boosts make voices sound boomy or chesty rather than just adding weight underneath them — the shelf is reaching into the vocal's body, exactly as it must, because a shelf has a transition band rather than a wall.
Knowing the corner also tells you something non-obvious about slowed edits: slowing a track lowers every frequency, so content that sat above 150 Hz can descend below it and get caught by a shelf that previously missed it. Set the speed first and judge the bass against the slowed version. A boost dialled in at normal speed usually reads as too much once the track is slowed.
6. Reasonable Starting Points
With the caveat that these are opening positions to be checked against your file's headroom, not targets:
- Light warmth — +2 to +4 dB. Suits vocals, acoustic material, and anything already well mastered.
- Hip-hop, trap, phonk — +5 to +8 dB, headroom permitting. This is where the boost is doing genre work rather than correction.
- Slowed edits — +2 to +4 dB, set after the speed. The slowdown has already moved energy downward for you.
- Sped-up and nightcore edits — +4 to +6 dB. Here the boost is genuine compensation: speeding up moved the low end out of the range that provides weight, so you are replacing something rather than adding it.
- Speech and podcasts — +1 to +3 dB. Much beyond this and you are adding rumble rather than warmth.
- Already-loud master — +1 to +3 dB, or reduce the level first. Check the headroom figure before going further.
7. Mistakes Worth Avoiding
Most bass and pitch problems come from a short list.
- Boosting a file that is already clipped, expecting improvement. EQ cannot reconstruct a flattened waveform; the information is gone.
- Using pitch to brighten a slowed edit. On this tool that also speeds the track up, quietly walking the tempo out of the range you wanted. Use speed instead.
- Applying one setting to every track. Headroom varies enormously between recordings, so a fixed boost is a gamble taken repeatedly.
- Judging bass on laptop or phone speakers. They cannot reproduce the range, so you will overshoot.
- Expecting a transposition to preserve tempo. On varispeed it never does — check the running time before you assume it worked.
- Exporting MP3 from an MP3 source when the file is heading into a video editor. That is two lossy generations before the editor adds a third.
The Short Version
Pitch shifting here is varispeed, so every semitone also moves the tempo — exactly 2^(n/12) — and voices break down past about ±3 because formants scale along with the note. If tempo must not move, varispeed is the wrong tool.
Bass is limited by your file's headroom, not by your headphones. Measure it with the volume booster, treat the number as a budget, and remember the shelf sits at 150 Hz so it will reach into a low vocal before it reaches the sub-bass. Then boost with a number you worked out rather than one you copied.
Method
The running-time table is computed directly from the varispeed relationship, rate = 2^(n/12), applied to a 210-second track. The headroom demonstration was measured by running a 60 Hz tone at amplitude 0.95 through this site's own render path with a +6 dB low shelf and analysing the result with its loudness analyser: −0.45 dBFS before, +5.39 dBFS after. The 150 Hz shelf corner and the ±12 semitone range are read from the tool's source rather than estimated. The formant explanation is standard acoustics.
An earlier version of this page recommended bass levels by playback device. That framing was wrong — the constraint is the file's headroom, not the speakers — and it has been replaced rather than adjusted.