Upload your audio file
Choose an MP3, WAV, FLAC, or OGG file from your device.
Free Online Pitch Shifter — Change Song Key & Transpose Audio
Upload a song, shift the pitch up or down, preview the result, and export your edited audio as MP3 or WAV. No signup required. This pitch shifter online runs in your browser.
Drag and drop your audio here, or click to browse. Supports MP3, WAV, FLAC, OGG, and M4A input files.
🎵 Want slowed + reverb instead? Use the main Slowed Reverb Generator.
Open Generator →Choose an MP3, WAV, FLAC, or OGG file from your device.
Move the pitch control up or down in semitones to raise or lower the song key.
Listen before exporting so you can check if the vocals and instruments still sound natural.
Export your pitch-shifted track as MP3 or WAV.
| Use Case | Pitch Shift | Notes |
|---|---|---|
| Lower vocals slightly | -1 to -3 semitones | Deeper vocals and darker edits. Below about -3 a voice starts reading as an effect rather than a singer. |
| Raise vocals slightly | +1 to +3 semitones | Brighter vocals and pop edits. The most forgiving range, because the formants have not moved far enough to sound synthetic. |
| Nightcore-style edit | +3 to +6 semitones | The classic bright, fast edit. The tempo rise here is the point rather than a side effect. |
| Match a singer's range | Whatever the interval is | Count semitones from the original key to your target key and enter that. A fourth is 5, a fifth is 7. |
| Fix a slightly out-of-tune recording | -0.5 to +0.5 semitones | 0.1 semitone is 10 cents. Use the fine steps here rather than the presets. |
| Full octave | -12 or +12 semitones | Doubles or halves both the frequency and the playback rate. Dramatic, and rarely natural on a full mix. |
Small pitch changes usually sound more natural. Large shifts can create artifacts, especially on vocals or complex mixes.
Because this is a varispeed shift, every semitone you move also moves the playback rate. The relationship is exact: n semitones sets the rate to 2n/12, and the running time is the original divided by that. Nothing here is an estimate — the column on the right is a 3:30 track put through the arithmetic.
| Shift | Playback rate | 3:30 becomes | Difference |
|---|---|---|---|
| -12 st | 0.500x | 7:00 | Twice as long |
| -5 st | 0.749x | 4:40 | 70 seconds longer |
| -3 st | 0.841x | 4:09 | 40 seconds longer |
| -1 st | 0.944x | 3:42 | 12 seconds longer |
| 0 st | 1.000x | 3:30 | Unchanged |
| +1 st | 1.059x | 3:18 | 12 seconds shorter |
| +3 st | 1.189x | 2:56 | 33 seconds shorter |
| +5 st | 1.335x | 2:37 | 53 seconds shorter |
| +12 st | 2.000x | 1:45 | Half as long |
The practical consequence: a shift big enough to be interesting is also big enough to be obvious in the tempo. At ±1 to ±3 semitones the timing change is small enough that most listeners will not register it as a speed edit. Past about ±5 they will.
This is a varispeed shift — pitch and tempo move together, like changing the speed of a tape. Most free shifters do the same thing without telling you, which is why your export came out shorter than the original.
−12 to +12 in 0.1 steps, so +7 is a fifth and +12 is an octave. You can transpose to a target key instead of nudging until it sounds about right.
0.1 semitone is 10 cents. That is fine enough to nudge a recording into tune with another one rather than only moving it in musical intervals.
The playback rate is exactly 2^(n/12) for n semitones, so the change in running time is arithmetic, not a surprise. The table below gives it for every common shift.
Large shifts fall apart on vocals long before they do on a synth line. Hearing it on your actual track beats guessing from a number.
Lowering the pitch makes the track longer — a −12 st shift doubles it. The render is sized from the actual combined rate, so the extra length is really there instead of being cut off at the original duration.
A human voice carries two separate kinds of pitch information. There is the fundamental — the rate the vocal folds vibrate at, which is the note being sung. And there are the formants — resonant peaks created by the fixed size and shape of the singer’s throat, mouth, and nasal cavity. The formants are what make a voice recognisable as that person, and as an adult rather than a child. Crucially, they do not move when a singer changes note: a soprano and a bass singing the same pitch still sound like different people, because their resonant cavities are different sizes.
Varispeed cannot tell those two apart. Speeding the audio up scales every frequency by the same factor, so the formants rise with the fundamental. The result implies a singer whose head physically shrank, which is not a thing that happens, and your ear knows it. That is the chipmunk effect — not distortion or a quality problem, but an acoustic contradiction. The same thing runs in reverse going down: lower the pitch far enough and the formants imply a vocal tract the size of a doorway, which is exactly the cavernous quality slowed edits are after.
This is why the tolerance is so much narrower on voices than on instruments. A synth pad or a guitar has no formant structure your ear is auditing against a mental model of a human body, so it can take ±12 and still sound intentional. A lead vocal usually starts giving the game away somewhere past ±3. If you are shifting a full mix, the vocal is the part that will break first — judge the setting by it, not by the drums.
There are two ways to move the pitch of a recording, and they fail in different directions.
Varispeed — what this page does — is the tape approach: play the samples faster or slower. Pitch and tempo move together because they are the same operation. Its great advantage is that it is exact. No sample is invented, nothing is spliced, and the waveform that comes out is the waveform that went in, read at a different rate. It cannot produce artefacts, because it is not making anything up. The price is that you do not get to choose whether the tempo comes along.
Time-stretching is the other approach: cut the audio into short overlapping windows and lay them back down at a different spacing, so the running time changes without the pitch following. Combine that with a varispeed shift of the same ratio and the tempo change cancels, leaving pitch alone — pitch shifting with the tempo preserved. The price is the mirror image: the algorithm is now deciding where to overlap and crossfade, and on transient-heavy material — drums, plucked strings, consonants — those decisions are audible as flamming or smearing.
So the honest answer to “which is better” is that varispeed is better when you want a slowed or nightcore edit, where the tempo change is part of the effect, or when you want a mathematically clean result. Time-stretching is better when the tempo genuinely must not move — transposing a backing track for a singer, or matching two songs for a mix.
This site does have the tempo-preserving path: it is built into the 432 Hz converter, which offers both engines side by side and reports the tempo cost of each on your file. It is organised around reference-pitch targets rather than arbitrary semitones, though, so it covers a fixed set of intervals rather than the full ±12 range here. For deliberate tempo work, use the audio speed changer; for the two combined into a finished edit, the slowed reverb generator or the nightcore maker.
Shift the tuning reference itself rather than transposing by whole semitones.
Slow songs and add reverb online.
Slow down or speed up songs online.
Make fast, high-pitched nightcore edits online.
Increase low-end bass in songs and audio files.
Master pitch shifting, semitones, key changes, and audio transposition.
In this tool, yes. The shift is varispeed, so raising the pitch speeds the track up and lowering it slows the track down. The rate is exactly 2^(n/12) for n semitones — +12 plays at double speed and halves the running time, -12 plays at half speed and doubles it. This is stated plainly because many free pitch shifters do the same thing and describe themselves as though they do not.
Not on this page. Doing so requires time-stretching, which cuts the audio into overlapping windows and re-spaces them, and that introduces its own artefacts on drums and consonants. The site does offer that path in the 432 Hz converter, which runs both engines and shows the tempo cost of each — but it is built around reference-pitch targets rather than arbitrary semitone amounts, so it covers a fixed set of intervals.
For a result that still sounds like a recording rather than an effect, ±1 to ±3. Past about ±5 on a vocal, the formants have moved far enough that the voice implies a differently sized human and the ear notices. Instruments tolerate much more than voices. If you are transposing to a specific key rather than chasing a sound, use the actual interval: a fourth is 5 semitones, a fifth is 7, an octave is 12.
Because varispeed scales every frequency by the same factor, including the formants — the resonances set by the physical size of the singer's throat and mouth. Those normally stay put when a singer changes note, so moving them implies a vocal tract that changed size, which your ear reads as unnatural. It is not distortion or a quality fault; it is what uniform frequency scaling does to a voice. Smaller shifts keep it below the threshold where it registers.
Pitch is how high or low a single sound is. Key is the tonal centre the whole piece is organised around. Shifting every note by the same number of semitones moves the key by that interval while leaving the relationships between notes intact — which is why the song still sounds like itself, just higher or lower. Shifting by a non-integer amount, such as 0.5 semitones, does not land in a recognised key at all; that setting is for tuning corrections, not transposition.
Yes — the control moves in 0.1 semitone steps, which is 10 cents. That resolution is aimed at tuning rather than transposition: nudging a recording into agreement with another one, or correcting a source that was not at standard pitch to begin with. For reference, the gap between A=440 and A=432 is about 32 cents, a little over three of these steps.