
Spotify deletes big chunks of a song, bets you won't notice. Your ears are the alibi
A CD carries 1,411 kilobits of sound every second, while Spotify and other streaming services send a fraction of that. This instalment of Science of Sound explains masking, the quirk of human hearing that lets a machine delete most of a song's data without anyone noticing.

Take any carriage of the last Metro out of Rajiv Chowk. It is nearly midnight, the doors sigh, the rails hum a low, steady note beneath everything. A woman by the window has closed her eyes, one earphone in, a song she has loved since college, folded into the rattle of the train. The song reaching her is mostly absent.
Somewhere between a studio and that earphone, a machine went through the song moment by moment and removed most of the information a CD would have carried, so carefully that she would never miss it.
In this instalment of Science of Sound, we follow a song from the studio to her ear, and find out what was removed and why her own ears let it get away with it.
THE FILE THAT SHRINKS
Back on the Metro, the woman by the window has no idea how much of her song is missing. We can count it, and the count is worth the effort, because it measures how much her ears are willing to forgive.
Sound is a river, and a CD cannot carry a river. It carries snapshots of it, taken so fast that the ear hears a flow. Think of a flipbook, where each page is a still picture and the pages turn so quickly that the picture seems to move. A CD's flipbook turns 44,100 pages or snapshots every second.
Why 44,100 pages? The human ear hears sounds up to about 20,000 vibrations a second. To capture any vibration faithfully, you must take at least two snapshots for every one of its cycles. Engineers call this the Nyquist sampling rule. Twice 20,000 is 40,000. Adding a small margin of safety gave the CD its figure of 44,100.
Here is how much information those pages hold.
Step 1: One snapshot.
Each snapshot stores how loud the sound is at that instant, using 16 on-or-off switches called bits.
Each switch has two settings, so one switch gives 2 choices, two switches give 4, three give 8, and the total doubles every time you add another.
We multiply rather than add, because each new switch can be on or off for every combination already there, the way a shirt can be worn with each pair of trousers, so the choices multiply instead of piling up.
Doubling 16 times is the same as multiplying 2 by itself 16 times, which equals 65,536 loudness levels.
Step 2: One second, one ear.
Take 44,100 snapshots and give each one 16 bits.
44,100 multiplied by 16 equals 7,05,600 bits.
Step 3: One second, two ears.
Music on a CD has a left channel and a right channel, so everything doubles.
7,05,600 multiplied by 2 equals 14,11,200 bits every second.
That is where the figure of 1,411 kilobits per second comes from.
Step 4: What Spotify Premium sends.
At its best setting, Spotify's lossy stream sends 320 kilobits, or 320,000 bits, every second. Set that against the CD.
3,20,000 divided by 14,11,200 equals 0.227, or about 23 per cent.
Step 5: What the free tier sends.
The free tier sends about 1,60,000 bits or 160 kilobits every second.
1,60,000 divided by 14,11,200 equals 0.113, or about 11 per cent.
Step 6: What is left behind.
Take each answer away from 100 per cent.
100 minus 23 equals 77 per cent gone on Premium.
100 minus 11 equals 89 per cent gone on the free tier.
Step 7: What it means for a phone.
One hour is 3,600 seconds. Eight bits make one byte, the unit your phone uses to count storage. A small b means bit and a capital B means byte, so 320 kbps is kilobits, while 635 MB is megabytes.
On a CD: 14,11,200 multiplied by 3,600, divided by 8, equals 63,50,40,000 bytes, or about 635 megabytes per hour.
On Premium: 3,20,000 multiplied by 3,600, divided by 8, equals 14,40,00,000 bytes, or 144 megabytes.
On the free tier: 1,60,000 multiplied by 3,600, divided by 8, equals 7,20,00,000 bytes, or 72 megabytes.
Those figures use a CD as the yardstick, but a song does not begin life on a CD. In a studio, each snapshot is usually written with 24 switches instead of 16.
Remember that each extra switch doubles the loudness levels. Sixteen switches give a CD 65,536 levels, while 24 switches give a studio recording 1,67,77,216. That is like a volume knob with far more notches, so it captures loudness much more finely.
Now do the sums again with 24 bits. A total of 44,100 snapshots multiplied by 24 equals 10,58,400 bits for one ear. Double that for the left and right ears, and you get 21,16,800 bits every second, about 2,117 kilobits. The CD's figure was 1,411.
Set Spotify against it: 320 divided by 2,117 is about 15 per cent, and 160 divided by 2,117 is about 8 per cent. Take those from 100, and at least 85 to 92 per cent of the studio recording never reaches you. The CD, it turns out, makes streaming look better than it is.
Spotify is not alone. Google says YouTube Music Premium members can stream at up to 256 kilobits per second, using codecs called AAC (Advanced Audio Coding) and Opus, while the default setting tops out at 128.
Now set the sums beside the woman on the Metro. Between 77 and 89 per cent of what a CD would have carried never reached her earphone, and the song moves her all the same. The only witness is a listener who did not miss it. The arithmetic tells us how much was taken. It cannot tell us why she never felt it go.
This is lossy compression. Compression means the file is made smaller, and lossy means something is lost along the way. Its opposite, lossless compression, is like a zipped folder, which shrinks and then restores everything exactly.
Amrit Srivastava, a musicologist based in Rome, puts it plainly. "You make the bitrate smaller to get a smaller file," he told India Today Digital, "and what you are doing is squeezing out some parts of the sound."
So the sums leave us with a puzzle about people rather than machines. The deletion is chosen, so that it falls only on what the ear would have missed anyway. How can a machine know that?
WHEN ONE SOUND BLINDS ANOTHER
The answer is masking, and everyone has experienced it. Stand beside a roadside drill and try to hear a friend whisper. The whisper still travels through the air. The ear simply cannot register it under the roar.
Hearing science measures this. A loud sound raises the level a quieter sound must reach to be heard, and the effect is strongest for sounds close to it in pitch. Pitch is how high or low a sound is, counted in vibrations per second, or hertz.
This is frequency masking. The ear sorts pitch into about 24 bands, called critical bands, a division first set out by Eberhard Zwicker in a 1961 paper in the Journal of the Acoustical Society of America.
Masking also works across time. After a loud burst, the ear stays partly deaf to quieter sounds for a fraction of a second. This is temporal masking. A technical report on the theory behind MP3 describes how both effects are used to decide what a compressed file can drop.
So, when a cymbal crashes, a soft violin note at a similar pitch in that instant is still in the air, yet effectively gone from hearing.
THE EAR'S OWN BLIND SPOTS
Masking is not the only gap. The ear is also unevenly sensitive. It hears best between about 2,000 and 5,000 hertz, and grows steadily deafer towards deep bass and very high pitches, a pattern first mapped by Harvey Fletcher and Wilden Munson in 1933 and known as the equal-loudness contours.
The software that does this is called a codec, short for coder and decoder.
It slices the song into frames, each lasting a tiny fraction of a second. For every frame, it splits the sound into frequency bands that mirror the ear's own. It finds the loudest components and calculates how much each one masks the sounds beside it. The result is an invisible line called the masking threshold.
Everything above the line is kept. Everything below it is judged inaudible and is not stored.
This is the idea behind MP3. The audio standard it belongs to, MPEG-1, was published in 1993, and its engineers, Karlheinz Brandenburg and Gerhard Stoll, described the scheme in a 1994 paper in the Journal of the Audio Engineering Society. Spotify's apps use a relative called Ogg Vorbis, an open, patent-free format built on the same principle.
Brandenburg has recalled how he tuned his encoder on an unaccompanied recording, which is a single solo voice playing without any background music, of Suzanne Vega singing Tom's Diner. He knew it would be nearly impossible to compress that warm a cappella voice. A lone voice leaves nothing to hide behind, which made it the hardest possible test.
THE LONG ROAD TO THE EAR
The codec is only one stop on a long road. Jaideep Giridhar, a former editor of the music magazine Rave, who mostly listens to hi-res audio in 24 bits, told India Today Digital that the full journey should be mapped, from the studio to the ear. Think of it as a parcel travelling from a studio to a doorstep.
The parcel begins as the master, the finished studio recording and the richest version of the song that will ever exist. Before it is sent, an encoder packs it. This is the codec we met above, and it packs the song the way you pack a suitcase, leaving behind whatever you were never going to wear. The packed file, called a bitstream, then travels across the internet to your phone.
There, a decoder opens the suitcase. It rebuilds the song as a plain list of snapshot numbers, the same flipbook a CD holds, which engineers call PCM, short for pulse-code modulation. But what was left at home stays at home, and no decoder can bring it back.
Next, the phone tidies the list. Software called DSP, short for digital signal processing, adjusts the volume or the tone, much as an equaliser does. A trick called oversampling adds extra snapshots in between the real ones, so the next step has a smoother picture to work from.
Now the song must cross from numbers back into sound. A converter called a DAC, short for digital-to-analogue converter, turns the list into a smoothly rising and falling electrical signal, the river the snapshots were first taken from. That signal is faint, so an amplifier strengthens it enough to move a speaker. Inside the speaker or earphone, a coil and a magnet shake a thin diaphragm, which pushes the air. The air carries the song into the ear, where masking finally decides what you notice.
Every link can change the sound, so what reaches the ear depends on the whole chain and not on Spotify alone. A wireless earphone adds one more compression step between the decoder and the speaker, as we will see below.
WHAT YOU LOSE, AND WHO NOTICES
What disappears is subtlety. Srivastava compares the effect to hanging four carpets on your wall when the neighbours are having a party. The party is still there, but its finer textures are muffled.
Whether you notice, he says, depends on how you listen. The difference shows only if you listen to a lot of good, well-mixed music on great headphones or hi-fi. A casual listener has long been used to it.
Apple makes a similar point. It says the difference between AAC, its compressed format, and lossless is virtually indistinguishable, though it still offers lossless as an option for those who want it.
Srivastava's advice for anyone who wants the subtleties back is simple: a good pair of speakers at home, a soundbar for a television, and better headphones on the move.
WHERE THE TRICK STRUGGLES
Masking works beautifully inside a dense rock chorus, where plenty of loud sound is available to hide the deletions. It has less disguise when a single voice or a single flute holds a note.
Indian music asks a particular question here. A gamakam is an ornament, a continuous glide or oscillation around a note, and it is the soul of a raga. Professor M. Ramanathan of IIT Madras, who works on teaching machines to recognise ragas, says the music resists easy rules. Each performer can render a raga in a different manner, he tells India Today Digital, so it is difficult to formulate rules, and only broad guidelines are available.
We asked him two questions. Could compression make gamakams harder to catch? And might compression built mainly for simpler, Western-style music handle such variations badly? He said he would have to test both before answering properly.
His short answers were these. To the first: "No, we don't usually compress." To the second: "Maybe yes."
On a related point, he was more reassuring. To identify a raga, he said, "it is probably sufficient to have a global feature recognition system without going into nuances like gamakam." In plain terms, a machine can name a raga from its overall shape, the notes it favours and the way it moves, without tracking every glide.
Recognising a raga is a different task from savouring one, and the first may survive compression better than the second. The question for the next experiment is whether a method tuned to hide deletions in dense music treats a gliding note with the same care.
THE LOSSLESS TWIST
The industry has an answer of its own. After announcing changes to its Premium plans in August 2025, Spotify added a lossless tier in September, streaming FLAC, short for Free Lossless Audio Codec, at up to 24-bit and 44.1 kilohertz, depending on the track, with nothing deleted. Spotify's own announcement says each file shows as Lossless 16-bit or Lossless 24-bit.
Giridhar says Spotify now offers 24-bit audio. Strictly, Spotify stops at 44.1 kilohertz, below the 96 kilohertz or higher that is usually called hi-res. Spotify says its own listening tests at a research lab in Stockholm found that, beyond 24 bits and 44.1 kilohertz, even trained listeners could not reliably hear a difference.
It arrived late. In 2021, Apple announced that Apple Music offers lossless audio in a format called ALAC (Apple Lossless Audio Codec), from CD quality up to 24-bit and 192 kilohertz. In the same year, Amazon Music said its HD tier would stream lossless audio at CD quality, with Ultra HD going up to 24-bit and 192 kilohertz, at no extra cost to Unlimited subscribers.
There is one more catch, and Apple states it itself. AirPods and Beats headphones use Apple's AAC Bluetooth codec, and Bluetooth connections are not lossless. So, even on a lossless stream, the last few metres to a wireless earphone are compressed again.
Every song on a phone is a quiet collaboration between the recording and the limits of the person listening. The ear has blind spots, and a machine learnt to hide inside them.
Music has always been finished by the listener. Streaming has simply learnt where the listener stops.





