APC Science Union Explainer: Audio Compression
Author: Shiguang
Reviewer: Baiyan
Before we begin, here is a question: why compress a file? Unsure of the answer, you put the question aside and decide to download something. You open NetEase Cloud Music, find several songs you love in the daily recommendations, and excitedly decide to claim them for your own. The download menu offers three settings: Highest, Very High, and Standard. Naturally, you want to treat your ears well, so you choose Highest, only for the app to ask you to buy a membership. Outraged, you… choose Very High instead.
Now the original question comes back to you. Is the file at Very High simply a compressed version of the one at Highest? After all, it is smaller. The answer is yes. It may not sound quite as good as the Highest version, but it takes up less disk space and is easier to transmit. Compression is therefore important for audio. Of course, it can also tempt people into paying for a membership…
But enough of that digression. Let us look at the basic approaches to audio compression, beginning with the formula for file size:
File size = duration sampling rate bit depth * number of channels
Because we cannot change the duration of an audio recording, we have three options: lower the sampling rate, reduce the bit depth, or use fewer channels.
First, consider the sampling rate. In general, a higher sampling rate produces better audio quality. Some common rates are shown below:

The lowest rate shown is 11,025 Hz, used for speech and amplitude-modulation (AM) radio. Frequency-modulation (FM) radio uses twice the sampling rate of AM. That is one reason drivers generally tune to an FM station rather than an AM one: FM sounds better. If you are curious, compare the two the next time you are in a car.
Another way to compress audio is to reduce its bit depth. Common bit depths are 8 and 16 bits. Converting an approximately 10 MB file from 16-bit to 8-bit audio can cut its size by roughly 5 MB. Ordinary speech, which does not demand high fidelity, generally uses 8 bits. Music, where sound quality matters more, usually uses 16 bits; nobody wants to hear a song buried in noise.

What does the number of channels mean? Stereophonic audio is a method of sound reproduction that generally uses at least two channels. The listener hears sound coming from two directions, creating a sense of space and making the recording feel more natural. Remove one channel from a two-channel stereo recording, and the file size is cut in half. The sound also suffers, so this technique is suitable only for short sound effects or speech, not music.
Audio files are also a poor fit for lossless compression because consecutive samples rarely have the same value. Lossy formats such as MP3 are more common. MP3 offers a good compression ratio while keeping the sound quality acceptable.






