The Beautiful Mathematics Behind Every JPEG
I have spent four decades working with images, and one idea still gives me goosebumps every time I watch it work. Take a tiny square of a photograph, just 8 × 8 pixels, 64 numbers. Rewrite those 64 numbers in a different language, the language of waves, and something remarkable happens: almost all of the information collapses into a handful of values in one corner, and the rest become so small you can throw them away. Your eye will never notice.
That rewriting is the Discrete Cosine Transform, or DCT. It sits inside every JPEG on your phone, in the video calls you make, in streaming movies. This article shows it working on real numbers. Every table and picture below was computed by the same mathematics, not drawn by hand.
1. A heretical idea about heat (1807)
In 1807 Joseph Fourier, studying how heat flows through a metal plate, presented a startling claim to the Institut de France: any function, even one with sharp corners, can be written as a sum of simple sine and cosine waves. Lagrange, among the greatest mathematicians alive, objected. It seemed absurd that smooth waves could add up to a jagged edge. Fourier was right, and his 1822 book Théorie analytique de la chaleur gave the world a new way to see.
The insight is this: a signal and its frequencies are two descriptions of the same thing. One says "what is the value here?" The other says "how much of each rhythm is present?" Nothing is lost in translation. But the second description can be far more economical, and economy is what compression is all about.
2. The 64 patterns that build every photograph
Applied to an 8 × 8 block of pixels, Fourier's idea becomes concrete. There are exactly 64 elementary patterns, made by multiplying a horizontal cosine wave by a vertical one. Here they are:
Here is the astonishing part. Every possible 8 × 8 image, whether a patch of sky, an eyelash or a letter "A", is a weighted mix of these 64 patterns. Exactly 64 weights, no more, no less. The DCT is simply the recipe that finds those weights.
The patterns are orthogonal: each one is perfectly independent of all the others, like the north, east and up axes of space, just in 64 dimensions. That is why the weights can be computed one at a time, and why the transform loses nothing: by Parseval's theorem, the total "energy" (sum of squares) is exactly the same before and after.
3. The formula
For a block of pixels f(x, y), the weight of pattern (u, v) is:
F(u,v) = ¼ · C(u) · C(v) · Σx Σy f(x,y) · cos[(2x+1)uπ/16] · cos[(2y+1)vπ/16]
where C(0) = 1/√2 and C(k) = 1 otherwise; x, y, u, v run from 0 to 7
Read it slowly and it is almost poetry. For each pattern, walk across the block, multiply each pixel by the pattern's value at that spot, and add everything up. If the block "resembles" the pattern, the sum is large. If not, positives and negatives cancel to nearly zero. The DCT is 64 questions of the form "how much do you look like this?"
4. Watch it happen on a real block
Below is an 8 × 8 block of brightness values from a photograph (0 = black, 255 = white). Before transforming, JPEG subtracts 128 so values are centred on zero.
| 52 | 55 | 61 | 66 | 70 | 61 | 64 | 73 |
| 63 | 59 | 55 | 90 | 109 | 85 | 69 | 72 |
| 62 | 59 | 68 | 113 | 144 | 104 | 66 | 73 |
| 63 | 58 | 71 | 122 | 154 | 106 | 70 | 69 |
| 67 | 61 | 68 | 104 | 126 | 88 | 68 | 70 |
| 79 | 65 | 60 | 70 | 77 | 68 | 58 | 75 |
| 85 | 71 | 64 | 59 | 55 | 61 | 65 | 83 |
| 87 | 79 | 69 | 68 | 65 | 76 | 78 | 94 |
Now the DCT. The shading shows the size of each weight:
| -415 | -30 | -61 | 27 | 56 | -20 | -2 | 0 |
| 4 | -22 | -61 | 10 | 13 | -7 | -9 | 5 |
| -47 | 7 | 77 | -25 | -29 | 10 | 5 | -6 |
| -49 | 12 | 34 | -15 | -10 | 6 | 2 | 2 |
| 12 | -7 | -13 | -4 | -2 | 2 | -3 | 3 |
| -8 | 3 | 2 | -6 | -2 | 1 | 4 | 2 |
| -1 | 0 | 0 | -2 | -1 | -3 | 4 | -1 |
| 0 | 0 | -1 | -4 | -1 | 0 | 1 | 2 |
Look where the energy went. The single top-left value, the DC coefficient (the block's average brightness), holds 86% of the block's total energy. Just 14 of the 64 coefficients carry 99% of it. The bottom-right region, the fast, fine-grained patterns, is a sea of small numbers.
This is energy compaction, and it is not luck. Neighbouring pixels in real photographs are strongly correlated: the pixel next to a sky-blue pixel is almost always sky blue. Smooth variation lives in low frequencies. The DCT turns that statistical fact into a lopsided list of numbers, and lopsided lists compress beautifully.
5. Where the loss happens and why you don't see it
The DCT itself loses nothing. Compression comes from the next step, quantization: dividing each coefficient by a number from a table and rounding. Here is the standard luminance table published with the JPEG standard:
| 16 | 11 | 10 | 16 | 24 | 40 | 51 | 61 |
| 12 | 12 | 14 | 19 | 26 | 58 | 60 | 55 |
| 14 | 13 | 16 | 24 | 40 | 57 | 69 | 56 |
| 14 | 17 | 22 | 29 | 51 | 87 | 80 | 62 |
| 18 | 22 | 37 | 56 | 68 | 109 | 103 | 77 |
| 24 | 35 | 55 | 64 | 81 | 104 | 113 | 92 |
| 49 | 64 | 78 | 87 | 103 | 121 | 120 | 101 |
| 72 | 92 | 95 | 98 | 112 | 100 | 103 | 99 |
Notice the shape. Small divisors in the top-left preserve the coarse structure carefully. Large divisors in the bottom-right crush fine detail. This table encodes a fact about human vision: our eyes are far less sensitive to rapid, fine-grained variation than to broad changes in brightness. Mathematics meets biology. Divide and round:
| -26 | -3 | -6 | 2 | 2 | -1 | 0 | 0 |
| 0 | -2 | -4 | 1 | 1 | 0 | 0 | 0 |
| -3 | 1 | 5 | -1 | -1 | 0 | 0 | 0 |
| -3 | 1 | 2 | -1 | 0 | 0 | 0 | 0 |
| 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
44 of the 64 values are now zero. Read in zig-zag order from the top-left, the long run of trailing zeros is stored as a single "end of block" symbol, and the remaining numbers are packed with Huffman coding. 64 pixel values shrink to a small fraction of their original size.
Reverse the process (multiply back, inverse DCT, add 128) and the decoded block differs from the original by an average of just 4.9 brightness levels out of 255 (largest single error: 15). Side by side:
6. The goosebumps moment: an image assembling itself
Now rebuild the block using only the first few coefficients, in zig-zag order from coarse to fine. Watch the picture emerge from nothing:
error 20.9
error 20.6
error 18.0
error 14.7
error 5.4
error 0.0
One number gives a grey square of exactly the right average brightness. Each further wave adds structure: the bright shape in the centre gathers itself, and by 21 coefficients, a third of the total, the average error has fallen to about 5 brightness levels out of 255. The remaining 43 coefficients add only fine texture. This is exactly what a progressive JPEG does as it loads over a slow connection: it sends the coarse waves first.
7. Why cosines, and not Fourier's full series?
Fourier's original transform treats a signal as if it repeats forever. Chop out an 8-pixel row whose left end is dark and right end is bright, repeat it, and you create an artificial cliff at every seam. Cliffs need many high-frequency waves to describe, which is precisely what we want to avoid.
The cosine transform quietly assumes something smarter: that the row is mirrored at its edges. A mirror image joins itself smoothly, with no cliff and no wasted waves. That one choice is why the DCT packs energy so tightly.
There is deeper beauty still. The theoretically perfect transform for compacting energy, the Karhunen–Loève transform, depends on the statistics of each image and is costly to compute. But for signals where neighbours are highly correlated, as in natural photographs, the fixed, universal cosine transform turns out to be almost indistinguishable from that optimum. A single elegant formula is nearly as good as the ideal custom solution for every photograph ever taken.
8. How compression evolved: a short history
| Year | Milestone |
|---|---|
| 1807 | Joseph Fourier proposes that any function can be built from sine and cosine waves. |
| 1948 | Claude Shannon's A Mathematical Theory of Communication defines information and entropy, the limits of lossless compression. |
| 1952 | David Huffman, a graduate student at MIT, invents optimal prefix codes, still used inside every JPEG. |
| 1965 | Cooley and Tukey publish the Fast Fourier Transform, making frequency analysis practical on computers. |
| 1974 | Nasir Ahmed, with T. Natarajan and K. R. Rao, publishes the Discrete Cosine Transform. Ahmed later recalled that his earlier funding proposal for the idea had been turned down as too simple. |
| 1977 | Chen, Smith and Fralick publish a fast DCT algorithm; Lempel and Ziv publish LZ77, ancestor of ZIP and PNG. |
| 1986 | The Joint Photographic Experts Group (JPEG) committee is formed. |
| 1988 | H.261, the first widely adopted digital video standard, puts the DCT at the heart of video coding. |
| 1992 | The JPEG standard (ITU-T T.81) is published. Digital photography now has a common language. |
| 1996 | PNG becomes a W3C recommendation: lossless compression for graphics and screenshots. |
| 2000 | JPEG 2000 swaps the DCT for wavelets. Technically impressive, but it never displaced JPEG on the web. |
| 2010 | Google introduces WebP, built on VP8 video coding with block transforms and prediction. |
| 2014 | Mozilla releases MozJPEG, squeezing smaller files from the 1992 format with smarter encoding. CompJPG uses it today. |
| 2019–2022 | AVIF (from AV1 video) and JPEG XL push efficiency further, and the cosine transform's descendants are still at their core. |
It is a humbling arc. An idea born from the physics of heat waited 150 years for computers, was nearly dismissed as too simple, and then quietly became one of the most-used pieces of mathematics in human history. Billions of photographs a day are, underneath, sums of cosines.
9. What this means when you compress a photo
- The quality slider scales the quantization table. Lower quality means bigger divisors, more zeros and more crushed high frequencies.
- Busy images cost more. Foliage, hair and noise have energy in high frequencies that cannot be zeroed without visible damage.
- Blocking artifacts are 8 × 8 tiles whose fine coefficients were all quantized away, leaving each block a flat or gently shaded square.
- Resizing helps more than you expect. Fewer pixels means fewer blocks sharing the same byte budget. See our 20KB–200KB size guide.
For the full encoding pipeline (colour conversion, chroma subsampling and Huffman coding) read How JPEG Compression Works. Then open the compressor, drag the quality slider, and know that every movement is reshaping a quantization table and choosing which cosines survive.