From JPG to WebP: The Quiet Mathematics of Guessing Well
JPEG describes a picture. WebP tries to predict it, and then writes down only how wrong the prediction was.
That one shift in attitude is the heart of this story. Think of how a friend gives you directions to their house. A stranger needs every turn spelled out. But if you already know the neighbourhood, they just say "same as the temple road, but turn left at the tea stall." They send only the surprise.
WebP's lossy mode is borrowed from VP8, a video codec Google acquired in 2010. Video codecs live and die by guessing, because frame 201 looks almost exactly like frame 200. WebP takes the trick a video codec uses inside a single frame and turns it into a still-image format. Google reports that lossy WebP files are 25–34% smaller than comparable JPEGs at equivalent SSIM quality.
Where does that saving come from? Not from one clever idea, but from four small, elegant pieces of mathematics working together. If you have read The Beautiful Mathematics Behind Every JPEG, you will meet an old friend here, and we will even reuse the same 8 × 8 block of photograph. As before, every table and picture below is computed, not drawn by hand.
1. What "converting" really means
There is no shortcut from JPG to WebP. The converter fully decodes the JPEG back into pixels, then encodes those pixels again from scratch.
- Decode the JPEG. Undo the Huffman coding, multiply the DCT coefficients back up, run the inverse DCT, and turn YCbCr back into RGB.
- Convert to YUV again. WebP also separates brightness (Y) from colour (U, V) and keeps colour at half resolution in each direction, because our eyes are far less fussy about colour detail.
- Predict, transform, quantise and entropy-code each block. These are the steps the rest of this article is about.
Here is the humbling bit. Whatever the JPEG threw away is gone forever. The WebP encoder sees JPEG's faint 8 × 8 block edges and ringing as if they were real detail, and dutifully spends bits preserving them. Every lossy re-encode is a photocopy of a photocopy.
So the practical rule is simple: if you still have the original PNG or camera file, encode WebP from that. Convert from JPG when the JPG is all you have, and then use a slightly higher quality setting than you would from a clean source.
2. Idea one: guess the block from its neighbours
WebP cuts the image into 16 × 16 macroblocks and works through them like reading a page: left to right, top to bottom. By the time it reaches a block, the row of pixels just above it and the column just to its left are already decoded. Both the encoder and the decoder can see them, so both can make exactly the same guess about what the new block probably looks like.
There is a small menu of guessing styles:
- DC: fill the block with the average of the neighbours. Good for flat sky.
- Vertical: copy the row above straight down. Good for pillars and tree trunks.
- Horizontal: copy the left column straight across. Good for horizons and shelves.
- TrueMotion: the clever one, below.
For detailed areas the encoder can instead predict each of the sixteen 4 × 4 sub-blocks separately, with ten choices each, including six diagonal directions. The encoder tries the options and keeps whichever guess lands closest.
TrueMotion: a gradient in one line
Call the pixel at the top-left corner C, the pixels above the block A, and the pixels to its left L. TrueMotion predicts the pixel in row r, column c as:
P(r, c) = clamp( L[r] + A[c] − C ) clamp keeps the result between 0 and 255
Read it slowly. A[c] − C is how much brightness changes as you move right along the top edge. TrueMotion takes the left edge and adds that same change. It assumes the picture is a gently tilted surface, and extends the tilt into the unknown block.
A tiny example. Suppose the top row is 100, 110, 120, 130, the left column is 100, 105, 110, 115, and the corner is 95. TrueMotion fills in:
| 105 | 115 | 125 | 135 |
| 110 | 120 | 130 | 140 |
| 115 | 125 | 135 | 145 |
| 120 | 130 | 140 | 150 |
The guess ramps smoothly in both directions, exactly like light falling across a wall. Nothing about the inside of the block was transmitted; it was inferred from its edges.
Now on a real photograph
Real images are not perfect ramps, so let's be honest and try it on real data: a 4 × 4 patch taken from the same photograph block we used in the DCT article, with its true neighbours above and to the left. Here are all four guesses side by side, with the total error of each (the sum of how far every predicted pixel is from the truth):
error 428
error 294
error 480
error 248
TrueMotion wins, with a total error of 248. For comparison, with no guess at all (just centring every pixel on mid-grey, as JPEG does) the errors would total 594. The prediction has removed 58% of what would otherwise need to be described. What remains is the difference between guess and reality:
| 10 | 22 | 6 | -16 |
| 16 | 29 | 5 | -15 |
| 1 | 4 | -10 | -14 |
| -25 | -37 | -22 | -16 |
That difference is called the residual, and it is all the encoder has to send for this patch. Smaller numbers, centred on zero, are far cheaper to store than the pixels themselves. In smooth regions of a photo (sky, walls, skin) the residual is smaller still.
3. Idea two: the DCT's practical little cousin
The DCT re-describes a block of pixels as a mix of smooth waves, so that most of the energy collects in a few low-frequency numbers. WebP does the same thing to the residual, with two thoughtful changes.
It uses 4 × 4 blocks instead of 8 × 8. The residual is already mostly leftovers, not a whole picture. Smaller blocks keep any error close to where it happened, so a sharp edge doesn't smear ripples across a wide area.
It uses whole numbers instead of real cosines. The true DCT needs values like √2·sin(π/8) ≈ 0.541196, which a computer cannot store exactly. The VP8 decoder replaces them with fractions over 65,536:
| Ideal value | Exact | Stored as | Which equals | Error |
|---|---|---|---|---|
| √2 · sin(π/8) | 0.541196 | 35468 / 65536 | 0.541199 | 0.0000026 |
| √2 · cos(π/8) | 1.306563 | 1 + 20091 / 65536 | 1.306564 | 0.0000014 |
The second constant is bigger than 1, so it is stored as "one plus a fraction" and only the fraction, 20091, needs to fit in 16 bits. The errors are in the sixth decimal place: invisible to your eye, and exact in a way floating point never is.
Why does exactness matter so much? Because of prediction. The decoder builds each guess from pixels it reconstructed. If its arithmetic differed from the encoder's by even one unit, every later guess would stand on a slightly different foundation, and the error would snowball down the image. Integer mathematics guarantees that every phone, browser and laptop decodes exactly the same pixels. A little loss of mathematical purity buys perfect agreement.
After the transform comes the one truly lossy step, just as in JPEG: quantisation, dividing each coefficient by a step size and rounding. WebP adds a nice twist. The encoder may split the image into up to four segments and give each its own step size, so a detailed face can keep fine quantisation while a blurry background is squeezed hard.
4. Idea three: a transform made only of plus and minus
Here is a subtle waste. When a macroblock is predicted as one 16 × 16 piece, its residual is still transformed as sixteen 4 × 4 blocks, and each of those gets its own DC coefficient, its average level. In a smooth area those sixteen averages are nearly identical. Storing all sixteen separately is like writing "₹500" on sixteen envelopes.
So WebP gathers the sixteen DC values into their own little 4 × 4 grid and transforms them a second time, using the Walsh–Hadamard transform. Its matrix contains nothing but +1 and −1:
⎡ 1 1 1 1 ⎤
H₄ = ⎢ 1 −1 1 −1 ⎥
⎢ 1 1 −1 −1 ⎥
⎣ 1 −1 −1 1 ⎦
No cosines, no fractions, no multiplications at all: a computer can do it with additions and subtractions alone. And it has a lovely self-similar structure. Each Hadamard matrix is built from the one half its size, copied three times and flipped once:
H₂ₙ = ⎡ Hₙ Hₙ ⎤ starting from H₁ = [ 1 ]
⎣ Hₙ −Hₙ ⎦
The rows are square waves instead of cosine waves. The first row asks "what is the total?" The others ask "how does the left differ from the right?", "the top from the bottom?", and so on.
Let's try it. Here are sixteen block averages from a gently brightening patch of sky:
| 150 | 151 | 153 | 154 |
| 149 | 151 | 152 | 154 |
| 148 | 150 | 151 | 153 |
| 147 | 149 | 151 | 152 |
Apply H₄ to the rows and then to the columns, using only additions and subtractions:
| 2415 | -13 | -25 | -1 |
| 5 | 1 | 1 | 1 |
| 13 | 1 | 1 | 1 |
| -1 | 3 | -1 | -1 |
One big number, 2415, is simply the sum of all sixteen. The other fifteen describe the gentle slope, and they are tiny. Quantise with a step of 8 and 11 of the 16 values become zero. If the sixteen averages had been exactly equal, all fifteen would have been zero: sixteen numbers described by one. This is the whole format in miniature: find the redundancy, describe it once.
5. Idea four: spending less than one bit
Everyday JPEGs finish with Huffman coding, which gives common symbols short codes and rare symbols long ones. It is wonderful, but it has a floor: every symbol costs at least one whole bit. You cannot write half a letter.
Now imagine a question the encoder asks millions of times: "Is this coefficient zero?" In a well-predicted image the answer is very often yes. Claude Shannon showed in 1948 that an event with probability p truly carries:
cost = −log₂ p bits
| Probability of "yes" | Honest cost (bits) | Huffman minimum (bits) |
|---|---|---|
| 50% | 1.000 | 1 |
| 75% | 0.415 | 1 |
| 90% | 0.152 | 1 |
| 95% | 0.074 | 1 |
| 99% | 0.014 | 1 |
At 95% certainty, the honest price is about 0.074 bits, and Huffman would charge more than thirteen times that.
WebP uses a boolean arithmetic coder, which gets very close to the honest price. The idea is beautiful. Picture the number line from 0 to 1. For each yes/no decision, split the current interval in proportion to the probabilities and step into the part that actually happened. Here is a run of six decisions where "zero" has probability 0.9:
After all six decisions the interval has shrunk to a width of 0.0590. Writing down any number inside it takes about −log₂ of that width:
total bits ≈ −Σ log₂ pᵢ = 4.08 bits (Huffman: at least 6)
The whole image becomes a single, very precise fraction. Likely events barely shrink the interval, so they cost almost nothing. Surprises shrink it a lot, so they cost more. The bill is always fair.
WebP sharpens the probabilities further by looking at context: which frequency the coefficient belongs to, and whether the neighbouring blocks had anything non-zero. Better guesses about the guesses mean even fewer bits. (The JPEG standard also allows arithmetic coding, but for patent reasons almost nobody used it.)
6. A bonus: the lossless side has its own tricks
WebP has a second, entirely separate mode that throws nothing away, a rival to PNG that Google reports is 26% smaller on average. It uses a completely different bitstream, and its ideas are charming in their own right.
Subtract green. In most photographs, red, green and blue rise and fall together; a bright pixel is usually bright in all three. So WebP can store red and blue as differences from green:
R′ = (R − G) mod 256 B′ = (B − G) mod 256
A grey pixel (140, 140, 140) becomes (0, 140, 0). Two of the three numbers collapse towards zero, and zeros compress beautifully. The decoder simply adds green back.
Thinking in two dimensions. Classic compressors like ZIP look back along a one-dimensional stream for repeated data. But an image is a grid, and the most likely match for a pixel is often the one directly above it, which in a stream is a whole row-width away. WebP reserves its shortest distance codes for the nearby 2D neighbourhood, so "same as the pixel above" costs very little however wide the image is.
A colour cache with a hash. WebP keeps a small table of recently seen colours, and finds each colour's slot by multiplying it by a fixed odd constant and keeping the top bits:
slot = (0x1e35a7bd × colour) >> (32 − k) a table of 2ᵏ recent colours
This is multiplicative hashing: the multiplication scatters similar colours to very different slots, so they rarely collide. When a colour repeats, the encoder sends a short slot number instead of four full bytes.
7. The quiet lesson
JPEG's genius was noticing what our eyes don't care about. WebP keeps that insight and adds a second one: most of a picture can be guessed, and you only need to pay for your surprises.
Every idea above is a version of that sentence. Prediction guesses a block from its neighbours. The Walsh–Hadamard transform notices sixteen averages are really one. Arithmetic coding charges each event exactly what its surprise is worth. Subtract green notices the three colours are mostly telling the same story. None of it is magic. It is patience with redundancy, written in the language of mathematics.
A few practical takeaways when you convert:
- Start from the cleanest source you have. A WebP made from an original beats one made from a JPEG.
- Quality 75–85 is a sensible range for photos. Much lower, and smooth gradients start to look waxy, because prediction plus heavy quantisation flattens fine texture.
- Prefer lossless for screenshots, logos and diagrams. Flat colours and sharp text are where lossy blocks struggle. Use PNG or lossless WebP for those; CompJPG's converter makes lossy WebP for photographs.
- Compare by eye, not just by file size. WebP is not always smaller, and a smaller file that looks worse is not a saving.
The next time you drop a photo into the JPG to WebP converter and it comes back smaller, you will know what happened inside: a small machine made millions of careful guesses, and wrote down only where it was wrong.
Further reading
- RFC 6386: VP8 Data Format and Decoding Guide, the full specification behind lossy WebP
- RFC 9649: WebP Image Format, the container and the lossless bitstream
- Google: WebP compression study
- Google: lossless and alpha study
- On this site: The Beautiful Mathematics Behind Every JPEG and JPEG vs PNG vs WebP