Compress static file responses behind an opt-in (Fix #2545) (#2572)

* Drop the claim that small bodies skip compression

There is no size threshold anywhere in the compression path.
encoding_type() gates on the content type and Accept-Encoding only, and
apply_ranges() compresses whatever body it is given, so a two-byte
text/plain response comes back gzipped at 22 bytes.

Say what actually happens and leave the decision to the handler.

* Compress static file responses behind an opt-in (Fix #2545)

apply_ranges() runs the compressor inside the branch it takes when
res.body is non-empty. A response served from a file leaves res.body
empty and sets content_length_, so it took the other branch, which
writes Content-Length and returns; encoding_type() was computed before
the split and never consulted on that side. The same bytes handed to
set_content() came back gzipped, which left set_mount_point() and
Response::set_file_content() as the one path that missed out.

Add Server::set_static_file_compression(), off by default so nothing
about an existing server changes. When it is on, the file-backed
provider is run through the compressor into res.body ahead of the rest
of apply_ranges(), so the response is framed the way set_content()
already frames one: it keeps its Content-Length, and HEAD still reports
the size a GET would return.

Ranges are answered from the identity representation, since RFC 9110
applies Range after content coding and slicing a compressed body would
mean compressing the whole file first. The ETag carries the coding it
belongs to, so a client that cached the compressed form revalidates
against its own validator rather than the identity one. Both the ETag
and the body take their coding from static_file_encoding(), so the two
cannot disagree.

Providers registered with set_content_provider() are left alone. zlib
buffers until its window fills, so running one through a compressor
would hold back writes that a caller expects to reach the peer as they
are produced.

The compressed bytes stay in memory until the response has been
written, so the peak cost scales with requests in flight.
set_static_file_compression_max_length() bounds it, defaulting to 4MB.

* Add a minimum size for static file compression

Compressing a file that already fits in a single 1500-byte MTU does not
get it to the client any sooner, and a file of a few bytes comes back
larger than it went in once gzip's header and trailer are added. Every
other server draws this line: nginx's gzip_min_length, Caddy's
minimum_length, IIS's minFileSizeForComp, CloudFront's 1000-byte floor.

The note this replaces told callers to decide in the handler. A response
served through set_mount_point() has no handler to decide in, so the
floor has to live in the server. It defaults to 1400 bytes, the size
that fits inside one MTU with room for headers.

set_static_file_compression_min_length() moves it, and
CPPHTTPLIB_STATIC_FILE_COMPRESSION_MIN_LENGTH sets the default at
compile time. The empty-file case keeps its own early-out so that a zero
floor still cannot turn an empty body into a 20-byte gzip stream.

The two bounds now read as a pair, so the documentation says what each
one is for: the lower bound is about what is worth compressing, the
upper bound about what one request is allowed to cost.

Every file under test/www except 1MB.txt is below the default floor, so
the tests that need a small file compressed lower it explicitly.
This commit is contained in:
yhirose
2026-08-27 17:19:43 -04:00
committed by GitHub
parent b4ec1bb1de
commit 139f30e0f1
5 changed files with 637 additions and 24 deletions

View File

@@ -447,6 +447,8 @@ svr.set_pre_compression_logger([](const httplib::Request& req, const httplib::Re
The pre-compression logger is only called when compression would be applied. For responses without compression, only the access logger is called.
For a static file response (see [Static file compression](#static-file-compression)), `res.body` is empty when the logger runs. The bytes are still on disk at that point, not in memory.
#### Error Logging
Error loggers capture failed requests and connection issues. Unlike access loggers, error loggers only receive the Error and Request information, as errors typically occur before a meaningful Response can be generated.
@@ -1466,6 +1468,32 @@ The server can apply compression to the following MIME type contents:
- application/protobuf
- application/xhtml+xml
### Static file compression
Responses served from a file, whether through `set_mount_point()` or `Response::set_file_content()`, are sent as is by default. Turn compression on for them with:
```c++
svr.set_static_file_compression(true);
```
Only files within a size range are compressed, and both ends of it can be moved:
```c++
svr.set_static_file_compression_min_length(512);
svr.set_static_file_compression_max_length(1024 * 1024);
```
The lower bound defaults to 1400 bytes. A response that already fits in a single 1500-byte MTU is not delivered any faster for being smaller, and a file of a few bytes comes back larger than it went in, since gzip's header and trailer outweigh what deflate saves. `0` compresses everything down to a single byte, and `CPPHTTPLIB_STATIC_FILE_COMPRESSION_MIN_LENGTH` sets the default at compile time. An empty file is never compressed regardless.
The upper bound defaults to 4MB, and exists for a different reason: the file is compressed per request, and the compressed bytes are held in memory until the response has been written, so the peak cost scales with the number of requests in flight. It is a bound on what one request can cost, not a statement about how well large files compress, which is why raising it is reasonable when the files are known and the traffic is not. `0` removes the limit, and `CPPHTTPLIB_STATIC_FILE_COMPRESSION_MAX_LENGTH` sets the default at compile time.
A compressed response keeps its `Content-Length`, so `HEAD` still reports the size a `GET` would return. Two details are worth knowing:
- Range requests are answered from the uncompressed representation, so `Content-Range` keeps naming the file's own bytes.
- The `ETag` carries the coding it belongs to (`W/"...-gzip"`), so a client that cached the compressed form revalidates against the right validator.
Content providers registered with `set_content_provider()` are not covered. Feeding one through a compressor would hold each write back until the compressor's window filled, which breaks providers that produce their body incrementally. Use `set_chunked_content_provider()` to compress a generated body.
### Zlib Support
'gzip' compression is available with `CPPHTTPLIB_ZLIB_SUPPORT`. `libz` should be linked.