846 Commits

Author SHA1 Message Date
metsw24-max
edc9760005 use the SAN type tag in Mbed TLS verify_hostname and get_cert_sans (#2614)
* use the SAN type tag in Mbed TLS verify_hostname and get_cert_sans

Mbed TLS keeps a subjectAltName entry's GeneralName tag in buf.tag and the bare value in buf.p / buf.len. verify_hostname ignored the tag, so a dNSName whose bytes equal an address authenticated that IP host, and an iPAddress or rfc822Name was matched as a DNS pattern. get_cert_sans looked for the tag inside the value, so it reported no entries for an ordinary certificate, or part of a dNSName as an entry of its own.

* Shorten the SAN type comments

---------

Co-authored-by: yhirose <yuji.hirose.bug@gmail.com>
2026-10-05 21:13:50 -04:00
yhirose
6d1049462d Ignore a subprotocol the WebSocket client did not offer
A SubProtocolSelector could return a value outside the client's
Sec-WebSocket-Protocol list, and the server sent it back as is. The
server now treats such a value as no selection.
2026-10-04 18:34:21 -04:00
metsw24-max
4fcbc08f2d reject an unoffered subprotocol in read_websocket_upgrade_response (#2595)
* reject an unoffered subprotocol in the ws client handshake

* test the subprotocol check through WebSocketClient so the split build compiles

* Trim comments in the subprotocol check

---------

Co-authored-by: yhirose <yuji.hirose.bug@gmail.com>
2026-10-04 18:21:39 -04:00
yhirose
09fa6820fe Fix 303 redirect request bodies and Location path decoding (Fix #2606, #2607)
A 303 response turns the follow-up request into a GET, but only the buffered body and headers were cleared. A content provider (sized or chunked) stayed on the request, so the original payload was sent again, and for a chunked provider the unframed chunks also broke the keep-alive connection.

The redirect also percent-decoded the Location path before sending it. That turned %23, %3F and %25 into a fragment, a query delimiter and a different octet, so the client requested a different resource than the one named, and made set_path_encode(false) fail on any Location containing %20. The path is now sent as given.
2026-10-03 19:06:45 -04:00
yhirose
eda9a10bfe Fix SSE client not clearing Last-Event-ID on an empty id field (Fix #2611)
An event with an empty id field must reset the last event ID, so that
no Last-Event-ID header is sent on reconnect. run_event_loop only
updated last_event_id_ when the id was non-empty, so it could not tell
an empty id field from an event with no id field, and kept sending the
stale ID.

Track whether an id field was seen with a has_id flag, as has_data
does for the data field, and add a regression test.
2026-10-03 17:29:33 -04:00
yhirose
4fd9ae8f42 Serve pipelined requests without waiting for the keep-alive timeout
A client may pipeline its requests (RFC 9112 9.3.2). The server read the
following request(s) into a per-request SocketStream buffer, discarded them
with the stream after the first response, and then waited in keep_alive()
for socket data that never came, closing the connection after the
keep-alive timeout. Over TLS the bytes stayed decrypted in the TLS library,
where keep_alive() could not see them either.

Create one stream per connection and serve a request that is already
buffered (Stream::is_readable()) without waiting in keep_alive().

Keeping the buffer means an extra CRLF that some clients send after a
request body is now parsed as the next request line, which answered 400
and closed the connection. Ignore one empty line before the request-line,
as RFC 9112 2.2 recommends.

Fixes #2599
2026-10-02 21:47:18 -04:00
yhirose
3d40dfc727 Fix SSE parsing of CRLF field names and data-less events
A field line without a colon kept the \r of a CRLF line ending in its
name, so "data\r\n" was not recognized. Strip the \r once per line
before parsing instead of from each value.

An event without a data field was neither dispatched nor cleared, so
its event type leaked into the next event and its id only reached
last_event_id after a later event was dispatched. Reset the message on
every blank line and record the id even when nothing is dispatched.

The parsing tests exercised a copy of parse_sse_line that had drifted
from the real one. Run them through SSEClient against a local server
instead, and add regression tests for the cases above.
2026-10-02 21:38:39 -04:00
metsw24-max
dd71728110 escape quoted-string auth-params in make_digest_authentication_header (#2597) 2026-10-02 21:29:35 -04:00
DosX
7255a7e979 Preserve empty SSE data fields (#2594)
SSE events may contain an empty data field, and a data field without a colon also has an empty value. Track whether a data field was seen separately from the accumulated payload so empty events are dispatched and leading empty lines are preserved. Add an integration regression test for both forms.
2026-10-02 21:17:47 -04:00
KBS
10aadd57f7 Reject trailing characters in URL port numbers (#2593)
parse_port checked only the error code of from_chars, which stops at
the first non-digit, so http://host:80abc was accepted as port 80, and
a redirect Location with such a port was followed. RFC 3986 defines
port as *DIGIT. Require the whole string to be consumed, as #2590 does
for quality values.
2026-10-02 21:07:34 -04:00
yhirose
43863e1f67 Make CryptoAPI the only chain verifier when Windows verification is on (#2604)
Every backend loads the Windows ROOT and CA stores, but Windows adds a
root to them only when CryptoAPI needs it to build a chain. On a machine
that had not needed a root yet, the backend rejected the chain before
the CryptoAPI check ran, so Windows never fetched the root (for example
OpenSSL error 20 for accounts.spotify.com under Starfield Root G2).

With Windows verification enabled, the backend's chain verdict is no
longer used. Mbed TLS and wolfSSL skip chain verification during the
handshake, and the post-handshake verify result is ignored. The
CryptoAPI check becomes mandatory: a leaf that cannot be encoded now
fails the connection instead of skipping the check, and the chain must
allow server authentication, which the backends used to check.

A server certificate verifier works on the backend's chain verification,
so when one is set the backend still decides and CryptoAPI only adds its
own check, as before.

ClientCertMissing now disables hostname verification: cert.pem does not
match HOST, and on Windows the hostname check fails before the chain
check.

Refs #2596
2026-10-02 20:48:47 -04:00
yhirose
0db1df7cf2 Pass the server's intermediates to Windows certificate verification (#2602)
CryptoAPI got only the leaf, so it fetched an issuer from the leaf's AIA
URL instead of using the intermediates the server sent. For
accounts.spotify.com that issuer chains to Certainly Root R1, which
Windows does not trust, while the server's own chain ends at Starfield
Root G2.

Add tls::get_peer_certs(), which returns the certificates the peer sent
in the same way get_ca_certs() returns the CA certificates, for every
backend. The CryptoAPI check puts them into a memory store that it
passes to CertGetCertificateChain(), skipping any certificate that
cannot be added. wolfSSL keeps the received chain only when built with
SESSION_CERTS; without it, CryptoAPI still gets the leaf alone.

Refs #2596
2026-10-02 20:48:10 -04:00
yhirose
639391ad7f Reject an invalid Content-Length and honor Connection: close
A request whose Content-Length was present but not a valid decimal length
(e.g. "42, 42", "+42", "0x2e" or empty) was treated as having no body
unless a handler read it: no 400 was returned, the body was not drained,
and the bytes after the header block were parsed as the next request on
the keep-alive connection. Reject such a request with 400 and close the
connection before routing, as RFC 9112 Section 6.3 requires.

The server also kept reading after a response that announced
Connection: close. A rejected request line or header block left the rest
of the message to be parsed as a new request, and an error response to a
bodyless request did the same with whatever followed. Close the
connection whenever the final response carries Connection: close
(RFC 9112 Section 9.6), and mark the two request-head rejection paths
closed explicitly as the other rejection paths already do.
2026-10-02 00:22:24 -04:00
yhirose
0715c2739e Enforce a minimum SSE reconnect wait to avoid a busy loop (#2592)
SSEClient::wait_for_reconnect() sleeps in 100ms steps until the
reconnect interval has elapsed. With an interval of 0 (for example
"retry: 0" from the server, or set_reconnect_interval(0)) it never
slept at all, so a server that sends "retry: 0" and closes the stream
made the client reconnect in a tight loop. set_max_reconnect_attempts()
does not stop this either, because each successful connection resets
the attempt counter.

Always wait at least one step (100ms). Intervals of 1-99ms already
waited 100ms because of the step size, so only 0 and negative values
change behavior.
2026-09-28 17:25:37 -04:00
KBS
3330d0eb06 Ignore an SSE retry field that is not all digits (#2591)
parse_sse_line checked only the error code of from_chars, which accepts
a leading '-' and stops at the first non-digit, so retry: -1 made the
client reconnect without waiting and retry: 10s set 10 ms. The SSE spec
ignores a retry value that is not all ASCII digits.
2026-09-28 15:31:20 -04:00
yhirose
c1c2b1f4b4 Apply the request-target check to the client and encode control chars
Share the server's request-target check as fields::is_request_target()
and use it in write_request_line too. The client previously used
is_field_value(), which let an embedded SP or HTAB through.

encode_path() only escaped CR/LF among the control characters, so with
path encoding enabled a path like "/a\tb" would now be rejected instead
of sent. Percent-encode every control character (0x00-0x1F, 0x7F).
2026-09-28 03:37:11 -04:00
yhirose
e11dbec7b3 Reject control characters in the request-target
RFC 9112 §3.2 does not allow control characters in the request-target,
and §2.2 requires a bare CR to be treated as invalid. parse_request_line
accepted them, so e.g. "GET /a\rb HTTP/1.1" was routed normally. Reject
any byte that is not VCHAR or obs-text with 400 Bad Request. obs-text is
still allowed since some clients send raw UTF-8 in the target.
2026-09-27 23:35:09 -04:00
DosX
57c4f7f385 Reject trailing characters in HTTP quality values (#2590)
parse_quality accepted values such as q=0.5junk because it checked only the conversion error and ignored the returned end pointer. Require the numeric parser to consume the complete q parameter so malformed Accept values are rejected and invalid Accept-Encoding weights are ignored. Add regression cases for both headers.
2026-09-27 19:20:14 -04:00
VecSzn
8b6ab24159 Resolve relative Location references on redirect (#2586) 2026-09-26 19:49:21 -04:00
yhirose
6d59d1e2df Build open_stream's request head in memory before sending
open_stream wrote the request line and each header straight to the
socket, so a header rejected by check_and_write_headers left the request
line and the headers before it on the wire. Build them in a BufferStream
first and flush once, as write_request and the WebSocket handshake do.
2026-09-21 19:15:39 -04:00
yhirose
9386b25dd7 Don't wait for a response to a request rejected before sending
process_request reads the response even after write_request fails, so
that an early response (e.g. 413/414) sent while the body is still being
uploaded is not lost. A request line or header rejected while the
request is being built in memory never reaches the socket, though, so
no response will come and the client blocked until the read timeout (or
until the server closed the idle connection).

write_request now reports such a local rejection, and process_request
returns immediately in that case. Socket write failures still read the
response as before.
2026-09-21 19:15:30 -04:00
yhirose
ad88645a83 Reject non-token methods in write_request_line
write_request_line checked the request target for CR/LF but concatenated
the method verbatim. A method carrying CR/LF could smuggle a whole
request ahead of the real one, and the client would take the smuggled
request's response as its own. A method with a space or an empty method
put a malformed request line on the wire.

Require the method to be a token (RFC 9110 Section 9.1) before anything
is written. All three callers (the buffered client path, open_stream and
the WebSocket handshake) go through this function and fail with
Error::Write, as they already do for a rejected target.

Claude-Session: https://claude.ai/code/session_01NTDesJQTQPEuu4o4XCu69g
2026-09-21 12:25:12 -04:00
metsw24-max
8b872605e0 reject control characters in chunk extensions in read_payload (#2585)
* reject control characters in chunk extensions in read_payload

* Bound every chunk-size line scan by the line terminator

read_payload() ended its scans of one line buffer two different ways: the
hex-size parse and the space skip that follows stopped on the NUL that
stream_line_reader::append() writes, while the new chunk-ext check walked
to an explicit end pointer. Compute that end pointer first and bound all
of them by it, so no scan depends on the buffer's NUL and the terminator
can never be read as line content.

The bare-LF branch is reachable only under
CPPHTTPLIB_ALLOW_LF_AS_LINE_TERMINATOR, where getline() ends the line on
an LF that is the terminator rather than extension text. Say so: the
comment below it explains why a bare LF inside the line is rejected, and
without that note the two read as contradictory. Its guard no longer
depends on the scan cursor either, since all it ever needed was a check
that there is a byte to look at.

* Reuse the chunked-body helper in the chunk-ext acceptance test

AcceptsChunkExtension repeated expect_chunked_body_rejected()'s body
verbatim apart from the expected status, so parameterise the helper on
the status and keep the rejection wrapper for the existing callers. The
decoded body is already checked by the /chunked handler, so asserting
the status is all the new test needs.

Also record why the control-character literal stays split: a hex escape
consumes every hex digit that follows it, so "\x01b" would be the single
byte \x1b rather than \x01 followed by 'b', and joining the halves would
quietly change what the test sends.

---------

Co-authored-by: yhirose <yuji.hirose.bug@gmail.com>
2026-09-20 20:22:52 -04:00
yhirose
52f214bf2e Don't wait for the peer's close_notify on OpenSSL shutdown
tls::shutdown() on OpenSSL called SSL_shutdown() a second time to wait
for the peer's close_notify. An idle keep-alive client never sends one,
so closing its connection held the worker thread until the read timeout,
and Server::stop() waited for it. Send close_notify and return, as the
Mbed TLS and wolfSSL backends already do.
2026-09-19 17:28:09 -04:00
yhirose
09c02f1335 Send 100 Continue only when the request body is read
The server wrote 100 Continue as soon as it saw the expectation, before
pre_routing_handler, pre_request_handler, or routing ran. A request
those handlers rejected, or one that matched no route, still invited
the client to send a body the server would never read.

Defer the interim response until the body is about to be read. If the
request is answered without reading the body, 100 Continue is never
sent and the connection is closed, since whether and when the client
sends the body is unknown.

Also treat a 417 returned by expect_100_continue_handler as the final
response. It used to be written as a bare status line, after which the
request was processed and a second response was written.
2026-09-19 17:28:09 -04:00
yhirose
4cb363e3f2 Run pre_request_handler for WebSocket routes
The WebSocket upgrade path matched the route and switched protocols
without setting req.matched_route or calling pre_request_handler, so a
check placed there (e.g. authentication) never ran for WebSocket
routes. Set matched_route and run the handler before the upgrade; if it
handles the request, reply with a regular HTTP response instead of 101.

Also write rejected upgrade responses (from pre_routing_handler too)
with write_response_with_content, so they carry Content-Length. Without
it, a client reading the body waited until the keep-alive timeout.
2026-09-19 17:28:09 -04:00
yhirose
f37a5b1407 Fix #2583 2026-09-14 17:31:07 -04:00
yhirose
6b8c3f5387 Report only a caller-set WebSocket read timeout as Timeout
199d7ee made read() return the new ReadResult::Timeout for every read
timeout and leave the connection open. The compile-time server default
(CPPHTTPLIB_WEBSOCKET_SERVER_READ_TIMEOUT_SECOND, 300s) is always in
effect, so a handler written as `while (ws.read(msg))`, the form the
README's Quick Start uses, no longer ended when a peer went quiet:
Timeout is non-zero, so the loop ran its body again with the previous
message still in `msg`, and the worker the backstop is meant to reclaim
was never released. Nothing caught it because every test of the new
result set a timeout explicitly and checked the result by value, and the
heartbeat tests keep the connection alive with pings.

The two timeouts mean different things. One the caller sets through
set_read_timeout() is a request for control back, and is reported as
Timeout on a still-open connection. The compile-time default is a
backstop against a peer that has gone quiet, and elapsing it is now a
failure again: read() returns Fail and closes the connection, as it did
before 199d7ee. WebSocket tracks whether set_read_timeout() was called,
and WebSocketClient carries the same flag over to the WebSocket it
creates on connect().

Tests use the heartbeat binary, which compiles both defaults down to 3s:
a `while (ws.read(msg))` server handler runs its body once and exits
when the client falls silent, and a client that never set a timeout gets
Fail with the connection closed. The README and cookbook now say which
timeout produces Timeout.

Claude-Session: https://claude.ai/code/session_01EF5uZ1X2kaHhqJ8VgfjVaQ
2026-09-11 17:29:20 -04:00
yhirose
73a4092f8a Point the remaining httpbingo.org proxy tests at the self-hosted httpbin
ProxyTest, RedirectTest.HTTPBin*, KeepAliveTest and ProxyTest.SSLOpenStream
still sent their requests through the squid proxies to the external
httpbingo.org, so an upstream hiccup there failed CI with no code change
involved (KeepAliveTest.SSLWithDigest got a 502 on its first /get).

Switch them to the "httpbin" container (nginx + go-httpbin) that
BaseAuthTest/DigestAuthTest already use. go-httpbin serves /get,
/redirect/n and /digest-auth the same way, so the test logic is
unchanged; the SSL variants disable certificate verification for the
self-signed test cert, as BaseAuthTest.SSL does.

RedirectTest.YouTube* is left pointing at youtube.com since it exercises
a real cross-host, http -> https redirect chain.

Claude-Session: https://claude.ai/code/session_0148ZAzsuYRYXkwcA7UFeh95
2026-09-11 16:35:11 -04:00
metsw24-max
0480ff77b8 reject ambiguously framed responses in client read paths (#2581)
* reject ambiguously framed responses in client read paths

* Accept non-chunked Transfer-Encoding responses in the client framing guard

RFC 9112 §6.3 treats requests and responses differently when the final
transfer coding is not chunked: a request's body length cannot be
determined and the server must answer 400, but a response's body simply
runs until the server closes the connection. read_content() and the
open_stream() body reader already do that, so such a response is not
ambiguous and rejecting it broke valid responses such as
"Transfer-Encoding: gzip" followed by a close.

Keep rejecting a Transfer-Encoding paired with a non-zero Content-Length,
which is the actual ambiguity, and drop the non-chunked clause from both
client read paths.

Tests: check that rejection surfaces as Error::Read, that a non-chunked
Transfer-Encoding response is read until close on both paths, and that
HEAD, 204 and 304 responses with both framing headers are not rejected.

Claude-Session: https://claude.ai/code/session_01JYPWKpbp4a881EdpEf2xSi

* Share the framing check and reuse existing test helpers

Factor "Transfer-Encoding with a non-zero Content-Length" into
detail::has_conflicting_content_length() next to
is_chunked_transfer_encoding(), and call it from the server request
guard and both client read paths so the rule and its RFC 9112 §6.3
rationale live in one place.

In the tests, drop the POSIX-only raw socket helper in favour of the
existing serve_single_response() and read_all(), which also lets the
tests run on Windows. Fold the stream-only test into the buffered one so
each case checks both Get() and open_stream(), and cover the HEAD/204/304
exclusion on the open_stream() path too.

Claude-Session: https://claude.ai/code/session_01JYPWKpbp4a881EdpEf2xSi

---------

Co-authored-by: yhirose <yuji.hirose.bug@gmail.com>
2026-09-11 15:55:28 -04:00
KBS
8d25b6a3ac Reject a Range first-byte-pos that overflows ssize_t (#2580)
* Reject a Range first-byte-pos that overflows ssize_t

parse_range_header initializes first to the -1 sentinel that means "no
first-byte-pos" and only overwrites it when detail::from_chars succeeds.
On std::errc::result_out_of_range the assignment is skipped and -1
survives, so "bytes=9223372036854775808-100" is parsed as the suffix
range "bytes=-100" and range_error serves the last 100 bytes instead of
returning 416.

Before the parser was rewritten onto detail::from_chars, std::stoll threw
std::out_of_range on the same input, the catch arm added in 8f8761e for
issue #705 returned false, and the request was answered with 416. The
catch arm is still there but from_chars reports through an error code, so
nothing reaches it any more.

get_header_value_u64 and parse_port already reject an out-of-range value
at their from_chars call sites; this was the remaining one that dropped
the error.

The last-byte-pos side is deliberately unchanged: -1 there is the
documented RFC 9110 14.1.2 "remainder of the representation" value, so an
oversized last-byte-pos stays accepted.

* Simplify the Range first-byte-pos overflow check

Parse the first-byte-pos straight into first, since a failed parse now
returns before first is read, and fold the overflow test into the
existing batch of rejected ranges. Also note on the last-byte-pos side
why an overflow there deliberately keeps -1.

Claude-Session: https://claude.ai/code/session_01JYPWKpbp4a881EdpEf2xSi

---------

Co-authored-by: yhirose <yuji.hirose.bug@gmail.com>
2026-09-11 13:48:04 -04:00
metsw24-max
515b8f84af send each credential only to its own hop in write_request (#2579)
* send each credential only to its own hop in write_request

An SSLClient behind a proxy sent Proxy-Authorization inside the TLS tunnel, where the origin reads it, and sent the origin's Authorization on the CONNECT request the proxy reads. Attach each only on the message its hop actually reads.

* Keep default headers off the CONNECT request

set_default_headers() is typically used for origin credentials such as
Authorization, Cookie or API keys, but they were also attached to the
CONNECT request an SSLClient sends to its proxy, in plaintext before the
TLS tunnel exists. Default headers now go only on requests the origin
reads, the same split the previous commit makes for set_basic_auth and
set_bearer_token_auth.

Claude-Session: https://claude.ai/code/session_01JYPWKpbp4a881EdpEf2xSi

* Simplify per-hop credential handling and its tests

Flatten the Authorization insertion in write_request into one guard with
an else-if (Basic already took precedence over Bearer), and shorten the
comments around it. Fold DefaultHeadersStayOffConnect into the
CredentialsStayWithTheirHop helper, which now takes the list of headers
that must reach only the origin.

Claude-Session: https://claude.ai/code/session_01JYPWKpbp4a881EdpEf2xSi

---------

Co-authored-by: yhirose <yuji.hirose.bug@gmail.com>
2026-09-11 13:25:48 -04:00
yhirose
88956ccad8 Self-host the httpbin auth-testing backend for BaseAuthTest/DigestAuthTest
These tests exercise the squid proxies by hitting /basic-auth and
/digest-auth on an external httpbin-style site. That site's identity has
already moved twice (httpbin.org -> httpcan.org, per #2300) chasing
uptime, and httpcan.org itself is now down (Cloudflare 502 from its
origin), failing CI with no code change involved.

Adds two containers to the existing squid docker-compose stack instead:
go-httpbin (mccutchen/go-httpbin) as the backend, and an nginx sidecar in
front of it under the single "httpbin" hostname so both the NoSSL tests
(port 80) and the SSL tests, which CONNECT-tunnel through the proxy to
port 443, resolve the same name -- go-httpbin only listens on one port at
a time, so it can't serve both protocols itself. nginx uses the repo's
existing self-signed test cert; the SSL client tests already disable
verification for it like other self-signed-cert tests in this suite.

go-httpbin was picked over the more feature-complete kennethreitz/httpbin
after finding the latter accepts a wrong digest-auth username as long as
the password matches -- confirmed with a direct curl against the
container, unrelated to anything in this repo. go-httpbin correctly
rejects both. The trade-off is losing SHA-512 digest-auth coverage here,
since go-httpbin only implements MD5 and SHA-256; nothing else in the
suite exercises SHA-512 digest auth against a live server. Response body
assertions are adjusted to go-httpbin's actual JSON shape (an added
"authorized" field, no "algorithm" field), and the domain changes from
httpcan.org to the self-hosted "httpbin".

This only affects 'make proxy'/'make proxy_mbedtls'/'make proxy_wolfssl'
and the Proxy Test CI workflow -- the default 'make' target is untouched.
2026-09-07 21:49:00 -04:00
yhirose
6303a99ce1 Run clang-format on httplib.h and test.cc
Fixes formatting introduced in 199d7ee and d5f8858 that clang-format
18.1.3 (the version used by CI) disagrees with.
2026-09-07 16:29:08 -04:00
yhirose
199d7ee248 Tell a WebSocket read timeout apart from a closed connection
read() collapsed every failure into Fail and marked the connection closed with
it, so a read timeout could not be used to get control back and send on the
same connection -- it killed the connection instead. The information was
already there and thrown away: SocketStream::read records Error::Timeout, and
read_websocket_frame flattened it into a bool.

ReadResult gains Timeout, reported only when the timeout elapsed on a frame
boundary with nothing consumed, which is the only case where the stream can be
read again. Every multi-byte field now loops until it has its bytes, which also
fixes a frame header straddling the read buffer's boundary failing the frame:
Stream::read is allowed to return less than asked for, and only the payload
was reading in a loop.

ws::WebSocket::set_read_timeout() lets a server handler bound its own reads,
and WebSocketClient::set_read_timeout() now reaches an already-open connection
instead of only seeding the next connect().

The read timeout macro splits in two. A client waits forever by default -- a
read timeout is the caller's tool for taking back control, not a liveness
check, which is ping/pong's job -- while a server keeps the 300s that reclaims
a worker from a peer gone quiet. Defining the old name still sets both.

Also record a reason on the two WebSocketSSLStream::read failure paths that
returned -1 without one, so get_error() cannot report a previous call's
timeout, and make SocketStream's read timeout atomic now that it can be
changed while a read is in flight.
2026-09-03 14:32:13 -04:00
Avionic Harshit
9e2e33da56 let EXTRA_CXXFLAGS override -fsanitize=address (#2577) 2026-09-03 10:22:00 -04:00
Jhen-Jie Hong
7d53a31d23 Don't compress a response whose handler already set Content-Encoding (#2575)
* Don't compress a response whose handler already set Content-Encoding

* Don't compress a pre-encoded response served from a file

The guard that stands down when a response already names a content coding
covered the responses that settle their coding in `apply_ranges()`, but a
file-backed one settles it in `static_file_encoding()`, which asked the
content-type overload and so never saw the field. With static file
compression enabled, a mount point naming the coding for a tree of
build-time compressed assets, and a handler setting the field on a
`set_file_content()` response, both had their stored bytes compressed a
second time and a second `Content-Encoding` field line appended.

A file-backed response has not been given a content type by the time its
coding is decided, which is the only reason it could not go through
`encoding_type()`. It takes the type as an argument now, so both paths share
the one guard instead of carrying a copy each.

`Response::content_encoding_` becomes `content_coding_`, after what it
holds. It names the coding chosen for the body, which is what its own
comment already called it, while the old name read as the value of the
`Content-Encoding` field whose presence is exactly what forces the coding to
`None`.

README gains the behaviour, including the part that stays with the handler:
`Vary` is added only to a coding the server chose, so a handler that picks a
representation from `Accept-Encoding` has to add the field itself.

---------

Co-authored-by: yhirose <yuji.hirose.bug@gmail.com>
2026-08-31 20:44:12 -04:00
yhirose
9ce15a14e5 Reject Digest challenges missing realm or nonce (RFC 7616 §3.3)
parse_www_authenticate() accepted any WWW-Authenticate: Digest
challenge that carried at least one auth-param, so a server sending
e.g. Digest qop="auth" with no realm/nonce would make it through.
make_digest_authentication_header() then dereferences auth.at("realm")
and auth.at("nonce") unconditionally, throwing std::out_of_range with
no try/catch on the retry path, which terminates the client process.

Now require both realm and nonce before treating a Digest challenge as
usable, same as if no Digest challenge were present at all.
2026-08-28 23:36:40 -04:00
yhirose
139f30e0f1 Compress static file responses behind an opt-in (Fix #2545) (#2572)
* Drop the claim that small bodies skip compression

There is no size threshold anywhere in the compression path.
encoding_type() gates on the content type and Accept-Encoding only, and
apply_ranges() compresses whatever body it is given, so a two-byte
text/plain response comes back gzipped at 22 bytes.

Say what actually happens and leave the decision to the handler.

* Compress static file responses behind an opt-in (Fix #2545)

apply_ranges() runs the compressor inside the branch it takes when
res.body is non-empty. A response served from a file leaves res.body
empty and sets content_length_, so it took the other branch, which
writes Content-Length and returns; encoding_type() was computed before
the split and never consulted on that side. The same bytes handed to
set_content() came back gzipped, which left set_mount_point() and
Response::set_file_content() as the one path that missed out.

Add Server::set_static_file_compression(), off by default so nothing
about an existing server changes. When it is on, the file-backed
provider is run through the compressor into res.body ahead of the rest
of apply_ranges(), so the response is framed the way set_content()
already frames one: it keeps its Content-Length, and HEAD still reports
the size a GET would return.

Ranges are answered from the identity representation, since RFC 9110
applies Range after content coding and slicing a compressed body would
mean compressing the whole file first. The ETag carries the coding it
belongs to, so a client that cached the compressed form revalidates
against its own validator rather than the identity one. Both the ETag
and the body take their coding from static_file_encoding(), so the two
cannot disagree.

Providers registered with set_content_provider() are left alone. zlib
buffers until its window fills, so running one through a compressor
would hold back writes that a caller expects to reach the peer as they
are produced.

The compressed bytes stay in memory until the response has been
written, so the peak cost scales with requests in flight.
set_static_file_compression_max_length() bounds it, defaulting to 4MB.

* Add a minimum size for static file compression

Compressing a file that already fits in a single 1500-byte MTU does not
get it to the client any sooner, and a file of a few bytes comes back
larger than it went in once gzip's header and trailer are added. Every
other server draws this line: nginx's gzip_min_length, Caddy's
minimum_length, IIS's minFileSizeForComp, CloudFront's 1000-byte floor.

The note this replaces told callers to decide in the handler. A response
served through set_mount_point() has no handler to decide in, so the
floor has to live in the server. It defaults to 1400 bytes, the size
that fits inside one MTU with room for headers.

set_static_file_compression_min_length() moves it, and
CPPHTTPLIB_STATIC_FILE_COMPRESSION_MIN_LENGTH sets the default at
compile time. The empty-file case keeps its own early-out so that a zero
floor still cannot turn an empty body into a 20-byte gzip stream.

The two bounds now read as a pair, so the documentation says what each
one is for: the lower bound is about what is worth compressing, the
upper bound about what one request is allowed to cost.

Every file under test/www except 1MB.txt is below the default floor, so
the tests that need a small file compressed lower it explicitly.
2026-08-27 17:19:43 -04:00
yhirose
b4ec1bb1de Respect quoted-strings when splitting header parameters (Fix #2568) (#2573)
parse_disposition_params() and extract_media_type() both split on every
';' and then on every '=', with no idea that a parameter value can be a
quoted-string. RFC 9110 5.6.6 allows ';' and '=' inside one, so
filename="report=v2.pdf" came out as v2.pdf", and filename="a;b.txt" was
truncated at the semicolon and left a bogus parameter behind.

The same defect reached the boundary. RFC 2046 5.1.1 allows '=' in a
boundary, which forces a sender to quote it, so the common MIME form
boundary="----=_NextPart_000_0000_01D9" parsed as
_NextPart_000_0000_01D9".

Add split_unquoted(), which is split() with the one extra rule that a
delimiter inside a quoted-string is not a delimiter, and route both
parameter parsers through it. The key/value split, duplicated verbatim
in the two of them, moves into divide_param_pair(). That one divides at
the first '=' without tracking quotes: 5.6.6 makes the key a token, so
no quote can precede the separator, and reusing divide() keeps this off
the per-byte scan.

A backslash stays an ordinary character here. Both browsers and
httplib's own sender percent-encode '"' rather than escaping it, and
recognizing a quoted-pair without also unescaping it would just trade
one wrong value for another.
2026-08-27 17:18:29 -04:00
yhirose
c58061ea81 Fail a content provider that makes no progress
write_content_with_progress() advances its offset only by what the provider
writes, so a provider that reported success without writing anything and
without calling done() was handed the same offset and length again on the next
pass. With the peer still connected it spun there, re-entering the provider as
fast as the loop could run.

make_file_body()'s provider was one way to reach this and was fixed in #2566,
but any user-supplied provider can do the same. Treat a pass that makes no
progress as a short body, which is how done() called early is already handled.
2026-08-26 23:05:19 -04:00
Robert Miller
794a997d8c Fail make_file_body()'s provider when the file is short (#2566)
make_file_body() measures the file once and that length is already the
response's Content-Length. The provider re-opens the file by path on each
call, so if the file has been truncated since, the read comes up empty and
the provider returned true without writing. write_content_with_progress()
advances its offset only by what was written, so it called the provider
again, got nothing again, and kept spinning until the peer gave up.

Return false instead, as every other failure in this provider does.
2026-08-26 22:44:04 -04:00
yhirose
e96a52e9dd Ignore empty list elements in the Accept header (Fix #2567)
parse_accept_header() rejected any Accept value with a leading, trailing
or doubled comma, and Server::process_request() validates Accept before
routing, so "Accept: text/html," was answered 400 Bad Request on every
route.

RFC 9110 Section 5.6.1.2 requires a recipient to parse and ignore empty
list elements in a #rule list, so those values are legal. split() already
trims each element and skips the empty ones, which made the guard inside
the callback unreachable as well; drop both and let the empty elements
fall away. The header length limit bounds how many a sender can send, so
ignoring all of them cannot be used as a denial-of-service vector.

get_combined_header_value() keeps skipping empty field lines, but that
skip is no longer observable through a request now that a stray comma
parses cleanly, so it gets its own test.
2026-08-26 22:37:22 -04:00
yhirose
2addb41089 Stop a throwing user callback from terminating the server (#2564)
Server::process_request() wraps only routing() in a try/catch.
Everything else the user supplies runs outside it:

- the content provider, from write_response_core()
- post_routing_handler_, error_handler_, logger_
- expect_100_continue_handler_
- a WebSocket handler, and pre_routing_handler_ on the upgrade path

An exception from any of those unwinds out of process_and_close_socket()
into the task queue, which calls the job without a catch, so it reaches
the top of a pool thread and terminates the process. One handler that
throws takes down every other connection the server is holding.

Add Server::serve_guarded() and run the serving loop through it in both
process_and_close_socket() overloads. The exception is not turned into a
500: by the time a content provider runs, the status line and headers
are already on the wire, so there is nothing left to replace. Report it
through the error logger as Error::UserCallbackException and drop the
connection, which is what the peer observes regardless. Requests on
other connections are unaffected, and the socket is still drained and
closed - which unwinding used to skip on the non-SSL path, since
drain_and_close_socket() sits after the call rather than in a scope
guard.

The error logger is a user callback too, so the report inside the guard
is itself wrapped: a throwing logger must not be able to open the guard
back up.

Adds ServerExceptionTest: a throwing content provider, post-routing
handler, WebSocket handler and error logger, plus the content provider
case against SSLServer, each checking that a later request on a new
connection still succeeds. Every test runs the server on a single worker
thread, so a guard that catches the exception but still loses the thread
shows up as the follow-up request never being served. Note that all of
them abort the test binary without this change - which is the bug, but
it means a regression here fails the run rather than one test.
2026-08-26 01:16:48 -04:00
yhirose
ae417b405a Do not let a zero-length write end a chunked body (#2563)
write_content_chunked()'s sink treated "the provider wrote nothing" as
"the provider has finished":

    data_available = l > 0;

so sink.write(p, 0) ended the loop. Only done()/done_with_trailer()
emit the terminating zero-length chunk, so the body was left
unterminated - and the function still returned Success, because the
post-loop check only reports the is_shutting_down() case. The peer waits
for a last chunk that never arrives, and on a keep-alive connection
anything written next is parsed as a chunk-size line.

A provider reaching a pass with nothing to hand over is ordinary:
popping an empty buffer off a queue, or a compressor that has consumed
its input without producing output yet. It is not the end of the
message.

Ignore zero-length writes instead. A zero-length chunk is the terminator
in chunked coding, so it must never be emitted mid-body either way, and
data_available is now controlled only by done()/done_with_trailer().
This matches write_content_without_length(), where the sink's write
never ends the body.

The old behaviour cannot have been relied on: it produced an
unterminated response, so a provider using it never worked in the first
place.
2026-08-26 01:16:38 -04:00
yhirose
bc7e51dbb9 Give DataSink's optional callbacks safe defaults (#2562)
DataSink has four callbacks, but only write is assigned by every writer
that hands a sink to a content provider:

  write_content_with_progress()    write, is_writable
  write_content_without_length()   write, is_writable, done
  write_content_chunked()          all four
  send_with_content_provider...()  write
  get_multipart_content_provider() write, done  (cur_sink)

A provider that calls one of the unassigned ones invokes an empty
std::function and throws std::bad_function_call. Nothing on that path
catches it, so it unwinds out of the thread running the provider and
terminates the process. The README's own idiom is enough to hit it:
sink.done() is documented for the without-length overload, but a
provider registered through set_content_provider() with a length gets a
sink where done is empty.

Default the three optional callbacks instead. A sink is writable unless
a writer says otherwise, and a sink that cannot carry trailers still has
to finish, so done_with_trailer() falls back to done(). Capturing this
for that is safe because DataSink is neither copyable nor movable.

A no-op done() alone would only trade the crash for a hang on the two
length-framed paths: both loop until offset reaches the promised length,
so a provider that reports itself done without writing would be called
again immediately, forever. Both now record that the provider finished
and stop, and the short body is reported as a write error. The client
path gains that check for the compressor-failure exit as well, which
used to send a truncated request body without reporting anything.

cur_sink in get_multipart_content_provider() now forwards is_writable
from the outer sink, so a provider item asking whether it may keep going
gets the stream's answer rather than the default.
2026-08-26 01:16:17 -04:00
yhirose
f9c205632d Fix accept() error handling on Windows (#2561)
The accept loop in Server::listen_internal() classified accept() failures
by reading errno, but Winsock reports them through WSAGetLastError() and
never touches the CRT errno. Both retry branches were therefore dead code
on Windows, and every accept() failure fell through to the fatal path,
which closes the listening socket and ends listen().

That is reachable in normal operation: a peer resetting a pending
connection before it is accepted is enough, and descriptor or buffer
exhaustion shows up under load. One such event stopped the server from
accepting anything again.

Add is_accept_resource_error() and is_accept_transient_error() next to
is_connection_error(), which already abstracts the same errno vs
WSAGetLastError() difference, and use them in the accept loop.

The POSIX sets are widened to match the Windows ones rather than being
left as they were: ECONNABORTED is the POSIX spelling of the aborted
pending connection that motivates this, and ENFILE, ENOBUFS and ENOMEM
are resource exhaustion in the same sense as EMFILE.
2026-08-26 01:16:01 -04:00
yhirose
19352ae929 Cap the received multipart boundary at RFC 2046's 70 characters (#2565)
parse_multipart_boundary only rejected an empty boundary, so a request could
declare one as long as a header line is allowed to be. A stock server accepts
up to 8146 bytes there, which is what CPPHTTPLIB_HEADER_MAX_LENGTH leaves after
"Content-Type: multipart/form-data; boundary=".

FormDataParser searches the body for "--" + boundary + CRLF with a plain
substring scan. buf_find scans for that delimiter's first byte, always '-', and
at every position that matches calls start_with, which compares until the first
mismatch. A body of '-' makes every position a candidate, and a boundary of '-'
makes each candidate compare the whole delimiter before failing at the CRLF. The
worst case is the product of the body length and the boundary length, and only
the first factor was bounded.

Measured by driving the parser directly in 16 KB reads, Apple clang 17 at
-O2 -DNDEBUG, best of three runs on an otherwise idle machine. 100 MB of '-',
the default payload limit, costs 2.59 s of CPU with a 70 byte boundary and
281.83 s with an 8147 byte one, a factor of 109. The same shape shows at 8 MB:
0.211 s, 3.081 s, 11.359 s and 22.091 s for boundaries of 70, 1024, 4096 and
8147 bytes.

RFC 2046 5.1.1 caps a boundary at 70 characters, so honoring that limit bounds
the multiplier too. The limit applies to the value after unquoting, so a quoted
70 character boundary stays valid. Only the server receive path parses a
boundary out of a Content-Type, so what clients may send is unaffected, and the
boundaries the library generates itself are 45 characters.
2026-08-26 01:15:22 -04:00
yhirose
2afe933103 Send the Connection: close header the multipart test comment describes
expect_split_multipart_ok() carries a comment saying the request sends
"Connection: close" so the response drain ends as soon as the server has
answered, but the header itself never made it into the request, so both
callers kept idling until the read timeout instead.

Add the header. EpilogueSplitAcrossReadsIsIgnored and
InitialBoundarySplitAfterLongPreamble each drop from about 3.1s to about
0.11s.
2026-08-26 01:06:21 -04:00
yhirose
bc58e6e9ac Bound the multipart parser's buffer while it waits for a boundary (#2557)
* Bound the multipart parser's buffer while it waits for a boundary

FormDataParser accumulated the entire request body whenever the declared
boundary never appeared in it. State 0 returned without erasing anything, so
the buffer grew to the full payload (100 MB by default) and buf_find rescanned
all of it on every 16 KB read. The cost grew with the square of the body size:
50 MB of '-' took 198 s of CPU on one core, and the buffer pinned the body in
memory for the whole request. One unauthenticated request was enough, and the
parser runs for any multipart request even when the handler never looks at the
parsed result.

State 0 now keeps only the last dash_boundary_crlf_.size() - 1 bytes while it
waits, which bounds both the memory and the rescan without capping how long a
preamble may be. The same 50 MB body now takes 0.14 s and the buffer stays at
one read plus the boundary. A boundary split across reads still parses, which
is what de5a255 (#2159) gave up this erase for.

State 4 buffered without bound in the same way when a boundary was followed by
neither CRLF nor "--". No further data can make such a body valid, so it now
fails right away. That is only safe because the close-delimiter branch moves to
a new state 5 that discards the epilogue: it used to stay in state 4, so an
epilogue arriving in a later read fell into this same branch. An epilogue
beginning with CRLF was then parsed as a new part and the request was rejected
with 400, which state 5 fixes as well.

Affected since v0.23.0, where de5a255 replaced the erase that had kept the
buffer in check.

* Skip buffering the multipart epilogue

Once the close delimiter has been parsed the parser is in state 5 and discards
whatever follows, but it still copied each epilogue read into the buffer before
erasing it. Return before buffering so a large epilogue spread across several
reads is dropped without being copied in at all.

* Clean up the multipart parser tests and the state 4 branch

Review follow-ups on top of the previous two commits, no behavior change.

- Move the four new tests next to the rest of MultipartFormDataTest. They
  had landed in the middle of the RedirectTest block.
- Use bind_to_any_port instead of the fixed PORT, as AGENTS.md requires for
  newly added servers. NoInitialBoundaryParsingIsNotQuadratic holds its port
  for a couple of seconds, which matters when the suite is run sharded.
- Send "Connection: close" from expect_split_multipart_ok. The server kept
  the connection alive after answering, so the response drain idled until the
  client read timeout; both tests drop from about 3s to about 0.11s.
- Drop the dead `dash_.size() > buf_size()` guard in state 4 and flatten the
  nested else. The check above it already guarantees two buffered bytes, and
  both CRLF and "--" are two bytes, so it can never fire. Removing it is what
  makes the new comment's claim readable straight off the code.

* Rename the timing test's locals to avoid a Windows macro

MSVC's <rpcndr.h>, pulled in by <windows.h>, defines `small` as `char`, so
`auto small = ...` failed to compile on the Windows jobs. Same class of
problem as the std::min / std::max collision.
2026-08-25 22:55:52 -04:00