These commits are when the Protocol Buffers files have changed: (only the last 100 relevant commits are shown)
| Commit: | 0f5430c | |
|---|---|---|
| Author: | Florentin DUBOIS | |
| Committer: | GitHub | |
test(e2e): cover real gRPC lifecycles (#1850) Exercise unary and bidirectional gRPC through HTTPS/H2 and transparent TCP. Validate real deadline and cancellation cleanup plus fresh frontend reconnection in a dedicated workflow. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | d2cea3f | |
|---|---|---|
| Author: | Florentin Dubois | |
test(e2e): cover real gRPC lifecycles Exercise unary and bidirectional gRPC through HTTPS/H2 and transparent TCP. Validate real deadline and cancellation cleanup plus fresh frontend reconnection in a dedicated workflow. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 7e2ebc0 | |
|---|---|---|
| Author: | Florentin DUBOIS | |
| Committer: | GitHub | |
fix(command): transfer pending commands during main upgrade (#1846) The main-process hot upgrade rebuilt an empty command Hub while worker work could still be in flight. The original command client then received EOF, and a later worker response reached the replacement as an unknown task even though its data-plane request could still finish. This change transfers the complete control-plane ownership graph to the replacement main: connected clients, client and worker channel buffers, active and queued tasks, worker-response routes, subscribers, worker pending queues, stopped workers still referenced by a task or route, task deadlines, audit clocks, and future-ID counters. The replacement restores task state directly without replaying operator requests or re-scattering work. The handoff uses an explicit V2 `PREPARED` / `COMMIT` / `ACTIVATED` protocol: - the replacement validates and registers inherited descriptors while paused, without command or worker data I/O; - pre-commit rejection closes and reaps the candidate, restores `CLOEXEC`, and lets the unchanged old Hub continue; - the first attempt to send `COMMIT` irreversibly fences the old Hub, so an ambiguous post-commit failure cannot execute work twice; - after a complete commit, the replacement activates the Hub, schedules restored sessions once, publishes PID and systemd readiness, preserves the initiating client's terminal response under backpressure, and acknowledges activation before the old main exits. Compatibility fails closed before adoption. A running main from before this protocol cannot export the missing Hub state, so the first deployment requires a controlled service restart after pending commands are drained and state is saved. Subsequent V2-to-V2 hot upgrades transfer pending commands. A V2 sender rejects a legacy candidate before freezing live ownership; a legacy sender remains subject to its existing generation, audit, and child-reaping limitations when a V2 candidate refuses its old envelope. `SIGTERM` observed before commit aborts and reaps the candidate, then the old event loop consumes the preserved stop intention. Operators should signal the service control group. A signal aimed only at the old PID after the final pre-commit observation remains outside this contract, and a failure after commit is fail-stop rather than crash recovery. The branch is based on `9561b24e61c23511326ffeb547b0b7fe51ee7055`, including the deadline and cancelled-route fixes from #1838 and the PROXY protocol corrections merged in #1844. Channel snapshot validation accepts all five defined `Ready` bits in both serialized readiness and interest, preserves them exactly, and rejects the first unknown bit `0x20`. This covers the Linux readiness word `WRITABLE | HUP | WRITE_CLOSED` (`0x001a`) without special-casing that observed value. Descriptor-cleanup tests observe the original command-channel and SCM socket peers instead of probing released numeric descriptors or requiring immediate EOF. EOF is the only success outcome within a fixed one-second bound: data, other errors, and expiration fail. This tolerates transient descriptor copies across a concurrent fork-to-exec window while retaining a durable duplicate of either endpoint still fails the test. The exact transient holder in the original CI jobs was not instrumented, so that scheduler-level cause remains unverified. Validation completed locally with Rust 1.93.1 and at most four build jobs: - `cargo test -p sozu-command-lib --locked`: 244 unit tests, 9 integration tests, and 2 doctests passed; 3 existing doctests were ignored. - `cargo test -p sozu --locked --test upgrade_fd_inheritance_e2e -- --ignored`: failed 0/1 on the earlier snapshot mask with `invalid channel snapshot readiness bits: 0x001a`, then passed 1/1 after accepting `WRITE_CLOSED`. - the snapshot regression exercises all 32 subsets of the five defined bits in both readiness and interest; it rejects `0x20` in each field. Removing `WRITE_CLOSED` makes the test fail on `0x10`. - `cargo test -p sozu --locked --no-default-features --features crypto-openssl,opentelemetry,splice,simd --verbose`: 152/152 library tests passed and every integration and doctest target completed successfully. - `cargo test -p sozu --locked --no-default-features --features fips,opentelemetry,splice,simd --verbose`: 152/152 library tests passed and every integration and doctest target completed successfully. - retaining durable duplicates of both the command channel and SCM endpoint makes each corrected cleanup oracle fail with both EOF deadlines; restoring the candidate makes each targeted test pass 1/1. - the accepted-handoff, rejected-candidate, SIGTERM-during-PREPARE, and descriptor-inheritance process tests pass 1/1 on the rebased head. - the compatibility matrix passes 1/1 in both directions between the rebased candidate and the frozen pre-protocol `d7b5f14d` binary. This verifies those two transition paths rather than certifying every historical version pair; the legacy generation, audit, and child-reaping limitations described above remain observable. - `cargo clippy --all-targets --locked --no-default-features --features crypto-openssl,opentelemetry,splice,simd -- -D warnings` and the equivalent FIPS matrix passed. - `cargo +nightly fmt --all --check` passed. - `.github/scripts/check_doc_citations.py --base 9561b24e61c23511326ffeb547b0b7fe51ee7055` passed all 37 citations, 1,078 test names, and 15 pinned excerpts. - `env -u RUST_LOG ... cargo +1.93.1 test -p sozu --locked --features opentelemetry,splice,simd --test sigterm_soft_stop_e2e -- --ignored`: 1/1 passed. The local 27/27 missing-access-log result was caused by ambient `RUST_LOG=warn`, which GitHub CI does not set; the test requires its configured `info` access records. GitHub CI checks are pending on the rebased commit `a8bdd54b8833060576674a8c2aff0374f6e8083b`. Fixes #1832 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 7a508dd | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
fix(command): transfer pending commands across main upgrade Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 3fdf940 | |
|---|---|---|
| Author: | Florentin DUBOIS | |
| Committer: | GitHub | |
fix(metrics): expire metric-detail leases from the worker event loop (#1834) Fixes #1831 ## Cause CAUSE: `lib/src/server.rs` `Server::notify` (the only production `lease_tick` call, reached only from command-channel messages) -> no event-loop path ran the janitor -> SYMPTOM: an abandoned `SetMetricDetail` lease outlived TTL + 5 s until the next worker command, regardless of data-plane traffic. ## Fix - The janitor body moves into `tick_metric_detail_leases(now)` (`lib/src/server.rs`), called once per `Server::run` iteration (before `send_queue`, so the `lease_tick_expired` event leaves on the same turn) and still from `notify` (keeps the single clock read shared with `lease_apply`). `lease_tick_due` still gates the table walk to once per 5 s. The loop wakes at least once per 1 s `poll_timeout`, so expiry is bounded by TTL + 5 s + 1 s with no traffic and no command. - Docs: `lib/src/metrics/mod.rs` comments, `command.proto` lifecycle comment, `doc/configure.md` (the old "at most one `ttl` window" was not true), `doc/sozu-top.md`, CHANGELOG (Unreleased / Fixed). ## Regression test `test_abandoned_lease_expires_without_a_command` (`e2e/src/tests/metrics_lifecycle_tests.rs`): real worker, apply a 1 s `DETAIL_BACKEND` lease, then write nothing more on the command channel while HTTP requests keep flowing; passively read the channel until a `lease_tick_expired` `METRIC_DETAIL_CHANGED` event (backend -> cluster) arrives within TTL + 8 s. - Red on `bd19a78e` (unfixed): `test ... test_abandoned_lease_expires_without_a_command ... FAILED` (`stability check FAILED: iteration 1 of 3`) - Green with fix: `test result: ok. 4 passed; 0 failed` (all `metrics_lifecycle` e2e, `--features opentelemetry,splice,simd`) ## Validation - `cargo +nightly fmt --check` -> exit 0 - `cargo +1.93.1 clippy --all-features --all-targets --locked -- -D warnings` -> exit 0 - `RUSTDOCFLAGS="-D warnings" cargo +1.93.1 doc --no-deps --locked -p sozu-lib -p sozu-command-lib` -> exit 0 - `cargo +1.93.1 test -p sozu-command-lib --all-features --locked` -> exit 0 - `cargo +1.93.1 test -p sozu-lib --lib --all-features --locked` -> 1591 passed, 1 failed: `protocol::pipe::tests::splice_readable_eof_drains_kernel_pipe_bytes_before_closing` hit its own host precondition guard (`splice pipe capacity is 8192 bytes ... exceed fs.pipe-user-pages-soft (16384) ... this is a host limit, not a Sozu regression`) on a shared, heavily loaded host; three isolated reruns: pass / same precondition / pass. `pipe.rs` is untouched here. - `python3 .github/scripts/check_doc_citations.py --root . --base origin/main` -> exit 0 - Full e2e suite not run locally (left to CI). #1828 (same-ID cluster incarnation) is deliberately not in this PR: it needs a design decision, see the issue. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 579af63 | |
|---|---|---|
| Author: | Florentin Dubois | |
fix(metrics): expire metric-detail leases from the worker event loop The lease janitor ran only at the top of `Server::notify`, which only a worker command reaches. A lease whose owner disappeared (a crashed `sozu top`) therefore kept the worker's metric cardinality elevated past its TTL until the next control request, regardless of data-plane traffic. Move the janitor into `tick_metric_detail_leases` and call it once per event-loop iteration, before the response queue is flushed, as well as from `notify`. The loop wakes at least once per one-second poll timeout, so an abandoned lease is retired and its `lease_tick_expired` event is pushed within TTL + 5 s + one poll timeout with no further command. The new real-worker regression applies a one-second lease, keeps requests flowing, writes nothing more on the command channel and waits for the expiry event; it fails on the previous code. Fixes #1831 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | d2d95fd | |
|---|---|---|
| Author: | Florentin DUBOIS | |
| Committer: | GitHub | |
docs: fix stale flood-window, idle-timeout and rejected-per-IP statements (#1816) Docs and comments only, no behaviour change. Each statement was checked against the code on `main` (6c3b8e89). | Statement | Code | Fix | |---|---|---| | `h2_flood_detector.rs` comments size the connection window at 1 MiB | `ENLARGED_CONNECTION_WINDOW = 16 * 1024 * 1024` (`h2.rs`), stream window `DEFAULT_INITIAL_WINDOW_SIZE = 65535` | Stream-0 example: Sōzu acks every 8 MiB of 16 MiB. Glitch budget: a cancelled upload is bounded by the 64 KiB stream window, about four glitches in 16 KiB frames (same arithmetic as the refused-stream comment in `h2.rs`), not 64. Chained-proxy test doc now describes the 1 MiB peer it models, and notes that Sōzu's 16 MiB window is sparser. The test body is unchanged. | | `command.proto` (×2) and `command/src/config.rs`: `h2_stream_idle_timeout_seconds` "Default: 30" | `get_h2_stream_idle_timeout` (`http.rs`, `https.rs`): `unwrap_or(back_timeout.max(30))`, explicit `s.max(1)` | Unset, it inherits `back_timeout` with a floor of 30 (so 30 with the default `back_timeout`); an explicit value wins, and `0` means 1. `doc/configure.md` already said this. | | `connections.rejected_per_cluster_ip` counts TCP rejections (`doc/configure.md`, `doc/rate-limit-design.md` §3.5/§4), and e2e asserts it (`rate-limit-design.md` §5, `cluster_ip_limit_tests.rs` header) | Sole increment: `answers.rs` `set_default_answer`, status `429`. The TCP gate (`tcp.rs` `connect_to_backend` → `TooManyConnectionsPerIp` → `handle_connection_result` → `Close`) increments nothing, and no e2e test reads the counter | Docs now say it counts HTTP/HTTPS 429 answers only (per-IP or per-subnet gate), and that TCP rejections are not counted. The false "confirms the metric" claims are dropped. | Not changed: - The TCP per-(cluster, IP) rejection path has no counter. Adding one would be a behaviour change and needs its own decision. - The remaining `CONN_RETRIES` mentions are historical ("replaces the former constant") and accurate. - `BackendMap::max_failures` (`lib/src/backends.rs`) is set to 3 but nothing in the tree reads it. It is a dead public field and is left untouched here. Validation: `cargo +nightly fmt --check`, `cargo +1.93.1 clippy --all-features --all-targets --locked -- -D warnings`, `RUSTDOCFLAGS="-D warnings" cargo +1.93.1 doc --no-deps --locked -p sozu-lib -p sozu-command-lib`, `cargo +1.93.1 test -p sozu-command-lib --all-features --locked` and `check_doc_citations.py --base origin/main` all pass. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 65ab6af | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
docs: fix stale flood-window, idle-timeout and rejected-per-IP statements - h2_flood_detector.rs: the connection window is 16 MiB (acked every 8 MiB), not 1 MiB; a cancelled upload's in-flight DATA is bounded by the 64 KiB stream window, about four glitches in 16 KiB frames. - command.proto, command/src/config.rs: h2_stream_idle_timeout_seconds inherits back_timeout floored at 30 when unset; 0 means 1. - doc/configure.md, doc/rate-limit-design.md, cluster_ip_limit_tests.rs: connections.rejected_per_cluster_ip is incremented only by the HTTP/HTTPS 429 answer; a TCP rejection is not counted and no e2e test asserts the counter. No behaviour change. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | bd787ac | |
|---|---|---|
| Author: | Florentin DUBOIS | |
| Committer: | GitHub | |
docs: align the example configuration and docs with recent changes (#1809) Audit of the docs, `bin/config.toml` and `os-build/config.toml` against everything merged on `main` from 2026-10-01 to 2026-10-03 (43 commits, #1724 through #1808). Most of the new keys were already documented by their own PRs (`max_connection_attempts`, `h2_stream_refusal_percent`, health check `mode`/`accepted_statuses`, the raised flood defaults, the 16 MiB connection window). This PR fixes what was still stale. | Area | Where | Was | Now | |---|---|---|---| | #1806 | `bin/config.toml`, `os-build/config.toml` answer comments | 503 only for "no backend available"; "504 means the backend timed out" | 503 also when every attempt failed; 504 = connected backend silent or frontend timeout while waiting; connect timeout fails over | | #1806 | `bin/config.toml` `max_connection_attempts` | connect timeout always fails over | a TCP session whose connect times out is closed | | #1806 | `doc/debugging_strategies.md` | 503 after a circuit breaker of 3 failed connections; ERROR log lines, one that no longer exists | `max_connection_attempts` (default 5), `backend.connect.retries_exhausted`, real TCP log levels and text | | #1806 | `doc/configure.md` `front_timeout` / `connect_timeout` rows | "inactivity for a request to connect" | time to establish a backend connection; HTTP fails over, TCP closes; H1 failover bound by `front_timeout` | | #1806 | `doc/configure.md` `client_timeout_during_response`, `backend_timeout`, `http.503.errors` | "connection-level slowness", "circuit breaker" | failover case included; established connection only; attempt budget | | #1806 | `LIFECYCLE.md` backend timer | 504 for any linked stream | established connection only | | #1806 | `cli.rs` help, `doc/configure_cli.md` | "answered 503" | "or a TCP session is closed" | | #1756 | `doc/configure.md` flood metrics | RST lifetime rows described as absolute ceilings; no `window_update_stream0_window` row | floor-and-ratio semantics; row added | | #1756/#1798 | `doc/architecture.md`, `doc/h2_mux_internals.md` | 6 / 13 `H2FloodConfig` thresholds | 14, with `h2_stream_refusal_percent` | | #1756 | `doc/configure_admin_ops.md` | "halve" example 50/25 (old defaults); H2 knobs HTTPS only | 1000/500; HTTP and HTTPS; `h2.flood.stream_refused` | | #1752 | `doc/configure.md` glitch tuning | "1 MiB window ... about 64" frames | 64 KiB stream window, about four | | #1805 | `doc/health_checks.md` | `remove` resets backends to healthy; DOWN logs an error; `load_balancing.rs::BackendList::healthy` | backends keep their last state (see note); warning; `BackendList::select_tiers` in `backends.rs` | | #1805 | `doc/metrics.md` | health-check counters "per cluster" | unlabelled | | #1805 | `doc/configure_cli.md` | no health-check section | short section, `--mode`, `--probe-timeout` vs global `--timeout` | | #1746 | `cli.rs` `--load-balancing-policy`, `bin/config.toml`, proto comment | HRW/MAGLEV key = UDP flow key | UDP affinity key (source IP, or IP and port) | | UDP | `doc/configure.md` | hot upgrade drains flows bounded by idle timeout; `udp.flows.evicted` reasons | soft-stop closes flows at once; listener removal/soft-stop/abort added | | #1770 | `doc/udp_simulation.md` | no per-source limits | in reconfig storms and the invariant list | | #1743 | `doc/lifetime_of_a_session.md` | keep-alive backend always reused | unless its backend was removed | | older | `command.proto` update messages | `h2_stream_shrink_ratio >= 1` | `>= 2`, as validated in `state.rs` | | #1806 | e2e shuffle-sharding doc comment | "three connection attempts" | `max_connection_attempts` (5 by default) | Reported here, not changed in this PR: - `sozu cluster health-check remove` does not reset backends a probe marked DOWN (`remove_health_check_state` skips the reset that `set_health_check_config(None)` does), so they stay out of rotation until the cluster is added again. The doc now describes this. It looks like a code bug. - `sozu top`'s H2 pane reads `h2.flood.violation.{rapid_reset,made_you_reset,ping,settings,priority,continuation}`, which the detector never emits (it emits `rst_stream_pre_response_lifetime`, `ping_window`, ...). Only `glitch_window` matches, and `h2.flood.stream_refused` is not shown. - `h2_flood_detector.rs` comments still refer to a 1 MiB connection window, and the `h2_stream_idle_timeout_seconds` proto comment says "Default: 30" where the value inherits `back_timeout`. - `connections.rejected_per_cluster_ip` is documented as also counting TCP rejections; only the HTTP 429 path increments it. Validation: - `sozu config check` passes on `bin/config.toml`, on a copy with every commented key under review uncommented (global and cluster `max_connection_attempts`, the H2 block, the health check block), and on `os-build/config.toml` with local state/socket paths. Negative controls (`max_connection_attempts = 0`; `expected_status = 200` next to `accepted_statuses`) are rejected. - `cargo +1.93.1 test -p sozu-command-lib --all-features --locked`: 235 passed. - `cargo +1.93.1 test -p sozu --bin sozu --locked`: 124 passed (including the `--load-balancing-policy` help tests). - `cargo +nightly fmt --all --check` and `python3 .github/scripts/check_doc_citations.py --root . --base origin/main` are clean. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 832d77c | |
|---|---|---|
| Author: | Florentin Dubois | |
docs: align the example configuration and docs with recent changes Audit the documentation, bin/config.toml and os-build/config.toml against the changes merged on main from 2026-10-01 to 2026-10-03 and fix every statement that no longer matches the code: - connection attempts (#1806): 503 and 504 answer comments in both example configurations, the TCP close on connect timeout, the front_timeout and connect_timeout rows, the access-log timeout tokens, http.503.errors, the three-attempt circuit breaker and its stale log lines in doc/debugging_strategies.md, the TCP close in the CLI help; - H2 flood protection (#1756, #1798): floor-and-ratio semantics of the RST_STREAM lifetime violation metrics, the missing window_update_stream0_window metric row, fourteen H2FloodConfig thresholds, the admin-ops worked example against the current defaults, h2.flood.stream_refused, h2_stream_refusal_percent among the listener-update fields; - health checks (#1805): health-check remove keeps the last health state, DOWN logs a warning, the fail-open citation, unlabelled health_check counters, a configure_cli.md section, the TCP-mode timeout comment; - UDP (#1746, #1770): HRW/MAGLEV key on the affinity key rather than the flow key, soft-stop closes flows at once, udp.flows.evicted reasons, per-source limits in the simulation doc; - pooled connections to a removed backend are not reused (#1743); - the h2_stream_shrink_ratio floor of 2 in the update-message comments. Documentation only, no behaviour change. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | e959838 | |
|---|---|---|
| Author: | Florentin DUBOIS | |
| Committer: | GitHub | |
fix(mux): fail over to another backend on connect timeout (#1806) Fixes #1800 ## Summary A backend that accepts no connection (blackholed: SYN dropped, accept queue full) hit `connect_timeout` and the request was answered `504`. The timeout path marked the connection `Disconnecting`, so the dead-backend branch of `Mux::ready_inner`, the only place a connect failure is counted, never saw it. The backend's retry policy was never armed and the load balancer kept picking it. - **Connect timeout counts as a connect failure.** `Mux::timeout_inner` now flags a still-`Connecting` connection's readiness `ERROR`, the way the kernel flags a refused dial. `Server::timeout` then runs the `ready` pass that only it can trigger (new `ProxySession::needs_ready_pass`, default `false`). That pass reuses the existing dead-backend branch: it counts the failure (`failures`, `retry_policy`, `backend.connections.error`, `backend.down`), closes the connection, and re-links the request through `EndStreamAction::Reconnect`. The request sent no byte, so it is safe to resend. A response timeout on an established connection still answers `504`. - **A WRITABLE event no longer marks a connecting socket connected.** `SO_ERROR` (`take_error`) is read first, then `getpeername(2)`. An error is a connect failure. `ENOTCONN` means the WRITABLE was raised by Sōzu itself, not by the kernel, so the connection waits for the kernel's completion edge. This fixes a separate defect you can observe on its own: on `main`, a dial from an `http2 = true` cluster to a blackholed backend was marked connected and then hung until `back_timeout`. - **A synchronous `connect(2)` failure is retried** instead of being answered `503` at once. - **Each retry skips the backends that already failed for this request** (`Stream::tried_backends`), for every load-balancing policy. The exclusion narrows the candidate set (`BackendList::select_with_key_excluding`). If every selectable backend was already tried, selection falls back to the plain behaviour, fail-open included, so the request retries one of them instead of being refused. A sticky cookie that names a tried backend is not followed. - **The retry budget is configurable.** `max_connection_attempts` counts every attempt, the first included. It replaces the hard-coded `CONN_RETRIES = 3`, and **the default is now 5**. It can be set: - globally: TOML and `ServerConfig` field 29; - per cluster: `Cluster` field 23, TOML, `sozu cluster add --max-connection-attempts`, and `sozu cluster connection-attempts set|unset` (query, then re-send with the field changed, like `cluster h2`). A cluster table column shows the value; - at runtime on every worker: `sozu connection-attempts set` (`Request.set_max_connection_attempts`, field 62). Like `connection-limit set`, this is not saved. Allowed range is `1..=255`. It is validated at config load, on `AddCluster` (main and worker), and on the runtime set. A cluster value wins over the global one, for both HTTP and TCP clusters. - **Failover is bounded by `front_timeout`.** The frontend's nominal timeout is set only on a request's first link, not on every re-link. The deadline armed then therefore bounds the whole failover, so a request can no longer run for `max_connection_attempts × connect_timeout`. When the deadline expires during failover, the request gets the frontend timeout's existing answer: `504`, with access-log message `client_timeout_during_response`. `doc/configure.md` recommends keeping `connect_timeout` well below `front_timeout / max_connection_attempts`, and explains how back-off grows. - **Socket error logged.** The socket error read by `take_error()` is logged at `debug!`, the same level as the synchronous-refusal path. **Status change for alerting:** when every attempted backend is unreachable (connect timeout or refused), the answer is now `503`. A connect timeout used to answer `504`. Requests that arrive during a back-off window can also fail fast with `503`. Operators alerting on `504` rates should watch `503`, `backend.connections.error` and `backend.connect.retries_exhausted` instead. With the default budget at 5, an idempotent request on a stale keep-alive connection can be replayed to up to 4 backends (`n - 1`), against 2 before. **Coordination:** `Request` field 62 is also used by the open PR #1300. That PR already needs renumbering, so this one keeps 62. **Library API break:** `sozu_lib::server::CONN_RETRIES` is removed, and new fields are added to `Cluster`, `ServerConfig`, `RequestType`, `FileConfig`, `FileClusterConfig`, `HttpClusterConfig`, `TcpClusterConfig`, `Stream` and `SessionManager`. Upgrade the main process, the workers and the CLI together before setting a per-cluster budget: an older CLI that patches clusters by read-modify-write drops the field. Docs: `doc/configure.md` (new "Backend connection failover" section, global key, metrics, shard and replay budget text), `doc/configure_cli.md`, `doc/health_checks.md`, `doc/lifetime_of_a_session.md`, `doc/rate-limit-design.md`, `lib/src/protocol/mux/LIFECYCLE.md` §7.3, `bin/config.toml`, and `CHANGELOG.md`. ## Existing tests touched Two existing tests now call `mark_connected(h1)` in their fixtures: `a_backend_timeout_leaves_no_linked_stream_behind_on_the_close_path` and `a_completed_backend_response_timeout_unlinks_through_end_stream_alone`. Both describe a backend that has already sent response bytes, but their fixture left it `Connecting`. With this change, a `Connecting` connection is treated as a connect timeout. No assertion changed, and neither test encodes the old `504`; they cover the response-timeout arms. An alternative would be to narrow the timeout branch to `Connecting && no linked stream has back.consumed`. I kept the fixture fix instead, because that extra condition cannot be reached in production. ## Tests New: `e2e/src/tests/backend_failover_tests.rs`, plus `connect_outcome_tests`, `exclusion_tests`, `a_connect_timeout_flags_the_dial_dead_instead_of_answering_504`, the `max_connection_attempts` config tests, and the state validation test. What each test proves: - The connect-timeout e2e tests use 3 healthy and 3 blackholed backends with `connect_timeout = 1`. The blackhole is a backlog-0 listener with a full accept queue; `test_blackholed_listener_times_out` checks that it really times out. Every request must answer 200, and the cluster-level `backend.connections.error` must be greater than 0. That counter shows the timeouts are now counted against the backend, which arms its retry policy. It does not measure a drop in how often dead backends are picked. - The refused-backend e2e tests use a refused/healthy alternation that drives the first request onto three refused backends in a row. The budget raise is what fixes them. The tried set itself is covered by `exclusion_tests`. ## CI fix: leaked mock backend threads CI on `9bd8bfa6` and `379b4a4e` failed `tests::h2_tests::test_h2_rapid_reset_triggers_goaway` in every cell ("received 0 frames"). The cause was in this PR's test module, not in production code. **Cause:** `Backends` in `e2e/src/tests/backend_failover_tests.rs` dropped its `AsyncBackend` handles without stopping them. That backend's thread (`e2e/src/mock/async_backend.rs`, the spawn loop) only exits on a stop message, so every dropped handle left a thread spinning for the rest of the test process. That was about two dozen threads, all started before `h2_tests` in name order. On a small CI runner they starved the rapid-reset worker: CI logged its 999 resets over 12 s, in batches of about 17 every 110 ms. Main's CI run took 0.16 s for the same resets, with identical per-stream timings. **Fix:** `impl Drop for Backends` now stops those backends. The rapid-reset test is unchanged. **Evidence:** I ran `taskset -c 0 cargo test -p sozu-e2e --locked -- backend_failover tests::h2_tests::test_h2_rapid_reset_triggers_goaway --test-threads=4 --nocapture` before and after the fix. - Before: about 24 threads were busy during the rapid-reset test (`ps -L`), and the 999 resets took 1.57 s instead of 0.1 s. - After: no busy leftover threads, and the 999 resets took 0.32 s. - The full e2e suite dropped from 3098 s to 1963 s. ## CI fix: worker events read as answers At `dff0285e`, `test_connect_timeout_fails_over_h2` failed in the crypto-openssl cell even though every request got its 200. The test's metrics query read a worker event (`content_type: Some("event")`) in place of its answer. The failover tests make backends fail on purpose, and a backend whose retry policy gives up on it is reported `BACKEND_DOWN` on the same command channel. Reads in these tests now skip events (`read_answer`). That event's arrival depends on randomized back-off, so I did not reproduce the race locally. After the fix, both connect-timeout tests passed 20 runs out of 20 under `taskset -c 0`, and 20 out of 20 on 2 CPUs loaded with 6 busy processes. Validation at head `6c7dd668`, after merging origin/main `320e86d6` (#1805; conflicts in `command/src/config.rs` and `doc/health_checks.md` resolved keeping both features): Verified: `cargo +nightly fmt --all -- --check` -> exit 0 Verified: `cargo +1.93.1 clippy --all-features --all-targets --locked -- -D warnings` -> exit 0 Verified: `cargo test -p sozu-lib --lib --all-features` -> `test result: ok. 1586 passed; 0 failed` Verified: `cargo test -p sozu-command-lib -p sozu --all-features --locked` -> exit 0, all `test result: ok` (227 command-lib unit tests) Verified: `cargo test -p sozu-e2e --features opentelemetry,splice,simd -- --test-threads=4` -> `test result: ok. 621 passed; 0 failed` Verified: `RUSTDOCFLAGS="-D warnings" cargo +1.93.1 doc --no-deps --locked -p sozu-lib -p sozu-command-lib -p sozu` -> exit 0 (An earlier run at `f8d4c18e` had one timing failure under load: `test_h1_chunked_encoding_edge_cases`. That test passed 3 times out of 3 in isolation, and the full `h1_security_tests` module passed with `--test-threads=4`.) Verified: `python3 .github/scripts/check_doc_citations.py --root . --base origin/main` -> exit 0 Red without the fix: - New e2e tests on `main` (`dc25d2ff`): - `test_connect_timeout_fails_over_h1`/`_h2`: `[200, 504, 200, 504, …]` (12×200, 12×504) and `backend.connections.error: 0`. - `test_refused_backends_never_exhaust_attempts_h1`/`_h2`: first request `503`. - `test_h2_backend_connect_failure_fails_over_h1`/`_h2`: `[200, 0, 200, 200, 0, …]`, where `0` is no answer within 10 s. - `test_failover_is_bounded_by_front_timeout_h1`/`_h2` (6 blackholed backends, budget 20, `connect_timeout` 1 s, `front_timeout` 3 s): - in a scratch copy of this branch with the per-attempt re-arm restored: no answer within 10 s on both frontends; - with the fix: `[504]` after about 3.09 s on both. - `try_connect_timeout_fails_over` now asserts `backend.connections.error >= 3`, one failure per blackholed backend. Round robin hits each of the six backends in the first six requests, so this is deterministic. - Budget and unit tests, in scratch copies of this branch with one part of the fix reverted at a time: - default set back to 3: `test_default_connection_attempt_budget_is_five` fails (`503`), and so does the config default test; - cluster override ignored: `…overrides_global` fails (`503`), and so do the runtime cluster update and the default test's per-cluster `four`; - runtime setter disabled: `…changes_at_runtime` fails (`200` where `503` is expected); - exclusion, fallback or sticky filter removed: the matching `exclusion_tests` fail; - `ENOTCONN` treated as connected: `a_dial_nobody_answers_is_pending` and the H2-backend e2e tests fail; - `Connecting` early return removed: `a_connect_timeout_flags_the_dial_dead_instead_of_answering_504` fails; - validation removed: the config and state range tests fail. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 6c7dd66 | |
|---|---|---|
| Author: | Florentin Dubois | |
Merge remote-tracking branch 'origin/main' into fix/backend-connect-timeout-failover Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 320e86d | |
|---|---|---|
| Author: | Florentin DUBOIS | |
| Committer: | GitHub | |
feat(health-check): add a TCP connect probe mode and configurable accepted HTTP statuses (#1805) Fixes #1801 ## What - **Probe mode** — `HealthCheckConfig.mode` (proto field 7, `HealthCheckMode`, default `HTTP`): - `HTTP`: same behaviour as before. - `TCP`: the backend is healthy iff the TCP connection is established within `timeout` (the socket becomes writable with no `SO_ERROR`, `peer_addr` succeeds). Nothing is sent, and the stream is deregistered and dropped once connected. That sends a FIN on an empty receive queue, or an RST if a banner-first backend (SMTP, SSH) has already sent bytes the probe never reads. - `uri` is ignored. A non-zero `expected_status` or a non-empty `accepted_statuses` is refused in TCP mode at every validation layer. - It runs on the existing `HealthChecker` loop, so tokens, timeout, thresholds, metrics, events and fail-open are unchanged. - **Accepted statuses** — `HealthCheckConfig.accepted_statuses` (proto field 8, `repeated HttpStatusRange { start, end }`, inclusive): - TOML and CLI entries: `"404"`, `"200-399"`, `"2xx"`, `"any"` (= `100-599`). Bounds must lie in `100-599` (RFC 9110 §15). - When the list is set, it replaces `expected_status`. A non-zero `expected_status` together with a list is refused by `validate_health_check_config`, which the CLI, the TOML loader, `ConfigState` and the worker all call. - **Final status only** — HTTP probes skip interim `1xx` responses (except `101`), on HTTP/1.1 and on h2c. The h2c walk keeps one HPACK decoder across header blocks, so accepting `1xx` or `any` never lets a `100 Continue` stand in for a final 500. - The 4096-byte response cap covers interim responses too: a `103 Early Hints` with more than about 4 KB of headers fails the probe. - **Compatibility**: - `uri` stays `required` on the wire and may be empty in TCP mode. The TOML `uri` and `--uri` default to `/`. - Both new fields carry `#[serde(default)]`, so JSON saved states written before them still load. `state_compat_v1_1_1` stays green. - A message without field 7 decodes as `HTTP`. - **CLI** — `sozu cluster health-check set --mode http|tcp --accepted-statuses …`. - `list` output changes: a new `mode` column, `expected status` becomes `accepted statuses`, and a legacy `0` is shown as `2xx` instead of `any 2xx`. - In TOML, `mode` is `"HTTP"`/`"TCP"`, case-sensitive like the other protocol enum keys. It is distinct from the UDP `udp.health.mode = "TCP_PROBE"`. ## Upgrade warning (also in CHANGELOG and `doc/health_checks.md`) **Upgrade the main process and every worker, and use the new CLI, before enabling `tcp` mode or `accepted_statuses`.** Otherwise the default probe silently comes back: `GET <uri>` (`/` if unset) accepting only 2xx, which marks down every 3xx/404/500 backend. This happens with: - an older worker running between `sozu upgrade main` and `upgrade worker`, which ignores fields 7 and 8; - an older `sozu` CLI doing read-modify-write (`cluster h2 enable|disable` re-sends `AddCluster` without fields 7 and 8); - an older binary loading a new saved state. Downgrading while a cluster uses either field is not supported. **Library API (BREAKING):** - `HealthCheckConfig` gains the public fields `mode` and `accepted_statuses`. - `FileHealthCheckConfig` gains `mode` and `accepted_statuses`. - `FileHealthCheckConfig::to_proto` now returns `Result`. - `ConfigError` gains `InvalidHealthCheck`. ## Pre-existing bug fixed: `health-check set` always panicked The subcommand's `--timeout` (u32) had the same clap id as the global `-t/--timeout` command timeout (u64), so every invocation panicked before sending anything: ``` $ sozu cluster health-check set --id app --uri /x thread 'main' panicked at bin/src/cli.rs:25:18: Mismatch between definition and access of `timeout`. Could not downcast to u64, need to downcast to u32 ``` The probe timeout flag is now `--probe-timeout` (clap id `probe_timeout`). **`--timeout` on `health-check set` is now the global command timeout, in milliseconds, not the probe timeout.** `cluster_health_check_set_parses_mode_and_accepted_statuses` parses both timeouts together; it failed with this panic before the rename. ## Protocol / security impact - Two new optional `HealthCheckConfig` fields; no existing field number or type changes. - The TCP probe writes nothing to the backend. - URI validation (CR/LF/NUL/C0) still applies in every mode. The leading `/` is only required in HTTP mode, the only mode that puts the URI on the wire. ## Tests | Test | What it covers | | --- | --- | | `test_health_check_tcp_mode_marks_blackholed_and_refused_backends_down` | Blackhole: full accept queue, so the probe times out. Refused: closed port, so the probe sees the socket error. Both go DOWN. This guards the connect verdict. | | `test_health_check_tcp_mode_keeps_500_and_404_backends_up` | Backends that accept but answer 500/404 stay up. This guards the mode switch. | | `test_health_check_http_mode_accepted_statuses_keep_404_backend_up` | `200-499` keeps a backend that answers 404 on the probe path up. | | `test_health_check_legacy_config_marks_404_backend_down` | A legacy config still accepts only 2xx. | | Unit tests in `lib/src/health_check.rs` | Status matching for lists, ranges and `any`; the 2xx default being replaced; `tcp_connect_outcome`; `http1_interim_responses_are_skipped_for_the_final_status`; `h2c_interim_headers_are_skipped_for_the_final_status`. | | Unit tests in `command/src/config.rs` | Range parser and its `Display` round trip; validation (mode, TCP URI exemption, TCP status refusal, bounds, exclusivity); TOML to `AddCluster`. | | Unit test in `command/src/state.rs` | `set_health_check_tcp_mode_with_status_field_rejected`. | | Unit test in `bin/src/cli.rs` | CLI parsing of the new flags and of both timeouts. | The e2e tests are in `e2e/src/tests/health_check_mode_tests.rs`. They read the worker's `health_check.{success,failure,down}` counters and skip the event responses the checker queues on the command channel. Note: the `Blackhole` helper duplicates the one in the in-flight #1800 failover branch (`e2e/src/tests/backend_failover_tests.rs`). Whichever PR lands second should deduplicate it. **Seen red.** On `main` the new tests do not compile, because the fields are absent. I also ran mutation rounds on a `git archive` copy of this branch: 1. TCP mode forced back to the HTTP probe: `tcp_mode_keeps_500_and_404_backends_up` fails. 2. `accepted_statuses` ignored: the accepted e2e test and the four `accepted_statuses_*` unit tests fail. 3. Exclusivity check removed: the `command/src/config.rs` health-check tests fail. 4. `tcp_connect_outcome` ignores socket errors: the blackholed/refused e2e test and `tcp_connect_outcome_reports_established_and_refused` fail. 5. `expected_status = 0` accepts any status: the legacy e2e test, `test_is_status_healthy_any_2xx` and two parse tests fail. 6. Interim skip removed: both interim tests fail. A decoder rebuilt per h2c block also fails `h2c_interim_headers_are_skipped_for_the_final_status` (`left: Some(false) right: Some(true)`): its `x-hint` header is outside the HPACK static table, so the final block references the dynamic table. 7. TCP-mode status refusal removed: the state test and the two `command/src/config.rs` health-check tests fail. ## Commands run (all exit 0 on `d743a112`, after merging `main` at `3aec7035`) ``` cargo +nightly fmt --all --check cargo +1.93.1 clippy --all-features --all-targets --locked -- -D warnings cargo +1.93.1 test -p sozu-lib --lib --all-features --locked # 1578 passed python3 .github/scripts/check_doc_citations.py --root . --base origin/main ``` On `4a9290f6`, the previous head: `sozu-command-lib` + `sozu` tests all passed and the health_check e2e tests passed (10). The full e2e suite passed (592/592) on `9f7d62a4`. Neither was rerun on `d743a112`, which changes only a unit test and docs on this branch, plus the merge of `main`. ## Docs `doc/health_checks.md` (modes, parameters, accepted-status syntax, interim responses, TOML case and the UDP `TCP_PROBE` naming, upgrade warning, `--probe-timeout` vs global `--timeout`, list output), `doc/configure.md`, `doc/metrics.md`, `bin/config.toml`, `CHANGELOG.md`. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 4a9290f | |
|---|---|---|
| Author: | Florentin Dubois | |
fix(health-check): judge the final status, refuse status fields in TCP mode, document upgrades HTTP probes skip interim 1xx responses (except 101) on HTTP/1.1 and h2c, with one HPACK decoder for the whole h2c walk, so accepting 1xx or any never lets a 100 Continue stand in for a final status. validate_health_check_config refuses a non-zero expected_status or a non-empty accepted_statuses with mode TCP, which judges no status. Document that an older worker, an older CLI patching a cluster by read-modify-write, or an older binary loading a saved state ignores fields 7 and 8 and runs GET <uri> accepting only 2xx; that --timeout on health-check set is the global command timeout; the library API break; the TOML case of mode against the UDP TCP_PROBE; and the RST a banner backend can see when the probe closes. Refs #1801 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | aefec78 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
Merge remote-tracking branch 'origin/main' into feat/health-check-tcp-mode Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | c005f03 | |
|---|---|---|
| Author: | Florentin Dubois | |
Merge remote-tracking branch 'origin/main' into fix/backend-connect-timeout-failover Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | c81ded1 | |
|---|---|---|
| Author: | Florentin DUBOIS | |
| Committer: | GitHub | |
feat(mux-h2): refuse new streams before the pre-response RST_STREAM cap (#1798) Fixes #1797 ## What The pre-response RST_STREAM cap (`h2_max_rst_stream_abusive_lifetime`) ends the connection with `GOAWAY(ENHANCE_YOUR_CALM)`, and every stream in flight on it goes with it. This adds a soft state below the cap. New client streams are refused with `RST_STREAM(REFUSED_STREAM)`, which RFC 9113 §8.7 makes safe to retry, and the streams already open keep being served. The cap itself is unchanged. ## Design `H2FloodDetector::refuses_new_streams` returns true only when all three of these hold: - the pre-response resets are past `h2_stream_refusal_percent` of the cap's floor (default 50); - pre-response resets already outnumber the backend-routed streams. This is the cap's own ratio, so a client that would never trip the cap is never refused; - the last pre-response reset is less than one flood window (1 s) old. The pre-response count never decays, and refused streams cannot raise the answered count, so this condition is what ends the refusal once the client stops cancelling. `ConnectionH2::handle_header_state` checks it after the drain check and before the MAX_CONCURRENT_STREAMS check, on server connections only. The refusal goes through `enqueue_rst` with `RstOrigin::Local`. It records no glitch and does not trigger the SETTINGS back-pressure, which would halve MAX_CONCURRENT_STREAMS for the rest of the connection over a state that decays. A refused stream is never registered, so its id never enters `rst_sent`. The refusal therefore feeds no flood counter, and in particular not the provoked-RST (MadeYouReset) cap. A refused stream never got a response. `ConnectionH2::flood_refused_streams` keeps the latest contiguous run of refused ids. When the client resets one of them, the reset counts as a pre-response reset, and `flood_refusal_ignored` ends soft refusal for that connection. A client that ignores REFUSED_STREAM then meets the cap exactly as it does today, with no extra RST_STREAM frames in front of the GOAWAY. Config: `h2_stream_refusal_percent` (`u32`; `0` disables; `100` or more never refuses before the cap). It is available in TOML, in `command.proto` (`HttpListenerConfig` 37, `HttpsListenerConfig` 50, `UpdateHttpListenerConfig` 43, `UpdateHttpsListenerConfig` 44) and as `--h2-stream-refusal-percent` on `sozu listener http|https update`. Every value has a meaning, so it has no floor and no clamp. New counter: `h2.flood.stream_refused`. `H2FloodConfig::new` / `from_optional` take a fourteenth argument (library API). ## Trade-offs - **The per-window RST_STREAM rate has no soft state.** I implemented it first and dropped it. That counter counts every reset, answered streams included, and it half-decays every window, so a client sustaining half the rate never reaches the limit. In practice it refused streams from `test_h2_cancels_during_a_backend_outage_keep_the_connection`, a client the limit never cuts off. - **PING, SETTINGS, empty DATA, CONTINUATION and stream-0 WINDOW_UPDATE floods are unchanged.** They open no streams. - **The provoked-RST cap has no soft threshold of its own.** Provoked resets still count toward the shared pre-response ratio. - **A cancel that races its own refusal ends soft refusal for that connection.** This happens when the client resets a stream before the REFUSED_STREAM reaches it. The connection then falls back to today's behaviour. - **While refusing, Sōzu still decodes each refused HEADERS block (HPACK) and queues one RST_STREAM.** This is bounded by the pending-RST queue bound per pass and by the one-window lifetime of the soft state. ## Tests Red runs were made against a `git archive origin/main` copy with only the test file copied in, or against a scratch copy carrying the named mutation. Green runs were made in the branch. - `test_h2_cancels_past_the_soft_threshold_refuse_new_streams` (e2e, new). The client keeps an upload open, cancels one past the soft threshold and under the floor, then opens a probe stream. Red on origin/main: `probe RST_STREAM code=None`. Green: `code=Some(7)`, the upload completes with 200, there is no GOAWAY, and a stream opened after 1.5 s is served. - `resets_of_streams_refused_under_flood_count_as_pre_response` (h2.rs unit). Red when refused-stream resets are not counted as pre-response. With the same mutation, `test_h2_rapid_reset_triggers_goaway` and `e2e_h2_flood_rapid_reset_abusive_lifetime` both go red (no GOAWAY). - Detector unit tests: entering and leaving the soft state, the ratio condition, the hard limits unchanged, `0` and `100` inert, and the RST rate ignored. - Rapid-reset attacker tests still end in `GOAWAY(ENHANCE_YOUR_CALM)`. ## Validation - `cargo +nightly fmt --all -- --check`: ok - `cargo +1.93.1 clippy --all-features --all-targets --locked -- -D warnings`: ok - `RUSTDOCFLAGS="-D warnings" cargo doc --all-features --no-deps`: ok - `cargo test -p sozu-lib --all-features`: 1536 passed - `cargo test -p sozu-command-lib -p sozu --all-features --locked`: 527 passed - `cargo test -p sozu-e2e --locked --features opentelemetry,splice,simd -- --test-threads=4 flood rapid_reset goaway h2spec`: 43 passed - `cargo +nightly fuzz build`: ok - `check_doc_citations.py` `--self-test`, `--root .`, `--root . --base origin/main`, `--audit --root .`: all ok Docs: `doc/configure.md` (new section, reference row, metric), `doc/h2_mux_internals.md`, `lib/src/protocol/mux/LIFECYCLE.md`, `bin/config.toml`, `CHANGELOG.md`. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 9f7d62a | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(health-check): add a TCP connect probe mode and configurable accepted HTTP statuses `HealthCheckConfig` gains a `mode` (proto field 7): `HTTP`, the default and the previous behaviour, or `TCP`, where a backend passes as soon as the TCP connection is established within `timeout`. Nothing is sent, the connection is closed once established, and `uri` and the status fields are ignored. A platform can then enable health checks on every cluster to take blackholed backends out of rotation without marking down applications that answer 3xx, 401, 404 or 500 on a fixed path. HTTP mode gains `accepted_statuses` (proto field 8, inclusive `HttpStatusRange` entries; TOML/CLI entries "404", "200-399", "2xx" or "any", bounds within 100-599). It replaces `expected_status` when set; setting both a non-zero `expected_status` and a list is refused by `validate_health_check_config` at every boundary. `uri` stays a required proto field (may be empty in TCP mode); the TOML key and `--uri` default to "/". Messages and saved states without the new fields decode as HTTP mode with no list, so existing configurations keep their meaning. `sozu cluster health-check set` gains `--mode` and `--accepted-statuses`; `list` shows the mode and accepted statuses. The probe timeout flag of `health-check set` becomes `--probe-timeout`: its `--timeout` shared the clap id of the global `-t/--timeout` command timeout (u64 against u32), so every `health-check set` invocation panicked on the downcast before sending anything. Fixes #1801 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | b63963d | |
|---|---|---|
| Author: | Florentin Dubois | |
fix(mux): fail over to another backend on connect timeout A backend that accepted no connection hit `connect_timeout` and the request was answered 504. The timeout path marked the connection `Disconnecting`, so the dead-backend branch of `Mux::ready_inner`, the only place a connect failure is counted, never saw it: the backend's retry policy was never armed and the load balancer kept choosing it. A connect timeout is now a backend connect failure like a refused dial. `Mux::timeout_inner` flags the connecting connection's readiness ERROR and `Server::timeout` runs the `ready` pass only it can trigger (`ProxySession::needs_ready_pass`); the dead-backend branch counts the failure, closes the connection and re-links the request, which sent no byte, through `EndStreamAction::Reconnect`. 504 stays the answer to a response timeout on an established connection. On the same path: - a WRITABLE event on a connecting socket reads `SO_ERROR` and `getpeername(2)` before the connection is marked established: an error is a connect failure, a dial still in flight (a WRITABLE sozu raised itself, as for an HTTP/2 backend's preface) waits for the kernel edge; - a synchronous `connect(2)` failure is retried instead of answered 503; - each retry skips the backends that already failed for the request, under every load-balancing policy, falling back to the plain selection when every selectable backend was tried; - the retry budget replaces the hard-coded `CONN_RETRIES = 3` with `max_connection_attempts`: global (default 5) and per cluster, in TOML, the protocol and the CLI, changeable at runtime with `sozu connection-attempts set` and `sozu cluster connection-attempts`. Fixes #1800 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 4be6a69 | |
|---|---|---|
| Author: | Florentin Dubois | |
feat(mux-h2): refuse new streams before the pre-response RST_STREAM cap The pre-response RST_STREAM cap (h2_max_rst_stream_abusive_lifetime) ends the connection with GOAWAY(ENHANCE_YOUR_CALM), and every stream in flight on it. Add a soft state below it: once a connection's pre-response resets are past h2_stream_refusal_percent (default 50, 0 disables) of the cap's floor, already outnumber its backend-routed streams, and the last one is less than a flood window old, each new client stream is refused with RST_STREAM(REFUSED_STREAM), retryable per RFC 9113 §8.7, while the open streams keep being served. The refusal feeds no flood counter: it is a local reset, records no glitch, and skips the SETTINGS back-pressure. A reset of a refused stream counts as a pre-response reset and ends the refusals on that connection, so a client that ignores them meets the cap exactly as before. The per-window RST_STREAM rate has no soft state: it counts every reset, and with its half-decay a client at half the rate never reaches the limit. New listener key h2_stream_refusal_percent (TOML, protobuf, sozu ctl) and counter h2.flood.stream_refused. Fixes #1797 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 933eba0 | |
|---|---|---|
| Author: | Florentin DUBOIS | |
| Committer: | GitHub | |
fix(mux-h2): raise H2 flood-protection defaults and scale RST caps with answered streams (#1756) Raise H2 flood-protection defaults that tripped on legitimate traffic; scale RST caps with the streams a backend answered; stop counting Sōzu-initiated resets. Fixes #1749 ## Changes - **Stream-0 WINDOW_UPDATE**: an update answering DATA Sōzu sent is credited (two per DATA frame, after Envoy's `max_inbound_window_update_frames_per_data_frame_sent`); stored credit is capped by the DATA sent recently and decays once DATA stops. `h2_max_window_update_stream0_per_window` bounds only unsolicited updates (RFC 9113 §6.9 leaves the cadence to the peer). - **RST caps relative to backend-routed streams**: the three `h2_max_rst_stream_*_lifetime` knobs become floors. Backend-routed streams are those a backend answered, those Sōzu answered 502/503/504 after selecting a backend, and those it answered 503 because every backend of the cluster was failing; refused streams and streams answered without a backend for another reason (no route, a cluster with no backend, a session or buffer limit) do not count. Pre-response resets received and peer-provoked resets emitted share one count and trip once they exceed the streams a backend answered, i.e. more than half of the backend-routed streams (Envoy's 50 % line). `h2_max_rst_stream_lifetime` compares only the resets received after a response started or on a closed stream with the answered streams; a stream reset after its response counts as answered before the check, so cancelling answered streams never trips it. No absolute per-connection reset ceiling is added. - **Sōzu-initiated resets** (idle-reaper `CANCEL`, `REFUSED_STREAM`, `STREAM_CLOSED`, converter errors on a backend failure) no longer feed the CVE-2025-8671 cap. Each peer-provoked reset also counts as a glitch. - **Refused streams**: counted as glitches once the client has acknowledged our SETTINGS (buffer-pool refusals excepted); their ids are no longer kept in the per-connection reset set. - **Backend connections**: a `RST_STREAM` a backend sends before its response is not counted as Rapid Reset. - **Pending RST queue**: bounded on what is pending (at least 4000, or four per advertised `h2_max_concurrent_streams`); GOAWAY only when an RST could not be queued. - **Defaults ×20** for every count threshold: per-window RST/PING/empty-DATA/stream-0 WINDOW_UPDATE 2000, SETTINGS 1000, glitch 2000, RST floors 200 000 / 1000 / 10 000, PING/SETTINGS lifetime 200 000, pending RST floor 4000, back-pressure 1000 refusals per 60 s. - **Memory bounds unchanged**: `h2_max_header_list_size`, `h2_max_header_fields` (128), `h2_max_continuation_frames`, `h2_max_header_table_size`, the PRIORITY map size and buffer sizes. - `doc/configure.md`, `doc/h2_mux_internals.md`, `LIFECYCLE.md`, `bin/config.toml`, proto/config/CLI help and `CHANGELOG.md` updated. Reference defaults compared: nghttp2 (reset burst 1000 / 33 s⁻¹, glitch 10 000 / 330 s⁻¹, outbound queue 1000), Go x/net/http2 (queued control frames 10 000), hyper h2 (local error resets 1024), Envoy (premature reset ≥ 50 % of ≥ 500 streams, WINDOW_UPDATE per DATA frame sent, outbound control frames 1000). ## Tests New, each seen failing without its fix and passing with it: - e2e `test_h2_pre_response_cancels_keep_the_connection`, `test_h2_per_frame_connection_window_updates_do_not_trip` - `local_resets_are_not_charged_to_the_peer`, `backend_resets_before_the_response_are_not_rapid_reset` - e2e `test_h2_cancels_during_a_backend_outage_keep_the_connection`; `test_flood_detector_pre_response_cap_trips_only_above_half`, `test_flood_detector_window_update_credit_covers_data_in_flight` - `refused_streams_keep_rst_sent_bounded_and_trip_the_glitch_budget`, `peer_provoked_resets_count_as_glitches`, `unanswered_streams_do_not_dilute_the_reset_caps`, `test_flood_detector_pre_response_caps_share_one_count` - `resets_after_the_response_never_trip_the_lifetime_cap`, `pre_response_resets_do_not_count_toward_the_lifetime_cap`, `a_5xx_answered_without_a_backend_is_not_routed`, `a_5xx_answered_during_a_backend_outage_is_routed`; guard `resets_of_streams_already_gone_trip_the_lifetime_cap` - detector and control-queue unit tests (ratio caps, WINDOW_UPDATE credit and its cap, pending bound) Existing flood tests keep their assertions; only their frame counts rise above the new thresholds. ## Validation - `cargo +nightly fmt --all -- --check` - `cargo +1.93.1 clippy --all-features --all-targets --locked -- -D warnings` - `RUSTDOCFLAGS="-D warnings" cargo doc --all-features --no-deps` - `cargo test --all-features --locked -p sozu-command-lib -p sozu-lib -p sozu` - targeted `cargo test -p sozu-e2e --locked --features opentelemetry,splice,simd` (flood, rapid reset, GOAWAY, h2spec) - `cargo +nightly fuzz build` - `check_doc_citations.py` in its four modes Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 931d7bf | |
|---|---|---|
| Author: | Florentin Dubois | |
Merge branch 'main' into fix/h2-flood-thresholds Resolve the overlap with #1753 in the closed-stream arms of ConnectionH2::handle_frame. Main's HEADERS arm for a stream this endpoint reset (discard the payload, continue) is kept as is. On the DATA arm, the RST_STREAM(STREAM_CLOSED) sent for a stream not reset locally takes this branch's third enqueue_rst argument, RstOrigin::Local: that DATA is already counted as a glitch just above, which bounds it, and charging it to the provoked-RST cap too would count it twice. The old comment's other reason, that such DATA is usually the tail of a stream sozu reset, no longer applies: main now ignores that case before this call. doc/configure.md keeps this branch's flood-threshold rows and main's h2_initial_connection_window default; doc/h2_mux_internals.md citations are re-anchored to the merged code. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 99e6aee | |
|---|---|---|
| Author: | Florentin DUBOIS | |
| Committer: | GitHub | |
feat(udp): cap concurrent flows per source address and subnet (#1770) A UDP cluster can now cap concurrent flows per source address and per masked subnet with its own `max_connections_per_ip` and `max_connections_per_subnet` (with the worker's `subnet_ipv4_prefix` / `subnet_ipv6_prefix`). The limit is opt-in: the global defaults do not apply to UDP, so existing configurations admit the same UDP flows as before. - `UdpManager` counts live flows per `(cluster, source IP)` and `(cluster, masked subnet)`: incremented on admission, decremented in `close_flow`, checked before the `max_flows` cap. A datagram that would open a flow over either limit is dropped (`DropReason::Shed`), counted in `udp.flows.shed` and the new `udp.flows.shed.source_limit`. Datagrams of existing flows are not affected. - Counting is unconditional, so a cluster limit set at runtime (an `AddCluster` upsert) applies to the flows already open. - Counters are per UDP listener and separate from the TCP/HTTP connection counters. - Library API: `udp::ClusterConfig` gains four fields and `udp::MetricEvent` gains `FlowShedSourceLimit`. Tests: `test_udp_per_ip_flow_limit` (e2e), per-IP / per-subnet / runtime-limit / slot-release unit tests in `manager.rs`, `udp_flow_limits_come_from_the_cluster_not_the_global_defaults` in `server.rs`, the UDP simulation and fuzz harnesses with random per-source caps, and the flow-count invariant checked after every manager input. Docs: `doc/configure.md`, `doc/rate-limit-design.md` §3.5.1, `LIFECYCLE.md`, CHANGELOG. Fixes #1769 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 9117603 | |
|---|---|---|
| Author: | Florentin DUBOIS | |
| Committer: | GitHub | |
fix(mux-h2): emit fewer connection-level WINDOW_UPDATE frames when receiving DATA (#1752) Sōzu sends one stream-0 `WINDOW_UPDATE` per half of `h2_initial_connection_window` it receives (RFC 9113 §6.9). With the 1 MiB default, a 64 MiB transfer costs 125 of these frames. That happens whether Sōzu is the H2 server receiving a request body or the H2 client receiving a backend response. At LAN speed this rate can trip a peer's stream-0 `WINDOW_UPDATE` flood detection, including Sōzu's own. This raises the default to 16 MiB. For comparison, Chromium uses 15 MB, Envoy 24 MiB and Go's HTTP/2 client 1 GiB. The same transfer now costs 8 frames. The window is advertised but not enforced: DATA is read only when its stream buffer has room, so the buffer pool still bounds memory. Listeners that set the key explicitly keep their value. User-space memory stays bounded by the buffer pool. The kernel cost does grow: a connection Sōzu stops reading may now hold up to min(window, `h2_max_concurrent_streams` × 65535) octets, about 6.25 MiB by default instead of about 1 MiB, in its socket receive buffer. This is documented in `doc/configure.md`. On a backend connection, the one-shot connection-window grant fired on every SETTINGS frame while Sōzu's own send window was at most 65535, not only on the first. A backend re-sending SETTINGS could thus be pushed past 2^31-1 (RFC 9113 §6.9.1). The grant is now tracked by a flag and sent once per connection; `repeated_backend_settings_enlarge_the_connection_window_once` covers it. Stream-level credit is still returned per DATA frame. At the default 65535-octet stream window, returning it per half window halved single-stream throughput on loopback. A larger stream window is a separate decision. Tests: `e2e/src/tests/h2_window_update_tests.rs` moves 64 MiB through Sōzu in each position, with a sender that honours the windows Sōzu grants. Each test asserts that the transfer completes and that Sōzu sends at most 1 + 64 MiB / 8 MiB stream-0 `WINDOW_UPDATE` frames. With the 1 MiB default both tests fail (125 > 9); with 16 MiB both pass (8). Commands run: `cargo +nightly fmt --all -- --check`, `cargo +1.93.1 clippy --all-features --all-targets --locked -- -D warnings`, `RUSTDOCFLAGS="-D warnings" cargo doc --all-features --no-deps`, `cargo test -p sozu-lib --all-features -- protocol::mux`, targeted `cargo test -p sozu-e2e` (window / flow-control / large-response / upload), `cargo +nightly fuzz build`, `check_doc_citations.py` (self-test, root, base origin/main, audit). Docs: `doc/configure.md`, `doc/h2_mux_internals.md`, `bin/config.toml`, proto and config comments, CHANGELOG. Fixes #1744 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 5867edb | |
|---|---|---|
| Author: | Florentin Dubois | |
Merge branch 'main' into fix/h2-window-update-coalescing Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | f891b52 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
fix(mux-h2): stop the received-RST cap tripping on answered-stream cancels A frontend stream reset after its response started now counts as answered before the RST caps are checked, and the received-RST lifetime cap compares only the resets outside the pre-response count (after a response, or on a stream already gone) with the answered streams. A client that cancels every stream after its response, or mixes pre- and post-response cancels, no longer trips it past its floor. A 502/503/504 stream counts toward the caps' denominator only when a backend was selected for it, or when its cluster has backends and every one of them is failing (the outage case, flagged on the request context). A 503 for a cluster with no backend, a session or buffer limit hit before selection, or a routing error no longer dilutes the ratio. Docs, CLI help and the configuration example now state the exact ratios. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 63b0f60 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
fix(mux-h2): trip pre-response RST caps above half of backend-routed streams - Trip the pre-response caps once pre-response resets exceed the answered streams, i.e. more than half of all backend-routed streams were reset before their response. - Count a stream answered 502/503/504 by Sōzu, or with a selected backend, as backend-routed. - Size the stream-0 WINDOW_UPDATE credit cap on the DATA frames sent recently, decaying with the flood window. Refs #1749 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 9575925 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
fix(mux-h2): bound refused-stream state and share the pre-response RST ratio - Do not record never-registered stream ids (refused streams, DATA on a closed stream) in the RST dedupe set; only stream-table eviction removes an id from it. - Count every refused stream as a glitch once the peer has acknowledged our SETTINGS (buffer-pool refusals excepted). - Compare pre-response resets received and peer-provoked resets emitted as one count against half of the answered streams. - Count each peer-provoked reset as a glitch. - Count toward the RST caps' denominator only the frontend streams a backend answered. - Cap the stored stream-0 WINDOW_UPDATE credit. Refs #1749 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 17acaf0 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
fix(mux-h2): raise H2 flood-protection defaults and make RST caps decay Raise H2 flood-protection defaults that tripped on legitimate traffic; make RST caps decay; stop counting Sōzu-initiated resets. - Stream-0 WINDOW_UPDATEs answering DATA sent are credited (two per DATA frame); h2_max_window_update_stream0_per_window bounds only unsolicited ones. - The received, pre-response and peer-provoked emitted RST caps trip only once their count also exceeds the streams opened (received) or half of them (the other two); the configured values become floors. - Resets Sōzu decides on its own (idle CANCEL, REFUSED_STREAM from its concurrency limit, back-pressure or buffer pool, STREAM_CLOSED, converter errors) no longer count toward the CVE-2025-8671 cap. - On a backend connection, backend resets before the response are not counted as Rapid Reset. - The pending RST_STREAM queue bound limits what is pending, not the connection's lifetime; GOAWAY only when an RST could not be queued. - Count defaults x20: per-window RST/PING/empty-DATA/WINDOW_UPDATE 2000, SETTINGS 1000, glitch 2000, RST lifetime floors 200000/1000/10000, PING/SETTINGS lifetime 200000, pending RST floor 4000, back-pressure 1000 refusals per 60 s. Memory bounds are unchanged. - A stream beyond the advertised max_concurrent_streams counts as a glitch (RFC 9113 §5.1.2). Fixes #1749 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 81218ca | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(udp): cap concurrent flows per source address and subnet Each client source IP and port is its own UDP flow, with its own connected upstream socket and max_flows slot, and nothing bounded the flows one source address held. UdpManager now counts live flows per (cluster, source IP) and per (cluster, masked source subnet), incremented on admission and decremented in close_flow, and sheds a datagram that would open a flow over the cluster's own max_connections_per_ip / max_connections_per_subnet (with the worker's subnet_ipv4_prefix / subnet_ipv6_prefix) through DropReason::Shed, counted in udp.flows.shed and the new udp.flows.shed.source_limit. Datagrams of existing flows are never refused. The limit is opt-in per cluster: the global max_connections_per_ip / max_connections_per_subnet defaults are not inherited by UDP, so an existing configuration admits the same UDP flows as before. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | fcdec35 | |
|---|---|---|
| Author: | Florentin DUBOIS | |
| Committer: | GitHub | |
fix(udp): key flows on the client source address, not on the affinity key (#1746) UDP flows were keyed on the affinity key: under the default `affinity_key = SOURCE_IP`, `FlowKey::from_src` zeroed the source port, so every socket of one client IP shared one flow and one connected upstream socket, and backend replies went to the port of the socket that opened the flow. The flow table is now keyed on the client source IP and port whatever the affinity key, as the #1273 design states (virtual 4-tuple flows, one connected upstream socket per flow). `affinity_key` only feeds the backend-selection hash, so the sockets of one IP still land on one backend under HRW / Maglev. - `FlowKey::from_src` loses its `with_port` argument; `UdpManager::affinity_with_port` and the shell's port normalisation are removed. - Each source port of a client IP now holds its own flow, upstream socket and `max_flows` slot. - The idle-expiry part of the report does not reproduce on `main`: it matches the lost timer wakeup fixed in 596b3ec0, which is not in a release yet. This adds the end-to-end idle-eviction test that fix lacked. - Docs (`doc/configure.md`, `LIFECYCLE.md`, proto comment) and CHANGELOG updated. Tests: `same_ip_clients_get_distinct_flows_and_their_own_replies` (unit), `test_udp_same_ip_clients_are_distinct_flows` and `test_udp_idle_flow_is_torn_down` (e2e). Fixes #1732 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 8614c7e | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
fix(udp): key flows on the client source address, not on the affinity key Under the default affinity_key = SOURCE_IP, FlowKey::from_src zeroed the source port, so every socket of one client IP shared one flow and one connected upstream socket, and every backend reply went to the port of the socket that opened the flow. The flow table is now keyed on the client source IP and port whatever the affinity key, as the #1273 design states; the affinity key only feeds the backend-selection hash, so the sockets of one IP still land on one backend under HRW / MAGLEV. FlowKey::from_src loses its with_port argument, UdpManager::affinity_with_port and the shell's client_key normalisation are removed. Also add the end-to-end idle-eviction test the lost-wakeup fix (596b3ec0) lacked: a flow silent past its idle timeout is torn down and the next datagram opens a new one. Fixes #1732 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | d244853 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
fix(mux-h2): emit fewer connection-level WINDOW_UPDATE frames when receiving DATA Sōzu returns connection credit in one stream-0 WINDOW_UPDATE per half of h2_initial_connection_window received (RFC 9113 §6.9). With the 1 MiB default that is one frame per 512 KiB: a 64 MiB transfer cost 125 of them, enough at LAN speed to trip a peer's stream-0 WINDOW_UPDATE flood detection, Sōzu's own included. Raise the default to 16 MiB (Chromium 15 MB, Envoy 24 MiB): the same transfer now costs 8, as H2 server and as H2 client. The window stays advertised, not enforced; DATA is read only when its stream buffer has room, so the buffer pool still bounds memory. Stream-level credit is still returned per DATA frame: at the 65535-octet stream window, returning it per half window halved single-stream throughput on loopback. Fixes #1744 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | b33416c | |
|---|---|---|
| Author: | Florentin Dubois | |
fix(mux-h2): trip pre-response RST caps above half of backend-routed streams - Trip the pre-response caps once pre-response resets exceed the answered streams, i.e. more than half of all backend-routed streams were reset before their response. - Count a stream answered 502/503/504 by Sōzu, or with a selected backend, as backend-routed. - Size the stream-0 WINDOW_UPDATE credit cap on the DATA frames sent recently, decaying with the flood window. Refs #1749 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | eabfa45 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
fix(udp): key flows on the client source address, not on the affinity key Under the default affinity_key = SOURCE_IP, FlowKey::from_src zeroed the source port, so every socket of one client IP shared one flow and one connected upstream socket, and every backend reply went to the port of the socket that opened the flow. The flow table is now keyed on the client source IP and port whatever the affinity key, as the #1273 design states; the affinity key only feeds the backend-selection hash, so the sockets of one IP still land on one backend under HRW / MAGLEV. FlowKey::from_src loses its with_port argument, UdpManager::affinity_with_port and the shell's client_key normalisation are removed. Also add the end-to-end idle-eviction test the lost-wakeup fix (596b3ec0) lacked: a flow silent past its idle timeout is torn down and the next datagram opens a new one. Fixes #1732 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 1c48744 | |
|---|---|---|
| Author: | Florentin Dubois | |
fix(mux-h2): bound refused-stream state and share the pre-response RST ratio - Do not record never-registered stream ids (refused streams, DATA on a closed stream) in the RST dedupe set; only stream-table eviction removes an id from it. - Count every refused stream as a glitch once the peer has acknowledged our SETTINGS (buffer-pool refusals excepted). - Compare pre-response resets received and peer-provoked resets emitted as one count against half of the answered streams. - Count each peer-provoked reset as a glitch. - Count toward the RST caps' denominator only the frontend streams a backend answered. - Cap the stored stream-0 WINDOW_UPDATE credit. Refs #1749 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | e140306 | |
|---|---|---|
| Author: | Florentin Dubois | |
fix(mux-h2): raise H2 flood-protection defaults and make RST caps decay Raise H2 flood-protection defaults that tripped on legitimate traffic; make RST caps decay; stop counting Sōzu-initiated resets. - Stream-0 WINDOW_UPDATEs answering DATA sent are credited (two per DATA frame); h2_max_window_update_stream0_per_window bounds only unsolicited ones. - The received, pre-response and peer-provoked emitted RST caps trip only once their count also exceeds the streams opened (received) or half of them (the other two); the configured values become floors. - Resets Sōzu decides on its own (idle CANCEL, REFUSED_STREAM from its concurrency limit, back-pressure or buffer pool, STREAM_CLOSED, converter errors) no longer count toward the CVE-2025-8671 cap. - On a backend connection, backend resets before the response are not counted as Rapid Reset. - The pending RST_STREAM queue bound limits what is pending, not the connection's lifetime; GOAWAY only when an RST could not be queued. - Count defaults x20: per-window RST/PING/empty-DATA/WINDOW_UPDATE 2000, SETTINGS 1000, glitch 2000, RST lifetime floors 200000/1000/10000, PING/SETTINGS lifetime 200000, pending RST floor 4000, back-pressure 1000 refusals per 60 s. Memory bounds are unchanged. - A stream beyond the advertised max_concurrent_streams counts as a glitch (RFC 9113 §5.1.2). Fixes #1749 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 99fe7cb | |
|---|---|---|
| Author: | Florentin Dubois | |
fix(mux-h2): emit fewer connection-level WINDOW_UPDATE frames when receiving DATA Sōzu returns connection credit in one stream-0 WINDOW_UPDATE per half of h2_initial_connection_window received (RFC 9113 §6.9). With the 1 MiB default that is one frame per 512 KiB: a 64 MiB transfer cost 125 of them, enough at LAN speed to trip a peer's stream-0 WINDOW_UPDATE flood detection, Sōzu's own included. Raise the default to 16 MiB (Chromium 15 MB, Envoy 24 MiB): the same transfer now costs 8, as H2 server and as H2 client. The window stays advertised, not enforced; DATA is read only when its stream buffer has room, so the buffer pool still bounds memory. Stream-level credit is still returned per DATA frame: at the 65535-octet stream window, returning it per half window halved single-stream throughput on loopback. Fixes #1744 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | c2cad33 | |
|---|---|---|
| Author: | Florentin DUBOIS | |
| Committer: | GitHub | |
feat(state): remove a cluster's frontends and backends with it (#1729) Closes #1723 ## What changes Removing a cluster (`RemoveCluster`, `sozu cluster remove`) now removes every frontend and backend that names it, for HTTP, HTTPS, TCP and UDP, in the main process state and in every worker. A frontend without a `cluster_id` (a deny or answer route) stays. - **State.** `ConfigState::remove_cluster` also removes the cluster's `backends`, `tcp_fronts` and `udp_fronts` buckets, plus the matching `http_fronts`/`https_fronts` entries. New debug assertions check two things: nothing naming the cluster remains, and no other cluster's objects are touched. - **Worker.** Before the worker applies `RemoveCluster` to its own state, it calls `ConfigState::cluster_removal_requests` to get the cluster's objects as removal orders. It then drops each one from its proxies through the same removal the object's own order takes: `HttpProxy::remove_http_frontend`, `HttpsProxy::remove_https_frontend`, `TcpProxy::remove_tcp_front`, `UdpProxy::remove_udp_front` and `Server::remove_backend`. Routes, tags, the TCP and UDP listener→cluster mappings and per-backend metrics are therefore cleaned the same way. Finally `BackendMap::remove_cluster` drops the cluster's backend list, health-check config and http2 hint. The worker still sends exactly one answer to the `RemoveCluster`. - **Live sessions.** An established session keeps its backend connection and drains on it. A new request gets the listener's 404, and a new TCP connection is closed. - **Mixed versions.** A worker of this release may receive, from an older main process, the removal of an object its `RemoveCluster` already removed. It answers ok instead of failing on a missing pre/post route. - **`diff()`.** `ConfigState::diff` now computes its frontend and backend diffs against `self` with the cascade already applied. For a removed cluster it emits `RemoveCluster` alone, with no per-object `Remove*` after it. It also adds back, after the `RemoveCluster`, any object the target state still holds for that cluster. The existing `diff` test now expects the new orders. ## Protocol and upgrade impact This is a behaviour change, recorded in CHANGELOG (Unreleased, Changed): - Upgrade the main process and every worker before relying on the cascade. An older worker removes only the cluster definition and keeps routing the cluster's frontends. - A `Remove*Frontend`/`RemoveBackend` sent after the `RemoveCluster` of the same cluster, whose object went with the cluster, is answered ok without any change by the main process and by the workers (`ConfigState::removed_with_its_cluster`), so `sozu cluster remove … && sozu frontend http remove …` keeps working. - An HTTP/HTTPS frontend removal matches only a frontend of the cluster it names. The route key has no cluster in it, so before this change a late `RemoveHttpFrontend(A, K)` could evict the route that cluster B took over under key K. A worker never applies to its proxies a frontend removal that its state refused (NotFound or NoChange). TCP and UDP proxies strip a listener's cluster by address alone, so a stale TCP/UDP removal naming a re-added cluster could otherwise unmap another cluster's listener. - The ok answer for an object gone with its cluster also covers a typo in the cluster id. The main process answer says the cluster does not exist and names the cluster, or the route with no cluster, that now holds the key. - An older main process keeps the removed cluster's objects in its own state, while workers of this release drop them. - A request to a removed cluster's hostname can still be served by another cluster's wildcard or catch-all route that covers the same hostname. - A caller that re-adds frontends or backends after removing a cluster must re-add the cluster first. ## Tests added (each one seen red before the fix) - `command/src/state.rs`: `remove_cluster_removes_its_frontends_and_backends_for_every_protocol`, `diff_removing_a_cluster_relies_on_the_cascade`, `diff_re_adds_a_frontend_the_target_keeps_for_a_removed_cluster`, `a_removed_cluster_leaves_nothing_for_load_state_to_replay`; `diff` updated. - State and worker: `a_stale_frontend_removal_never_removes_another_clusters_route`, which replays the sequence RemoveCluster(A), AddCluster(B), AddHttpFrontend(B, K), then a stale RemoveHttpFrontend(A, K). - Worker: `a_stale_tcp_or_udp_frontend_removal_never_unmaps_another_clusters_listener`, which replays RemoveCluster(A), B takes X, A is re-added on Y, then a stale Remove(A, X), for both TCP and UDP. - `lib/src/server.rs`: `remove_cluster_removes_its_routes_and_backends_from_the_worker`, covering routes, the deny route, the other cluster, the `BackendMap`, a single answer, and ok answers to follow-up old-style removals. - `e2e/src/tests/remove_cluster_tests.rs`: HTTP (404 on new and kept-alive connections) and TCP (new connection closed, established session drains). ## Commands run (all exit 0 unless noted) - `cargo +nightly fmt --all --check` - `cargo +1.93.1 clippy --all-features --all-targets --locked -- -D warnings` - `RUSTDOCFLAGS="-D warnings" cargo doc --no-deps --all-features --locked` - `cargo test -p sozu-command-lib` / `-p sozu-lib` / `-p sozu` - `cargo test -p sozu-e2e -- remove_cluster_tests metrics_lifecycle` - Full `cargo test -p sozu-e2e`: exit 101. 538 passed and 3 failed under full-suite load: `test_h1_global_limit_emits_429`, `test_h2_graceful_shutdown_completes_large_transfer`, `test_proxy_protocol_keepalive_hup`. None of the three uses `RemoveCluster`, and each passed 3 isolated runs out of 3. - `cd fuzz && cargo +nightly fuzz build` - `RUSTFLAGS='--cfg tokio_unstable' cargo test -p sozu-sim` - `check_doc_citations.py` in all four modes: `--self-test`, `--root .`, `--root . --base origin/main`, `--audit --root .` ## Docs `doc/configure.md` (Clusters), `doc/configure_cli.md` (new "Remove a cluster" section), `doc/observability.md`, the `sozu cluster remove` help, and the `remove_cluster` comment in `command/src/command.proto`. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | a810211 | |
|---|---|---|
| Author: | Florentin DUBOIS | |
| Committer: | GitHub | |
feat(lb): shuffle-shard clients over the HRW ranking of their affinity key (#1712) Refs #524 (PR 2 of 2: shuffle sharding over the HRW ranking; PR 1 was #1706). ## What A cluster can serve each client from a **shard** of its backends instead of all of them. A client that overloads or poisons its backends then takes down only the clients whose shard lies inside its own. References: - Colm MacCárthaigh, [*Workload isolation using shuffle sharding*](https://aws.amazon.com/builders-library/workload-isolation-using-shuffle-sharding/), Amazon Builders' Library, 2019. It gives `C(N, k)` shards against `N / k` partitions: 28 against 4 for eight backends in shards of two. - The shard is the top of the rendezvous ranking: Thaler & Ravishankar, *Using name-based mappings to increase hit rates*, IEEE/ACM ToN 1998, [doi:10.1109/90.663936](https://doi.org/10.1109/90.663936). ## Decisions implemented (maintainer comment on #524) - **Key:** the affinity key from #1706. That is the source IP (post PROXY protocol), or the `affinity_header` / `affinity_cookie` value, falling back to the IP when absent. - **Shard:** `k = max(2, ceil(shard_percent × N / 100))`, capped at `N`. It is the top `k` of the weighted HRW ranking (`hrw_score`, now public) over the cluster's **configured primary** backends. - Backends that are down stay members, so a shard can be exhausted. - It is deterministic across workers and restarts. - A change of `N` moves a client only at the tail of its ranking. - **Activation:** only when `N >= shard_min_backends`, which defaults to 8. Off until `shard_percent` is set. - **Exhausted shard:** - `FALLBACK` (default) spills over to the rest of the cluster, backup tier and fail-open included, and counts `backend.shard.spillover`. - `STRICT` selects nothing (HTTP 503, TCP close) and counts `backend.shard.exhausted`. - In the fail-open regime a strict shard fails open inside itself only. - **Priority:** a sticky cookie naming a live backend wins, even outside the shard. Connection reuse keeps priority. - **Upgrade note:** explicit in `CHANGELOG.md` and `doc/configure.md`. - Nothing changes until `shard_percent` is set. - Once set, the cluster derives a key under every policy. - Every client behind one NAT, or behind an L4 balancer without PROXY, shares one shard. ## Design - `BackendList::select_with_key` (`lib/src/backends.rs`) narrows the primary candidates to the shard before the cluster's `load_balancing` policy runs. Any policy therefore picks inside the shard (`LEAST_LOADED` balances within it, `HRW` pins to its top member with in-shard failover), and retries stay in it. - `next_available_backend_with_key` keeps its signature. `BackendMap::select_backend` (HTTP reserve and TCP connect) and the UDP keyed path read the `ShardOutcome` and count it under the cluster id. - **Zero allocation:** two scratch buffers (`shard_scores`, `shard`) are reserved in `add_backend` beside `candidates`. `select_nth_unstable_by` partitions the ranking in O(N), and `retain` keeps the candidates in list order for `Candidates::new`. - **Key derivation widened:** `cluster_reads_affinity_key` (`router.rs`, shared by the TCP proxy) is true for `HRW`/`MAGLEV` **or** `shard_percent`. Without this, a sharded `ROUND_ROBIN` cluster would silently not shard. - **Config:** - `Cluster` proto fields `shard_percent = 20`, `shard_min_backends = 21`, `shard_mode = 22` (new enum `ShardMode { FALLBACK, STRICT }`). - TOML keys; `--shard-percent` / `--shard-min-backends` / `--shard-strict` on `sozu cluster add`; a `shuffle_sharding` column in the cluster table; the fields in `Debug`. - `validate_shuffle_sharding` runs at config load and on every `AddCluster`: percent in `1..=100`, min ≥ 2, and the knobs are refused without `shard_percent`. - Hot-updatable through `BackendMap::set_shuffle_sharding_for_cluster` from `server.rs::add_cluster`. - **Retry budget interaction (documented):** `CONN_RETRIES = 3`. - With `k = 2`, a request whose two members refuse spills over, or is refused, on its third attempt. - With `k ≥ 3`, that request exhausts its attempts inside the shard (503), and the next one, finding the members in back-off, spills over or is refused at once. - **UDP:** sharding lives in the backend list, so a sharded UDP cluster shards its flows by their flow key. Protocol/security impact: selection narrowing only; no parser, framing or TLS change. `STRICT` can answer 503 while other backends are healthy, by design. ## Red → green Each fix was reverted, the test run, the fix restored, and the test run again. | Mutation | Fails | Evidence | |---|---|---| | `compute_shard` never shards | e2e in-shard, e2e strict, `a_sharded_selection_stays_in_the_shard`, `an_exhausted_shard_refuses_in_strict_and_spills_in_fallback` | `shard ["pong0", "pong2"], reached ["pong0", "pong1", "pong2", "pong3", …]`, `strict answered Some((200, "pong2"))` | | `STRICT` spills over | e2e strict, unit strict | `shard [2, 3] down, strict answered Some((200, "pong0"))`, `key 0: strict never leaves the shard` | | `FALLBACK` refuses | e2e fallback | `shard [0, 3] down, fallback answered Some((503, …))` | | no key for a sharded non-HRW cluster | `only_hrw_maglev_and_sharded_clusters_derive_a_key` | `a sharded cluster derives a key whatever its policy` | | shard collected into a fresh `Vec` | `a_sharded_selection_allocates_nothing` | `256 sharded selections made 256 allocations` | Green: every test above passes after restoring. ## Tests added - **Unit (`lib/src/backends.rs`):** - `a_shard_is_the_top_of_the_hrw_ranking` - `a_shard_changes_only_at_its_tail_when_backends_change`: add with `k` fixed, add with `k` growing, remove a non-member, remove a member - `a_sharded_selection_stays_in_the_shard` (ROUND_ROBIN, LEAST_LOADED, RANDOM, HRW) - `an_exhausted_shard_refuses_in_strict_and_spills_in_fallback` - `a_strict_shard_fails_open_inside_itself` - `a_sticky_backend_outside_the_shard_wins` - `a_sharded_selection_allocates_nothing` - **Unit (`load_balancing.rs`):** `a_shard_size_follows_the_decided_formula`. - **Unit (`router.rs`):** `only_hrw_maglev_and_sharded_clusters_derive_a_key`. - **Command crate:** `shuffle_sharding_reaches_the_add_cluster_order_and_is_validated`, `add_cluster_with_invalid_shuffle_sharding_rejected`. - **e2e (`e2e/src/tests/shuffle_sharding_tests.rs`):** four backends, `shard_percent = 50`, `shard_min_backends = 4`, with the shard predicted via `hrw_score`. The shard's members are left without a listener, so connections to them are refused. - in-shard: every request lands in the predicted shard - FALLBACK: 200 from a non-member - STRICT: 503 ## Validation (head 2cda4f01 on main 2e55942d) - `cargo +nightly fmt --all -- --check`: 0. - `cargo clippy --all-features --all-targets --locked -- -D warnings`: 0. - `cargo +1.93.1 clippy --all-features --all-targets --locked -- -D warnings`: 0. - `RUSTDOCFLAGS="-D warnings" cargo doc --no-deps --all-features --locked`: 0. - `cargo test -p sozu-lib`: 1420 passed. `-p sozu-command-lib`: 217. `-p sozu`: 123. `RUSTFLAGS="--cfg tokio_unstable" cargo test -p sozu-sim`: ok. - `cargo test -p sozu-e2e -- shuffle_sharding_tests affinity_key_tests cluster_ip_limit`: 15/15. The full e2e suite (before the rebase onto #1708): 532 passed, 0 failed. - `cd fuzz && cargo +nightly fuzz build`: 0. - `check_doc_citations.py`: `--self-test`, `--root .`, `--root . --base origin/main`, `--audit --root .` all 0. ## Docs - `doc/configure.md`: - a new "Shuffle sharding" section: literature, formula, stability, keys, TOML example, composition with policy, retries, sticky, reuse and backups, UDP, and the enable-time note - metrics rows - a "Client affinity" cross-reference - `doc/metrics.md`, `doc/configure_cli.md`, `doc/architecture.md`, `lib/src/protocol/mux/LIFECYCLE.md`, `doc/testing.md`, `bin/config.toml`, `CHANGELOG.md` (Added, with the before-enabling note). - The line citations of `cluster_h2_command` in `doc/configure_admin_ops.md` became prose, because they drift with every edit of `request_builder.rs`. ## Review fixes (FIX-FIRST on a029ac63) - **M1:** `STRICT` now decides between refusing and failing open on the **primary** tier only. The `cluster_can_open` that also counted backups is replaced by `primaries_can_open`, so a healthy backup no longer turns a strict shard's fail-open into a 503. Test: `a_healthy_backup_does_not_turn_a_strict_fail_open_into_a_refusal`. Red with backups counted: `panicked … a strict shard fails open inside itself`. `doc/configure.md` "Backups are outside every shard" now states the STRICT/backup interaction explicitly. - **M2:** the upgrade note in `CHANGELOG.md` and in the `doc/configure.md` "Before enabling it" note now says: - Upgrade the main process and every worker before setting `shard_*`: an older worker ignores fields 20–22 and serves unsharded. - An older CLI's `QueryClusterById` → `AddCluster` read-modify-write (`cluster h2`) drops the fields and silently disables sharding. - **L2:** a `FALLBACK` fail-open pick that lands back in the shard is reported `InShard` (`settle_spill_outcome`) and not counted as `backend.shard.spillover`. Test: `a_fallback_fail_open_pick_inside_the_shard_is_not_a_spill_over`. Red without the settle: `127.0.0.1:20004 is a shard member, left: SpilledOver right: InShard`. - **L3:** tests added. - `a_strict_shard_failing_health_checks_refuses_while_others_are_healthy`. Red with the `primaries_can_open` guard forced false: `key 0: strict must refuse, not fail open`. - Non-uniform weights in `a_shard_is_the_top_of_the_hrw_ranking`, which also asserts that the weights move some shard. - e2e `test_re_adding_a_cluster_without_shard_percent_turns_sharding_off`, re-added under `HRW` so the key is still derived. Red when `Server::add_cluster` skips `set_shuffle_sharding_for_cluster` without `shard_percent`: `sharded Some((503, …)), then unsharded Some((503, …))`. - The backup-tier interaction is covered by M1's test. - **L4:** `compute_shard` counts the primaries and checks `shard_size` / `min_backends` before hashing, so a cluster below the threshold pays no hash. The `has_tier_candidate` walk is gone, including for unsharded clusters. - **L5:** - The weight wording now reads "ranks higher for more keys, not proportional once `k > 1`". - The UDP `DefaultHasher` cross-toolchain caveat is on "across restarts". - The `backend.shard.*` rows are aligned between `configure.md` and `metrics.md`. - The `--affinity-header` / `--affinity-cookie` help mentions that sharding reads the key under any policy. - **L1 (done):** the worker mirrors `validate_shuffle_sharding` on `AddCluster` next to the health-check mirror in `server.rs`. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 1ca4b05 | |
|---|---|---|
| Author: | Florentin DUBOIS | |
| Committer: | GitHub | |
feat(lb): pass a client affinity key to Maglev and HRW on HTTP and TCP (#1706) Refs #524 (PR 1 of 2: the affinity key; shuffle sharding follows in PR 2). ## Problem Only the UDP datapath handed `HRW` / `MAGLEV` a key. HTTP, HTTPS and TCP selected with `None`, so a cluster configured with either policy silently fell back to round-robin: a client's requests visited every backend in turn, contrary to the documentation. ## Design - **Key derivation.** `affinity_key_from_ip` and `affinity_key_from_value` (`lib/src/load_balancing.rs`) use the `DEFAULT_HASH_SEED`-seeded FNV-1a + splitmix64 hash the policies already use, with distinct domain tags for addresses and values. - Deterministic across workers and restarts; the values are pinned by a unit test. - `::ffff:a.b.c.d` keys as `a.b.c.d`. - **HTTP/HTTPS (H1 and H2).** `Router::plan_connect` derives the key right after routing (`affinity_key` in `router.rs`), for `HRW`/`MAGLEV` clusters only; every other policy derives none and selects exactly as before. - Source: the cluster's `affinity_header` value (case-insensitive name, first occurrence) or `affinity_cookie` value (exact name), else the source IP after the PROXY protocol. - Header blocks and the cookie jar are read in place and hashed without copying. Values are hashed raw: a quoted cookie and its unquoted spelling are two keys. - An `affinity_cookie` named like the listener's sticky cookie reads the value captured before that cookie is elided. - Stored in `HttpContext::affinity_key`, so a replay reuses it like the routed cluster. - **TCP.** `TcpSession::connect_to_backend` keys on `effective_session_address`, which is the PROXY-v2 source when the cluster expects one. - **Where the key enters selection** (on top of #1709's shapes): - HTTP: `BackendSelector::select(cluster, affinity, key)` → `BackendMap::reserve_backend(cluster, key, now)` / `reserve_sticky_backend(cluster, cookie, key, now)` → `select_backend` → `next_available_backend_with_key`. - TCP: `BackendMap::backend_from_cluster_id(cluster, key, now)`. - **Precedence.** - A sticky cookie naming a live backend still wins over the key. - **Connection reuse keeps priority** (owner decision): the key applies only when a new backend connection is dialled. A request reusing a backend connection its session holds (H1 keep-alive, H2 multiplexed) follows that connection. Behind an upstream proxy or CDN multiplexing tenants over warm keep-alive/H2 connections, tenants follow the connection's backend. Documented and pinned by e2e tests. - **Config plumbing.** - Proto `Cluster.affinity_header = 18`, `affinity_cookie = 19`. - TOML keys, `--affinity-header` / `--affinity-cookie` on `sozu cluster add`, an `affinity_key` column in the cluster table, lengths-only `Debug`. - `validate_affinity_key` runs at config load and on every `AddCluster`: at most one of the two, each an RFC 9110 token. `Host` and `Cookie` are refused case-insensitively, since Sōzu never keeps them as headers. - The main state refuses a cluster that names a key and serves a TCP frontend, in both orders (`AddCluster` after `AddTcpFrontend`, and the reverse). The TOML loader refuses the keys on a `tcp` cluster. - Hot-updatable. Protocol/security impact: a header/cookie key is client-chosen, so a client can pick its backend. That is the same property as a sticky cookie, and it is documented. No parser, framing or TLS change. ## Check before upgrading HTTP, HTTPS and TCP clusters already configured with `HRW` or `MAGLEV` ran round-robin and now pin each client by its source IP. Clusters whose clients arrive from few addresses — behind a NAT, or an L4 balancer without the PROXY protocol — concentrate on few backends. Switch policy, enable the PROXY protocol, or key on a header or cookie. The note is in both `CHANGELOG.md` and `doc/configure.md`. ## Allocations - `deriving_an_affinity_key_allocates_nothing`: 256 derivations make 0 allocations. - `stamping_a_dialled_backend_allocates_nothing` is unchanged at 0. ## Red / green Each case was run with the fix reverted, then restored. - **e2e** (`e2e/src/tests/affinity_key_tests.rs`). Each test predicts every key's backend with the library's `Rendezvous`: provable outcomes, not statistical ones. - Key dropped in `RegistrySelector::select` (after the #1709 rebase): HTTP, HTTP header, HTTPS, H2 header/cookie and H2-reuse all fail. For example: `source 127.0.0.2: predicted pong0 …, reached ["pong0","pong1","pong2","pong3"]`. - Key dropped in `TcpSession::connect_to_backend`: `source 10.0.0.1: predicted backend-0 for every session, reached [backend-0 … backend-3]`. - Header/cookie keying disabled: `h2 X-Tenant: tenant-0: predicted pong0 …, reached ["pong3" ×4]`. - Reuse pinning, with reuse disabled in `Router::decide_after_gate` (`reuse_token.filter(|_| false)`): - `one keep-alive connection: predicted Some("pong1") for both … reached [Some("pong1"), Some("pong0")]` - `one H2 connection: predicted Some("pong2") for both … reached [Some("pong2"), Some("pong3")]` - Green: 7/7 (HTTP IP, HTTP header, HTTPS IP+cookie, H2 header+cookie, TCP PROXY-v2, H1 keep-alive reuse, H2 reuse), 3 consecutive runs. - **Unit.** - `key 0 must stay on one backend under HRW` → green. - `256 derivations made 128 heap allocations, expected none` → green. - Sticky-named cookie: `left: Some(1686…) right: Some(1733…)` with the capture arm removed → green. - Also added: pinned hash values, derivation rules, `Host`/`Cookie` refusal, the TCP-frontend/key refusal in both orders. ## Validation (head c70f65ab, rebased on main 4420d132 = #1709) - `cargo +nightly fmt --all -- --check`: 0. - `cargo clippy --all-features --all-targets --locked -- -D warnings`: 0. - `cargo +1.93.1 clippy --all-features --all-targets --locked -- -D warnings`: 0. - `RUSTDOCFLAGS="-D warnings" cargo doc --no-deps --all-features --locked`: 0. - `cargo test -p sozu-lib`: 1407 passed. `-p sozu-command-lib`: 212. `-p sozu`: passed. - `cargo test -p sozu-e2e -- affinity_key_tests cluster_ip_limit`: 11/11. - `cd fuzz && cargo +nightly fuzz build`: 0. - `check_doc_citations.py` `--self-test`, `--root .`, `--root . --base origin/main`, `--audit --root .`: all 0. ## Docs `doc/configure.md` ("Client affinity", reuse behaviour, upgrade check, raw values, refused names, forwarding-policy interactions), `doc/configure_cli.md`, `doc/architecture.md`, `doc/lifetime_of_a_session.md`, `lib/src/protocol/mux/LIFECYCLE.md`, `bin/config.toml`, `doc/testing.md`, `CHANGELOG.md` (Fixed with the upgrade check, Added). Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 5471bfb | |
|---|---|---|
| Author: | Florentin DUBOIS | |
| Committer: | GitHub | |
feat(listener): make forwarded headers configurable per listener (#1685) ## Summary Sōzu added both forwarding header families (`X-Forwarded-*` and RFC 7239 `Forwarded`) to every forwarded request, unconditionally. This PR adds a per-listener enum for HTTP and HTTPS listeners: ```toml forwarded_headers = "both" # both (default) | x_forwarded | rfc7239 | none ``` | Mode | Sōzu adds | Client-supplied headers | | --- | --- | --- | | `both` (default) | `X-Forwarded-For`, `-Proto`, `-Port`, `Forwarded` | same as before, byte for byte: XFF and `Forwarded` chains extended, client `-Proto`/`-Port` kept | | `x_forwarded` | `X-Forwarded-For`, `-Proto`, `-Port` | same handling as `both` for that family; `Forwarded` passes through | | `rfc7239` | `Forwarded` only | `Forwarded` extended (RFC 7239 §4); `X-Forwarded-For/-Proto/-Port/-Host` **removed** | | `none` | nothing | everything passes through untouched (documented trust caveat) | ## Design - **Wire and config**: `enum ForwardedHeaders { BOTH = 0; X_FORWARDED; RFC7239; NONE }` (`command/src/command.proto:824`). `BOTH = 0` means an absent field, or a state file written before this change, keeps today's behaviour. The field is `optional ForwardedHeaders forwarded_headers` on `HttpListenerConfig` (36), `HttpsListenerConfig` (49), `UpdateHttpListenerConfig` (42) and `UpdateHttpsListenerConfig` (43). The TOML-side `ForwardedHeadersMode` (`command/src/config.rs:1773`) is `snake_case` and uses the `MetricDetailLevel` ↔ `MetricDetail` pattern, because `build.rs` puts `SCREAMING_SNAKE_CASE` on every prost enum. - **Hot update**: it follows the `send_x_real_ip` path through `L7ListenerHandler::get_forwarded_headers` (`lib/src/lib.rs:661`, `lib/src/http.rs:975`, `lib/src/https.rs:1423`) and `update_config`, and is captured per stream in `Context::create_stream` (`lib/src/protocol/mux/mod.rs:1206`), like `strict_sni_binding`. An H2 stream picks up a patch at its next HEADERS, an H1 connection when it opens (`mux/LIFECYCLE.md` §2.5). Unlike `send_x_real_ip`/`elide_x_real_ip`, which `ConfigState::update_http(s)_listener` does not apply (see Residual risks), the main-process state records the patched mode. The state (`command/src/state.rs:874,999`) and the worker (`lib/src/http.rs:1460`, `lib/src/https.rs:1744`) both run `validate_forwarded_headers` (`command/src/state.rs:3226`) and refuse an unknown wire value. - **Editor** (`lib/src/protocol/kawa_h1/editor.rs`, shared by H1 and H2 frontends): `emits_x_forwarded` / `emits_rfc7239` / `strips_x_forwarded` (L240–L254) only gate *emission*: the chain extensions (L1324), the synthesis (L1343, L1395) and the `rfc7239` elision branch (L1147). `ForwardingHop` stays mode-agnostic and renders every value once per connection. That keeps `send_x_real_ip` working in every mode, keeps the `renders()` cache and `reset()` invariants unchanged, and keeps the allocation-count tests for the default path unchanged. `X-Real-IP` is injected at the same position as before, so `both` stays byte-exact (`header_editing_output_is_byte_exact_across_keep_alive_requests` is unchanged and green). A debug postcondition (L1478) checks that `rfc7239` forwards no managed `X-Forwarded-*` header. `xff_chain` (access log) is still recorded in every mode. The `http.trusting.x_{proto,port}` counters count only in `both` and `x_forwarded`. - **`X-Forwarded-Host` on `rewrite_host`** (`lib/src/protocol/mux/router.rs:1539`): injected, replacing a client one, only in `both` and `x_forwarded`. In `rfc7239` and `none` nothing is injected, and under `none` a client `X-Forwarded-Host` passes through as sent. Sōzu's `Forwarded` element keeps its current parameters (`proto`, `for`, `by`); adding RFC 7239 §5.3 `host=` would change the wire format for every rewrite, so it is left as a possible follow-up. As a result the pre-rewrite authority is not conveyed in `rfc7239`, which the docs state. - **`X-Real-IP` independence**: `elide_x_real_ip` / `send_x_real_ip` are untouched by the mode. `X-Real-IP` belongs to neither family, and `send_x_real_ip = true` still injects it under `none` and `rfc7239`. Unit and e2e tests assert this. - **H2 trailers**: no change. `pkawa::handle_trailer` already drops `X-Forwarded-For`, `Forwarded` and `X-Real-IP` from trailer blocks in every mode. - **CLI**: `sozu listener {http,https} {add,update} --forwarded-headers <both|x_forwarded|rfc7239|none>` (clap `ValueEnum` with the TOML spellings). `listener list` shows the mode, and the patch audit diff renders it by name. ## What "strict" means for `rfc7239` In `rfc7239` mode, client-supplied `X-Forwarded-For`, `X-Forwarded-Proto`, `X-Forwarded-Port` and `X-Forwarded-Host` (the four Sōzu itself manages) are **elided**. They are neither passed through nor converted. The reasons: 1. **RFC 7239 §8.1** (Header Validity and Integrity): the header "cannot be relied upon to be correct, as it may be modified, whether mistakenly or for malicious reasons, by every node on the way to the server, including the client making the request". Its caveat that "the chain of IP addresses listed before the request came to the proxy cannot be trusted" applies equally to `X-Forwarded-*`. 2. **Pass-through would be weaker than today.** In this mode Sōzu no longer appends its hop to `X-Forwarded-For`, so a backend still reading the rightmost `X-Forwarded-For` would trust a client-forged value with nothing Sōzu wrote after it. Elision leaves the backend one forwarding family, whose last element is Sōzu's. 3. **No §7.4 conversion.** RFC 7239 §7.4 (Transition) only *encourages* converting `X-Forwarded-*` into `Forwarded` "if it can be done in a sensible way". It warns that when several `X-Forwarded-*` fields coexist "it may not be possible to do a conversion, since it is not possible to know in which order the already existing fields were added". Converted elements would also be exactly as untrusted as the originals. 4. **RFC 7239 §4** lets a proxy remove forwarding fields ("A proxy MAY remove all "Forwarded" header fields from a request"). Removing the non-standard family in a mode that promises the standard one is within that intent. The client `Forwarded` chain itself is preserved and extended, following §4 ("append it to the last existing "Forwarded" header field after a comma separator") and §7.2 ("A directly forwarded request should preserve and possibly extend it"). 5. **RFC 7239 §7.1** (HTTP Lists): a list "may be split over multiple header fields", so several `Forwarded` lines form one list. Appending Sōzu's element to the *last* client `Forwarded` line therefore extends the whole list, which is why the editor keeps only the last `Forwarded` header it sees. The same list semantics are why `rfc7239` elides *every* client `X-Forwarded-For` line, not only the last (the editor unit test sends two). Other `X-Forwarded-*` names (for example `X-Forwarded-Prefix`) are not attestations Sōzu manages and pass through in every mode. `none` is the RFC 7239 §8.2 (Information Leak) option of "preventing the internal nodes from updating the HTTP header field", for operators who do not want Sōzu to disclose the peer or its own address. In `x_forwarded` mode a client `Forwarded` passes through untouched. The owner's decision was that this mode keeps exactly the current handling of the family it emits. The docs carry the same trust caveat for it as for `none`. ## Review follow-ups (folded into this commit) 1. **Malformed client `Forwarded`: MEDIUM, fixed.** In `both` and `rfc7239`, a client line such as `Forwarded: for="6.6.6.6` (unclosed quoted-string) was extended into `for="6.6.6.6, proto=http;for="10.0.0.1:54321";by=127.0.0.1`. That swallowed Sōzu's element, and I reproduced it in both modes, so the bug predates this PR in `both`. The client value is now checked against the RFC 7239 §4 grammar by `is_valid_forwarded` (`lib/src/protocol/kawa_h1/editor.rs`): a list of `;`-separated `token=value` pairs, where a value is a token or an RFC 9110 §5.6.4 quoted-string. Empty list elements and OWS around `,`/`;` are tolerated, as RFC 9110 §5.6.1 asks of recipients. A malformed line is elided, which RFC 7239 §4 permits. Sōzu's element then goes to an earlier well-formed line or into a synthesised `Forwarded`, so it is always parseable and last. - **Also applied in `both`:** it is the other mode that extends the chain and had the same corruption. Well-formed input is byte-identical; every existing byte-exact editor test is unchanged and green. - **Not applied in `x_forwarded` and `none`:** they do not touch `Forwarded`, so the line passes through as sent. - **Tests:** `a_malformed_client_forwarded_is_elided_where_the_chain_is_extended` (the probe input, plus a well-formed line followed by a malformed one) and `is_valid_forwarded_follows_the_rfc7239_grammar` (unit), and `test_forwarded_headers_rfc7239_malformed_client_forwarded_h1` (e2e). 2. **H2 trailers: LOW, fixed.** `pkawa::handle_trailer` now also drops `x-forwarded-proto`, `x-forwarded-port` and `x-forwarded-host`. The drop is unconditional, matching the existing semantics there: every header the initial-HEADERS pass rewrites is dropped from trailers regardless of listener flags, and in every mode each of these three is either synthesised (`both`/`x_forwarded`) or removed (`rfc7239`), so a trailer copy is a spoof vector in all of them. Test: `handle_trailer_drops_each_spoof_header_individually`, extended. 3. **Add-path validation: LOW, fixed.** `validate_forwarded_headers` now runs on `AddHttpListener`/`AddHttpsListener` in `ConfigState::add_http(s)_listener` (before anything is stored) and in the worker's `HttpListener::new` / `HttpsListener::try_new`. Tests: `add_listener_rejects_an_unknown_forwarded_headers_value` (state), and `a_listener_with_an_unknown_forwarded_headers_value_is_refused` (http.rs and https.rs). Follow-up RED, then GREEN: - Editor, against a stub validator returning `true` (the pre-fix behaviour): `a_malformed_client_forwarded_is_elided_where_the_chain_is_extended ... FAILED` with left `["for=\"6.6.6.6, proto=http;for=\"10.0.0.1:54321\";by=127.0.0.1"]`, right `["proto=http;for=\"10.0.0.1:54321\";by=127.0.0.1"]`; the grammar test also FAILED. With the fix: both `ok`. - e2e, with the validity check disabled: `test_forwarded_headers_rfc7239_malformed_client_forwarded_h1 ... FAILED`. Restored: `ok`. - Trailer, before extending the drop list: `trailer with only x-forwarded-proto should leave no surviving block ... FAILED`. After: `ok` (all 4 `handle_trailer` tests). - State add, before the fix: `called Result::unwrap_err() on an Ok value ... FAILED`. After: `ok`. - Worker add, before the fix: both `a_listener_with_an_unknown_forwarded_headers_value_is_refused` FAILED (`an unknown mode must be refused`). After: `ok`. ## Tests (seen RED, then GREEN) RED runs were produced by removing the behaviour under test (for the editor, before the gating existed; otherwise by temporarily reverting the named line), never by a compile error. Editor, before gating (`both` passes from the start because it pins today's behaviour): ``` test ...editor::tests::forwarded_headers_survives_reset ... ok test ...editor::tests::forwarded_headers_x_forwarded_adds_only_the_x_forwarded_family ... FAILED test ...editor::tests::forwarded_headers_none_adds_nothing_and_passes_client_headers_through ... FAILED test ...editor::tests::forwarded_headers_rfc7239_adds_only_forwarded_and_strips_client_x_forwarded ... FAILED test ...editor::tests::forwarded_headers_both_adds_both_families_and_extends_client_chains ... ok test result: FAILED. 2 passed; 3 failed ``` After gating: `test result: ok. 5 passed; 0 failed`. Router, with `synthesises_xfh = rewriting_host` alone: `a_host_rewrite_discloses_x_forwarded_host_only_in_x_forwarded_modes ... FAILED` (left: `[b"rewrite.example.com"]`, right: `[]`). Restored: `ok`. Config, with `to_http`/`to_tls` not setting the field: `forwarded_headers_parses_every_mode_defaults_to_both_and_rejects_others ... FAILED` (left `(None, None)`, right `(Some(0), Some(0))`). Restored: `ok`. The test also covers every spelling, the absent key (`both`) and serde's `unknown variant` error for `RFC7239`, `x-forwarded`, `forwarded` and `""`. State, without the patch arm: `update_listener_forwarded_headers_applies_and_rejects_unknown_values ... FAILED` (left `Both`, right `Rfc7239`). Restored: `ok`. CLI, without `#[value(name = "x_forwarded")]`: `listener_https_update_parses_forwarded_headers ... FAILED`. Restored: `ok`. e2e, without `http_context.forwarded_headers = listener.get_forwarded_headers()` in `Context::create_stream`: ``` test tests::tests::test_forwarded_headers_none_h1 ... FAILED test tests::tests::test_forwarded_headers_rfc7239_h1 ... FAILED test tests::listener_update_tests::test_forwarded_headers_hot_change ... FAILED test tests::protocol_pair_matrix::forwarded_headers_rfc7239::test_h1_h1 ... ok test tests::protocol_pair_matrix::forwarded_headers_rfc7239::test_h2_h2 ... FAILED test tests::protocol_pair_matrix::forwarded_headers_rfc7239::test_h1_h2 ... FAILED test tests::protocol_pair_matrix::forwarded_headers_rfc7239::test_h2_h1 ... ok test result: FAILED. 2 passed; 5 failed ``` The `*_h1` backend cells pass here because, as for the existing X-Real-IP matrix feature, `AsyncBackend` exposes no request bytes; those cells only check forwarding. The H2 frontend is proved by `h2_h2`. Without the worker's `update_config` arm, `test_forwarded_headers_hot_change` fails with `patch_ok=true before_ok=true after_ok=false`. Restored: `test result: ok. 7 passed; 0 failed`. ## Validation Run on head `1076a246`, rebased onto `origin/main` `d9ce2a02` (after #1686): - `cargo build --all-features`: exit 0 - `cargo clippy --all-features --all-targets -- -D warnings`: exit 0 - `cargo +nightly fmt --all -- --check`: exit 0 - `RUSTDOCFLAGS="-D warnings" cargo doc --no-deps --all-features --locked`: exit 0 - `cargo test -p sozu-lib`: exit 0 (`1368 passed; 0 failed`) - `cargo test -p sozu`: exit 0 (`122 passed`) - `cargo test -p sozu-command-lib`: exit 0 (`202 passed`) - e2e groups, each exit 0: `forward` (23), `x_real_ip` (9), `listener_update` (15), `listener` (28), `protocol_pair_matrix` (28), `trailer` (5) - `python3 .github/scripts/check_doc_citations.py` `--self-test`, `--root .`, `--base origin/main` and `--audit --root .`: all exit 0 - the full e2e suite (`508 passed; 0 failed`) was run on the previous head `a339a456`, before #1686 Rebase onto #1686 (listener `interface`): - `command.proto`: #1686 took field numbers 35 (`HttpListenerConfig`) and 48 (`HttpsListenerConfig`) for `interface`, the numbers this PR had used. `interface` keeps them, and `forwarded_headers` moves to 36 and 49. The Update messages keep 42 and 43. No release has shipped the old numbers. - `state.rs`: `add_http(s)_listener` keeps #1686's `validated_listener_key` and runs `validate_forwarded_headers` before it. - `cli.rs` / `request_builder.rs`: both new CLI tests are kept, and the `add` builders chain `.with_interface(interface)` after the optional `with_forwarded_headers`. - `CHANGELOG.md`: both new entries are kept. - `doc/configure_admin_ops.md`: kept this PR's symbol citation in place of the renumbered `https.rs:515`. - e2e: the new H1 helper gains `interface: None` in its `ActivateListener`, as #1686 did for the sibling helpers. ## Docs `doc/configure.md` (new "Forwarding headers" section with the security note and trust caveat, an updatable-fields row, and the `send_proxy` paragraph), `doc/configure_admin_ops.md` (patch boundary list), `doc/lifetime_of_a_session.md`, `lib/src/protocol/kawa_h1/LIFECYCLE.md` §2 and §4, `lib/src/protocol/mux/LIFECYCLE.md` §2.5 capture table, `bin/config.toml` (both listener examples), `e2e/COVERAGE.md`, and `CHANGELOG.md` (Unreleased). ## Residual risks - In `rfc7239` and `none`, a `rewrite_host` conveys no pre-rewrite authority. The RFC 7239 §5.3 `host=` parameter is the natural follow-up. - The trust caveat for `none` (and for `Forwarded` under `x_forwarded`) is documented, not enforced. - Existing gap, not changed here: H1 chunked trailers have no spoof-header filter at all (`X-Real-IP`, `X-Forwarded-For` and `Forwarded` included); only H2 trailers are filtered. - Existing gap, not changed here: `ConfigState::update_http(s)_listener` does not apply `elide_x_real_ip` / `send_x_real_ip` patches, so after a hot update the main-process state (`listener list`, `SaveState`, upgrade replay) still shows the old values for those two knobs. `forwarded_headers` is applied there. Closes #322 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | d9ce2a0 | |
|---|---|---|
| Author: | Florentin DUBOIS | |
| Committer: | GitHub | |
feat(listener): bind listeners to a network interface (#1686) ## What A listener can be bound to a network interface by name, for an interface whose address is not known ahead or not stable (a WireGuard `wg0`, the case in #719): ```toml [[listeners]] protocol = "https" address = "0.0.0.0:443" interface = "wg0" ``` - New optional `interface` on `HttpListenerConfig`, `HttpsListenerConfig`, `TcpListenerConfig`, `UdpListenerConfig` (`command/src/command.proto:530`, `:681`, `:715`, `:739`), on `ActivateListener` / `DeactivateListener` / `RemoveListener` (`:829`, `:839`, `:848`), on the TOML `ListenerBuilder`, and `--interface` on `sozu listener <protocol> add|remove|activate|deactivate`. - Applied with socket2 `Socket::bind_device(Some(name))` **before** `bind(2)` by `bind_to_device` (`lib/src/socket.rs:1634`), called from `server_bind` (`:1667`, HTTP/HTTPS/TCP) and `udp_bind` (`:1749`). Setting it before `bind` is what lets two listeners share an address under `SO_REUSEADDR`/`SO_REUSEPORT`. - **UDP included**: `SO_BINDTODEVICE` is a `SOL_SOCKET` option, so `udp_bind` takes it exactly like `server_bind`. Replies to clients leave through that same listener socket. The per-flow upstream sockets (`udp_connect`) talk to backends, not to clients, so they are left alone: binding them would tie backend traffic to the frontend interface. - **Not in `Update*ListenerConfig`**: changing the interface needs a rebind (remove + add). ## Design: the interface is part of the listener identity `ListenerKey { address, interface: Option<String> }` (`command/src/listener_key.rs:120`) replaces the bare `SocketAddr` wherever a listener is stored or looked up: | Surface | Change | |---|---| | `ConfigState::{http,https,tcp,udp}_listeners` | `BTreeMap<ListenerKey, _>` (`command/src/state.rs:206`); add, remove, activate, deactivate, diff and `generate_*_requests` carry the interface | | `ActivateListener` / `DeactivateListener` / `RemoveListener` | optional `interface` selector; `listener_key()` on each message and config (`command/src/listener_key.rs`) | | SCM hand-off (`Listeners`, `ListenersCount` manifest) | `Vec<(ListenerKey, RawFd)>` (`command/src/scm_socket.rs:501`), manifest entries parsed as keys (`:622`) | | Worker proxies | `inherited_socket_fate`, `activate_listener`, `listener_token`, `give_back_listener(s)`, `remove_listener` match on the key (`HttpListener::is` and siblings in `lib/src/{http,https,tcp,udp}.rs`); `Server` builds the key from the request (`lib/src/server.rs:2706`, `:3142`, `:3398`) | | Main rollback of a failed listener add | `compute_rollback` carries the interface (`bin/src/command/requests.rs:2445`) | | `ListenersList` / `sozu listener list` / `sozu top` / audit targets | keyed and shown as the key text; `interface` column/row in the tables | **Backward compatibility.** The key's text form is the bare `SocketAddr` text for a listener without interface and `<address>%<interface>` otherwise (`FromStr`/`Serialize` at `command/src/listener_key.rs:163`, `:184`). `%` is refused in interface names (`validate_interface_name`, `:63`) and a `SocketAddr` only prints `%` inside IPv6 brackets, so parsing is unambiguous. Consequences: - `ConfigState` JSON crossing a **main-process upgrade** (`UpgradeData`) keeps identical **map keys** for every existing listener (the listener values now carry `"interface": null`, which an older Sōzu ignores as an unknown field); an older state without `interface` fields deserializes to the same state. - **Saved state files** (`WorkerRequest` JSON) without `interface` load unchanged (`Option` defaults to `None`). - The **SCM manifest** from an older Sōzu carries bare addresses, which parse as keys without interface. - A listener without `interface` behaves exactly as before. **Frontends and certificates stay keyed by address** and apply to **every** listener of their protocol on that address (`listeners_at` in `lib/src/http.rs:1134`, `lib/src/https.rs:2338`, `lib/src/tcp.rs:2905`; `listener_tokens_at` in `lib/src/udp.rs:763`). This keeps `RequestHttpFrontend`, `RequestTcpFrontend`, `RequestUdpFrontend`, `AddCertificate` & co. untouched. - **A listener added later converges with its siblings.** Each proxy's `add_listener` copies the routes of a listener already on the address when it creates the new one: `Router` + tags for HTTP, `Router` + tags + a deep copy of the `CertificateResolver` into the new listener's own resolver (the one its rustls config already points at) + an HSTS-inheritance refresh for HTTPS, cluster/SNI routes + tags for TCP, cluster + tags + flow-manager `SetCluster` for UDP (`adopt_sibling_route`). This covers `sozu listener <protocol> add --interface` and a configuration reload, which both send only the new listener's add/activate (`reload_adding_an_interface_listener_sends_only_that_listener`). I chose the worker-side copy over a main-side replay of the address's frontends/certificates: frontend and certificate requests name an address, so a replay would reach the existing siblings too, which already hold those routes (duplicate adds would fail the scatter and trigger the main's rollback). Replaying would need an interface selector on every frontend/certificate message, which is the per-interface-routing change this PR keeps out of scope. `Router`, `TrieNode` and `CertificateResolver` gain `#[derive(Clone)]` for this. - **Removals reach every listener.** Frontend removal (HTTP, HTTPS) and certificate removal/replacement now apply to every listener on the address even when one fails, and return the first failure in token order (`listeners_at` sorts by token), so no sibling keeps a route or certificate the others dropped. TCP/UDP removals were already infallible per listener. - In the configuration file, listeners sharing an address must use the same protocol (`ConfigError::ListenerProtocolMismatchOnAddress`), because a file frontend finds its listener's protocol from its address. At runtime, frontend requests are typed per protocol, so I do not enforce this in `ConfigState`: an HTTP listener on `wg0` and a TCP listener on the bare address are unambiguous and useful. **Per-interface routing (different frontends per interface on one address) is out of scope**: it would need an interface on every frontend/certificate message, state key and CLI verb. **Update patches** still name their listener by address only. When exactly one listener sits on that address (with or without interface) the patch applies to it; when several do, the main process refuses it with `StateError::AmbiguousListenerAddress` (`command/src/state.rs:139`, `listener_key_for_update` at `:154`) instead of patching an arbitrary one. If you prefer an optional key-only `interface` selector on the four `Update*ListenerConfig` messages, that is a small follow-up; I kept them untouched per the "not part of the Update messages" decision. ## Mixed versions The main process tracks no worker version (`WorkerInfo` carries none), so it cannot refuse an interface listener while an older worker runs. An older worker ignores the `interface` field and would bind on every interface, and matches Activate/Deactivate/Remove by address only. `doc/configure.md` and the CHANGELOG say: upgrade the main process and every worker before using `interface`; downgrading while a listener names an interface is unsupported (an older Sōzu cannot parse `<address>%<interface>` keys in a saved state or an upgrade hand-off). ## Errors and platform - Invalid names (empty, > 15 bytes, `.`/`..`, `/`, `:`, `%`, whitespace, NUL) are refused at config load (`ListenerBuilder::validated_interface`, `command/src/config.rs:944`) and at state dispatch (`validated_listener_key`, `command/src/state.rs:144`). - **Non-Linux refuses the key** with `InterfaceError::Unsupported` ("… SO_BINDTODEVICE, which only Linux provides") at config load and dispatch (`validate_listener_interface`, `command/src/listener_key.rs:90`); `bind_to_device` also returns `BindToDeviceUnsupported` off Linux. - A failed `SO_BINDTODEVICE` returns `ServerBindError::BindToDevice`: `could not bind the listener <addr> to the network interface '<iface>' (SO_BINDTODEVICE): <os error>. The interface must exist when the listener is activated, and on Linux kernels before 5.7 binding a socket to an interface requires the CAP_NET_RAW capability`. - **systemd**: `os-build/systemd/sozu.service` and `sozu@.service` set no `User=`, `CapabilityBoundingSet=` nor `AmbientCapabilities=`, so Sōzu runs as root with its full capability set and `SO_BINDTODEVICE` works on every kernel: no functional change. Both units gain a comment saying a drop-in that restricts capabilities on a kernel before 5.7 must keep `CAP_NET_RAW`; `doc/configure.md` says the same. - The binding survives upgrades: it belongs to the socket, and the socket crosses SCM under its key. ## Tests (each seen red, then green) | Test | RED (before the fix, or with the fix mutated out) | GREEN | |---|---|---| | `state::tests::listeners_on_one_address_with_different_interfaces_coexist` (`command/src/state.rs:6978`) | `an HTTP listener on another interface must not collide: StateError(HttpListener with id_bytes=12 already exists; …)` | ok | | `state::tests::listener_diff_distinguishes_interfaces` (`:7123`) | `left: [None] right: [Some("eth0")]` | ok | | `state::tests::update_listener_is_refused_when_the_address_is_ambiguous` (`:7162`) | `no entry found for key` | ok | | `config::tests::listener_interface_round_trips_from_the_config_file` (`command/src/config.rs:5359`), `invalid_listener_interface_rejected_at_config_load` (`:5470`) | `unknown field \`interface\`, expected one of \`address\`, …` | ok | | `socket::tests::bind_sets_so_bindtodevice_to_the_listener_interface` (`lib/src/socket.rs:2172`): binds TCP and UDP to `lo`, reads back `Socket::device()`; skips with a printed reason only on `EPERM` (it did not skip here) | `the TCP listener must be bound to lo — left: None` | ok | | `socket::tests::bind_to_a_missing_interface_names_the_interface_and_the_capability` (`:2222`) | `binding to a missing interface must fail: ()` | ok | | `listener_key::tests::listener_interface_is_refused_without_bind_to_device` (platform check mutated out) + cfg-gated `listener_interface_is_refused_off_linux` (compiles only off Linux, not run here) | `called Result::unwrap_err() on an Ok value: ()` | ok | | `state_compat_v1_1_1::legacy_config_state_json_without_interface_deserializes` (bare-address key parsing mutated out), `legacy_listener_records_without_interface_load_as_bare_address_listeners` | `a pre-#719 ConfigState must deserialize: Error("invalid listener address '127.0.0.1:8080'")` | ok | | `tcp::sni_routing_tests::listeners_sharing_an_address_on_different_interfaces_coexist` (`lib/src/tcp.rs:4417`): two listeners, `lo` + none, in one worker; key lookups, frontend fan-out, `SO_BINDTODEVICE` on the handed-back sockets (fan-out mutated to first listener only) | `the frontend must reach the listener at Token(1) — left: None right: Some("cluster-a")` | ok | | `scm_socket::tests::send_and_receive_listeners_keyed_by_interface` (`command/src/scm_socket.rs:755`) | new behaviour (keys on the wire) | ok | | e2e `test_upgrade_interface_bound_listener` (`e2e/src/tests/tests.rs:1374`): `lo`-bound HTTP listener serves, worker upgrade, 8 fresh connections served (manifest mutated to drop the interface) | `connection 3 after the upgrade was not served` | ok ×3 | | `rollback_inverse_of_a_listener_add_carries_its_interface` (`bin/src/command/requests.rs:5288`) — found in review: the rollback dropped the interface | `left: None right: Some("wg0")` | ok | | `cli::tests::listener_commands_parse_the_interface_flag` (`bin/src/cli.rs:1985`) | new flag | ok | | `http::listener_sibling_tests::a_listener_added_next_to_a_sibling_takes_over_its_frontends` (review finding 1) | `the listener added after the frontend must route it too` | ok | | `http::listener_sibling_tests::removing_a_frontend_reaches_every_listener_and_reports_a_failure` (review finding 1, removal) | `the failure on one listener must not spare its sibling's route` (HashMap order, not every run) | ok | | `https::listener_sibling_tests::a_listener_added_next_to_a_sibling_takes_over_its_certificates_and_frontends` | `the listener added after the certificate must serve it too` | ok | | `tcp::sni_routing_tests::a_listener_added_next_to_a_sibling_takes_over_its_routes` (no-SNI cluster and SNI routes) | `the listener added after the frontend must route to its cluster — left: None` | ok | | `udp::listener_sibling_tests::a_listener_added_next_to_a_sibling_takes_over_its_cluster` | `left: None right: Some("cluster-a")` | ok | | e2e `test_listener_added_on_an_interface_serves_existing_frontends`: bare `0.0.0.0:P` routes a frontend, then a `lo`-bound `0.0.0.0:P` is added + activated at runtime (what CLI add and reload send); 4 requests through `127.0.0.1` (route takeover mutated out) | `request 0 through lo was not routed: Some("HTTP/1.1 404 Not Found …")`, the reviewer's outage | ok ×3 | | `tcp::sni_routing_tests::traffic_through_an_interface_reaches_the_listener_bound_to_it` (review finding 3): bare and `lo`-bound `0.0.0.0:P`, 8 connections through `127.0.0.1`, all accepted by the `lo` listener's token, none by the bare one (`SO_BINDTODEVICE` mutated out) | `the lo listener must accept every connection, got 4 of 8: WouldBlock` (3/3 runs red) | ok | | `state::tests::reload_adding_an_interface_listener_sends_only_that_listener` | characterization of the reload diff (passes before and after; it is why the fix lives in the worker) | ok | ## Validation (head after the review fixes) - `Verified: cargo build --all-features -> exit 0` - `Verified: cargo clippy --all-features --all-targets -- -D warnings -> exit 0` - `Verified: cargo +nightly fmt --all -- --check -> exit 0` - `Verified: cargo test -p sozu-lib -> 1367 passed, 0 failed` (all test binaries) - `Verified: cargo test -p sozu -> 242 passed, 0 failed` (both test binaries) - `Verified: cargo test -p sozu-command-lib -> 210 passed, 0 failed` - `Verified: cargo test -p sozu-e2e -- listener upgrade tcp_tests udp_tests tls_tests https hsts_tests command_channel_security -> 78 passed, 0 failed` (the full e2e suite, 495 passed, ran on the previous head) - `Verified: python3 .github/scripts/check_doc_citations.py --self-test | --root . | --base origin/main | --audit --root . -> exit 0 each` File:line references above point at the initial head; the symbols named next to them are the stable reference. ## Residual risks - Per-interface routing is not supported (see design); certificate listings show one entry per listener, so an address shared by two listeners appears twice. - HTTPS HSTS: frontends copied from a sibling that inherited the sibling's listener-default HSTS are refreshed to the new listener's default; a frontend added while the sibling had no default is copied as explicit, so a new listener whose default differs from its sibling's does not apply its default to those copies. - Cluster-level answer templates (`AddCluster.answers`) are applied to the listeners that exist when the cluster is added; this PR does not change that pre-existing behaviour for later listeners. - With the fan-out, `TcpProxy::fronts` (cluster → listener token) records the last listener on the address only. It is write-only bookkeeping today (no reader in the proxy), so nothing changes; a future reader would need a set of tokens. - Downgrading while a listener has an `interface` is unsupported (documented): an older Sōzu fails to parse that listener's key in the state / SCM manifest; listeners without interface are unaffected. - Mixed versions: see above; not enforced by the main process. - The non-Linux refusal is proven through the platform-parameterized validator on Linux; the `cfg(not(target_os = "linux"))` test and the non-Linux branch of `bind_to_device` were not compiled here (no non-Linux target installed). Closes #719 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 274268a | |
|---|---|---|
| Author: | Florentin DUBOIS | |
| Committer: | GitHub | |
feat(server): add a per-(cluster, source-subnet) connection limit (#1538) ## The bypass The per-`(cluster, source-IP)` limiter has no concept of a subnet. Someone on an IPv6 `/64` — the standard residential allocation — bypasses it **entirely** by taking a fresh address for every connection: the per-address counter never reaches 2. Reported by @jpds in #1270 from reading the release notes and the limiter code. ## The shape: two independent limiters Per @FlorentinDUBOIS in the issue: > I think we need both as the current rate limit is for handling rate limit by > ip as an api gateway, and the subnet aware will be there for ddos purposes. So this adds a **second, independent counter** rather than changing the first. `cluster_ip_at_limit`, `track_cluster_ip`, `effective_max_connections_per_ip` and both per-IP maps are untouched — same fields, same semantics, same bodies. What did change around them: the two call sites now go through a combined gate, and `untrack_all_cluster_ip` also drains the subnet slots (it is the single release point all three protocols already call on close, and a second entry point would be one more thing each could forget). Both gates are consulted on every admission and **both must admit**, which is what makes "10 per IP **and** 100 per /64" expressible: ```toml max_connections_per_ip = 10 # existing, untouched max_connections_per_subnet = 100 # new subnet_ipv4_prefix = 24 subnet_ipv6_prefix = 56 ``` The conjunction lives in exactly one place, `SessionManager::cluster_connection_at_limit`, so the H1+H2 mux and raw TCP cannot drift on what "admitted" means. ## The defaults leave current behaviour intact `max_connections_per_subnet` defaults to `0` (disabled) and the prefixes default to `/32` and `/128`, which mask nothing. An upgrade is inert until an operator opts in. `/24` and `/56` are **documented recommendations**, not defaults — matching what the reporter did in the equivalent Caddy module (default left at single IPs, those prefixes highly recommended in the docs). They are written up in `doc/configure.md` and `doc/rate-limit-design.md` §7.1.1. Stronger than inert: while the cap is `0` the subnet bookkeeping is **not written at all**, so a deployment that never enables the feature allocates nothing for it. `an_unset_subnet_cap_leaves_the_admission_path_untouched` asserts both maps stay empty. This is a **deliberate divergence from the per-IP twin**, which tracks unconditionally — flagging it explicitly as the point where you may want to decide otherwise. The cost is that enabling the cap at runtime does not retroactively count sessions already in flight; the counter converges as they close. ## New configuration surface Global TOML keys (`command/src/config.rs`): | Key | Type | Default | |---|---|---| | `max_connections_per_subnet` | integer, 0 = unlimited | `0` | | `subnet_ipv4_prefix` | integer 0-32 | `32` | | `subnet_ipv6_prefix` | integer 0-128 | `128` | Plus a per-cluster `max_connections_per_subnet` override with the same `None` inherits / `Some(0)` explicit-unlimited / `Some(n)` overrides semantics as its per-IP twin. Out-of-range prefixes are **rejected** at config-load time (`ConfigError::InvalidSubnetPrefix`), never silently clamped. ### Additions to `command.proto`, named explicitly | Message / field | Tag | |---|---| | `Request.set_max_connections_per_subnet` (uint64) | 60 | | `Request.query_max_connections_per_subnet` | 61 | | `message QueryMaxConnectionsPerSubnet {}` | — | | `message MaxConnectionsPerSubnetLimit { limit, ipv4_prefix, ipv6_prefix }` | — | | `ResponseContent.max_connections_per_subnet_limit` | 18 | | `Cluster.max_connections_per_subnet` (optional uint64) | 17 | | `ServerConfig.max_connections_per_subnet` | 26 | | `ServerConfig.subnet_ipv4_prefix` | 27 | | `ServerConfig.subnet_ipv6_prefix` | 28 | `command/src/proto/command.rs` is generated and was not edited. ### CLI New `sozu subnet-connection-limit {set|remove|show}`, mirroring `sozu connection-limit`. `show` reports the two prefixes alongside the limit: they are boot-time only, so that reply is the only runtime surface that can tell an operator what the live cap is a cap *on*. Like its twin the setter is non-sticky — workers reset to the TOML value on restart. The prefixes have no runtime setter on purpose: the mask is part of the counter key, so patching it live would re-key every counter in flight. ## Masking `Ipv4Addr` / `Ipv6Addr` have no standard masking method and no dependency was added. Both shift-overflowing bounds are branched explicitly — `u128::MAX << 128` panics under `debug_assertions` and wraps silently in release, so a full-width prefix must never reach the shift; a zero prefix is branched from the other end. **IPv4-mapped IPv6 sources are canonicalised first.** A dual-stack `[::]` listener reports an IPv4 client as `::ffff:a.b.c.d`; masking that with an IPv6 prefix would collapse the entire IPv4 internet into one `::/56` bucket, so it is converted to IPv4 and charged to `subnet_ipv4_prefix`. The per-IP counter is deliberately left un-canonicalised — part of what keeps its behaviour untouched. ## To SEE THIS RED Delete the `|| self.cluster_subnet_at_limit(...)` disjunct from `SessionManager::cluster_connection_at_limit` in `lib/src/server.rs`, leaving every other piece — config keys, proto fields, masking, tracking — in place: ``` cargo test -p sozu-lib --locked --no-default-features \ --features crypto-ring,opentelemetry,splice,simd subnet ``` ``` test ...::a_subnet_cap_refuses_a_fresh_address_from_an_already_counted_subnet ... FAILED panicked at lib/src/protocol/mux/router.rs: a fresh address from an already-counted /24 must be refused — an Admitted verdict here is the #1270 bypass: a new address per connection defeats the cap ``` The test drives the real `plan_connect` → `consult_ip_gate` → `plan_connect_resume` sequence for four **distinct** addresses in one `/24`, under a per-IP cap of 10 that no single address comes near — so only the subnet counter can refuse the fourth. A fifth address in a *different* `/24` is asserted admitted, so the refusal is the subnet key and not a blanket cap on new addresses. Restoring the disjunct returns it to green; the cycle was run twice. ## Tests added - `a_subnet_cap_refuses_a_fresh_address_from_an_already_counted_subnet` — the regression above. - `an_unset_subnet_cap_leaves_the_admission_path_untouched` — the default changes nothing, and allocates nothing. - `a_cluster_override_dropping_to_zero_after_a_track_is_a_no_op` — a cluster override can drop to `Some(0)` after a session was tracked under a positive limit (`AddCluster` is an upsert and clears no counters; only the *global* `SetMaxConnectionsPerSubnet(0)` drains). An earlier revision asserted that could not happen and panicked on it. - `masking_handles_both_shift_overflowing_prefix_bounds` — prefixes 0, full width, and over-wide. - `masking_groups_addresses_at_the_recommended_prefixes` — /24 and /56, including the issue's own `/64`-holder scenario. - `an_ipv4_mapped_source_is_charged_to_the_ipv4_prefix` — including the negative case where two mapped `/24`s must not collide. - `subnet_connection_limit_keys_round_trip_from_toml` and `an_out_of_range_subnet_prefix_is_rejected_not_clamped` — the config plumbing and its validation. ## Known limitation A subnet refusal currently surfaces as `BackendConnectionError::TooManyConnectionsPerIp` and increments the shared `connections.rejected_per_cluster_ip` counter, so an operator cannot tell the two caps apart in metrics. Kept shared to hold this changeset to the feature; a distinct error variant and counter are a reasonable follow-up. ## Validation All run locally at the tip of this branch, rebased onto `main` at `42ed1111` (after #1537 and #1540 landed); CI is the verdict. The rebase over #1537 conflicted only on line citations, in `doc/testing.md` and `lib/src/protocol/mux/LIFECYCLE.md` (13 hunks). Each was resolved by locating the text the citation named at the merge base inside the merged tree, never by arithmetic; the resulting diff on those two files is citations only — every removed line has an identical added counterpart once the digits are stripped, so none of #1537's prose was altered. | Command | Result | |---|---| | `cargo build --locked --no-default-features --features crypto-ring,opentelemetry,splice,simd` | pass | | `cargo +nightly fmt --all -- --check` | pass | | `cargo clippy --all-targets --locked --no-default-features --features crypto-ring,opentelemetry,splice,simd -- -D warnings` | pass | | `cargo test -p sozu-lib …` | 1147 passed, 0 failed | | `cargo test -p sozu …` | pass | | `cargo test -p sozu-command-lib --locked --no-default-features` | 171 passed, 0 failed | | `cargo test -p sozu-e2e … -- --skip tests::fuzz_tests::` | 449 passed, 0 failed | | `RUSTFLAGS="--cfg tokio_unstable" cargo test -p sozu-sim --locked` | pass | | `RUSTFLAGS="--cfg tokio_unstable" cargo clippy -p sozu-sim --all-targets --locked -- -D warnings` | pass | | `check_doc_citations.py --self-test / --root . / --base / --audit --root .` | all pass | | `RUSTDOCFLAGS="-D warnings" cargo doc --no-deps --all-features --locked` | pass | > **Note on the two config tests — resolved since this PR was opened.** > `subnet_connection_limit_keys_round_trip_from_toml` and > `an_out_of_range_subnet_prefix_is_rejected_not_clamped` live in > `command/src/config.rs`. When this PR was first pushed, CI ran `cargo test` for > `sozu-lib`, `sozu`, `sozu-e2e` and `sozu-sim` but **never** for > `sozu-command-lib`, and there was no workspace-wide `cargo test` anywhere — so > these two would have landed without ever running. That gap was reported as > [#1539](https://github.com/sozu-proxy/sozu/issues/1539), and > [#1540](https://github.com/sozu-proxy/sozu/pull/1540) landed the step that shut it; > it is now in `main`. > This branch is rebased on top of it, so CI does exercise both tests. Docs updated in the same changeset: `doc/configure.md`, `doc/rate-limit-design.md` (§3.1, §3.2, §3.3, §3.7, §3.8, §7.1.1), `bin/config.toml`, `CHANGELOG.md`. Line citations displaced by the insertions were renumbered and each renumbered bound verified against the text it named at the base. > **On the renumbering, for `check_doc_citations.py`'s own record.** The first > automated pass produced **three wrong ranges** on lines whose text is ambiguous > (a bare `}`), two of them inverted. The checker caught two. The third — > `lib/src/server.rs:1060-1068`, which should have been `1056-1064` — **passed > silently, because a citation this changeset re-anchored is exempt from the > base comparison and is never checked at all** (the run reports this honestly: > "39 re-anchored spans were exempt from the comparison"). It was only found by > auditing all 29 renumbered bounds content-against-content, resolving each one > through the text it named at the base. Another agent hit the same mechanism on > a different changeset today, so this is an independent confirmation that the > re-anchoring exemption is a real hole, not a theoretical one: the rule cannot > guard precisely the citations a changeset touches, which are the ones most > likely to be wrong. Closes #1270 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | eac0409 | |
|---|---|---|
| Author: | Florentin DUBOIS | |
| Committer: | GitHub | |
fix(parser): match the HTTP method token case-sensitively (#1474) `Method::new` compared with `compare_no_case`, so a request line reading `get /path HTTP/1.1` produced `Method::Get`. RFC 9110 §9.1 makes the method token case-sensitive and the origin receives the original `get` bytes, so sozu and the backend could disagree about what method a request carries. Long harmless, made load-bearing by the stale-upstream replay (#1442): `is_idempotent` gates whether a request already written to a pooled upstream may be replayed on a fresh one, and `Method::Custom` is deliberately not idempotent. A lowercase `get` was therefore judged replayable while the origin was free to treat the token as an extension method with side effects of its own. The narrow scoping is not available: `context.method` is already normalised by the time the replay decision runs, and the raw token is gone. The match is now exact against the eight canonical spellings; anything else is `Method::Custom`, carrying the token verbatim. No case-insensitive fallback and no alias table: either would reintroduce the same divergence in a form that is harder to see. Metric cardinality was measured before changing, not after. No metric key or label is derived from the method: the `http.status.*` keys come only from `context.status: u16` through a closed eighteen-code list plus buckets, and the only labels the emission sites pass are `cluster_id` and `backend_id`, both operator-declared. A peer cannot mint label values by varying case, before or after this change. Routing cuts both ways, because `MethodRule::new` classifies a frontend's declared `method` through the same `Method::new`. A frontend declared without a `method` is unaffected. A frontend declared `method = "GET"` no longer matches a client sending `get`; symmetrically one declared `method = "get"` stops answering canonical `GET` requests, so a method declared in non-canonical case must be corrected to the case the client sends. Response framing changes in the direction worth having. `Method::Head` is what terminates the response parse after the headers, and `head` no longer reaches that arm. Before, an RFC-conforming origin answering `405` with a body left that body unread on a pooled keep-alive upstream, desyncing the next request on it — the common pairing. After, the cost lands only on the rare non-conforming pairing, a case-folding origin answering `head` bodiless with a `Content-Length`, where sozu serves 504 at `configured_backend_timeout`: a bounded stall, not a hang. Two smaller consequences: the access log echoes the token verbatim, as it already did for every custom method, and `Method`'s `Debug` now renders a lowercase verb as `Custom(bytes=N)` at the three sites that print `context.method` with `{method:?}`. `doc/configure.md` documented the old lenient behaviour explicitly, in both the routing-rule and replay-eligibility sections, and is updated alongside `mux/LIFECYCLE.md`, the `RequestHttpFrontend.method` proto field and the `--method` CLI help. Closes #1451. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 16e26b7 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
fix(parser): match the HTTP method token case-sensitively `Method::new` compared with `compare_no_case`, so a request line reading `get /path HTTP/1.1` produced `Method::Get`. RFC 9110 §9.1 makes the method token case-sensitive and the origin receives the original `get` bytes, so sozu and the backend could disagree about what method a request carries. Long harmless, made load-bearing by the stale-upstream replay (#1442): `is_idempotent` gates whether a request already written to a pooled upstream may be replayed on a fresh one, and `Method::Custom` is deliberately not idempotent. A lowercase `get` was therefore judged replayable while the origin was free to treat the token as an extension method with side effects of its own. The narrow scoping is not available: `context.method` is already normalised by the time the replay decision runs, and the raw token is gone. The match is now exact against the eight canonical spellings; anything else is `Method::Custom`, carrying the token verbatim. No case-insensitive fallback and no alias table: either would reintroduce the same divergence in a form that is harder to see. Metric cardinality was measured before changing, not after. No metric key or label is derived from the method: the `http.status.*` keys come only from `context.status: u16` through a closed eighteen-code list plus buckets, and the only labels the emission sites pass are `cluster_id` and `backend_id`, both operator-declared. A peer cannot mint label values by varying case, before or after this change. Routing cuts both ways, because `MethodRule::new` classifies a frontend's declared `method` through the same `Method::new`. A frontend declared without a `method` is unaffected. A frontend declared `method = "GET"` no longer matches a client sending `get`; symmetrically one declared `method = "get"` stops answering canonical `GET` requests, so a method declared in non-canonical case must be corrected to the case the client sends. Response framing changes in the direction worth having. `Method::Head` is what terminates the response parse after the headers, and `head` no longer reaches that arm. Before, an RFC-conforming origin answering `405` with a body left that body unread on a pooled keep-alive upstream, desyncing the next request on it — the common pairing. After, the cost lands only on the rare non-conforming pairing, a case-folding origin answering `head` bodiless with a `Content-Length`, where sozu serves 504 at `configured_backend_timeout`: a bounded stall, not a hang. Two smaller consequences: the access log echoes the token verbatim, as it already did for every custom method, and `Method`'s `Debug` now renders a lowercase verb as `Custom(bytes=N)` at the three sites that print `context.method` with `{method:?}`. `doc/configure.md` documented the old lenient behaviour explicitly, in both the routing-rule and replay-eligibility sections, and is updated alongside `mux/LIFECYCLE.md`, the `RequestHttpFrontend.method` proto field and the `--method` CLI help. Closes #1451. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 32b8e48 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
fix(command): resolve `certificate list --domain` through the SNI trie `sozu certificate list --domain foo.example.com` returned nothing when the certificate actually serving that host was registered as `*.example.com`, while TLS handshakes for it kept succeeding. The inventory contradicted the running proxy exactly when an operator was debugging a certificate problem and trusting the query, and the wildcard case is the common one. It fails closed rather than open, so this is a diagnostic defect, not a security one. The main process answered the filter from `ConfigState::get_certificates`, which compares the SAN list with `Vec::contains` — exact string equality — whereas the certificate that serves a handshake is chosen by `CertificateResolver::domain_lookup`, a `TrieNode` lookup resolving `*.` wildcard and, since sozu#1378, regex labels. The two disagree for every certificate not bound by an exact name. `--domain` is now resolved through that same trie, rebuilt from the control plane's own record, in `certificates_serving_domain` (`bin/src/command/requests.rs`). It is fixed at the caller because `TrieNode` lives in `sozu-lib`, which already depends on `sozu-command-lib`: the reverse edge is a cargo cycle, moving the trie into `command/` would grow a `regex` production dependency that crate does not have, and reimplementing hostname matching there would duplicate it and drift from the resolver — which is this bug. `bin/` depends on both crates and holds the only `--domain` consumer of `get_certificates`. Three properties of the rebuilt index, each pinned by a test. One trie per listener address, unioned, because `ConfigState::certificates` is keyed by listener and `TrieNode::insert` answers `InsertResult::Existing` and keeps the incumbent, so a single global trie would silently drop a second listener's certificate for the same name. Both sides are ASCII-lowercased, which is the sole derivation difference between the two name sets — `add_certificate` resolves the SAN/CN set exactly as `CertifiedKeyWrapper::try_from` does, but only the resolver folds case. And a name longer than `MAX_HOSTNAME_LENGTH`, which the resolver refuses outright, is skipped rather than indexed. This changes the shape of the answer: a trie lookup resolves to one entry, so `--domain` returns at most one certificate per HTTPS listener where the exact-equality filter could return several. The main process does not retain `AddCertificate::expired_at`, so it cannot reproduce the resolver's longest-lived tie-break and reports a reproducible choice instead; `--workers` reports what each worker's resolver actually selected. `doc/configure_cli.md` gains the section stating which question `--domain` answers — the documentation said nothing at all before — and the CLI help text and the `QueryCertificatesFilters.domain` proto comment now state it too. One divergence is deliberately left in place and commented where it lives: `lib/src/server.rs` intercepts a `QueryCertificatesFromWorkers` request whose filter carries a fingerprint and hands it to `get_certificates`, whose domain arm is tested first, so `--fingerprint <hex> --domain <host> --workers` still takes the exact-SAN arm. The two filters were never a conjunction on any path. Closes #1383 Repairs the citations this changeset moved. The added imports, `certificates_serving_domain` and the test module shift every `bin/src/command/requests.rs` line citation below them; twenty citation endpoints across twelve groups in five files drifted, of which the CI resolver reported one, because it fails only on a blank landing and every other one had drifted onto a different non-blank line. Each is re-anchored on the symbol the surrounding prose already names — `Server::handle_client_request`, `Config::command_allowed_uids`, `audit_record_to_json`, `validate_frontend_request`, `should_rollback_fanout`, `load_state`, and the `RequestType::Update*Listener` variants — rather than on a fresh line number. A symbol survives every edit above it, and these citations had already rotted once without anyone noticing: `requests.rs:1071`, cited three times in `listener_update_tests.rs` as the merge-only `LoadState` path, named a `PeerCredUnavailable` doc comment before this branch existed. No range is kept where the enclosing item was already cited in the same sentence, and the import list in `doc/configure_admin_ops.md` is replaced by the three `RequestType` variants it was standing in for. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 9c8feed | |
|---|---|---|
| Author: | kannar | |
feat(ctl): filter `frontend list` by cluster id `sozu frontend list` could only be narrowed by protocol and by domain, so answering "which frontends belong to this cluster, and what access-log tags do they carry?" meant dumping every frontend and filtering client-side. `FrontendFilters` gains an optional `cluster_id` (proto tag `5`), wired to `sozu frontend list -i/--cluster-id <id>`. `ConfigState::list_frontends` applies it to HTTP, HTTPS, TCP and UDP frontends, ANDed with the existing protocol and `--domain` filters. TCP and UDP frontends are keyed by cluster id in the state, so the filter is applied to the map key; HTTP/HTTPS frontends carry an optional cluster id, and a frontend that denies traffic (no cluster id) therefore never matches. Covered by `list_frontends_filters_by_cluster_id`, which asserts both the positive and the negative space: the wanted frontend of every protocol is listed, no other cluster's frontend leaks through, an unknown id matches nothing, a `Deny` frontend is never selected, and `cluster_id` composes with (rather than replaces) the protocol and domain filters. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: kannar <kannarfr@gmail.com>
| Commit: | 2ed35c8 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin DUBOIS | |
feat(tcp): SNI+ALPN ClientHello preread passthrough routing (#1279) Route TCP connections by the SNI (and offered ALPN) read from the TLS ClientHello without terminating TLS: buffer and parse the cleartext ClientHello on an SNI-enabled TCP listener, select a backend cluster by exact/wildcard SNI + client-preference-order ALPN, and forward the original encrypted bytes byte-for-byte, preserving end-to-end TLS and client-cert mTLS (e.g. many Kubernetes API servers multiplexed behind one :443). Architecture mirrors the UDP sans-io work: a pure ClientHello preread core (no I/O, injected time/config, invariant sweeps, no panic on adversarial input) in lib/src/protocol/tcp_preread/, driven by an impure mio shell; a deterministic simulation in sozu-sim, a cargo-fuzz target, and focused e2e. Config: TCP frontends gain optional `sni`/`alpn`; TCP listeners gain `sni_preread_timeout` (default 5s) and `sni_preread_max_bytes` (default 16 KiB, clamped to buffer_size, floored at the 5-byte TLS record header). SNI-enabled listeners reject non-TLS/no-SNI/malformed first bytes; non-SNI TCP listeners are byte-identical to before. Validation is enforced both at config load and at the worker request boundary (dynamic AddTcpFrontend / state replay). Also fixes a general Pipe close-before-flush data-truncation shared by H1/H2/WebSocket/TCP: on frontend EOF with bytes still queued for the backend, Pipe::readable returned Close unconditionally and dropped them, silently truncating any upload racing a frontend close. It now transitions to the half-closed WriteOpen state, arms the backend writable, and defers teardown to check_connections so the queue drains first. Tests: tcp_preread unit + parser fuzz (fuzz_tcp_clienthello) + deterministic simulation (tcp_preread_sim) + SNI/ALPN/mTLS/reject-matrix/PROXY/splice/ counter-conservation e2e. All current TCP e2e unchanged and green; full e2e suite (all protocols) green. CI wired for the new fuzz and simulation targets. Docs: doc/configure.md, doc/testing.md, CHANGELOG.md, and a new lib/src/protocol/tcp_preread/LIFECYCLE.md. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 9114ca2 | |
|---|---|---|
| Author: | Florentin Dubois | |
feat(tcp): SNI+ALPN ClientHello preread passthrough routing (#1279) Route TCP connections by the SNI (and offered ALPN) read from the TLS ClientHello without terminating TLS: buffer and parse the cleartext ClientHello on an SNI-enabled TCP listener, select a backend cluster by exact/wildcard SNI + client-preference-order ALPN, and forward the original encrypted bytes byte-for-byte, preserving end-to-end TLS and client-cert mTLS (e.g. many Kubernetes API servers multiplexed behind one :443). Architecture mirrors the UDP sans-io work: a pure ClientHello preread core (no I/O, injected time/config, invariant sweeps, no panic on adversarial input) in lib/src/protocol/tcp_preread/, driven by an impure mio shell; a deterministic simulation in sozu-sim, a cargo-fuzz target, and focused e2e. Config: TCP frontends gain optional `sni`/`alpn`; TCP listeners gain `sni_preread_timeout` (default 5s) and `sni_preread_max_bytes` (default 16 KiB, clamped to buffer_size, floored at the 5-byte TLS record header). SNI-enabled listeners reject non-TLS/no-SNI/malformed first bytes; non-SNI TCP listeners are byte-identical to before. Validation is enforced both at config load and at the worker request boundary (dynamic AddTcpFrontend / state replay). Also fixes a general Pipe close-before-flush data-truncation shared by H1/H2/WebSocket/TCP: on frontend EOF with bytes still queued for the backend, Pipe::readable returned Close unconditionally and dropped them, silently truncating any upload racing a frontend close. It now transitions to the half-closed WriteOpen state, arms the backend writable, and defers teardown to check_connections so the queue drains first. Tests: tcp_preread unit + parser fuzz (fuzz_tcp_clienthello) + deterministic simulation (tcp_preread_sim) + SNI/ALPN/mTLS/reject-matrix/PROXY/splice/ counter-conservation e2e. All current TCP e2e unchanged and green; full e2e suite (all protocols) green. CI wired for the new fuzz and simulation targets. Docs: doc/configure.md, doc/testing.md, CHANGELOG.md, and a new lib/src/protocol/tcp_preread/LIFECYCLE.md. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | a923c00 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin DUBOIS | |
feat(udp): add first-class UDP load balancing (#1273) Add a `udp` listener type alongside tcp/http/https so Sōzu can load-balance datagram services (DNS, syslog, NTP, generic UDP) without a second load balancer in front of it, reusing the control plane, metrics/logging and e2e harness. Architecture: a two-level sans-io core (lib/src/protocol/udp/: UdpManager admission/flow-table/generation-token timer-wheel over per-flow UdpFlow, pure and property-tested) driven by an I/O shell (lib/src/udp.rs) on the existing single-threaded mio loop. Static per-listener worker ownership; scale across cores by running multiple listeners. Flows reset on hot-upgrade (listener fd handed off via SCM, flow state not migrated). - Control plane: UdpListenerConfig + RequestUdpFrontend proto, TOML config parsing, ConfigState with diff/generate_requests/activate replay symmetry, SCM listener-fd hand-off, ctl CLI + master-process routing, and sozu top/listener-list/frontend-list/cluster display parity. - Datapath: virtual 4-tuple flows, three-knob teardown (responses / idle timeout / requests), symmetric NAT return via a per-flow connected upstream socket, bounded per-flow/per-listener write-queue egress backpressure with writable re-arm, max_flows fd-pressure shedding. No panic on network input (EMFILE/ENFILE/oversize/invalid -> drop + metric). - Load balancing: round-robin + HRW/rendezvous (default) + Maglev (opt-in), selectable per cluster via a key: Option<u64> threaded through the LB trait; the Maglev table is rebuilt only on backend-set change, never per datagram. - PROXY protocol v2 to the backend (first-datagram default, per-cluster every-datagram opt-in; DGRAM transport byte 0x12/0x22). - Active health checks bound to the endpoint: companion TCP probe (primary) + app-level UDP probe (secondary), rise/fall hysteresis, fail-open. - Metrics (udp.*) with no gauge underflow on any teardown path, and per-flow access logs under the UDP / UDP-FLOW log tag. - Tests: sans-io quickcheck/state-machine core tests, 14 e2e scenarios (RR/HRW/Maglev affinity, NAT return, teardown, truncation, PPv2, health/fail-open, hot reconfig, upgrade fd hand-off) with a mock UDP backend/client, and a fuzz_udp_flow cargo-fuzz target. - Docs: doc/configure.md UDP listener/cluster/health section + metrics, CHANGELOG, README, architecture. Plaintext only. Single-listener multi-core / eBPF SK_REUSEPORT / userland dispatcher / DSR / io_uring / DTLS / QUIC / UDP-over-HTTP / flow-state migration across upgrade are explicit non-goals. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 6d79213 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin DUBOIS | |
fix(mux): mitigate the HTTP/2 bomb (HPACK header-field cap + window-stall reaping) Closes both halves of the HTTP/2 memory-amplification DoS class (calif.io disclosure; same family as Apache CVE-2026-49975). HPACK header bomb ----------------- Thousands of 1-byte HPACK indexed references — or a single `cookie` header split into many crumbs (RFC 9113 §8.2.3) — each materialize a `Pair` of per-entry bookkeeping, amplifying wire bytes into allocation. `SETTINGS_MAX_HEADER_LIST_SIZE` was accounted as name+value bytes only, omitting the RFC 9113 §6.5.2 mandated 32-octet per-field overhead, and nothing capped the field count. - `pkawa.rs`: account `+32` octets/field per §6.5.2 and enforce a per-block materialized-field-count cap in `decode_headers_with_budget` and `handle_trailer`; cookie crumbs count individually. Over-limit blocks are rejected with `RST_STREAM` / `GOAWAY(ENHANCE_YOUR_CALM)`. - New per-listener `h2_max_header_fields` knob (default 128), plumbed through `H2FloodConfig`, `command.proto`, `config.rs`, the http/https flood-config getters, live-update patch-apply + state replay + validation, the sozuctl CLI, and listener `Display`. Window-stall ------------ A peer that holds its receive window shut while a response is buffered can pin the stream and its `MAX_CONCURRENT_STREAMS` slot. The bidirectional liveness timer (`stream_last_activity_at`) is refreshed by any inbound frame, so a 1-byte inbound DATA drip kept a stalled stream warm indefinitely. - New per-stream flow-control-stall deadline (`stream_fc_stalled_since`): armed only when a stream holds buffered response data it cannot send because its effective send window `min(stream.window, connection.window)` is exhausted, cleared on real outbound progress (`consumed > 0`), and never refreshed by inbound DATA/HEADERS — so the inbound-drip vector cannot keep it warm. Reaped by `cancel_timed_out_streams` (governed by `h2_stream_idle_timeout_seconds`) via a deduped union with the idle guard (`collect_timed_out_streams`). - The reaper runs from both `readable()` and `MuxState::timeout`, so even a fully-silent peer's stalled stream is reaped (slot freed + queued `RST_STREAM(CANCEL)` flushed). Direction-correct: a slow upload has no buffered response, a socket-bound slow link keeps a positive window, and a legitimate slow download clears the deadline on every window grant — none arm it. No new config knob (reuses `h2_stream_idle_timeout_seconds`). Tests: `pkawa` unit tests (`+32` accounting, field cap, cookie-crumb cap, legit request) and `collect_timed_out_streams` unit tests; e2e `e2e_h2_parser_reject_header_field_bomb`, `test_h2_window_stall_response_reaped`, `test_h2_window_stall_inbound_drip_reaped`, `test_h2_window_stall_silent_reaped`, with `test_h2_active_upload_survives_idle_timeout` as a false-positive guard. Validated: `cargo build --all-features`, `clippy -D warnings`, nightly fmt, the full h2 e2e suite (222), lib + command unit tests, and the cargo-fuzz hpack / frame targets. Docs: `doc/configure.md`, `CHANGELOG.md`, mux `LIFECYCLE.md`. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | f73efb7 | |
|---|---|---|
| Author: | Florentin Dubois | |
feat(udp): add first-class UDP load balancing (#1273) Add a `udp` listener type alongside tcp/http/https so Sōzu can load-balance datagram services (DNS, syslog, NTP, generic UDP) without a second load balancer in front of it, reusing the control plane, metrics/logging and e2e harness. Architecture: a two-level sans-io core (lib/src/protocol/udp/: UdpManager admission/flow-table/generation-token timer-wheel over per-flow UdpFlow, pure and property-tested) driven by an I/O shell (lib/src/udp.rs) on the existing single-threaded mio loop. Static per-listener worker ownership; scale across cores by running multiple listeners. Flows reset on hot-upgrade (listener fd handed off via SCM, flow state not migrated). - Control plane: UdpListenerConfig + RequestUdpFrontend proto, TOML config parsing, ConfigState with diff/generate_requests/activate replay symmetry, SCM listener-fd hand-off, ctl CLI + master-process routing, and sozu top/listener-list/frontend-list/cluster display parity. - Datapath: virtual 4-tuple flows, three-knob teardown (responses / idle timeout / requests), symmetric NAT return via a per-flow connected upstream socket, bounded per-flow/per-listener write-queue egress backpressure with writable re-arm, max_flows fd-pressure shedding. No panic on network input (EMFILE/ENFILE/oversize/invalid -> drop + metric). - Load balancing: round-robin + HRW/rendezvous (default) + Maglev (opt-in), selectable per cluster via a key: Option<u64> threaded through the LB trait; the Maglev table is rebuilt only on backend-set change, never per datagram. - PROXY protocol v2 to the backend (first-datagram default, per-cluster every-datagram opt-in; DGRAM transport byte 0x12/0x22). - Active health checks bound to the endpoint: companion TCP probe (primary) + app-level UDP probe (secondary), rise/fall hysteresis, fail-open. - Metrics (udp.*) with no gauge underflow on any teardown path, and per-flow access logs under the UDP / UDP-FLOW log tag. - Tests: sans-io quickcheck/state-machine core tests, 14 e2e scenarios (RR/HRW/Maglev affinity, NAT return, teardown, truncation, PPv2, health/fail-open, hot reconfig, upgrade fd hand-off) with a mock UDP backend/client, and a fuzz_udp_flow cargo-fuzz target. - Docs: doc/configure.md UDP listener/cluster/health section + metrics, CHANGELOG, README, architecture. Plaintext only. Single-listener multi-core / eBPF SK_REUSEPORT / userland dispatcher / DSR / io_uring / DTLS / QUIC / UDP-over-HTTP / flow-state migration across upgrade are explicit non-goals. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | f59de3e | |
|---|---|---|
| Author: | Florentin Dubois | |
fix(mux): mitigate the HTTP/2 bomb (HPACK header-field cap + window-stall reaping) Closes both halves of the HTTP/2 memory-amplification DoS class (calif.io disclosure; same family as Apache CVE-2026-49975). HPACK header bomb ----------------- Thousands of 1-byte HPACK indexed references — or a single `cookie` header split into many crumbs (RFC 9113 §8.2.3) — each materialize a `Pair` of per-entry bookkeeping, amplifying wire bytes into allocation. `SETTINGS_MAX_HEADER_LIST_SIZE` was accounted as name+value bytes only, omitting the RFC 9113 §6.5.2 mandated 32-octet per-field overhead, and nothing capped the field count. - `pkawa.rs`: account `+32` octets/field per §6.5.2 and enforce a per-block materialized-field-count cap in `decode_headers_with_budget` and `handle_trailer`; cookie crumbs count individually. Over-limit blocks are rejected with `RST_STREAM` / `GOAWAY(ENHANCE_YOUR_CALM)`. - New per-listener `h2_max_header_fields` knob (default 128), plumbed through `H2FloodConfig`, `command.proto`, `config.rs`, the http/https flood-config getters, live-update patch-apply + state replay + validation, the sozuctl CLI, and listener `Display`. Window-stall ------------ A peer that holds its receive window shut while a response is buffered can pin the stream and its `MAX_CONCURRENT_STREAMS` slot. The bidirectional liveness timer (`stream_last_activity_at`) is refreshed by any inbound frame, so a 1-byte inbound DATA drip kept a stalled stream warm indefinitely. - New per-stream flow-control-stall deadline (`stream_fc_stalled_since`): armed only when a stream holds buffered response data it cannot send because its effective send window `min(stream.window, connection.window)` is exhausted, cleared on real outbound progress (`consumed > 0`), and never refreshed by inbound DATA/HEADERS — so the inbound-drip vector cannot keep it warm. Reaped by `cancel_timed_out_streams` (governed by `h2_stream_idle_timeout_seconds`) via a deduped union with the idle guard (`collect_timed_out_streams`). - The reaper runs from both `readable()` and `MuxState::timeout`, so even a fully-silent peer's stalled stream is reaped (slot freed + queued `RST_STREAM(CANCEL)` flushed). Direction-correct: a slow upload has no buffered response, a socket-bound slow link keeps a positive window, and a legitimate slow download clears the deadline on every window grant — none arm it. No new config knob (reuses `h2_stream_idle_timeout_seconds`). Tests: `pkawa` unit tests (`+32` accounting, field cap, cookie-crumb cap, legit request) and `collect_timed_out_streams` unit tests; e2e `e2e_h2_parser_reject_header_field_bomb`, `test_h2_window_stall_response_reaped`, `test_h2_window_stall_inbound_drip_reaped`, `test_h2_window_stall_silent_reaped`, with `test_h2_active_upload_survives_idle_timeout` as a false-positive guard. Validated: `cargo build --all-features`, `clippy -D warnings`, nightly fmt, the full h2 e2e suite (222), lib + command unit tests, and the cargo-fuzz hpack / frame targets. Docs: `doc/configure.md`, `CHANGELOG.md`, mux `LIFECYCLE.md`. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 1bd27bc | |
|---|---|---|
| Author: | miton18 | |
| Committer: | Florentin DUBOIS | |
fix(otel): capture wall-clock start time for accurate span reconstruction Access-log consumers reconstructing OTel spans computed the request start timestamp as `precise_time - request_time`. This subtraction mixes two unsynchronised clock sources: - `precise_time` is a wall-clock reading (`CLOCK_REALTIME` via `OffsetDateTime::now_utc()`), captured at log-emission time. - `request_time` is a monotonic duration (`CLOCK_MONOTONIC` via `Instant::now() - start`), captured in `SessionMetrics` *before* the wall-clock reading. The two clocks drift independently (NTP adjustments affect CLOCK_REALTIME but not CLOCK_MONOTONIC), and they are sampled at different instants within the log-emission path. On short-lived requests the resulting error can exceed the request duration itself, producing spans whose computed start time falls after a subsequent request's start — visibly broken traces. This commit adds a `start_wall: Option<SystemTime>` field to `SessionMetrics`, captured at the same point as the existing monotonic `start: Option<Instant>`. A new `start_wall_ns()` accessor converts it to nanoseconds-since-epoch for the access log. Wire changes: - `ProtobufAccessLog.start_time` (field 30, optional `Uint128`) in `command.proto` — backwards-compatible: old consumers ignore it, new consumers with old producers see `None` and can fall back to the subtraction. - `RequestRecord.start_time_ns: Option<i128>` plumbed through all four access-log emission sites (kawa_h1, mux/stream, pipe, tcp). Consumers should now prefer `start_time` over `time - request_time` whenever the field is present. Signed-off-by: miton18 <remi@collignon-ducret.fr>
| Commit: | 3f41a83 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
refactor(command,server): remove the proto_version capability handshake Drop the proto_version field and the SetMetricDetail capability partition that gated the dispatch on worker proto version. Production deployments keep master + workers in sync via the existing UpgradeMain hot-upgrade flow, so the mixed-version-fleet state the field was designed for does not occur. The implementation was structurally broken anyway: the master stamped every WorkerInfo.proto_version from its own SOZU_PROTO_VERSION constant at fork time, so the field always read as the master's version and MIN_PROTO_VERSION_FOR_SET_METRIC_DETAIL was effectively unconditional. The proto contract is additive-only; a worker that does not recognise tag 55 (SetMetricDetail) returns WorkerResponse::error("unknown request type") which already surfaces in the standard fan-out error tally (extras.fanout.workers_err counter). MetricDetailStatus.unsupported_workers becomes redundant and is also removed. Deletions: - WorkerInfo.proto_version (tag 4) and MetricDetailStatus.unsupported_workers (tag 5); both replaced by `reserved` markers per project convention. - sozu_command_lib::SOZU_PROTO_VERSION constant. - WorkerSession.proto_version field. - MIN_PROTO_VERSION_FOR_SET_METRIC_DETAIL constant and the capability partition in set_metric_detail_request; SetMetricDetail now fans out unconditionally via the standard scatter path. - Capability-handshake paragraphs in CHANGELOG.md and doc/sozu-top.md. - print_metric_detail_status block that rendered the unsupported_workers table row. Old workers without tag 55 surface as 'succeeded with errors' (the normal fan-out failure shape) rather than as a dedicated capability-skip list. Operator visibility is preserved. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | fe3ff05 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(command,lib): per-worker WorkerMetricDetailStatus payload Final piece of the PR #1256 follow-through. Previously the capability-aware dispatcher in `bin/src/command/requests.rs` (`SetMetricDetailTask::on_finish`) synthesised `MetricDetailStatus.workers[<worker_id>]` using the master's aggregator view as a stand-in for each worker because workers replied with `WorkerResponse::ok(message.id)` carrying no payload. Each worker holds an independent `Aggregator` with its own lease table, so that stand-in obscured real per-worker drift (different configured floors, different active lease counts after a partial fan-out). Wire it properly end-to-end: - New `ResponseContent::WorkerMetricDetailStatus` oneof variant (tag 17 — proto additive). Carries the worker's own `(configured, effective, previous_effective, active_lease_count)` quartet, semantically distinct from the aggregated `MetricDetailStatus` at tag 16. - New `lib/src/server.rs::worker_metric_detail_status_content` helper that builds the response payload from a `(configured, effective, previous_effective, lease_count)` snapshot captured BEFORE the `METRICS.borrow_mut` scope ends (so the per- request snapshot is consistent with the transition that just happened). - The three ok-paths in the worker's SetMetricDetail arm (clear-Cleared, clear-NotFound, apply-Applied) now reply via `WorkerResponse:: ok_with_content` with the freshly-built payload instead of the payload-less `ok`. The `clear-NotFound` path reports `previous_effective == effective` (no transition). - Master-side `SetMetricDetailTask::on_finish` collects the per-worker payload from `response.content` and only falls back to skipping the worker entry when the response has no payload (e.g. an older worker that never went through `ok_with_content`). Removes the master-view stand-in noted as a follow-up in commit `70cd24af` (`set_metric_detail_request`). - `command/src/proto/display.rs` adds a silent OK match arm for the new variant — the per-worker payload flows master-side and is never printed directly on the operator's terminal. Build/clippy clean; 1075/1075 workspace tests pass. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | b24fba3 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(metrics,command): bind lease ownership to peer credentials Replace the unauthenticated `lease_clear(client_id)` API with a binding- aware variant so one same-UID operator cannot clear another operator's lease by guessing the `client_id` format. The binding pairs the master- side `actor_pid` (captured via `SO_PEERCRED` at command-socket accept) with the per-connection session ULID; both halves must match the apply- time binding for the worker to authorise the clear. Implementation: - New `PeerBinding { pid, session_ulid }` + `LeaseEntry` + `LeaseClearOutcome` in `lib/src/metrics/mod.rs`. `lease_apply` records the binding alongside `(level, expires_at)`; `lease_clear` returns `Cleared`/`NotFound`/ `Unauthorized` depending on the apply-time binding vs the presented one. - A "binding unknown" apply (pre-binding caller or platform without `SO_PEERCRED`) preserves backward compat: the worker accepts any clear. A fully-known apply rejects every clear whose presented binding does not match, including the default ("unknown") clear. - New proto fields `SetMetricDetail.peer_pid` (tag 6) and `SetMetricDetail.peer_session_ulid` (tag 7), additive. Clients leave them empty; the master populates them in `bin/src/command/requests.rs::worker_request` from the connecting `ClientSession` before fan-out. - Worker parses the presented ULID via `rusty_ulid::Ulid::from_str` with a `0x…` hex fallback, then routes the `LeaseClearOutcome::Unauthorized` outcome to `WorkerResponse::error` so the operator gets a loud failure rather than a silent no-op. - Six new unit tests cover authorised clear, unauthorised mismatch, unknown-apply accepts-any, known-apply rejects default clear, the existing apply/clear/tick paths through the new signature. Build/clippy clean; 26/26 metric tests pass. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 8720d28 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(command): worker proto version handshake, scaffolding for unsupported_workers Wire the proto-capability layer that the `MetricDetailStatus.unsupported_workers` field needs to be usefully populated. Previously the field was declared but always empty because the master had no way to know which workers could decode the new SetMetricDetail (tag 55) verb. This commit lays the rails; the dispatch-time gating itself remains a follow-up because it requires a dedicated `WorkerTask` impl for SetMetricDetail that synthesises a MetricDetailStatus reply rather than the current generic worker_request flow. The scaffolding shipped here: - New `sozu_command_lib::SOZU_PROTO_VERSION = 1` constant, baked into every Sōzu binary at compile time. Bumped any time a new wire- affecting `RequestType` or proto field needs capability gating. Version 1 covers `SetMetricDetail`, `METRIC_DETAIL_CHANGED`, the worker→master audit IPC, and per-lease peer-credential binding. - New `WorkerInfo.proto_version` proto field (tag 4, optional uint32) so TUI / status consumers can observe each worker's version directly. - New `WorkerSession.proto_version` master-side field, snapshotted at fork time from the binary's `SOZU_PROTO_VERSION` constant. The `to_info()` / `list_workers` paths populate the new proto field. - Proto comment on `MetricDetailStatus.unsupported_workers` updated to describe the capability-gate model AND the inherited-after-UpgradeMain caveat: a re-exec master does NOT yet re-query inherited workers' versions, so they retain the new master's compile-time value until the planned per-worker capability handshake on Status reply lands. Build/clippy clean; 53/53 TUI unit tests pass. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 65c9ed0 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(server,command): wire worker→master audit IPC for METRIC_DETAIL_CHANGED Previously, worker-local cardinality-lease transitions (TTL janitor expiry, worker-arm apply/clear) left no audit trail because the worker had no IPC path to the master's audit pipeline. Only operator-initiated transitions audited from `bin/src/command/requests.rs::worker_request` produced an audit row. A SOC analyst correlating "who elevated metrics cardinality" could see the apply but not the implicit clear. Close the gap by reusing the existing worker→master `Event` channel: - New proto `MetricDetailTransition` carrying `previous_effective`, `effective`, `transition_kind`, and an optional `client_id` for explicit apply/clear. Folded into `Event.metric_detail` (tag 5) so `EventKind::METRIC_DETAIL_CHANGED` events now carry their full payload through the existing fan-out plumbing. - Worker emits the event from three sites in `lib/src/server.rs::notify`: the polled `lease_tick` janitor (transition_kind = "lease_tick_expired" and `client_id = None` because the janitor may retire multiple leases at once), the SetMetricDetail worker arm on apply ("lease_apply"), and on clear ("lease_clear"). All three callers gate on `previous != effective` so the helper itself is a defence-in-depth no-op when nothing actually changed. - Master's `handle_worker_response` recognises METRIC_DETAIL_CHANGED events and routes them through a dedicated `audit_worker_metric_detail_transition` helper that writes to both audit sinks (text + JSON). The worker is its own actor — `worker_id` takes the `client_id` slot in the envelope; `actor_role=worker`, `actor_comm=sozu-worker`. Subscriber fan-out is unchanged. - The 11 existing `push_event(Event { … })` constructors get `metric_detail: None` to stay compatible with the new field; only the worker's lease-transition emitter populates it. Build/clippy clean; 53/53 TUI unit tests pass. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 09de4ce | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
style(top,proto): silence the last clippy and rustdoc warnings Clears the 14 warnings remaining after the simplify pass: - Five dead-code items: drop `render_placeholder` (all panes ship, no placeholder needed); drop `ClusterRow.errors_5xx_total` (the renderer reads `error_rate_pct`, never the raw count); drop `ThresholdTable.conn_warn_pct` (no consumer); drop `BackendRow.bytes_in` / `bytes_out` (the BACKENDS pane reads `back_bytes_in` / `back_bytes_out`). `Skin.categorical` gets a short doc + `#[allow(dead_code)]` because the field is read by tests and is the TOML surface for operator skins — the production consumer (cluster-row categorical tinting) lands in a follow-up. - Nine rustdoc warnings in the prost-generated `command.rs`: my earlier H2 commit's numbered-list comment in the `SetMetricDetail` preamble used 4-space indentation on the continuation lines, which prost-build expanded to overindented in the rendered docstring. Trim to 3 spaces so the rendered comment lands on the rustdoc list-item-alignment rule (3 spaces aligns with the text after `1. ` markers). - One rustdoc warning in `bin/src/ctl/top/mod.rs::run_top`: the doc comment had a `+` at line-start that markdown parsed as a list bullet, then complained the next two lines were unindented. Rephrase to join with "and" so the `+` no longer leads a line. Build is now warning-free under `cargo clippy --all-features --all-targets --locked`. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 1c26057 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
docs(proto,changelog): scope METRIC_DETAIL_CHANGED to operator-initiated transitions The proto documented "every effective-level transition emits an `EventKind::METRIC_DETAIL_CHANGED` event", but only operator-initiated transitions emit today: the master-side audit-log wiring covers the SetMetricDetail fan-out path, but worker-local transitions (lease expiry on the polled janitor, post-fan-out apply/clear in the worker arm) leave a silent gap. Replicating the audit emission inside the worker `notify` arm needs a new IPC back to the master and is deferred to a follow-up. Update the `SetMetricDetail` doc comment and the `EventKind::METRIC_DETAIL_CHANGED` declaration to scope the contract to operator-initiated transitions, and add a CHANGELOG caveat so audit-log consumers know which transitions are not yet surfaced. The three previous `// TODO(sozu-top week 2):` markers in the worker arm now point to the proto doc as the source of truth for the scope. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 631d4a1 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(proto,command,server): SetMetricDetail TTL lease verb + plumbing Adds a runtime cardinality lease verb so `sozu top` can elevate the metrics drain to `MetricDetailLevel::Backend` for the duration of an interactive session. The lease design (TTL-bounded, `client_id`-keyed, self-expiring) is crash-safe and composes with multiple concurrent clients — see `Aggregator` lease bookkeeping in the previous commit. Proto (additive, backwards-compat): - `Request.request_type::SetMetricDetail = 55` carries `{ client_id, detail?, ttl_seconds?, clear?, reason? }`. - `ResponseContent::metric_detail_status = 16` returns `MetricDetailStatus { configured, effective, previous_effective, workers: map<id, WorkerMetricDetailStatus>, unsupported_workers[] }` for mixed-version-fleet safety. - `EventKind::METRIC_DETAIL_CHANGED = 30` on the `SubscribeEvents` audit stream; distinct from `METRICS_CONFIGURED` (Enabled/Disabled /Clear) since the cause is different. - `command/build.rs` re-attaches `Hash, Eq` for `MetricDetailStatus` (the embedded `map<string, WorkerMetricDetailStatus>` strips the prost auto-derive, which propagates to `ResponseContent.content_type` and `Request.request_type`). - `command/src/proto/display.rs` adds arms for `RequestType:: SetMetricDetail`, `ContentType::MetricDetailStatus` (with a prettytable renderer that lists per-worker configured/effective /previous_effective + unsupported workers), and `EventKind:: MetricDetailChanged`. - `command/src/request.rs` routes `SetMetricDetail` through the worker-level dispatch group (mirrors `ConfigureMetrics`). Master + worker plumbing: - `bin/src/command/requests.rs::is_mutating_verb` learns the new verb so the master brackets it with `RELOADING=1`/`READY=1` systemd hints. - The dispatch match routes through the existing `worker_request` fan-out path (same shape as `ConfigureMetrics`); per-worker `MetricDetailStatus` aggregation lands in week 2 when the TUI starts consuming it. - `lib/src/server.rs::notify` adds two hooks: a polled lease-expiry janitor at the top of every dispatch (gated by `lease_tick_due` so it only walks the lease table every 5 s), and the `SetMetricDetail` arm itself. The arm clamps the TTL via the `Aggregator` setter, decodes the `MetricDetail` enum defensively, and acks with `WorkerResponse:: ok()` for now (week 2 will return the per-worker `WorkerMetricDetailStatus` payload). `MetricDetailChanged` audit emission is left as a `TODO(sozu-top week 2)` at both `lease_apply`/`lease_clear` call sites and the janitor — plumbed through `bin/src/command/requests.rs::audit_emit_inline` once the master collects per-worker effective levels back. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | afe7871 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
fix(command): default missing repeated/map fields when loading older state files `bin/src/command/requests.rs::load_state` reads each `\n\0`-separated JSON record via `command::parser::parse_several_requests::<WorkerRequest>`, which calls `serde_json::from_slice` per record. The prost-build config in `command/build.rs` attaches `Serialize`/`Deserialize` derives to every generated message but did not attach `#[serde(default)]` anywhere, so missing `repeated`/`map` fields rejected the record (`Vec<T>` and `BTreeMap<K, V>` are required-by-serde without an explicit default). Post-1.1.1 schema additions (`Cluster.answers`, `Cluster.authorized_hashes`, `RequestHttpFrontend.headers`, plus the listener-level `answers` / `alpn_protocols` fields) therefore broke `LoadState` for any older client (e.g. proxy-manager pinned to `sozu-command-lib = "1.1.1"`): the first `AddCluster` or `AddHttpFrontend` failed to deserialize, `parse_several_requests`'s `many0(complete(...))` left the unparsed bytes as the remainder, and the read loop reported `"Error consuming load state message"` to the client at EOF. Each post-1.1.1 `repeated`/`map` field on a state-file-emittable message now carries `#[serde(default)]` via a `field_attribute(... )` line in `command/build.rs`, so missing fields default to empty (mirroring the protobuf wire-format default). Required scalars stay strict on purpose: `ConfigState::add_cluster` keys the cluster map by `cluster_id` without a non-empty check, so a struct-level blanket would silently insert a bogus `""`-keyed entry. The field-level annotation preserves that defense-in-depth. The new contract is documented at the top of `command/src/command.proto` and inline in `command/build.rs`: any new `repeated` or `map` field on a message reachable from a `SaveState`/`LoadState` JSON file (anything emitted by `ConfigState::generate_requests`, or carried in a `RequestType` an external client may build and feed through `LoadState`) must add the matching `field_attribute(...)` line. Regression tests in `command/tests/state_compat_v1_1_1.rs` pin both the 1.1.1-shaped fixture round-trip and the missing-required-scalar rejection contract. Signed-off-by: Florentin Dubois <florentin.dubois@clever-cloud.com> Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 7138746 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(hsts): force-replace mode + Set replaces Delete Adds an operator opt-in to override backend-supplied `Strict-Transport-Security` instead of preserving it. Also folds in the broader cleanup that drops `HeaderEditMode::Delete` (zero producers in the diff) in favour of the new `Set` mode that subsumes its semantics via delete-then-insert. - New `HstsConfig.force_replace_backend` proto field (tag 5). Default RFC 6797 §6.1 backend-wins behaviour is unchanged (`HeaderEditMode::SetIfAbsent`); operators flip the field to `true` when a stale or weak upstream HSTS policy needs hardening at the proxy edge (`HeaderEditMode::Set`). - New `FileHstsConfig.force_replace_backend` field, plumbed through `to_proto` to the proto value; documented in `bin/config.toml` and `doc/configure.md`. - New `--hsts-force-replace-backend` CLI flag on `sozu frontend https/http add`. Treated as an enabling flag — `--hsts-force-replace-backend` alone enables HSTS with the canonical default `max-age = DEFAULT_HSTS_MAX_AGE`. Mutually exclusive with `--hsts-disabled`. Two new unit tests cover the force-replace cells. - New `HeaderEditMode::Set` (drop `Delete`). `Set` is delete-then-insert in one entry: the retain pass drops every header with the matching name, then the insert pass appends the new value unconditionally. The legacy empty-`val` Append delete encoding is preserved verbatim for backwards compatibility. - `Frontend::new` chooses `Set` vs `SetIfAbsent` based on `cfg.force_replace_backend`. The single materialiser site is the only consumer; the helper handles both modes uniformly. - Two new `apply_response_header_edits` unit tests cover the `Set` path (replaces existing header / inserts when absent). Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 86e32b2 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(command,lib): add HstsConfig proto + HeaderEditMode for SetIfAbsent Introduce the HSTS (RFC 6797) typed-config foundation: - new HstsConfig proto message with enabled/max_age/include_subdomains/ preload, attached as optional field on HttpsListenerConfig (tag 46), UpdateHttpsListenerConfig (tag 41), and RequestHttpFrontend (tag 16); inline rustdoc cites RFC 6797 §6.1, §7.2, §8.1, §11.4, §14.2 - new HeaderEditMode { Append, Delete, SetIfAbsent } and a `mode` field on HeaderEditSnapshot; apply_response_header_edits learns SetIfAbsent so upstream-supplied Strict-Transport-Security passes through unchanged (RFC 6797 §6.1 single-header requirement) - legacy empty-val Delete encoding preserved via Append+empty fallback so no existing call site needs updating in this commit - unit tests cover SetIfAbsent skip-when-present and insert-when-absent No behaviour change on the wire yet — the materialisation path (router → headers_response) lands in a follow-up commit. Existing struct literals updated only to add 'hsts: None' / 'mode: HeaderEditMode::Append'. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 408ffe8 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(events): add ClusterRecovered event (proto tag 29) Pairs with the existing NoAvailableBackends event (tag 2) so dashboards can plot per-cluster recovery as well as the all-down transition. Highest existing tag was 28 (HealthCheckUnhealthy); 29 was free. Refs #892 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 98ca1a6 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(redirect): add 302 (Found) and 308 (Permanent Redirect) policies Closes #1009. Sōzu's `RedirectPolicy` previously only supported `Permanent` (301) and `Unauthorized` (401) at the frontend. Two new variants: - `RedirectPolicy::Found` → 302 Found (RFC 9110 §15.4.3) — a temporary redirect; user agents MAY rewrite POST → GET on follow. - `RedirectPolicy::PermanentRedirect` → 308 Permanent Redirect (RFC 9110 §15.4.9) — like 301 but the HTTP method MUST be preserved on follow (no GET-rewrite on POST). Wire-level changes: - `command/src/command.proto::RedirectPolicy` gains `FOUND = 3` and `PERMANENT_REDIRECT = 4`. Tags 0..2 unchanged so v1.x clients still decode `FORWARD` / `PERMANENT` / `UNAUTHORIZED` correctly. - `HttpContext` gains `redirect_status: Option<u16>` stashed by `Router::route_from_request` per resolved policy. The answer engine reads it in `mux/mod.rs` to pick the matching `http.{301,302,308}.redirection` template; the legacy `cluster.https_redirect = true` path keeps its 301 default (the field is `None`). - `DefaultAnswer` gains `Answer302` and `Answer308` variants alongside `Answer301`; `set_default_answer_with_retry_after` dispatches all three through one redirect-shaped path. - Default templates `default_302` and `default_308` ship in-tree — `Connection: close`, `Sozu-Id: %REQUEST_ID`, single `Location: %REDIRECT_LOCATION`. Operator-supplied templates flow through the renamed `HttpAnswers::render_inline_redirect(code, …)`; `render_inline_301` is preserved as a thin wrapper for binary / source compatibility. - Per-status counters `http.302.redirection` and `http.308.redirection` fire alongside the existing `http.301.redirection`, both labelled by cluster + backend, mirroring the 301 emission path. Two e2e regression tests in `redirect_rewrite_auth_tests.rs`: - `try_redirect_found_h1_emits_302`: H1 frontend with `RedirectPolicy::Found` must emit 302 + HTTPS Location, no backend contact. - `try_redirect_permanent_redirect_h1_emits_308`: same but for 308. H2 frontend cells, plus the inline-template H2 path, remain to be backfilled in a follow-up; the H1 path covers the routing decision + answer-engine plumbing end-to-end. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | e50bf24 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
fix(router,mux,answers): close PR #1162 follow-up gaps on main Four gaps identified after the redirect/rewrite/answer-template/auth stack landed on `main`. PR #1162 ("Redirection and URL rewrite", @Wonshtrum, base `main`) is structurally superseded by the unified mux implementation but its last commit `448ece33 Rewriting fixes` carried two genuine correctness fixes that did not make it to `main`. The clusterless-redirect ordering bug and the doc/code drift on `%STATUS_CODE` are independent gaps surfaced by the same review pass. * Clusterless `RedirectPolicy::Permanent` now reachable. In `lib/src/protocol/mux/router.rs::route_from_request` the `RedirectPolicy::Permanent` branch is moved ahead of the `Unauthorized || cluster_id.is_none()` deny, so a frontend declared with `redirect = permanent` and no backing cluster (the canonical "this hostname has moved, no service remains" shape from #1161) emits 301 instead of 401. The `Permanent` block does not read `cluster_id`; the cluster-derived knobs already default to safe sentinels at the cluster lookup when `cluster_id` is `None`. The `let Some(cluster_id) = cluster_id else { unreachable!() }` binding moves below the deny block; its invariant still holds. * Non-trie host regex anchored at both ends. In `lib/src/router/mod.rs::convert_regex_domain_rule` the compiled regex now opens with `\A` and closes with `\z`. Without anchors, `Regex::is_match` is unanchored, so an operator's `/example\.com/` matched any hostname containing `example.com` as a substring, including `attacker.example.com.evil.org` — letting an attacker-controlled domain reach a frontend that should only serve `example.com` (CWE-1023, routing bypass). The trie path (`lib/src/router/pattern_trie.rs`) was already anchored. * Inner `/`-finding loop terminates on first match. Same function; the inner loop now `break`s after `found = true`, so a multi-segment regex hostname like `/seg1/.foo./seg2/.com` no longer overwrites `index` on every later `/` and the literal `.` separators between regex segments are kept in their correct position (`\Aseg1\.foo\.seg2\.com\z`). * `%STATUS_CODE` doc/code drift removed. The placeholder was advertised in `doc/configure.md`, the `redirect-template` CLI help in `bin/src/cli.rs`, and the `redirect_template` doc comments in `command/src/command.proto` and `command/src/config.rs`, but `lib/src/protocol/kawa_h1/answers.rs::HttpAnswers::template` never defined a `STATUS_CODE` variable. References are stripped so operators no longer see a promised variable that silently no-ops. Implementation deferred to the same change that adds 302/308 for #1009. Tests: `convert_regex` (updated for the anchored shape) plus three new unit / e2e regressions — `regex_domain_rule_rejects_suffix_and_prefix`, `regex_domain_rule_multi_segment_segments_are_isolated`, and `try_clusterless_permanent_redirect_emits_301` (asserts 301 + Location on the rewritten host). Closes parts of #1161 (proposal — clusterless permanent redirect was the last reachable gap) and #1154 (modify Host header — already addressed by `rewrite_host`, now closeable). Leaves #1009 partially open: 302 (`Temporary`) and 308 are still missing on `main`. Validation: * cargo build --all-features --locked * cargo clippy --all-targets --locked * cargo +nightly fmt --all -- --check * cargo test --workspace --locked (955 passed, 7 ignored) * cargo test -p sozu-e2e -- redirect (19 passed) Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud> Co-authored-by: Eloi Démolis <43861898+Wonshtrum@users.noreply.github.com> Co-authored-by: hcaumeil <78665596+hcaumeil@users.noreply.github.com>
| Commit: | a3d00d5 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
chore(comments,docs): drop review-process leakage; align stale h2c doc comments Inline comments and the proto file referenced PR numbers and 'Codex finding' from the cross-model review pipeline. Replace each with durable technical rationale so the committed text stands on its own merit and ages with the code rather than the review log. Also realign the doc comments on FileClusterConfig.health_check and HttpClusterConfig.health_check that still claimed HTTP/1.1-only probes — this PR adds h2c support, the proto already documents 'probe wire follows cluster.http2', the Rust-side config doc comments should match. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | caa8b1e | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
refactor(health-check): derive h2c probe from cluster.http2; drop is_h2c The probe wire format and the data-plane backend connection both need to choose between HTTP/1.1 and HTTP/2 (h2c) for the same backends. Carrying that decision twice — once on `Cluster.http2` (read by the mux router at `protocol/mux/router.rs::Router::connect`) and once on `HealthCheckConfig.is_h2c` — invites the two flags to drift, so an h2c-only backend gets probed with HTTP/1.1 (or vice versa) the moment an operator updates one without the other. Collapse to a single source of truth: derive the probe wire from `cluster.http2` directly. The probe and the data-plane backend connection now share one switch and cannot diverge. Wire / API surface - `HealthCheckConfig.is_h2c` (proto field 7) is removed. Field 7 remains reserved-by-omission; new fields will start at 8. - `FileHealthCheckConfig.is_h2c` removed; `to_proto` no longer carries the field. - `--h2c` CLI flag on `sozu cluster health-check set` removed. - `BackendMap` gains a `cluster_http2: HashMap<ClusterId, bool>` populated from `Cluster.http2` on every `AddCluster`. The health-check probe reads the entry at probe-creation time and records it on the `InFlightCheck` (`h2c: bool`) so the response parser stays consistent even if the operator flips the flag mid-probe. Documentation - `doc/health_checks.md` rewrites the wire-format paragraph: probe follows `cluster.http2`; no `is_h2c` knob. - The configuration-parameters table drops the `is_h2c` row and gains a paragraph explaining the lockstep with `cluster.http2`. - `CHANGELOG.md` Added bullet rephrased: "cluster-derived h2c probes" instead of "opt-in h2c probes". Validation - cargo build --all-features --locked: pass. - cargo +nightly fmt --all -- --check: pass. - cargo clippy --all-targets --all-features --locked: 0 errors, 0 warnings. - cargo test -p sozu-lib health_check::: 12 passed (existing h2c unit tests still cover the parser; they construct configs without the dropped field). Signed-off-by: Florentin Dubois <florentin.dubois@clever-cloud.com> Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | c6f02b2 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(health-check): h2c prior-knowledge probe support Add an opt-in HTTP/2 prior-knowledge (cleartext, h2c) probe path so operators can health-check backends configured with `cluster.http2 = true` without co-locating an HTTP/1.1 endpoint. Wire / API surface: - `HealthCheckConfig.is_h2c` (proto field 7, optional bool, default false). The TOML key on `[clusters.<id>.health_check]` and the `FileHealthCheckConfig` struct gain a matching field. - `sozu cluster health-check set --h2c` flag forwards the bool through the request-builder validator. Probe wire: - 24-byte connection preface `PRI * HTTP/2.0\r\n\r\nSM\r\n\r\n`. - Empty client SETTINGS frame (0-byte payload, type 0x04, flags 0, stream 0). - HEADERS frame on stream 1 with END_STREAM | END_HEADERS (flags 0x05) carrying a hand-rolled HPACK header block: * `:method GET` indexed (static idx 2 → 0x82). * `:scheme http` indexed (static idx 6 → 0x86). * `:path <uri>` literal w/o indexing, name idx 4 (0x04). * `:authority <host:port>` literal w/o indexing, name idx 1 (0x01). Length octets use the HPACK 7-bit-prefix integer form, including the multi-byte continuation chain for values > 127 bytes. Response parser (`try_parse_h2c_status`): - Walks frames in the buffered response, ignoring SETTINGS, SETTINGS ACK, DATA, etc. until it finds a HEADERS frame on stream 1. - Strips PADDED and PRIORITY prefixes correctly. - Decodes `:status` from either: * Static-table indexed forms 0x88..0x8E (200, 204, 206, 304, 400, 404, 500), or * Literal-with-indexed-name forms 0x08 / 0x18 (literal w/o indexing or never-indexed for name index 8) followed by length + 3-byte ASCII status code. - A GOAWAY frame is treated as a probe failure. - Returns `None` while the buffer is truncated mid-frame so the caller keeps reading. Dispatch in `progress_checks`: a single `parse_probe_response` helper chooses HTTP/1.1 status-line parsing or h2c frame walking based on `config.is_h2c`. The HTTP/1.1 path is unchanged. Tests (in `lib/src/health_check.rs::tests`): - `build_h2c_probe_starts_with_preface_and_settings` - `h2c_indexed_status_200_is_healthy_for_any_2xx` - `h2c_indexed_status_500_fails_default_2xx_check` - `h2c_literal_status_503_matches_expected_503` - `h2c_goaway_marks_unhealthy` - `h2c_truncated_buffer_returns_none` - `h2c_padded_headers_strips_pad_length_octet` Documentation: `doc/health_checks.md` updates the "HTTP/1.1 only" note to describe the new `is_h2c` opt-in, the wire shape, and the parser coverage. HTTPS (h2 over TLS) probes remain a follow-up. Validation: - cargo build --no-default-features --features crypto-ring: pass. - cargo +nightly fmt --all -- --check: pass. - cargo clippy --all-targets --no-default-features --features crypto-ring: 0 errors, 1 cosmetic warning unrelated. - cargo test -p sozu-lib --no-default-features --features crypto-ring health_check::: 11 passed (5 prior + 6 new h2c). Signed-off-by: Florentin Dubois <florentin.dubois@clever-cloud.com> Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 8a1af13 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(health-check): non-blocking HTTP health checks with fail-open and CLI Add HTTP/1.1 backend health-check probing that runs inside the existing single-threaded mio event loop. No additional threads, no async runtime. Key changes: - New lib/src/health_check.rs (~480 LOC) with HealthChecker, threshold- based state machine, jittered intervals, response size cap (4 KB), CRLF-sanitised URIs, per-backend in-flight tracking. Sockets are registered with mio in a dedicated bounded token namespace [HEALTH_CHECK_TOKEN_BASE, HEALTH_CHECK_TOKEN_BASE + HEALTH_CHECK_TOKEN_CAPACITY) = [1<<24, 1<<24 + 1<<16) so the upper bound never falsely claims the mux GOAWAY sentinel Token(usize::MAX). The allocator picks slot offsets modulo the capacity and skips offsets matching in-flight checks; if the table is full it logs an error and returns None rather than silently colliding. - Fail-open routing in lib/src/load_balancing.rs: when ALL backends for a cluster are unhealthy, route to the Normal backends rather than returning 503 (Amazon health-check paper recommendation). - Backend.health::HealthState machine in lib/src/backends.rs with consecutive success/failure counters and Up/Down event emission. - Server event loop integration in lib/src/server.rs: poll the health checker each iteration; dispatch ready tokens; pass Registry for register/deregister. - Server-side validation (command/src/request.rs) rejects zero interval/timeout/thresholds and bad URIs. - Proto: HealthCheckConfig (fields 1-6), SetHealthCheck (cluster_id, config), QueryHealthChecks (optional cluster_id), HealthChecksList (map<cluster, config>) at command/src/command.proto. - CLI: cluster health-check {set,remove,list} subcommands in bin/src/cli.rs with --uri/--interval/--timeout/--healthy-threshold /--unhealthy-threshold/--expected-status flags. CLI-side validation rejects URIs missing a leading '/' and rejects \r, \n, NUL, and any C0 control byte (RFC 9110 §5.1) — not just CR/LF. - Master dispatch in bin/src/command/requests.rs broadcasts state-mutating health-check requests and aggregates worker responses on query. - ConfigState dispatch + persistence in command/src/state.rs. - Tabular display in command/src/proto/display.rs. - Metrics health_check.{success,failure,up,down,healthy_backends}. - Log envelope: every emit in health_check.rs prefixes its format string with "{}, log_context!()" so the lib/tests/log_layout.rs regression guard recognises the canonical HEALTH-CHECK tag (mirrors the hyphenated MUX-H2 / PROXY-RELAY / TLS-RESOLVER convention). Adaptations vs the original PR #1191 commit (rebase onto post-1209 main): - Proto field renumbering. Post-1209 main occupies several carriers PR #1191 originally used; the rebased commit retargets: * Request.set_health_check 47 -> 52 * Request.remove_health_check 48 -> 53 * Request.query_health_checks 49 -> 54 * Cluster.health_check 8 -> 15 * ResponseContent.health_checks_list 14 -> 15 * EventKind.HEALTH_CHECK_HEALTHY 5 -> 27 * EventKind.HEALTH_CHECK_UNHEALTHY 4 -> 28 - Token namespace tightened (bounded modulo allocator + bounded owns_token range) to avoid claiming the mux GOAWAY sentinel. - URI sanitisation hardened to reject every C0 control byte. - ClusterCmd::HealthCheck variant placed alongside ClusterCmd::H2 (PR #1191's structure), not nested under ClusterH2Cmd. Validation: - cargo build --no-default-features --features crypto-ring: pass - cargo +nightly fmt --all -- --check: pass - cargo clippy --all-targets --no-default-features --features crypto-ring: 0 errors, 1 cosmetic warning (redundant_closure on Criterion bench setup, unrelated to this commit). - cargo build --tests --no-default-features --features crypto-ring: pass. Signed-off-by: Florentin Dubois <florentin.dubois@clever-cloud.com> Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 14ea96c | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(splice): operator-tunable kernel-pipe capacity Make the splice(2) kernel-pipe capacity per direction configurable via a new `splice_pipe_capacity_bytes` field on ServerConfig, instead of the hard-coded 64 KiB constant. The override flows through the same path as the existing `basic_auth_max_credential_bytes` knob: declared in `command/src/command.proto` (ServerConfig field 22), wired through `FileConfig`, the `ConfigBuilder` mapper, the runtime `Config`, and the proto round-trip via `From<&Config> for ServerConfig`. Linux-only storage at the lib layer: a `OnceLock<usize>` in `lib/src/splice.rs` populated once per worker boot from `Server::try_new_from_config` (cfg-gated), with a setter that no-ops on `0` so an explicit zero does not collapse the pipe to PAGE_SIZE. `SplicePipe::new` now applies the configured capacity via `fcntl(F_SETPIPE_SZ)` on each pipe and reads the realised value back with `fcntl(F_GETPIPE_SZ)`, storing the smaller of the two sizes in a new `capacity` field. This handles the kernel's behaviour: it rounds up to PAGE_SIZE and clamps at `/proc/sys/fs/pipe-max-size` (default 1 MiB unprivileged; CAP_SYS_RESOURCE goes higher). On `F_SETPIPE_SZ` failure the kernel keeps the previous capacity (typically 64 KiB), the failure is logged at `warn!`, and SplicePipe continues with that realised value — splice still works, just at the kernel default. `splice_in` gains a `len: usize` parameter so callers thread the realised capacity through; pipe.rs reads `splice_pipe.capacity` for both the per-call `len` and the "pipe is full" backpressure check (replaces the deleted `SPLICE_PIPE_CAPACITY` const). CHANGELOG, doc/getting_started.md (feature-flag table), and doc/configure.md (global parameters table) updated to document the new knob and its kernel-side limits. Signed-off-by: Florentin Dubois <florentin.dubois@clever-cloud.com> Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 1a9e951 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(http,https,mux): add X-Real-IP injection and anti-spoof elision (H1+H2) Two listener-scoped opt-in bool flags, both default false and independently combinable. Configurable on `HttpListenerConfig` / `HttpsListenerConfig` and runtime-patchable through the new `Update*ListenerConfig` partial-update verbs. - elide_x_real_ip: when true, any client-supplied `X-Real-IP` header is stripped from the request before forwarding (anti-spoofing). - send_x_real_ip: when true, a proxy-generated `X-Real-IP` header carrying the connection peer IP is appended. The IP is read from the post-PROXY-v2 `session_address`, so deployments terminating PROXY v2 surface the original client IP, not the upstream proxy's. Both H1 and H2 are covered by a single elision branch in `HttpContext::on_request_headers`, dispatched from `pkawa::handle_header` for H2 initial HEADERS frames. H2 trailer HEADERS frames take a separate path through `pkawa::handle_trailer` and would otherwise bypass the elision; `handle_trailer` now drops `x-real-ip` trailer pairs when the listener flag is set, closing that gap. Five e2e tests cover the four-flag matrix on H1 (including PROXY-v2 unwrap with original client IP) plus an H2 trailer regression placeholder. Mirrors the `strict_sni_binding` precedent for trait accessor / mux Context propagation / Update*ListenerConfig apply path. Closes #1113 Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud> Co-authored-by: Eloi Démolis <43861898+Wonshtrum@users.noreply.github.com> Co-authored-by: hcaumeil <78665596+hcaumeil@users.noreply.github.com>
| Commit: | e49b731 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(server,mux,tcp,cli): per-cluster per-IP connection limit with HTTP 429 + graceful TCP close Caps the number of simultaneous frontend connections one source IP may hold against a given cluster. Replaces the simpler per-source-IP variant the previous PR #1193 commits proposed (those did not survive the H2 mux unification). Closes #890, #1057. Wire shape (proto): - ServerConfig.max_connections_per_ip = 22 [optional, default 0] - ServerConfig.retry_after = 23 [optional, default 60] - Cluster.max_connections_per_ip = 13 [optional, override] - Cluster.retry_after = 14 [optional, override] - CustomHttpAnswers.answer_429 = 12 [optional, custom 429 template] - Request.set_max_connections_per_ip = 50 (uint64) - Request.query_max_connections_per_ip = 51 (QueryMaxConnectionsPerIp {}) - ResponseContent.max_connections_per_ip_limit = 14 (MaxConnectionsPerIpLimit { limit: u64 }) Override semantics: cluster `None` inherits the global default, `Some(0)` is explicit "unlimited for this cluster", `Some(n > 0)` overrides. Source IP is the parsed PROXY-protocol source when present, else `peer_addr`. Enforcement: - HTTP/HTTPS via the unified mux at protocol/mux/router.rs::connect: after cluster resolution, before backend selection. The check fires AFTER auth/redirect/SNI decisions so a 401/421/redirect frontend never trips the limit. On hit: stash the resolved Retry-After on the stream context and return BackendConnectionError::TooManyConnectionsPerIp, which the mux converts into a 429 default answer through the new set_default_answer_with_retry_after path. Covers H1 and H2 — the unified mux serves both protocols. - H2 multiplex semantics: SessionManager keeps a per-token HashSet<(cluster, ip)>. Multiple streams to the same (cluster, ip) from the same H2 connection share one slot — the limit governs distinct frontend connections, not requests. Decrement is wholesale on session close (untrack_all_cluster_ip). - TCP at tcp.rs::connect_to_backend: rejection produces a graceful FIN via SessionResult::Close (no SO_LINGER trick). Answer engine (lib/src/protocol/kawa_h1/answers.rs + lib/src/protocol/mux/answers.rs): - Answer429 variant on DefaultAnswer with retry_after: Option<u32>. - 429 template registered alongside 421/503/etc. - New %RETRY_AFTER variable with `or_elide_header = true`: when the resolved value is 0/None, the engine drops the entire `Retry-After:` line. `Retry-After: 0` would invite an immediate retry that defeats the limit, so we omit instead of rendering literal 0. - legacy_to_map flattens CustomHttpAnswers.answer_429 into the per-listener answers map. - default_answer_for_code(429, ...) builds the variant. - set_default_answer_with_retry_after threads the resolved retry value through to the rendered answer. Proxy-protocol fixes (lib/src/protocol/proxy_protocol/{expect,relay}.rs): - ExpectProxyProtocol::into_pipe and RelayProxyProtocol::into_pipe now use ProxyAddr::source() from the parsed v2 header instead of the raw TCP peer_addr. Without this fix the pipe phase records the upstream PROXY-emitter (an LB / edge proxy / health-check probe), not the originating client — which means the per-(cluster, source-IP) limit would have keyed on the LB's IP. Relay mode previously discarded the parsed addresses entirely; we now stash them on the struct. SessionManager (lib/src/server.rs): - max_connections_per_ip: u64 + retry_after: u32 fields. - connections_per_cluster_ip: HashMap<(String, IpAddr), usize> + cluster_ip_tracks: HashMap<Token, HashSet<(String, IpAddr)>>. - cluster_ip_at_limit (token-aware: a token already holding a slot is never at the limit), track_cluster_ip (idempotent within a token), untrack_all_cluster_ip (drains on session close), effective_max_connections_per_ip / effective_retry_after (override resolution), clear_cluster_ip_tracking (runtime disable). - HttpProxy / HttpsProxy / TcpSession close paths drain via untrack_all_cluster_ip on the frontend token before the slab slot is reused. ProxySession trait (lib/src/lib.rs): - New cluster_id() and session_address() default-implemented as None. - New L7Proxy::sessions() so the mux router can reach the SessionManager from a Rc<RefCell<dyn L7Proxy>>. - New BackendConnectionError::TooManyConnectionsPerIp variant with cluster_id payload. CLI (bin/src/{cli,ctl/mod,ctl/request_builder}.rs): - `sozu connection-limit set <N>` / `remove` / `show` patches the global limit at runtime. The setter is non-sticky (workers reset to the TOML-configured value on restart); operators must mirror the change in the config to make it durable. - `--answer-429 <path>` flag added to `sozu listener {http,https} update`. - New build_http_answers helper threads answer_429 through the existing CustomHttpAnswers builder. Config (command/src/config.rs): - DEFAULT_MAX_CONNECTIONS_PER_IP = 0 (disabled), DEFAULT_RETRY_AFTER = 60. - FileConfig + Config + ServerConfig From conversion threaded. - FileClusterConfig + HttpClusterConfig + TcpClusterConfig + ListenerBuilder + CustomHttpAnswers gain the new fields. Metric: connections.rejected_per_cluster_ip (counter, labelled by cluster_id, backend_id always empty — rejection happens before backend selection). Replaces the simpler `connections.rejected_per_ip` from the older PR #1193 commits — single name covers H1, H2, TCP. Tests: - e2e/src/tests/cluster_ip_limit_tests.rs covers H1 global limit (429 + Retry-After), H1 Retry-After: 0 elides the header, H1 per-cluster override (unlimited cluster coexists with capped), TCP graceful close on limit. All four pass `cargo test -p sozu-e2e cluster_ip_limit`. Validation: - cargo build --all-features --locked - cargo +nightly fmt --all -- --check - cargo clippy --all-features --all-targets --locked -- -D warnings - cargo test --workspace --locked Docs: bin/config.toml, doc/configure.md (parameters + metric), CHANGELOG.md. Co-authored implementation guidance from Codex on the H2 multiplex semantics, Retry-After: 0 elision, PROXY-protocol Relay-mode address stashing, and rolling-upgrade safety (optional proto fields, not required). Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud> Co-authored-by: Eloi Démolis <43861898+Wonshtrum@users.noreply.github.com> Co-authored-by: hcaumeil <78665596+hcaumeil@users.noreply.github.com>
| Commit: | fd7b77f | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(server): evict least-active sessions when accept queue is full Add an opt-in `evict_on_queue_full` knob (proto `ServerConfig` field 22, default `false`) that, when set, evicts the oldest 1% of non-listener sessions to make room for queued sockets when `SessionManager::check_limits` refuses a new accept. Selection uses `select_nth_unstable_by_key` (introselect, O(n) average) over `Session::last_event()` to partition the candidate slab in place, avoiding an O(n log n) sort. The cap loop in `Server::create_sessions` keeps the existing `incr!("listener.connection_capped")` counter (so dashboards stay meaningful regardless of eviction outcome), runs the eviction batch, and re-checks limits before continuing. A new `sessions.evicted` counter is emitted only when the mitigation fires. Eviction is deliberately skipped during graceful `shutting_down`: forcing sessions closed there defeats the shutdown semantics and is wasted work since the worker is winding down anyway. Default `false` because during a DDoS the active sessions are more likely to be legitimate clients than the queue contents — evicting them would serve attackers. Operators dominated by normal traffic spikes can flip it on. Closes #644, closes #916. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | daf4dd3 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(config,auth): expose basic_auth_max_credential_bytes + warn at 33% buffer The maximum length of a base64-decoded `Authorization: Basic` payload the worker accepts was a hard-coded 4 KiB. Exposing it in the main TOML config lets operators on hardened tenants lower it (256 / 512 typical) to bound the per-failed-auth allocation tighter against hostile peers sending large tokens. Surfaces: - New `ServerConfig.basic_auth_max_credential_bytes` proto field (tag 21, optional uint64), threaded from `FileConfig.basic_auth_max_credential_bytes` through `Config` so the value rides the existing config-load and state-restore plumbing. - `lib::protocol::mux::auth` replaces the const cap with a `OnceLock<usize>` overlay over a built-in default (4096). The override is committed once on each worker at boot via `lib::server::Server::try_new_from_config` calling `set_max_decoded_credential_bytes` — a single set-once handoff with no per-request atomic. Subsequent attempts to mutate are no-ops (`OnceLock::set` rejects), and an explicit `0` is treated as "use default" so a typo cannot disable the cap by accident. New unit test pins the zero-is-noop semantics. - Config validator emits a `warn!` at boot when the operator-set cap is `>= buffer_size / 3`. At that point a single failed-auth attempt can pin ~33% of the per-frontend buffer's worth of bytes; combined with in-flight request/response framing the buffer trends toward back-pressure under load. Informational only — operators with a deliberate threat model can keep the value, but the surprise stays visible in the boot log. - `doc/configure.md` "HTTP Basic authentication" section gains a "Tuning the credential decode cap" subsection covering both the knob and the 33% warning. `CHANGELOG.md` mentions both. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | b02fccb | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
chore(proto): drop preemptive `reserved` blocks on Cluster and RequestHttpFrontend The `reserved 13 to 19` (Cluster) and `reserved 16 to 19` (RequestHttpFrontend) blocks were defensive "reserve room for future fields" — not idiomatic protobuf. The `reserved` keyword is meant to guard against tag REUSE after a field is deleted, so an old serialised message doesn't alias into a new field with an incompatible type. We didn't delete anything; we only added new fields. The next person adding a field would just pick the next free tag (13 / 16) anyway, and removing the `reserved` line would be a single-line edit if those numbers were genuinely needed. Drops the blocks; the regenerated `command.rs` shrinks accordingly. No wire-format change, no downstream impact. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | b844c6b | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
fix(cli,router,auth): clear /review's deferred Medium/Low/Nit backlog Closes the items the prior `fix(answers,auth,router)` pass deferred: * M2 — `sozu frontend {http,https} add` gains the per-frontend policy flags (`--redirect`, `--redirect-scheme`, `--redirect-template`, `--rewrite-host`, `--rewrite-path`, `--rewrite-port`, `--required-auth`, `--header <position>=<name>=<value>`). The two `frontend …` siblings now share a single `build_http_frontend_add` helper so they cannot drift; each policy field is validated up-front (range-checked, scheme/policy parsed against the proto enum) so a malformed input surfaces as a typed `CtlError::ArgsNeeded` instead of reaching the worker. * M3 — `Route::Frontend(Rc<Frontend>)` is reachable. `HttpFrontend` carries the new policy fields all the way through `RequestHttpFrontend::to_frontend`, and `Router::add_http_front` flips onto the rich path whenever any policy field is non-default (otherwise still produces the legacy `Route::ClusterId` / `Route::Deny` shapes). The dead-arm warning is no longer hypothetical — the variant is exercised end-to-end on the live routing path. * L4 — `mux::auth::canonicalize_basic_credentials` tightens the whitespace grammar to `1*SP` per RFC 7235 §2.1 / RFC 9110 §11.4. SP before and between scheme and token68 is tolerated; HTAB and zero-spaces are rejected. * L5 — `ConfigError::InvalidHeaderPosition` becomes `{ index, position }` so a multi-entry config pinpoints the bad row. `parse_header_edit` takes the array index from `entries.iter().enumerate()`. * L6 — new `template_fill_adjacent_body_variables` regression test pins the `body_size` accounting against an `[%ROUTE%ROUTE]` body so a future tweak to the inter-chunk arithmetic cannot silently break back-to-back placeholders. * N2 — `Cluster` reserves field numbers 13-19 and `RequestHttpFrontend` reserves 16-19. A future migration that accidentally aliases an old field number now fails at proto compile time. Knock-on: every test fixture and example builder constructing `HttpFrontend { … }` directly inherits the eight new `Option`/`Vec` fields; existing in-tree call sites are populated with `None` / `Vec::new()` (see `lib/src/http.rs::frontend_from_request_test`). The 4 cluster-add CLI tests added in c1e0b1a6 plus the new `template_fill_adjacent_body_variables` bring sozu-lib to 436 passing tests; sozu-command-lib stays at 66/5 ignored. Build, clippy `-D warnings`, fmt --check, log-layout regression all clean. Out of scope for this pass * The header-injection runtime path on the mux side still needs to consume `Frontend.headers_request` / `headers_response` at `mux/h1.rs::writable` and `mux/h2.rs::write_streams`. The plumbing is in place (the routing layer builds and stashes the slices) — the actual emission is the next round of work. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | a94208e | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
fix(answers,auth,router): act on /review findings Apply the High and tractable Medium/Low/Nit items the read-only review surfaced. Build, clippy --all-targets -D warnings, fmt --check, sozu-lib (435 pass) and sozu-command-lib (66 pass) all green. High * answers.rs: stop unconditionally splicing a synthetic `Content-Length` header into rendered responses. The engine now detects an operator-supplied `Content-Length:` line during the parse pass and routes its value through the existing `ContentLength` placeholder, so the rendered response carries exactly one header with the auto-computed body size. Previously a template with a literal `Content-Length:` line ended up with two headers — RFC 9110 §8.6 / RFC 7230 §3.3.2 request-smuggling vector. * mux/auth.rs: pad both candidate and stored hash into a fixed `[u8; AUTH_COMPARE_PAD_LEN + 8]` envelope (256-byte body + 8-byte little-endian length) before `subtle::ConstantTimeEq::ct_eq`, so the per-entry compare loop iterates the full padded length even when lengths differ. `subtle`'s slice `ct_eq` short-circuits on length mismatch — the padding here defeats that leak. Closes the realm-size and matching-username-length channels (CWE-208 family). * kawa_h1/editor.rs: `HttpContext::reset()` now clears `redirect_location` and `www_authenticate` between pipelined H1 requests so a future 301/401 default-answer path that bypasses routing cannot inherit a stale Location / realm from a prior request. Backed by the existing `test_reset_clears_request_response_state` regression, extended with the two new field assertions. Medium * command.proto: `HeaderPosition` gains an explicit `HEADER_POSITION_UNSPECIFIED = 0` variant; the proto-default-encoded shape now deserialises into a typed "unset" instead of failing `HeaderPosition::try_from(0)`. The runtime drops `Header { position: Unspecified, … }` entries with a `warn!` rather than guessing a position. * mux/auth.rs: replaces the hand-rolled `to_hex` (which carried two `expect("hi/lo nibble")` panics on an auth path) with `hex::encode` from the existing `hex` workspace dep. Drops the corresponding unit test and the `_METHOD_USED` workaround that kept the otherwise- unused `Method` import alive. * kawa_h1/answers.rs: `HttpAnswers::new` propagates `TemplateError` from the bundled-default-template parse path via `?` instead of `expect(...)`. The new `default_templates_all_parse` regression test guards the invariant. * router/pattern_trie.rs: new `segment_regex_rejects_partial_matches` test pins the `\A...\z` anchoring contract — `cdn[0-9]+` matches `cdn1` and `cdn123` but rejects `cdn1xxx`, `xxxcdn1`, and `cdnabc`. Without the test the next clippy or refactor pass could silently revert anchoring. Low / Nit * request_builder.rs: validate `--https-redirect-port` against `1..=65535` before sending so a typo doesn't render `Location: https://host:70000/...` on the wire. Lifts `looks_like_authorized_hash` to module scope and adds 6 unit tests covering canonical form, missing colon, short hex, uppercase, empty username, non-alnum username. * router/mod.rs: drop the dead-defensive `from_utf8(b).unwrap_or_default()` in `RewritePart`; `pattern` is `template.as_bytes()` and the split is on the ASCII byte `$`, so `template[start..i]` lies on char boundaries and indexes safely. * mux/router.rs: gate the legacy `cluster.https_redirect` `redirect_location` stash on `proxy.kind() == ListenerType::Http` so an HTTPS listener never carries a stale URL into a downstream default-answer path. Replaces `cluster_id.expect(...)` with a `let-else { unreachable!() }` form per project style. * router/mod.rs: tighten the `log_module_context!` doc-comment to reflect the single-call-site reality (`Frontend::new`'s warn). Out of scope (deferred to follow-up) * /review M2: per-frontend `--redirect`, `--rewrite-*`, `--header` CLI flags (cluster-level flags shipped in c1e0b1a6). * /review M3: making the `Route::Frontend(Rc<Frontend>)` variant reachable from `add_http_front` — needs the Wave 1c bridge that converts `RequestHttpFrontend` policy fields into a `Frontend`. * /review L4: `Authorization` whitespace tolerance (RFC 7235 §2.1 permits SP only; current accepts SP and HT). * /review L5: `parse_header_edit` failure does not include array index — minor diagnostic polish. * /review L6: `body_size -= variable.name.len() + 1` accounting fuzz test for adjacent `%VAR%VAR` patterns. * /review N2: `reserved` declarations in `Cluster` / `RequestHttpFrontend` / listener configs. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | e519939 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(proto): add answer templates, redirect, rewrite, headers, auth schema Lay the wire schema for the bundled feature work that brings PR #1206..#1210 onto current main. Wave 1 of a multi-wave implementation; subsequent waves land the runtime engine, mux integration, auth helper, e2e tests, and docs. Cluster gains four fields: per-cluster `answers` (status code → template body) at field 9, `https_redirect_port` at 10, `authorized_hashes` at 11, and `www_authenticate` realm at 12. RequestHttpFrontend gains eight: `redirect` policy at 8, `required_auth` at 9, `redirect_scheme` at 10, `redirect_template` at 11, `rewrite_host`/`rewrite_path`/`rewrite_port` at 12-14, and a repeated `headers` list at 15. Listener configs gain a parallel `answers` map at 31 (HttpListenerConfig) and 43 (HttpsListenerConfig); the legacy `CustomHttpAnswers http_answers` field is preserved on the wire so existing state files round-trip — the runtime will read both for one minor. New top-level enums RedirectPolicy / RedirectScheme / HeaderPosition and a Header message land alongside. An empty `Header.val` deletes the named header (HAProxy `del-header` parity). prost stops auto-deriving Hash/Eq once a message holds a map field, so the build script gains explicit `#[derive(Hash, Eq)]` for Cluster, HttpListenerConfig, HttpsListenerConfig, UpdateHttpListenerConfig, and UpdateHttpsListenerConfig. Existing initializers grow a `..Default::default()` fallthrough so this commit is a pure additive scaffolding step — no runtime behaviour changes yet. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 321197d | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(config): expose slab_entries_per_connection knob (default 4, [2,32]) The slab capacity multiplier was a private constant SLAB_ENTRIES_PER_CONNECTION = 4 introduced by the H2 mux work to accommodate stream multiplexing (1 frontend + up to 3 backend connections per session). Operators with topologies that fan out across more than 4 backends per session had no recourse short of a recompile, and slab exhaustion presents as "accept refused" with no telemetry pointing at the slab. Adds an optional uint64 slab_entries_per_connection field to ServerConfig (proto tag 20) and a matching Config field. Effective value flows through ServerConfig::effective_slab_entries_per_connection which clamps to [MIN_SLAB_ENTRIES_PER_CONNECTION = 2, MAX_SLAB_ENTRIES_PER_CONNECTION = 32]; absent or 0 falls back to DEFAULT_SLAB_ENTRIES_PER_CONNECTION = 4 so existing deployments keep their current capacity. Documented in doc/configure.md alongside max_buffers / buffer_size. A `worker.slab.utilization_pct` operator-facing gauge is left as a follow-up (gauge wiring lives in the worker hot path; this commit is config plumbing only). PR #1209 consolidated review (MED-6). Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 2118dfa | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
refactor(command): finish audit log enrichment — ts, actor_role, connect_ts, boot_generation, build_git_sha, JSON sink Closes the deferred items from 6b21c9c4. After this commit, the only audit-log gaps left are architectural decisions that need their own designs: OTel trace_id/span_id (proto bump on Request to carry tracing context end-to-end) and Linux auditd integration (new kernel-side sink, large scope). `source_ip=` stays permanently N/A while the command channel is unix-only. New mandatory fields on every `Command(...)` line ------------------------------------------------- * `ts=<RFC3339 UTC microseconds>` — in-body timestamp. Lets operators extract a single audit line and still know when it fired without cross-referencing the outer logger prefix. Implemented with a std-only Hinnant `civil_from_days` converter — no `time` or `chrono` dep added. * `actor_role=root|system|user|unknown` — buckets the actor uid so SOC dashboards can route on it without parsing UID values per host (uid=0 → root, 1..1000 → system, 1000+ → user, missing → unknown). * `connect_ts=<RFC3339 UTC>` — wall-clock `accept(2)` time of the client connection, stamped on `ClientSession.connect_ts` via `SystemTime::now()`. Forensic windowing — "all verbs from connections that opened in the 30s before the incident". * `boot_generation=<u32>` — counter incremented at every `MAIN_UPGRADED` re-exec, persisted across the re-exec via `UpgradeData.boot_generation`. Disambiguates post-upgrade sessions from pre-upgrade ones — PIDs reset, but `(boot_generation, session_ulid)` is a durable correlation pair. * `build_git_sha=<12hex>` — short git SHA embedded by `bin/build.rs` via a `cargo:rustc-env=SOZU_BUILD_GIT_SHA=…` directive. Falls back to `unknown` outside a git tree (vendored tarballs, sysroots). `bin/build.rs` re-runs only when `.git/HEAD` or `.git/refs` move so cached cargo builds stay fast. JSON sink — `audit_logs_json_target` ------------------------------------ Mirrors every audit line as a single-line JSON object to a dedicated file. Same `O_APPEND | O_CREAT | 0o640` lifecycle as the human sink. Schema is stable; missing values are JSON `null` so SIEM pipelines can flatten without conditional fields. The `audit` block groups actor identity (uid/gid/pid/user/comm/role) and the `extras` block groups completion-time fields (elapsed_ms, fanout, error_code, reason, request_sha256). Both sinks (text + JSON) are independent — operators can set both for tail-friendly + machine-parseable. Plumbed through `FileConfig` → `Config` → `ServerConfig` (proto tag 19) so workers see the field even though only the main process writes to it. Server gains `audit_log_json_writer: Option<RefCell<File>>` opened at boot via the same `open_audit_log_file` helper as the text sink. `Server.boot_generation` + `UpgradeData.boot_generation` ------------------------------------------------------- * New `pub boot_generation: u32` on `Server`, `0` on first boot, `saturating_add(1)` at every `upgrade_main` invocation BEFORE the re-exec serialises `UpgradeData` so the new main starts at the bumped value. * `UpgradeData.boot_generation` carries it across the re-exec boundary (`#[serde(default)]` so old upgrade payloads decode as `0`). * The `MainUpgraded` audit line records the new generation in its `target=` field for the audit-trail itself. Documentation + retention policy (LISA-013) ------------------------------------------ * `doc/observability.md` documents the JSON schema with a worked example. * New retention-policy section: PCI-DSS 10.7 calls for ≥ 1 year of audit retention with the most recent 3 months immediately available. Recommended shape: `logrotate` daily compress, off-host archive after 90 days, 400-day cold retention. Sōzu does NOT rotate audit files itself — delegated to the OS-level rotator with `copytruncate` since sōzu keeps the file handle open. (`SIGHUP` re-open is a TODO.) * Sample configs (`bin/config.toml`, `os-build/config.toml`) document `audit_logs_json_target` next to `audit_logs_target`. Drift-guard ----------- * `audit_format_tests` extended with three new tests: `actor_role_buckets`, `rfc3339_utc_round_numbers` (epoch + Y2K+1day with microsecond fraction), and `build_git_sha_format` (either 12 hex chars or `unknown`). * Regex updated to anchor on the new mandatory fields (`ts`, `actor_role`, `connect_ts`, `build_git_sha`, `boot_generation`) and accept the optional extras in any order between `result=…` and `sozu_version=…`. * `TestServer` stub mirrors `Server.boot_generation` so the macro expansion compiles in tests without the full Poll / listener ceremony. Macro signature change ---------------------- `audit_log_context!` now takes `($server, $client, $request_id, $entry, $result)` instead of `($client, $request_id, $entry, $result)`. The only in-tree caller is `audit_emit` which already has `server: &mut Server` in scope; updated trivially. Test stubs got a matching `TestServer { boot_generation }` minimal struct. Followups (truly architectural — not in this commit) ---------------------------------------------------- * OTel `trace_id` / `span_id` propagation needs a `Request` proto bump to carry tracing context end-to-end from sozuctl through the command channel. Worth its own design. * Linux auditd subsystem integration would be a third sink with kernel-side persistence — new dep + integration design. * `source_ip=` stays permanently N/A while the command channel is unix-only. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | da38ca3 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
refactor(command): audit log P2+P3 — ucred/user/socket, dedicated sink, diff before→after, request sha256, cert-replace new fingerprint Follow-up to 53b51e03. Takes the audit log from "MUX-layout line with the core five fields" to an operationally-complete control-plane trail that satisfies PCI-DSS 10.5 routing and records everything a SOC analyst needs to reconstruct a control-plane incident without external context. New fields on `Command(...)` ---------------------------- * `actor_user=<username>` — resolved NSS account name (`getpwuid_r(uid)` at accept time via nix `user` feature). Distinct from `actor_uid` which stays the primary attribution key. `unknown` when lookup fails. * `socket=<path>` — command-socket path the client connected through. Propagated via `Arc<str>` on `CommandHub` and cloned onto each accepted `ClientSession`. Disambiguates multi-instance deployments that share a SIEM sink. * `request_sha256=<16hex>` — truncated (64-bit, first 16 hex chars) SHA-256 of the proto `Request` wire-encoding. Set on verbs that flow through `worker_request`; useful for replay / dedupe detection. Uses the workspace's existing `sha2` dep (same crate already hashes certificates in `command/src/certificate.rs`). Dedicated sink (`audit_logs_target`) ------------------------------------ New optional config field on `FileConfig` / `Config` / `ServerConfig` (+ proto tag 18). When set to a filesystem path (e.g. `/var/log/sozu/audit.log`), every audit line is *also* appended to that file opened `O_APPEND | O_CREAT` with mode `0o640` so granting an `audit` group tail-only access via filesystem ACL is one step away. ANSI escape sequences are stripped before writing to the dedicated sink via a single-pass `strip_ansi` helper so the file stays SIEM-parseable regardless of `log_colored`. Write failures log a warning and never block the mutation. `None` (default) keeps audit lines routed only through `log_target`. PCI-DSS 10.5 ("protect audit trails") can now be met with a simple filesystem ACL and logrotate config. Sample configs (`bin/config.toml`, `os-build/config.toml`) document the option under `[General]` next to `access_logs_target`. Smarter `target=` contents -------------------------- * **`UpdateHttp/Https/TcpListener`** — `format_patch_diff_*` now takes an optional snapshot of the pre-patch listener state and emits each patched field as `field=old→new` instead of just `field=new`. Falls back to `field=?→new` when no current listener is known (e.g. the patch arrived before the listener was registered), so the audit line never swallows a change. * **`ReplaceCertificate`** — the new cert's fingerprint is computed at audit time via `sozu_command_lib::certificate::calculate_fingerprint` and included alongside the old one: `target=certificate:<addr>:old=<fp>:new=<fp>`. Forensic win: cert rotation patterns + substituted-cert detection. Dep --- * `nix` gains the `user` feature (for `User::from_uid`). * `sha2` added as an explicit `bin/` dep (was already a workspace dep used by `command/src/certificate.rs`). Drift-guard ----------- Regex extended to cover the four new mandatory fields (`actor_user`, `socket`) and the new optional `request_sha256` extras slot. All four existing tests still green. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | df2a4f4 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
refactor(command): enrich audit log — MUX layout, full ucred, fanout, injection defence Follow-up to 0c2dc869 — takes the audit log from a proof-of-concept into a compliance-grade control-plane log. Driven by parallel codex / lisa analyses (`tasks/audit-log-enhancements/*.md`). Layout ------ * Drop the trailing `\t >>>` continuation marker — the line is self-contained, nothing follows the closing paren. * Rename the keyword from `Session(...)` to `Command(...)` — the payload describes a control-plane command, not a proxy session. Retains the MUX-family bracket + tab + uppercase `AUDIT` tag layout. Taxonomy (proto EventKind additions — 10 new variants) ------------------------------------------------------ * `STATE_LOADED`, `STATE_SAVED` — `LoadState`/`SaveState` emit at task completion with `ok:<n> errors:<n>` in `target=`. * `LISTENER_ADDED`, `LISTENER_REMOVED` — `AddHttp/Https/TcpListener` and `RemoveListener`. * `SOZU_STOP_REQUESTED` — SoftStop / HardStop; `target=stop:soft|hard`. * `MAIN_UPGRADED`, `WORKER_UPGRADED` — hot-upgrade entry points. * `EVENTS_SUBSCRIBED` — SubscribeEvents subscribers. Actor identification -------------------- * `peer_cred_from_stream` captures the full SO_PEERCRED triple (uid, gid, pid), not just uid. * `peer_comm(pid)` reads `/proc/<pid>/comm` at accept time so operators can tell `sozuctl` apart from ad-hoc shells sharing a UID. * `ClientSession` gains `actor_gid`, `actor_pid`, `actor_comm`. Completion-time audit + timing ------------------------------ Worker-fanning verbs emit two audit lines — attempt-time (accepted by main state) and completion-time (applied across workers). The completion line carries `fanout=ok|partial|timeout|local_only`, `workers=<ok>/<err>/<expected>`, `elapsed_ms`, and on err paths `error_code` + `reason`. Structured error taxonomy ------------------------- New `AuditErrorCode` enum (dispatch_error, worker_failure, worker_timeout, peer_cred_unavailable, invalid_input, io_error, other). Wired at dispatch rejection, worker fan-out failures/timeouts, logging filter parse errors, and state save/load I/O errors. Log-injection defence --------------------- New `sanitize_for_audit` helper replaces ASCII control chars with `?`. Applied at render time to every attacker-influenced field (target, actor_comm, reason) so embedded tabs/newlines/ANSI escapes cannot forge additional audit lines. Drift-guard test covers the injection scenario. Dependency ---------- `rustls-webpki` 0.103.12 → 0.103.13 (RUSTSEC-2026-0104 reachable panic in CRL parsing). Out of audit codepath, flagged by CVE sweep. Followups (not in this commit) ------------------------------ * Dedicated tamper-resistant audit sink (`audit_logs_target` config + O_APPEND file) — PCI-DSS 10.5. * `socket_path=` field for multi-instance SIEM disambiguation. * `actor=<username>` via getpwuid_r cached at accept. * `request_sha256=` for dedupe / tamper-detection hints. * `ReplaceCertificate` target currently carries old fingerprint only — add new fingerprint too. * Patch-diff formatters log new values only — include old values. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 1b1a718 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(command): sozu listener {http,https,tcp} update runtime-patch verb Introduces an in-place update verb for non-bind-only listener settings, so operators can tune CVE-related H2 flood thresholds, SNI binding, disable_http11, ALPN, graceful-shutdown deadline, stream-0 WINDOW_UPDATE cap, sozu_id_header, custom HTTP answers, and timeouts under attack without cycling the listening socket. Wire protocol: - UpdateHttpListenerConfig, UpdateHttpsListenerConfig, UpdateTcpListenerConfig (+ AlpnProtocols wrapper so absent/empty is unambiguous) in command.proto, RequestType tags 47/48/49, EventKind::LISTENER_UPDATED = 18. Control plane (command/src/state.rs): - ConfigState::update_{http,https,tcp}_listener with field-mask merge. - merge_custom_http_answers preserves per-field http_answers so a patch that sets only answer_503 does NOT wipe answer_401/404/etc (regression of the earlier HttpAnswers blocker). - Server-side validation (validate_h2_flood_knobs_*, validate_alpn_*, validate_sozu_id_header) enforces H2 flood knobs >= 1, h2_stream_shrink_ratio >= 2, lifetime/header caps >= 1, ALPN values in {h2, http/1.1}, and RFC 9110-approximating token grammar on sozu_id_header. Raw protobuf clients cannot bypass via LoadState. Worker plane (lib/src/http.rs, https.rs, tcp.rs, server.rs): - *Listener::update_config + *Proxy::update_listener + Server-level notify_update_* routing. - HttpsListener rebuilds rustls ServerConfig on ALPN patch via the existing create_rustls_context (pure over (&config, resolver)) so the MioTcpListener/token/resolver are preserved. - HttpAnswers::replace_defaults rewrites listener-default templates per-field and leaves cluster_custom_answers untouched, fixing the silent-data-loss blocker that naive HttpAnswers::new swap would cause. - Defense-in-depth: worker-side update_config also runs the pub validators so a raw protobuf client or state replay cannot bypass. - ListenerError::InvalidValue maps StateError::InvalidValue through to the WorkerResponse failure path. CLI + audit (bin/src/cli.rs, ctl/request_builder.rs, command/requests.rs): - Per-protocol Update variants mirroring add/remove/activate/deactivate. - Paired --foo / --no-foo boolean flags via ArgAction::SetTrue + overrides_with; --alpn-protocols / --reset-alpn pair; answer-file paths loaded client-side. - worker_request dispatch + audit_entry_for emit EventKind::LISTENER_ UPDATED with the field-diff string on the audit log only (Event has no free-form field). Tests (e2e/src/tests/listener_update_tests.rs): - 14 e2e tests (5 passing, 9 currently #[ignore] with TODO notes) plus 26 new state.rs unit tests + replace_defaults_preserves_* in answers.rs. - Passing: test_flood_knob_validation, test_http_answers_replace_ preserves_cluster_overrides (the codex HIGH must-pass), test_not_ found, test_disable_http11_toggle, test_disable_http11_inflight_ keepalive. Ignored tests document follow-up needs on the e2e command-channel query path and timing harness. Docs: - doc/configure.md: new "Runtime patch" section with mutability-class table, CVE references, and worked examples. - CHANGELOG.md: Added + Changed entries. Known follow-ups (review notes in the commit, not blockers): - ALPN rebuild master/worker divergence rollback (M-1 codex review). - Full RFC 9110 token tokenizer for sozu_id_header. - Un-ignoring the 9 e2e tests once command-channel query + timing harness land. Verification: - cargo build --locked --all-features: clean - cargo +nightly fmt --all: clean - cargo clippy -p sozu-command-lib -p sozu-lib -p sozu --all-targets --locked: clean - cargo test -p sozu-command-lib --lib: 63/63 - cargo test -p sozu-lib --lib --test-threads=1: 382/382 - cargo test -p sozu --bin sozu: 2/2 - cargo test -p sozu-e2e listener_update --test-threads=1: 5/5 passed, 9 ignored with TODO rationale. Plan: /home/florentin/.claude/plans/i-want-you-to-precious-sun.md Cross-reviewed by codex + /review + /guidelines + /simplify. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 7394d29 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(listener): configurable Sozu-Id correlation header name Adds a new listener knob `sozu_id_header` that lets operators rename the per-request correlation header Sozu injects into every request and response. Default stays `"Sozu-Id"`; common rebrands include `"X-Request-Trace"` or `"X-Edge-Id"`. The value is plumbed end-to-end: * `command/src/command.proto` — new field on both `HttpListenerConfig` (= 30) and `HttpsListenerConfig` (= 42), `optional string` so the default is only applied when absent. * `command/src/config.rs` — matching `Option<String>` on the builder with plumbing through `to_http` / `to_tls`. * `command/src/proto/display.rs` — `sozuctl`-style output shows the knob when set. * `lib/src/lib.rs` — `L7ListenerHandler` trait gains `fn get_sozu_id_header(&self) -> &str` with default `"Sozu-Id"` so call sites that haven't been updated keep the legacy value. * `lib/src/http.rs` + `lib/src/https.rs` — concrete listener implementations honour the config, falling back to the literal `"Sozu-Id"` when the field is `None` or empty. * `lib/src/protocol/kawa_h1/editor.rs` — `HttpContext` gains `sozu_id_header: String`, initialised from the listener at stream creation. Both the request-side and response-side writers use it instead of a hard-coded `kawa::Store::Static(b"Sozu-Id")`. * `lib/src/protocol/mux/mod.rs` — H2 stream creation reads the name from the listener and passes it to `HttpContext::new`. * `lib/src/protocol/kawa_h1/mod.rs` — H1 `Http::new` reads from the listener and passes it to `HttpContext::new` (the listener handle is already in scope, no upstream call-site change). * `doc/configure.md` — documents the knob. Tests: * `test_sozu_id_header_default_name_stored_on_context` — default path. * `test_sozu_id_header_custom_name_stored_on_context` — operator override stored verbatim. Hot-reload semantics: the value is cached on `HttpContext` per connection, so a config change takes effect on NEW connections only. Existing keep-alive connections continue to emit the old header name until they close — consistent with how sozu handles other listener-level config (connect_timeout, sticky_name, etc.) and with typical HTTP-proxy expectations. Addresses issue #1145 (§B3-t in the 2026-04-21 H2 triage plan). Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 62674a0 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(h2): add flood counter for connection-level WINDOW_UPDATE (stream 0) Adds a new per-sliding-window counter in `H2FloodDetector` that tracks non-zero stream-0 WINDOW_UPDATE frames and triggers GOAWAY(ENHANCE_YOUR_CALM) when the configurable threshold is exceeded. Context: the pre-existing detector tracked RST_STREAM, PING, SETTINGS, empty DATA, and CONTINUATION floods plus a generic glitch counter, but non-zero stream-0 WINDOW_UPDATE frames were uncounted. Zero-increment stream-0 WINDOW_UPDATEs already short-circuit into GOAWAY(PROTOCOL_ERROR) per RFC 9113 §6.9, but legal non-zero increments have no per-frame cost limit and a peer could burn proxy CPU by sending millions of them. Changes: * `command/src/command.proto` — new optional field `h2_max_window_update_stream0_per_window` on both `HttpListenerConfig` (= 29) and `HttpsListenerConfig` (= 41). * `command/src/config.rs` — matching `Option<u32>` on `ListenerBuilder` with plumbing through `to_http` / `to_tls`. * `command/src/proto/display.rs` — extends `add_h2_flood_rows` with the new field so `sozuctl`-style output shows it when set. * `lib/src/protocol/mux/h2.rs`: - new constant `DEFAULT_MAX_WINDOW_UPDATE_STREAM0_PER_WINDOW = 100` (mirrors the other per-window defaults). - `H2FloodConfig` gains `max_window_update_stream0_per_window: u32` with matching default, `new()` arg, and `.max(1)` clamp. - `H2FloodDetector` gains `window_update_stream0_count: u32`, initialised to 0, halved in `maybe_reset_window`, and checked in `check_flood` with metric key `h2.flood.violation.window_update_stream0_window`. - `handle_window_update_frame` increments the counter (saturating) on every non-zero stream-0 WINDOW_UPDATE before the arithmetic, and calls `check_flood` so a burst is stopped before we pay the cost. * `lib/src/{http,https}.rs` — wire the listener-config field into `get_h2_flood_config` with defaults fallback. * `doc/configure.md` — document the knob in the flood-thresholds table and the TOML example. Tests: * `test_flood_detector_window_update_stream0_trips_at_threshold` asserts strict greater-than semantics + correct violation metadata. * `test_flood_detector_window_update_stream0_honours_default` asserts the default counter config value matches the documented constant. * `test_flood_detector_half_decay_on_window_expiry` extended to exercise the new counter alongside the existing four. Addresses Codex finding G2 from the 2026-04-21 H2 triage (~/.claude/plans/ask-h2-issues-triage-plan.md §Un-ticketed Gaps Matrix and §Implementation Outcome follow-up list). Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud> Co-authored-by: Eloi Démolis <43861898+Wonshtrum@users.noreply.github.com> Co-authored-by: hcaumeil <78665596+hcaumeil@users.noreply.github.com>
| Commit: | 4eb871f | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(h2): make soft_stop graceful-shutdown deadline configurable Introduces per-listener knob `h2_graceful_shutdown_deadline_seconds` (proto field on `HttpListenerConfig` and `HttpsListenerConfig`, TOML key on `[[listeners]]`). When a worker receives `soft_stop`, the existing `Mux::shutting_down()` now arms a forced-close deadline at the moment `graceful_goaway()` first transitions the connection into draining; once the budget elapses the session is torn down even with Linked/Unlinked streams still in flight. Default: 5 seconds (preserves historic intent). `0` maps to `None`, which disables the forced-close branch entirely and reverts to the old behavior of waiting indefinitely for streams to drain. Wiring: - `command.proto`: `h2_graceful_shutdown_deadline_seconds = 28` on `HttpListenerConfig`, `= 40` on `HttpsListenerConfig`. - `command/src/config.rs`, `proto/display.rs`: ListenerBuilder field + `to_http`/`to_tls` propagation + display row. Mirrors the sibling H2 knob wiring (`h2_max_rst_stream_per_window` template). - `lib/src/lib.rs`: new trait method `L7ListenerHandler::get_h2_graceful_shutdown_deadline` with a `Some(Duration::from_secs(5))` default. - `lib/src/http.rs`, `lib/src/https.rs`: HttpListener / HttpsListener implementations. Value `0` maps to `None`. - `lib/src/protocol/mux/h2.rs`: `H2DrainState` gains `started_at` (armed once by `graceful_goaway`) and `graceful_shutdown_deadline`. Peer-initiated GOAWAYs (via `handle_goaway_frame`) deliberately do NOT arm the timer — the budget applies only to the proxy's own soft-stop. Adds `ConnectionH2::graceful_shutdown_deadline_elapsed`. - `lib/src/protocol/mux/connection.rs`: `new_h2_server` / `new_h2_client` plumb the deadline; enum-level `Connection::graceful_shutdown_deadline_elapsed` (H1 returns `false` — no multiplex to drain). - `lib/src/protocol/mux/mod.rs::Mux::shutting_down`: checks `graceful_shutdown_deadline_elapsed` right after `drive_frontend_shutdown_io` and returns `true` on expiry so the server loop closes the session. - `lib/src/protocol/mux/router.rs`, `lib/src/https.rs`: plumb the deadline through `new_h2_client` / `new_h2_server` call sites. - `lib/src/protocol/mux/LIFECYCLE.md`: documents the armed timer and forced-close branch in §8.3 Session drain. - `doc/configure.md`: adds the knob to the H2 connection tuning table and TOML example. Tests (e2e/src/tests/h2_tests.rs): - `test_h2_graceful_shutdown_timeout_forces_close` — default 5 s budget fires between 3 s and 15 s after soft_stop with a held request. - `test_h2_graceful_shutdown_deadline_configurable_short` (deadline=1 s) — worker stops in under 4 s. - `test_h2_graceful_shutdown_deadline_configurable_long` (deadline=60 s) — no premature close within 10 s, then completes on release. Uses `resolve_request_timeout(60 s)` so hyper does not abort the in-flight stream inside the assertion window. Validation: `cargo build --all-features --locked`, `cargo clippy --all-targets --locked`, `cargo +nightly fmt --all -- --check`, `cargo test --workspace --locked --lib`, focused `cargo test -p sozu-e2e --locked -- --test-threads=1 test_h2_graceful` — all pass. Resolves feat/h2-mux audit Q1 (2026-04-21): the soft_stop deadline is intentional, must remain configurable, default preserved. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud> Co-authored-by: Eloi Démolis <43861898+Wonshtrum@users.noreply.github.com> Co-authored-by: hcaumeil <78665596+hcaumeil@users.noreply.github.com>
| Commit: | 847ed93 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
docs(cluster): clarify http2 flag is a backend-capability hint cluster.http2 = true signals that the BACKEND speaks HTTP/2 (h2c or h2+TLS). It does NOT gate H2 acceptance at the frontend — frontend H2 is negotiated via TLS ALPN on the listener (alpn_protocols) and is fully independent of per-cluster configuration. Per user decision: feat/h2-mux / PR #1209 planning, 2026-04-21 (audit Q8 — clarify http2 field semantics). Files updated: - doc/configure.md: added callout box under "HTTP/2 backend connections" - command/src/command.proto: expanded Cluster.http2 doc comment - command/src/config.rs: expanded ClusterConfig.http2 doc comment - CLAUDE.md: tightened H2 config knobs bullet Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 1e7a05c | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(metrics): metrics.detail cardinality knob (foundation) Foundation for per-listener / per-cluster / per-backend metric label opt-in, mirroring HAProxy's `process|frontend|backend|server` extra-counters knob. Adding labels to the StatsD keyspace under load (e.g. per-listener bytes_in on a host with many listeners) can blow up the keyspace of any statsd aggregator; HAProxy solved this by making the detail level an explicit operator choice. This commit lands the same control surface on Sōzu. Added: - `MetricDetail` proto enum in `command/src/command.proto` (`DETAIL_PROCESS=0 | FRONTEND=1 | CLUSTER=2 | BACKEND=3`, each a superset of the previous). `ServerMetricsConfig` gains an optional `detail` field (tag 4) so workers built before this lands default to DETAIL_CLUSTER on the lib side and preserve historical behaviour. - `MetricDetailLevel` Rust enum in `command/src/config.rs` with `serde(rename_all = "lowercase")` so operators write `detail = "frontend"` in the TOML config. Doc-comment calls out the superset relationship and the HAProxy analogue. Deferred (intentional — scope hard stop per the brief): - Wiring the new macros / labels into the existing `incr!`/`gauge!` call sites. That's >50 call sites and risks either a breaking API change to the macro or a parallel `incr_listener!` family — either shape wants a dedicated follow-up MR so the macro surface stays reviewable. - Touching `lib/src/metrics/network_drain.rs` to actually honour the detail level on the wire. Depends on the macro decision. - The accept-path telemetry (commit 4378101c) already labels by listener address regardless of this knob; once the wiring lands, that path becomes opt-in via DETAIL_FRONTEND and collapses to DETAIL_PROCESS otherwise. Shipping the proto + config enum first means the follow-up MR is a pure labelling change against a stable config shape — operators can already set the knob in their TOML files today; the value is simply ignored by the current drain until wiring lands. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | b9d1616 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(logging): surface TLS metadata and XFF chain on access logs Extend the `ProtobufAccessLog` wire schema with five additional optional fields (tags 25–29) capturing the per-connection TLS handshake metadata and the upstream-attested forwarded chain. Follows the same pattern as `x_request_id` (c9ec90cf): plumb through `mux::Context` → `HttpContext` → `RequestRecord` → proto, populated at every access-log emit site. Fields: - `tls_version` (`&'static str`): short label from `rustls_version_label` (e.g. `TLSv1.3`). New helper alongside the existing metric-prefixed `rustls_version_str` so the log records `TLSv1.3` rather than the dotted metric key. - `tls_cipher` (`&'static str`): short label from `rustls_ciphersuite_label` (e.g. `TLS_AES_128_GCM_SHA256`). Paired with `tls_version` in `lib/src/https.rs::upgrade_handshake`. - `tls_sni` (`&str`): borrowed from `HttpContext.tls_server_name`, same pre-lowercased value the routing layer uses to enforce the SNI ↔ `:authority` binding (CWE-346 / CWE-444). - `tls_alpn` (`&'static str`): on-the-wire ALPN label (`h2`, `http/1.1`) captured alongside the existing `AlpnProtocol` match. - `xff_chain` (`&str`): verbatim `X-Forwarded-For` value snapshotted in `editor.rs::on_request_headers` *before* Sōzu appends its own peer hop — the log records the upstream-attested chain, not the rewritten header Sōzu forwards. TLS fields are connection-scoped (stamped once on `mux::Context` at handshake completion, propagated to every per-stream `HttpContext` via `Context::create_stream`); `HttpContext::reset()` intentionally preserves them across H1 keep-alive so request N+1 still carries the handshake metadata. `xff_chain` is per-request and resets. Coverage across session types: - H1 + H2 mux: populated from the shared `HttpContext`. - WSS post-upgrade pipe: `Pipe` grows `set_tls_metadata`, called from `https.rs::upgrade_mux` so the WebSocket access log inherits the handshake metadata. - Plain TCP / WS / TCP+proxy-protocol: always `None` (no TLS termination on those paths). Also: - `HttpContext` gains corresponding fields + reset-preservation test. - `doc/configure.md` documents the five new access-log fields with their wire tags and source of truth. Signed-off-by: Florentin Dubois <florentin.dubois@clever-cloud.com> Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 4f505c4 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(command): extend EventKind with control-plane mutation variants Foundation for a control-plane audit trail. Today the SubscribeEvents bus carries only four backend-health events (BACKEND_DOWN, BACKEND_UP, NO_AVAILABLE_BACKENDS, REMOVED_BACKEND_HAS_NO_CONNECTIONS). Add 14 new EventKind variants for the mutation surface so subscribers (audit shims, SIEMs, compliance ingestion) can react to cluster / frontend / certificate / listener / configuration / worker / logging changes without polling state diffs. New variants (numeric tags 4..17, existing tags unchanged for wire compatibility): - CLUSTER_ADDED, CLUSTER_REMOVED - FRONTEND_ADDED, FRONTEND_REMOVED - CERTIFICATE_ADDED, CERTIFICATE_REMOVED, CERTIFICATE_REPLACED - LISTENER_ACTIVATED, LISTENER_DEACTIVATED - CONFIGURATION_RELOADED - WORKER_KILLED, WORKER_RELAUNCHED - LOGGING_LEVEL_CHANGED, METRICS_CONFIGURED `Display for Event` (`command/src/proto/display.rs`) gains the matching human-readable strings. `bin/Cargo.toml` enables the `nix` `socket` feature so the follow-up that wires emit sites can use `nix::sys::socket::sockopt::PeerCredentials` to capture the SO_PEERCRED actor UID off the unix socket. The feature flag has no runtime cost when unused. Deferred to follow-up (intentionally not in this commit): - Emit sites in `bin/src/command/{requests,server}.rs`. Each mutating request handler needs to push an `Event` with the matching kind and log a structured audit line. Touches ~10 handlers. - SO_PEERCRED capture at the unix-socket accept site, plumbed through `ClientSession` into the audit log line. - Per-verb `incr!("config.<verb>", ...)` counters (depend on a cardinality decision: per-actor labels need scoping rules first). - doc/configure.md update describing the new event taxonomy. Compliance regimes (PCI-DSS 10.2, ISO 27001 A.8.15, SOC 2) require an immutable audit trail of privileged mutations; this commit is the wire- format prerequisite for that work to land without breaking subscribers. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | cf676e8 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(http): propagate x-request-id and log it on access log Add end-to-end `x-request-id` handling. The header is the de-facto correlation key used by every modern LB (Envoy, HAProxy, most cloud load-balancers) — before this change Sōzu dropped incoming values and never synthesised one, so request flows couldn't be correlated across the proxy. Behaviour: - Incoming request with `x-request-id`: value preserved verbatim in `HttpContext.x_request_id`, forwarded unchanged to the backend, `incr!("http.x_request_id.propagated")`. - Incoming request without `x-request-id`: generate a header value from the request ULID (`self.id`), inject it into the block list, store the value in `HttpContext.x_request_id`, `incr!("http.x_request_id.generated")`. Works for both H1 and H2 because `pkawa.rs` decodes HPACK into a kawa-H1 representation and dispatches into the shared `on_headers` callback in `editor.rs`. Also surfaced on access logs: - `x_request_id: Option<&str>` added to `RequestRecord`. - `optional string x_request_id = 24;` added to `ProtobufAccessLog` (wire-compatible append). - Populated from `HttpContext.x_request_id` on H1 and H2 mux paths; always None on pure-TCP / WebSocket paths. - Reset test updated. doc/configure.md gains a `Request-ID propagation` subsection. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | de505bc | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(mux): propagate per-session ULID through the protocol stack + log-context schema Introduces a stable per-connection identity that survives protocol upgrades (ExpectProxy → TLS handshake → H1/H2) and rewrites the access log-context block so operators can grep a whole TCP/TLS session (`session_id`) independently of a single HTTP exchange (`request_id`). **Log-context schema** (`command/src/logging/access_logs.rs`, `display.rs`): - `LogContext { session_id: Ulid, request_id: Option<Ulid>, cluster_id, backend_id }`. Rendered as `[<session_id> <request_id_or_-> <cluster_id_or_-> <backend_id_or_->]`. - `session_id` is minted once per accepted socket and copied across every protocol state change. - `request_id` becomes optional: present for H1 keep-alive exchanges and per H2 stream, absent for pre-upgrade events that belong to the session but not to any single request. - `RequestRecord` serialisation now emits both ids so downstream access-log consumers can correlate either axis. **`SessionTcpStream` wrapper** (`lib/src/socket.rs`, +287 lines): New `SocketHandler::session_ulid() -> Option<Ulid>` method with a default `None` for raw `mio::TcpStream`. `SessionTcpStream` is a thin `mio::TcpStream` wrapper that carries the owning `Ulid` and returns it from `session_ulid()` — used by every frontend socket path so error logs inside `SocketHandler` implementations can stamp the session prefix without threading it through every call site. **Propagation**: the session ULID flows from the outer `HttpSession` / `HttpsSession` constructors (`lib/src/{http,https}.rs`) through: - `Connection::new_h1_server(session_ulid, socket, …)`, `new_h2_{server,client}` - `ConnectionH1 { …, session_ulid: Ulid }` and the per-stream generation of `request_id` - `Pipe { …, session_id, request_id }` (pipe protocol now tracks both) - `Context::new(session_ulid, pool, listener, …)` on the mux side - ProxyProtocol `expect`/`relay`/`send` handlers pick up the ULID when the upgrade transfers the socket forward. **`log_context!` macros** (`mux/{h1,h2,mod,router}.rs`, `kawa_h1/{mod,editor}.rs`, `pipe.rs`, `rustls.rs`, `tcp.rs`): read the new `LogContext.session_id`/`request_id.unwrap_or(session_id)` fields when formatting the bracketed prefix. No behaviour change for logs that were already session-scoped; adds the second ULID slot everywhere else. **proto** (`command/src/command.proto`): access-log wire format gains a `session_id` field alongside the existing `request_id`. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud> Co-authored-by: Eloi Démolis <43861898+Wonshtrum@users.noreply.github.com> Co-authored-by: hcaumeil <78665596+hcaumeil@users.noreply.github.com>
| Commit: | 77854c1 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(h2): expose h2_max_rst_stream_emitted_lifetime as a listener config knob The MadeYouReset (CVE-2025-8671) mitigation added in 7953696d hard-coded the 500 ceiling as a compile-time `DEFAULT_MAX_RST_STREAM_EMITTED_LIFETIME`. Operators on very busy or very quiet frontends may want to tune it — mirror the plumbing used by the existing `h2_max_rst_stream_lifetime` and `h2_max_rst_stream_abusive_lifetime` knobs: - `command/src/command.proto`: new optional `h2_max_rst_stream_emitted_lifetime` on `HttpListenerConfig` (#27) and `HttpsListenerConfig` (#39); prost-build regenerates the proto module at build time (gitignored). - `command/src/proto/display.rs`: new argument on `add_h2_flood_rows` + the two listener-Display callers, with a dedicated row label. - `command/src/config.rs::ListenerBuilder`: new `Option<u64>` field wired through the `None` default constructor and both `to_http` / `to_https` forwarders. - `lib/src/http.rs` + `lib/src/https.rs` `get_h2_flood_config`: read the listener value, fall back to `defaults.max_rst_stream_emitted_lifetime`. - `doc/configure.md`: renamed the "RST_STREAM lifetime caps" section to cover both directions, added the new row + a MadeYouReset note + updated the TOML example. - `bin/config.toml`: commented block showing all three lifetime caps. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud> Co-authored-by: Eloi Démolis <43861898+Wonshtrum@users.noreply.github.com> Co-authored-by: hcaumeil <78665596+hcaumeil@users.noreply.github.com>
| Commit: | 0c4a739 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
fix(h2): convert per-stream idle timeout from lifetime cap to true idle timeout The h2_stream_idle_timeout_seconds guard (introduced in 77260267) measured time since stream creation, not time since last activity. This caused active uploads exceeding 30s to be RST_STREAM(CANCEL)'d even with data actively flowing — matching a customer report of PHP uploads interrupted at ~30s. Rename stream_opened_at to stream_last_activity_at and refresh the timestamp on each non-empty inbound DATA frame and on HEADERS for existing streams (response headers, trailers). Empty DATA frames (CVE-2019-9518 vector) do NOT reset the timer, preserving the slow-multiplex Slowloris defense. Add two e2e tests: - active upload trickles DATA past timeout → stream survives (200 response) - idle stream with only PINGs → RST_STREAM(CANCEL) after timeout Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud> Co-authored-by: Eloi Démolis <43861898+Wonshtrum@users.noreply.github.com> Co-authored-by: hcaumeil <78665596+hcaumeil@users.noreply.github.com>
| Commit: | 65afe06 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
feat(config): add h2_max_header_table_size listener knob Make the SETTINGS_HEADER_TABLE_SIZE cap configurable per-listener. Wire through config.rs, command.proto, display.rs, http.rs, https.rs, H2FloodConfig, and doc/configure.md. Default: 65536 (64 KB). Also documents previously missing H2 knobs: h2_max_rst_stream_lifetime, h2_max_rst_stream_abusive_lifetime, h2_stream_idle_timeout_seconds, strict_sni_binding, disable_http11. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud>
| Commit: | 8233860 | |
|---|---|---|
| Author: | Florentin Dubois | |
| Committer: | Florentin Dubois | |
fix(h2): add per-stream idle timeout to prevent slow-multiplex Slowloris The connection-level TimeoutContainer resets on every HTTP/2 frame, so a peer that sends periodic DATA/HEADERS across many streams can keep the connection alive indefinitely while pinning up to the listener's configured `h2_max_concurrent_streams` slots — a slow-multiplex Slowloris variant (audit Pass 4 Medium #3 / Pass 3 Low #6). Introduce a per-stream deadline: - Timestamp each stream on open (server `create_stream`, client `start_stream`) via `ConnectionH2::stream_opened_at`. - Clear the entry everywhere the stream leaves the map (`remove_dead_stream`, RST_STREAM handler, GOAWAY retry path, discard on oversized data). - `cancel_timed_out_streams()` scans the open streams on every `readable()` call; any stream older than `stream_idle_timeout` is queued as RST_STREAM(CANCEL) through the existing `pending_rst_streams` path and marked in `rst_sent` to prevent duplicate frames. Expose the knob as `h2_stream_idle_timeout_seconds` on both the HTTP and HTTPS listener configs (proto fields 25 / 37), plumb it through `ListenerBuilder`, `get_h2_stream_idle_timeout()` (default 30 seconds), and `Connection::new_h2_server/client`. Unset behavior matches the previous default-30s ceiling, so existing deployments stay within a sensible safety envelope without any config change. Signed-off-by: Florentin Dubois <florentin.dubois@clever.cloud> Co-authored-by: Eloi Démolis <43861898+Wonshtrum@users.noreply.github.com> Co-authored-by: hcaumeil <78665596+hcaumeil@users.noreply.github.com>