These commits are when the Protocol Buffers files have changed: (only the last 100 relevant commits are shown)
| Commit: | 00a49b0 | |
|---|---|---|
| Author: | Abhinab Saha | |
| Committer: | GitHub | |
[BACKPORT 2026.1][CLOUDGA-37243] DocDB: Use one socket per yb_thin_client connection so a load balancer sends calls to one tserver (#34451) (#34706) ## Summary A yb_thin_client session exists only on the tserver that issued its id. Each connection has its own messenger, and `rpc::Proxy` spreads a messenger's calls over `--num_connections_to_server` sockets per address (8 by default). Behind a load balancer that picks a backend per TCP connection, such as a Kubernetes ClusterIP Service, those sockets reach different tservers. A session opened over one socket then has its Performs sent over others, to tservers that answer `Unknown session`, and paged reads keep failing. Changes: - Each connection's messenger keeps one socket to its host, however many addresses the client is given: any of them may be a load balancer, such as one Service per zone. Separate connections still open their own sockets, so load still spreads. - One socket does not survive a reconnect, and with one address the connection never moves host, so #34374's failover leaves its stale sessions in place. Each connection now has an epoch. A `NetworkError` on a call's own RPC, including `OpenTable`, or a lost-session reply for the session's current incarnation bumps it once, and sessions opened at an older epoch reopen before their next use. Errors the tserver returns about its own work, such as a fenced write or an unreachable tablet leader, bump nothing. - The epoch needs an exact test for a lost session. `FromStatus` searched the status text for "ession" and "nknown" or "xpired", but the text names the source file, so a fenced multi-row write's `Operation expired (yb/tserver/pg_client_session_util.cc:...)` matches it too. ThinClientService now attaches a new `TabletServerErrorPB::SESSION_LOST` code to both statuses for a session it no longer serves: `InvalidArgument "Unknown session N"` and `ShutdownInProgress "Session is shutting down"`. Both the epoch and `FromStatus` match the code. Statuses from an older tserver carry no code, so for those the client matches the two statuses by their exact text. This changes two cases: - A Perform that finds its session just as the session expires gets `Session is shutting down`. It used to return `YBTHIN_OTHER` and leave the session open until a later call got `Unknown session`. It now returns `YBTHIN_NETWORK`, closes the session and bumps the epoch, as `Unknown session` does. - `FromStatus` only applied the text search to `YBTHIN_OTHER` and `YBTHIN_INVALID`, so a multi-row write's non-fence `Expired`, such as a stale retryable request id, was `YBTHIN_NETWORK`. It is now `YBTHIN_OTHER`, like a single-row one. - The `rpc::Proxy` class comment documents the connections-per-server behavior. ## Alternatives considered - Setting `--num_connections_to_server=1`. The library runs inside the caller's process, which never parses YB flags, and the C ABI has no way to set one. It also does nothing about reconnects. - One socket only when the client is given one address, and the default 8 with several. Several addresses do not mean several tservers, the socket count is a messenger setting shared by all of a connection's proxies, and the epoch already treats a connection as one socket. - Reusing `PgsqlError(YB_PG_CONNECTION_DOES_NOT_EXIST)`, which PgClientService's unknown-session status already carries. It would tie the thin client to a PostgreSQL error code. - Matching only the code. A newer client would then report an older tserver's `Unknown session` as `YBTHIN_INVALID`, and leave the session open until the keepalive closed it. - `sessionAffinity: ClientIP` on the Service. kube-proxy would send all of a client's connections to one tserver, so its load would stop spreading. When that tserver restarts, the affinity moves and the stale sessions come back. ## Performance We ran pgbench, the standard Postgres benchmark (8 clients running small bank-style transactions), against an application that stores its data through yb_thin_client. The application reached a 3-tserver YugabyteDB cluster on GKE through one ClusterIP address, the setup this PR fixes. "Before" is the old client with 8 sockets per connection; "after" uses 1. These runs used a 2026.1 release build with only the one-socket change; the other changes only affect failed calls. We tested two data sizes: - Small data, about 150 MB (pgbench scale 10): it fits in the application's memory, so the test mostly exercises writes. - Large data, about 750 MB (pgbench scale 50): it does not fit, so almost every read goes through the thin client. Each version ran on several identical, independent deployments, each with its own cluster, application and benchmark client, and each deployment ran the benchmark 5 times for a minute. The numbers are medians. Runs with `--rate` hold the load at a fixed tps (0.6 x the slowest deployment's); their `latency average` counts from each transaction's scheduled start. | Data | Measure | Before | After | Change | |---|---|---|---|---| | Small | tps | 914 | 922 | +1% | | Small | latency average | 8.7 ms | 8.6 ms | -1% | | Small | latency average, `--rate` | 6.4 ms | 6.6 ms | +3% | | Large | tps | 705 | 739 | +5% | | Large | latency average | 11.3 ms | 10.7 ms | -5% | | Large | latency average, `--rate` | 6.3 ms | 7.4 ms | +18% | Identical deployments differ from each other by 10-20%, so these changes are within noise. With several tserver addresses, connections also drop from 8 sockets to 1. We did not benchmark that setup. A caller that needs more parallelism can open more connections, with a smaller `sessions_per_conn`. The +18% `--rate` latency on large data compares runs at different rates, since the "after" runs used a lower target, against a single usable "before" run. It needs more runs before it can be read as a regression. ## "Unknown session" errors How often the client hit the bug: the "Unknown session" retries the application logged during the benchmark, counted per deployment. | Data | Before | After | |---|---|---| | Small | 13, 3405+ and 3475+ retries in the 3 deployments measured; 4 operations gave up after 10 retries | 0 in all 9 deployments | | Large | 67 to 3623+ retries per deployment across 6 deployments; 40 operations gave up; 4 page reads failed in Postgres | 0 in all 8 deployments | A "+" means at least that many, because the log collector keeps only the last 5000 lines. No benchmark transaction failed in either version, because the application retries failed operations. ## Upgrade/Rollback safety No C ABI or RPC change. The client now uses 1 socket per connection, whatever the address count. The tserver adds a new `TabletServerErrorPB::SESSION_LOST` code to its unknown-session and session-shutting-down statuses, and keeps their text. Clients and tservers can be upgraded in either order: - An older client already decodes the tablet server error category and ignores the new value. It still finds `Unknown session` by its text, and still reports `Session is shutting down` as `YBTHIN_OTHER`, as before. - A newer client against an older tserver gets both statuses without the code. It matches them by their exact text, which older tservers can no longer change, so it handles them as it does with a newer tserver. A test flag that makes the tserver leave the code off covers this. Original commit: 4abbab4be911fd43bb0697d194f6240c0b5c0be0 / #34451 ## Test plan New tests in `yb_thin_client-itest`: - [x] `PgThinClientLoadBalancerTest.ConnectionStaysOnOneTserverBehindALoadBalancer`: through a round-robin TCP forwarder over two tservers, a connection's Performs stay on one tserver. Without the one-socket default, a scan fails with `Unknown session 10`. - [x] `PgThinClientLoadBalancerTest.SessionsReopenTogetherAfterTheSocketDrops`: the forwarder drops its sockets and the reconnect reaches the other tserver; the read session reopens with the write session. Without the epoch, the scan fails with `Unknown session 10`. - [x] `PgThinClientLoadBalancerTest.SessionsReopenTogetherBehindOlderTservers`: the same, with the tservers leaving off `SESSION_LOST`. Without the text fallback, the upsert after the drop returns `YBTHIN_INVALID`. - [x] `PgThinClientTest.FencedBatchKeepsOtherSessionsOpen`: a fenced 2-row upsert returns `YBTHIN_FENCED`, and a scan paged on the other session continues. It fails if the epoch's lost-session check uses the old substring search, which matches the fenced status's source path. - [x] `PgThinClientSessionExpiryTest.ShuttingDownSessionIsLost`: with a short session lifetime, a Perform held past its session's expiry gets `Session is shutting down`. It returns `YBTHIN_NETWORK`, a scan pinned to the connection's other session restarts, and a retried upsert lands. - [x] `PgThinClientSessionExpiryTest.ShuttingDownSessionIsLostOnOlderTservers`: the same, with the tserver leaving off `SESSION_LOST`. - [x] `PgThinClientTest.OneSocketPerConnection`: 1 socket with one address and with two. The previous revision's address-count rule gave the two-address case 8. Also: - [x] The full `yb_thin_client-itest`, including #34374's tests, passed 25/25 on an earlier revision that still had the socket count as a pool option. - [x] In debug, on this branch: `OneSocketPerConnection`, both earlier load-balancer tests, `FencedBatchKeepsOtherSessionsOpen` and `LateFailureKeepsTheReopenedSession` passed before the switch from the PgsqlError code to `SESSION_LOST`. `ShuttingDownSessionIsLost`, `ShuttingDownSessionIsLostOnOlderTservers` and `SessionsReopenTogetherBehindOlderTservers` passed after it. - [ ] The `FromStatus` change for a multi-row non-fence `Expired` has no test: nothing in the tests can force one. - [x] On a local cluster behind one load-balanced address, the client with the one-socket change recovers from a dead tserver in 28-56 s; without it, it does not recover within 300 s. <!-- Reviewable:start --> - - - This change is [<img src="https://reviewable.io/review_button.svg" height="34" align="absmiddle" alt="Reviewable"/>](https://reviewable.io/reviews/yugabyte/yugabyte-db/34706) <!-- Reviewable:end -->
| Commit: | 4abbab4 | |
|---|---|---|
| Author: | Abhinab Saha | |
| Committer: | GitHub | |
[CLOUDGA-37243] DocDB: Use one socket per yb_thin_client connection so a load balancer sends its calls to one tserver (#34451) ## Summary A yb_thin_client session exists only on the tserver that issued its id. Each connection has its own messenger, and `rpc::Proxy` spreads a messenger's calls over `--num_connections_to_server` sockets per address (8 by default). Behind a load balancer that picks a backend per TCP connection, such as a Kubernetes ClusterIP Service, those sockets reach different tservers. A session opened over one socket then has its Performs sent over others, to tservers that answer `Unknown session`, and paged reads keep failing. Changes: - Each connection's messenger keeps one socket to its host, however many addresses the client is given: any of them may be a load balancer, such as one Service per zone. Separate connections still open their own sockets, so load still spreads. - One socket does not survive a reconnect, and with one address the connection never moves host, so #34374's failover leaves its stale sessions in place. Each connection now has an epoch. A `NetworkError` on a call's own RPC, including `OpenTable`, or a lost-session reply for the session's current incarnation bumps it once, and sessions opened at an older epoch reopen before their next use. Errors the tserver returns about its own work, such as a fenced write or an unreachable tablet leader, bump nothing. - The epoch needs an exact test for a lost session. `FromStatus` searched the status text for "ession" and "nknown" or "xpired", but the text names the source file, so a fenced multi-row write's `Operation expired (yb/tserver/pg_client_session_util.cc:...)` matches it too. ThinClientService now attaches a new `TabletServerErrorPB::SESSION_LOST` code to both statuses for a session it no longer serves: `InvalidArgument "Unknown session N"` and `ShutdownInProgress "Session is shutting down"`. Both the epoch and `FromStatus` match the code. Statuses from an older tserver carry no code, so for those the client matches the two statuses by their exact text. This changes two cases: - A Perform that finds its session just as the session expires gets `Session is shutting down`. It used to return `YBTHIN_OTHER` and leave the session open until a later call got `Unknown session`. It now returns `YBTHIN_NETWORK`, closes the session and bumps the epoch, as `Unknown session` does. - `FromStatus` only applied the text search to `YBTHIN_OTHER` and `YBTHIN_INVALID`, so a multi-row write's non-fence `Expired`, such as a stale retryable request id, was `YBTHIN_NETWORK`. It is now `YBTHIN_OTHER`, like a single-row one. - The `rpc::Proxy` class comment documents the connections-per-server behavior. ## Alternatives considered - Setting `--num_connections_to_server=1`. The library runs inside the caller's process, which never parses YB flags, and the C ABI has no way to set one. It also does nothing about reconnects. - One socket only when the client is given one address, and the default 8 with several. Several addresses do not mean several tservers, the socket count is a messenger setting shared by all of a connection's proxies, and the epoch already treats a connection as one socket. - Reusing `PgsqlError(YB_PG_CONNECTION_DOES_NOT_EXIST)`, which PgClientService's unknown-session status already carries. It would tie the thin client to a PostgreSQL error code. - Matching only the code. A newer client would then report an older tserver's `Unknown session` as `YBTHIN_INVALID`, and leave the session open until the keepalive closed it. - `sessionAffinity: ClientIP` on the Service. kube-proxy would send all of a client's connections to one tserver, so its load would stop spreading. When that tserver restarts, the affinity moves and the stale sessions come back. ## Performance We ran pgbench, the standard Postgres benchmark (8 clients running small bank-style transactions), against an application that stores its data through yb_thin_client. The application reached a 3-tserver YugabyteDB cluster on GKE through one ClusterIP address, the setup this PR fixes. "Before" is the old client with 8 sockets per connection; "after" uses 1. These runs used a 2026.1 release build with only the one-socket change; the other changes only affect failed calls. We tested two data sizes: - Small data, about 150 MB (pgbench scale 10): it fits in the application's memory, so the test mostly exercises writes. - Large data, about 750 MB (pgbench scale 50): it does not fit, so almost every read goes through the thin client. Each version ran on several identical, independent deployments, each with its own cluster, application and benchmark client, and each deployment ran the benchmark 5 times for a minute. The numbers are medians. Runs with `--rate` hold the load at a fixed tps (0.6 x the slowest deployment's); their `latency average` counts from each transaction's scheduled start. | Data | Measure | Before | After | Change | |---|---|---|---|---| | Small | tps | 914 | 922 | +1% | | Small | latency average | 8.7 ms | 8.6 ms | -1% | | Small | latency average, `--rate` | 6.4 ms | 6.6 ms | +3% | | Large | tps | 705 | 739 | +5% | | Large | latency average | 11.3 ms | 10.7 ms | -5% | | Large | latency average, `--rate` | 6.3 ms | 7.4 ms | +18% | Identical deployments differ from each other by 10-20%, so these changes are within noise. With several tserver addresses, connections also drop from 8 sockets to 1. We did not benchmark that setup. A caller that needs more parallelism can open more connections, with a smaller `sessions_per_conn`. The +18% `--rate` latency on large data compares runs at different rates, since the "after" runs used a lower target, against a single usable "before" run. It needs more runs before it can be read as a regression. ## "Unknown session" errors How often the client hit the bug: the "Unknown session" retries the application logged during the benchmark, counted per deployment. | Data | Before | After | |---|---|---| | Small | 13, 3405+ and 3475+ retries in the 3 deployments measured; 4 operations gave up after 10 retries | 0 in all 9 deployments | | Large | 67 to 3623+ retries per deployment across 6 deployments; 40 operations gave up; 4 page reads failed in Postgres | 0 in all 8 deployments | A "+" means at least that many, because the log collector keeps only the last 5000 lines. No benchmark transaction failed in either version, because the application retries failed operations. ## Upgrade/Rollback safety No C ABI or RPC change. The client now uses 1 socket per connection, whatever the address count. The tserver adds a new `TabletServerErrorPB::SESSION_LOST` code to its unknown-session and session-shutting-down statuses, and keeps their text. Clients and tservers can be upgraded in either order: - An older client already decodes the tablet server error category and ignores the new value. It still finds `Unknown session` by its text, and still reports `Session is shutting down` as `YBTHIN_OTHER`, as before. - A newer client against an older tserver gets both statuses without the code. It matches them by their exact text, which older tservers can no longer change, so it handles them as it does with a newer tserver. A test flag that makes the tserver leave the code off covers this. ## Test plan New tests in `yb_thin_client-itest`: - [x] `PgThinClientLoadBalancerTest.ConnectionStaysOnOneTserverBehindALoadBalancer`: through a round-robin TCP forwarder over two tservers, a connection's Performs stay on one tserver. Without the one-socket default, a scan fails with `Unknown session 10`. - [x] `PgThinClientLoadBalancerTest.SessionsReopenTogetherAfterTheSocketDrops`: the forwarder drops its sockets and the reconnect reaches the other tserver; the read session reopens with the write session. Without the epoch, the scan fails with `Unknown session 10`. - [x] `PgThinClientLoadBalancerTest.SessionsReopenTogetherBehindOlderTservers`: the same, with the tservers leaving off `SESSION_LOST`. Without the text fallback, the upsert after the drop returns `YBTHIN_INVALID`. - [x] `PgThinClientTest.FencedBatchKeepsOtherSessionsOpen`: a fenced 2-row upsert returns `YBTHIN_FENCED`, and a scan paged on the other session continues. It fails if the epoch's lost-session check uses the old substring search, which matches the fenced status's source path. - [x] `PgThinClientSessionExpiryTest.ShuttingDownSessionIsLost`: with a short session lifetime, a Perform held past its session's expiry gets `Session is shutting down`. It returns `YBTHIN_NETWORK`, a scan pinned to the connection's other session restarts, and a retried upsert lands. - [x] `PgThinClientSessionExpiryTest.ShuttingDownSessionIsLostOnOlderTservers`: the same, with the tserver leaving off `SESSION_LOST`. - [x] `PgThinClientTest.OneSocketPerConnection`: 1 socket with one address and with two. The previous revision's address-count rule gave the two-address case 8. Also: - [x] The full `yb_thin_client-itest`, including #34374's tests, passed 25/25 on an earlier revision that still had the socket count as a pool option. - [x] In debug, on this branch: `OneSocketPerConnection`, both earlier load-balancer tests, `FencedBatchKeepsOtherSessionsOpen` and `LateFailureKeepsTheReopenedSession` passed before the switch from the PgsqlError code to `SESSION_LOST`. `ShuttingDownSessionIsLost`, `ShuttingDownSessionIsLostOnOlderTservers` and `SessionsReopenTogetherBehindOlderTservers` passed after it. - [ ] The `FromStatus` change for a multi-row non-fence `Expired` has no test: nothing in the tests can force one. - [x] On a local cluster behind one load-balanced address, the client with the one-socket change recovers from a dead tserver in 28-56 s; without it, it does not recover within 300 s. <!-- Reviewable:start --> - - - This change is [<img src="https://reviewable.io/review_button.svg" height="34" align="absmiddle" alt="Reviewable"/>](https://reviewable.io/reviews/yugabyte/yugabyte-db/34451) <!-- Reviewable:end -->
The documentation is generated from this commit.
| Commit: | 2c44e0b | |
|---|---|---|
| Author: | Mark Lillibridge | |
| Committer: | Mark Lillibridge | |
[#34578] DocDB: Record the backend's connected database from an explicit Perform option Summary: PG client service records the database a YSQL backend is connected to (`PgClientSession::database_oid_`, used for the session's QoS cgroup, its `YBSession` pool tag, and history-retention pin attribution) from the `namespace_id` of the first Perform that carries one. pggate sets `namespace_id` only when all of a Perform's relations belong to one database other than template1, so a backend can issue many requests and create transactions before its database is recorded, e.g., a separate DDL transaction created by a Perform that reads only shared catalogs. The upcoming xCluster StopPersisting/StartPersisting work stamps each transaction with its database when the transaction is created, so it needs the database recorded first. Postgres now hands `MyDatabaseId` to pggate as soon as `InitPostgres` sets it (next to `YBCSetupPgBackendCgroup`). pggate sends it in a new `PgPerformOptionsPB.connected_database_oid` on every Perform, legacy catalog session reads included, and on the DDL RPCs. PG client service records the database from that field (`EnsureSessionDatabase`) instead of from `namespace_id`: in `DoPerform` before the response-cache lookup, and in `SetupSession` / `SetupSessionForDdl`. Template1 backends stay unrecorded, as before. The recorded value is unchanged; only when it is recorded changes, so the cgroup and retention-pin consumers see the same database, earlier. **Upgrade / rollback safety** `connected_database_oid` is a new field on `PgPerformOptionsPB` (pg_client.proto), on the pggate <-> PG client service protocol only. A backend and its TServer always run the same version, and a missing field reads as 0 (not yet known). No new gflags. **Update (v2):** recording the database and moving the session into the database's cgroup are now separate. `EnsureSessionDatabase` only records, from any request that carries `connected_database_oid`; the cgroup move happens only in `DoPerform`, once per session (`moved_to_database_cgroup_`), and only for a Perform that arrived through the shared memory exchange (the only `DoPerform` caller without an `RpcContext`), since that is the one path that runs on the session's own exchange thread. Previously the move rode on the first recording request, which v1 had widened to DDL and snapshot read-time RPCs that run on RPC worker threads shared by every session; a worker thread could then end up in a database's cgroup while the exchange thread never did. Behavior change for QoS: with `pg_client_use_shared_memory=false` no cgroup move happens at all, where before the first Perform moved whichever worker thread it ran on. Test Plan: Both ran clean in release and fastdebug. ``` ./yb_build.sh release --cxx-test pgwrapper_pg_mini-test --gtest_filter \ PgMiniTest.SessionRecordsConnectedDatabase ``` New. Fails with the old `namespace_id`-based recording: the fresh backend's session has recorded no database. ``` ./yb_build.sh release --cxx-test pgwrapper_pg_db_pin_tracker-test ``` Includes the new `Template1BackendDoesNotPin`. UPDATE: ``` ./yb_build.sh release --cxx-test pgwrapper_pg_cgroups-test ``` Existing; to be run on Jenkins, since the local machine does not delegate the CPU cgroup controller (every test in the file fails in `SetUp` with "CPU controller not delegated"). The relevant test is `PgCgroupsTest.TestQosBackend`, which connects and checks that a `shmem_exchange_` thread of the new backend's session is in the database's cgroup, i.e., that the move now done only from `DoPerform` for shared-memory Performs still lands on the session's exchange thread. The remaining tests in the file check per-request pool tagging of RPC worker, wait-queue and callback threads, which this diff does not touch. Reviewers: esheng Reviewed By: esheng Subscribers: kfranz, ybase, yql Differential Revision: https://phorge.dev.yugabyte.com/D58942
| Commit: | 9b83021 | |
|---|---|---|
| Author: | Sanketh I | |
| Committer: | Sanketh I | |
[BACKPORT 2025.2][#31589] DocDB: Serve YSQL lease RPCs on the master high-priority thread pool Summary: YSQL lease heartbeats from tservers to the master were handled by the regular-priority RPC thread pool via the MasterDdl service. Under an overload of regular-priority master RPCs (e.g. catalog cache reads or heartbeats or GetTableSchema or GetTabletLocations) the lease refresh could be backed up, and a lease failure kills all PG sessions on that tserver. This moves RefreshYsqlLease and RelinquishYsqlLease into a dedicated MasterYsqlLease service and registers it on the master's high-priority RPC thread pool (rpc::ServicePriority::kHigh), isolating lease refreshes from regular master RPC load. (Consensus is the other user of the high-priority pool.) **Upgrade/Rollback safety:** The new service keeps custom_service_name = "yb.master.MasterService" and the same method names, so the on-the-wire remote method is identical to when these RPCs lived in MasterDdl. This makes the change compatible with any master/tserver upgrade order: an old tserver's MasterDdlProxy and a new tserver's MasterYsqlLeaseProxy produce byte-identical requests, resolved by whichever service the master has registered. The only visible change is the RPC's metrics group name (yb.master.MasterDdl -> yb.master.MasterYsqlLease). Original commit: d212939a4eee2554679515671f1aef4381dde1c8 / D56573 Cherry-picked from the 2026.1 backport 9dc7699ddbf033aa966679ad638ac1841e7e6496 / D56921 ## Merge conflicts One conflict, context drift only. No logic from this change was altered. - src/yb/tserver/ysql_lease_poller.h:15 - the commit replaces the master_ddl.fwd.h include with master_ysql_lease.fwd.h, because RefreshYsqlLeaseInfoPB (the only master:: type this header names) moved to the new proto. On 2025.2 the include block also has a preceding #include <future>, which 2026.1 does not have; that line is pre-existing 2025.2 context and is kept, as std::future is used by RelinquishLease() at line 39. Result: <future> kept, master_ddl.fwd.h replaced by master_ysql_lease.fwd.h. The other 23 files applied with hunks byte-identical to the 2026.1 commit. Verified after resolution: the five lease messages moved out of master_ddl.proto are byte-identical to the ones added in master_ysql_lease.proto, so the wire format is unchanged; all four imports of the new proto exist on 2025.2; RpcServerBase::RegisterService accepts the rpc::ServicePriority argument (src/yb/server/server_base.h:126) and DEFINE_NON_RUNTIME_int32 exists (src/yb/util/flags/flag_tags.h:353) on 2025.2; catalog_manager.{h,cc}, ts_descriptor.h and tablet_server.h still name the lease types and are in the same include state as on 2026.1, so no extra includes are needed. Test Plan: build-support/lint.sh --rev origin/2025.2 reports 0 errors, 0 warnings. Existing tests exercise the lease RPCs through the new service/proxy/client: ./yb_build.sh fastdebug --clang21 --cxx-test object_lock-test ./yb_build.sh fastdebug --clang21 --cxx-test master-test Reviewers: #db-approvers, bkolagani, zdrudi Reviewed By: zdrudi Subscribers: ybase Differential Revision: https://phorge.dev.yugabyte.com/D57299
| Commit: | 25e71bd | |
|---|---|---|
| Author: | Mark Lillibridge | |
| Committer: | Mark Lillibridge | |
[#33941] xCluster: Add PropagateXClusterGuardedInfo Summary: ## Overview This is diff 3 of 3 to implement the PropagateXClusterGuardedInfo primitive; it adds the primitive itself. The first diff grouped the xCluster-guarded information into a single versioned protobuf carried by heartbeat responses and applied on TServers only if newer; the second added the xCluster-guarded lease that lets master know when a TServer can no longer be giving out that information. The primitive: ``` ~/code/yugabyte-db/src/yb/master/xcluster/xcluster_manager.h:81: // Ensures that on successful return no TServer will ever give out xCluster-guarded information // less recent than master had when this was called. For example, after SetXClusterRole followed // by a successful call to this method, no TServer will ever again give out the previous role. // // TServers that have not heartbeated for longer than the xCluster lease duration do not slow this // down; those unresponsive for less than that can make it fail by timing out. This method may // also fail if master leadership changes during the call; callers wanting to survive that should // retry. Safe to call concurrently. Status PropagateXClusterGuardedInfo(MonoTime deadline) ``` ## How it works Master makes a copy of the xCluster-guarded information the same way it does for a heartbeat response (under the version mutex, bumping the version count) and pushes it to TServers via a new TServer admin RPC: ``` ~/code/yugabyte-db/src/yb/tserver/tserver_admin.proto:513: // Applies the xCluster-guarded information unless the TServer already holds a copy with an equal // or newer version; succeeds either way. Master uses this to push information to TServers // without waiting for their next heartbeat. rpc PropagateXClusterGuardedInfo(PropagateXClusterGuardedInfoRequestPB) returns (PropagateXClusterGuardedInfoResponsePB); ``` The TServer handler calls the same apply-if-newer routine the heartbeater uses, so the version ordering established in the first diff resolves any race between the two paths. The RPCs are sent in parallel using the master async task framework (one `AsyncPropagateXClusterGuardedInfo` task per TServer), retrying until the caller's deadline. Only TServers that may hold a lease are sent to: a TServer master knows lacks a lease must acquire information more recent than the copy before it can get a new lease, so it can be skipped. A TServer that loses its lease while we are trying is likewise excused from having to answer. The call fails if any other TServer cannot be reached by the deadline. The primitive requires the TServer registry to be persisted (otherwise master cannot enumerate every TServer that might hold a lease) and does not check master leadership; per the specification, a leadership change during the call is the caller's problem to retry. No callers are added in this diff; they will come with the automatic mode switchover/setup and import snapshot work. **Upgrade/Rollback safety:** Mid-upgrade, some TServers do not implement the new RPC. Rather than introduce another auto flag, the primitive is gated on the existing `enforce_xcluster_guarded_lease` auto flag from the previous diff: while it is off, `PropagateXClusterGuardedInfo` returns OK immediately without sending anything. This is safe because nothing depends on the information having propagated until leases are enforced, and enforcement only begins once the upgrade is finalized, at which point every TServer implements the RPC. Nothing else in this diff changes wire behavior: the new RPC is only ever sent when the flag is on, and old masters never send it. Rollback before finalization therefore needs no special handling. Fixes #33941 Test Plan: Added two tests to XClusterPropagateGuardedInfoTest. This test stops heartbeats, changes the xCluster role and OID cache invalidation count on master, and verifies that neither reaches any TServer until PropagateXClusterGuardedInfo is called and that both reach every TServer once it is. ``` ybd release --cxx-test xcluster_guarded_lease-test --gtest_filter '*.PropagatesWithoutHeartbeats' ``` This test shuts down a TServer and verifies that `PropagateXClusterGuardedInfo` fails by its deadline naming that TServer while master still considers it as possibly holding a lease; succeeds when the TServer loses its lease during the call; and, once master has marked it DEFINITELY_NO_LEASE, skips it from the start (succeeding with a deadline that was too short before): ``` ybd release --cxx-test xcluster_guarded_lease-test --gtest_filter '*.PropagateEvenWithDeadTServer' ``` Reviewers: xCluster, hsunder, zdrudi Reviewed By: zdrudi Subscribers: svc_phabricator, ybase Differential Revision: https://phorge.dev.yugabyte.com/D58157
| Commit: | 42e62c1 | |
|---|---|---|
| Author: | Mark Lillibridge | |
| Committer: | Mark Lillibridge | |
[#30820] xCluster: Create xCluster role lease Summary: As part of various operations like xCluster automatic-mode switchover, we need to be able to change the xCluster role and ensure it is propagated so all TServers either have the new role or know they don't have the role and will get a sufficiently up-to-date version of the role before using it. This diff is to allow handling the second case; the function to ensure propagation (PropagateXClusterGuardedInfo) will be in the next diff. In particular, we need to ensure that * TServers not heart beating for long enough will eventually return UNAVAILABLE for the xCluster role * Master is able to determine when enough time has passed without heartbeats that a TServer is guaranteed to be returning UNAVAILABLE * when such a TServer starts heart beating again, it starts returning role information at least as current as the new heartbeat response it got Here we do this by creating a new xCluster-guarded information lease granted by master via heartbeat responses. The TServer will return role UNAVAILABLE if it does not have such a current lease. A master gflag controls the duration of this lease: ``` ~/code/yugabyte-db/src/yb/master/master_heartbeat_service.cc:78: DEFINE_NON_RUNTIME_uint32(xcluster_guarded_lease_duration_ms, 2 * 60 * 1000, "Duration of xCluster-guarded information lease in milliseconds; not safe to lower."); ``` Master offers this lease by including a lease expiration duration in its heartbeat response: ``` ~/code/yugabyte-db/src/yb/master/master_heartbeat.proto:378: // If present, this heartbeat response grants the TServer a xCluster-guarded information lease // of this duration from when the TServer sent the heartbeat request. optional uint32 xcluster_guarded_lease_duration_ms = 36; ``` I am sending the lease duration so all the calculations using lease duration will occur on master; I hope this will make changing the duration flag value easier if that ever becomes necessary. (Raising it is safe; lowering it is not, since a master running with the lower value would consider leases granted under the old, longer duration expired while TServers still honor them.) To avoid the need for wall time, the TServer computes its lease expiration time independently. In particular, it records a MonoTime just before sending the heartbeat request and then adds the returned duration to that time to figure out the lease expiration time. (This avoids the need for master to take network transit time into account when estimating when the TServer will expire.) Each TServer remembers the last lease expiration time (a local MonoTime) it has seen. If now is beyond that (or no lease expiration time has been seen) then it considers itself as having no lease and returns UNAVAILABLE for the xCluster role when asked. The TServer-side checks that forbid DDLs and sequence-manipulation functions on an automatic-mode target treat role UNAVAILABLE the same as AUTOMATIC_TARGET (with a distinct error message saying the role is currently unavailable), so losing the lease fails closed rather than letting those operations through. This matches the DDL replication extension, which already errors out on UNAVAILABLE. xcluster_target_manual_override still bypasses both checks. We want master to (effectively) track persistently the latest lease it has given to each TServer so it can determine when that TServer must no longer have a lease. Here we are going to do something similar to the already persisted TServer responsiveness state: ``` message SysTabletServerEntryPB { ... enum State { ... // A heartbeat from the TServer has been ACK'd recently. LIVE = 1; // No heartbeat from the TServer has been ACK'd within at least tserver_unresponsive_timeout_ms. UNRESPONSIVE = 2; ``` tserver_unresponsive_timeout_ms might be smaller than our desired xCluster-guarded information lease duration and in general we want to avoid coupling that duration to the xCluster lease duration so we are going to add a parallel field just for the xCluster-guarded information lease: ``` enum XClusterGuardedLeaseState { // Proto best practices are to include a sentinel as the first enum value. // See https://protobuf.dev/programming-guides/dos-donts/#unspecified-enum XClusterGuardedLeaseState_UNSPECIFIED = 0; // It is possible that this TServer has a xCluster-guarded information lease. MAYBE_HAS_LEASE = 1; // It is definitely not possible that this TServer has a xCluster-guarded information lease. // // Moreover if this TServer does require such a lease in the future, the associated information // (e.g., the xCluster role) from master will date at least from the point at which this state // was changed to MAYBE_HAS_LEASE. DEFINITELY_NO_LEASE = 2; } ... optional XClusterGuardedLeaseState xcluster_guarded_lease_state = 11; ``` This allows us to create the following function: ``` ~/code/yugabyte-db/src/yb/master/ts_descriptor.h:105: // If this returns false then we can be sure the TServer does not currently have a // xCluster-guarded information lease. False positives (i.e., returning true when it does not // have a lease) are possible. bool MaybeHasXClusterGuardedLease() const; ``` The remaining problem is to ensure when the TServer re-acquires the lease, the associated information (e.g., the role) is at least as current as the information in the heartbeat response conveying the new lease. (This allows master to set the role, determine that the TServer does not currently have a lease, then move on confident that when the TServer does get a new lease, it will have information at least as current as the new role.) We do this similarly to how memory barriers work: * on master, in the middle of computing the heartbeat response we save the start time for the next lease as now * we then add the lease time and current information to the heartbeat response * on TServer, we copy the heartbeat response information into place * we update the lease time Thus, if the master sees that MaybeHasXClusterGuardedLease() is false then it can be sure that any future lease that TServer gets must be based on a heartbeat response whose information was copied after the function was called. Possibly things are clearer with a timeline of a heartbeat: ``` [TServer] just before sending new heartbeat [TServer] record now as potential lease start time (T_t) [TServer] send new heartbeat [master] begin preparing heartbeat response [master] save now as potential lease start time (T_m) [master] move xcluster_guarded_lease_state to MAYBE_HAS_LEASE [master] copy data covered by lease (e.g., OID invalidation) into heartbeat [master] send response [TServer] save covered data by lease from response [TServer] update lease expiration based on duration and T_t ``` If MaybeHasXClusterGuardedLease() returns false then xcluster_guarded_lease_state was (at least during the call) DEFINITELY_NO_LEASE, which means all T_m's are at least duration + drift slack before now. Since T_t < T_m and the two monotonic clocks can have drifted apart by at most the drift slack over that interval, all T_t's are at least duration before (now on the TServer) => the TServer does not currently have a valid lease and all heartbeats that have already reached the MAYBE_HAS_LEASE step cannot confer a valid lease. Thus the first heartbeat to give such a TServer a new lease must have copied data after now and moreover that data will be available on the TServer before the lease is extended. The above assumes master knows about every heartbeat the TServer has sent. After a master leadership change it does not: the previous leader may have granted a lease the new leader never saw. A master only grants leases while it holds the Raft leader lease, so the previous leader cannot have granted one starting later than when the new master acquires its lease. The new leader therefore assumes every TServer heartbeated when it got its Raft leader lease. Concretely, the new leader loads its TServer descriptors from the persisted TServer registry after acquiring its lease and treats the load time as each one's last heartbeat. This relies on `persist_tserver_registry`; without it the new leader does not know about TServers that heartbeated only the previous leader, so PropagateXClusterGuardedInfo (next diff) will refuse to run when that flag is off. A TServer that master first learns of from a Raft config (`RegisterTsFromRaftConfig`) rather than a heartbeat starts out DEFINITELY_NO_LEASE, just as it starts out UNRESPONSIVE: it has not heartbeated any master leader whose registry was loaded, so it cannot hold a lease. `RemoveTabletServer` (yb-admin `remove_tablet_server`) now additionally requires the TServer to be in DEFINITELY_NO_LEASE. A removed TServer is invisible to the forthcoming propagation primitive, so it must not be able to hold a lease; previously a TServer could be removed after only tserver_unresponsive_timeout_ms (60 s) of silence, less than the xCluster-guarded lease duration. Also made some of the usual tidying fixes while at it. **Upgrade/Rollback safety:** We guard the enforcement of the lease via an auto flag: ``` ~/code/yugabyte-db/src/yb/server/server_common_flags.cc:72: DEFINE_RUNTIME_AUTO_bool(enforce_xcluster_guarded_lease, kLocalPersisted, false, true, "Should the xCluster-guarded information lease be enforced? If so, the xCluster role for a " "namespace will be reported as UNAVAILABLE if the TServer does not have a current lease."); ``` If this flag is false, the code to determine the xCluster role ignores the lease and falls back on previous behavior (i.e., return UNAVAILABLE only if we have never received a heartbeat). All the other new code is on regardless of the auto flag. That is, masters with new code will be sending leases and adding the field xcluster_guarded_lease_state. TServers with new code will be saving the lease expiration information if it is provided (it's an optional field). Masters with old code will not provide the lease and ignore xcluster_guarded_lease_state field if already present. Old TServers will ignore the lease expiration field if present. The consequences of this is that mid upgrade, the master leader will be providing the lease information and updating/creating the xcluster_guarded_lease_state field as needed. Likewise the TServers will be receiving and upgrading the lease information. However, the get role code will be unchanged, behaving as before the upgrade. Once the upgrade is finalized, enforcement of the lease will begin, with the role returning UNAVAILABLE when a lease is unavailable. Note that universes should not be finalized while nodes are unresponsive so switching the gflag should cause no problems. UPDATE: * note that heartbeat responses requesting registration do not include leases or xCluster-guarded information * the TServer keeps its current copy and lease (if any) when it receives such a response, so it never runs on empty information; normally this only occurs briefly at TServer startup anyway * the reason for persisting the lease tracking information is that if we do not, access to xCluster roles will be blocked for the lease duration on master leadership changes * e.g., xCluster switch over and clone (which is going to depend indirectly on the role) would not be possible for two minutes after a master leadership change * minor behavior change: there is a corner case where the master can never construct a valid heartbeat response because of unable to decrypt or parse universe key registry; this used to not result in the TServer being marked LIVE. After this diff it does. * I think this is acceptable and in general LIVE does not promise but the TServer is fully working, just that it is sending heartbeats to master. UPDATE: One case is knowingly not handled. When a TServer registers at the same host and port as an existing one with a higher instance sequence number, master marks the old descriptor REPLACED, and PropagateXClusterGuardedInfo does not consider REPLACED descriptors even if they are still MAYBE_HAS_LEASE. The RPC could not be sent to such a TServer anyway: its address now belongs to the replacement, which would apply the copy and answer OK on the old TServer's behalf. Normally replacement means the old TServer process is gone (e.g., restarted with a new UUID), so no lease exists to worry about. The gap only matters if two live TServers advertise the same address, which is a misconfiguration. If that becomes a concern, master could instead wait for such descriptors to reach DEFINITELY_NO_LEASE, which the background liveness task brings about one lease duration after their last heartbeat. Fixes #30820 Test Plan: I have added the following new tests (plus one existing test updated, see the end): This integration test verifies that MaybeHasXClusterGuardedLease works correctly: ``` ybd release --cxx-test master_heartbeat-itest --gtest_filter '*.MaybeHasXClusterGuardedLease' ``` This end to end test verifies that the lease running out actually causes the role to switch to UNAVAILABLE and that the role switches back when the lease is reacquired. ``` ybd release --cxx-test xcluster_ddl_replication-test --gtest_filter '*.GuardedLeaseExpiration' ``` RemoveTabletServer now also requires the TServer to be in DEFINITELY_NO_LEASE. This test verifies that removing a TServer that has been dead only briefly fails for that reason: ``` ybd release --cxx-test master_cluster_test --gtest_filter 'RemoveTabletServerTest.MayStillHoldLease' ``` And this existing test verifies that TServer removal still succeeds once the lease has expired (it now waits for DEFINITELY_NO_LEASE first): ``` ybd release --cxx-test master_cluster_test --gtest_filter 'RemoveTabletServerTest.HappyPath' ``` These new tests verify that a TServer without a lease (which reports the xCluster role as UNAVAILABLE) blocks sequence bumps and DDLs on a plain database, and that both work again once the lease is reacquired. ``` ybd release --cxx-test xcluster_guarded_lease-test --gtest_filter 'XClusterGuardedLeaseTest.*' ``` Reviewers: xCluster, hsunder, jhe, zdrudi Reviewed By: zdrudi Subscribers: yql, asrivastava, sergei, ybase Differential Revision: https://phorge.dev.yugabyte.com/D51544
| Commit: | 7a751f8 | |
|---|---|---|
| Author: | Minghui Yang | |
| Committer: | Minghui Yang | |
[BACKPORT 2026.1][#33572] YSQL: Measure catalog version staleness by master read time Summary: A stress test crashed a tserver on the staleness check added for this issue: ``` F0921 15:33:42 ../../src/yb/tserver/tablet_server.cc:1694] Ignoring ysql db 16642 catalog version update: new version too old. New: 19636, Old: 19637, stale_for: 36.802s, threshold: 36.000s ``` This one was a false positive. `pg_yb_catalog_version` did hold 19637: another tserver received it by heartbeat 16 seconds after this FATAL, and the restarted tserver was seeded with it. What was behind was the master's heartbeat cache, frozen for ~88 seconds because both of its writers failed together: ``` W0921 15:32:30.068879 1953073 object_lock_info_manager.cc:1107] Couldn't populate catalog version on exclusive lock release: Timed out (yb/rpc/outbound_call.cc:790): GetTransactionStatus RPC (request call id 7807806) to 172.151.28.2:9100 timed out after 4.915s W0921 15:32:30.068957 1952523 catalog_manager.cc:14483] Catalog versions refresh failed: Timed out (yb/rpc/outbound_call.cc:790): GetTransactionStatus RPC (request call id 7807806) to 172.151.28.2:9100 timed out after 4.915s ``` Reading `pg_yb_catalog_version` has to resolve intents, so one slow transaction status tablet stalls every reader of it at once. a43409fcce8c349a656a6f6adc6f34786da6f8b0/D57779 does not cover this. Its invariant is that the cache holds the newest snapshot either path has *read*. Here neither path could read, so the cache kept answering with what it had, and the tserver's wall-clock budget ran out. This diff takes master's failing to read `pg_yb_catalog_version` into account: if master tells the tserver that it could not read `pg_yb_catalog_version`, tserver will not FATAL and will continue to wait until master is able to read `pg_yb_catalog_version` successfully. A master that is reading normally and still reports a version below the tserver's is unchanged: that is real divergence, as in #34074 where a PITR restore drops an already-committed DDL transaction, and it must keep aborting. **Upgrade / rollback safety** `catalog_versions_read_time` is a new optional field on an existing message. A new tserver against an old master sees it unset and falls back to the wall-clock measure, which is today's behaviour; an old tserver against a new master ignores it. Rollback needs no action. Original commit: 3caa00251170118674a7bba7937265cbbc798ea0 / D58473 Test Plan: All three below ran clean; release/clang21, macOS arm64. ``` ./yb_build.sh release --cxx-test pg_catalog_version-test --gtest_filter \ PgCatalogVersionFrozenCacheNoBroadcastTest.FrozenCacheDoesNotCrashTserverAfterLocalInstall ``` New. Freezes the master's cache with no DDL-commit broadcast, so the committing backend's own tserver is the only one holding the new version and every heartbeat afterwards reports the frozen, lower one. A `LogWaiter` asserts the staleness path was entered, so it cannot pass vacuously. Observed `master_read_ht` identical across all 16 stale heartbeats, `stale_for` pinned at 0.000s while `elapsed` reached 14.062s against a 5.000s threshold. Before this change the tserver FATAL at 5s. ``` ./yb_build.sh release --cxx-test pg_catalog_version-test --gtest_filter \ PgCatalogVersionStaleFatalTest.TableBehindTserverCrashesTserver ``` New, rolls `pg_yb_catalog_version` back below what the tservers hold. Observed `master_read_ht` advancing and eventually all three tservers FATAL on the staleness check. ``` ./yb_build.sh release --cxx-test pg_catalog_version-test --gtest_filter \ PgCatalogVersionFrozenCacheTest.BroadcastKeepsFrozenCacheFromCrashingTserver ``` Regression check that the broadcast path of D57779 is unaffected. Reviewers: sanketh, kfranz Reviewed By: kfranz Subscribers: yql, ybase Differential Revision: https://phorge.dev.yugabyte.com/D58899
| Commit: | b16a4aa | |
|---|---|---|
| Author: | Zachary Drudi | |
| Committer: | Zachary Drudi | |
[BACKPORT 2025.2][#31935] DocDB: Fix cloned tablet index_map for vector indexes Summary: Clone applied the source tablet's index_map onto the child. Restore was supposed to rewrite those IDs, but it does not for vector indexes, so the child kept the source's index entries. Send the target table's indexes on CloneTabletRequestPB (already remapped in the catalog). Remap snapshot index IDs on colocated tables the same way we already remap index_info. The tserver uses those IDs instead of the source index_map. **Upgrade/Rollback safety:** CloneTabletRequestPB gains optional repeated `target_indexes` (field 18). - Mixed cluster: old tservers ignore the field and keep inheriting the source index_map (pre-fix behavior). New tservers talking to an old master see an empty list and write an empty index_map instead of stale source IDs. - Rollback: new field is ignored; clone returns to inheriting the source index_map. Original commit: 29c5c8fab570e0d49a96ab7dd816eb0415586070 / D58453 Test Plan: ``` ./yb_build.sh release --cxx-test clone_state_manager-test --gtest_filter CloneStateManagerTest.ScheduleCloneOps ./yb_build.sh release --cxx-test pg_vector_index-test --gtest_filter 'PgVectorIndexColocationOnlyTest.CloneRemapsVectorIndexMap*' ``` Reviewers: mhaddad Reviewed By: mhaddad Subscribers: yql, ybase Differential Revision: https://phorge.dev.yugabyte.com/D58818
| Commit: | 65aeeb3 | |
|---|---|---|
| Author: | austenLacy | |
| Committer: | austenLacy | |
[#34038] YSQL: Pin eligible authentication attempts to follower-read snapshots Keep one fresh snapshot through password checks, HBA membership, database CONNECT authorization, and role settings. Prefetch resets and leader fallback must not replace it. Exclude shared-cache auth attempts, profiles, connection-manager and internal paths; resume normal invalidation before ordinary SQL. Routing remains default-off and experimental. --- _automated · pi_
| Commit: | 60e8afc | |
|---|---|---|
| Author: | austenLacy | |
| Committer: | austenLacy | |
[#34038] DocDB: Admit fixed-snapshot authentication reads on master followers Keep ordinary master reads leader-only. Admit only explicitly marked, allowlisted physical catalog reads on a replica that has applied the permanent reservation and reached the requested timestamp. Bound follower waiting without changing the snapshot. --- _automated · pi_
| Commit: | d13aa53 | |
|---|---|---|
| Author: | austenLacy | |
| Committer: | austenLacy | |
[#34038] YSQL: Establish fresh authentication catalog snapshots Follower safe time alone cannot establish authentication freshness. Have a ready leader select a clock-skew-covering timestamp and wait under its lease. Isolate bounded waits from ordinary RPC workers, preserve request deadlines, and reject overload. Snapshot acquisition requires the default-off routing gate and durable PITR exclusion; this layer does not route authentication reads. --- _automated · pi_
| Commit: | 33618b3 | |
|---|---|---|
| Author: | austenLacy | |
[#34038] DocDB: Allow startup PITR exclusion on existing universes A coordinated master restart can establish permanent PITR exclusion without live reservation arbitration. Capture startup intent before RPCs start, reject retained PITR or unfinished restore state, and preserve the existing configuration while committing the mode before leader readiness. --- _automated - pi_
| Commit: | 56312de | |
|---|---|---|
| Author: | Zachary Drudi | |
| Committer: | Zachary Drudi | |
[BACKPORT 2026.1][#31935] DocDB: Fix cloned tablet index_map for vector indexes Summary: Clone applied the source tablet's index_map onto the child. Restore was supposed to rewrite those IDs, but it does not for vector indexes, so the child kept the source's index entries. Send the target table's indexes on CloneTabletRequestPB (already remapped in the catalog). Remap snapshot index IDs on colocated tables the same way we already remap index_info. The tserver uses those IDs instead of the source index_map. **Upgrade/Rollback safety:** CloneTabletRequestPB gains optional repeated `target_indexes` (field 18). - Mixed cluster: old tservers ignore the field and keep inheriting the source index_map (pre-fix behavior). New tservers talking to an old master see an empty list and write an empty index_map instead of stale source IDs. - Rollback: new field is ignored; clone returns to inheriting the source index_map. Original commit: 29c5c8fab570e0d49a96ab7dd816eb0415586070 / D58453 Test Plan: ``` ./yb_build.sh release --cxx-test clone_state_manager-test --gtest_filter CloneStateManagerTest.ScheduleCloneOps ./yb_build.sh release --cxx-test pg_vector_index-test --gtest_filter 'PgVectorIndexColocationOnlyTest.CloneRemapsVectorIndexMap*' ``` Reviewers: mhaddad Reviewed By: mhaddad Subscribers: yql, ybase Differential Revision: https://phorge.dev.yugabyte.com/D58817
| Commit: | 3caa002 | |
|---|---|---|
| Author: | Minghui Yang | |
| Committer: | Minghui Yang | |
[#33572] YSQL: Measure catalog version staleness by master read time Summary: A stress test crashed a tserver on the staleness check added for this issue: ``` F0921 15:33:42 ../../src/yb/tserver/tablet_server.cc:1694] Ignoring ysql db 16642 catalog version update: new version too old. New: 19636, Old: 19637, stale_for: 36.802s, threshold: 36.000s ``` This one was a false positive. `pg_yb_catalog_version` did hold 19637: another tserver received it by heartbeat 16 seconds after this FATAL, and the restarted tserver was seeded with it. What was behind was the master's heartbeat cache, frozen for ~88 seconds because both of its writers failed together: ``` W0921 15:32:30.068879 1953073 object_lock_info_manager.cc:1107] Couldn't populate catalog version on exclusive lock release: Timed out (yb/rpc/outbound_call.cc:790): GetTransactionStatus RPC (request call id 7807806) to 172.151.28.2:9100 timed out after 4.915s W0921 15:32:30.068957 1952523 catalog_manager.cc:14483] Catalog versions refresh failed: Timed out (yb/rpc/outbound_call.cc:790): GetTransactionStatus RPC (request call id 7807806) to 172.151.28.2:9100 timed out after 4.915s ``` Reading `pg_yb_catalog_version` has to resolve intents, so one slow transaction status tablet stalls every reader of it at once. a43409fcce8c349a656a6f6adc6f34786da6f8b0/D57779 does not cover this. Its invariant is that the cache holds the newest snapshot either path has *read*. Here neither path could read, so the cache kept answering with what it had, and the tserver's wall-clock budget ran out. This diff takes master's failing to read `pg_yb_catalog_version` into account: if master tells the tserver that it could not read `pg_yb_catalog_version`, tserver will not FATAL and will continue to wait until master is able to read `pg_yb_catalog_version` successfully. A master that is reading normally and still reports a version below the tserver's is unchanged: that is real divergence, as in #34074 where a PITR restore drops an already-committed DDL transaction, and it must keep aborting. **Upgrade / rollback safety** `catalog_versions_read_time` is a new optional field on an existing message. A new tserver against an old master sees it unset and falls back to the wall-clock measure, which is today's behaviour; an old tserver against a new master ignores it. Rollback needs no action. Test Plan: All three below ran clean; release/clang21, macOS arm64. ``` ./yb_build.sh release --cxx-test pg_catalog_version-test --gtest_filter \ PgCatalogVersionFrozenCacheNoBroadcastTest.FrozenCacheDoesNotCrashTserverAfterLocalInstall ``` New. Freezes the master's cache with no DDL-commit broadcast, so the committing backend's own tserver is the only one holding the new version and every heartbeat afterwards reports the frozen, lower one. A `LogWaiter` asserts the staleness path was entered, so it cannot pass vacuously. Observed `master_read_ht` identical across all 16 stale heartbeats, `stale_for` pinned at 0.000s while `elapsed` reached 14.062s against a 5.000s threshold. Before this change the tserver FATAL at 5s. ``` ./yb_build.sh release --cxx-test pg_catalog_version-test --gtest_filter \ PgCatalogVersionStaleFatalTest.TableBehindTserverCrashesTserver ``` New, rolls `pg_yb_catalog_version` back below what the tservers hold. Observed `master_read_ht` advancing and eventually all three tservers FATAL on the staleness check. ``` ./yb_build.sh release --cxx-test pg_catalog_version-test --gtest_filter \ PgCatalogVersionFrozenCacheTest.BroadcastKeepsFrozenCacheFromCrashingTserver ``` Regression check that the broadcast path of D57779 is unaffected. Reviewers: sanketh, kfranz Reviewed By: kfranz Subscribers: ybase, yql Differential Revision: https://phorge.dev.yugabyte.com/D58473
| Commit: | a7bc7bf | |
|---|---|---|
| Author: | austenLacy | |
| Committer: | austenLacy | |
[#34038] YSQL: Pin eligible authentication attempts to follower-read snapshots Keep one fresh snapshot through password checks, HBA membership, database CONNECT authorization, and role settings. Prefetch resets and leader fallback must not replace it. Exclude shared-cache auth attempts, profiles, connection-manager and internal paths; resume normal invalidation before ordinary SQL. Routing remains default-off and experimental. --- _automated · pi_
| Commit: | 0f8f8ec | |
|---|---|---|
| Author: | austenLacy | |
| Committer: | austenLacy | |
[#34038] DocDB: Admit fixed-snapshot authentication reads on master followers Keep ordinary master reads leader-only. Admit only explicitly marked, allowlisted physical catalog reads on a replica that has applied the permanent reservation and reached the requested timestamp. Bound follower waiting without changing the snapshot. --- _automated · pi_
| Commit: | fc96cc6 | |
|---|---|---|
| Author: | austenLacy | |
| Committer: | austenLacy | |
[#34038] YSQL: Establish fresh authentication catalog snapshots Follower safe time alone cannot establish authentication freshness. Have a ready leader select a clock-skew-covering timestamp and wait under its lease. Isolate bounded waits from ordinary RPC workers, preserve request deadlines, and reject overload. Snapshot acquisition requires the default-off routing gate and durable PITR exclusion; this layer does not route authentication reads. --- _automated · pi_
| Commit: | 9a7438a | |
|---|---|---|
| Author: | austenLacy | |
[#34038] DocDB: Use creation-time PITR exclusion for auth follower reads Restrict opt-in to new universes created with `--disable_pitr=true`. A replicated, immutable mode removes live reservation arbitration while preserving the exclusion required by fixed authentication snapshots. Existing universes cannot opt in, and routing rollback does not enable PITR. --- _automated - pi_
| Commit: | a4373c9 | |
|---|---|---|
| Author: | Mark Lillibridge | |
| Committer: | Mark Lillibridge | |
[#33906] xCluster: group xCluster guarded info, set up to handle races Summary: ## Overview This is diff 1 of 3 to implement the PropagateXClusterGuardedInfo primitive. This diff pulls together the xCluster guarded info into a single protobuf structure, which is carried by the master-TServer heartbeat starting in this diff and will be carried by the PropagateXClusterGuardedInfo RPC in the third diff as well. There are two tricky things being addressed here: backward compatibility for upgrades and handling races between the heartbeat and the (coming) PropagateXClusterGuardedInfo RPC. ## New protobuf contents ``` ~/code/yugabyte-db/src/yb/common/common_types.proto:248: message XClusterGuardedInfoPB { // Version of the xCluster-guarded information in this copy. required XClusterGuardedInfoVersionPB xcluster_guarded_info_version = 1; // Local NamespaceId -> xCluster info for that namespace. map<string, XClusterNamespaceInfoPB> xcluster_info_per_namespace = 2; optional uint32 oid_cache_invalidations_count = 3; } ``` We plan to add an additional new field for the `StartPersisting/StopPersisting` mechanism in the future. ## Resolving races Because copies of the guarded information will be able to arrive via 2 paths, an older copy could arrive after a newer copy. To prevent the TServer from having its information go backwards, we use a version number on each copy: ``` ~/code/yugabyte-db/src/yb/common/common_types.proto:237: // Identifies how recent a copy of xCluster-guarded information is. Master copies that information // into heartbeat responses and into PropagateXClusterGuardedInfo RPCs; a TServer keeps the copy // with the highest (term, count), ordered lexicographically. // // term is the Raft term of the master leader that made the copy; count is number of times a copy // has been made on that master. message XClusterGuardedInfoVersionPB { required int64 term = 1; required uint64 count = 2; } ``` The TServer only updates its copy if the version number of the incoming copy is higher. Note that like today there is no guarantee that the guarded fields were all read or updated at the same time; while they cannot go backwards, one of them could be ahead of the other. It is thus incorrect to think of this mechanism as atomically updating the entire set of guarded fields as a unit. ## When we make the above copy for the heartbeat response For reasons that will be clear in the next diff, we need to make the copy of the guarded information after: ``` ~/code/yugabyte-db/src/yb/master/master_heartbeat_service.cc:480: auto desc_result = UpdateAndReturnTSDescriptorOrRespond(l.epoch(), *req, resp, &rpc); ``` To make sure this will work, we do this in this diff. This does mean that the protobuf information is not copied if the heartbeat response is "needs registration". Note that such responses are very rare in the presence of the persisted TServer registry -- when TServers start up they automatically send registration, which avoids that pathway. **Upgrade/Rollback safety:** An auto flag is used to move the information from its existing heartbeat fields into the new structure: ``` ~/code/yugabyte-db/src/yb/server/server_common_flags.cc:68: DEFINE_RUNTIME_AUTO_bool(skip_fields_moved_to_xcluster_guarded_info, kLocalVolatile, false, true, "Skip sending in master heartbeat responses the fields that were moved into " "xcluster_guarded_info, and skip applying them on TServers."); ``` The existing fields that are being moved into the new protobuf are renamed to start with `DEPRECATED_`. The behavior of the code in the diff is as follows: * sending the now deprecated fields: * only done if flag is off * sending the new protobuf * always done unless registration is required * processing the now deprecated fields on the TServer: * only done if the new protobuf is not present in the heartbeat and the flag is off * (the check if new is present is to avoid processing both of them) * processing the new protobuf on the TServer: * always done (when present) Conceptually the journey of an upgrade looks like: * pre-upgrade: send only old, process only old * mid-upgrade: send both, process new if present otherwise old * post upgrade: send only new, process only new Fixes #33906 Test Plan: Added new test to ensure out of order copies are handled correctly: ``` ybd release --cxx-test tablet_server-test --gtest_filter '*.ApplyXClusterGuardedInfoIfNewer' ``` Ran locally the existing tests that check the guarded information reaches the TServers (OID cache invalidation count and xCluster roles respectively). Test clusters promote all auto flags, so this exercises the post-upgrade path (new protobuf only): ``` ybd release --cxx-test xcluster_secondary_oid_space-test --gtest_filter OidAllocationTest.CacheInvalidation ybd release --cxx-test xcluster_ddl_replication-test --gtest_filter XClusterDDLReplicationTest.ExtensionRoleUpdating ``` The above plus all the other tests with the promoted auto flags will be run by Jenkins. Manually ran the same two tests with the flag off to exercise the pre-upgrade path (deprecated fields only). This was done with a temporary patch (not included) that demotes `use_xcluster_guarded_info_pb` on all processes via the master's `DemoteSingleAutoFlag` after the test clusters start, then waits until every process has applied the demotion. Reviewers: xCluster, zdrudi, hsunder Reviewed By: zdrudi Subscribers: ybase Differential Revision: https://phorge.dev.yugabyte.com/D58122
| Commit: | 04a2a70 | |
|---|---|---|
| Author: | Basava | |
| Committer: | Basava | |
[#33734] DocDB: TableLocks: Register WaitForLockers wait dependencies with the deadlock detector Summary: `WaitForLockers` is executed from YSQL on path like CREATE INDEX CONCURRENTLY/ALTER TABLE DETACH PARTITION CONCURRENTLY etc in one of the phases. These DDLs are executed in phases, and `WaitForLockers` is generally invoked at the start of a phase/txn when either no object locks are held or session object locks alone are held. Support for this construct was introduced in https://github.com/yugabyte/yugabyte-db/commit/40435579fdb6c143f448ac52855fc6bf356cca68, where the request is fanned-out to all tserver with a live lease, and we wait on active transactions which hold conflicting lock modes for the desired object. The problem is that these wait-for dependencies weren't registered with the `local_waiting_txn_registry` which forwards the wait-for probes to the central deadlock detector. Hence something like below could lead to an undetected deadlock, and stall until the statement timeout, because the waiting happens at both the Master's object lock manager and the Tserver's object lock managers (OLM) ``` CREATE TABLE parent(k INT PRIMARY KEY); CREATE TABLE child(k INT PRIMARY KEY, parent_k INT REFERENCES parent(k)); <session1> begin; select * from child; <session2> create index child_parent_k_idx ON child(parent_k); // acquire SHARE lock on child // pause right before the first waitforlockers call <session3> DROP TABLE parent CASCADE // acquires exclusive lock on parent, // waits for exclusive lock on child // at the YB-master's object lock manager // session3 -> session2 at YB-master's OLM insert into child ...; // acquires RowExclusive on // child, but needs // AccessShare on parent, // and waits on session3 // at the YB-tserver's OLM // session1 -> session3 // resume waitforlockers, waits on insert at the // tserver's lock manager hosting the DML // session2 -> session1 at the YB-tserver's OLM ``` The deadlock `session1 -> session3 -> session2 -> session1` is never detected because the dependency `session2 -> session1` that came out of `WaitForLockers` wasn't being forwarded to the central deadlock detector. This revision fixes the issue by forwarding the wait-for probes arising out of the `WaitForLockers` request to the central deadlock detector. **Upgrade / Downgrade safety** Adding a new proto field to `WaitForLockersMultipleGlobalRequestPB` message which is only used when table locking is enabled, and the feature hasn't been shipped on in any official release yet. Test Plan: Jenkins ./yb_build.sh --cxx-test pg_object_locks-test --gtest_filter PgObjectLocksTest.WaitForLockersParticipatesInDeadlockDetection -n 5 --tp 1 Without the changes in the diff, the new test fails with the following ``` ../../src/yb/yql/pgwrapper/pg_object_locks-test.cc:2515: Failure Value of: drop_status.ok() ^ insert_status.ok() Actual: false Expected: true drop_status: Network error (yb/yql/pgwrapper/libpq_utils.cc:685): Execute of 'DROP TABLE parent CASCADE' failed: 7, message: ERROR: canceling statement due to statement timeout (pgsql error 57014) (aux msg ERROR: canceling statement due to statement timeout), insert_status: Network error (yb/yql/pgwrapper/libpq_utils.cc:685): Execute of 'INSERT INTO child VALUES (1, 1)' failed: 7, message: ERROR: canceling statement due to statement timeout (pgsql error 57014) (aux msg ERROR: canceling statement due to statement timeout) ``` Reviewers: patnaik.balivada, amitanand Reviewed By: patnaik.balivada Subscribers: yql, ybase Differential Revision: https://phorge.dev.yugabyte.com/D58268
| Commit: | 3364778 | |
|---|---|---|
| Author: | Minghui Yang | |
| Committer: | Minghui Yang | |
[#34309] YSQL: Bound concurrent YSQL catalog prefetches on the master leader Summary: A DDL invalidates every tserver's cached catalog prefetch results, so the next connection on each of them prefetches the catalog from the master leader. The catalog version is part of the response cache key, so a burst of DDLs produces not one fill per node but one fill per version per node: a connection pins the version it wants at startup, and with connections arriving every second and a fill taking tens of seconds, several versions are in flight on each node at once. Thirteen DDLs on a 200-node cluster put on the order of a thousand concurrent full-catalog prefetch reads on a single master leader. Serving those prefetch RPCs is CPU bound, so beyond the leader's capacity each read slows in proportion to the number queued: a preload taking 22s against a quiet leader took 280s, long enough for callers to give up and reconnect, which in turn added additional load. The fan-in scales with cluster size, so it worsens as more tserver nodes are added. This diff adds admission control for catalog prefetch RPCs at master side. When master processes a read RPC, master_max_concurrent_ysql_catalog_prefetches caps how many prefetch RPCs the master leader serves at once; the rest are rejected with ServiceUnavailable, which reaches the client as ERROR_SERVER_TOO_BUSY and is retried against the same master on the client's own exponential backoff. Each admitted prefetch then runs at the speed the leader can actually serve, instead of every prefetch slowing in proportion to how many are in flight, and the waiting moves to the callers, where each rejection doubles the wait before the next attempt. The retrier jitters that by only a few milliseconds, so a fleet rejected together comes back together; what the doubling bounds is how often it comes back. The flag master_max_concurrent_ysql_catalog_prefetches defaults to -1, resolved at startup to twice the core count available to the process, following the AutoInitServiceFlags pattern the tserver already uses. Serving a prefetch is CPU bound, so a limit near the core count keeps the leader fully busy while bounding the queue; the doubling covers the part of each request that waits on RocksDB rather than on CPU. 0 disables the bound. That default is a guess at how many prefetches the leader can serve at once, and doubling the core count makes it a deliberately generous guess. Guessing too generously leaves the leader no better off than it is today, because a bound that is never reached never takes effect. Guessing too tightly leaves it worse off than today: prefetches the leader could have served are rejected, and a client rejected over and over runs out its RPC deadline -- ten minutes by default -- and fails a connection that works now. Only prefetches are bounded. PgsqlReadRequestPB gains catalog_prefetch_kind, set by the system table prefetcher. A catalog cache miss, a relcache build and a systable scan cost a fraction of a prefetch and run inside a query that is already executing; delaying those would add latency where it is most visible, and would let prefetches crowd out the cheap reads they contend with. NONE is never bounded, so an older client against a newer master keeps working (e.g., during upgrade). There are two thresholds, not one: the full limit above, and a lower one set by master_new_ysql_catalog_prefetch_pct as a percentage of it. Only the opening read of a connection-startup prefetch (i.e., the first prefetch RPC) is held to the lower threshold; every other read may use the full limit. A prefetch is a sequence of paged reads, a hundred or more of them on a large catalog, and refusing one late in the sequence wastes every read before it -- the connection starts over. If every read faced the same bound, that waste would be the norm under pressure: the leader busy, few prefetches finishing. So a read carrying a paging state is continuing a sequence, and gets the full limit. Same for a backend refreshing its catalog cache, which is why the request has the reason why it is prefetching rather than merely that it is. Such a backend is already serving a session and can run nothing at all until the prefetch completes, where a connection still starting up has little invested. The refresh case is also the one that shows up when things are already going badly: with invalidation messages available a backend refreshes incrementally and never prefetches, and falls back to a full prefetch only when the message chain cannot carry it -- a cache reset, a gap wider than the queue, a version whose messages are missing. Those are the bad times to queue it behind new connection startups. That leaves the opening read of a startup prefetch, held to the lower threshold and floored at a single admission so that a small limit still lets connections begin. The two kinds come from the two entry points that already distinguish them: YBPreloadRelCache for a refresh, YbPrefetchRequiredData for connection startup. Admission is a new hook on ReadTabletProvider, the interface the read path already uses to ask its owner for behaviour that differs between master and tserver -- it is how the master supplies its system tablet where a tserver looks up a tablet peer. Putting it there keeps the mechanism in generic code and the policy in the master, and the tserver's default admits everything. The hook returns a handle that occupies one unit of the limit until it is destroyed, and ReadQuery holds it. ReadQuery lives until the read has produced its response, not merely until the service method returns, so the unit stays occupied for the scan and encode that make a prefetch expensive. A read can reschedule itself while waiting for safe time, so releasing on return would free the unit while the read was still running. TEST_ysql_catalog_prefetch_rejections_to_inject injects rejections through the same path, so the client-side behaviour can be covered without a cluster large enough to saturate a leader. New tests added. They cover that a rejected prefetch is retried rather than failing the connection, that admissions are released when a read completes, and that the new-prefetch limit has a floor so new prefetch RPCs are not 100% rejected. Upgrade / rollback safety: Nothing here is persisted, so there is no migration and no on-disk format change. enable_ysql_catalog_prefetch_admission is an AutoFlag, so the bound does not engage until the upgrade is finalized. catalog_prefetch_kind is a new field on a message a tserver sends its master. An old master ignores it; a new master reads NONE from an old tserver and leaves that tserver unbounded. Test Plan: ./yb_build.sh release --cxx-test pg_catalog_perf-test \ --gtest_filter 'PgCatalogPrefetchAdmissionTest.*' Reviewers: sanketh, kfranz, esheng Reviewed By: kfranz Subscribers: ybase, yql Differential Revision: https://phorge.dev.yugabyte.com/D58730
| Commit: | 1c1c837 | |
|---|---|---|
| Author: | austenLacy | |
| Committer: | austenLacy | |
[#34038] YSQL: Pin eligible authentication attempts to follower-read snapshots Keep one fresh snapshot through password checks, HBA membership, database CONNECT authorization, and role settings. Prefetch resets and leader fallback must not replace it. Exclude shared-cache auth attempts, profiles, connection-manager and internal paths; resume normal invalidation before ordinary SQL. Routing remains default-off and experimental. --- _automated · pi_
| Commit: | af158f2 | |
|---|---|---|
| Author: | austenLacy | |
| Committer: | austenLacy | |
[#34038] YSQL: Establish fresh authentication catalog snapshots Follower safe time alone cannot establish authentication freshness. Have a ready leader select a clock-skew-covering timestamp and wait under its lease. Isolate bounded waits from ordinary RPC workers, preserve request deadlines, and reject overload. Snapshot acquisition requires the default-off routing gate and durable PITR exclusion; this layer does not route authentication reads. --- _automated · pi_
| Commit: | 9330652 | |
|---|---|---|
| Author: | austenLacy | |
| Committer: | austenLacy | |
[#34038] DocDB: Admit fixed-snapshot authentication reads on master followers Keep ordinary master reads leader-only. Admit only explicitly marked, allowlisted physical catalog reads on a replica that has applied the permanent reservation and reached the requested timestamp. Bound follower waiting without changing the snapshot. --- _automated · pi_
| Commit: | cf9c306 | |
|---|---|---|
| Author: | austenLacy | |
| Committer: | austenLacy | |
[#34038] DocDB: Reserve catalog follower reads without PITR System-catalog restore can change historical snapshots. Require an explicit, durable reservation that permanently excludes PITR before future catalog follower routing can be enabled. Ordinary backup snapshots and data-only restores remain supported. Serialize reservation and PITR admission, retaining term-local claims after failed or timed-out writes because those writes may still commit. Gate first reservation behind a persisted AutoFlag; demotion cannot release it. Downgrade to binaries that do not enforce a committed reservation is unsupported. Master read routing remains leader-only. --- _automated · pi_
| Commit: | 66cc180 | |
|---|---|---|
| Author: | austenLacy | |
| Committer: | austenLacy | |
[#34038] YSQL: Publish catalog-version snapshot times Catalog reads must not start before the snapshot that supplied their catalog-version cache key. Carry that time with version reports and advance the tserver clock before publishing the versions. Gate report timestamps with `ysql_enable_catalog_version_read_time`. --- _automated · pi_
| Commit: | 29c5c8f | |
|---|---|---|
| Author: | Zachary Drudi | |
| Committer: | Zachary Drudi | |
[#31935] DocDB: Fix cloned tablet index_map for vector indexes Summary: Clone applied the source tablet's index_map onto the child. Restore was supposed to rewrite those IDs, but it does not for vector indexes, so the child kept the source's index entries. Send the target table's indexes on CloneTabletRequestPB (already remapped in the catalog). Remap snapshot index IDs on colocated tables the same way we already remap index_info. The tserver uses those IDs instead of the source index_map. **Upgrade/Rollback safety:** CloneTabletRequestPB gains optional repeated `target_indexes` (field 18). - Mixed cluster: old tservers ignore the field and keep inheriting the source index_map (pre-fix behavior). New tservers talking to an old master see an empty list and write an empty index_map instead of stale source IDs. - Rollback: new field is ignored; clone returns to inheriting the source index_map. Test Plan: ``` ./yb_build.sh release --cxx-test clone_state_manager-test --gtest_filter CloneStateManagerTest.ScheduleCloneOps ./yb_build.sh release --cxx-test pg_vector_index-test --gtest_filter 'PgVectorIndexColocationOnlyTest.CloneRemapsVectorIndexMap*' ``` Reviewers: mhaddad Reviewed By: mhaddad Subscribers: ybase, yql Differential Revision: https://phorge.dev.yugabyte.com/D58453
| Commit: | 7daec94 | |
|---|---|---|
| Author: | skhilar | |
| Committer: | skhilar | |
[PLAT-22611] Dual direction federated IAM for on-prem universe Summary: An on-prem universe can have nodes on different clouds, so cross-cloud federation direction is a per-node property. It was hardcoded: on-prem was always treated as AWS-backed, so a GCP-backed on-prem node could never reach S3. PLAT-22355 already detects each node's cloud; nothing read it. This makes each node pick its direction from its detected cloud, and replaces the direction-pair model with one keyed by target cloud, so adding Azure/OCI later costs one entry per cloud instead of one per ordered pair. | Area | Change | |---|---| | Model | `CrossCloudFederationDirection` (pair enum) deleted. All three providers now carry the same `crossCloudFederationTargets` list via two `CloudInfoInterface` defaults; flat `federatedIam*` fields deprecated and read through. | | Node | `ManageCloudFederation` resolves `(detected cloud, target cloud)` per node. No fallback - a node whose cloud cannot be determined fails the task rather than being configured with a guess. | | Proto | `FlowDirection` enum replaced by `sourceCloud`/`targetCloud`; `GcsOnAwsConfig`/`S3OnGcpConfig` renamed `GcsConfig`/`S3Config`. | | Multicloud | Federation on a multi-provider cluster rejected at config time (`UniverseCRUDHandler`, v2 enable API) instead of warned about at backup time. | | YBA access | GCS storage config gains `USE_CROSS_CLOUD_FEDERATION`, mirroring S3. Both flags describe *YBA's own* access, never the DB nodes'. Fixes YBA-on-GCP -> GCS, which built an AWS-IMDS credential for a same-cloud bucket. Also fixes `AWSUtil.validate` falling into the default AWS chain when saving a federated S3 config. | | UI | Federated IAM fields on the on-prem provider page, modelled as "which storage clouds must my nodes reach" - the node's own cloud is detected, never asked for. | | Rollout | All federated IAM UI gated behind `yb.ui.feature_flags.enable_cross_cloud_federated_iam` (GLOBAL, default false). Gates display only; an already-configured provider or storage config keeps working. | Reviewers must note: - Renames: `federationTargets` -> `crossCloudFederationTargets`, `federationAudience/RoleArn` -> `crossCloudFederation*`. Backup-snapshot fields keep their names deliberately - they are persisted, and renaming would orphan existing backups. All PREVIEW; no shipped API is broken. - Behaviour change: an on-prem node whose cloud cannot be probed now fails the task instead of defaulting to GCS-on-AWS. Deduplicated on the way: `GCPProviderValidator` and `OnPremValidator` shared identical ARN/audience regexes (now one definition on `CrossCloudFederationTarget`), and `BackupHelper` had two near-identical resolve methods (now one). Test Plan: - Unit tests: 228 tests across the federation suites, 0 failures. New: `OnPremValidatorTest` (12, class did not exist), `CloudInfoInterfaceFederationTest` (8, incl. legacy flat-field back-compat), `BackupHelperTest` +4 pinning each YBA-cloud/bucket-cloud combination. - Build and codegen: `sbt compile`, `sbt swaggerGen`, node-agent build (4 binaries on the new proto), `go test --tags testonly ./app/task/`, UI `tsc --noEmit`, `google-java-format` - all clean. - Manual E2E on a mixed on-prem universe (1 AWS node, 1 GCP node, YBA on GCP): each node detected its own cloud and was configured for the opposite storage cloud; GCS backup succeeded from both. Verified the on-node artifacts and that YBC ran with the credential. - Rollout flag: with `yb.ui.feature_flags.enable_cross_cloud_federated_iam` off, the provider and backup config pages render as before. - Not yet exercised: restore, and the retroactive `POST .../cross-cloud-federation` path. Reviewers: #yba-api-review!, amalyshev, nsingh, yshchetinin, kkannan Reviewed By: amalyshev, kkannan Subscribers: svc_phabricator, yugaware Differential Revision: https://phorge.dev.yugabyte.com/D58606
| Commit: | c4f4678 | |
|---|---|---|
| Author: | Saurav Tiwary | |
| Committer: | Saurav Tiwary | |
[#30813] YSQL: Support Logical Replication with Transactional DDL via historical catalog reads Summary: This revision introduces support for Logical Replication with Transactional DDL. As part of D56763 / 46a9f7a, we had added support to detect DDLs in the virtual WAL by treating DMLs on PG catalog tables (`pg_class` and `pg_attribute`) as DDLs. This revision builds on that and uses the DDL records sent by the virtual WAL in the walsender to correctly stream transactions that involve interleaved DDLs and DMLs. ##### PG Background In PG, the walsender maintains a historical snapshot of the catalog while reading the WAL serially. For every transaction that it sees, it does the following: 1. DML records: The heap tuples are mem copied from the WAL record and queued to reorder buffer. Note the memcpy i.e. **PG does not need the Relation description to queue entries into the reorder buffer** 2. Invalidation message: Queued to the reorder buffer 3. COMMIT: All collected records of the transaction are iterated and streamed in order. Invalidation message records are used to invalidate the relcache and lead to a new RELATION message being sent on the next DML on that table. The output plugin also ignores records that are not of use, for e.g. publication only asks for UPDATE records, therefore DELETE or INSERT will be discarded. ##### YB vs PG differences In YB, the biggest difference is in our handling of DML records. The virtual WAL returns a `QLValuePB` for the DML record. Therefore, we need to first convert it into a PG heap tuple before queueing into the reorder buffer. This means that we need the schema of the relation in two places (vs just 1 in PG): 1. Converting from proto to heap tuple before inserting to reorder buffer 2. In the reorder buffer & output plugin: Sending the RELATION message to the client on the first DML after a DDL, filtering not needed records in output plugin These 2 places fall in different areas of the code and require us to move the **read time backwards in the presence of interleaved DDLs**. Consider this example: ``` create table test (a int primary key); begin; insert into test values (1); alter table test add column b int; insert into test values (2, 2); commit; ``` The logic of walsender will be: 1. Get the batch of changes from virtual WAL which will be: BEGIN, INSERT (1), DDL, INSERT (2, 2), COMMIT 2. For each record of the transaction: - Begin: start a new txn - Insert(1): set read_time_context. Get the schema of `test` at this time to create the heap tuple and queue into reorder buffer - DDL: invalidate relcache. Append an `YB_REORDER_BUFFER_CHANGE_DDL` entry into the reorder buffer for the DDL record - Insert(2, 2): Update read_time_context. The in_txn_limit will be updated to the record time of this Insert record. Get the schema of `test` at this time to create the heap tuple and queue into reorder buffer. Note that the relcache was invalidated in the previous DDL, therefore, a fresh read request will be made to get the schema - COMMIT: Call `ReorderBufferCommit` to start streaming the txn records to the client and **move the read time back to what it was at the start of the transaction** (not required in PG). 3. In `ReorderBufferCommit` -> `ReorderBufferProcessTXN`. Iterate over all queued records: - Begin: no action required - Insert(1): Get the schema of `test` at the correct time and stream the record to the output plugin. To process this record, we needed the schema of 'test' before the DDL of the next step which is why we had to move the read time back to what it was at the start of the txn. - DDL: update the read time context, invalidate relcache - Insert(2, 2): Get the schema of `test` at this time and stream the record to the output plugin - commit: send commit record to the output plugin and cleanup --- To be able to read the PG catalog within the context of a historical transaction, this revision introduces a new session type: `kHistoricalRead`. This is only used by the walsender PG backend. Properties: 1. The PG backend (walsender) passes the DocDB txn id instead of the session creating one: `historical_read_transaction_id` 2. When receiving a request for this session, a YBTransaction object is fabricated with the given txn id without any registration or heartbeat to the status tablet 3. All Write ops are rejected 4. A flag `is_read_only_historical_committed_txn` is sent in the Read request to the tablet servers 5. When set, the transaction participant skips transaction status resolution and uses the passed `read_time` & `in_txn_limit` to resolve the intents. Note that it is the responsibility of the caller to **ensure that the intents are present** i.e. not GC'ed. In this feature i.e. logical replication we stop intent from getting cleaned up. So that requirement is satisfied. 6. `read_query` additionally disables the hybrid-time file filter for such historical reads: otherwise the intents DB is filtered out once no transaction is running any more. For step 1, the walsender tracks the read_time and in_txn_limit of the current record in a historical read context (`YBCSetHistoricalReadContext`). The stored read_time, in_txn_limit and the transaction id is sent in the Perform RPC while reading the catalog. Note that the existing `yb_read_time` feature was insufficient because we also need to read DDLs done prior to a statement within the txn. This also needs the `in_txn_limit`. --- **Upgrade/Rollback safety:** The entire change is gated on `TEST_ysql_yb_enable_replication_slot_transactional_ddl`. No production users are expected to use this feature. After testing, this flag will be eventually converted to an autoflag. Test Plan: New tests, each run individually: ``` ./yb_build.sh --java-test 'org.yb.pgsql.TestPgReplicationSlotWithTxnDdl#addColumnInsideTransaction' ./yb_build.sh --java-test 'org.yb.pgsql.TestPgReplicationSlotWithTxnDdl#dropColumnInsideTransaction' ./yb_build.sh --java-test 'org.yb.pgsql.TestPgReplicationSlotWithTxnDdl#multipleDdlsInsideTransaction' ./yb_build.sh --java-test 'org.yb.pgsql.TestPgReplicationSlotWithTxnDdl#rolledBackDdlInsideTransaction' ./yb_build.sh --java-test 'org.yb.pgsql.TestPgReplicationSlotWithTxnDdl#ddlOnlyTransaction' ./yb_build.sh --java-test 'org.yb.pgsql.TestPgReplicationSlotWithTxnDdl#createTableAndDropColumnInsideTransaction' ./yb_build.sh --java-test 'org.yb.pgsql.TestPgReplicationSlotWithTxnDdl#singleShardInsertAfterTransactionalDdl' ``` Reviewers: sumukh.phalgaonkar, asrinivasan, pjain, patnaik.balivada, myang Reviewed By: pjain Subscribers: jason, ybase, yql Differential Revision: https://phorge.dev.yugabyte.com/D57627
| Commit: | 99b85af | |
|---|---|---|
| Author: | Sergei Politov | |
| Committer: | Sergei Politov | |
[#33912] DocDB: Fix the vector payload mode when the index is created Summary: #33733 made a vector index store the ybctid as the vector payload, deciding it from the `vector_index_store_ybctid` flag every time the index was opened. Such an index could mix chunks written with and without payload, so its vectors had to be resolved through two different paths, and turning the flag off back was not supported at all. Whether the vector reverse mapping is written was decided from the vector indexes of the tablet, although the reverse mapping is owned by the table: creating an index could stop the writes for a table which already had entries, and dropping it could start them again. Both decisions are now made by master when the object is created and are never changed afterwards: - `TablePropertiesPB::skip_vector_reverse_mapping` is set at table creation from the `vector_index_store_payload` flag. Such a table writes no reverse mapping at all, neither an entry on insert nor a tombstone on delete, and all its vector indexes store a payload whatever the flag says later. - `PgVectorIdxOptionsPB::store_payload` is set at index creation, so an index stores a payload either for all its vectors or for none of them. `DocVectorIndex::StoresPayload` forces it for an index of a table which writes no reverse mapping, so the option cannot contradict the table. A table without `skip_vector_reverse_mapping` always writes the reverse mapping and its tombstones, whether or not its vector indexes store a payload. It supports vector indexes of both kinds and keeps behaving exactly as before, so an index created there with a payload just resolves its vectors without reading the reverse mapping. Tables and indexes created before this change have both options unset, which is that same behaviour. The write path decides per table rather than per tablet: reverse mapping entries are gated by `CompactionSchemaInfo::table_writes_vector_reverse_mapping` for a packed row, and by `VectorIndexesUpdater::TableWritesVectorReverseMapping` for a column-keyed one. A deleted vector id is not produced at all for a table which writes no reverse mapping (`PgsqlWriteOperation::FillRemovedVectorId`), so nothing downstream writes a tombstone for it. Search of an index which stores payloads never takes the no-deletions shortcut in `PgsqlVectorFilter::Init`: such an index keeps its deleted vectors, and they are detected by fetching the row the stored ybctid points to. Mixing the two modes within an index is not possible anymore, so the machinery which supported it is gone: - `VectorLSMMergeFilter::RestorePayload` and `MergingIterator::RestorePayloadIfMissing`, which attached the ybctids to payload-less vectors during compaction. - `DocVectorIndexContext::CreateReverseMappingReaderAtHistoryCutoff`, which existed only to serve the payload restore. The flag moved to master accordingly, and is renamed to `vector_index_store_payload`, since the payload is going to carry more than the ybctid later, e.g. columns for covering indexes. **Upgrade/Rollback safety:** - Both new fields, `TablePropertiesPB::skip_vector_reverse_mapping` and `PgVectorIdxOptionsPB::store_payload`, are additive `optional` fields which default to false. An old binary ignores them, and a new binary reading an object created by an old master sees them unset, which is the pre-change behaviour: the table writes the reverse mapping and the index resolves its vectors through it. - Both are set by master at creation only. `TableProperties::AlterFromTablePropertiesPB` deliberately does not merge `skip_vector_reverse_mapping`, so ALTER cannot change it, and it is restored from backup metadata next to `owns_vector_reverse_mapping`. No migration of existing tables or indexes is involved. - The feature is guarded by the runtime flag `vector_index_store_payload`, disabled by default, so an upgrade changes nothing by itself. - It is a runtime flag and not an AutoFlag on purpose: a table created while the flag is on has no reverse mapping to fall back to, so an older binary, which ignores both new fields, cannot serve its vector indexes. The flag stays off until deleted vectors are removed from an index which stores payloads, see #33912, and enabling it is not rollback safe until then. Test Plan: ./yb_build.sh --cxx-test doc_vector_id-test --gtest_filter DocVectorIdTest.VectorIndexPayload ./yb_build.sh --cxx-test schema-test ./yb_build.sh --cxx-test non_transactional_batch_writer-test ./yb_build.sh --cxx-test pg_vector_index-test --gtest_filter PgVectorIndexUtilTest.StoredYbctidFixedForTable ./yb_build.sh --cxx-test pg_vector_index-test --gtest_filter PgVectorIndexUtilTest.SstDumpStoredYbctid ./yb_build.sh --cxx-test pg_vector_index-test --gtest_filter PgVectorIndexUtilTest.SearchSkipsDeletedVectorsWithStoredYbctid ./yb_build.sh --cxx-test pg_vector_index-test --gtest_filter PgVectorIndexUtilTest.SearchSkipsTombstonedReverseMapping ./yb_build.sh --cxx-test pg_vector_index-test --gtest_filter PgVectorIndexUtilTest.ReverseMappingPostSplitCompaction ./yb_build.sh --cxx-test pg_vector_index-test --gtest_filter PgVectorIndexUtilTest.SstDump ./yb_build.sh --cxx-test pg_vector_index-test --gtest_filter PgVectorIndexUtilTest.BackfillSkipsReverseMapping ./yb_build.sh --cxx-test pg_vector_index-test --gtest_filter PgVectorIndexUtilTest.BackfillWritesReverseMapping Reviewers: arybochkin Reviewed By: arybochkin Subscribers: ybase, yql Tags: #jenkins-ready Differential Revision: https://phorge.dev.yugabyte.com/D58227
| Commit: | 78c1bf5 | |
|---|---|---|
| Author: | Ella Baron | |
[#16670] DocDB: dist-trace: drop dead tablet_id tags from remote_bootstrap.proto remote_bootstrap_proto is built without lightweight messages, so the tags emitted nothing. Even with MESSAGES TRUE they would stay silent: remote bootstrap and snapshot transfer run on threads with no active trace context.
| Commit: | 660994c | |
|---|---|---|
| Author: | Anton Rybochkin | |
| Committer: | Anton Rybochkin | |
[BACKPORT 2025.2][#31065] docdb: Deny split until the parent tablet vector index is post-split compacted Summary: **Background** After a tablet split, each child tablet's vector indexes inherit all vectors from the parent, including roughly half that no longer belong to the child's key range. Vector search must filter these irrelevant vectors via reverse-mapping lookups into RegularDB (two reads per vector), which is significantly slower than the vector search itself. If splits continue before the inherited vectors are compacted away, the ratio of irrelevant vectors grows with each generation, degrading search latency. To control this, tablet splits should be postponed until co-hosted vector indexes have completed post-split compaction. **Changes** 1. **Splits wait for vector index post-split compaction.** `Tablet::StillHasOrphanedPostSplitDataAbortable()` now also reports leftover vector index parent data, so a tablet is not split again (and the load balancer does not move it) until that data is gone. New advanced gflag `vector_index_require_parent_data_compacted_before_split` (default `true`) turns the wait off. The existing hidden gflag `vector_index_include_into_post_split_compaction` turns the vector index side off entirely. 2. **Scheduling a post-split compaction uses its own check.** `Tablet::NeedPostSplitCompaction()` answers a different question: is a post-split compaction still owed for this tablet? `TriggerPostSplitCompactionIfNeeded()` calls it instead of the split check. Otherwise setting `vector_index_require_parent_data_compacted_before_split=false` would also stop an interrupted vector index compaction from being retried when the tablet reopens. 3. **Each vector index tracks its own progress.** Two new `ConsensusFrontier` fields: `split_generation` and `split_min_chunk_serial_no`. Chunks with a serial number below the bound came from the parent tablet. At open time `InitFrontiers` compares the tablet's `split_generation` with the frontier's; if a new split happened it records `LastSerialNo() + 1` as the bound and flushes. After that, `ParentDataCompacted()` is simply `MinSerialNo() >= bound`, or true when there are no chunks. Nothing about vector index compaction is stored in tablet metadata. 4. **New `split_generation` in `KvStoreInfoPB`.** A tserver-local counter, parent + 1 in `CreateSplitChildMetadata`, which is how a vector index notices that a new split happened. Restored from the vector index manifests during bootstrap if a rollback dropped it -- see Upgrade/Rollback safety. 5. **Reverse mapping lookups now respect tablet key bounds.** `IndexReverseMappingReader::Fetch` returns nothing for a mapping whose ybctid is outside the tablet's range. That lets a split child's compaction drop the vectors that now belong to its sibling. Tombstones are still returned unchanged (will be fix in a follow-up revision). 6. **`parent_data_compacted` renamed to `rocksdb_parent_data_compacted`** in `KvStoreInfoPB`, `TabletStatusPB` and everywhere it is used, because it only covers RocksDB and is now easy to confuse with the vector index state. 7. **`CreateSubtablet` renamed to `CreateSplitChildTablet`** (and `CreateSubtabletMetadata` to `CreateSplitChildMetadata`), since both are only used for tablet splitting. 8. **RocksDB and vector indexes are compacted independently.** `TriggerManualCompactionIfNeeded` takes an `IncludeVectorIndexes` argument and skips whichever side is already done; `VectorIndexList::Compact` skips indexes that are already compacted. If the RocksDB compaction fails, the vector index compaction is skipped: its merge filter needs the obsolete reverse mappings to be gone first. 9. **New `vector_indexes_parent_data_compacted` in `TabletStatusPB`.** A diagnostic field for tests and the status page. Always reported by the tserver when the tablet is available, and never used as persistent state. **Upgrade/Rollback safety** All new fields are optional and default to pre-change behavior: `split_generation` defaults to 0 (matching a never-split tablet) and `split_min_chunk_serial_no` to 0 (read as "no inherited parent data"), so pre-existing split children are treated as having nothing to compact and the gate applies only to splits performed by the new binary; the renames keep their field numbers (`KvStoreInfoPB` 8, `TabletStatusPB` 18), so on-disk and wire formats are unchanged. Rollback drops `KvStoreInfoPB.split_generation` from the superblock while the vector index manifests keep it, so on the next upgrade `TabletBootstrap::MaybeUpdateMetaAfterTabletHasBeenOpened` restores it from the largest generation persisted across the tablet's vector indexes -- without that, the next split child would get generation 1, which is not greater than the persisted value, and would skip stamping its own parent data boundary. Merge conflicts: Resolved on 2026.1 and inherited here via chaining: - src/yb/tablet/metadata.proto:136 - 2026.1 has no tiered storage, so the hunk's context (`tier_paths = 12` and the `TierPathPB` message) conflicted. Added only this commit's `split_generation` field and dropped that context. Kept field number 13, leaving 12 unused, so `KvStoreInfoPB` stays wire-compatible with master and a later tiered-storage backport still gets 12. - src/yb/tablet/tablet_metadata.cc:2158 - same cause: master's `child_rocksdb_dir` local exists only to feed `BuildTierPaths`, which 2026.1 does not have. Kept the branch's `set_rocksdb_dir(GetSubRaftGroupDataDir(...))` line and applied only this commit's two changes (`set_rocksdb_parent_data_compacted`, `set_split_generation`). Resolved on 2025.2: - src/yb/docdb/non_transactional_batch_writer-test.cc:33 - dropped the whole `CountingVectorIndex` test double the hunk wanted to add; this commit only adds two overrides to it. The class comes from GH#31899 (5b930664ec8 on master), whose 2025.2 backport (c6346fda8c2) landed without it, so neither it nor `ExternalApplyGatesVectorIndexFeed` exists here and nothing else in the file references them. The real `DocVectorIndex` implementor in doc_vector_index.cc already got both new pure-virtual overrides via clean auto-merge, so this file is unchanged by the backport on 2025.2. - src/yb/docdb/consensus_frontier-test.cc:201 - added the new `SplitGenerationAndMinChunkSerialNo` test but kept 2025.2's `} // namespace docdb` / `} // namespace yb` closing instead of master's `} // namespace yb::docdb`. - src/yb/yql/pgwrapper/pg_vector_index-test.cc:74 - the hunk re-added three `DECLARE_bool`s that 2025.2 already declares elsewhere in the list. Added only this commit's new `vector_index_require_parent_data_compacted_before_split`, in sorted position. Apart from the dropped test double above, the added and removed lines match the original commit. Original commit: ca6fd80065b / D52469 Test Plan: Jenkins: jobs: full ./yb_build.sh --cxx-test consensus_frontier-test --gtest_filter=ConsensusFrontierTest. SplitGenerationAndMinChunkSerialNo ./yb_build.sh --cxx-test tablet-split-test --gtest_filter=TabletSplitTest.StillHasOrphanedPostSplitDataVectorIndex ./yb_build.sh --cxx-test tablet-split-test --gtest_filter=TabletSplitTest.SplitGenerationOnSubtablet ./yb_build.sh --cxx-test tablet-split-itest --gtest_filter=TabletSplitSingleServerITest.PostSplitCompactionScheduledOnTabletOpen ./yb_build.sh --cxx-test pg_vector_index-test --gtest_filter=PgDistributedVectorIndexTest.SplitBlockedWithOrphanedPostSplitData/* ./yb_build.sh --cxx-test pg_vector_index-test --gtest_filter=PgDistributedVectorIndexTest.ManualSplitSimple/* ./yb_build.sh --cxx-test pg_vector_index-test --gtest_filter=PgDistributedVectorIndexTest.SplitGenerationRestoredAfterSuperblockReset/* ./yb_build.sh --cxx-test non_transactional_batch_writer-test --gtest_filter=NonTransactionalBatchWriterTest.ExternalApplyGatesVectorIndexFeed ./yb_build.sh --cxx-test tablet-split-test --gtest_filter=TabletSplitTest.SplitTablet ./yb_build.sh --cxx-test tablet-split-itest --gtest_filter=TabletSplitSingleServerITest.TabletServerOrphanedPostSplitData ./yb_build.sh --cxx-test pg_tablet_split-test --gtest_filter=PgTabletSplitTest.SplitDuringLongRunningTransaction Reviewers: timur, zdrudi, sergei Reviewed By: zdrudi Subscribers: hbhanawat, svc_phabricator, ybase Differential Revision: https://phorge.dev.yugabyte.com/D58578
| Commit: | 9c16bc7 | |
|---|---|---|
| Author: | Anton Rybochkin | |
| Committer: | Anton Rybochkin | |
[BACKPORT 2026.1][#31065] docdb: Deny split until the parent tablet vector index is post-split compacted Summary: **Background** After a tablet split, each child tablet's vector indexes inherit all vectors from the parent, including roughly half that no longer belong to the child's key range. Vector search must filter these irrelevant vectors via reverse-mapping lookups into RegularDB (two reads per vector), which is significantly slower than the vector search itself. If splits continue before the inherited vectors are compacted away, the ratio of irrelevant vectors grows with each generation, degrading search latency. To control this, tablet splits should be postponed until co-hosted vector indexes have completed post-split compaction. **Changes** 1. **Splits wait for vector index post-split compaction.** `Tablet::StillHasOrphanedPostSplitDataAbortable()` now also reports leftover vector index parent data, so a tablet is not split again (and the load balancer does not move it) until that data is gone. New advanced gflag `vector_index_require_parent_data_compacted_before_split` (default `true`) turns the wait off. The existing hidden gflag `vector_index_include_into_post_split_compaction` turns the vector index side off entirely. 2. **Scheduling a post-split compaction uses its own check.** `Tablet::NeedPostSplitCompaction()` answers a different question: is a post-split compaction still owed for this tablet? `TriggerPostSplitCompactionIfNeeded()` calls it instead of the split check. Otherwise setting `vector_index_require_parent_data_compacted_before_split=false` would also stop an interrupted vector index compaction from being retried when the tablet reopens. 3. **Each vector index tracks its own progress.** Two new `ConsensusFrontier` fields: `split_generation` and `split_min_chunk_serial_no`. Chunks with a serial number below the bound came from the parent tablet. At open time `InitFrontiers` compares the tablet's `split_generation` with the frontier's; if a new split happened it records `LastSerialNo() + 1` as the bound and flushes. After that, `ParentDataCompacted()` is simply `MinSerialNo() >= bound`, or true when there are no chunks. Nothing about vector index compaction is stored in tablet metadata. 4. **New `split_generation` in `KvStoreInfoPB`.** A tserver-local counter, parent + 1 in `CreateSplitChildMetadata`, which is how a vector index notices that a new split happened. Restored from the vector index manifests during bootstrap if a rollback dropped it -- see Upgrade/Rollback safety. 5. **Reverse mapping lookups now respect tablet key bounds.** `IndexReverseMappingReader::Fetch` returns nothing for a mapping whose ybctid is outside the tablet's range. That lets a split child's compaction drop the vectors that now belong to its sibling. Tombstones are still returned unchanged (will be fix in a follow-up revision). 6. **`parent_data_compacted` renamed to `rocksdb_parent_data_compacted`** in `KvStoreInfoPB`, `TabletStatusPB` and everywhere it is used, because it only covers RocksDB and is now easy to confuse with the vector index state. 7. **`CreateSubtablet` renamed to `CreateSplitChildTablet`** (and `CreateSubtabletMetadata` to `CreateSplitChildMetadata`), since both are only used for tablet splitting. 8. **RocksDB and vector indexes are compacted independently.** `TriggerManualCompactionIfNeeded` takes an `IncludeVectorIndexes` argument and skips whichever side is already done; `VectorIndexList::Compact` skips indexes that are already compacted. If the RocksDB compaction fails, the vector index compaction is skipped: its merge filter needs the obsolete reverse mappings to be gone first. 9. **New `vector_indexes_parent_data_compacted` in `TabletStatusPB`.** A diagnostic field for tests and the status page. Always reported by the tserver when the tablet is available, and never used as persistent state. **Upgrade/Rollback safety** All new fields are optional and default to pre-change behavior: `split_generation` defaults to 0 (matching a never-split tablet) and `split_min_chunk_serial_no` to 0 (read as "no inherited parent data"), so pre-existing split children are treated as having nothing to compact and the gate applies only to splits performed by the new binary; the renames keep their field numbers (`KvStoreInfoPB` 8, `TabletStatusPB` 18), so on-disk and wire formats are unchanged. Rollback drops `KvStoreInfoPB.split_generation` from the superblock while the vector index manifests keep it, so on the next upgrade `TabletBootstrap::MaybeUpdateMetaAfterTabletHasBeenOpened` restores it from the largest generation persisted across the tablet's vector indexes -- without that, the next split child would get generation 1, which is not greater than the persisted value, and would skip stamping its own parent data boundary. Merge conflicts: - src/yb/tablet/metadata.proto:136 - 2026.1 has no tiered storage, so the hunk's context (`tier_paths = 12` and the `TierPathPB` message) conflicted. Added only this commit's `split_generation` field and dropped that context. Kept field number 13, leaving 12 unused, so `KvStoreInfoPB` stays wire-compatible with master and a later tiered-storage backport still gets 12. - src/yb/tablet/tablet_metadata.cc:2158 - same cause: master's `child_rocksdb_dir` local exists only to feed `BuildTierPaths`, which 2026.1 does not have. Kept the branch's `set_rocksdb_dir(GetSubRaftGroupDataDir(...))` line and applied only this commit's two changes (`set_rocksdb_parent_data_compacted`, `set_split_generation`). The resolved commit's added and removed lines are byte-identical to the original commit. Original commit: ca6fd80065b / D52469 Test Plan: Jenkins: jobs: full ./yb_build.sh --cxx-test consensus_frontier-test --gtest_filter=ConsensusFrontierTest. SplitGenerationAndMinChunkSerialNo ./yb_build.sh --cxx-test tablet-split-test --gtest_filter=TabletSplitTest.StillHasOrphanedPostSplitDataVectorIndex ./yb_build.sh --cxx-test tablet-split-test --gtest_filter=TabletSplitTest.SplitGenerationOnSubtablet ./yb_build.sh --cxx-test tablet-split-itest --gtest_filter=TabletSplitSingleServerITest.PostSplitCompactionScheduledOnTabletOpen ./yb_build.sh --cxx-test pg_vector_index-test --gtest_filter=PgDistributedVectorIndexTest.SplitBlockedWithOrphanedPostSplitData/* ./yb_build.sh --cxx-test pg_vector_index-test --gtest_filter=PgDistributedVectorIndexTest.ManualSplitSimple/* ./yb_build.sh --cxx-test pg_vector_index-test --gtest_filter=PgDistributedVectorIndexTest.SplitGenerationRestoredAfterSuperblockReset/* ./yb_build.sh --cxx-test non_transactional_batch_writer-test --gtest_filter=NonTransactionalBatchWriterTest.ExternalApplyGatesVectorIndexFeed ./yb_build.sh --cxx-test tablet-split-test --gtest_filter=TabletSplitTest.SplitTablet ./yb_build.sh --cxx-test tablet-split-itest --gtest_filter=TabletSplitSingleServerITest.TabletServerOrphanedPostSplitData ./yb_build.sh --cxx-test pg_tablet_split-test --gtest_filter=PgTabletSplitTest.SplitDuringLongRunningTransaction Reviewers: timur, zdrudi, sergei Reviewed By: zdrudi Subscribers: ybase, svc_phabricator, hbhanawat Differential Revision: https://phorge.dev.yugabyte.com/D58577
| Commit: | d5dcc08 | |
|---|---|---|
| Author: | Nikhil Londhe | |
| Committer: | Nikhil | |
[BACKPORT 2026.1][#33494] YSQL: Fix upgrade unable to roll back or resume after a master restart Summary: **What is the issue?** `PREPARING` marks two unrelated things: a database being created for a client, and a live database whose new-version catalog a major YSQL upgrade is rebuilding. - A new master leader rebuilds its in-memory namespace maps from the sys catalog tablet. It reads `PREPARING` off disk with nothing to say which operation wrote the value. A live database holding user data therefore looks the same as an abandoned creation. - A live database read as an abandoned creation is left out of the by-name map and queued for deletion. The catalog rollback rejects `DELETING`. The upgrade then has no exit. - The on-disk value never changes, so every new master leader repeats the same reap decision. Restarting masters does not recover. A leader step-down with no process restart is enough to queue the deletion. **What needs fixing?** A namespace read in `PREPARING` needs different handling depending on which operation wrote it: - The upgrade's rebuild leaves the database usable, resolvable by name, and eligible for the catalog rollback. - An abandoned creation is marked for deletion and queued for cleanup, whether or not a major upgrade is in progress. - A half-built new-version catalog is marked failed. The catalog rollback accepts a failed new-version catalog. **How are we fixing it?** - The upgrade-in-progress flag cannot separate the two cases. The flag describes the cluster at the moment it is read. The `PREPARING` value on disk can be much older. A creation that failed before any upgrade began leaves `PREPARING` behind, and the loader reads the leftover namespace as a database the upgrade is rebuilding. - Add `NEXT_VER_PREPARING` to `SysNamespaceEntryPB.YsqlNextMajorVersionState`. The value means a database whose new major version catalog is being built. - Set `NEXT_VER_PREPARING` in `CreateNamespace` during a major YSQL upgrade. The write goes in the same upsert that sets `PREPARING`, so both values become durable together and cannot disagree. Restrict the write to YSQL databases: `CreateNamespace`'s YCQL path moves a namespace out of `PREPARING` without clearing `NEXT_VER_PREPARING`. - Branch on `NEXT_VER_PREPARING` in `NamespaceLoader::Visit`. Place the branch ahead of the fallthrough into the `FAILED` handler, which is where the reap decision is made. A namespace without `NEXT_VER_PREPARING` is reaped. - Load the namespace as `PREPARING`, and set `ysql_next_major_version_state` to `NEXT_VER_FAILED`. An interrupted catalog copy comes to rest in the same pair of values. Add the namespace to the id and name maps. - Add an upgrade test that kills the master leader inside the catalog copy window and asserts the rollback succeeds. Add a second that leaves a database in `PREPARING` from before the upgrade and asserts the database is still reaped. **Upgrade/Rollback safety:** `NEXT_VER_PREPARING` is written only by `CreateNamespace` while a major YSQL catalog upgrade is in flight, and only for YSQL databases. Every exit from the catalog copy clears it: `ProcessPendingNamespace` writes `NEXT_VER_RUNNING` on success and `NEXT_VER_FAILED` on failure, and the catalog rollback writes `NEXT_VER_RUNNING` and upserts. No database carries `NEXT_VER_PREPARING` once the upgrade finishes or is rolled back. Every yb-master runs the new major version for the duration of the catalog upgrade, so no yb-master reads the value without recognizing it. A yb-master on the older major version that wins an election mid-upgrade already aborts in `YsqlInitDBAndMajorUpgradeHandler::Load`, and the rollback clears the field before a downgrade is allowed. No AutoFlag, preview flag or test flag guards the change. `NEXT_VER_PREPARING` is not a format a peer process must understand; it is per-database, transient, and written and read inside a single catalog upgrade. Backport-through: 2025.2 Original commit: f4942f7ecafda14e1183028cd49d4dfeb37d4bb3 / D57574 Test Plan: ``` ./yb_build.sh release --cxx-test ysql_major_upgrade_rpcs-test ``` Also, `TestHardKillCatalogMigration` as created in D57358 ``` itest: TEST_SUITE=db_cross_functional_direct_upgrade, ITEST_REGEX=TestHardKillCatalogMigration, 2024.2.4.0 -> 2026.1.1.0-b88, KEEP_UNIVERSE_RESOURCES=true, RETRIES=0. ``` Reviewers: jason, fizaa, zdrudi Reviewed By: zdrudi Subscribers: ybase Differential Revision: https://phorge.dev.yugabyte.com/D58635
| Commit: | b077b6d | |
|---|---|---|
| Author: | Nikhil Londhe | |
| Committer: | Nikhil | |
[BACKPORT 2025.2][#33494] YSQL: Fix upgrade unable to roll back or resume after a master restart Summary: **What is the issue?** `PREPARING` marks two unrelated things: a database being created for a client, and a live database whose new-version catalog a major YSQL upgrade is rebuilding. - A new master leader rebuilds its in-memory namespace maps from the sys catalog tablet. It reads `PREPARING` off disk with nothing to say which operation wrote the value. A live database holding user data therefore looks the same as an abandoned creation. - A live database read as an abandoned creation is left out of the by-name map and queued for deletion. The catalog rollback rejects `DELETING`. The upgrade then has no exit. - The on-disk value never changes, so every new master leader repeats the same reap decision. Restarting masters does not recover. A leader step-down with no process restart is enough to queue the deletion. **What needs fixing?** A namespace read in `PREPARING` needs different handling depending on which operation wrote it: - The upgrade's rebuild leaves the database usable, resolvable by name, and eligible for the catalog rollback. - An abandoned creation is marked for deletion and queued for cleanup, whether or not a major upgrade is in progress. - A half-built new-version catalog is marked failed. The catalog rollback accepts a failed new-version catalog. **How are we fixing it?** - The upgrade-in-progress flag cannot separate the two cases. The flag describes the cluster at the moment it is read. The `PREPARING` value on disk can be much older. A creation that failed before any upgrade began leaves `PREPARING` behind, and the loader reads the leftover namespace as a database the upgrade is rebuilding. - Add `NEXT_VER_PREPARING` to `SysNamespaceEntryPB.YsqlNextMajorVersionState`. The value means a database whose new major version catalog is being built. - Set `NEXT_VER_PREPARING` in `CreateNamespace` during a major YSQL upgrade. The write goes in the same upsert that sets `PREPARING`, so both values become durable together and cannot disagree. Restrict the write to YSQL databases: `CreateNamespace`'s YCQL path moves a namespace out of `PREPARING` without clearing `NEXT_VER_PREPARING`. - Branch on `NEXT_VER_PREPARING` in `NamespaceLoader::Visit`. Place the branch ahead of the fallthrough into the `FAILED` handler, which is where the reap decision is made. A namespace without `NEXT_VER_PREPARING` is reaped. - Load the namespace as `PREPARING`, and set `ysql_next_major_version_state` to `NEXT_VER_FAILED`. An interrupted catalog copy comes to rest in the same pair of values. Add the namespace to the id and name maps. - Add an upgrade test that kills the master leader inside the catalog copy window and asserts the rollback succeeds. Add a second that leaves a database in `PREPARING` from before the upgrade and asserts the database is still reaped. **Upgrade/Rollback safety:** `NEXT_VER_PREPARING` is written only by `CreateNamespace` while a major YSQL catalog upgrade is in flight, and only for YSQL databases. Every exit from the catalog copy clears it: `ProcessPendingNamespace` writes `NEXT_VER_RUNNING` on success and `NEXT_VER_FAILED` on failure, and the catalog rollback writes `NEXT_VER_RUNNING` and upserts. No database carries `NEXT_VER_PREPARING` once the upgrade finishes or is rolled back. Every yb-master runs the new major version for the duration of the catalog upgrade, so no yb-master reads the value without recognizing it. A yb-master on the older major version that wins an election mid-upgrade already aborts in `YsqlInitDBAndMajorUpgradeHandler::Load`, and the rollback clears the field before a downgrade is allowed. No AutoFlag, preview flag or test flag guards the change. `NEXT_VER_PREPARING` is not a format a peer process must understand; it is per-database, transient, and written and read inside a single catalog upgrade. Backport-through: 2025.2 Original commit: f4942f7ecafda14e1183028cd49d4dfeb37d4bb3 / D57574 Test Plan: ``` ./yb_build.sh release --cxx-test ysql_major_upgrade_rpcs-test ``` Also, `TestHardKillCatalogMigration` as created in D57358 ``` itest: TEST_SUITE=db_cross_functional_direct_upgrade, ITEST_REGEX=TestHardKillCatalogMigration, 2024.2.4.0 -> 2026.1.1.0-b88, KEEP_UNIVERSE_RESOURCES=true, RETRIES=0. ``` Reviewers: jason, fizaa, zdrudi Reviewed By: zdrudi Subscribers: ybase Differential Revision: https://phorge.dev.yugabyte.com/D58636
| Commit: | 6a65cdc | |
|---|---|---|
| Author: | Dmitry Uspenskiy | |
| Committer: | Dmitry Uspenskiy | |
[#31547] YSQL: Individual Perform requests sequence for each worker Summary: `PgClientSession` executes incoming Perform requests in same order like they were sent by `pggate`. For this purpose each request protobuf has the `serial_no` field. In case `PgClientSession` received request with `serial_no` == `N` it waits to execute request with `serial_no` == `N - 1` first. Because both `pggate` and `PgClientSession` resides on single node it is not expected that some request might be lost (due to some transport issues) and block further requests. But because multiple `pggate` processes (Postgres workers) uses same `PgClientSession` and `serial_no` uses shared counter for next value it is possible that some request will be lost due to kill of `pggate` process prior to sending the message. Example: - main YSQL process prepares and sends request with serial_no == 1 - worker process prepares request with serial_no == 2 - main process kills worker process prior it sends the request - main YSQL process prepares and sends request with serial_no == 3 - PgClientSession receives requests with serial_no == 1 and serial_no == 3 only. Execution of request with serial_no == 3 is postponed till the processing of request with serial_no == 2, but it will never come. Solution is to handle request sequences for each worker separately. For this purpose serial_no + its predecessor serial_no is sent within protobuf. ``` message PgRequestSequenceNumPB { uint64 serial_no = 1; OptionalUint64PB predecessor_serial_no = 2; } ``` **Upgrade/Rollback safety:** The diff modifies the `pg_client.proto` API only. This API is used for communication only, so any modification of this API are safe for upgrade/rollback Test Plan: New unit tests are introduced ``` ./yb_build.sh --gtest_filter 'PgDebugReadRestartsTest.ParallelWorkersReadRestarts' -n 24 --tp 8 ./yb_build.sh --gtest_filter 'PgRequestSequencerTest.*' ``` Reviewers: sergei, pjain, patnaik.balivada, bkolagani Reviewed By: pjain Subscribers: yql, ybase, jason Tags: #jenkins-ready Differential Revision: https://phorge.dev.yugabyte.com/D55610
| Commit: | ddc5bb9 | |
|---|---|---|
| Author: | jhe | |
| Committer: | jhe | |
[BACKPORT 2026.1][#33972] xClusterDDLRepl: Fix ddl_queue error reporting Summary: Fixing ddl_queue handler to not report any errors when it hits TryAgain errors, since those are expected during normal operations (eg waiting for other pollers to catch up to our safe time). Adding a new error status REPLICATION_DDL_QUEUE_PAUSED to cover the case that ddl_queue does end up pausing due to a stuck DDL. Also propagating this error message up to the master leader, so this should now be visible in the master leader's /xcluster UI page / the get_xcluster_status API. This allows us to go to the master for all the info and not have to dig around tservers for that info. **Upgrade/Rollback safety:** Adding new error status and new error_detail field. Masters are upgraded before tservers, so will be able to process the old and new rpcs from their tservers. Sample image from master leader's /xcluster page on a target with a paused ddl_queue: {F566148} Original commit: 2315ce970a8 / D58317 Test Plan: ybd --cxx-test xcluster_consumer_replication_error-test --gtest_filter XClusterConsumerReplicationError.TestCollector ybd --cxx-test xcluster_ddl_replication-test --gtest_filter XClusterDDLReplicationTest.DDLQueuePollerPreservesOriginalError ybd --cxx-test xcluster_ddl_replication-test --gtest_filter XClusterDDLReplicationTest.DDLQueueWaitingForSafeTimeIsNotReplicationError Reviewers: xCluster, mlillibridge, yyan, hsunder, #db-approvers, neera.mital Reviewed By: yyan, #db-approvers, neera.mital Subscribers: ybase Differential Revision: https://phorge.dev.yugabyte.com/D58528
| Commit: | f4942f7 | |
|---|---|---|
| Author: | Nikhil Londhe | |
| Committer: | Nikhil Londhe | |
[#33494] YSQL: Fix upgrade unable to roll back or resume after a master restart Summary: **What is the issue?** `PREPARING` marks two unrelated things: a database being created for a client, and a live database whose new-version catalog a major YSQL upgrade is rebuilding. - A new master leader rebuilds its in-memory namespace maps from the sys catalog tablet. It reads `PREPARING` off disk with nothing to say which operation wrote the value. A live database holding user data therefore looks the same as an abandoned creation. - A live database read as an abandoned creation is left out of the by-name map and queued for deletion. The catalog rollback rejects `DELETING`. The upgrade then has no exit. - The on-disk value never changes, so every new master leader repeats the same reap decision. Restarting masters does not recover. A leader step-down with no process restart is enough to queue the deletion. **What needs fixing?** A namespace read in `PREPARING` needs different handling depending on which operation wrote it: - The upgrade's rebuild leaves the database usable, resolvable by name, and eligible for the catalog rollback. - An abandoned creation is marked for deletion and queued for cleanup, whether or not a major upgrade is in progress. - A half-built new-version catalog is marked failed. The catalog rollback accepts a failed new-version catalog. **How are we fixing it?** - The upgrade-in-progress flag cannot separate the two cases. The flag describes the cluster at the moment it is read. The `PREPARING` value on disk can be much older. A creation that failed before any upgrade began leaves `PREPARING` behind, and the loader reads the leftover namespace as a database the upgrade is rebuilding. - Add `NEXT_VER_PREPARING` to `SysNamespaceEntryPB.YsqlNextMajorVersionState`. The value means a database whose new major version catalog is being built. - Set `NEXT_VER_PREPARING` in `CreateNamespace` during a major YSQL upgrade. The write goes in the same upsert that sets `PREPARING`, so both values become durable together and cannot disagree. Restrict the write to YSQL databases: `CreateNamespace`'s YCQL path moves a namespace out of `PREPARING` without clearing `NEXT_VER_PREPARING`. - Branch on `NEXT_VER_PREPARING` in `NamespaceLoader::Visit`. Place the branch ahead of the fallthrough into the `FAILED` handler, which is where the reap decision is made. A namespace without `NEXT_VER_PREPARING` is reaped. - Load the namespace as `PREPARING`, and set `ysql_next_major_version_state` to `NEXT_VER_FAILED`. An interrupted catalog copy comes to rest in the same pair of values. Add the namespace to the id and name maps. - Add an upgrade test that kills the master leader inside the catalog copy window and asserts the rollback succeeds. Add a second that leaves a database in `PREPARING` from before the upgrade and asserts the database is still reaped. **Upgrade/Rollback safety:** `NEXT_VER_PREPARING` is written only by `CreateNamespace` while a major YSQL catalog upgrade is in flight, and only for YSQL databases. Every exit from the catalog copy clears it: `ProcessPendingNamespace` writes `NEXT_VER_RUNNING` on success and `NEXT_VER_FAILED` on failure, and the catalog rollback writes `NEXT_VER_RUNNING` and upserts. No database carries `NEXT_VER_PREPARING` once the upgrade finishes or is rolled back. Every yb-master runs the new major version for the duration of the catalog upgrade, so no yb-master reads the value without recognizing it. A yb-master on the older major version that wins an election mid-upgrade already aborts in `YsqlInitDBAndMajorUpgradeHandler::Load`, and the rollback clears the field before a downgrade is allowed. No AutoFlag, preview flag or test flag guards the change. `NEXT_VER_PREPARING` is not a format a peer process must understand; it is per-database, transient, and written and read inside a single catalog upgrade. Backport-through: 2025.2 Test Plan: ``` ./yb_build.sh release --cxx-test ysql_major_upgrade_rpcs-test ``` Also, `TestHardKillCatalogMigration` as created in D57358 ``` itest: TEST_SUITE=db_cross_functional_direct_upgrade, ITEST_REGEX=TestHardKillCatalogMigration, 2024.2.4.0 -> 2026.1.1.0-b88, KEEP_UNIVERSE_RESOURCES=true, RETRIES=0. ``` Reviewers: jason, fizaa, zdrudi Reviewed By: zdrudi Subscribers: ybase Differential Revision: https://phorge.dev.yugabyte.com/D57574
| Commit: | a76fe97 | |
|---|---|---|
| Author: | Aleksandr Malyshev | |
| Committer: | Aleksandr Malyshev | |
[BACKPORT 2026.1][PLAT-22410] YBA: Archive whole multi-line audit records in log retention Summary: zip_purge_yb_logs.sh extracts audit records out of the plain YSQL/YCQL logs into audit/<flavor>/ before gzipping the source, and those archives are what the audit retention window keeps. The extractor used a line-oriented "grep -F <marker>". A pgaudit (and YCQL) record spans multiple physical lines whenever its statement text contains newlines, and only the first line carries the AUDIT marker, so every continuation line was dropped and the retained audit archive held truncated records. Extract by record instead. awk tracks record boundaries with a record-start pattern and emits a whole record - the marker line plus its continuation lines up to the next start - whenever the record begins with the marker. The marker is still matched as a fixed string (index()), so grep -F semantics are preserved. The boundary is the log_line_prefix, which is a gflag (ysql_pg_conf_csv / log_line_prefix), so a hardcoded timestamp anchor would be wrong for a customized prefix. YBA derives the POSIX-ERE start pattern from the resolved prefix - the same prefix the collector's multiline receiver already uses - and ships it in log_cleanup_env through a new InstallOtelCollectorInput.ysqlAuditLineStartRegex field. The script falls back to a default-prefix anchor when the field is absent (an env file written by an older YBA). Interval quantifiers are loosened to '+' so the pattern works under both mawk and gawk. Keeping it correct when the gflag changes: GFlagsUpgrade already re-runs ManageOtelCollector when the collector is enabled, and that rebuilds the payload from the current gflags, so a log_line_prefix change regenerates the pattern on the node with no extra wiring. YCQL audit is glog, whose header is fixed and independent of log_line_prefix, so its boundary stays a constant in the script. The proto is the single source for the Go and Java stubs; both regenerate at build, so this diff carries only the .proto change, not the generated code. Original commit: 1820368d45ea9d00662deb1d48e77b76505ab58d / D58552 Original diff: https://phorge.dev.yugabyte.com/D58552 Test Plan: New Go tests exercise the real archive functions extracted from the template: cd managed/node-agent go test -tags testonly ./app/task/ -run TestArchive -v - 4 passed - TestArchiveKeepsMultiLineYsqlAuditRecords - TestArchiveKeepsMultiLineYcqlAuditRecords - TestArchiveNoAuditRecordsProducesNoArchive - TestArchiveKeepsMultiLineYsqlAuditRecordsForACustomLogLinePrefix, which feeds the pattern YBA generates for "%t [%p] %u@%d " through awk and checks a multi-line record archives whole. That is where the custom-prefix pattern is exercised against a real engine; the Java test pins the exact string the two share. New Java tests cover the ERE generator, including a custom (non-default) prefix. They assert the generated pattern as text rather than compiling it: it is POSIX ERE for awk, and java.util.regex is a different dialect that rejects the portable "[[]" spelling of a literal '[' outright, reading the inner bracket as a nested character class. An earlier revision of this diff did compile it and failed for exactly that reason. sbt "testOnly com.yugabyte.yw.common.audit.otel.OtelCollectorConfigGeneratorTest" - 47 passed - generateAuditLineStartEreForDefaultPrefix - generateAuditLineStartEreForCustomPrefix - re2ToPosixEreTranslatesPcreConstructs Manually verified end to end: sourced the archive functions against fixtures with default and custom prefixes and YCQL glog, confirmed multi-line records archive whole while query records and decoys are excluded, and confirmed the generated printf command writes a single-quoted log_cleanup_env line that sources back to the exact regex. Tested locally to make sure multi-line logs are being retained: ``` [yugabyte@ip-10-9-113/home/yugabyte/tserver/logs/audit/ysql/postgresql-2026-09-23_134750.log.audit.log.gzg.audit.log.gz 2026-09-23 14:08:22.722 UTC [14114] LOG: AUDIT: SESSION,1,1,ROLE,DROP ROLE,,,DROP USER tp_user,<none> 2026-09-23 14:08:56.827 UTC [14704] LOG: AUDIT: SESSION,1,1,ROLE,GRANT ROLE,,,"GRANT pg_read_all_data TO ""tp_user""",<none> 2026-09-23 14:29:10.871 UTC [26671] LOG: AUDIT: SESSION,1,1,DDL,CREATE TABLE,,,"create table test( id bigserial, data text );",<none> ``` Reviewers: vbansal Reviewed By: vbansal Subscribers: yugaware Differential Revision: https://phorge.dev.yugabyte.com/D58607
| Commit: | ce4968e | |
|---|---|---|
| Author: | Naorem Khogendro Singh | |
| Committer: | Naorem Khogendro Singh | |
[BACKPORT 2026.1][PLAT-22528] Clock skew input param acceptable_clock_skew_sec is not passed from YBA to node agent Summary: Some issues detected by AI. 1. Clock sync related params are not passed down to node agent. Not very important as there is a default but good to cover in case we want to change. 2. process_plain_files in zip_purge_yb_logs.sh.j2 does not use the input size parameter. Not a big issue as permitted_disk_usage_plain_kb is anyways used. 3. collect_metrics_wrapper.sh.j2 ignores the detected node-exporter. Some old universes may complain. Original commit: da866a053e69eeff268b2c298cb51cd7844cdd73 (D58224) Test Plan: Itests must pass. Otherwise, these are good to fix issues but should not change behavior Reviewers: amalyshev, hzare Reviewed By: hzare Subscribers: yugaware Differential Revision: https://phorge.dev.yugabyte.com/D58585
| Commit: | f522338 | |
|---|---|---|
| Author: | yusong-yan | |
| Committer: | yusong-yan | |
[BACKPORT 2025.2][#26678] xCluster: DDL Replication - Add support for vector indexing Summary: Original commit: b5ddf8d06f45243d00c33aa1549fad4cb869974d / D50522 (Note that D50522 already partially applied some changes) Background A vector index is created via an ADD TABLE change-metadata op and is colocated on each indexed table’s tablets. Writes to the vector index happen during apply-intent, we write to the vector index storage engine and to regular DB (including reverse mapping). Backfill runs on the tserver with MVCC snapshots(read time is the index creation time hybrid_time) backfill reads both committed intents and regularDB. Vector index metadata has a colocation_id but we don’t use it. Solution On the target, vector index writes happen when we apply external intents(replicated from the source) We skip the ADD TABLE CMOPS for the vector index on target. This is safe because backfill covers writes that happened before the index existed; it runs with read time equal to index creation time. Writes after the index creation on the target go through the normal apply-external-intent path and are applied to the vector index via VectorIndexesUpdater::Feed. We added delete_vector_ids in the CDC record and in CombineExternalIntents so that UPDATE/DELETE can tombstone old vector IDs on the target during external intent apply. Current backfill design can cause writes to be skipped When updating the vector index during apply intent we skip the row if commit_ht < vector_index.hybrid_time (vector index creation time). This check is to prevent we double-apply what backfill will already see.(Backfill can read both committed intent and regulardb.) This can cause issue on an xCluster target, the commit_ht from the source can be less than the target’s vector index hybrid_time in certain case. So the problem is the row is skipped in Feed and never appears in the target’s vector index. XClusterDDLReplicationTest.VectorIndexLateWriteAfterBackfillMissing new added c++ test is able to repro the issue The fix is on external-intents apply we do not apply the commit_ht < vector_index.hybrid_time() check on the target, so those writes are still feed to the vector index. The purpose of the check is to prevent double write, but backfill on target doesn't see committed intent, so it's safe to skip the intent. Upgrade/Rollback safety delete_vector_ids is an optional CDC field: old producers never send it and old consumers ignore it, so mixed versions are safe. Test Plan: ./yb_build.sh --cxx-test xcluster_ddl_replication-test --gtest_filter XClusterDDLReplicationTest.VectorIndex ./yb_build.sh --cxx-test xcluster_ysql-test --gtest_filter XClusterYsqlTest.VectorIndex ./yb_build.sh --cxx-test xcluster_ddl_replication-test --gtest_filter XClusterDDLReplicationTest.VectorIndexLateWriteAfterBackfillMissing Reviewers: jhe, xCluster, hsunder, zdrudi Reviewed By: zdrudi Subscribers: ybase Differential Revision: https://phorge.dev.yugabyte.com/D58461
| Commit: | 1ee535f | |
|---|---|---|
| Author: | skhilar | |
| Committer: | skhilar | |
[PLAT-21148] YBA: Support cross-cloud federated IAM for GCP VMs backing up to S3 Summary: S3-on-GCP cross-cloud federated IAM: a GCP DB node backs up to S3 using its own GCE identity, with no static AWS keys stored. This is the second direction of PLAT-21147 (GCS-on-AWS) and reuses its RPC, subtask and env-file plumbing. How it works: node-agent renders a credential_process script that exchanges the node's GCE identity token for temporary AWS credentials via STS AssumeRoleWithWebIdentity, and installs an AWS_PROFILE pointing at it. YBC picks it up from the federation env file. YBA performs the same exchange in-process for the paths it must run itself (preflight, backup deletion). | Area | Change | | --- | --- | | node-agent | `S3_ON_GCP` + `S3OnGcpConfig` on the existing ConfigureCloudFederation RPC (reserved field 6). New `aws_credential_process.sh`, built on curl + sed only - no awscli/jq/unzip, so airgap and on-prem need no package install. Profile written to `~/.aws/config` as a marker-bracketed managed block. | | Provider | `GCPCloudInfo`: `enableFederatedIam`, `federatedIamRoleArn`, `federatedIamAudience`, editable on an in-use provider. Resolved through `CloudInfoInterface`. | | Storage config | An S3 config is federated when `IAM_CONFIGURATION.CREDENTIAL_SOURCE=WEB_TOKEN`. Role ARN and audience are transient, resolved from the universe's provider at backup time, never persisted on the config. | | Task wiring | `createConfigureCloudFederationTasks` now accepts gcp; same trigger points as GCS-on-AWS, so nodes added later are covered. | | YBA-side | `AWSUtil`: in-process AssumeRoleWithWebIdentity, narrow validation skips, `USE_AWS_IAM` for YBC, region derived from `SIGNING_REGION` / host base. | | UI | Federated IAM toggle on the S3 storage config; the three fields on the GCP provider create/edit forms. | **Notes for reviewers** - The credential_process error path prints only the STS `Code`/`Message`, never the response body. The guard keys on "did we parse an AccessKeyId", not "did STS fail", so an unparseable success response would otherwise have written live credentials to stderr, which lands in the YBC log and from there in support bundles. - Script renamed from `gen_aws_credentials.sh`: it generates nothing, it implements the AWS credential_process contract. The name now matches the `credential_process` directive that points at it. - `StorageConfiguration.jsx` also sets `IAM_INSTANCE_PROFILE=true`. Deliberate back-compat: helpers predating federation still read that flag. - The validation skips are narrow by design. Only a metadata/identity failure is skipped; real S3 errors (NoSuchBucket, AccessDenied) still fail the precheck. - `createCredsMapYbc` derives the bucket region instead of calling `getBucketRegion`, which needs a signed client a federation config has no credentials for. YBC has no cross-region redirect handling, so a wrong region surfaces as an opaque HTTP 301. - Fixes a pre-existing gap: scheduled backups never had the federation identity stamped, so they became undeletable once their universe was gone. **Setup prerequisites** (non-obvious, both cost real debugging time) - The AWS role's trust policy must condition on `accounts.google.com:sub` (the service account's numeric unique ID). AWS maps the GCE token's `azp` claim onto `accounts.google.com:aud`, so conditioning `:aud` on the audience URL always fails with AccessDenied even though the token decodes correctly. - Universe VMs need a GCP service account attached. YBA's GCP provider does not attach one today. Test Plan: Unit: - `AWSUtilTest` - federation detection, the three validation skips, host-base region extraction (including that `s3-website-us-west-2` does not yield "website"), and creds-map region resolution. - `CustomerConfigValidatorTest` - a federation S3 config skips credential validation. - node-agent `configure_cloud_federation_test.go` - desired-state hash, S3 input validation, managed-block strip is inclusive of markers, idempotent, and a no-op when the file is absent. Manual (GCP universe, S3 bucket in us-west-2): - GCP provider: enable Federated IAM with role ARN + audience; save on an in-use provider. - S3 storage config: create with Federated IAM on; no access keys required, config saves. - Create a universe on that provider; ManageCloudFederation runs per node, and each node has `~/.yugabyte/federation/aws_credential_process.sh` plus the managed block in `~/.aws/config`. - Run `aws_credential_process.sh` on a node; it returns temporary STS credentials. - Backup to S3, then delete the backup (exercises the YBA-side in-process exchange). Reviewers: #yba-api-review!, nsingh, vkumar, yshchetinin, kkannan, amalyshev Reviewed By: kkannan, amalyshev Subscribers: svc_phabricator, nsingh, yugaware Differential Revision: https://phorge.dev.yugabyte.com/D57288
| Commit: | 1820368 | |
|---|---|---|
| Author: | Aleksandr Malyshev | |
| Committer: | Aleksandr Malyshev | |
[PLAT-22410] YBA: Archive whole multi-line audit records in log retention Summary: zip_purge_yb_logs.sh extracts audit records out of the plain YSQL/YCQL logs into audit/<flavor>/ before gzipping the source, and those archives are what the audit retention window keeps. The extractor used a line-oriented "grep -F <marker>". A pgaudit (and YCQL) record spans multiple physical lines whenever its statement text contains newlines, and only the first line carries the AUDIT marker, so every continuation line was dropped and the retained audit archive held truncated records. Extract by record instead. awk tracks record boundaries with a record-start pattern and emits a whole record - the marker line plus its continuation lines up to the next start - whenever the record begins with the marker. The marker is still matched as a fixed string (index()), so grep -F semantics are preserved. The boundary is the log_line_prefix, which is a gflag (ysql_pg_conf_csv / log_line_prefix), so a hardcoded timestamp anchor would be wrong for a customized prefix. YBA derives the POSIX-ERE start pattern from the resolved prefix - the same prefix the collector's multiline receiver already uses - and ships it in log_cleanup_env through a new InstallOtelCollectorInput.ysqlAuditLineStartRegex field. The script falls back to a default-prefix anchor when the field is absent (an env file written by an older YBA). Interval quantifiers are loosened to '+' so the pattern works under both mawk and gawk. Keeping it correct when the gflag changes: GFlagsUpgrade already re-runs ManageOtelCollector when the collector is enabled, and that rebuilds the payload from the current gflags, so a log_line_prefix change regenerates the pattern on the node with no extra wiring. YCQL audit is glog, whose header is fixed and independent of log_line_prefix, so its boundary stays a constant in the script. The proto is the single source for the Go and Java stubs; both regenerate at build, so this diff carries only the .proto change, not the generated code. Test Plan: New Go tests exercise the real archive functions extracted from the template: cd managed/node-agent go test -tags testonly ./app/task/ -run TestArchive -v - 4 passed - TestArchiveKeepsMultiLineYsqlAuditRecords - TestArchiveKeepsMultiLineYcqlAuditRecords - TestArchiveNoAuditRecordsProducesNoArchive - TestArchiveKeepsMultiLineYsqlAuditRecordsForACustomLogLinePrefix, which feeds the pattern YBA generates for "%t [%p] %u@%d " through awk and checks a multi-line record archives whole. That is where the custom-prefix pattern is exercised against a real engine; the Java test pins the exact string the two share. New Java tests cover the ERE generator, including a custom (non-default) prefix. They assert the generated pattern as text rather than compiling it: it is POSIX ERE for awk, and java.util.regex is a different dialect that rejects the portable "[[]" spelling of a literal '[' outright, reading the inner bracket as a nested character class. An earlier revision of this diff did compile it and failed for exactly that reason. sbt "testOnly com.yugabyte.yw.common.audit.otel.OtelCollectorConfigGeneratorTest" - 47 passed - generateAuditLineStartEreForDefaultPrefix - generateAuditLineStartEreForCustomPrefix - re2ToPosixEreTranslatesPcreConstructs Manually verified end to end: sourced the archive functions against fixtures with default and custom prefixes and YCQL glog, confirmed multi-line records archive whole while query records and decoys are excluded, and confirmed the generated printf command writes a single-quoted log_cleanup_env line that sources back to the exact regex. Tested locally to make sure multi-line logs are being retained: ``` [yugabyte@ip-10-9-113/home/yugabyte/tserver/logs/audit/ysql/postgresql-2026-09-23_134750.log.audit.log.gzg.audit.log.gz 2026-09-23 14:08:22.722 UTC [14114] LOG: AUDIT: SESSION,1,1,ROLE,DROP ROLE,,,DROP USER tp_user,<none> 2026-09-23 14:08:56.827 UTC [14704] LOG: AUDIT: SESSION,1,1,ROLE,GRANT ROLE,,,"GRANT pg_read_all_data TO ""tp_user""",<none> 2026-09-23 14:29:10.871 UTC [26671] LOG: AUDIT: SESSION,1,1,DDL,CREATE TABLE,,,"create table test( id bigserial, data text );",<none> ``` Reviewers: vbansal Subscribers: yugaware Differential Revision: https://phorge.dev.yugabyte.com/D58552
| Commit: | 96985b8 | |
|---|---|---|
| Author: | Brandon McNama | |
| Committer: | GitHub | |
[#10492] DocDB: Support maximum replicas per placement block (#32828) ## Summary Add an optional `max_num_replicas` constraint to placement blocks used by cluster, table, and tablespace policies. When omitted, the effective maximum is the placement replication factor, preserving existing behavior and metadata. Parse and validate the constraint for primary and read-replica placements, extend `yb-admin` placement syntax to `cloud.region.zone[:min[:max]]`, and enforce the bound during initial replica assignment and ongoing balancing. The maximum is a hard cap on placement decisions: table creation fails if the caps make a quorum infeasible, and the balancer never adds a replica above a block's maximum, except transiently during same-block replacement (for example, draining a blacklisted server), where the add precedes the remove. A placement that specifies any maximum must consist solely of fully-qualified, non-duplicate placement blocks, so every tserver is unambiguously attributed to one block; wildcard blocks cannot be combined with maxima. The balancer treats under-replication, over-replication, and blocks above their maximum as mutually exclusive handling states: under-replication is repaired first, blocks above their maximum are repaired add-before-remove, and removals for over-replicated tablets are steered to blocks above their maximum. The load balancer's time-to-balance estimate does not account for maxima and can be low when they are binding. Closes #10492 ## Upgrade/Rollback safety `PlacementBlockPB.max_num_replicas` is an additive optional proto2 field. No catalog migration or metadata backfill is required; absent values continue to mean the placement replication factor. Explicit maxima can only be written by new binaries (the new `yb-admin` syntax and tablespace JSON field); old binaries preserve and ignore the unknown field. In a mixed-version cluster an old master leader does not enforce stored maxima, so placement can transiently exceed a cap until a new-version leader repairs it on a subsequent balancer run — degraded but self-healing, with no correctness impact. Rollback remains wire-compatible: older binaries ignore the field and revert to legacy placement behavior without rewriting metadata. ## Test plan - [x] `build-support/lint.sh --rev upstream/master` - [x] `TablespaceParserTest.MaxNumReplicasParsing` - [x] `TablespaceParserTest.MaxNumReplicasValidation` - [x] `CatalogManagerUtilTest.MaxNumReplicasValidation` - [x] `TestPlacementInfoContainsPlacementInfo.TestMaxReplicas` - [x] `LoadBalancerRF5MaxReplicasMockedTest.RepairPlacementAboveMaximum` - [x] `LoadBalancerRF5MaxReplicasMockedTest.UnderReplicationPrioritizedOverMaxPlacement` - [x] `LoadBalancerMaxReplicasBlacklistMockedTest.BlacklistedMoveAllowedAtMaxPlacement` - [x] `LoadBalancerRF5MaxReplicasManyTabletsMockedTest.BalancedLoadStillRepairsOverMaxPlacement` - [x] `OptimalLoadDistributionTest.Slack` - [x] `OptimalLoadDistributionTest.SlackManyTservers` - [x] `load_balancer_mocked-test` (all tests) - [x] `load_balancer_placement_policy-test` (all tests, including `CreateTableRespectsMaxNumReplicas` and `MaxNumReplicasIsAHardCap`) - [x] `yb-admin_client-test` (all tests) - [x] `common_net-test` (all tests) - [x] `AdminCliTest.TestModifyPlacementPolicy` <!-- Reviewable:start --> - - - This change is [<img src="https://reviewable.io/review_button.svg" height="34" align="absmiddle" alt="Reviewable"/>](https://reviewable.io/reviews/yugabyte/yugabyte-db/32828) <!-- Reviewable:end --> --------- Co-authored-by: Craig Soules <craig.soules@shopify.com>
| Commit: | 7d5edec | |
|---|---|---|
| Author: | jhe | |
| Committer: | jhe | |
[BACKPORT 2025.2][#33972] xClusterDDLRepl: Fix ddl_queue error reporting Summary: Fixing ddl_queue handler to not report any errors when it hits TryAgain errors, since those are expected during normal operations (eg waiting for other pollers to catch up to our safe time). Adding a new error status REPLICATION_DDL_QUEUE_PAUSED to cover the case that ddl_queue does end up pausing due to a stuck DDL. Also propagating this error message up to the master leader, so this should now be visible in the master leader's /xcluster UI page / the get_xcluster_status API. This allows us to go to the master for all the info and not have to dig around tservers for that info. **Upgrade/Rollback safety:** Adding new error status and new error_detail field. Masters are upgraded before tservers, so will be able to process the old and new rpcs from their tservers. Sample image from master leader's /xcluster page on a target with a paused ddl_queue: {F566148} Original commit: 2315ce970a8 / D58317 Test Plan: ybd --cxx-test xcluster_consumer_replication_error-test --gtest_filter XClusterConsumerReplicationError.TestCollector ybd --cxx-test xcluster_ddl_replication-test --gtest_filter XClusterDDLReplicationTest.DDLQueuePollerPreservesOriginalError ybd --cxx-test xcluster_ddl_replication-test --gtest_filter XClusterDDLReplicationTest.DDLQueueWaitingForSafeTimeIsNotReplicationError Reviewers: xCluster, mlillibridge, yyan, hsunder Reviewed By: yyan Subscribers: ybase Differential Revision: https://phorge.dev.yugabyte.com/D58529
| Commit: | 850d5ed | |
|---|---|---|
| Author: | Ella Baron | |
| Committer: | Ella Baron | |
[#16670] DocDB: dist-trace: tag tablet, table and object-lock request fields Tablet and table ids on tserver, master, consensus, CDC and backup requests, plus the object-lock oids and mode on PgClient, become req.* attributes on the client span. The proxy drain calls TracingAttributes(req) from the .messages.cc, so protos that did not generate messages (master_client, master_admin, tserver_admin, cdc_service) move to MESSAGES_PROTO_FILES; without that the tag compiles but emits nothing. Design: https://docs.google.com/document/d/1nwb-V1dIa-8mLM5dvZwsDgoQ1ODLUA9TUFFqjXUeplo/edit?tab=t.0#heading=h.k8r70zn36j4j
| Commit: | 96e26e4 | |
|---|---|---|
| Author: | Ella Baron | |
[#16670] DocDB: dist-trace: (yb.rpc.trace) field option emits request fields as span attributes Tagging a request field with [(yb.rpc.trace) = {enabled: true}] makes gen_yrpc emit TracingAttributes() for the message and drain it onto the client span in the proxy stub, gated on an active trace context. The shared-memory exchange has no stub, so pggate drains by hand before starting the simulated client span. No production proto is tagged here; only rtest.proto, for the lwproto-test coverage of every supported field type. Design: https://docs.google.com/document/d/1nwb-V1dIa-8mLM5dvZwsDgoQ1ODLUA9TUFFqjXUeplo/edit?tab=t.0#heading=h.k8r70zn36j4j
| Commit: | 5590eb2 | |
|---|---|---|
| Author: | ellabaron-code | |
| Committer: | IshanChhangani | |
[BACKPORT 2025.2][#16670] DocDB: dist-trace: Propagate trace context across the RPC boundary (#33112) Summary: ## Summary Carry a distributed trace from an RPC caller to its callee. The request header gains an optional trace-context field; the caller writes its active span's context into it while the header is being written, and the callee decodes it and opens a server span parented under the caller's span, so one trace spans both processes. Short-circuited local calls take the parent directly from the outbound call, since no header exists. Decoding is best-effort — a malformed context is logged and the call proceeds untraced. Master and tserver also start their process-wide tracer here, without which none of these spans are exported. The wire field, its writer and its reader land together because none of them is observable on its own. ## Upgrade/Rollback safety `src/yb/rpc/rpc_header.proto` gains one new optional field: `RequestHeader.trace_context`, tag 8, of a new message type `TraceContextPB`. Tag 8 was previously unused on this branch, and matches the tag the field carries on master and on 2026.1, so the encoding is identical across all three. No existing field is renumbered, retyped, renamed, or removed, and no default changes. **Forward compatibility (new client → old server).** A new node only writes the field when distributed tracing is enabled (`otel_collector_traces_endpoint` set) *and* a trace is active for the call; otherwise the header bytes are exactly what they are today. When it is written, an old server's protobuf parser does not know tag 8 and `SkipField`s it into unknown fields, then handles the RPC normally. The field is purely observability metadata — no request semantics depend on it — so the old server simply serves an untraced call. **Backward compatibility (old client → new server).** The field is absent. `ParsedRequestHeader.trace_context` stays empty, `YBInboundCall::ParseFrom` leaves `parent_span_context_` invalid, and `CreateServerSpan` starts no server span. Same behavior as tracing being disabled. **Malformed input.** `ParseTraceContext` / `ToSpanContext` reject a zero or otherwise invalid context. Parse failure is logged and the call proceeds untraced — a bad or hostile trace context can never fail an RPC. **Rollback.** Nothing is persisted: the field exists only in in-flight RPC headers, so there is no on-disk or catalog state to migrate or clean up. Rolling a node back to a build without this change makes it fall into the "old server" case above from the first RPC onward; rolling the whole cluster back leaves no trace of the field. ## Backport notes Five files conflicted on cherry-pick — `outbound_call.cc` (two hunks), `rpc_header.proto`, `serialization.h`, `pg_client_session.cc` and `pg_wrapper-test.cc`. Each was resolved by taking only this commit's own additions and leaving out master-only surrounding code that 2025.2 does not carry. The applied change is byte-identical to the 2026.1 backport of the same commit (D58339) across all 16 files. Two resolutions worth calling out, both matching D58339: - `pg_client_session.cc` — kept the new `TEST_perform_async_error` flag; left out `TEST_shared_exchange_big_response_delay_ms`, which belongs to master-only code. - `pg_wrapper-test.cc` — no change on this branch, so the file is absent from the diff. Upstream this commit only renames `otel_batch_max_queue_size` to `otel_ysql_batch_max_queue_size` inside `ValidateCoDependentFlags`, a test 2025.2 does not have. That accounts for 16 files here against 17 upstream. Depends on D58438 Original commit: 1b14fded83d / #33112 ### Stack 1. #33116 — span/scope API and its direct callers 2. #33112 — propagate trace context across the RPC boundary 3. #33113 — propagate trace context over the PG↔tserver shared memory exchange 4. #33114 — carry trace context across thread hops 5. #33115 — propagate traceparent to the PG backend via the startup packet Test Plan: - [x] `./yb_build.sh release --cxx-test dist_trace-test` — including `TestRpcSpanReachesTabletServer`. Reviewers: asaha Reviewed By: asaha Differential Revision: https://phorge.dev.yugabyte.com/D58439
| Commit: | 1428a64 | |
|---|---|---|
| Author: | Naorem Khogendro Singh | |
| Committer: | Naorem Khogendro Singh | |
[BACKPORT 2.31.0.6390][PLAT-22528][PLAT-22441][PLAT-22433][PLAT-22369][PLAT-22369] Show previous stdout before task failure on UI Summary: Stderr does not have enough information on failure as most output lines are in stdout. Buffer the last N stdout lines into a ring buffer and spit out at the end on failure along with stderr. [PLAT-22433] Forward stderr or stdout of commands run by node agent to YBA if debug is enabled This change adds the ability to forward stdout and stderr to YBA. The forwarding already exists but command output was not routed. But, this is done only when debug mode is enabled to avoid too many unnecessary log lines for stdout. For stderr, it is always forwarded. We already have the global runtime config to dynamically turn it on ``` yb.node_agent.server.request_log_level ``` [PLAT-22441] node-agent clean_cores.sh renders empty and deletes nothing Wrong template substitution variable. Corrected and added more checks to fail with non-zero exit code. Also used AI to check for similar issues in other templates. This is the only hit. [PLAT-22528] Clock skew input param acceptable_clock_skew_sec is not passed from YBA to node agent Some issues detected by AI. 1. Clock sync related params are not passed down to node agent. Not very important as there is a default but good to cover in case we want to change. 2. process_plain_files in zip_purge_yb_logs.sh.j2 does not use the input size parameter. Not a big issue as permitted_disk_usage_plain_kb is anyways used. 3. collect_metrics_wrapper.sh.j2 ignores the detected node-exporter. Some old universes may complain. [PLAT-22368] Always enable num volume check before tserver or master is started Motivation Cloud START previously did a one-shot df | grep /mnt/dN count. That is brittle when volumes are slow to appear (e.g. after maintenance). Disk readiness needs a wait/retry path, shared with the existing systemd pre-start checks, without changing the normal clock-sync startup flow for on-prem. What changed 1. Split disk checks out of clock-sync.sh.j2 which still includes the disk-check.sh.j2 during generation to avoid duplicating the effort. New disk-check.sh.j2: wait for configured mount points (fstab/findmnt), then assert /mnt/dN count vs cloud_num_volumes. Existing clock-sync.sh.j2 Jinja-includes that template (as_include=true) and calls run_disk_checks before clock skew. Server Configure subtask still installs only clock-sync.sh (self-contained after render). Server Control subtask regenerates and runs standalone disk-check.sh such that it is already updated. 2. Server control (node-agent) Replaces the one-shot df check with: render disk-check.sh → run it more reliably. For CSPs, mount paths are derived as /mnt/d0…/mnt/d{N-1} from numVolumes (no mount paths on the RPC). Skips when numVolumes == 0 (on-prem / STOP). 3. YBA payload Cloud START always sets ServerControlInput.numVolumes from deviceInfo (no longer gated on checkVolumesAttached). Configure also sends ConfigureServerInput.numVolumes so rendered clock-sync.sh gets cloud_num_volumes. [PLAT-22368] Always enable num volume check before tserver or master is started Multiple Backports (Some are showing errors correctly for tasks). 1. 9a5d5fe4f9520b1a7c9f1a6fbe1edb73f54ec676 (https://phorge.dev.yugabyte.com/D57966) 2. da866a053e69eeff268b2c298cb51cd7844cdd73 (https://phorge.dev.yugabyte.com/D58224) 3. e439b095f9caab8fa5b47e70eaffb9c77314a225 (https://phorge.dev.yugabyte.com/D58129) 4. 6788cdedaafb0d34ccf398115619507114d3888f (https://phorge.dev.yugabyte.com/D58119) 5. 68595ac0e333dec5455b282364c744cadde63334 (https://phorge.dev.yugabyte.com/D57971) Test Plan: Manually tested by failing systemctl start. Manually tested. We have full visibility on the remote commands when debug is turned on ``` 2026-09-11T18:17:48.250Z [info] a78c4897-1c5b-4e5d-b735-f3884fed4d45 NodeAgentClient.java:649 [NodeAgentGrpcPool-3] com.yugabyte.yw.common.NodeAgentClient Download command curl -L -o /tmp/yugabyte-2025.2.0.0-b131-centos-x86_64.tar.gz https://s3.us-west-2.amazonaws.com/uploads.dev.yugabyte.com/nkhogen/yugabyte-2025.2.0.0-b131-centos-x86_64.tar.gz Dowloading software Running shell command for download-software % Total % Received % Xferd Average Speed Time Time Time Current Dload Upload Total Spent Left Speed 0 0 0 0 0 0 0 0 --:--:-- --:--:-- --:--:-- 0 1 457M 1 8182k 0 0 28.2M 0 0:00:16 --:--:-- 0:00:16 28.1M 0 457M 0 4243k 0 0 54.5M 0 0:00:08 --:--:-- 0:00:08 53.8M 3 457M 3 14.2M 0 0 46.0M 0 0:00:09 --:--:-- 0:00:09 46.0M 13 457M 13 63.6M 0 0 52.4M 0 0:00:08 0:00:01 0:00:07 52.4M 8 457M 8 38.2M 0 0 33.2M 0 0:00:13 0:00:01 0:00:12 33.1M 10 457M 10 50.3M 0 0 39.1M 0 0:00:11 0:00:01 0:00:10 39.1M 24 457M 24 111M 0 0 48.8M 0 0:00:09 0:00:02 0:00:07 48.8M 19 457M 19 88.4M 0 0 42.5M 0 0:00:10 0:00:02 0:00:08 42.5M 26 457M 26 119M 0 0 52.4M 0 0:00:08 0:00:02 0:00:06 52.3M 33 457M 33 152M 0 0 46.6M 0 0:00:09 0:00:03 0:00:06 46.6M 33 457M 33 154M 0 0 49.8M 0 0:00:09 0:00:03 0:00:06 49.7M 35 457M 35 164M 0 0 50.0M 0 0:00:09 0:00:03 0:00:06 50.0M 43 457M 43 199M 0 0 47.3M 0 0:00:09 0:00:04 0:00:05 47.3M 45 457M 45 208M 0 0 50.6M 0 0:00:09 0:00:04 0:00:05 50.6M 46 457M 46 214M 0 0 47.4M 0 0:00:09 0:00:04 0:00:05 47.4M 54 457M 54 248M 0 0 46.5M 0 0:00:09 0:00:05 0:00:04 47.5M 54 457M 54 248M 0 0 48.8M 0 0:00:09 0:00:05 0:00:04 48.7M 57 457M 57 264M 0 0 49.3M 0 0:00:09 0:00:05 0:00:04 49.5M 2026-09-11T18:17:53.942Z [info] a78c4897-1c5b-4e5d-b735-f3884fed4d45 NodeAgentClient.java:1389 [TaskPool-CreateUniverse(a6b098cd-48c3-4cae-8106-4ef84ae53586)-1] com.yugabyte.yw.common.NodeAgentClient Reconnecting to node agent [uuid=393bb29 ``` ``` Running step: post-install OpenSSL binary: /home/yugabyte/yb-software/yugabyte-2025.2.0.0-b131-centos-x86_64/bin/../bin/../bin/openssl FIPS module: /home/yugabyte/yb-software/yugabyte-2025.2.0.0-b131-centos-x86_64/bin/../bin/../lib/ossl-modules/fips.so HMAC : (Module_Integrity) : Pass SHA1 : (KAT_Digest) : Pass SHA2 : (KAT_Digest) : Pass SHA3 : (KAT_Digest) : Pass TDES : (KAT_Cipher) : Pass AES_GCM : (KAT_Cipher) : Pass AES_ECB_Decrypt : (KAT_Cipher) : Pass RSA : (KAT_Signature) : RNG : (Continuous_RNG_Test) : Pass Pass ECDSA : (PCT_Signature) : Pass ECDSA : (PCT_Signature) : Pass DSA : (PCT_Signature) : Pass TLS13_KDF_EXTRACT : (KAT_KDF) : Pass TLS13_KDF_EXPAND : (KAT_KDF) : Pass TLS12_PRF : (KAT_KDF) : Pass PBKDF2 : (KAT_KDF) : Pass SSHKDF : (KAT_KDF) : Pass KBKDF : (KAT_KDF) : Pass HKDF : (KAT_KDF) : Pass SSKDF : (KAT_KDF) : Pass X963KDF : (KAT_KDF) : Pass X942KDF : (KAT_KDF) : Pass HASH : (DRBG) : Pass CTR : (DRBG) : Pass HMAC : (DRBG) : Pass DH : (KAT_KA) : Pass ECDH : (KAT_KA) : Pass RSA_Encrypt : (KAT_AsymmetricCipher) : Pass RSA_Decrypt : (KAT_AsymmetricCipher) : Pass RSA_Decrypt : (KAT_AsymmetricCipher) : Pass INSTALL PASSED Running step: remove-older-release ``` Set it to a non-default value using the runtime config and see it was reflected. Also ran the script for errors. ``` [yugabyte@ip-10-9-124-125 bin]$ head -n 50 /home/yugabyte/bin/clean_cores.sh #!/usr/bin/env bash # # Copyright 2019 YugabyteDB, Inc. and Contributors # # Licensed under the Polyform Free Trial License 1.0.0 (the "License"); you # may not use this file except in compliance with the License. You # may obtain a copy of the License at # # https://github.com/YugaByte/yugabyte-db/blob/master/licenses/POLYFORM-FREE-TRIAL-LICENSE-1.0.0.txt set -euo pipefail print_help() { cat <<EOT Usage: ${0##*/} [<options>] Options: -n, --num_corestokeep <numcorestokeep> number of latest core files to keep (default: 5). -h, --help Show usage EOT } num_cores_to_keep=6 YB_CRASH_DIR="/home/yugabyte/cores/" if [[ -n "/home/yugabyte/cores" ]]; then YB_CRASH_DIR="/home/yugabyte/cores/" fi while [[ $# -gt 0 ]]; do case "$1" in -n|--num_corestokeep) num_cores_to_keep=$2 shift ;; -h|--help) print_help exit 0 ;; *) echo "Invalid option: $1" >&2 print_help exit 1 esac shift done if [[ -z "${num_cores_to_keep}" || ! "${num_cores_to_keep}" =~ ^[0-9]+$ ]]; then echo "num_cores_to_keep must be a non-negative integer" >&2 exit 1 fi [yugabyte@ip-10-9-124-125 bin]$ /home/yugabyte/bin/clean_cores.sh [yugabyte@ip-10-9-124-125 bin]$ echo $? 0 ``` Also tested by generating 3 coredumps as ``` kill -ABRT $(pgrep -f '/yb-master\b') ``` and changed the keep-number to 1 to see 2 are deleted. ``` [yugabyte@ip-10-9-111-206 bin]$ ./clean_cores.sh Deleting core file /home/yugabyte/cores/core_yb.1789274361.sigterm_loopxxx.19022.24784.gz Deleting core file /home/yugabyte/cores/core_yb.1789274391.sigterm_loopxxx.24919.25242.gz ``` Itests must pass. Otherwise, these are good to fix issues but should not change behavior Manually tested. 1. sudo unmount -l /mnt/d0 after entering a node in maintenance. 2. Exit maintenance failed after 5 mins timeout. ``` Failed to execute task {"nodeExporterUser":"prometheus","universeUUID":"0b361fca-7b67-4219-b6f7-019de6451216","enableYbc":true,"ybcSoftwareVersion":"2.2.0.4-b11","installYbc":false,"ybcInstalled":true,"encryptionAtRestConfig":{"encryptionAtRestEnabled":false,"opType":"UNDEFINED","type":"DATA_KEY"},"communicationPorts":{"masterHttpPort":7000,"masterRpcPort":7100,"tserverHttpPort":9000,"tserverRpcPort":9100,"ybControllerHttpPort":14000,"ybControllerrRpcPort":18018,"redisServerHttpPort":11000,"redisServerRpcPort":6379..., hit error: Code: 1, Error: exit status 1, State: Failed, Output: Running shell command for clock-sync Failed to run shell command for clock-sync: exit status 1 -. ``` 3. sudo mount /mnt/d0 and retry the task which succeeded. Screenshot is when /etc/fstab is modified to remove the data mount. {F563489} Reviewers: #yba-api-review!, amalyshev, skhilar, nbhatia, spothuraju, dshubin, hzare, yshchetinin, vbansal, sanketh, anijhawan Reviewed By: amalyshev, spothuraju Subscribers: nikhil, yugaware Differential Revision: https://phorge.dev.yugabyte.com/D58456
| Commit: | e57389b | |
|---|---|---|
| Author: | Samson Shaji | |
| Committer: | Samson Shaji | |
[#32248] DocDB: Add storage tier for tablespace DDL Summary: - Adds an optional `storage_tier` field to the `replica_placement` tablespace JSON which allows users to specify whether the table's replicas should be placed in SSD or HDD storage. - The storage tier is specified at the table-level, so a table cannot mix SSD and HDD placements. - Users can specify storage tiers for read replicas as well. - If the user does not specify a storage tier, replicas will be constrained to SSD storage only. SSD is always assumed to exist. Example: ``` { "num_replicas": 3, "storage_tier": "hdd", "placement_blocks": [ {"cloud": "cloud1", "region": "region1", "zone": "zone1", "min_num_replicas": 3} ] } ``` **Upgrade/Rollback safety:** - `storage_tier` is a new `optional` field on an existing message -- a standard additive proto change. Old readers ignore fields they don't know about; new readers see it unset (default: no tiering preference) when talking to old peers. - The parser only checks for known JSON keys via `HasMember` and does not reject unrecognized ones, so an old master reading a tablespace's `spcoptions` containing `"storage_tier":"..."` (created by a new master) silently ignores that key and behaves exactly as before -- `CREATE TABLE`/tablet placement is unaffected on rollback. - No persisted-protobuf migration concern for tablespaces specifically: a tablespace's `ReplicationInfoPB` is never serialized to disk as-is -- it's re-derived from the `spcoptions` text array in `pg_tablespace` on every read, so there's no stored-bytes compatibility issue there. - No AutoFlag/runtime flag is used to gate this, since the field is inert until a follow-up change starts consuming it -- there is no observable behavior difference on either binary version during a mixed-version upgrade or rollback. A guard flag should be revisited once the follow-up work actually acts on `storage_tier`. Test Plan: Added `StorageTierParsing` in `tablespace_parser-test.cc` and a new invalid case in `ReadReplicaPlacementParsingErrors`: - `StorageTierParsing` -- valid string is accepted and stored (`has_storage_tier()`/`storage_tier()`), omitted field leaves it unset rather than defaulted, non-string value is rejected with a clear type-mismatch error, empty string is rejected. - `ReadReplicaPlacementParsingErrors` (new case) -- `storage_tier` on a read-replica placement is rejected. Added a new block to `yb.orig.tablespaces.sql`/`.out` covering `CREATE TABLESPACE` with `storage_tier`: wrong-type rejection, empty-string rejection, and accept + persist in `spcoptions` + use in a `CREATE TABLE`. Ran the following commands to test: ``` ./yb_build.sh release --cxx-test tablespace_parser-test ``` ``` ./yb_build.sh release --java-test 'org.yb.pgsql.TestPgRegressTablespaces#testPgRegressTablespaces' ``` Reviewers: mhaddad, kfranz, sanketh Reviewed By: mhaddad, kfranz Subscribers: svc_phabricator, yql, ybase Differential Revision: https://phorge.dev.yugabyte.com/D56789
| Commit: | f68d2d7 | |
|---|---|---|
| Author: | ellabaron-code | |
| Committer: | IshanChhangani | |
[BACKPORT 2026.1][#16670] DocDB: dist-trace: Propagate trace context across the RPC boundary (#33112) Summary: ## Summary Carry a distributed trace from an RPC caller to its callee. The request header gains an optional trace-context field; the caller writes its active span's context into it while the header is being written, and the callee decodes it and opens a server span parented under the caller's span, so one trace spans both processes. Short-circuited local calls take the parent directly from the outbound call, since no header exists. Decoding is best-effort — a malformed context is logged and the call proceeds untraced. Master and tserver also start their process-wide tracer here, without which none of these spans are exported. The wire field, its writer and its reader land together because none of them is observable on its own. ## Upgrade/Rollback safety `src/yb/rpc/rpc_header.proto` gains one new optional field: `RequestHeader.trace_context` (tag 12), of a new message type `TraceContextPB`. No existing field is renumbered, retyped, renamed, or removed, and no default changes. **Forward compatibility (new client → old server).** A new node only writes the field when distributed tracing is enabled (`otel_collector_traces_endpoint` set) *and* a trace is active for the call; otherwise the header bytes are exactly what they are today. When it is written, an old server's protobuf parser does not know tag 12 and `SkipField`s it into unknown fields, then handles the RPC normally. The field is purely observability metadata — no request semantics depend on it — so the old server simply serves an untraced call. **Backward compatibility (old client → new server).** The field is absent. `ParsedRequestHeader.trace_context` stays empty, `YBInboundCall::ParseFrom` leaves `parent_span_context_` invalid, and `CreateServerSpan` starts no server span. Same behavior as tracing being disabled. **Malformed input.** `ParseTraceContext` / `ToSpanContext` reject a zero or otherwise invalid context. Parse failure is logged and the call proceeds untraced — a bad or hostile trace context can never fail an RPC. **Rollback.** Nothing is persisted: the field exists only in in-flight RPC headers, so there is no on-disk or catalog state to migrate or clean up. Rolling a node back to a build without this change makes it fall into the "old server" case above from the first RPC onward; rolling the whole cluster back leaves no trace of the field. Enabling or disabling `otel_collector_traces_endpoint` is likewise a pure runtime toggle with no persistent effect. ## Backport notes Two cherry-pick conflicts, both resolved by taking only this commit's own additions: - `src/yb/tserver/pg_client_session.cc` — kept the new `TEST_perform_async_error` flag; dropped `TEST_shared_exchange_big_response_delay_ms`, which is master-only content this branch does not carry. - `src/yb/yql/pgwrapper/pg_wrapper-test.cc` — no change on this branch. Upstream this commit only renames `otel_batch_max_queue_size` to `otel_ysql_batch_max_queue_size` inside `ValidateCoDependentFlags`, and 2026.1 has neither that test nor any other otel flag assertion in this file. Depends on D58338 Original commit: 1b14fded83d / #33112 ### Stack 1. #33116 — span/scope API and its direct callers 2. #33112 — propagate trace context across the RPC boundary 3. #33113 — propagate trace context over the PG↔tserver shared memory exchange 4. #33114 — carry trace context across thread hops 5. #33115 — propagate traceparent to the PG backend via the startup packet Test Plan: - [x] `./yb_build.sh release --cxx-test dist_trace-test` — full binary, green on this commit, including the new `TestRpcSpanReachesTabletServer`. Reviewers: asaha Reviewed By: asaha Differential Revision: https://phorge.dev.yugabyte.com/D58339
| Commit: | 2315ce9 | |
|---|---|---|
| Author: | jhe | |
| Committer: | jhe | |
[#33972] xClusterDDLRepl: Fix ddl_queue error reporting Summary: Fixing ddl_queue handler to not report any errors when it hits TryAgain errors, since those are expected during normal operations (eg waiting for other pollers to catch up to our safe time). Adding a new error status REPLICATION_DDL_QUEUE_PAUSED to cover the case that ddl_queue does end up pausing due to a stuck DDL. Also propagating this error message up to the master leader, so this should now be visible in the master leader's /xcluster UI page / the get_xcluster_status API. This allows us to go to the master for all the info and not have to dig around tservers for that info. **Upgrade/Rollback safety:** Adding new error status and new error_detail field. Masters are upgraded before tservers, so will be able to process the old and new rpcs from their tservers. Sample image from master leader's /xcluster page on a target with a paused ddl_queue: {F566148} Test Plan: ybd --cxx-test xcluster_consumer_replication_error-test --gtest_filter XClusterConsumerReplicationError.TestCollector ybd --cxx-test xcluster_ddl_replication-test --gtest_filter XClusterDDLReplicationTest.DDLQueuePollerPreservesOriginalError ybd --cxx-test xcluster_ddl_replication-test --gtest_filter XClusterDDLReplicationTest.DDLQueueWaitingForSafeTimeIsNotReplicationError Reviewers: xCluster, mlillibridge, yyan, hsunder Reviewed By: yyan Subscribers: ybase Differential Revision: https://phorge.dev.yugabyte.com/D58317
| Commit: | 1eba34a | |
|---|---|---|
| Author: | yusong-yan | |
| Committer: | yusong-yan | |
[#23495] xCluster: Add wal anchor stream for handling target DDL retry Summary: **Issue:** A create table ddl that fails on the target rolls back and deletes the stream, so the retry has nothing to replay from and replication gets stuck. However, even if we keep the stream, the source might already GC some of the WAL, and this is not acceptable as target should recieve WAL from the begining. How it breaks: * Source runs and commits the DDL * Target creates the table and adds it to replication. New stream, poller starts. * Poller moves the checkpoint forward, and the source can now drop the early WAL. * The DDL fails on the target before it commits. The transaction rolls back. * The rollback drops the target table, as well as delete the stream * The retry recreates the table and adds it again, but the source stream is gone so it failed with the stream not found. Another problem is the early WAL already be dropped. Solution: Use two streams per new table. On first add, the source creates a main stream and a WAL anchor stream, both at op id zero. The anchor is never consumed, so its checkpoint stays at zero and the existing WAL GC min across streams logic keeps WAL retained from the start for the life of the DDL. On failure, the target deletes only the main stream. On retry, `GetXClusterStreams(create_stream_if_missing)` sees the remaining WAL_ANCHOR and creates a fresh main stream. The retry finds it by source table id and replays WAL from zero. Cost: extra source WAL and one extra stream per in flight DDL. WAL anchor stream deletion: After the DDL commits, the target persists the source table id into its system catalog, a background task drains those ids, sends `DeleteXClusterWalAnchorStreams` to the source, and clears the id on success. With the source table id being persisted and bg task, the anchor stream deletion can survive lost RPCs and target leader crashes/changes, which would otherwise leak the anchor stream forever. Dropped tables are the exception: `BEGIN; CREATE t; DROP t` is ambiguous (commit vs rollback look the same), so the target does not delete that anchor; the source reaps it with the hidden table after `cdc_wal_retention_time_secs` tracked by https://github.com/yugabyte/yugabyte-db/issues/33446 Other cases: No tablet split while an anchor exists (splits require a bootstrapped stream, and the anchor never bootstraps). Colocated table creation reuses the existing stream, so no anchor is made. Upgrade/Rollback safety: Guarded by a new kExternal auto flag (`enable_xcluster_wal_anchor_stream_infra`) plus a runtime kill switch (`enable_xcluster_wal_anchor_stream`). If the target finalizes first, the old source makes no anchor and the new target finds none to delete. If the source finalizes first, auto flag incompatibility pauses replication (GetChanges rejected, no DDL reaches the target) and WAL sits at zero until the target finalizes. UI: Here is an example of how wal anchor stream looks like in `:7000/xcluster` {F532003} Others: For backfill, when target do its local backfill instead of relying on source sending the backfill write, source in this case advance its stream checkpoint at the end of the WAL once it mark backfill as complete. The hole here is, if target CREATE INDEX fail and retry, it recreate the table stream on the source and new checkpoint start at 0. This is tracked by https://github.com/yugabyte/yugabyte-db/issues/33443 Test Plan: Jenkins: jobs: full `XClusterWalAnchorStreamTest XClusterWalAnchorStreamTxnBlockTest` Reviewers: jhe, mlillibridge, xCluster, hsunder Reviewed By: jhe Subscribers: svc_phabricator, ybase Differential Revision: https://phorge.dev.yugabyte.com/D56239
| Commit: | ea7fcef | |
|---|---|---|
| Author: | Naorem Khogendro Singh | |
| Committer: | Naorem Khogendro Singh | |
[BACKPORT 2025.2][PLAT-21967][PLAT-22099][PLAT-22133][PLAT-22130] Implement the API handler node Agent custom cert upgrade in YBA Summary: This is in preparation to allow custom certs for Node Agent. The current issue is that server.key needs to be copied to YBA as it is used in signing the JWT and our cert storage for encryption in transit (EIT) which is going to be used for NA does not store it. Morever, customers are not willing to share it to YBA. This change creates a signer public and private key which are independently managed by YBA and is not tied to the custom cert. [PLAT-22099] Create a V2 API to upgrade node agent including custom cert Add the V2 APIs to invoke node agent upgrade which can also perform custom cert replacement. The implementation will follow. Sending this out to make the diff smaller. [PLAT-22133] Add alert for node agent certs getting expired Add two alerts for expiring certificates 1. For universe nodes. 2. Nodes in node_instance table but not in universe. [PLAT-22130] Implement the API handler node Agent custom cert upgrade in YBA This implements the API handler to upgrade node agents belonging to universes. The request payload exposes an option to perform only the cert upgrade. The following changes are made. 1. API handler is fully implemented. 2. YNP node-agent-provision.sh now allows the certificate_name to be specified. [PLAT-22151] Wire up node agent installation or reinstallation with universe custom certs This passes down the certificate from the rootCA of universe under the below cases 1. New node agent installation. 2. NodeToNode EIT is enabled for the universe. For rotation, the UI task (to follow) will handle. Here is the doc https://docs.google.com/document/d/1d_4DIjsu5EpoCgnrPUMrJ0syCAw14vvC3VrU-_0hQVM/edit?usp=sharing [PLAT-22289] Fix double upgrade and add more checks on certificate type Follow-up changes from https://phorge.dev.yugabyte.com/D57304 [PLAT-22290] Skip cert update for custom certs on node agent upgrade when YBA is upgraded For custom certs, it does not make sense to update the cert directory for node agents when it is upgraded because YBA has upgraded. This also a control parameter in DeployType to selectively choose only binary upgrade. This introduces three types of deployment. 1. FULL (both binary and cert). 2. BINARY_ONLY 3. CERTS_ONLY For CERTS_ONLY, neither the certs are generated (for self-managed) nor the new cert dir is set. No change needed in node agent golang code. [PLAT-22132] Add UI task for node agent upgrade for custom certs UI changes to add "Update Node Agent Certificate" task. 1. Add the task dialog and handlers. 2. Revamp /nodeagent page to link universe, provider and certificate to their respective pages. 3. Add certificate column and adjust the column widths. 4. Add tool tip for truncated names. [PLAT-22334] Re-register node-agent if certificate_name has changed in the config yaml This solves the issue of silently ignoring certificate_name if node agent is already running. Also, the config generator is enhanced to pull the data from node agent if it is available to generate the config correctly. [PLAT-22549] Node agent metrics are not removed after a node agent is deregistered or re-registered Guage metrics need to be cleared as they track state. Every scrap gets the last updated value irrespective of the node agent entry. [PLAT-22130] Implement the API handler node Agent custom cert upgrade in YBA conflict resolution https://phorge.dev.yugabyte.com/D57011 (c84c1624f9b9f16218265cfa1f135fbb78be01d0) https://phorge.dev.yugabyte.com/D57086 (8a6d091deb237fdf7b2e8a8bcffde66a0897867e) https://phorge.dev.yugabyte.com/D57267 (925fd3f2ac6f7db56b05d2b84b3957b03c4520bf) https://phorge.dev.yugabyte.com/D57304 (bca4d6523ac9217819bc575e4fdc42e6d1a07861) https://phorge.dev.yugabyte.com/D57490 (9d3741c1a1bc25468a6c8c1f85e401564a2081ef) https://phorge.dev.yugabyte.com/D57665 (03682bb9052b2327dbf4e50b5dbaeb93c5a0a5bc) https://phorge.dev.yugabyte.com/D57733 (9b23519e79171a61b360526db9361d963ee7e027) https://phorge.dev.yugabyte.com/D57470 (69d4bb77d3bd6565e4cacdb63815473022acabc9) https://phorge.dev.yugabyte.com/D57841 (db317cb1dc052f7789a349226ad56c36ca6d0a9e) https://phorge.dev.yugabyte.com/D58309 (fe85ac11758963e3b965abc22dafd2a9481eeebd) Test Plan: Manually tested. 1. Upgrade NA of nodes in a universe before this change. 2. New universe creation 3. Previous itests passed too. Build passed. Manually tested by changing the expiry threshold runtime config. {F550626} {F550627} 1. Tested NA upgrade multiple times - sef-signed to custom, custom to custom, custom to self-signed. 2. YNP provisioning for on-prem manual nodes. Some UTs are also added to test the components. UT added + manually tested. Manually tested (minor) 1. UTs added. 2. Manually did a. Reinstall b. Update API call to update only certs. c. Force upgrade by changing the version in the postgres record. c. Update API again. At each step, verified that the cert path remains the same in case of binary only. Tested manually. Video attached. {F553337} UTs added. Manually tested. 1. YNP config with first and without again for existing node agent. 2. Non-existing node-agent case. UT added. Verified manually by unregistering a node agent. Reviewers: #yba-api-review!, amalyshev, skhilar, amindrov, nbhatia, jmak, spothuraju, ayush.kushwaha Reviewed By: amalyshev Subscribers: yugaware Differential Revision: https://phorge.dev.yugabyte.com/D58326
| Commit: | 9d92ffb | |
|---|---|---|
| Author: | Naorem Khogendro Singh | |
| Committer: | Naorem Khogendro Singh | |
[BACKPORT 2025.2.6][PLAT-21967][PLAT-22099][PLAT-22133][PLAT-22130] Implement the API handler node Agent custom cert upgrade in YBA Summary: This is in preparation to allow custom certs for Node Agent. The current issue is that server.key needs to be copied to YBA as it is used in signing the JWT and our cert storage for encryption in transit (EIT) which is going to be used for NA does not store it. Morever, customers are not willing to share it to YBA. This change creates a signer public and private key which are independently managed by YBA and is not tied to the custom cert. [PLAT-22099] Create a V2 API to upgrade node agent including custom cert Add the V2 APIs to invoke node agent upgrade which can also perform custom cert replacement. The implementation will follow. Sending this out to make the diff smaller. [PLAT-22133] Add alert for node agent certs getting expired Add two alerts for expiring certificates 1. For universe nodes. 2. Nodes in node_instance table but not in universe. [PLAT-22130] Implement the API handler node Agent custom cert upgrade in YBA This implements the API handler to upgrade node agents belonging to universes. The request payload exposes an option to perform only the cert upgrade. The following changes are made. 1. API handler is fully implemented. 2. YNP node-agent-provision.sh now allows the certificate_name to be specified. [PLAT-22151] Wire up node agent installation or reinstallation with universe custom certs This passes down the certificate from the rootCA of universe under the below cases 1. New node agent installation. 2. NodeToNode EIT is enabled for the universe. For rotation, the UI task (to follow) will handle. Here is the doc https://docs.google.com/document/d/1d_4DIjsu5EpoCgnrPUMrJ0syCAw14vvC3VrU-_0hQVM/edit?usp=sharing [PLAT-22289] Fix double upgrade and add more checks on certificate type Follow-up changes from https://phorge.dev.yugabyte.com/D57304 [PLAT-22290] Skip cert update for custom certs on node agent upgrade when YBA is upgraded For custom certs, it does not make sense to update the cert directory for node agents when it is upgraded because YBA has upgraded. This also a control parameter in DeployType to selectively choose only binary upgrade. This introduces three types of deployment. 1. FULL (both binary and cert). 2. BINARY_ONLY 3. CERTS_ONLY For CERTS_ONLY, neither the certs are generated (for self-managed) nor the new cert dir is set. No change needed in node agent golang code. [PLAT-22132] Add UI task for node agent upgrade for custom certs UI changes to add "Update Node Agent Certificate" task. 1. Add the task dialog and handlers. 2. Revamp /nodeagent page to link universe, provider and certificate to their respective pages. 3. Add certificate column and adjust the column widths. 4. Add tool tip for truncated names. [PLAT-22334] Re-register node-agent if certificate_name has changed in the config yaml This solves the issue of silently ignoring certificate_name if node agent is already running. Also, the config generator is enhanced to pull the data from node agent if it is available to generate the config correctly. [PLAT-22549] Node agent metrics are not removed after a node agent is deregistered or re-registered Guage metrics need to be cleared as they track state. Every scrap gets the last updated value irrespective of the node agent entry. [PLAT-22130] Implement the API handler node Agent custom cert upgrade in YBA Original commits: https://phorge.dev.yugabyte.com/D57011 (c84c1624f9b9f16218265cfa1f135fbb78be01d0) https://phorge.dev.yugabyte.com/D57086 (8a6d091deb237fdf7b2e8a8bcffde66a0897867e) https://phorge.dev.yugabyte.com/D57267 (925fd3f2ac6f7db56b05d2b84b3957b03c4520bf) https://phorge.dev.yugabyte.com/D57304 (bca4d6523ac9217819bc575e4fdc42e6d1a07861) https://phorge.dev.yugabyte.com/D57490 (9d3741c1a1bc25468a6c8c1f85e401564a2081ef) https://phorge.dev.yugabyte.com/D57665 (03682bb9052b2327dbf4e50b5dbaeb93c5a0a5bc) https://phorge.dev.yugabyte.com/D57733 (9b23519e79171a61b360526db9361d963ee7e027) https://phorge.dev.yugabyte.com/D57470 (69d4bb77d3bd6565e4cacdb63815473022acabc9) https://phorge.dev.yugabyte.com/D57841 (db317cb1dc052f7789a349226ad56c36ca6d0a9e) https://phorge.dev.yugabyte.com/D58309 (fe85ac11758963e3b965abc22dafd2a9481eeebd) Test Plan: Manually tested. 1. Upgrade NA of nodes in a universe before this change. 2. New universe creation 3. Previous itests passed too. Build passed. Manually tested by changing the expiry threshold runtime config. {F550626} {F550627} 1. Tested NA upgrade multiple times - sef-signed to custom, custom to custom, custom to self-signed. 2. YNP provisioning for on-prem manual nodes. Some UTs are also added to test the components. UT added + manually tested. Manually tested (minor) 1. UTs added. 2. Manually did a. Reinstall b. Update API call to update only certs. c. Force upgrade by changing the version in the postgres record. c. Update API again. At each step, verified that the cert path remains the same in case of binary only. Tested manually. Video attached. {F553337} UTs added. Manually tested. 1. YNP config with first and without again for existing node agent. 2. Non-existing node-agent case. UT added. Verified manually by unregistering a node agent. Reviewers: #yba-api-review!, amalyshev, skhilar, amindrov, nbhatia, jmak, spothuraju, ayush.kushwaha Reviewed By: amalyshev Subscribers: yugaware Differential Revision: https://phorge.dev.yugabyte.com/D58324
| Commit: | b1a97b1 | |
|---|---|---|
| Author: | Naorem Khogendro Singh | |
| Committer: | Naorem Khogendro Singh | |
[BACKPORT 2026.1][PLAT-21967][PLAT-22099][PLAT-22133][PLAT-22130] Implement the API handler node Agent custom cert upgrade in YBA Summary: This is in preparation to allow custom certs for Node Agent. The current issue is that server.key needs to be copied to YBA as it is used in signing the JWT and our cert storage for encryption in transit (EIT) which is going to be used for NA does not store it. Morever, customers are not willing to share it to YBA. This change creates a signer public and private key which are independently managed by YBA and is not tied to the custom cert. [PLAT-22099] Create a V2 API to upgrade node agent including custom cert Add the V2 APIs to invoke node agent upgrade which can also perform custom cert replacement. The implementation will follow. Sending this out to make the diff smaller. [PLAT-22133] Add alert for node agent certs getting expired Add two alerts for expiring certificates 1. For universe nodes. 2. Nodes in node_instance table but not in universe. [PLAT-22130] Implement the API handler node Agent custom cert upgrade in YBA This implements the API handler to upgrade node agents belonging to universes. The request payload exposes an option to perform only the cert upgrade. The following changes are made. 1. API handler is fully implemented. 2. YNP node-agent-provision.sh now allows the certificate_name to be specified. [PLAT-22151] Wire up node agent installation or reinstallation with universe custom certs This passes down the certificate from the rootCA of universe under the below cases 1. New node agent installation. 2. NodeToNode EIT is enabled for the universe. For rotation, the UI task (to follow) will handle. Here is the doc https://docs.google.com/document/d/1d_4DIjsu5EpoCgnrPUMrJ0syCAw14vvC3VrU-_0hQVM/edit?usp=sharing [PLAT-22289] Fix double upgrade and add more checks on certificate type Follow-up changes from https://phorge.dev.yugabyte.com/D57304 [PLAT-22290] Skip cert update for custom certs on node agent upgrade when YBA is upgraded For custom certs, it does not make sense to update the cert directory for node agents when it is upgraded because YBA has upgraded. This also a control parameter in DeployType to selectively choose only binary upgrade. This introduces three types of deployment. 1. FULL (both binary and cert). 2. BINARY_ONLY 3. CERTS_ONLY For CERTS_ONLY, neither the certs are generated (for self-managed) nor the new cert dir is set. No change needed in node agent golang code. [PLAT-22132] Add UI task for node agent upgrade for custom certs UI changes to add "Update Node Agent Certificate" task. 1. Add the task dialog and handlers. 2. Revamp /nodeagent page to link universe, provider and certificate to their respective pages. 3. Add certificate column and adjust the column widths. 4. Add tool tip for truncated names. [PLAT-22334] Re-register node-agent if certificate_name has changed in the config yaml This solves the issue of silently ignoring certificate_name if node agent is already running. Also, the config generator is enhanced to pull the data from node agent if it is available to generate the config correctly. [PLAT-22549] Node agent metrics are not removed after a node agent is deregistered or re-registered Guage metrics need to be cleared as they track state. Every scrap gets the last updated value irrespective of the node agent entry. [PLAT-22130] Implement the API handler node Agent custom cert upgrade in YBA conflict resolution. https://phorge.dev.yugabyte.com/D57011 (c84c1624f9b9f16218265cfa1f135fbb78be01d0) https://phorge.dev.yugabyte.com/D57086 (8a6d091deb237fdf7b2e8a8bcffde66a0897867e) https://phorge.dev.yugabyte.com/D57267 (925fd3f2ac6f7db56b05d2b84b3957b03c4520bf) https://phorge.dev.yugabyte.com/D57304 (bca4d6523ac9217819bc575e4fdc42e6d1a07861) https://phorge.dev.yugabyte.com/D57490 (9d3741c1a1bc25468a6c8c1f85e401564a2081ef) https://phorge.dev.yugabyte.com/D57665 (03682bb9052b2327dbf4e50b5dbaeb93c5a0a5bc) https://phorge.dev.yugabyte.com/D57733 (9b23519e79171a61b360526db9361d963ee7e027) https://phorge.dev.yugabyte.com/D57470 (69d4bb77d3bd6565e4cacdb63815473022acabc9) https://phorge.dev.yugabyte.com/D57841 (db317cb1dc052f7789a349226ad56c36ca6d0a9e) https://phorge.dev.yugabyte.com/D58309 (fe85ac11758963e3b965abc22dafd2a9481eeebd) Test Plan: Manually tested. 1. Upgrade NA of nodes in a universe before this change. 2. New universe creation 3. Previous itests passed too. Build passed. Manually tested by changing the expiry threshold runtime config. {F550626} {F550627} 1. Tested NA upgrade multiple times - sef-signed to custom, custom to custom, custom to self-signed. 2. YNP provisioning for on-prem manual nodes. Some UTs are also added to test the components. UT added + manually tested. Manually tested (minor) 1. UTs added. 2. Manually did a. Reinstall b. Update API call to update only certs. c. Force upgrade by changing the version in the postgres record. c. Update API again. At each step, verified that the cert path remains the same in case of binary only. Tested manually. Video attached. {F553337} UTs added. Manually tested. 1. YNP config with first and without again for existing node agent. 2. Non-existing node-agent case. UT added. Verified manually by unregistering a node agent. Reviewers: #yba-api-review!, amalyshev, skhilar, amindrov, nbhatia, jmak, spothuraju, ayush.kushwaha Reviewed By: amalyshev Subscribers: yugaware Differential Revision: https://phorge.dev.yugabyte.com/D58328
| Commit: | 9a5d5fe | |
|---|---|---|
| Author: | Naorem Khogendro Singh | |
| Committer: | Naorem Khogendro Singh | |
[PLAT-22368] Always enable num volume check before tserver or master is started Summary: Motivation Cloud START previously did a one-shot df | grep /mnt/dN count. That is brittle when volumes are slow to appear (e.g. after maintenance). Disk readiness needs a wait/retry path, shared with the existing systemd pre-start checks, without changing the normal clock-sync startup flow for on-prem. What changed 1. Split disk checks out of clock-sync.sh.j2 which still includes the disk-check.sh.j2 during generation to avoid duplicating the effort. New disk-check.sh.j2: wait for configured mount points (fstab/findmnt), then assert /mnt/dN count vs cloud_num_volumes. Existing clock-sync.sh.j2 Jinja-includes that template (as_include=true) and calls run_disk_checks before clock skew. Server Configure subtask still installs only clock-sync.sh (self-contained after render). Server Control subtask regenerates and runs standalone disk-check.sh such that it is already updated. 2. Server control (node-agent) Replaces the one-shot df check with: render disk-check.sh → run it more reliably. For CSPs, mount paths are derived as /mnt/d0…/mnt/d{N-1} from numVolumes (no mount paths on the RPC). Skips when numVolumes == 0 (on-prem / STOP). 3. YBA payload Cloud START always sets ServerControlInput.numVolumes from deviceInfo (no longer gated on checkVolumesAttached). Configure also sends ConfigureServerInput.numVolumes so rendered clock-sync.sh gets cloud_num_volumes. Test Plan: Manually tested. 1. sudo unmount -l /mnt/d0 after entering a node in maintenance. 2. Exit maintenance failed after 5 mins timeout. ``` Failed to execute task {"nodeExporterUser":"prometheus","universeUUID":"0b361fca-7b67-4219-b6f7-019de6451216","enableYbc":true,"ybcSoftwareVersion":"2.2.0.4-b11","installYbc":false,"ybcInstalled":true,"encryptionAtRestConfig":{"encryptionAtRestEnabled":false,"opType":"UNDEFINED","type":"DATA_KEY"},"communicationPorts":{"masterHttpPort":7000,"masterRpcPort":7100,"tserverHttpPort":9000,"tserverRpcPort":9100,"ybControllerHttpPort":14000,"ybControllerrRpcPort":18018,"redisServerHttpPort":11000,"redisServerRpcPort":6379..., hit error: Code: 1, Error: exit status 1, State: Failed, Output: Running shell command for clock-sync Failed to run shell command for clock-sync: exit status 1 -. ``` 3. sudo mount /mnt/d0 and retry the task which succeeded. Screenshot is when /etc/fstab is modified to remove the data mount. {F563489} Reviewers: yshchetinin, skhilar, nbhatia, vbansal, sanketh, #yba-api-review!, anijhawan Reviewed By: yshchetinin Subscribers: nikhil, yugaware Differential Revision: https://phorge.dev.yugabyte.com/D57966
| Commit: | 1b14fde | |
|---|---|---|
| Author: | ellabaron-code | |
| Committer: | GitHub | |
[#16670] DocDB: dist-trace: Propagate trace context across the RPC boundary (#33112) ## Summary Carry a distributed trace from an RPC caller to its callee. The request header gains an optional trace-context field; the caller writes its active span's context into it while the header is being written, and the callee decodes it and opens a server span parented under the caller's span, so one trace spans both processes. Short-circuited local calls take the parent directly from the outbound call, since no header exists. Decoding is best-effort — a malformed context is logged and the call proceeds untraced. Master and tserver also start their process-wide tracer here, without which none of these spans are exported. The wire field, its writer and its reader land together because none of them is observable on its own. ## Upgrade/Rollback safety `src/yb/rpc/rpc_header.proto` gains one new optional field: `RequestHeader.trace_context` (tag 12), of a new message type `TraceContextPB`. No existing field is renumbered, retyped, renamed, or removed, and no default changes. **Forward compatibility (new client → old server).** A new node only writes the field when distributed tracing is enabled (`otel_collector_traces_endpoint` set) *and* a trace is active for the call; otherwise the header bytes are exactly what they are today. When it is written, an old server's protobuf parser does not know tag 12 and `SkipField`s it into unknown fields, then handles the RPC normally. The field is purely observability metadata — no request semantics depend on it — so the old server simply serves an untraced call. **Backward compatibility (old client → new server).** The field is absent. `ParsedRequestHeader.trace_context` stays empty, `YBInboundCall::ParseFrom` leaves `parent_span_context_` invalid, and `CreateServerSpan` starts no server span. Same behavior as tracing being disabled. **Malformed input.** `ParseTraceContext` / `ToSpanContext` reject a zero or otherwise invalid context. Parse failure is logged and the call proceeds untraced — a bad or hostile trace context can never fail an RPC. **Rollback.** Nothing is persisted: the field exists only in in-flight RPC headers, so there is no on-disk or catalog state to migrate or clean up. Rolling a node back to a build without this change makes it fall into the "old server" case above from the first RPC onward; rolling the whole cluster back leaves no trace of the field. Enabling or disabling `otel_collector_traces_endpoint` is likewise a pure runtime toggle with no persistent effect. ## Test plan - [x] `./yb_build.sh release --cxx-test dist_trace-test` — full binary, green on this commit, including the new `TestRpcSpanReachesTabletServer`. `TestRpcSpanReachesTabletServer` runs a `SELECT` under a known traceparent and asserts that the tserver's inbound `rpc yb.tserver.PgClientService.Perform` span lands in that trace as a child of the ysql backend's outbound span, with the expected `rpc.service`/`rpc.method` attributes. The new `WaitForRemoteChildSpan` collector helper does the caller/callee pairing check. Mixed-version wire behavior is argued from the encoding rather than exercised by a test — see the Upgrade/Rollback section. An old peer `SkipField`s the unknown tag, and the field is only written when tracing is on and a trace is active. ### Stack 1. #33116 — span/scope API and its direct callers 2. #33112 — propagate trace context across the RPC boundary 3. #33113 — propagate trace context over the PG↔tserver shared memory exchange 4. #33114 — carry trace context across thread hops 5. #33115 — propagate traceparent to the PG backend via the startup packet
| Commit: | bae5985 | |
|---|---|---|
| Author: | ellabaron-code | |
| Committer: | GitHub | |
[#33837] YSQL: Report index backfill failures with the PostgreSQL SQLSTATE instead of XX000 (#33748) ## Summary When a DDL from another session bumps a table's DocDB schema version while an index backfill is running, the BACKFILL INDEX statement fails with "schema version mismatch". This is unrelated to the preview concurrent-DDL feature: with `ysql_enable_concurrent_ddl` / object locking enabled the DDL waits on the table lock, so the race cannot happen, and strictly serial DDLs cannot hit it either since CREATE INDEX only returns after the backfill completes. The exposure is the default configuration (feature off), where nothing stops a second session from running a DDL against the table mid-backfill. The error is flattened to plain text on its way to the master (`IndexInfoPB.backfill_error_message`), and the YBClient inside the tserver hosting the CREATE INDEX backend re-raised it as an `Aborted` status with no error codes attached (`Aborted` is the `Status` code, not `TransactionErrorCode::kAborted`, and plays no part in the resulting SQLSTATE). The session running CREATE INDEX therefore sees XX000 (`internal_error`) instead of the retryable 40001 (`serialization_failure`) that every other schema-mismatch path (read and write DML) already delivers. Fix: propagate the real failure status end-to-end instead of re-tagging by message text: - The tserver's `QueryPostgresToDoBackfill` keeps the PostgreSQL error code from the libpq error (e.g. 40001 for schema version mismatch) on the status it returns, instead of dropping it when rebuilding the status. - The master persists the full status in a new `IndexInfoPB.backfill_status` field, alongside `backfill_error_message`, which stays populated for old readers. - `YBClient::Data::IsBackfillIndexInProgress` re-raises the persisted status (as `Aborted`, keeping its error codes), so pggate's `FetchErrorCode` maps the schema mismatch to 40001. Against an older master that doesn't populate `backfill_status`, it falls back to the message string as before. - The backend's unique-violation message rewrite (`YBCMakeStatusErrorData`) now skips statuses without a relation OID. Backfill duplicate-key failures now reach it with error code 23505 but no relation OID attached, and the rewrite crashed the backend dereferencing the missing relation. Such statuses already carry a fully formatted PG error message, which is now raised unchanged — so CREATE UNIQUE INDEX failing on duplicate data surfaces as 23505 (`unique_violation`, matching vanilla PostgreSQL) instead of XX000. ## Upgrade/Rollback safety `IndexInfoPB.backfill_status` is a new optional field. - New master, old tserver/client: old readers ignore the field; `backfill_error_message` is still populated, so behavior is unchanged. - Old master, new client: `backfill_status` is never set; the client uses the `backfill_error_message` fallback (pre-change behavior, XX000). - Rollback: the field is ignored again and `backfill_error_message` remains authoritative. No migration needed either way. ## Test plan - [x] `PgSchemaVersionMismatchBackfillTest.BackfillSurfacesAsSerializationFailure` — new regression test; fails with XX000 without the fix. Passed 20 consecutive iterations. - [x] `pg_index_backfill-test` — full suite. Four duplicate-key tests caught the backend crash and pass with the `YBCMakeStatusErrorData` guard. - [x] `pg_op_buffering-test` — covers the DML duplicate-key path through the changed unique-violation rewrite. - [x] `cassandra_cpp_driver-test` — YCQL backfill failure path through the changed master code.
| Commit: | 1b46764 | |
|---|---|---|
| Author: | Fizaa Luthra | |
[pg19] merge: master commit c25b05ab5ee21ed0e73e0fa433f6f31bea7ec5b7 into pg19 Summary: Merge: master commit c25b05ab5ee21ed0e73e0fa433f6f31bea7ec5b7 into pg19. Merge base: ad7b6d8dd7cf0220a4a84fd70d3a78fdca9c6368. Merged-into side: origin/pg19 (a63dd7b1893). Attribution key: - "master commit ..." = YB master-branch commit, i.e. what independently changed on master after the pg19 fork point. - "upstream PG commit ..." = upstream PostgreSQL commit whose change is on the pg19 branch via the initial merge (YB pg19 commit 4c01ae04c6d8e5487679d2079e44a067a66febcd merged upstream tip 90630ec42939d074ecc7b6b959b48252eed32646). - "YB pg19 commit ..." = YB-authored commit on the pg19 branch (initial-merge resolutions and follow-up pg19 fixes). - src/lint/upstream_repositories.csv: - src/postgres row: - master commit 62a60a60f20ef1f18125532858b76ad315006731 updated the src/postgres pin to c6778c409141837799395e8ec73374427b05df14. - YB pg19 branch pins src/postgres at 6c32e69c415aaf1132b357c5559a31e0da695c7c. - Kept YB pg19's pin. - src/postgres/src/backend/utils/adt/jsonb.c + src/postgres/src/include/utils/jsonb.h: - yb_datum_to_jsonb_non_string_scalar_value (new at end of file) / its extern: - master commit d9add44123a7ca45bfd803c958b159c5e0d81b78 added the helper plus its extern. - upstream PG commit 3c152a27b06313fe27bd47079658f928e291986b replaced JsonbTypeCategory/jsonb_categorize_type with JsonTypeCategory/json_categorize_type in jsonfuncs.h; upstream PG commit b22391a2ff7bdfeff4438f7a9ab26de3e33fdeff renamed the 6-arg static datum_to_jsonb to datum_to_jsonb_internal; upstream PG commit 0986e95161cec929d8f39c01e9848f34526be421 moved JsonbInState into jsonb.h and changed pushJsonbValue to take a JsonbInState * and return void. YB pg19 has JsonbUnquote at the same end-of-file position where master appended the helper. - Kept both; ported the helper onto PG19's API (json_categorize_type with is_jsonb=true, JSONTYPE_BOOL/JSONTYPE_NUMERIC, datum_to_jsonb_internal). The parameter changed from master's JsonbParseState **pstate to JsonbInState *result: on PG19 the caller owns the JsonbInState, so the helper hands it straight to datum_to_jsonb_internal instead of shuffling parseState in and out of a local. Its caller yb_index_check.c auto-merged and was adapted to match (see the build-fix entries below). - src/postgres/src/include/utils/backend_status.h: - PgBackendStatus tail: - master commit f576915851219317ae42f8049b0a10f196494982 added yb_st_cm_client_addr / yb_st_cm_client_port / yb_st_cm_client_hostname plus the yb_pgstat_set_ycm_client_info extern. - YB pg19 commit 4c01ae04c6d8e5487679d2079e44a067a66febcd deleted the blank line directly above "} PgBackendStatus;", the same spot master's fields are appended to. - Took master's addition. backend_status.c auto-merged. - src/postgres/src/test/isolation/isolationtester.h: - PermutationStepBlockerType: - master commit 30773e0f6ea3688b8313dd5c194d904a0b6f84a8 added PSB_YB_NEVER_WAITS. - upstream PG added a trailing comma to the last enumerator. - Took master's enumerator with PG19's trailing-comma style. - src/postgres/src/bin/pg_upgrade/util.c: - prep_status: - master commit 1d5faf373ef0cc7beedc472a3973ea4b26d05f6c changed the format to keep two spaces between a long message and the appended status. - upstream PG commit 7652353d87a6753627a6b6b36d7acd68475ea7c7 introduced PG_REPORT_NONL and switched this pg_log call to it. - Combined: PG_REPORT_NONL with master's format string. - src/postgres/src/backend/nodes/copyfuncs.c: - _copyIndexStmt: - master commit 23b20dcac6c43bd2995b7c4e51f6fc34f4f34d78 added COPY_SCALAR_FIELD(yb_index_old_relfilenode). - upstream PG commit 2be87f092a2ac786264b2020797aafa837de5a8e made copyfuncs.c generated; the handwritten parse-node copy functions no longer exist. - Took the PG19 side. The field itself auto-merged into parsenodes.h, so the generated copy covers it. - src/postgres/yb-extensions/yb_pg_metrics/yb_pg_metrics.c: - ybpgm_ExecutorEnd top-level check: - master commit d9c6a4f783e7f5617f5a8c3255802fe89f2bf54b added !ybpgm_IsTserverInternalConn() to the guard (helper, include and the ProcessUtility site auto-merged). - YB pg19 commit 4c01ae04c6d8e5487679d2079e44a067a66febcd commented out the queryDesc->totaltime term under a YB_TODO_PG19MERGE because QueryDesc.totaltime no longer exists. - Combined: kept the YB_TODO_PG19MERGE and the commented-out totaltime term, added master's new term. - src/yb/yql/pgwrapper/pg_catalog_perf-test.cc: - AfterCacheRefreshRPCCountOnSelectMinPreload: - master commit 47d533d0fc401d7c4f8769d40df04295ba5afa83 raised the expected count 13 -> 14. - YB pg19 commit raised it 13 -> 15 (pg19 catalogs are bigger). - The two deltas have independent causes, so set it to 16. Every other master hunk in this file auto-merged. - AfterCacheRefreshRPCCountOnInsertMinPreload (auto-merged; CI fix): - master commit 47d533d0fc401d7c4f8769d40df04295ba5afa83 and YB pg19 commit ef2c04d3fad65a680c67ca97a5c3c83bd0976a0a both raised this count 6 -> 7 independently, so it auto-merged to 7 with no conflict. - CI measured 8 (pg_catalog_perf-test.cc:290). Same independent-delta arithmetic as OnSelectMinPreload above; set to 8. - AfterCacheRefreshRPCCountOnSelectWithAggregates{,Preload} (auto-merged; CI fix): - master commit 6d4cdf43baa19e84826af8720f7ac65accbd9f65 added these two tests with counts measured against master's catalog (16 and 12). - pg19's catalogs are larger, so each costs one extra paging-continuation read; CI measured 17 and 13. Set to 17 and 13. - src/postgres/src/backend/utils/time/snapmgr.c: - forward-declaration block after SnapshotResetXmin: - master commit 7b100f3d47063801a8fc9185f8cf9ca500d247b1 added five static forward decls (YbResetReadPoint, YbGetOldestReadPointHandle, YbNoteReadPointAdded, YbNoteReadPointRemoved, YbPublishOldestReadPointIfChanged); master commit f6bff4a024d3ae52052b894ce27207146b51f06c added the "/* YB declarations */" comment above them. - upstream PG commit b8bff07daa85c837a2747b4d35cd5a27e73fb7b2 added the ResourceOwner callback descriptor and the Remember/Forget wrappers at the same spot. - Kept both, PG's block first. - InvalidateCatalogSnapshot tail: - master commit 7b100f3d47063801a8fc9185f8cf9ca500d247b1 added YbNoteReadPointRemoved + YbPublishOldestReadPointIfChanged after SnapshotResetXmin; master commit f6bff4a024d3ae52052b894ce27207146b51f06c renamed the local read_point to yb_read_point. - upstream PG commit bc32a12e0db2df203a9cb2315461578e08568b9c added INJECTION_POINT("invalidate-catalog-snapshot-end") at the same spot. - Kept both, YB's calls first and the injection point last (it marks the end of the function). - src/postgres/src/backend/access/yb_access/yb_scan.c + src/postgres/src/backend/access/nbtree/nbtutils.c: - array-element sort in the SAOP bind path: - master commit b4235398d2dacba7a59cda4b0a607e681d6c8b83 moved the YB opfamily-comparator fallback out of nbtutils.c's _bt_sort_array_elements into a new self-contained ybSortAndUniqArrayElements in yb_scan.c, and made _bt_sort_array_elements static again. - upstream PG commit 597b1ffbf12352a3863a894f16741864aaf2242f moved _bt_sort_array_elements (and the array preprocessing) out of nbtutils.c into nbtpreprocesskeys.c with a different signature; YB pg19 commit 4c01ae04c6d8e5487679d2079e44a067a66febcd had parked the YB call under #if 0 with a YB_TODO_PG19MERGE. - Took master's ybSortAndUniqArrayElements call and dropped the YB_TODO_PG19MERGE, which master's rewrite makes moot. nbtutils.c: took the PG19 side; master's deletions (the YB fallback and the YB includes) are already absent there because PG19 relocated the function, and nbtpreprocesskeys.c never carried the fallback. - src/postgres/src/backend/catalog/index.c: - YBCCreateIndex call in index_create: - master commit 23b20dcac6c43bd2995b7c4e51f6fc34f4f34d78 passed yb_index_old_relfilenode instead of InvalidOid for oldRelfileNodeId. - upstream PG commit 23382b0f8b21e3f5330d765d1abfcef58d086111 renamed classObjectId to opclassIds. - Combined: master's argument with PG19's parameter name. - src/postgres/src/backend/commands/indexcmds.c: - DefineIndex, after YbDefineIndexHelper: - master commit 7dd192f0d83e7a4034f17478388c4a317413ab38 added the object-locking phase (relcache invalidation, commit/restart of the DDL txn, YbWaitForBackendsCatalogVersion); master commit d1cd3a64db8777a55b653a5baf9b97bbf7f06220 appended the LockRelationOid(ShareUpdateExclusiveLock) re-acquire at the end of it. - upstream PG commit 23382b0f8b21e3f5330d765d1abfcef58d086111 renamed DefineIndex's relationId parameter to tableId. - Took master's block with relationId rewritten to tableId. - src/postgres/src/backend/tcop/postgres.c: - include block: - master commit 6fb1464700d5cc212fdd04b21e16b21aed4a26a1 added common/pg_yb_conn_mgr_protocol.h. - YB pg19 has executor/executor.h at the same position. - Kept both in sorted order. - exec_parse_message call in PostgresMain: - master commit 6fb1464700d5cc212fdd04b21e16b21aed4a26a1 replaced the yb_firstchar argument with yb_echo/yb_echo_len. - YB pg19 commit 4c01ae04c6d8e5487679d2079e44a067a66febcd had rewritten the trailing comment for pg_fallthrough. - Took master's argument list; the comment it replaced is gone with the argument. Every other hunk of the ConnMgr refactor auto-merged. - src/postgres/src/backend/utils/adt/pgstatfuncs.c: - pg_stat_get_activity client address/port/hostname: - master commit f576915851219317ae42f8049b0a10f196494982 put the whole upstream client-address block behind an "if (beentry->yb_st_cm_client_addr[0] != '\0')" branch that reports the conn-mgr logical client instead. - upstream PG commit e8d59294282bd01bc06f2af13c79b9155024a917 replaced the memset/memcmp against a local zero_clientaddr with pg_memory_is_all_zeros, so zero_clientaddr no longer exists; upstream PG commit bcc8b14ef630b2ad9aae7813981fb248fbff9ed8 dropped the HAVE_IPV6 ifdef around the AF_INET6 test. - Took master's outer CM branch and rewrote the non-CM else branch in PG19 form. - pg_stat_get_backend_client_addr / _port auto-merged; both already use the PG19 pg_memory_is_all_zeros form. - src/postgres/contrib/postgres_fdw/postgres_fdw.c: - YbGlobalViewReadExecScan: - master commit a14e4eca6f3f076f9225a2b4b2cb6a94e12817f8 switched to the one-argument YBCPgResultFromPB, YBCPgGlobalViewReadGetError and YBCPgGlobalViewReadClearScanState, dropping the YbcRemotePgExecResult local. - upstream PG commit 7d8f5957792421ec3bb9d1b9b6ca25d689d974b7 added libpqsrv_PQwrap; YB pg19 wraps every PGresult produced here with it. - Combined: master's API with the libpqsrv_PQwrap wrapping on both result paths. The buffer-release hunk in postgresIterateForeignScan auto-merged. - src/postgres/src/bin/pg_dump/pg_dump.c: - dumpTableData parentTbinfo: - master commit 62a60a60f20ef1f18125532858b76ad315006731 added a "char *sanitized;" declaration immediately below parentTbinfo, which is what makes the adjacent declaration line conflict. - upstream PG commit 707f905399b4e47c295fe247f76fbbe53c737984 constified read-only TableInfo pointers. - Took PG19's const TableInfo *. - src/postgres/src/bin/pg_dump/pg_dumpall.c: - dumpUserConfig header, dumpDatabases locals, dumpDatabases header (three conflicts): - master commit 62a60a60f20ef1f18125532858b76ad315006731 (CVE import) added sanitize_line around the username and dbname printed into the dump header. - upstream PG added non-plain output formats to pg_dumpall: the headers are emitted only when archDumpFormat == archNull, and dumpDatabases gained an oid local for the map.dat path. - Took the PG19 side at all three. Nothing from master's import is lost: PG19 already carries the relevant changesnatively, from upstream's own fix that master back-ported, so the CVE change is subsumed by the PG19 side. - src/yb/yql/pgwrapper/pg_read_time-test.cc: - PgYbReadTimeTest/PgFollowerReadTest/PgClampDeferReadTest (three conflicts): - master commits a9018bc0f9c65231494466f01183a8cf10c2107f and dc119a8e350b6fea9eb516749e6ed0d6478c55f8 restructured the file (640 lines changed). - YB pg19 commit 97af892519dbed7b2b631faa608e7977bb700fff was the only change to this file: it renamed the GUC in three SET statements, following upstream PG commit 5352ca22e0012d48055453ca9992a9515d811291 (force_parallel_mode -> debug_parallel_query). - Took master's side at all three conflicts, then re-applied the rename over the whole file (6 sites). - src/postgres/src/backend/utils/misc/guc.c + guc_parameters.dat + guc_tables.c + guc_hooks.h + guc-file.l + guc_internal.h + guc_funcs.c + src/postgres/src/include/utils/guc.h + yb_ysql_conn_mgr_helper.h: - Background for this whole group: upstream PG commit 0a20ff54f5e66158930d5328f89f087d4e9ab400 split guc.c and moved ProcessConfigFile/ProcessConfigFileInternal/show_all_file_settings out of guc-file.l into guc.c and guc_funcs.c; upstream PG commit 63599896545c7869f7dd28cd593e8b548983d613 generates the GUC tables from guc_parameters.dat, so the ConfigureNames* arrays and PG's hook forward decls no longer live in guc.c (hooks are non-static and declared in guc_hooks.h; enum options arrays live in guc_tables.c). Six of the seven guc.c conflicts are master edits landing inside regions PG19 relocated wholesale; resolved by taking the PG19 side and re-porting each master delta into its live location, listed below. - CM logical-client GUC storage variables: - master commit f576915851219317ae42f8049b0a10f196494982 defined yb_conn_mgr_client_addr / yb_conn_mgr_client_port / yb_conn_mgr_client_hostname in guc.c. - YB pg19 has extern bool yb_conn_mgr_modifying_defaults at the same spot. - Kept both. Added externs for the three to yb_ysql_conn_mgr_helper.h, which guc_tables.c includes; on pg19 the generated table references them from another translation unit. - CM logical-client GUC entries (yb_conn_mgr_client_addr, yb_conn_mgr_client_hostname, yb_conn_mgr_client_port): - master commit f576915851219317ae42f8049b0a10f196494982 added them to ConfigureNamesString / ConfigureNamesInt. - Ported into guc_parameters.dat in alphabetical order, before yb_conn_mgr_selective_deallocate. - CM logical-client check/assign hooks: - master commit f576915851219317ae42f8049b0a10f196494982 added six static hooks at the end of guc.c. - Ported into guc.c's YB hook region as non-static functions and declared in guc_hooks.h. check_yb_conn_mgr_client_addr / _hostname called pg_clean_ascii(*newval) in place; upstream PG commit 45b1a67a0fcb3f1588df596431871de4c93cb76f made pg_clean_ascii return a new string, so both were rewritten in PG19's check_application_name form (in commands/variable.c: pg_clean_ascii with MCXT_ALLOC_NO_OOM, guc_strdup, guc_free). - This is the one place in the guc group where behaviour changes rather than just location. PG15's pg_clean_ascii overwrote each non-printable byte with "?" in place, so the value length never changed; PG19's escapes each as "\xNN" into a fresh buffer up to four times longer. A malformed yb_conn_mgr_client_addr therefore surfaces escaped rather than as "?", and the expansion can exceed the strlcpy targets in backend_status.h (NI_MAXHOST for addr, NAMEDATALEN for hostname), which truncates rather than overflows. Harmless for well-formed values from Odyssey. - assign_yb_conn_mgr_client_addr calls SetConfigOption("yb_conn_mgr_client_hostname", ...) while an outer set_config_option is still in flight, on the authentication path. That nesting existed on master too, but there the nested check hook only sanitized in place; after the PG19 rewrite it runs pg_clean_ascii + guc_strdup + guc_free inside the outer set. The nested call is a complete set_config_option over ordinary GUC memory, so it looks safe; an assert-enabled build would exercise it (guc_free asserts the chunk context, and cassert brings CLOBBER_FREED_MEMORY and MEMORY_CONTEXT_CHECKING). ASAN would not catch it, since this is palloc/pfree inside GUCMemoryContext. - yb_pg_stat_plans_show_max_exec_params default: - master commit 2cbe4f2b16b12bec114381691dfebd094535afa8 flipped the boot value false -> true. - Ported as a boot_val change on the guc_parameters.dat entry. - yb_db_history_retention_pin_mode: - master commit 7b100f3d47063801a8fc9185f8cf9ca500d247b1 added the enum GUC plus its yb_db_history_retention_pin_mode_options array. - Ported the entry into guc_parameters.dat and the options array into guc_tables.c. The variable and enum constants live in src/yb/yql/pggate/util/ybc_guc.{h,cc}, which auto-merged. - GUC config-file validation (ProcessConfigFileInternal signature, YbValidateConfigFile): - master commit 193b153576afa6f02f27fc20bcdcee32a0831389 added a leading yb_config_file parameter to ProcessConfigFileInternal, a yb_validating mode that parses at LOG and reports via set_config_option's elevel, and a new YbValidateConfigFile wrapper; yb_pg_conf_validator.c now calls it instead of its own yb_validate_guc_conf_file. - Ported the parameter and the yb_validating logic onto the PG19 copy of ProcessConfigFileInternal in guc.c, added YbValidateConfigFile to guc-file.l next to ProcessConfigFile, and updated the guc_internal.h declaration plus both callers (guc-file.l ProcessConfigFile and guc_funcs.c show_all_file_settings). yb_pg_conf_validator.c auto-merged. - src/postgres/src/include/utils/guc.h: - master commit 193b153576afa6f02f27fc20bcdcee32a0831389 collected the YB externs into a trailing "/* YB declarations */" block and added YbValidateConfigFile. - upstream PG commit 63599896545c7869f7dd28cd593e8b548983d613 moved the PG check/assign hook externs out of guc.h into guc_hooks.h. - Kept only the YB declarations block. Build fixes (auto-merged files; no conflict, so these are outside the resolved set): - src/postgres/src/backend/utils/activity/pgstat_backend.c + pgstat_io.c: - pgstat_tracks_backend_bktype / pgstat_tracks_io_bktype switches: - master commit ca702f7a7fbd9aec42175658d12e6b5c8e2dd5c1 added YB_XCLUSTER_DDL_QUEUE_BACKEND and YB_XCLUSTER_SETUP_BACKEND to BackendType. - Both switches are YB pg19 additions (they list every BackendType so a new one trips -Wswitch), so master never had to extend them. - Added both to the existing YB opt-out group in each. - Also added the YB_TODO_PG19MERGE disclaimer above the YB group in pgstat_io.c. YB pg19 commit 4c01ae04c6d8e5487679d2079e44a067a66febcd recorded that intent for both lists, but the comment only landed in pgstat_backend.c, so this switch read as a settled per-type classification rather than a provisional opt-out. The comment also records that the two functions are not symmetric: pgstat_count_backend_io_op returns early for an untracked type, while pgstat_count_io_op and pgstat_io_flush_cb assert, so false here is a claim that the process performs no tracked IO at all. - src/postgres/src/backend/utils/misc/yb_index_check.c: - yb_index_check tail: - master commit d9add44123a7ca45bfd803c958b159c5e0d81b78 calls tuplestore_donestoring. - upstream PG commit 75680c3d805e2323cd437ac567f0677fdfc7b680 retired it; it had already been a no-op macro. - Deleted the call. - col_to_jsonb_repr / populate_table_cols_data / populate_index_cols_data: - master commit d9add44123a7ca45bfd803c958b159c5e0d81b78 threads a JsonbParseState ** through these and reads pushJsonbValue's returned JsonbValue *. - upstream PG commit 0986e95161cec929d8f39c01e9848f34526be421 changed pushJsonbValue to take a JsonbInState * and return void. - Threaded JsonbInState through instead (by value in the two builders, by pointer into col_to_jsonb_repr) and took the finished value from pstate.result. - src/postgres/src/backend/utils/misc/guc.c: - include block: - The six conn-mgr hooks ported into this file call pg_clean_ascii and yb_pgstat_set_ycm_client_info, neither of which guc.c pulled in transitively on pg19. - Added common/string.h and utils/backend_status.h to the YB include group. - src/postgres/src/test/regress/expected/yb.orig.yb_index_check.out + yb.orig.yb_index_check_1.out: - EXPLAIN ANALYZE golden for yb_index_check: - master commit d9add44123a7ca45bfd803c958b159c5e0d81b78 made yb_index_check set-returning, so the plan gains a ProjectSet node above Result. - YB pg19 commit b2a5e6aa712ba17c547c8bc383d3663c37849afc reblessed these goldens for upstream PG's fractional EXPLAIN row counts (actual rows=1 -> actual rows=1.00). - Combined: kept master's ProjectSet plan shape and applied the fractional format to both nodes, recomputing psql's column width and separator.
| Commit: | da866a0 | |
|---|---|---|
| Author: | Naorem Khogendro Singh | |
| Committer: | Naorem Khogendro Singh | |
[PLAT-22528] Clock skew input param acceptable_clock_skew_sec is not passed from YBA to node agent Summary: Some issues detected by AI. 1. Clock sync related params are not passed down to node agent. Not very important as there is a default but good to cover in case we want to change. 2. process_plain_files in zip_purge_yb_logs.sh.j2 does not use the input size parameter. Not a big issue as permitted_disk_usage_plain_kb is anyways used. 3. collect_metrics_wrapper.sh.j2 ignores the detected node-exporter. Some old universes may complain. Test Plan: Itests must pass. Otherwise, these are good to fix issues but should not change behavior Reviewers: amalyshev, hzare Reviewed By: hzare Subscribers: yugaware Differential Revision: https://phorge.dev.yugabyte.com/D58224
| Commit: | 4552a3c | |
|---|---|---|
| Author: | Ella Baron | |
| Committer: | GitHub | |
dist-trace: Address review — drop AshMetadataPB mention, rename parent to traceparent
| Commit: | 1085ece | |
|---|---|---|
| Author: | Ella Baron | |
| Committer: | GitHub | |
dist-trace: propagate trace context across the RPC boundary Carry a distributed trace from an RPC caller to its callee: define the wire type, write it into the request header, decode it on the far side, and open a server span parented under the caller's client span. All of it lands together because none of the pieces is observable on its own -- a wire field nobody reads, or a reader with nothing on the wire, cannot be tested. Wire type (rpc_header.proto): - TraceContextPB (W3C trace-context) and RequestHeader.trace_context. This is the type shared by the RPC-header and shared-memory transports; the latter starts using it in a later commit. Writer (outbound_call.cc): - SetRequestParam serializes the active client span's SpanContext into the trace_context submessage while it is sizing and writing the header, so the header is still written in a single pass. ~30 bytes on the wire, and only when a trace is active; an older peer skips the unknown field. Reader (serialization.{cc,h}): - ParseTraceContext(Slice) decodes the RequestHeader.trace_context wire field; ToSpanContext(TraceContextPB) is the shared-memory equivalent; both go through BuildSpanContext and fail on a zero/invalid context. - ParseHeader captures RequestHeader.trace_context into ParsedRequestHeader. Server / inbound span (yb_rpc.{cc,h}): - YBInboundCall::ParseFrom parses the header trace_context into parent_span_context_ (best-effort: a bad context is logged, never fails the RPC). - CreateServerSpan starts a StartServerSpanWithScope child of that parent (remote) or of an explicitly passed context (local); RespondSuccess/Failure set status and end it; DropServerSpanScope releases the scope on the handler thread while the span ends later. Client local-call handoff (local_call.{cc,h}, rpc_context.cc): - LocalOutboundCall tags rpc.local_call and drops its client-span scope in the ctor (a local call never hops threads to drop it later). - otel_span_context() exposes the outbound span's context so the local inbound span can parent under it (no wire header exists locally). - RpcContext creates the server span: from the wire header for remote calls, from the outbound call's context for local calls. Process init (tserver/db_server_base.cc): - DbServerBase is the shared base of Master and TabletServer, so DbServerBase:: Init is the single place that initializes the process-wide tracer (service name = server name, node id = permanent uuid), with Shutdown tearing it down. Gated by IsDistTraceEnabled, so it is a no-op when tracing is off. Without this the server spans above are never exported, which is why the 10 lines ship here rather than as a commit of their own. Test: TestRpcSpanReachesTabletServer runs a SELECT under a known traceparent and asserts that the TabletServer's inbound "rpc yb.tserver.PgClientService.Perform" span lands in that trace as a child of the ysql backend's outbound span, with the expected rpc.service/rpc.method attributes. The new WaitForRemoteChildSpan collector helper does the caller/callee pairing check.
| Commit: | f1de8a1 | |
|---|---|---|
| Author: | Sergei Politov | |
| Committer: | Sergei Politov | |
[#33949] DocDB: Fix QLTabletTest.VerifyIndexRange stale verification snapshot Summary: When a `VerifyTableRowRange` request carries no `read_time`, `TabletServiceImpl::VerifyTableRowRange` verified at the replica's current safe time, which for `RequireLease::kFalse` is the propagated follower safe time and only advances as the leader ships it in a Raft heartbeat. Under TSAN `raft_heartbeat_interval_ms` is `1000`, so the snapshot could predate the just-completed index backfill, and every index lookup at that read time came back empty, reporting all 250 rows as mismatched. The request now verifies as of `server_->Clock()->MaxGlobalNow()` when no read time is supplied, so the existing `tablet->SafeTime()` call blocks until the replica catches up to that time; a caller-supplied `read_time` is unchanged. A replica that cannot reach that time before the deadline now logs a warning instead of `DFATAL`, since a lagging peer is an expected outcome rather than an invariant violation. ## Upgrade/Downgrade safety: Safe in both directions, no action required. No protobuf field is added or removed and no persisted format changes -- the only `tserver.proto` edit is a comment correcting the documented meaning of an empty `read_time`. The behavior change is confined to how a tserver picks a read time for a request it is serving, so in a mixed-version cluster each node independently uses its old or new selection and the RPC stays wire-compatible in both directions. The only non-test caller is the manual `ts-cli verify_tablet` admin command, which never sets `read_time`; against an upgraded tserver it verifies a snapshot that includes everything committed before the request instead of a possibly stale follower snapshot. --- _Produced fully automatically by csi-fix.py: Claude Opus (5) analyzed the failure logs, wrote the change, and verified it locally._ Test Plan: ./yb_build.sh tsan --clang21 --cxx-test ql-tablet-test --gtest-filter QLTabletTest.VerifyIndexRange -n 288 --stop-at-failure -- -p 8 Ran the above locally with no failures. Before the fix, this test reproduced locally (same test command and -p / env as above, without --stop-at-failure) -- with the CSI dashboard rate for reference: - clang21-tsan: 1/96 iterations failed locally; CSI 0/49 Machine: GCP n2-standard-32, AlmaLinux 8.10 (Cerulean Leopard) x86_64, 32 vCPU, 125 GiB RAM, Intel(R) Xeon(R) CPU @ 2.80GHz Reviewers: bkolagani Reviewed By: bkolagani Subscribers: ybase Tags: #jenkins-ready Differential Revision: https://phorge.dev.yugabyte.com/D58165
| Commit: | 72bfdc3 | |
|---|---|---|
| Author: | Balaji Subramanian | |
| Committer: | Balaji Subramanian | |
[#33557] xCluster: verify_xcluster_slice and verify_xcluster_group Summary: Add `verify_xcluster_slice` to compare one source/target key range at a common xCluster safe time, with schema checks before and after hashing. Its JSON result distinguishes a confirmed mismatch (`kDiverged`) from retryable, schema, and infrastructure failures. Add `verify_xcluster_group` to discover and sweep the tables in an automatic-DDL replication group. It derives structured source master addresses from the target, verifies logical source tablet ranges with positional row-limit and concurrency controls, and reports slice records followed by a group summary. Sequence data, vector indexes, and `replicated_ddls` are excluded because their rows are not comparable through this hash path. Upgrade/Rollback safety: `GetUniverseReplicationInfoResponsePB` retains legacy field 3 as `DEPRECATED_source_master_addresses` and adds structured `source_master_addrs` as field 7. Renaming the legacy source identifier does not change its wire tag or type. Old clients continue reading field 3 and ignore field 7; new clients use field 7 and return a clear upgrade error when an older target master does not provide it. The generated legacy accessor is used only in the C++ master and client adapter; YBA does not consume it. No data is persisted, so rollback has no storage impact. Test Plan: `./build-support/lint.sh --rev origin/master` (0 errors, 0 warnings) `xcluster_verify-test` (29 tests passed) `yb-admin-test --gtest_filter='*VerifyXCluster*'` (2 tests passed) `xcluster_db_scoped-test --gtest_filter='*VerifyXCluster*'` (14 tests passed) Reviewers: hideaki.kimura, jhe Reviewed By: hideaki.kimura, jhe Subscribers: svc_phabricator, ybase Differential Revision: https://phorge.dev.yugabyte.com/D57781
| Commit: | e9a9b5c | |
|---|---|---|
| Author: | Samson Shaji | |
| Committer: | Samson Shaji | |
[#28675] DocDB: Add lag column to list_all_master yb-admin Summary: - Added a `Lag(ms)` column to yb-admin `list_all_masters`: milliseconds since the Raft leader last had a successful consensus exchange with each follower master (same method used for `max_follower_heartbeat_delay`). - ServerEntryPB: optional `heartbeat_delay_ms` (wire field). - Leader ListMasters: fills lag per Raft-config peer UUID via `CatalogManager::GetMasterFollowerHeartbeatDelaysMs() -> sys-catalog GetFollowerCommunicationTimes()`. - yb-admin: prints Lag(ms) or N/A (leader / non-leader-served RPC). **Upgrade/Rollback safety:** Describe how this change handles upgrade and rollback of YugabyteDB. - `heartbeat_delay_ms` is optional on `ServerEntryPB`. Old masters/clients ignore unknown fields; old yb-admin ignores the new field. - New master + old yb-admin: unchanged columns; lag simply not shown until client is upgraded. Rollback: safe; no persisted schema change beyond optional RPC field. What Test/Preview/AutoFlag is used to guard the feature? Feature guard: No AutoFlag; backward compatible optional proto + CLI column only. Test Plan: Tested manually using following commands: ``` ./yb_build.sh release daemons --sj --skip-pg-parquet --no-odyssey --no-ybc ``` ``` ./build/latest/bin/yb-admin --master_addresses 127.0.0.1:7100,127.0.0.2:7100,127.0.0.3:7100 list_all_masters ``` Screenshots Before: {F516263} After: Notice that node 3 (the one taken down), has a high lag (`106662` ms). This should indicate/clarify to the user that this follower is catching up; thus the reason node 3 cannot be promoted to `LEADER` despite being an `ALIVE` + `FOLLOWER` from RAFT state. {F516619} Reviewers: balaji.subramanian, mhaddad Reviewed By: balaji.subramanian, mhaddad Subscribers: ybase Differential Revision: https://phorge.dev.yugabyte.com/D55098
| Commit: | dafbd4e | |
|---|---|---|
| Author: | Jason Kim | |
| Committer: | Jason Kim | |
[#33940] YSQL: document the NoMovementScanDirection contract Summary: YB overloads NoMovementScanDirection to mean "row order does not matter". create_index_path assigns it to unordered index scans of YB relations, IndexNext and IndexOnlyNext keep it, and yb_scan_core.c then skips YBCPgSetForwardScan, leaving is_forward_scan unset on the read request. Pggate takes an unset direction as permission to read tablets in parallel (CouldBeExecutedInParallel) and to skip preserving ybctid order on secondary index scans. Forward and Backward both force sequential, order-preserving reads. None of this is clear in the code. The full rationale exists only in the message of commit 1215c3a71a0a6283d402f08645f8f096ab3bb0ce (#13737). Upstream removed NoMovementScanDirection for unordered index scans in PG16 (commit e9aaf06328c7f962f8586618981e9763d31402a3) and asserts Forward or Backward in create_indexscan_plan. The lack of documentation caused the author to attempt to merge in upstream PG's favor when that removes the parallel optimizations for YB. Add comments at every point in the chain, with create_index_path as the canonical explanation and the others pointing at it. **Upgrade/Rollback safety:** none as this only changes comments. Test Plan: Comment-only change. arc lint is green. Close: #33940 Jenkins: compile only Reviewers: fizaa Reviewed By: fizaa Subscribers: ybase, yql Differential Revision: https://phorge.dev.yugabyte.com/D58156
| Commit: | dfd1753 | |
|---|---|---|
| Author: | Ella Baron | |
| Committer: | Ella Baron | |
dist-trace: propagate trace context across the RPC boundary Carry a distributed trace from an RPC caller to its callee: define the wire type, write it into the request header, decode it on the far side, and open a server span parented under the caller's client span. All of it lands together because none of the pieces is observable on its own -- a wire field nobody reads, or a reader with nothing on the wire, cannot be tested. Wire type (rpc_header.proto): - TraceContextPB (W3C trace-context) and RequestHeader.trace_context. This is the type shared by the RPC-header and shared-memory transports; the latter starts using it in a later commit. Writer (outbound_call.cc): - SetRequestParam serializes the active client span's SpanContext into the trace_context submessage while it is sizing and writing the header, so the header is still written in a single pass. ~30 bytes on the wire, and only when a trace is active; an older peer skips the unknown field. Reader (serialization.{cc,h}): - ParseTraceContext(Slice) decodes the RequestHeader.trace_context wire field; ToSpanContext(TraceContextPB) is the shared-memory equivalent; both go through BuildSpanContext and fail on a zero/invalid context. - ParseHeader captures RequestHeader.trace_context into ParsedRequestHeader. Server / inbound span (yb_rpc.{cc,h}): - YBInboundCall::ParseFrom parses the header trace_context into parent_span_context_ (best-effort: a bad context is logged, never fails the RPC). - CreateServerSpan starts a StartServerSpanWithScope child of that parent (remote) or of an explicitly passed context (local); RespondSuccess/Failure set status and end it; DropServerSpanScope releases the scope on the handler thread while the span ends later. Client local-call handoff (local_call.{cc,h}, rpc_context.cc): - LocalOutboundCall tags rpc.local_call and drops its client-span scope in the ctor (a local call never hops threads to drop it later). - otel_span_context() exposes the outbound span's context so the local inbound span can parent under it (no wire header exists locally). - RpcContext creates the server span: from the wire header for remote calls, from the outbound call's context for local calls. Process init (tserver/db_server_base.cc): - DbServerBase is the shared base of Master and TabletServer, so DbServerBase:: Init is the single place that initializes the process-wide tracer (service name = server name, node id = permanent uuid), with Shutdown tearing it down. Gated by IsDistTraceEnabled, so it is a no-op when tracing is off. Without this the server spans above are never exported, which is why the 10 lines ship here rather than as a commit of their own. Test: TestRpcSpanReachesTabletServer runs a SELECT under a known traceparent and asserts that the TabletServer's inbound "rpc yb.tserver.PgClientService.Perform" span lands in that trace as a child of the ysql backend's outbound span, with the expected rpc.service/rpc.method attributes. The new WaitForRemoteChildSpan collector helper does the caller/callee pairing check.
| Commit: | 1b0415d | |
|---|---|---|
| Author: | Ella Baron | |
| Committer: | Ella Baron | |
dist-trace: Address review — drop AshMetadataPB mention, rename parent to traceparent
| Commit: | ca6fd80 | |
|---|---|---|
| Author: | Anton Rybochkin | |
| Committer: | Anton Rybochkin | |
[#31065] docdb: Deny split until the parent tablet vector index is post-split compacted Summary: **Background** After a tablet split, each child tablet's vector indexes inherit all vectors from the parent, including roughly half that no longer belong to the child's key range. Vector search must filter these irrelevant vectors via reverse-mapping lookups into RegularDB (two reads per vector), which is significantly slower than the vector search itself. If splits continue before the inherited vectors are compacted away, the ratio of irrelevant vectors grows with each generation, degrading search latency. To control this, tablet splits should be postponed until co-hosted vector indexes have completed post-split compaction. **Changes** 1. **Splits wait for vector index post-split compaction.** `Tablet::StillHasOrphanedPostSplitDataAbortable()` now also reports leftover vector index parent data, so a tablet is not split again (and the load balancer does not move it) until that data is gone. New advanced gflag `vector_index_require_parent_data_compacted_before_split` (default `true`) turns the wait off. The existing hidden gflag `vector_index_include_into_post_split_compaction` turns the vector index side off entirely. 2. **Scheduling a post-split compaction uses its own check.** `Tablet::NeedPostSplitCompaction()` answers a different question: is a post-split compaction still owed for this tablet? `TriggerPostSplitCompactionIfNeeded()` calls it instead of the split check. Otherwise setting `vector_index_require_parent_data_compacted_before_split=false` would also stop an interrupted vector index compaction from being retried when the tablet reopens. 3. **Each vector index tracks its own progress.** Two new `ConsensusFrontier` fields: `split_generation` and `split_min_chunk_serial_no`. Chunks with a serial number below the bound came from the parent tablet. At open time `InitFrontiers` compares the tablet's `split_generation` with the frontier's; if a new split happened it records `LastSerialNo() + 1` as the bound and flushes. After that, `ParentDataCompacted()` is simply `MinSerialNo() >= bound`, or true when there are no chunks. Nothing about vector index compaction is stored in tablet metadata. 4. **New `split_generation` in `KvStoreInfoPB`.** A tserver-local counter, parent + 1 in `CreateSplitChildMetadata`, which is how a vector index notices that a new split happened. Restored from the vector index manifests during bootstrap if a rollback dropped it -- see Upgrade/Rollback safety. 5. **Reverse mapping lookups now respect tablet key bounds.** `IndexReverseMappingReader::Fetch` returns nothing for a mapping whose ybctid is outside the tablet's range. That lets a split child's compaction drop the vectors that now belong to its sibling. Tombstones are still returned unchanged (will be fix in a follow-up revision). 6. **`parent_data_compacted` renamed to `rocksdb_parent_data_compacted`** in `KvStoreInfoPB`, `TabletStatusPB` and everywhere it is used, because it only covers RocksDB and is now easy to confuse with the vector index state. 7. **`CreateSubtablet` renamed to `CreateSplitChildTablet`** (and `CreateSubtabletMetadata` to `CreateSplitChildMetadata`), since both are only used for tablet splitting. 8. **RocksDB and vector indexes are compacted independently.** `TriggerManualCompactionIfNeeded` takes an `IncludeVectorIndexes` argument and skips whichever side is already done; `VectorIndexList::Compact` skips indexes that are already compacted. If the RocksDB compaction fails, the vector index compaction is skipped: its merge filter needs the obsolete reverse mappings to be gone first. 9. **New `vector_indexes_parent_data_compacted` in `TabletStatusPB`.** A diagnostic field for tests and the status page. Always reported by the tserver when the tablet is available, and never used as persistent state. **Upgrade/Rollback safety** All new fields are optional and default to pre-change behavior: `split_generation` defaults to 0 (matching a never-split tablet) and `split_min_chunk_serial_no` to 0 (read as "no inherited parent data"), so pre-existing split children are treated as having nothing to compact and the gate applies only to splits performed by the new binary; the renames keep their field numbers (`KvStoreInfoPB` 8, `TabletStatusPB` 18), so on-disk and wire formats are unchanged. Rollback drops `KvStoreInfoPB.split_generation` from the superblock while the vector index manifests keep it, so on the next upgrade `TabletBootstrap::MaybeUpdateMetaAfterTabletHasBeenOpened` restores it from the largest generation persisted across the tablet's vector indexes -- without that, the next split child would get generation 1, which is not greater than the persisted value, and would skip stamping its own parent data boundary. Test Plan: Jenkins: jobs: full ./yb_build.sh --cxx-test consensus_frontier-test --gtest_filter=ConsensusFrontierTest. SplitGenerationAndMinChunkSerialNo ./yb_build.sh --cxx-test tablet-split-test --gtest_filter=TabletSplitTest.StillHasOrphanedPostSplitDataVectorIndex ./yb_build.sh --cxx-test tablet-split-test --gtest_filter=TabletSplitTest.SplitGenerationOnSubtablet ./yb_build.sh --cxx-test tablet-split-itest --gtest_filter=TabletSplitSingleServerITest.PostSplitCompactionScheduledOnTabletOpen ./yb_build.sh --cxx-test pg_vector_index-test --gtest_filter=PgDistributedVectorIndexTest.SplitBlockedWithOrphanedPostSplitData/* ./yb_build.sh --cxx-test pg_vector_index-test --gtest_filter=PgDistributedVectorIndexTest.ManualSplitSimple/* ./yb_build.sh --cxx-test pg_vector_index-test --gtest_filter=PgDistributedVectorIndexTest.SplitGenerationRestoredAfterSuperblockReset/* ./yb_build.sh --cxx-test non_transactional_batch_writer-test --gtest_filter=NonTransactionalBatchWriterTest.ExternalApplyGatesVectorIndexFeed ./yb_build.sh --cxx-test tablet-split-test --gtest_filter=TabletSplitTest.SplitTablet ./yb_build.sh --cxx-test tablet-split-itest --gtest_filter=TabletSplitSingleServerITest.TabletServerOrphanedPostSplitData ./yb_build.sh --cxx-test pg_tablet_split-test --gtest_filter=PgTabletSplitTest.SplitDuringLongRunningTransaction Reviewers: timur, zdrudi, sergei, #db-approvers Reviewed By: zdrudi, sergei, #db-approvers Subscribers: hbhanawat, svc_phabricator, ybase Differential Revision: https://phorge.dev.yugabyte.com/D52469
| Commit: | 121b033 | |
|---|---|---|
| Author: | Basava | |
| Committer: | Basava | |
[BACKPORT 2026.1][#33296] DocDB: Table locks: Fix CleanupExpiredLeaseEpochs logic at the master's object lock manager Summary: Expired-lease lock cleanup was writing the release against the wrong lease epoch in sys catalog. It stamped `max_lease_epoch_to_release + 1` on the request so tservers would fence out old acquires, but that same field is also the key used to delete locks from `lease_epochs`. The delete often became a no-op, so old locks stayed on disk after the in-memory release. This wasn't a problem prior to commit https://github.com/yugabyte/yugabyte-db/commit/c670b00d9da57a4af8b0be8850bf9b98391947da / D56278 as we were explicitly ignoring locks correspinding to expired ysql leases while exporting the object locks bootstrap payload. The commit changed the behavior for master bootstrap payload to load up all locks since it wanted to establish the semantic - Master's persistent lock state of active locks should process the release only after all tservers have released the lock. Despite the releases being relaunched, and released in the master's in-memory lock manager, the persisted state still remains (The space amplification problem existed prior to the commit). This could result in a false lock conflict on master failover when the new master tries replaying the expired locks as well as the active conflicting locks if any. This revision fixes the issue in the following manner - changing `lease_epoch` to reflect the actual epoch with which is the lock is associated to. - introduce `ignore_lease_epochs_before` to notify the tservers that they can drop any subsequent acquire requests from the source tserver with lease epoch lesser than `ignore_lease_epochs_before`. Note that this was the original intention for setting lease_epoch on the release request to `max_lease_epoch_to_release + 1`. The normal release operations don't populate `ignore_lease_epochs_before`. The tservers prefer `ignore_lease_epochs_before` when set, and fallback to use `lease_epoch` when tracking the max seen lease epoch for other tservers. **Upgrade/Downgrade safety** Table locking hasn't been shipped in any release yet, so there shouldn't be any implications. Original commit: ae5ffa985849cbc9d36389869b7713ffc585698b / D56930 Test Plan: Jenkins ``` ./yb_build.sh --cxx-test object_lock-test --gtest_filter ObjectLockTest.ExpiredLeaseReleaseClearsPersistedLocks ``` Reviewers: amitanand, zdrudi Reviewed By: zdrudi Subscribers: ybase Differential Revision: https://phorge.dev.yugabyte.com/D56975
| Commit: | 461a8a9 | |
|---|---|---|
| Author: | Balaji Subramanian | |
| Committer: | Balaji Subramanian | |
[#33556] DocDB: Phase 1 changes for supporting cross-cluster verification Summary: Comparing two clusters needs a table hash that actually means something, and the one dump_tablet_data produced did not. It XOR-ed every column's bytes together, so a row holding 'a' in col1 and 'b' in col2 hashed the same as one with those values swapped, and two genuinely different tables could look identical. It also always scanned to the end of the tablet, so on a large table it either ran past the RPC deadline or outlived its read snapshot. This change makes column placement affect the hash, lets a scan stop after a set number of rows and resume where it left off, and stamps each hash with the scheme that produced it so two hashes are only compared when they mean the same thing. yb-admin gets a get_table_hash command to drive it. Wire protocol (tserver_service.proto). Three new fields: max_rows = 8 on the request, next_key = 4 and hash_scheme_version = 5 on the response. The request now runs 5, 6, 7, 8 with master's max_wait_ms = 7 undisturbed, and the response runs 1 through 5, so no field number collides. Hashing core (tablet_dump_helper.{h,cc}). This is the substance. Row hashing now salts each column's contribution with the column's identity and seals each row before folding it into the table-wide XOR, so a value moved from one column to another no longer cancels out and reports a false match. Alongside that, a scan can be capped at max_rows and resumed from next_key, for both range- and hash-partitioned tables. MixHash64, the avalanche step the salt is built on, now lives in util/hash_util.h beside the other hash helpers, with its Murmur3 finalizer constants named rather than inline. The pagination subtlety worth a reviewer's attention: for a hash-partitioned table, next_key is an encoded row key rather than the 2-byte hash that partition bounds normally use, so a resumed scan restarts mid-band. The server distinguishes the two by length, and the client lifts both partition bounds and continuation keys into the same encoded space before comparing. That asymmetry is where this code is easiest to get wrong, so the server now also checks that a longer bound really is an encoded row key instead of taking any long byte string on faith -- a malformed one would otherwise match nothing and report an empty range as a successful hash of zero rows. Server plumbing is small (tablet_service.cc): pass the new request fields through and populate the response. It also gains a test-only flag, TEST_dump_tablet_data_hash_scheme_version, which makes one tserver report a different scheme version so the mixed-version case is reachable from a test instead of needing two builds. Client and CLI (yb-admin_client.{h,cc}, yb-admin_cli.cc, ts-cli.cc): the get_table_hash command, the max_rows_per_scan flag, tablet ordering and continuation across tablets, and hex fields on list_tablets. New header src/yb/tools/table_hash.h holds TableHashTotals and PartitionRangeOverlaps, which ComputeTableXorHash returns. ## Upgrade/Rollback safety No persisted state and no migration in either direction. A hash is computed on demand and returned in an RPC, so no binary ever reads back an artifact another binary wrote. All three new proto fields are optional, and unset reproduces the old behaviour: no max_rows is the uncapped scan the server already did. The mixed-binary case is a new yb-admin against a tserver that predates this change. That tserver ignores max_rows, scans to the end of the tablet and returns no next_key, so --max_rows_per_scan is silently not honoured. It also leaves hash_scheme_version unset, which the client reads as version 0 rather than as its own, and yb-admin refuses to fold together tablets whose schemes disagree. A rolling upgrade therefore cannot produce a hash that is half v0 and half v1: it either reports v0 throughout, or fails naming the two tablets. Old yb-admin against a new tserver never sets max_rows and ignores both new response fields. Rollback puts tservers back on v0. Hashes taken before and after a rollback differ, which the printed scheme version makes visible rather than silent. --max_rows_per_scan is a yb-admin client flag, non-runtime, default 0 (the previous behaviour). TEST_dump_tablet_data_hash_scheme_version is test-only. Test Plan: tablet_dump_helper-test, 11 cases. Beyond the placement cases (values swapped between columns, equal values that must not cancel, a value moved between rows, row order irrelevant, NULL distinct from zero and empty string): a golden-value case pinning the scheme to kTabletDataHashSchemeVersion, since every other case compares two hashes from the same build and so cannot see an unversioned change to the hashing; a direct column-identity case; a case pinning that collection element order reaches the hash, which holds because the DocDB read path emits map and set children in sorted order; and two randomized cases over generated tables, one asserting distinct tables hash distinctly and one that changing a single cell changes the table hash. Checked those are not vacuous by mutating the production hash -- first dropping the per-row seal, then the column-id salt -- and confirming the intended cases fail each time and pass again when reverted. yb-admin-test, 12 cases: pagination across tablets on range- and hash-partitioned tables; cap boundaries (cap landing exactly on a tablet boundary, cap equal to the row count, cap above it, and cap of 1 draining a row at a time); a capped scan resuming at a pinned read time, so rows written after it are not hashed; next_key returned as an encoded row key rather than a raw partition key when a cap on a hash-partitioned table lands exactly on a tablet boundary, then hashed onward from that key to confirm the counts sum and the hashes xor back to the whole-table values; a capped scan of one table in a colocated tablet; value placement reaching the hash end to end through the CLI; list_tablets JSON hex bounds, and those bounds fed back through get_table_hash to confirm they tile the key space; max_rows rejected for a colocation parent; a refusal to fold together hashes taken under different schemes, driven by the new test flag, including the version-0 pre-change tserver; and malformed hash bounds rejected. Re-ran the existing consumers of DumpTabletData and xor_hash for regressions: pg_libpq-test, yb-ts-cli-test, async_writes-test. ./yb_build.sh release --no-remote --target tablet_dump_helper-test --target yb-admin-test ./build/latest/tests-tablet/tablet_dump_helper-test ./build/latest/tests-tools/yb-admin-test --gtest_filter='*GetTableHash*:*ListTablets*Hex*' Lint clean. Co-authored-by: Cursor <cursoragent@cursor.com> Reviewers: hideaki.kimura, jhe Reviewed By: hideaki.kimura Subscribers: ybase Differential Revision: https://phorge.dev.yugabyte.com/D57082
| Commit: | 43568c7 | |
|---|---|---|
| Author: | William Wang | |
| Committer: | William Wang | |
[#33156] YSQL: Dynamic History Retention for sys catalog Summary: Dynamic DocDB history retention (#32314) currently only covers tablets hosted on tservers. Reads served by the **master's sys catalog tablet** (i.e. all YSQL catalog reads) are not covered, so a long-running transaction or DDL can still fail with Snapshot too old on a catalog read even while its user tablets are protected. We now consume the aggregated pin map on the master and merge it into the sys catalog cutoff, bounded by the original safety window and the hard cap used on tservers (defaulted to 24h). This aggregated pin is persisted through a new singleton sys catalog row (`SysRowEntryType::HISTORY_RETENTION_PIN` / `SysHistoryRetentionPinEntryPB`) that stores the raw oldest pinned read time. Raft-replicated like other sys catalog row. This aggregated pin is also cached in memory, in an atomic on HistoryRetentionPinInfo so compaction can read it without taking the cow lock (the writer holds that lock across the sys catalog write, which can wait on compaction). `CatalogManagerBgTasks::RunOnceAsLeader`(~1s) calls `PersistYsqlHistoryRetentionPin`, which writes to sys catalog only when the aggregated pin differs from the previously published value. It is only called by the master leader since the followers do not get heartbeat and cannot compute the aggregated pins. The followers runs the bg-task non-leader branch, which calls `RefreshYsqlHistoryRetentionPin` each loop to refresh the aggregated pin calculated by the master leader, meaning the staleness of the pin id worst-case one bg-task interval. Given the defaulted 4h retention window for sys catalog, there is plenty of time to update the pin before compaction kicks in. **Upgrade/Rollback safety:** A new `SysRowEntryType` has been added with a new PB message used to serialize and write the pin value to sys catalog. Any version mismatch between masters will default back to the previous behaviour since the persisted sys catalog pin would either be completely ignored or no value read at all, and no special handling is needed. Test Plan: `./yb_build.sh release --cxx-test master-test --gtest_filter 'MasterTest.SysCatalogHistoryCutoff*'` Reviewers: kfranz, xCluster, hsunder, sanketh Reviewed By: kfranz Subscribers: ybase, yql Differential Revision: https://phorge.dev.yugabyte.com/D57262
| Commit: | 8acb41b | |
|---|---|---|
| Author: | skhilar | |
| Committer: | skhilar | |
[PLAT-21147] Add support for setting up federated IAM for AWS VMs during Universe provision/configure via YNP Summary: Enable YB‑Controller to back up/restore to GCS from AWS‑backed DB nodes using each node's own AWS IAM identity via GCP Workload Identity Federation (WIF) — no static GCP keys on the nodes (PLAT‑21147). Supported for AWS and on‑prem providers (on‑prem VMs running on AWS). Config lives on the provider (no runtime flags): the AWS provider gets an "Enable Federated IAM" toggle + an audience field (the WIF pool/provider audience), mirroring the IAM toggle; on‑prem supplies the same enable + audience via the YNP yaml. The GCS storage config stays a plain USE_GCP_IAM config. Node setup: on universe create/edit/add‑node/replace‑node, YBA fans out a per‑node ManageCloudFederation subtask over node‑agent (AWS/on‑prem nodes only). YBA passes only the audience; the node renders the external_account creds + env + a yb-controller.service.d systemd drop‑in under ~/.yugabyte/ and restarts YBC once. No refresher daemon — the GCP client mints/refreshes short‑lived tokens in‑process. Persisted state: UserIntent.federationConfigured is set true only when every node in the universe is configured (avoids a mixed‑node edge case); edit/add‑node/replace‑node re‑apply only when it's set. Retrofit existing universes: a new v2 API POST .../universes/{uniUUID}/cross-cloud-federation {enabled} enables/disables federation on a running universe — enable precheck rejects when no audience resolves, fans out to all nodes, and persists the flag; disable stops applying config but does not tear down existing on‑node artifacts (v1). YBA side: backup preflight and backup deletion use YBA's own AWS identity (in‑process WIF from the same audience); the audience is snapshotted on each backup so deletes work even after the universe/provider is gone. Config‑time GCS checks are skipped for USE_GCP_IAM configs (no node context) and defer to YBC only when YBA genuinely can't federate — real GCS errors still fail. Scope: GCS‑on‑AWS only. S3‑on‑GCP (PLAT‑21148) is a follow‑up. Test Plan: **Unit**: GCPUtilTest, CustomerConfigValidatorTest, BackupsControllerTest pass (federation skip paths, USE_GCP_IAM validation, backup‑create validation). **Manual E2E — GCS‑on‑AWS, both provider type**s: - AWS provider (YBA on AWS IAM, AWS universe): enable Federated IAM + set audience on the provider → create universe → federation files land under ~/.yugabyte/ on all nodes → backup + restore succeed. - On‑prem provider (YBA on AWS IAM, on‑prem VMs on AWS, node‑agent/YNP‑provisioned): enable + audience via YNP yaml → same flow, incl. manually‑provisioned nodes → fan‑out configures them → backup + delete succeed (validated on AlmaLinux). Ubuntu surfaced a node‑side CA‑cert path issue (CentOS‑built YBC vs Debian cert paths) — a provisioning gap, not this change. - Retrofit existing universe: v2 cross-cloud-federation API {enabled:true} → all nodes configured + federationConfigured persisted → backup succeeds; {enabled:false} → stops applying (existing artifacts left in place). - Federation not enabled / audience unset → fan‑out is a no‑op (nothing written); enable precheck rejects with 400. Reviewers: #yba-api-review, nsingh, vkumar, yshchetinin, kkannan Reviewed By: #yba-api-review, nsingh, kkannan Subscribers: jmak, kkannan, svc_phabricator, yugaware Differential Revision: https://phorge.dev.yugabyte.com/D56165
| Commit: | 34cae40 | |
|---|---|---|
| Author: | William Wang | |
| Committer: | Angela Xu | |
[#33521] DocDB: Persist and Load Cluster YSQL DB Pin on Tserver Restart Summary: Depends on D57026 Tservers reads from `cluster_ysql_db_oldest_pinned_read_times_` to determine history retention cutoff, but that map is populated only from `TSHeartbeatResponsePB`. On tserver restart, it is empty until the first heartbeat response arrives, and compactions become eligible even earlier, before the heartbeater starts. During that window, per-database pins offer no protection, so a long-running transaction on another tserver could have its snapshot compacted away. This change has the tserver persist the cluster pin map to local disk on a timer (new flag `ysql_db_history_retention_pins_persist_interval_sec`, default 60s), and reload it on startup before tablets open, closing the gap. Transactions that start between the last persist and the restart are absent from the reloaded map, but they're still covered by the existing safety window. A missing or unreadable file just leaves the current behavior in place (compaction only held back by retention window). **Upgrade/Rollback safety:** A new PB message is added to persist the global cluster pin map to a file on disk, and is not used elsewhere. No upgrade/rollback handling needed. Test Plan: `./yb_build.sh release --cxx-test ts_tablet_manager-test --gtest_filter 'PersistedDbHistoryRetentionPinsTest.*'` Reviewers: kfranz Reviewed By: kfranz Subscribers: angela.xu, ybase, yql Differential Revision: https://phorge.dev.yugabyte.com/D57425
| Commit: | 86ef890 | |
|---|---|---|
| Author: | William Wang | |
| Committer: | Angela Xu | |
[#33204] DocDB: Stop updating db history retention pins on master failover Summary: Currently, on master failover, the new master will have an initial empty map of global history retention pins, and can send back incomplete maps before a heartbeat is received from all live tservers. We wish to load the previous master's list of live tservers from persist_tserver_registry on the new master's initialization and stop updating the global cluster pins on tservers until the new master has received 1 heartbeat from every live tserver (or until tserver gets dropped for not heartbeating master, whichever one comes first). Any subsequent pins get updated as normal. **Upgrade/Rollback safety:** The `cluster_ysql_db_pins_ready` field is added to `TSHeartbeatResponse` indicating whether tservers should update their map, and no special handling is needed: an old master does not set the field and tserver defaults it to true and apply the map immediately, identical to previous behavior. An old tserver ignores this field and behaves the same way. Test Plan: `./yb_build.sh release --cxx-test ts_tablet_manager-test --gtest_filter 'ComputeDbHistoryRetentionPinCutoffTest.PinsNotReadyKeepsPreviousClusterPins'` `./yb_build.sh release --cxx-test master-test --gtest_filter 'MasterTest.YsqlDbPinsWaitForAllLiveTserversAfterRestart'` Reviewers: kfranz Reviewed By: kfranz Subscribers: zdrudi, angela.xu, smishra, ybase, yql Differential Revision: https://phorge.dev.yugabyte.com/D57026
| Commit: | d55f928 | |
|---|---|---|
| Author: | John Meehan | |
| Committer: | John Meehan | |
[#33444] DocDB: Fence demotion of the deferred-verification capability flag The one unguarded downgrade path for deferred unique-index verification: binary downgrade to a pre-feature binary is already hard-fenced by the AutoFlags config validation while ysql_enable_deferred_unique_index_verification stays promoted, and RollbackAutoFlags refuses kLocalPersisted flags outright -- but DemoteSingleAutoFlag validates nothing beyond flag existence. Demoting the capability flag clears the startup fence while SKIP_ALL jobs or unflushed marked WAL entries exist; an old binary then bootstraps and replays marked writes with positional write IDs -- byte-divergent replicas. DemoteSingleAutoFlag now refuses to demote this flag (the demote-side first, mirroring the promote-side major-upgrade veto) while either holds: - Any table has an active SKIP_ALL backfill job (sys-catalog scan): its chunks may still be appending marked writes. - Any registered tserver reports marked-write state via the new CheckIndexBackfillDowngradeSafety admin RPC: a tablet with an active ordering generation, or a released one whose marked WAL entries are not yet below the regular-DB flushed frontier. Sequential synchronous fan-out -- demotion is a rare admin operation. An unreachable tserver FAILS the demotion (it may hold config-member replicas with unflushed marked WAL; wait for it or decommission it), with one exception: a node dead past follower_unavailable_considered_failed_sec is skipped -- its replicas have been evicted and re-replicated, and a returning node's stale local replay is confined to data that is tombstoned before serving. The probe itself skips tombstoned/deleted peers (WAL removed) and completed split parents (flushed before the split checkpoint), so husks on the master's cleanup schedule cannot block a drain. The per-tablet condition is made provable and convergent by two release changes: - Release persists released_marked_write_watermark in the generation superblock record: the highest marked-write Raft index the tablet has applied (tracked in memory, monotonic, re-seeded by bootstrap replay, which re-applies exactly the marked writes still above the flushed frontier). Anchoring to the last marked WRITE rather than the release op matters: metadata ops never advance the RocksDB frontier, so a release-op watermark would never converge on an idle tablet. - Release schedules an async regular-DB flush (FlushReason::kIndexBackfillGenerationRelease), pushing the marked writes below the frontier within seconds instead of waiting for an organic flush. The released record now survives in the superblock (previously only active generations were persisted) until the next activation. The tablet validator's straggler heal records the watermark too, deferring its heal to a later cycle if the tablet is not running (the watermark lives on the Tablet). The error text carries the drain procedure: finish or abort the builds, let the release flushes converge, retry. The veto is deliberately not bypassable: every blocking condition drains on its own, and demoting past it is precisely the divergence hazard. Test Plan: ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillVerifier.DowngradeFenceBlocksWhileGenerationsHeld/0' (release blocked by the fixture: demotion refused, message names the active ordering generation) ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillSkipAllRaftOrdering.DowngradeFenceBlocksWhileJobActive/0' (job held at DoBackfill: demotion refused before any tserver probe, message names the SKIP_ALL job's table) ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillVerifierReleased.DowngradeFenceClearsAfterRelease/0' (after the funnel releases and the release flush converges, demotion goes through -- this test caught the release-op-watermark convergence bug during development) ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillVerifierReleasedNoFlush.DowngradeFenceBlocksWhileMarkedWalUnflushed/0' (deterministic pin of the unflushed arm: release flush disabled via test flag, fence refuses with the not-yet-flushed message) ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.IndexBackfillOrderingGenerationPersistsAcrossReload' (extended: the released watermark survives reload; a never-released clear still persists as field absence) Regressions: full tablet_peer-test (27), full fixed_hybrid_time_write_id-itest (7), funnel release e2e (OrderingGenerationActivatedAndReleased), VerifyAfterReleaseFailsCleanly, RetentionPinMetricsClearOnRelease -- all green. Assisted-By: devx/08801cf6-2854-4c6f-a378-1d98d239b22d --- _automated · Claude Fable 5 (opencode)_
| Commit: | 2e95ed1 | |
|---|---|---|
| Author: | John Meehan | |
| Committer: | John Meehan | |
[#33444] DocDB: Add a keyed, per-job violation fingerprint to unique-index verification Verification failure reports have been value-free by design: encoding classes and counts, never the duplicate key. That satisfies the privacy contract but leaves support blind to correlation -- "is this the same duplicate as the last retry?" This adds the design document's keyed, scope-limited diagnostic fingerprint: a non-reversible token identifying the violating index key, safe for low-cardinality values because confirming a candidate requires the key, and the key's scope limits how long confirmation is possible at all. Mechanics: - Per-job secret (32 bytes, OpenSSL RAND_bytes) generated in the BackfillTable constructor for SKIP_ALL jobs (other modes never verify and carry no key), persisted in BackfillJobPB (field 9, latched-per-job like the mode and the gate), and deleted with the job: fingerprints correlate across the job's retries and master failover, and are unconfirmable once the job is gone. - The verifier exposes the first violating group's raw encoded DocKey in-process only (excluded from ToString; never serialized). The tserver handler renders the token -- "<jobtag8hex>-<digest16hex>", both truncated HMAC-SHA256 under the job key (yb/util/keyed_fingerprint.{h,cc}) -- into a structured response field and a "fingerprint=" suffix on the still-value-free reason. The job tag is derived from the key alone, so tokens from different jobs are visibly incomparable. - The reason rides the existing paths into the persisted verification state, the master outcome log, and (in gating mode) backfill_error_message -> the CREATE INDEX error. The structured token is also persisted (UniqueIndexVerificationStatePB field 5) for support tooling. - The key never leaves the master/tserver trust domain: GetBackfillJobs scrubs it from RPC responses (pinned by test), both request-log sites (master send VLOG, tserver receive DVLOG) share RedactedDebugString -- which also elides start_key, a raw index DocKey -- and the one log that printed a whole BackfillJobPB clears the key first. Test Plan: ./yb_build.sh release --cxx-test keyed_fingerprint-test --gtest_filter 'KeyedFingerprintTest.TokenFormatAndDeterminism' ./yb_build.sh release --cxx-test keyed_fingerprint-test --gtest_filter 'KeyedFingerprintTest.TagIdentifiesKeyScope' (token format, determinism under one key, shared scope tag per key, cross-key incomparability) ./yb_build.sh release --cxx-test unique_index_verifier-test --gtest_filter 'UniqueIndexVerifierTest.ViolationExposesGroupPrefixInMemoryOnly' (the privacy pin: the raw group prefix is exposed in memory for the caller and excluded from ToString) ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillFailClosedVerification.ViolationFailsCreateIndex/0' (extended: the CREATE INDEX error now carries "fingerprint=<token>"; verified rendering end-to-end: fingerprint=2d929cea-dbf259e3ff0d7297) Full unique_index_verifier-test 20/20. Regressions: shadow violation (waiter matches the enriched line), shadow clean multi-tablet, failover resume -- all green. Assisted-By: devx/08801cf6-2854-4c6f-a378-1d98d239b22d --- _automated · Claude Fable 5 (opencode)_
| Commit: | 1d70d52 | |
|---|---|---|
| Author: | John Meehan | |
| Committer: | John Meehan | |
[#33444] DocDB: Add verification scan accounting and master outcome counters Verification/job observability for deferred unique-index verification, plus the accounting fix deferred from the verifier-core review: versions_scanned was inconsistent on the bounded-memory fallback path (the forward pass stopped counting at buffer overflow, then the reverse walk re-counted its re-visits -- over- and under-counting the same group). The forward pass now owns the accounting -- every physical version in a scanned group is counted exactly once regardless of replay path -- and the new fallback_groups counter records the reverse walk's extra work instead (a fallback group is read roughly twice). Plumbing and aggregation: - VerifyUniqueIndexTabletResponsePB gains fallback_groups (field 8); the handler fills it from the verifier result. - The shadow coordinator aggregates per-index scan totals from every response and logs them with the recorded outcome: "... VERIFY_CLEAN [observational] (dockey_groups=300, versions=300, fallback_groups=0)". - Five cluster-level master counters (backfill_aborted pattern): unique_index_verification_outcome_{clean,violation,inconclusive} (one per recorded per-index outcome) and unique_index_verification_{versions_scanned,fallback_groups} (aggregated from responses; late responses dropped by the single-winner guard still count -- the scan happened). A persistently high fallback_groups rate is the signal to raise unique_index_verify_max_buffered_versions_per_group; the counter description says so. Test Plan: ./yb_build.sh release --cxx-test unique_index_verifier-test --gtest_filter 'UniqueIndexVerifierTest.FallbackAccountingMatchesBuffered' (the accounting pin: identical data verified through the buffered and fallback paths reports identical versions_scanned and dockey_groups_scanned; the fallback run reports fallback_groups = 1) ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillShadowVerification.CleanOutcomeRecordedMultiTablet/0' (extended: asserts the master outcome counter == 1 -- exact on a fresh cluster, catches double-counting -- and versions_scanned > 0 via the cluster metric entity) Full unique_index_verifier-test 19/19. Regressions: ViolationRecordedButDoesNotBlockPublication (waiter matches the enriched log line), ViolationFailsCreateIndex (gating path through Record), failover resume, paginated clean -- all green. Assisted-By: devx/08801cf6-2854-4c6f-a378-1d98d239b22d --- _automated · Claude Fable 5 (opencode)_
| Commit: | 2a0a4df | |
|---|---|---|
| Author: | John Meehan | |
| Committer: | John Meehan | |
[#33444] DocDB: Gate unique-index publication on the deferred verification outcome The last functional piece of deferred uniqueness verification: behind ysql_index_backfill_fail_closed_verification (default false), the verification phase's outcome decides publication instead of merely being recorded. The publication order becomes the design document's: backfill completes -> stay at DO_BACKFILL -> verify all index tablets -> persist the outcome -> RWD -> publish birth_time -> release generation + retention -> indisvalid In gating mode, backfill-success marking for unique indexes is deferred from chunk completion to verification completion: a clean index is then marked successful (including the xCluster backfill-completed notification, which must not fire for an index that may still fail) and proceeds through the terminal funnel to READ_WRITE_AND_DELETE and birth_time publication. A VIOLATION or unresolvable INCONCLUSIVE outcome fails the index through the existing backfill failure path -- backfill_error_message carries a value-free reason, CREATE INDEX returns it, the index is never READ_WRITE_AND_DELETE and never indisvalid, and the existing invalid-index cleanup applies. Coordinator failures fail closed too: whatever is still IN_PROGRESS when the phase fails is exactly the unverified set (the in-flight index, indexes the phase never reached, and indexes never selected because resolution itself failed) and is failed rather than published. Non-unique indexes in the job have nothing to verify and are marked successful at chunk completion as before. The gating decision is latched per job and persisted in BackfillJobPB, mirroring the uniqueness-check mode: failover, retries, and runtime flag flips cannot reinterpret an active job. Reading the flag live at each decision point would misbehave in both directions -- gating->observational after the deferred marking leaves the index IN_PROGRESS at the funnel; observational->gating mid-phase would fail an index whose success (and xCluster notification) already fired. The funnel logs DFATAL if it ever sees an IN_PROGRESS index (only reachable if terminal-state marking itself kept failing); the pre-existing mapping publishes such an index as failed. The gate implies the verification phase runs even when ysql_index_backfill_shadow_verification is off; with both flags off, nothing changes. Observational mode is untouched -- outcome logs now carry a [gating] / [observational] tag. The mode selector remains hardcoded to CHECK_ALL, so the gate is production-unreachable until activation; tests drive it through the master-side TEST override. Test Plan: ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillFailClosedVerification.CleanBuildPublishes/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillFailClosedVerification.ViolationFailsCreateIndex/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillFailClosedInconclusive.InconclusiveFailsCreateIndex/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillFailClosedPhaseFailure.CoordinatorFailureFailsCreateIndex/0' The violation test is the feature's end state: a real pre-existing duplicate survives the SKIP_ALL build and fails CREATE INDEX with a verification error; pg_index shows indisvalid = false, the base table stays usable, and the invalid index is droppable. The inconclusive test pins fail-closed = "not proven clean" via a forced-INCONCLUSIVE tserver response. The clean test pins publication: indisvalid = true and a consistent index. The phase-failure test pins the coordinator-failure path: an injected resolution failure fails CREATE INDEX with the unverified-index message on clean data. Regressions: shadow suite (violation still publishes in observational mode; multi-tablet clean; failover resume, which now also exercises the persisted gating latch on resume), skip_all funnel e2e, Drop -- all green. Assisted-By: devx/08801cf6-2854-4c6f-a378-1d98d239b22d --- _automated · Claude Fable 5 (opencode)_
| Commit: | 16f1287 | |
|---|---|---|
| Author: | John Meehan | |
| Committer: | John Meehan | |
[#33444] DocDB: Add the shadow verification coordinator with durable job state The first production caller of the verification RPC: after a SKIP_ALL unique-index backfill completes -- while the ordering generations are still active, before the terminal funnel releases them -- the master runs the deferred uniqueness verification scan observationally. The outcome is persisted and logged but never gates index publication; the fail-closed verify-before-publication gate is a separate, later capability. Disabled by default (ysql_index_backfill_shadow_verification). Durable state (BackfillJobPB.unique_index_verification, keyed by unique-index table id): verification state, verify_upper_ht, tablets confirmed clean, and a value-free reason. verify_upper_ht is chosen ONCE from cluster hybrid time (never wall clock) and persisted before any scan -- retries and failover resume verify the same window. Clean tablets are durable progress at tablet granularity; finer intra-tablet resume keys are deliberately not persisted, since re-scanning one tablet after failover is idempotent. All persistence goes through the indexed table's write lock and a sys-catalog upsert under the job's leader epoch. Coordinator mechanics (VerifyUniqueIndexForTablet, the GetSafeTimeForTablet task shape): - One index at a time; bounded per-tablet concurrency (index_backfill_shadow_verification_max_concurrent_tablets). - Pagination: index_backfill_verify_dockey_groups_per_rpc caps each RPC's group budget; a tablet resumes from the returned resume key without touching the join counts. - Short-circuit under a single-winner protocol: the terminal check and the launch-next / index-done decision are one critical section, so a CLEAN callback racing a short-circuiting VIOLATION can neither overwrite the recorded outcome nor double-advance the phase; the slow clean-tablet persist happens after the join decision. Late responses cannot change a recorded outcome. The join guards on job failure, not done() -- the whole phase runs in the job's kSuccess state. - Fail-open on coordinator errors: the phase is on the CREATE INDEX critical path, so every coordinator failure (task launch, persistence, pagination resume) degrades to VERIFY_INCONCLUSIVE and continues into publication -- no failure may strand the job short of the terminal funnel. The tserver returns typed error codes for deterministic failures (generation mismatch, non-unique-index tablet, cutoff violations) so the coordinator does not burn its retry budget on requests that cannot succeed. - An index with no tablets records VERIFY_INCONCLUSIVE: an empty tablet set is not evidence of uniqueness. - Master failover mid-phase resumes with the persisted window. Generic lost-backfill resume (load-time enqueue, resuming a job whose table has left ALTERING) is upstream since #33207/c830a71489 -- this part originally carried equivalent machinery and now relies on upstream's, adding only the verification-phase awareness: a resumed job whose indexes are all SUCCESS skips straight to the completion phase (and from there into verification) instead of relaunching empty backfill chunks. - No generation_base_op_index in the requests, on purpose: the base is per-tablet (the activation op's own Raft index) and a re-activated generation's higher base is still the generation this job's marked writes live under; active + index-table match is the semantic guard. Each verification RPC gets its own deadline (index_backfill_verify_rpc_timeout_ms, default 60s), mirroring the backfill chunks' ysql_index_backfill_rpc_timeout_ms instead of inheriting the generic 30s master_ts_rpc_timeout_ms: the deadline sizes one page of the scan (the tserver stops a grace margin early and returns a resume key), and pages should be sizable independently of unrelated master RPCs. Test Plan: ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillShadowVerification.CleanOutcomeRecordedMultiTablet/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillShadowVerification.ViolationRecordedButDoesNotBlockPublication/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillShadowVerificationPaginated.CleanAcrossManyRpcs/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillVerifier.GenerationMismatchFailsCleanly/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillShadowVerificationFailover.ResumesWithPersistedWindowAcrossFailover/0' The violation test is the shadow-semantics pin: a real duplicate is recorded (NOT CLEAN: VERIFY_VIOLATION in the master log) while CREATE INDEX still succeeds and the index serves reads. The multi-tablet test exercises the fan-out/join; the paginated fixture forces one DocKey group per RPC through the resume path. GenerationMismatchFailsCleanly covers the 4b review's L1 (table-id and base mismatches fail without scanning). The failover test holds the verification RPCs open with a retryable rejection, steps the master leader down mid-phase, and asserts the new leader resumes with the persisted window before completing the build. ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillShadowVerificationDeadlinePaginated.CleanAcrossDeadlinePages/0' The deadline-paginated test is the regression for the measurement-run defect: a slowed scan against a 3s verify budget and 1s grace margin must complete CLEAN through several deadline-bounded pages; before the margin (4b) and this budget existed, the page response always arrived after the coordinator's deadline and the tablet failed INCONCLUSIVE at the retry ceiling. Regressions: shadow-flag-off skip_all e2e, CHECK_ALL mode test, 4b verifier e2e -- all green. The failover-resume e2e now passes through upstream #33207's resume path. Assisted-By: devx/08801cf6-2854-4c6f-a378-1d98d239b22d --- _automated · Claude Fable 5 (opencode)_
| Commit: | 16d9f28 | |
|---|---|---|
| Author: | John Meehan | |
| Committer: | John Meehan | |
[#33444] DocDB: Expose the unique-index verification scan as a tserver admin RPC Wires the verifier core to the world: VerifyUniqueIndexTablet on the tablet server admin service, read-only and paginated, with the preconditions the scan's correctness rests on enforced at the entry point. No production caller yet -- the verification coordinator drives it in a later part; end-to-end tests drive it directly. The RPC (leader-only, BackfillIndex-style pagination): - Discharges the applied-intents precondition in two steps: waits for the tablet's safe time to pass verify_upper_ht (leader lease held; no future write can land at or below the bound), then TransactionParticipant::ResolveIntents(verify_upper_ht) -- safe time alone does not imply applied, since transactional apply is asynchronous, and a committed-but-unapplied foreground write is invisible to a regular-DB-only scan (the record missed could be exactly the foreground half of a duplicate). - Validates the active ordering generation against the request's expectation (index table, and base_op_index when set -- the coordinator always sets it; omitting it relaxes to "any active generation for this table" for test tooling). A released or replaced generation fails the RPC without scanning: the window no longer describes the tablet's marked writes. - Refuses a window that begins below the *applied* history cutoff, read from the regular DB's flushed frontier -- the cutoff past compactions actually used. Deliberately not the retention policy's directive: GetRetentionDirective mutates committed cutoff state (unacceptable on a read path) and only bounds future compactions, which the retention-hold part of this feature prevents from advancing anyway. Until that hold lands the pre-scan check is advisory, so the applied cutoff is re-checked after the scan: a compaction advancing it mid-scan fails the call rather than returning a result computed over possibly-incomplete history. Deterministic failures (generation mismatch, non-unique-index tablet, invalid or inverted window) return typed error codes, so callers do not burn a retry budget on requests that cannot succeed. The scan holds the RocksDB shutdown guard for its duration (the not-blocking variant: tablet shutdown aborts the scan, never the reverse), and an unset backfill_read_ht is rejected at the boundary -- it decodes to the invalid sentinel, which compares greater than any real hybrid time and would make the history-cutoff checks pass vacuously. Tablet::VerifyUniqueIndex resolves the scan options from the tablet: the primary table's ybidxbasectid column, the tablet metadata as the schema-packing provider, the regular DB within the tablet's key bounds, and the new runtime flag unique_index_verify_max_buffered_versions_per_group. TEST_block_index_backfill_ordering_generation_release (master) holds the terminal funnel's release open so tests can verify an index built by the real backfill path while its generation is still active. The e2e suite closes the encoding-drift loop the verifier-core review required: every physical shape asserted by the unit tests is re-asserted against tablets written by the production SKIP_ALL backfill and live pg DML. Deadline grace margin (found by the measurement runs): the scan paginates at the raw client deadline, so a deadline-bounded page's response was serialized exactly when the coordinator's RPC expired -- always discarded, with every retry rescanning the same page from scratch until the retry budget failed the tablet INCONCLUSIVE (deterministically ~1176s under the default 30s timeout and 20-retry budget; any tablet whose scan exceeds one RPC deadline hit it). The handler now stops unique_index_verify_deadline_grace_margin_ms (default 2s, clamped to half the remaining budget) before the client deadline, mirroring the backfill chunk path's backfill_index_timeout_grace_margin_ms. A new TEST_unique_index_verify_delay_per_group_ms flag makes deadline pagination deterministically reproducible with small data. Test Plan: ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillVerifier.CleanBuildWithForegroundDml/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillVerifier.PkUpdateAcrossBackfillBoundaryIsClean/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillVerifier.DetectsPreexistingDuplicates/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillVerifierPackedUpdate.PkUpdateAcrossBackfillBoundaryIsClean/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillVerifierReleased.VerifyAfterReleaseFailsCleanly/0' The PK-update regression runs under both physical shapes (standalone ybidxbasectid column with default flags; packed V2 + kIsUpdateFlag with pack_full_row_update + mark) and must be Clean in both. Regressions: unique_index_verifier-test (18), OrderingGenerationActivatedAndReleased (release still fires with the TEST flag off), OrderingGenerationChangeMetadataOp tablet_peer test -- all green. Assisted-By: devx/08801cf6-2854-4c6f-a378-1d98d239b22d --- _automated · Claude Fable 5 (opencode)_
| Commit: | 0065b32 | |
|---|---|---|
| Author: | John Meehan | |
| Committer: | John Meehan | |
[#33444] DocDB: Orchestrate the index-backfill ordering-generation lifecycle from the master The persisted ordering generation (#33580) had no production lifecycle: nothing activated it before marked backfill writes, nothing released it, and the master-side split suppression protecting a backfill is in-memory only (lost on failover). This part wires the full lifecycle and flips marked-write gating fail-closed. Activation rides ChangeMetadataOperation (a new `ChangeMetadataRequestPB.index_backfill_ordering_generation` variant, applied via `Tablet::UpdateIndexBackfillOrderingGeneration` exactly like `mark_backfill_done`): Raft-replicated, follower-applied, WAL-replayed, and `base_op_index` is the activation operation's own Raft index, so every replica and every replay derives the same base with no master-side bookkeeping. Re-activation (failover resume) is idempotent -- the base only moves up. Master flow, SKIP_ALL jobs on PGSQL indexed tables only (the mode rides YSQL chunk requests; YCQL backfill writes are never marked, so generations would only fence their splits without protecting anything): - Before the first chunk (`DoBackfill`), the job disables and drains index-table splitting, then fans out one waited `UpdateOrderingGenerationForTablet` task per index tablet (the `GetSafeTimeForTablet` join pattern). All acks -> chunks launch; any failure -> the job aborts before a single marked write exists, and CREATE INDEX fails cleanly. The drain closes the activation/split TOCTOU: consensus already rejects CHANGE_METADATA_OP while a split is pending, the generation fence rejects splits after activation applies, and a split appended in the append-to-apply window can at worst fail the activation task -- never corrupt a scan set that marked writes exist in. - Release is fire-and-forget from the terminal funnel (`UpdateIndexPermissionsForIndexes`), which every success/failure/abort path traverses. The tserver metadata validator converges stragglers, mirroring the retain_delete_markers machinery: tablets with an active generation join its existing GetBackfillStatus poll, and the generation is released locally (no Raft) once the master reports a terminal state. `IndexStatusPB` gains `BACKFILL_FAILED` (removal-path index permissions map to it) so a failed SKIP_ALL job whose funnel release was lost cannot leave an orphaned invalid index permanently split-fenced and retention-pinned; the retain_delete_markers heal itself stays success-only. - The tablet-split manager refuses to split an index whose indexed table has a durable SKIP_ALL backfill job. The manager's table validation takes no catalog locks itself -- the indexed table is a caller-resolved parameter, because the manual-split path enters through ValidateSplitCandidateUnlocked with the catalog mutex already held (a recursive shared acquisition is fatal under lock_debug) (`SysTablesEntryPB.backfill_jobs`, survives failover, cleared in the funnel), with an INFO-level skip reason -- today's two in-memory suppressions cover only the indexed table and evaporate on failover. The tablet-side generation fence stays as the fail-closed layer. Gating flip: `WriteOperation::ValidateLeaderOpId` now rejects marked writes with *no* active generation (previously deferred) -- a marked write outside a generation would store versions nothing tracks or releases (e.g. a stale chunk retry after the job's terminal state). Tests that drive marked writes directly now activate a generation first. `write_id_floor_version` numbering lands as `kIndexBackfillWriteIdFloorVersion = 1`. The mode selector remains hardcoded to CHECK_ALL (#33484), so this whole flow stays production-unreachable; tests drive it through the master-side TEST override. Test Plan: ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.OrderingGenerationChangeMetadataOpSetsBaseFromOwnRaftIndex' ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.FixedHybridTimeWriteRejectedAtOrBelowGenerationBase' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillSkipAllRaftOrdering.OrderingGenerationActivatedAndReleased/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillSkipAllActivationFailure.ActivationFailureFailsCreateIndexCleanly/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillSkipAllBlocked.SplitFencedDuringBackfill/0' Retrofitted (marked writes now require an active generation): ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.FixedHybridTimeWrite*' ./yb_build.sh release --cxx-test fixed_hybrid_time_write_id-itest Regressions: skip_all e2e + marker canary through the real activation flow, uniq-idx-1 mode tests, DuplicatesExistBeforeBackfill, 3b-i generation tests, 0a/0b abort tests, and drop-path tests for the permissions mapping change (PgIndexBackfillTest.Drop, PgIndexBackfillFastClientTimeout.DropWhileBackfilling) -- all green. Assisted-By: devx/08801cf6-2854-4c6f-a378-1d98d239b22d --- _automated · Claude Fable 5 (opencode)_
| Commit: | 944e3d8 | |
|---|---|---|
| Author: | John Meehan | |
| Committer: | John Meehan | |
[#33444] DocDB: Add persisted index-backfill ordering-generation record and tablet fences Deferred uniqueness verification scans the marked versions that unique-index backfill wrote at the fixed backfill hybrid time. That scan is only sound against a stable tablet set and a validated write-ID sequence, so each index tablet needs a durable record of the in-flight backfill's ordering state -- one that survives restart, remote bootstrap, and inheritance by split children. The record (`IndexBackfillOrderingGenerationPB`, superblock field 39, persistence following the `cdc_sdk_safe_time` pattern): `active`, the owning index `table_id`, `base_op_index`, `retention_barrier_ht`, and `write_id_floor_version` (numbering pinned in the proto: 1 = the initial floor, 0 = absent/unknown). The full combined shape is defined here; `retention_barrier_ht` is consumed by the verification read fence and identity/activation by the master orchestration, both in later parts. The durable form of an inactive generation is field absence: `LoadFromSuperBlock` clears the in-memory record when the field is absent (so a remote-bootstrap superblock replacement from a source without a generation cannot resurrect a released one), and the setter normalizes inactive input to the default -- the caller contract is "set an active generation or set {}", keeping in-memory and durable state identical. While a generation is active on a tablet: - Marked fixed-hybrid-time writes must carry Raft indexes strictly above `base_op_index` (`WriteOperation::ValidateLeaderOpId`). The base is the activation operation's own Raft index, so an index at or below it cannot belong to this generation -- it would indicate misrouting or replay of a foreign sequence. (Rejecting marked writes when *no* generation is active arrives with the master activation flow; until then marked writes remain production-unreachable.) - Tablet splitting is rejected before Raft append (`SplitOperation::ValidateLeaderOpId`) -- the fail-closed layer under the master-side split fence (next part). The master's split path retries and succeeds once the generation is released. - Tablet cloning is rejected the same way. Clone targets build fresh metadata (`RaftGroupMetadata::CreateNew`), not a superblock copy, so a clone would carry hard-linked snapshot data containing generation-scoped marked write IDs with no generation record describing them and no backfill job tracking their lifecycle. `TEST_bypass_index_backfill_ordering_generation_split_fence` forces the split fence open to demonstrate the property the fence is *not* needed for: split children inherit the parent's marked entries and continue drawing write IDs from their own Raft indexes, so per-key write-ID uniqueness holds even across a split (the fence exists for verification-scan stability and lifecycle tracking, not key correctness). Test Plan: ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.FixedHybridTimeWriteRejectedAtOrBelowGenerationBase' ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.IndexBackfillOrderingGenerationPersistsAcrossReload' ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.IndexBackfillOrderingGenerationClearedBySuperblockReplacement' ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.SplitRejectedWhileOrderingGenerationActive' ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.CloneRejectedWhileOrderingGenerationActive' ./yb_build.sh release --cxx-test fixed_hybrid_time_write_id-itest --gtest_filter 'FixedHybridTimeWriteIdITest.MasterSplitRefusedWhileOrderingGenerationActive' ./yb_build.sh release --cxx-test fixed_hybrid_time_write_id-itest --gtest_filter 'FixedHybridTimeWriteIdITest.OrderingGenerationSurvivesRemoteBootstrap' ./yb_build.sh release --cxx-test fixed_hybrid_time_write_id-itest --gtest_filter 'FixedHybridTimeWriteIdITest.SplitWithFenceBypassedPreservesPerKeyWriteIdUniqueness' Regressions: 3a fixed-hybrid-time tablet_peer tests and leader-change itest, 0a/0b abort tests, PgIndexBackfillSkipAllRaftOrdering e2e -- all green. Assisted-By: devx/08801cf6-2854-4c6f-a378-1d98d239b22d --- _automated · Claude Fable 5 (opencode)_
| Commit: | 2d761f6 | |
|---|---|---|
| Author: | John Meehan | |
| Committer: | John Meehan | |
[#33444] DocDB: Add marked fixed-hybrid-time write path with Raft-index write IDs SKIP_ALL unique-index backfill writes run no uniqueness checks, so distinct candidates sharing one unique-index key at the fixed backfill hybrid time must survive as distinct physical versions -- with the existing positional per-batch write IDs, both land at (key, backfill_ht, 0) and one silently overwrites the other, destroying the evidence deferred verification exists to inspect. Replicated write batches gain a self-describing marker (KeyValueWriteBatchPB.use_raft_index_for_write_id). For a marked batch, every replica derives the storage write ID from the Raft operation index as kBackfillWriteIdFloor | index -- one derivation site in Tablet::ApplyOperation covers leader apply, follower apply, and WAL-replay bootstrap, so replicas converge byte-identically and replay never consults runtime flags. The floor (an alias of kIntraTxnWriteIdLimit, #33499) keeps the marked domain disjoint from every unmarked writer, so a foreground write at exactly the backfill hybrid time cannot collide with a marked version; kBackfillWriteIdIndexMax preserves the reserved kMaxWriteId sentinel. Guards, all failing cleanly before WAL or storage effects: - WriteQuery::MaybeMarkRaftIndexWriteIdBatch marks batches whose ops carry UNIQUE_INDEX_BACKFILL_SKIP_ALL, and fails closed if a SKIP_ALL op is not an eligible batch (non-transactional, all-backfill PGSQL_INSERT, fixed hybrid time) or the deferred-verification capability AutoFlag is off -- writing SKIP_ALL unmarked would reintroduce the silent overwrite. - ValidateRaftIndexWriteIdBatch rejects marked batches that are transactional, apply external transactions (those records bypass the write-ID override), or contain duplicate unversioned keys, before Raft submission. A duplicate surviving to apply means corrupt or incompatible replicated data; the batch writer fail-stops with Corruption rather than silently collapsing a candidate. - WriteOperation::ValidateLeaderOpId (new Operation hook invoked from AddedToLeader) rejects Raft indexes above kBackfillWriteIdIndexMax before WAL append; the index is rolled back and reused, never skipped or wrapped (relies on the AddedToLeader cleanup fix, #33380). The master never selects SKIP_ALL yet (#33484's hardcoded CHECK_ALL selector), so the marked path remains production-unreachable; tests drive it through the master-side TEST_ysql_index_backfill_unique_check_mode override, exercising the full job-mode plumbing. Test Plan: ./yb_build.sh release --cxx-test doc_hybrid_time-test --gtest_filter 'DocHybridTimeTest.WriteIdEncodedSizeGrowth' ./yb_build.sh release --cxx-test non_transactional_batch_writer-test --gtest_filter 'NonTransactionalBatchWriterTest.SeparateBatchesAtSameFixedHybridTime' ./yb_build.sh release --cxx-test non_transactional_batch_writer-test --gtest_filter 'NonTransactionalBatchWriterTest.SeparateBatchesAtSameFixedHybridTimeWithWriteIdOverride' ./yb_build.sh release --cxx-test non_transactional_batch_writer-test --gtest_filter 'NonTransactionalBatchWriterTest.MarkedAndUnmarkedWritesWithSameHybridTimeAndWriteIdCollide' ./yb_build.sh release --cxx-test non_transactional_batch_writer-test --gtest_filter 'NonTransactionalBatchWriterTest.FloorSeparatesMarkedAndUnmarkedWriteIdDomains' ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.FixedHybridTimeWritesUseRaftIndexWriteId' ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.FixedHybridTimeWriteIdRequiresUniqueUnversionedKeys' ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.FixedHybridTimeWriteIdOverflowRejectedBeforeRaftAppend' ./yb_build.sh release --cxx-test fixed_hybrid_time_write_id-itest --gtest_filter 'FixedHybridTimeWriteIdITest.LeaderChangePreservesDistinctWriteIds' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillSkipAllRaftOrdering.UniqueIndexBackfillAndForegroundCheck/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillRaftOrderingActivation.MarkerReachesRaftIndexValidation/0' Regressions: uniq-idx-1 mode tests, DuplicatesExistBeforeBackfill, 0a/0b tablet_peer tests, write-ID cap e2e -- all green. Assisted-By: devx/f746c8e8-73e1-4ceb-bc27-579022f173cf --- _automated · Claude Fable 5 (opencode)_
| Commit: | 2203b35 | |
|---|---|---|
| Author: | John Meehan | |
| Committer: | John Meehan | |
[#33444] DocDB: Add immutable per-job uniqueness-check mode for unique index backfill First slice of deferred uniqueness verification (#33444): a typed `UniqueIndexBackfillMode` (`CHECK_ALL` = existing behavior, `SKIP_ALL` = skip both duplicate-check reads) that the master selects exactly once per backfill job, persists in `BackfillJobPB`, and carries immutably through the whole backfill write path: BackfillJobPB -> BackfillIndexRequestPB -> PgsqlBackfillSpecPB -> BACKFILL INDEX statement -> pggate -> PgsqlWriteRequestPB -> DocDB The mode travels with the work rather than being read from runtime flags, so master failover, retries, and flag changes cannot reinterpret an active job: the `BackfillTable` constructor reuses the persisted value on resume, and tservers act only on the mode carried by each write request. Not production-selectable yet: the selector returns `CHECK_ALL` unconditionally (SKIP_ALL requires the marked-write ordering and verification machinery from later parts of #33444), and the new capability AutoFlag `ysql_enable_deferred_unique_index_verification` (kLocalPersisted) ships unpromoted. Missing or unknown mode values resolve to `CHECK_ALL` at every consumer, so older or unaware components keep the fully checked path. Foreground writes never carry the field; their uniqueness enforcement is untouched (the mode is consulted only in the is_backfill branch of PGSQL_INSERT). `TEST_ysql_index_backfill_unique_check_mode` (test-only, check_all / skip_all) overrides behavior at the DocDB check site and, on the master, overrides selection, so tests can exercise both the check-site consumption in isolation and the full production plumbing. Test Plan: ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillTest.UniqueCheckModeTserverOverrideSkipsChecks/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillTest.UniqueCheckModePlumbedFromMasterToWritePath/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillMultiMaster.UniqueCheckModePersistedAcrossMasterFailover/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillBlockDoBackfill.DuplicatesExistBeforeBackfill/0' Assisted-By: devx/f746c8e8-73e1-4ceb-bc27-579022f173cf Assisted-By: devx/08801cf6-2854-4c6f-a378-1d98d239b22d --- _automated · Claude Fable 5 (opencode)_
| Commit: | 0e29aed | |
|---|---|---|
| Author: | Ella Baron | |
| Committer: | Ella Baron | |
dist-trace: Address review — drop AshMetadataPB mention, rename parent to traceparent
| Commit: | 4682144 | |
|---|---|---|
| Author: | Ella Baron | |
| Committer: | Ella Baron | |
dist-trace: propagate trace context across the RPC boundary Carry a distributed trace from an RPC caller to its callee: define the wire type, write it into the request header, decode it on the far side, and open a server span parented under the caller's client span. All of it lands together because none of the pieces is observable on its own -- a wire field nobody reads, or a reader with nothing on the wire, cannot be tested. Wire type (rpc_header.proto): - TraceContextPB (W3C trace-context) and RequestHeader.trace_context. This is the type shared by the RPC-header and shared-memory transports; the latter starts using it in a later commit. Writer (outbound_call.cc): - SetRequestParam serializes the active client span's SpanContext into the trace_context submessage while it is sizing and writing the header, so the header is still written in a single pass. ~30 bytes on the wire, and only when a trace is active; an older peer skips the unknown field. Reader (serialization.{cc,h}): - ParseTraceContext(Slice) decodes the RequestHeader.trace_context wire field; ToSpanContext(TraceContextPB) is the shared-memory equivalent; both go through BuildSpanContext and fail on a zero/invalid context. - ParseHeader captures RequestHeader.trace_context into ParsedRequestHeader. Server / inbound span (yb_rpc.{cc,h}): - YBInboundCall::ParseFrom parses the header trace_context into parent_span_context_ (best-effort: a bad context is logged, never fails the RPC). - CreateServerSpan starts a StartServerSpanWithScope child of that parent (remote) or of an explicitly passed context (local); RespondSuccess/Failure set status and end it; DropServerSpanScope releases the scope on the handler thread while the span ends later. Client local-call handoff (local_call.{cc,h}, rpc_context.cc): - LocalOutboundCall tags rpc.local_call and drops its client-span scope in the ctor (a local call never hops threads to drop it later). - otel_span_context() exposes the outbound span's context so the local inbound span can parent under it (no wire header exists locally). - RpcContext creates the server span: from the wire header for remote calls, from the outbound call's context for local calls. Process init (tserver/db_server_base.cc): - DbServerBase is the shared base of Master and TabletServer, so DbServerBase:: Init is the single place that initializes the process-wide tracer (service name = server name, node id = permanent uuid), with Shutdown tearing it down. Gated by IsDistTraceEnabled, so it is a no-op when tracing is off. Without this the server spans above are never exported, which is why the 10 lines ship here rather than as a commit of their own. Test: TestRpcSpanReachesTabletServer runs a SELECT under a known traceparent and asserts that the TabletServer's inbound "rpc yb.tserver.PgClientService.Perform" span lands in that trace as a child of the ysql backend's outbound span, with the expected rpc.service/rpc.method attributes. The new WaitForRemoteChildSpan collector helper does the caller/callee pairing check.
| Commit: | f06d4ec | |
|---|---|---|
| Author: | John Meehan | |
| Committer: | John Meehan | |
[#33444] DocDB: Fence demotion of the deferred-verification capability flag The one unguarded downgrade path for deferred unique-index verification: binary downgrade to a pre-feature binary is already hard-fenced by the AutoFlags config validation while ysql_enable_deferred_unique_index_verification stays promoted, and RollbackAutoFlags refuses kLocalPersisted flags outright -- but DemoteSingleAutoFlag validates nothing beyond flag existence. Demoting the capability flag clears the startup fence while SKIP_ALL jobs or unflushed marked WAL entries exist; an old binary then bootstraps and replays marked writes with positional write IDs -- byte-divergent replicas. DemoteSingleAutoFlag now refuses to demote this flag (the demote-side first, mirroring the promote-side major-upgrade veto) while either holds: - Any table has an active SKIP_ALL backfill job (sys-catalog scan): its chunks may still be appending marked writes. - Any registered tserver reports marked-write state via the new CheckIndexBackfillDowngradeSafety admin RPC: a tablet with an active ordering generation, or a released one whose marked WAL entries are not yet below the regular-DB flushed frontier. Sequential synchronous fan-out -- demotion is a rare admin operation. An unreachable tserver FAILS the demotion (it may hold config-member replicas with unflushed marked WAL; wait for it or decommission it), with one exception: a node dead past follower_unavailable_considered_failed_sec is skipped -- its replicas have been evicted and re-replicated, and a returning node's stale local replay is confined to data that is tombstoned before serving. The probe itself skips tombstoned/deleted peers (WAL removed) and completed split parents (flushed before the split checkpoint), so husks on the master's cleanup schedule cannot block a drain. The per-tablet condition is made provable and convergent by two release changes: - Release persists released_marked_write_watermark in the generation superblock record: the highest marked-write Raft index the tablet has applied (tracked in memory, monotonic, re-seeded by bootstrap replay, which re-applies exactly the marked writes still above the flushed frontier). Anchoring to the last marked WRITE rather than the release op matters: metadata ops never advance the RocksDB frontier, so a release-op watermark would never converge on an idle tablet. - Release schedules an async regular-DB flush (FlushReason::kIndexBackfillGenerationRelease), pushing the marked writes below the frontier within seconds instead of waiting for an organic flush. The released record now survives in the superblock (previously only active generations were persisted) until the next activation. The tablet validator's straggler heal records the watermark too, deferring its heal to a later cycle if the tablet is not running (the watermark lives on the Tablet). The error text carries the drain procedure: finish or abort the builds, let the release flushes converge, retry. The veto is deliberately not bypassable: every blocking condition drains on its own, and demoting past it is precisely the divergence hazard. Test Plan: ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillVerifier.DowngradeFenceBlocksWhileGenerationsHeld/0' (release blocked by the fixture: demotion refused, message names the active ordering generation) ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillSkipAllRaftOrdering.DowngradeFenceBlocksWhileJobActive/0' (job held at DoBackfill: demotion refused before any tserver probe, message names the SKIP_ALL job's table) ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillVerifierReleased.DowngradeFenceClearsAfterRelease/0' (after the funnel releases and the release flush converges, demotion goes through -- this test caught the release-op-watermark convergence bug during development) ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillVerifierReleasedNoFlush.DowngradeFenceBlocksWhileMarkedWalUnflushed/0' (deterministic pin of the unflushed arm: release flush disabled via test flag, fence refuses with the not-yet-flushed message) ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.IndexBackfillOrderingGenerationPersistsAcrossReload' (extended: the released watermark survives reload; a never-released clear still persists as field absence) Regressions: full tablet_peer-test (27), full fixed_hybrid_time_write_id-itest (7), funnel release e2e (OrderingGenerationActivatedAndReleased), VerifyAfterReleaseFailsCleanly, RetentionPinMetricsClearOnRelease -- all green. Assisted-By: devx/08801cf6-2854-4c6f-a378-1d98d239b22d --- _automated · Claude Fable 5 (opencode)_
| Commit: | ed37776 | |
|---|---|---|
| Author: | John Meehan | |
| Committer: | John Meehan | |
[#33444] DocDB: Add a keyed, per-job violation fingerprint to unique-index verification Verification failure reports have been value-free by design: encoding classes and counts, never the duplicate key. That satisfies the privacy contract but leaves support blind to correlation -- "is this the same duplicate as the last retry?" This adds the design document's keyed, scope-limited diagnostic fingerprint: a non-reversible token identifying the violating index key, safe for low-cardinality values because confirming a candidate requires the key, and the key's scope limits how long confirmation is possible at all. Mechanics: - Per-job secret (32 bytes, OpenSSL RAND_bytes) generated in the BackfillTable constructor for SKIP_ALL jobs (other modes never verify and carry no key), persisted in BackfillJobPB (field 9, latched-per-job like the mode and the gate), and deleted with the job: fingerprints correlate across the job's retries and master failover, and are unconfirmable once the job is gone. - The verifier exposes the first violating group's raw encoded DocKey in-process only (excluded from ToString; never serialized). The tserver handler renders the token -- "<jobtag8hex>-<digest16hex>", both truncated HMAC-SHA256 under the job key (yb/util/keyed_fingerprint.{h,cc}) -- into a structured response field and a "fingerprint=" suffix on the still-value-free reason. The job tag is derived from the key alone, so tokens from different jobs are visibly incomparable. - The reason rides the existing paths into the persisted verification state, the master outcome log, and (in gating mode) backfill_error_message -> the CREATE INDEX error. The structured token is also persisted (UniqueIndexVerificationStatePB field 5) for support tooling. - The key never leaves the master/tserver trust domain: GetBackfillJobs scrubs it from RPC responses (pinned by test), both request-log sites (master send VLOG, tserver receive DVLOG) share RedactedDebugString -- which also elides start_key, a raw index DocKey -- and the one log that printed a whole BackfillJobPB clears the key first. Test Plan: ./yb_build.sh release --cxx-test keyed_fingerprint-test --gtest_filter 'KeyedFingerprintTest.TokenFormatAndDeterminism' ./yb_build.sh release --cxx-test keyed_fingerprint-test --gtest_filter 'KeyedFingerprintTest.TagIdentifiesKeyScope' (token format, determinism under one key, shared scope tag per key, cross-key incomparability) ./yb_build.sh release --cxx-test unique_index_verifier-test --gtest_filter 'UniqueIndexVerifierTest.ViolationExposesGroupPrefixInMemoryOnly' (the privacy pin: the raw group prefix is exposed in memory for the caller and excluded from ToString) ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillFailClosedVerification.ViolationFailsCreateIndex/0' (extended: the CREATE INDEX error now carries "fingerprint=<token>"; verified rendering end-to-end: fingerprint=2d929cea-dbf259e3ff0d7297) Full unique_index_verifier-test 20/20. Regressions: shadow violation (waiter matches the enriched line), shadow clean multi-tablet, failover resume -- all green. Assisted-By: devx/08801cf6-2854-4c6f-a378-1d98d239b22d --- _automated · Claude Fable 5 (opencode)_
| Commit: | d4ee3dd | |
|---|---|---|
| Author: | John Meehan | |
| Committer: | John Meehan | |
[#33444] DocDB: Add verification scan accounting and master outcome counters Verification/job observability for deferred unique-index verification, plus the accounting fix deferred from the verifier-core review: versions_scanned was inconsistent on the bounded-memory fallback path (the forward pass stopped counting at buffer overflow, then the reverse walk re-counted its re-visits -- over- and under-counting the same group). The forward pass now owns the accounting -- every physical version in a scanned group is counted exactly once regardless of replay path -- and the new fallback_groups counter records the reverse walk's extra work instead (a fallback group is read roughly twice). Plumbing and aggregation: - VerifyUniqueIndexTabletResponsePB gains fallback_groups (field 8); the handler fills it from the verifier result. - The shadow coordinator aggregates per-index scan totals from every response and logs them with the recorded outcome: "... VERIFY_CLEAN [observational] (dockey_groups=300, versions=300, fallback_groups=0)". - Five cluster-level master counters (backfill_aborted pattern): unique_index_verification_outcome_{clean,violation,inconclusive} (one per recorded per-index outcome) and unique_index_verification_{versions_scanned,fallback_groups} (aggregated from responses; late responses dropped by the single-winner guard still count -- the scan happened). A persistently high fallback_groups rate is the signal to raise unique_index_verify_max_buffered_versions_per_group; the counter description says so. Test Plan: ./yb_build.sh release --cxx-test unique_index_verifier-test --gtest_filter 'UniqueIndexVerifierTest.FallbackAccountingMatchesBuffered' (the accounting pin: identical data verified through the buffered and fallback paths reports identical versions_scanned and dockey_groups_scanned; the fallback run reports fallback_groups = 1) ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillShadowVerification.CleanOutcomeRecordedMultiTablet/0' (extended: asserts the master outcome counter == 1 -- exact on a fresh cluster, catches double-counting -- and versions_scanned > 0 via the cluster metric entity) Full unique_index_verifier-test 19/19. Regressions: ViolationRecordedButDoesNotBlockPublication (waiter matches the enriched log line), ViolationFailsCreateIndex (gating path through Record), failover resume, paginated clean -- all green. Assisted-By: devx/08801cf6-2854-4c6f-a378-1d98d239b22d --- _automated · Claude Fable 5 (opencode)_
| Commit: | 4d6cbae | |
|---|---|---|
| Author: | Hari Krishna Sunder | |
| Committer: | GitHub | |
[BACKPORT 2026.1][#33175] DocDB: Fence writes past ignore_after_hybrid_time (#33507) (#33645) ## Summary A lease lets one component be the sole writer to some data for a bounded window without coordinating on every write: while it holds the lease it writes freely, and once the lease expires another holder may take over. That is only safe if the outgoing holder's writes cannot take effect after its window ends -- otherwise a delayed write from a former owner lands on top of the new owner's data. There is currently no way for a client to express that bound. An RPC deadline bounds how long the caller waits, not whether the write eventually applies: nothing cancels an operation once it has been handed to Raft, so a write the caller gave up on can still commit afterwards. A lease holder cannot tell a write that failed from one that will land later. The primitive exists for object locks (`AcquireObjectLockRequestPB`) but not for writes. This adds a per-write fence. The caller states the hybrid time after which the write must not take effect, and the tablet leader checks it when admitting the write, rejecting it if that moment has already passed. The check happens before the operation can have any effect, so a rejected write definitively did not apply -- the property a timeout cannot give. It is reported as `Expired` carrying `TabletServerErrorPB::WRITE_FENCE_EXPIRED`, distinguishable from the failures that may still have landed; the thin client surfaces it as `YBTHIN_FENCED`. The fence rides on `PgsqlWriteRequestPB.ignore_after_hybrid_time` per op, and is lifted onto `WritePB` as the earliest fence over the batch, that being the only message the leader's admission path sees. 0 means no fence. Two limits, both noted on the request field: each tablet leader judges the fence independently, so a batch spanning tablets can be partly applied if it straddles the boundary; and being a hybrid time comparison, the fence is only as tight as `max_clock_skew_usec`. ## Upgrade/Rollback safety The feature can only be used after the upgrade is completed. ## Merge conflicts Both conflicts were context-only -- the neighbouring field that master's hunk carried as context does not exist on 2026.1. No content differs from the original commit; the added fields keep the same tag numbers as master, deliberately leaving a gap rather than renumbering, so a tag means the same thing on both branches. - `src/yb/tablet/operations.proto:113` — `WritePB` ends at 22 here; master has `xcluster_target_applied = 23`. `ignore_after_hybrid_time` stays at **24**, 23 left unused. - `src/yb/tserver/tserver_types.proto:135` — `TabletServerErrorPB.Code` ends at 32 here; master has `READ_TIME_NOT_REACHED = 33`. `WRITE_FENCE_EXPIRED` stays at **34**, 33 left unused. `src/yb/common/pgsql_protocol.proto` applied cleanly: 2026.1 already carries `read_at_in_txn_limit = 31`, so the fence is at 32 as on master. Original commit: 26d436e257aea2d2142333484c643aeee751e367 / #33507 ## Test plan `yb_thin_client-itest` -- full suite on 2026.1, 11/11. `PgThinClientTest.WriteFencedByIgnoreAfterHybridTime` asserts on table contents as well as status: a past fence is rejected and leaves no row, a far-future fence applies, no fence behaves as before, and every peer reports zero running retryable requests afterwards. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
| Commit: | 6a00f9a | |
|---|---|---|
| Author: | John Meehan | |
| Committer: | John Meehan | |
[#33444] DocDB: Gate unique-index publication on the deferred verification outcome The last functional piece of deferred uniqueness verification: behind ysql_index_backfill_fail_closed_verification (default false), the verification phase's outcome decides publication instead of merely being recorded. The publication order becomes the design document's: backfill completes -> stay at DO_BACKFILL -> verify all index tablets -> persist the outcome -> RWD -> publish birth_time -> release generation + retention -> indisvalid In gating mode, backfill-success marking for unique indexes is deferred from chunk completion to verification completion: a clean index is then marked successful (including the xCluster backfill-completed notification, which must not fire for an index that may still fail) and proceeds through the terminal funnel to READ_WRITE_AND_DELETE and birth_time publication. A VIOLATION or unresolvable INCONCLUSIVE outcome fails the index through the existing backfill failure path -- backfill_error_message carries a value-free reason, CREATE INDEX returns it, the index is never READ_WRITE_AND_DELETE and never indisvalid, and the existing invalid-index cleanup applies. Coordinator failures fail closed too: whatever is still IN_PROGRESS when the phase fails is exactly the unverified set (the in-flight index, indexes the phase never reached, and indexes never selected because resolution itself failed) and is failed rather than published. Non-unique indexes in the job have nothing to verify and are marked successful at chunk completion as before. The gating decision is latched per job and persisted in BackfillJobPB, mirroring the uniqueness-check mode: failover, retries, and runtime flag flips cannot reinterpret an active job. Reading the flag live at each decision point would misbehave in both directions -- gating->observational after the deferred marking leaves the index IN_PROGRESS at the funnel; observational->gating mid-phase would fail an index whose success (and xCluster notification) already fired. The funnel logs DFATAL if it ever sees an IN_PROGRESS index (only reachable if terminal-state marking itself kept failing); the pre-existing mapping publishes such an index as failed. The gate implies the verification phase runs even when ysql_index_backfill_shadow_verification is off; with both flags off, nothing changes. Observational mode is untouched -- outcome logs now carry a [gating] / [observational] tag. The mode selector remains hardcoded to CHECK_ALL, so the gate is production-unreachable until activation; tests drive it through the master-side TEST override. Test Plan: ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillFailClosedVerification.CleanBuildPublishes/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillFailClosedVerification.ViolationFailsCreateIndex/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillFailClosedInconclusive.InconclusiveFailsCreateIndex/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillFailClosedPhaseFailure.CoordinatorFailureFailsCreateIndex/0' The violation test is the feature's end state: a real pre-existing duplicate survives the SKIP_ALL build and fails CREATE INDEX with a verification error; pg_index shows indisvalid = false, the base table stays usable, and the invalid index is droppable. The inconclusive test pins fail-closed = "not proven clean" via a forced-INCONCLUSIVE tserver response. The clean test pins publication: indisvalid = true and a consistent index. The phase-failure test pins the coordinator-failure path: an injected resolution failure fails CREATE INDEX with the unverified-index message on clean data. Regressions: shadow suite (violation still publishes in observational mode; multi-tablet clean; failover resume, which now also exercises the persisted gating latch on resume), skip_all funnel e2e, Drop -- all green. Assisted-By: devx/08801cf6-2854-4c6f-a378-1d98d239b22d --- _automated · Claude Fable 5 (opencode)_
| Commit: | d905f72 | |
|---|---|---|
| Author: | John Meehan | |
| Committer: | John Meehan | |
[#33444] DocDB: Add the shadow verification coordinator with durable job state The first production caller of the verification RPC: after a SKIP_ALL unique-index backfill completes -- while the ordering generations are still active, before the terminal funnel releases them -- the master runs the deferred uniqueness verification scan observationally. The outcome is persisted and logged but never gates index publication; the fail-closed verify-before-publication gate is a separate, later capability. Disabled by default (ysql_index_backfill_shadow_verification). Durable state (BackfillJobPB.unique_index_verification, keyed by unique-index table id): verification state, verify_upper_ht, tablets confirmed clean, and a value-free reason. verify_upper_ht is chosen ONCE from cluster hybrid time (never wall clock) and persisted before any scan -- retries and failover resume verify the same window. Clean tablets are durable progress at tablet granularity; finer intra-tablet resume keys are deliberately not persisted, since re-scanning one tablet after failover is idempotent. All persistence goes through the indexed table's write lock and a sys-catalog upsert under the job's leader epoch. Coordinator mechanics (VerifyUniqueIndexForTablet, the GetSafeTimeForTablet task shape): - One index at a time; bounded per-tablet concurrency (index_backfill_shadow_verification_max_concurrent_tablets). - Pagination: index_backfill_verify_dockey_groups_per_rpc caps each RPC's group budget; a tablet resumes from the returned resume key without touching the join counts. - Short-circuit under a single-winner protocol: the terminal check and the launch-next / index-done decision are one critical section, so a CLEAN callback racing a short-circuiting VIOLATION can neither overwrite the recorded outcome nor double-advance the phase; the slow clean-tablet persist happens after the join decision. Late responses cannot change a recorded outcome. The join guards on job failure, not done() -- the whole phase runs in the job's kSuccess state. - Fail-open on coordinator errors: the phase is on the CREATE INDEX critical path, so every coordinator failure (task launch, persistence, pagination resume) degrades to VERIFY_INCONCLUSIVE and continues into publication -- no failure may strand the job short of the terminal funnel. The tserver returns typed error codes for deterministic failures (generation mismatch, non-unique-index tablet, cutoff violations) so the coordinator does not burn its retry budget on requests that cannot succeed. - An index with no tablets records VERIFY_INCONCLUSIVE: an empty tablet set is not evidence of uniqueness. - Master failover mid-phase resumes with the persisted window. This required fixing a resume gap the failover test exposed, one that upstream also has in the (previously tiny) window between success-marking and the funnel: heartbeat-driven resume only fires on schema-version mismatches, which no longer occur once the chunks are done. Tables with a persisted backfill job are now queued for resume at sys-catalog load; the resume path drives the job directly when the table is no longer ALTERING; and a resumed job whose indexes are all SUCCESS skips straight to the completion phase instead of relaunching empty backfill chunks. - No generation_base_op_index in the requests, on purpose: the base is per-tablet (the activation op's own Raft index) and a re-activated generation's higher base is still the generation this job's marked writes live under; active + index-table match is the semantic guard. Test Plan: ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillShadowVerification.CleanOutcomeRecordedMultiTablet/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillShadowVerification.ViolationRecordedButDoesNotBlockPublication/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillShadowVerificationPaginated.CleanAcrossManyRpcs/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillVerifier.GenerationMismatchFailsCleanly/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillShadowVerificationFailover.ResumesWithPersistedWindowAcrossFailover/0' The violation test is the shadow-semantics pin: a real duplicate is recorded (NOT CLEAN: VERIFY_VIOLATION in the master log) while CREATE INDEX still succeeds and the index serves reads. The multi-tablet test exercises the fan-out/join; the paginated fixture forces one DocKey group per RPC through the resume path. GenerationMismatchFailsCleanly covers the 4b review's L1 (table-id and base mismatches fail without scanning). The failover test holds the verification RPCs open with a retryable rejection, steps the master leader down mid-phase, and asserts the new leader resumes with the persisted window before completing the build. Regressions: shadow-flag-off skip_all e2e, CHECK_ALL mode test, 4b verifier e2e -- all green. Assisted-By: devx/08801cf6-2854-4c6f-a378-1d98d239b22d --- _automated · Claude Fable 5 (opencode)_
| Commit: | 68db79a | |
|---|---|---|
| Author: | John Meehan | |
| Committer: | John Meehan | |
[#33444] DocDB: Expose the unique-index verification scan as a tserver admin RPC Wires the verifier core to the world: VerifyUniqueIndexTablet on the tablet server admin service, read-only and paginated, with the preconditions the scan's correctness rests on enforced at the entry point. No production caller yet -- the verification coordinator drives it in a later part; end-to-end tests drive it directly. The RPC (leader-only, BackfillIndex-style pagination): - Discharges the applied-intents precondition in two steps: waits for the tablet's safe time to pass verify_upper_ht (leader lease held; no future write can land at or below the bound), then TransactionParticipant::ResolveIntents(verify_upper_ht) -- safe time alone does not imply applied, since transactional apply is asynchronous, and a committed-but-unapplied foreground write is invisible to a regular-DB-only scan (the record missed could be exactly the foreground half of a duplicate). - Validates the active ordering generation against the request's expectation (index table, and base_op_index when set -- the coordinator always sets it; omitting it relaxes to "any active generation for this table" for test tooling). A released or replaced generation fails the RPC without scanning: the window no longer describes the tablet's marked writes. - Refuses a window that begins below the *applied* history cutoff, read from the regular DB's flushed frontier -- the cutoff past compactions actually used. Deliberately not the retention policy's directive: GetRetentionDirective mutates committed cutoff state (unacceptable on a read path) and only bounds future compactions, which the retention-hold part of this feature prevents from advancing anyway. Until that hold lands the pre-scan check is advisory, so the applied cutoff is re-checked after the scan: a compaction advancing it mid-scan fails the call rather than returning a result computed over possibly-incomplete history. Deterministic failures (generation mismatch, non-unique-index tablet, invalid or inverted window) return typed error codes, so callers do not burn a retry budget on requests that cannot succeed. The scan holds the RocksDB shutdown guard for its duration (the not-blocking variant: tablet shutdown aborts the scan, never the reverse), and an unset backfill_read_ht is rejected at the boundary -- it decodes to the invalid sentinel, which compares greater than any real hybrid time and would make the history-cutoff checks pass vacuously. Tablet::VerifyUniqueIndex resolves the scan options from the tablet: the primary table's ybidxbasectid column, the tablet metadata as the schema-packing provider, the regular DB within the tablet's key bounds, and the new runtime flag unique_index_verify_max_buffered_versions_per_group. TEST_block_index_backfill_ordering_generation_release (master) holds the terminal funnel's release open so tests can verify an index built by the real backfill path while its generation is still active. The e2e suite closes the encoding-drift loop the verifier-core review required: every physical shape asserted by the unit tests is re-asserted against tablets written by the production SKIP_ALL backfill and live pg DML. Test Plan: ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillVerifier.CleanBuildWithForegroundDml/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillVerifier.PkUpdateAcrossBackfillBoundaryIsClean/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillVerifier.DetectsPreexistingDuplicates/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillVerifierPackedUpdate.PkUpdateAcrossBackfillBoundaryIsClean/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillVerifierReleased.VerifyAfterReleaseFailsCleanly/0' The PK-update regression runs under both physical shapes (standalone ybidxbasectid column with default flags; packed V2 + kIsUpdateFlag with pack_full_row_update + mark) and must be Clean in both. Regressions: unique_index_verifier-test (18), OrderingGenerationActivatedAndReleased (release still fires with the TEST flag off), OrderingGenerationChangeMetadataOp tablet_peer test -- all green. Assisted-By: devx/08801cf6-2854-4c6f-a378-1d98d239b22d --- _automated · Claude Fable 5 (opencode)_
| Commit: | a0fa3ec | |
|---|---|---|
| Author: | John Meehan | |
| Committer: | John Meehan | |
[#33444] DocDB: Orchestrate the index-backfill ordering-generation lifecycle from the master The persisted ordering generation (#33580) had no production lifecycle: nothing activated it before marked backfill writes, nothing released it, and the master-side split suppression protecting a backfill is in-memory only (lost on failover). This part wires the full lifecycle and flips marked-write gating fail-closed. Activation rides ChangeMetadataOperation (a new `ChangeMetadataRequestPB.index_backfill_ordering_generation` variant, applied via `Tablet::UpdateIndexBackfillOrderingGeneration` exactly like `mark_backfill_done`): Raft-replicated, follower-applied, WAL-replayed, and `base_op_index` is the activation operation's own Raft index, so every replica and every replay derives the same base with no master-side bookkeeping. Re-activation (failover resume) is idempotent -- the base only moves up. Master flow, SKIP_ALL jobs on PGSQL indexed tables only (the mode rides YSQL chunk requests; YCQL backfill writes are never marked, so generations would only fence their splits without protecting anything): - Before the first chunk (`DoBackfill`), the job disables and drains index-table splitting, then fans out one waited `UpdateOrderingGenerationForTablet` task per index tablet (the `GetSafeTimeForTablet` join pattern). All acks -> chunks launch; any failure -> the job aborts before a single marked write exists, and CREATE INDEX fails cleanly. The drain closes the activation/split TOCTOU: consensus already rejects CHANGE_METADATA_OP while a split is pending, the generation fence rejects splits after activation applies, and a split appended in the append-to-apply window can at worst fail the activation task -- never corrupt a scan set that marked writes exist in. - Release is fire-and-forget from the terminal funnel (`UpdateIndexPermissionsForIndexes`), which every success/failure/abort path traverses. The tserver metadata validator converges stragglers, mirroring the retain_delete_markers machinery: tablets with an active generation join its existing GetBackfillStatus poll, and the generation is released locally (no Raft) once the master reports a terminal state. `IndexStatusPB` gains `BACKFILL_FAILED` (removal-path index permissions map to it) so a failed SKIP_ALL job whose funnel release was lost cannot leave an orphaned invalid index permanently split-fenced and retention-pinned; the retain_delete_markers heal itself stays success-only. - The tablet-split manager refuses to split an index whose indexed table has a durable SKIP_ALL backfill job. The manager's table validation takes no catalog locks itself -- the indexed table is a caller-resolved parameter, because the manual-split path enters through ValidateSplitCandidateUnlocked with the catalog mutex already held (a recursive shared acquisition is fatal under lock_debug) (`SysTablesEntryPB.backfill_jobs`, survives failover, cleared in the funnel), with an INFO-level skip reason -- today's two in-memory suppressions cover only the indexed table and evaporate on failover. The tablet-side generation fence stays as the fail-closed layer. Gating flip: `WriteOperation::ValidateLeaderOpId` now rejects marked writes with *no* active generation (previously deferred) -- a marked write outside a generation would store versions nothing tracks or releases (e.g. a stale chunk retry after the job's terminal state). Tests that drive marked writes directly now activate a generation first. `write_id_floor_version` numbering lands as `kIndexBackfillWriteIdFloorVersion = 1`. The mode selector remains hardcoded to CHECK_ALL (#33484), so this whole flow stays production-unreachable; tests drive it through the master-side TEST override. Test Plan: ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.OrderingGenerationChangeMetadataOpSetsBaseFromOwnRaftIndex' ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.FixedHybridTimeWriteRejectedAtOrBelowGenerationBase' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillSkipAllRaftOrdering.OrderingGenerationActivatedAndReleased/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillSkipAllActivationFailure.ActivationFailureFailsCreateIndexCleanly/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillSkipAllBlocked.SplitFencedDuringBackfill/0' Retrofitted (marked writes now require an active generation): ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.FixedHybridTimeWrite*' ./yb_build.sh release --cxx-test fixed_hybrid_time_write_id-itest Regressions: skip_all e2e + marker canary through the real activation flow, uniq-idx-1 mode tests, DuplicatesExistBeforeBackfill, 3b-i generation tests, 0a/0b abort tests, and drop-path tests for the permissions mapping change (PgIndexBackfillTest.Drop, PgIndexBackfillFastClientTimeout.DropWhileBackfilling) -- all green. Assisted-By: devx/08801cf6-2854-4c6f-a378-1d98d239b22d --- _automated · Claude Fable 5 (opencode)_
| Commit: | cc9ba40 | |
|---|---|---|
| Author: | John Meehan | |
| Committer: | John Meehan | |
[#33444] DocDB: Add persisted index-backfill ordering-generation record and tablet fences Deferred uniqueness verification scans the marked versions that unique-index backfill wrote at the fixed backfill hybrid time. That scan is only sound against a stable tablet set and a validated write-ID sequence, so each index tablet needs a durable record of the in-flight backfill's ordering state -- one that survives restart, remote bootstrap, and inheritance by split children. The record (`IndexBackfillOrderingGenerationPB`, superblock field 39, persistence following the `cdc_sdk_safe_time` pattern): `active`, the owning index `table_id`, `base_op_index`, `retention_barrier_ht`, and `write_id_floor_version` (numbering pinned in the proto: 1 = the initial floor, 0 = absent/unknown). The full combined shape is defined here; `retention_barrier_ht` is consumed by the verification read fence and identity/activation by the master orchestration, both in later parts. The durable form of an inactive generation is field absence: `LoadFromSuperBlock` clears the in-memory record when the field is absent (so a remote-bootstrap superblock replacement from a source without a generation cannot resurrect a released one), and the setter normalizes inactive input to the default -- the caller contract is "set an active generation or set {}", keeping in-memory and durable state identical. While a generation is active on a tablet: - Marked fixed-hybrid-time writes must carry Raft indexes strictly above `base_op_index` (`WriteOperation::ValidateLeaderOpId`). The base is the activation operation's own Raft index, so an index at or below it cannot belong to this generation -- it would indicate misrouting or replay of a foreign sequence. (Rejecting marked writes when *no* generation is active arrives with the master activation flow; until then marked writes remain production-unreachable.) - Tablet splitting is rejected before Raft append (`SplitOperation::ValidateLeaderOpId`) -- the fail-closed layer under the master-side split fence (next part). The master's split path retries and succeeds once the generation is released. - Tablet cloning is rejected the same way. Clone targets build fresh metadata (`RaftGroupMetadata::CreateNew`), not a superblock copy, so a clone would carry hard-linked snapshot data containing generation-scoped marked write IDs with no generation record describing them and no backfill job tracking their lifecycle. `TEST_bypass_index_backfill_ordering_generation_split_fence` forces the split fence open to demonstrate the property the fence is *not* needed for: split children inherit the parent's marked entries and continue drawing write IDs from their own Raft indexes, so per-key write-ID uniqueness holds even across a split (the fence exists for verification-scan stability and lifecycle tracking, not key correctness). Test Plan: ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.FixedHybridTimeWriteRejectedAtOrBelowGenerationBase' ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.IndexBackfillOrderingGenerationPersistsAcrossReload' ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.IndexBackfillOrderingGenerationClearedBySuperblockReplacement' ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.SplitRejectedWhileOrderingGenerationActive' ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.CloneRejectedWhileOrderingGenerationActive' ./yb_build.sh release --cxx-test fixed_hybrid_time_write_id-itest --gtest_filter 'FixedHybridTimeWriteIdITest.MasterSplitRefusedWhileOrderingGenerationActive' ./yb_build.sh release --cxx-test fixed_hybrid_time_write_id-itest --gtest_filter 'FixedHybridTimeWriteIdITest.OrderingGenerationSurvivesRemoteBootstrap' ./yb_build.sh release --cxx-test fixed_hybrid_time_write_id-itest --gtest_filter 'FixedHybridTimeWriteIdITest.SplitWithFenceBypassedPreservesPerKeyWriteIdUniqueness' Regressions: 3a fixed-hybrid-time tablet_peer tests and leader-change itest, 0a/0b abort tests, PgIndexBackfillSkipAllRaftOrdering e2e -- all green. Assisted-By: devx/08801cf6-2854-4c6f-a378-1d98d239b22d --- _automated · Claude Fable 5 (opencode)_
| Commit: | b953483 | |
|---|---|---|
| Author: | John Meehan | |
| Committer: | John Meehan | |
[#33444] DocDB: Add marked fixed-hybrid-time write path with Raft-index write IDs SKIP_ALL unique-index backfill writes run no uniqueness checks, so distinct candidates sharing one unique-index key at the fixed backfill hybrid time must survive as distinct physical versions -- with the existing positional per-batch write IDs, both land at (key, backfill_ht, 0) and one silently overwrites the other, destroying the evidence deferred verification exists to inspect. Replicated write batches gain a self-describing marker (KeyValueWriteBatchPB.use_raft_index_for_write_id). For a marked batch, every replica derives the storage write ID from the Raft operation index as kBackfillWriteIdFloor | index -- one derivation site in Tablet::ApplyOperation covers leader apply, follower apply, and WAL-replay bootstrap, so replicas converge byte-identically and replay never consults runtime flags. The floor (an alias of kIntraTxnWriteIdLimit, #33499) keeps the marked domain disjoint from every unmarked writer, so a foreground write at exactly the backfill hybrid time cannot collide with a marked version; kBackfillWriteIdIndexMax preserves the reserved kMaxWriteId sentinel. Guards, all failing cleanly before WAL or storage effects: - WriteQuery::MaybeMarkRaftIndexWriteIdBatch marks batches whose ops carry UNIQUE_INDEX_BACKFILL_SKIP_ALL, and fails closed if a SKIP_ALL op is not an eligible batch (non-transactional, all-backfill PGSQL_INSERT, fixed hybrid time) or the deferred-verification capability AutoFlag is off -- writing SKIP_ALL unmarked would reintroduce the silent overwrite. - ValidateRaftIndexWriteIdBatch rejects marked batches that are transactional, apply external transactions (those records bypass the write-ID override), or contain duplicate unversioned keys, before Raft submission. A duplicate surviving to apply means corrupt or incompatible replicated data; the batch writer fail-stops with Corruption rather than silently collapsing a candidate. - WriteOperation::ValidateLeaderOpId (new Operation hook invoked from AddedToLeader) rejects Raft indexes above kBackfillWriteIdIndexMax before WAL append; the index is rolled back and reused, never skipped or wrapped (relies on the AddedToLeader cleanup fix, #33380). The master never selects SKIP_ALL yet (#33484's hardcoded CHECK_ALL selector), so the marked path remains production-unreachable; tests drive it through the master-side TEST_ysql_index_backfill_unique_check_mode override, exercising the full job-mode plumbing. Test Plan: ./yb_build.sh release --cxx-test doc_hybrid_time-test --gtest_filter 'DocHybridTimeTest.WriteIdEncodedSizeGrowth' ./yb_build.sh release --cxx-test non_transactional_batch_writer-test --gtest_filter 'NonTransactionalBatchWriterTest.SeparateBatchesAtSameFixedHybridTime' ./yb_build.sh release --cxx-test non_transactional_batch_writer-test --gtest_filter 'NonTransactionalBatchWriterTest.SeparateBatchesAtSameFixedHybridTimeWithWriteIdOverride' ./yb_build.sh release --cxx-test non_transactional_batch_writer-test --gtest_filter 'NonTransactionalBatchWriterTest.MarkedAndUnmarkedWritesWithSameHybridTimeAndWriteIdCollide' ./yb_build.sh release --cxx-test non_transactional_batch_writer-test --gtest_filter 'NonTransactionalBatchWriterTest.FloorSeparatesMarkedAndUnmarkedWriteIdDomains' ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.FixedHybridTimeWritesUseRaftIndexWriteId' ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.FixedHybridTimeWriteIdRequiresUniqueUnversionedKeys' ./yb_build.sh release --cxx-test tablet_peer-test --gtest_filter 'TabletPeerTest.FixedHybridTimeWriteIdOverflowRejectedBeforeRaftAppend' ./yb_build.sh release --cxx-test fixed_hybrid_time_write_id-itest --gtest_filter 'FixedHybridTimeWriteIdITest.LeaderChangePreservesDistinctWriteIds' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillSkipAllRaftOrdering.UniqueIndexBackfillAndForegroundCheck/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillRaftOrderingActivation.MarkerReachesRaftIndexValidation/0' Regressions: uniq-idx-1 mode tests, DuplicatesExistBeforeBackfill, 0a/0b tablet_peer tests, write-ID cap e2e -- all green. Assisted-By: devx/f746c8e8-73e1-4ceb-bc27-579022f173cf --- _automated · Claude Fable 5 (opencode)_
| Commit: | 26d436e | |
|---|---|---|
| Author: | Hari Krishna Sunder | |
| Committer: | GitHub | |
[#33175] DocDB: Fence writes past ignore_after_hybrid_time (#33507) ## Summary A lease lets one client be the sole writer to some data for a bounded window without coordinating on every write: while it holds the lease it writes freely, and once the lease expires another holder may take over. That is only safe if the outgoing holder's writes cannot take effect after its window ends -- otherwise a delayed write from a former owner lands on top of the new owner's data. This change adds a per-write fence. The caller states the hybrid time after which the write must not take effect, and the tablet leader checks it when admitting the write, rejecting it if that moment has already passed. The check happens before the operation can have any effect, so a rejected write definitively did not apply. It is reported as `Expired` carrying `TabletServerErrorPB::WRITE_FENCE_EXPIRED`; the thin client surfaces it as `YBTHIN_FENCED`. The fence rides on `PgsqlWriteRequestPB.ignore_after_hybrid_time` per op, and is lifted onto `WritePB` as the earliest fence over the batch, that being the only message the leader's admission path sees. 0 means no fence. Two limits, both noted on the request field: each tablet leader judges the fence independently, so a batch spanning tablets can be partly applied if it straddles the boundary; and being a hybrid time comparison, the fence is only as tight as `max_clock_skew_usec`. ## Upgrade/Rollback safety The feature can only be used after the upgrade is completed. ## Test Plan `yb_thin_client-itest` -- full suite, 11/11. `PgThinClientTest.WriteFencedByIgnoreAfterHybridTime` asserts on table contents as well as status: a past fence is rejected and leaves no row, a far-future fence applies, no fence behaves as before, and every peer reports zero running retryable requests afterwards. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
| Commit: | ac54e8e | |
|---|---|---|
| Author: | John Meehan | |
| Committer: | John Meehan | |
[#33444] DocDB: Add immutable per-job uniqueness-check mode for unique index backfill First slice of deferred uniqueness verification (#33444): a typed `UniqueIndexBackfillMode` (`CHECK_ALL` = existing behavior, `SKIP_ALL` = skip both duplicate-check reads) that the master selects exactly once per backfill job, persists in `BackfillJobPB`, and carries immutably through the whole backfill write path: BackfillJobPB -> BackfillIndexRequestPB -> PgsqlBackfillSpecPB -> BACKFILL INDEX statement -> pggate -> PgsqlWriteRequestPB -> DocDB The mode travels with the work rather than being read from runtime flags, so master failover, retries, and flag changes cannot reinterpret an active job: the `BackfillTable` constructor reuses the persisted value on resume, and tservers act only on the mode carried by each write request. Not production-selectable yet: the selector returns `CHECK_ALL` unconditionally (SKIP_ALL requires the marked-write ordering and verification machinery from later parts of #33444), and the new capability AutoFlag `ysql_enable_deferred_unique_index_verification` (kLocalPersisted) ships unpromoted. Missing or unknown mode values resolve to `CHECK_ALL` at every consumer, so older or unaware components keep the fully checked path. Foreground writes never carry the field; their uniqueness enforcement is untouched (the mode is consulted only in the is_backfill branch of PGSQL_INSERT). `TEST_ysql_index_backfill_unique_check_mode` (test-only, check_all / skip_all) overrides behavior at the DocDB check site and, on the master, overrides selection, so tests can exercise both the check-site consumption in isolation and the full production plumbing. Test Plan: ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillTest.UniqueCheckModeTserverOverrideSkipsChecks/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillTest.UniqueCheckModePlumbedFromMasterToWritePath/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillMultiMaster.UniqueCheckModePersistedAcrossMasterFailover/0' ./yb_build.sh release --cxx-test pg_index_backfill-test --gtest_filter 'PgIndexBackfillBlockDoBackfill.DuplicatesExistBeforeBackfill/0' Assisted-By: devx/f746c8e8-73e1-4ceb-bc27-579022f173cf Assisted-By: devx/08801cf6-2854-4c6f-a378-1d98d239b22d --- _automated · Claude Fable 5 (opencode)_