Proto commits in chroma-core/chroma

These commits are when the Protocol Buffers files have changed: (only the last 100 relevant commits are shown)

Commit:e9c6a7c
Author:Tanuj Nayak

[CLN](gc): Append and deduplicate candidates

Commit:2bca3e1
Author:Tanuj Nayak

[BUG](gc): Round robin candidate policies

The documentation is generated from this commit.

Commit:14214ad
Author:Tanuj Nayak

[BUG](gc): Reserve capacity for deleted collections

Commit:1c0dccf
Author:tanujnay112
Committer:GitHub

[ENH](sysdb): Add tenant-scoped bulk database lookup (#7818) Add `POST /api/v2/tenants/{tenant}/databases/batch/get` with an `ids` array of at most 1,000 UUIDs. It returns matching active databases using one Go SysDB query: `WHERE tenant_id = ? AND id IN (...) AND is_deleted = false`. Missing IDs are omitted; duplicate IDs produce one result. Empty input returns an empty array. The endpoint requires tenant-wide `list_databases` permission before reading storage. Both HTTP and gRPC reject oversized batches. MCMR is unsupported; its server only has the required unimplemented RPC stub. SQLite supports the same lookup for single-node use. Used by https://github.com/chroma-core/hosted-chroma/pull/8433. Deploy Go SysDB, then the data-plane frontend, before deploying that caller. Validation: PostgreSQL lookup test, gRPC batch-limit test, frontend HTTP test, and Rust compilation checks. Tests cover tenant isolation, soft deletion, duplicate/missing IDs, empty input, and oversized batches.

The documentation is generated from this commit.

Commit:dbf3ea5
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](sysdb): Add tenant-scoped bulk database lookup

The documentation is generated from this commit.

Commit:31024f9
Author:tanujnay112
Committer:GitHub

[ENH](sysdb): Count databases without listing (#7815) Database-creation quota checks currently fetch every database row for a tenant just to count them. Add an internal `CountDatabases` RPC and Rust client method that return a scalar instead. Go SysDB executes `SELECT COUNT(*) FROM databases WHERE tenant_id = ... AND is_deleted = false`. The frontend client calls only Go SysDB; MCMR counting is intentionally out of scope. The Rust SysDB server has only the unimplemented stub required by the shared generated service trait, with no Spanner or backend changes. SQLite and the in-memory test backend implement the shared client interface. No public HTTP route or schema migration is needed. The caller change is https://github.com/chroma-core/hosted-chroma/pull/8432. Merge this first and deploy the new RPC on Go SysDB before deploying that caller. No enumeration fallback is introduced. Validation: - Real PostgreSQL DAO test passes: empty tenant, active rows, soft deletion, tenant isolation, and count/list agreement. - Go coordinator and gRPC packages compile. - `cargo check -p chroma-sysdb -p rust-sysdb` passes. - Formatting and commit hooks pass. - Focused SQLite count test passes (empty tenant, populated tenant, isolation, deletion).

Commit:cb6f4a2
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](sysdb): Count databases without listing

Commit:d18f081
Author:tanujnay112
Committer:GitHub

[HOTFIX] applying PR #7748 to release/2026-09-04 (#7756) ## Description of changes Backports #7748 to `release/2026-09-04`. - Add an optional MDAC token bucket to WQS `GetWork`. - Charge one token for each work item returned while preserving FIFO, failure filtering, and rendezvous-shard assignment. - Define `limit` as the maximum distinct function IDs and add `max_items` as the independent total-record cap. - Add `excluded_fn_ids` to `GetWorkRequest`; fn-consumers populate it from active functions and WQS skips those rows before charging tokens. - Return relative gRPC pushback on empty 429 responses and an absolute UTC `retry_at_unix_ms` deadline on partial successful responses. - Reject invalid limiter configuration during startup and bound `max_concurrent_workers` so exclusion requests stay within the configured gRPC message limit. - Emit `work_queue_get_work_rate_limited_count` and structured throttle logs. The source change was cherry-picked onto the release branch with conflict resolution. ## Test plan The original change was validated in #7748 with formatting, focused worker and MDAC tests, and clippy. This backport is additionally covered by this PR's CI. ## Migration plan The protobuf additions are backward-compatible. The limiter is optional and remains disabled when configuration is absent. Removing the config block disables it. The limiter is process-local, which is global with the current single WQS replica. Additional WQS replicas would each receive their own allowance. ## Observability plan Monitor `work_queue_get_work_rate_limited_count`, `returned_items`, `excluded_fn_count`, and the structured `rate_limited` field. Watch for shard starvation or a growing queue during rollout. ## Documentation changes No user-facing API changes. --------- Co-authored-by: Robert Escriva <robert@trychroma.com>

Commit:dec0673
Author:tanujnay112
Committer:Tanuj Nayak

Cherry-pick with conflicts: ec22d992d774d4e9c16743920c4883021cdd0d47

Commit:ec22d99
Author:tanujnay112
Committer:GitHub

[ENH](wqs): Pace GetWork results (#7748) ## Description of changes - Add an optional MDAC token bucket to WQS `GetWork`. - Charge one token for each work item returned while preserving existing FIFO, failure filtering, and rendezvous-shard assignment. - Define `limit` as the maximum distinct function IDs and add `max_items` as the independent total-record cap. - Add `excluded_fn_ids` to `GetWorkRequest`; fn-consumers populate it from active functions and WQS skips those rows before charging tokens. - Return relative gRPC pushback on empty 429 responses and an absolute UTC `retry_at_unix_ms` deadline on partial successful responses, and log delayed fn-consumer polls. - Reject invalid limiter configuration cleanly during startup. - Validate `max_concurrent_workers` at 55,188 so worst-case exclusion IDs use at most half of WQS's explicit 4 MiB gRPC decode limit. - Emit `work_queue_get_work_rate_limited_count` and structured throttle logs. This carries the exclusion contract from #7586 onto current `main` and composes it with the limiter. Token-bucket pacing is used here because this bounds new work handed to consumers; fn-consumer already owns the separate in-flight concurrency limit. ## Test plan - [x] `cargo fmt --check` - [x] `cargo test -p worker work_queue::config::tests --lib` - [x] `cargo test -p worker work_queue::work_queue_manager::tests --lib` - [x] `cargo test -p worker work_queue::work_queue_server::tests --lib` - [x] `cargo test -p worker fn_consumer::fn_consumer_manager::tests --lib` - [x] `cargo test -p worker fn_consumer::config::tests --lib` - [x] `cargo test -p mdac` - [x] `cargo clippy -p mdac -p worker --lib -- -D warnings` - [x] `git diff --check` The manager regression test uses a one-token bucket with an excluded in-progress function at the head and ready work behind it, verifying the ready function receives the token. Config tests verify the protobuf UUID wire-size assumption and both sides of the concurrency bound. Retry tests cover absolute-deadline conversion, stale deadlines, millisecond rounding, and the token bucket timestamp horizon. ## Migration plan The protobuf additions are backward-compatible. The limiter is optional and disabled when configuration is absent. chroma-core/k8s#3217 enables it with a 100-item burst and one token per 100 ms (10 items/second) after this image is available. Removing the config block disables it. Existing fn-consumer values are far below the new 55,188 maximum. Configurations above that value now fail deserialization rather than risking an oversized `GetWork` request. The limiter is process-local, which is global with the current single WQS replica. Additional WQS replicas would each receive their own allowance. Exclusions are advisory and request-scoped, preserving at-least-once delivery without durable leases. ## Observability plan Monitor `work_queue_get_work_rate_limited_count`, `returned_items`, `excluded_fn_count`, and the structured `rate_limited` field. Watch for shard starvation or a growing queue during staging rollout. ## Documentation changes No user-facing API changes.

Commit:830984d
Author:Tanuj Nayak

[ENH](wqs): Return absolute retry time

Commit:b481d0f
Author:Tanuj Nayak

[BUG](mdac): Preserve retry duration

Commit:6ea7a19
Author:Tanuj Nayak

[BUG](wqs): Preserve function batches

Commit:026af59
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](wqs): Pace GetWork results

Commit:168644a
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](wqs): Pace GetWork results

Commit:93ff040
Author:tanujnay112
Committer:GitHub

[ENH](sysdb): Add database ID lookup (#7709) ## Summary - extend the existing GetDatabase request with an optional database ID selector - implement the ID selector only in the Go SysDB, backed by the database primary key - expose a GET database-by-ID endpoint in the Rust frontend - retain tenant isolation by returning not found for cross-tenant IDs - leave the MCMR rust-sysdb implementation unchanged ## Why Cold database-scoped API-key verification currently resolves an ID by listing every database in the tenant. This adds the constant-time primitive needed to remove that tenant-wide scan for single-region databases. ## Validation - cargo check -p chroma-sysdb -p chroma-frontend -p rust-sysdb - cargo test -p chroma-sysdb test_get_database_by_id_is_tenant_scoped - go test ./pkg/sysdb/coordinator -run TestCatalog_GetDatabaseByIDIsTenantScoped - go test ./pkg/sysdb/grpc ./pkg/sysdb/coordinator -run ^$ The Docker-backed Go integration suites were not run because Docker is unavailable in the local environment.

Commit:61999fa
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](sysdb): Add database ID lookup

Commit:34f8e76
Author:tanujnay112
Committer:GitHub

[ENH](fn-consumer): Show collection IDs in list-in-progress-jobs (#7675) ## Summary Expose the input collection UUIDs for each active fn-consumer job. The fn-consumer now retains the collection IDs from each dispatched batch and returns them through the existing ListInProgressJobs RPC as a backward-compatible repeated field. ## Testing - cargo fmt --all --check - git diff --check - focused worker test build started locally; full validation is delegated to CI ## Compatibility The new protobuf field uses tag 3, so existing clients remain wire-compatible. No migration or deployment configuration changes are required.

Commit:e04a4f3
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](fn-consumer): Show collection IDs

Commit:bee9bb3
Author:Tanuj Nayak

[ENH](fn-consumer): Show collection IDs

Commit:863f232
Author:tanujnay112
Committer:GitHub

[ENH](fn-consumer): List in-progress jobs (#7668) ## Summary Adds ListInProgressJobs to the fn-consumer gRPC service so operators can inspect work running on an individual consumer. - returns the attached function ID and expiry time for each active job - reads the in-progress map through the FnConsumerManager message queue - exposes no mutation or cancellation behavior ## Testing - cargo test -p worker --lib fn_consumer::fn_consumer_manager::tests

Commit:e03f535
Author:Tanuj Nayak

[ENH](fn-consumer): List active jobs

Commit:c38b46f
Author:tanujnay112
Committer:GitHub

[BUG](fn-consumer): Retry unpublished boundary (#7652) ## Summary - gate fn-consumer execution on the committed collection compaction frontier - return a non-error RetryLater outcome while queued work is ahead of SysDB - defer temporarily unready work to the back of the FIFO queue without incrementing its failure count - allow ready work behind a deferred FIFO head to continue instead of being starved - wait for all inputs of a multi-input batch rather than executing a partial invocation ## Why Compaction durably enqueues async function work before registration publishes the corresponding version-file boundary. Fn-consumer can therefore observe a queued frontier while collection.log_position still points to the previous boundary. This is an expected race, but the existing path attempts boundary resolution, reports a workflow error, and increments the attached function failure count. Simply returning RetryLater leaves the same item at the FIFO head. If its boundary remains unpublished, fixed-size GetWork polls can repeatedly return that item and starve ready work behind it. RetryLater now asks the work queue to move the item to the back while preserving its offsets and failure count. Production example: completion offset 22162 was queued with compaction offset 22173. The first consumer attempt ran before registration and failed; the next poll succeeded after 22173 became visible. Failure trace: https://ui.honeycomb.io/chroma-/environments/production-aws--us-east-1/result/ty3EKNUCQPh/trace?trace_id=edd7bbf6f794a3186a2533b8d6bd9097 ## Validation - cargo fmt --all -- --check - cargo test -p worker function_execution --lib - cargo test -p worker fn_consumer_manager::tests --lib - cargo test -p worker work_queue::state::tests --lib - cargo clippy -p worker --lib -- -D warnings

Commit:76441eb
Author:Tanuj Nayak

[BUG](fn-consumer): Defer unpublished work

Commit:c1d5611
Author:Robert Escriva
Committer:Robert Escriva

[ENH](functions) Add memory-based admission for attached functions Introduce pod-local memory admission control so function consumers plan next-boundary execution against a cgroup memory limit instead of a fixed concurrency cap. Admission combines a per-window structural peak-memory estimate with configurable overheads and a safety multiplier, yielding Admit, Wait, or Unaddressable decisions. To feed the estimator, the compactor now records logical workload facts while materializing each compaction window. The new FunctionWorkload descriptor captures data shape (record counts, id, document, metadata, and embedding bytes) rather than a memory prediction, letting consumers apply an estimator matched to the function they run. Consecutive windows merge with saturating adds. These facts propagate from materialize_logs through CollectionFlushInfo and the FlushCollectionCompactionRequest proto into sysdb, and are surfaced on CollectionInfoMutable for consumers. Unknown format versions are treated as unavailable so compactor and consumer rollouts stay decoupled. Memory admission is off by default and gated behind MemoryAdmissionConfig, preserving existing behavior for deployments that do not opt in. Co-authored-by: AI

Commit:72343b8
Author:tanujnay112
Committer:GitHub

[ENH](wqs): Add manual function DLQ RPC (#7612)

Commit:71e7649
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](wqs): Add manual function DLQ RPC

Commit:d747acb
Author:chroma-droid
Committer:GitHub

[HOTFIX] applying PR #7582 to release/2026-08-18 (#7604) This PR cherry-picks the commit 290a4788fa773b49ce899c57eda52240659ac6a5 onto release/2026-08-18. If there are unresolved conflicts, please resolve them manually. Co-authored-by: tanujnay112 <tanujnay112@live.com>

Commit:24ea8c4
Author:tanujnay112
Committer:github-actions[bot]

[ENH](wqs): Add manual function DLQ RPC (#7582) ## Summary - Adds a SysDB RPC to set an attached function’s failure count to an exact non-negative value, scoped by `(fn_id, input_coll_id)`. - Adds WQS `SetFunctionFailureCount`: updates SysDB first, then durably mirrors the returned count into the WQS Parquet snapshot. - Returns `NotFound` when the WQS entry is absent, while leaving SysDB as the authoritative count. ## Testing - Go SysDB DAO test for absolute assignment. - Tilt-gated WQS integration test covering SysDB state and `GetWork` DLQ filtering. - Go package compile check.

Commit:290a478
Author:tanujnay112
Committer:GitHub

[ENH](wqs): Add manual function DLQ RPC (#7582) ## Summary - Adds a SysDB RPC to set an attached function’s failure count to an exact non-negative value, scoped by `(fn_id, input_coll_id)`. - Adds WQS `SetFunctionFailureCount`: updates SysDB first, then durably mirrors the returned count into the WQS Parquet snapshot. - Returns `NotFound` when the WQS entry is absent, while leaving SysDB as the authoritative count. ## Testing - Go SysDB DAO test for absolute assignment. - Tilt-gated WQS integration test covering SysDB state and `GetWork` DLQ filtering. - Go package compile check.

Commit:8408ea2
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](wqs): Add manual function DLQ RPC

Commit:cad0707
Author:Tanuj Nayak

[BUG](worker): Skip in-progress function work

Commit:78ced9b
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](wqs): Add manual work insertion RPC

Commit:1d66e67
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](wqs): Add manual work deletion RPC

Commit:a425453
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](wqs): Add manual function DLQ RPC

Commit:5879991
Author:tanujnay112
Committer:GitHub

[ENH](fn-consumer): Dead-letter fn work (#7572) ## Summary Adds fn-consumer DLQ behavior for attached-function work. - Reports a failed function batch to WQS for each input collection. - Before dispatch, reads the attached function from SysDB. - Logs and skips functions whose failure count has reached the configurable threshold (default: 5). - Leaves skipped queue records in WQS, per the current design.

Commit:c95f10e
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](wqs): Persist function failure counts

Commit:584eb4a
Author:tanujnay112
Committer:GitHub

[ENH](wqs): Report fn failures (#7571) ## Summary Adds the WQS failure-reporting boundary. - Adds `FailFunction(fn_id, input_coll_id)` to the WQS API. - Validates IDs and forwards failures to SysDB for durable count increments. - Does not add WQS-specific DLQ persistence; queue records remain in WQS. - Adds Kubernetes integration coverage for WQS → SysDB failure reporting.

Commit:46e9692
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](wqs): Report fn failures

Commit:1ca5cda
Author:tanujnay112
Committer:GitHub

[ENH](sysdb): Track fn failures (#7570) ## Summary Adds durable DLQ failure tracking for attached functions in SysDB. - Adds `attached_functions.failure_count` with an Atlas migration and updated checksum. - Adds the SysDB failure-reporting API used by WQS. - Resets `failure_count` in the existing successful invocation-completion SQL update; no separate reset query is issued. - Adds DAO coverage for failure-count increments.

Commit:66c6a82
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](wqs): Persist function failure counts

Commit:44be6b5
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](wqs): Report fn failures

Commit:870c402
Author:Tanuj Nayak

[ENH](sysdb): Return attached failure count

Commit:68db2e3
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](sysdb): Track fn failures

Commit:87b3f2b
Author:dbeglord

[CLN](sysdb): Remove min_records_for_invocation The field claimed to control how many new records accumulate before an attached function is invoked. Nothing ever compared it to a record count — every reference across Go and Rust was storage, a re-attach idempotency check, a pass-through, or a display line. Foundation set it to 100 and got the compactor's default of 10. The real gate is `should_compact_at` in the log service: uncompacted record count against the compactor's `min_compaction_size`, plus a reinsert threshold and a time-on-log timeout. None of that is per-function, so there was no knob here to keep. An inert setting is worse than a missing one — it invites tuning that does nothing and misleads anyone reading the config to understand cadence. Proto fields 8 and 9 are reserved rather than freed, so the wire numbers can never be recycled. Removal is safe in both directions during a rolling deploy precisely because the value was never read: an old client still sending it is ignored, and a new client omitting it changes no behavior. The DB column is deliberately left in place. Dropping it in the same release would break old sysdb pods mid-rollout, which still SELECT it; the column is NOT NULL with a default, so new writes fill it harmlessly. A follow-up migration should drop it once this release has rolled out. One unrelated fix was required to land: examples/task_api_example.py treated attach_function's `(function, created)` tuple as a single object, which mypy rejects. Staging the file for the comment update surfaced it. Verified: `cargo check --workspace --all-targets` and `cargo fmt --check` clean; `go build` and `go vet` clean over pkg/sysdb, and vet compiles the test files so the test edits are covered.

Commit:57b5710
Author:Tanuj Nayak
Committer:tanujnay112

[CHORE](sysdb): simplify finish API

Commit:1bc5f24
Author:tanujnay112
Committer:GitHub

[BUG](work-queue): expose compaction frontier to consumer (#7394) ## Description of changes This PR fixes a bug where the fn-consumer work queue was not exposing the **compaction frontier** — the offset up to which the input collection has been compacted. Without this, the fn-consumer couldn't tell whether a queued work item was stale (i.e., the attached function had already processed past the queued offset), leading to redundant or incorrect re-execution. The fix threads `compaction_offset` through the entire pipeline: 1. Proto: add `compaction_offset` to `WorkItemResult` 2. Work queue server: populate it when serving work items 3. Fn-consumer manager: use it to build `FunctionExecutionInput` (replacing the raw `(CollectionUuid, i64)` tuple) 4. Log fetch orchestrator: fetch attached function state first (when in fn-consumer mode), then use the resolved `completion_offset` from sysdb instead of the stale value from the work queue 5. Function execution context: after fetching, skip work items where the attached function has already reached or passed the queued frontier - Improvements & Bug fixes - ... - New functionality - ... ## Test plan _How are these changes tested?_ - [ ] Tests pass locally with `pytest` for python, `yarn test` for js, `cargo test` for rust ## Migration plan _Are there any migrations, or any forwards/backwards compatibility changes needed in order to make sure this change deploys reliably?_ ## Observability plan _What is the plan to instrument and monitor this change?_ ## Documentation Changes _Are all docstrings for user-facing APIs updated if required? Do we need to make documentation changes in the_ [_docs section](https://github.com/chroma-core/chroma/tree/main/docs/docs.trychroma.com)?_

Commit:78d56b6
Author:Tanuj Nayak
Committer:Tanuj Nayak

[BUG](work-queue): expose compaction frontier Return the queue compaction frontier from GetWork while completion_offset is no longer persisted in the WQS state file. Keep the field optional on the wire so older consumers remain compatible during rollout.

Commit:d6d11f0
Author:Tanuj Nayak
Committer:Tanuj Nayak

[BUG](work-queue): expose compaction frontier Return the queue compaction frontier from GetWork while completion_offset is no longer persisted in the WQS state file. Keep the field optional on the wire so older consumers remain compatible during rollout.

Commit:443b7ae
Author:tanujnay112
Committer:GitHub

[ENH](work-queue): require compaction offset (#7358) ## Description of changes This PR makes `compaction_offset` a required field throughout the work queue stack — removing the `Option<i64>` wrapper from the proto, Rust types, and all call sites. This also gets rid of sysdb-based filtering in get_work. - Improvements & Bug fixes - ... - New functionality - ... ## Test plan _How are these changes tested?_ - [ ] Tests pass locally with `pytest` for python, `yarn test` for js, `cargo test` for rust ## Migration plan _Are there any migrations, or any forwards/backwards compatibility changes needed in order to make sure this change deploys reliably?_ ## Observability plan _What is the plan to instrument and monitor this change?_ ## Documentation Changes _Are all docstrings for user-facing APIs updated if required? Do we need to make documentation changes in the_ [_docs section](https://github.com/chroma-core/chroma/tree/main/docs/docs.trychroma.com)?_

Commit:d26b423
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](work-queue): require compaction offset

Commit:ed328f3
Author:Tanuj Nayak
Committer:Tanuj Nayak

[BUG](work-queue): expose compaction frontier Return the queue compaction frontier from GetWork while completion_offset is no longer persisted in the WQS state file. Keep the field optional on the wire so older consumers remain compatible during rollout.

Commit:0f9751c
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](work-queue): require compaction offset

Commit:b4e4361
Author:Robert Escriva
Committer:GitHub

[ENH](logservice): add conditional push and insert offset to proto (#7317) ## Description of changes Extend PushLogsRequest and PushLogsResponse to support conditional log appends: - Add optional PushLogsCondition with observed_log_offset and read_ids - Add optional first_inserted_record_offset to PushLogsResponse Update all PushLogsRequest construction sites and the log-service response to populate the new fields with None defaults. ## Test plan CI ## Migration plan Backwards compatible ## Observability plan N/A ## Documentation Changes N/A Co-authored-by: AI

Commit:03e4a90
Author:Tanuj Nayak
Committer:Tanuj Nayak

[CHORE](sysdb): simplify finish API

Commit:7aaa20b
Author:Tanuj Nayak
Committer:Tanuj Nayak

[BUG](work-queue): expose compaction frontier Return the queue compaction frontier from GetWork while completion_offset is no longer persisted in the WQS state file. Keep the field optional on the wire so older consumers remain compatible during rollout.

Commit:33daa4d
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](work-queue): require compaction offset

Commit:65035af
Author:Tanuj Nayak
Committer:Tanuj Nayak

[CHORE](sysdb): simplify finish API

Commit:cfef885
Author:Tanuj Nayak
Committer:Tanuj Nayak

Revert "[BUG](fn-consumer): Read offsets from sysdb" This reverts commit f4ea0ff9c6e4c494a6ddf912cf985be639ade204.

Commit:2198048
Author:Tanuj Nayak
Committer:Tanuj Nayak

[BUG](fn-consumer): Read offsets from sysdb

Commit:d864335
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](work-queue): require compaction offset

Commit:5af81e3
Author:Tanuj Nayak
Committer:Tanuj Nayak

[BUG](work-queue): expose compaction frontier Return the queue compaction frontier from GetWork while completion_offset is no longer persisted in the WQS state file. Keep the field optional on the wire so older consumers remain compatible during rollout.

Commit:295fbdc
Author:Tanuj Nayak
Committer:Tanuj Nayak

[CHORE](sysdb): simplify finish API

Commit:1dd69d9
Author:Tanuj Nayak
Committer:Tanuj Nayak

[BUG](work-queue): expose compaction frontier Return the queue compaction frontier from GetWork while completion_offset is no longer persisted in the WQS state file. Keep the field optional on the wire so older consumers remain compatible during rollout.

Commit:35b47e4
Author:Tanuj Nayak
Committer:Tanuj Nayak

Revert "[BUG](fn-consumer): Read offsets from sysdb" This reverts commit f4ea0ff9c6e4c494a6ddf912cf985be639ade204.

Commit:1598f57
Author:Tanuj Nayak
Committer:Tanuj Nayak

[BUG](fn-consumer): Read offsets from sysdb

Commit:20ff80a
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](work-queue): require compaction offset

Commit:ac1778f
Author:tanujnay112
Committer:GitHub

[ENH](work-queue): add optional compaction offset (#7354) ## Description of changes Part 1 of a stack of changes to persist the input collection compaction offset instead of the function completion offset in the WQS file. This diff starts off by requiring this compaction offset to be sent in the PushWork API. - Improvements & Bug fixes - - New functionality - ... ## Test plan _How are these changes tested?_ - [ ] Tests pass locally with `pytest` for python, `yarn test` for js, `cargo test` for rust ## Migration plan _Are there any migrations, or any forwards/backwards compatibility changes needed in order to make sure this change deploys reliably?_ ## Observability plan _What is the plan to instrument and monitor this change?_ ## Documentation Changes _Are all docstrings for user-facing APIs updated if required? Do we need to make documentation changes in the_ [_docs section](https://github.com/chroma-core/chroma/tree/main/docs/docs.trychroma.com)?_

Commit:169b663
Author:Tanuj Nayak
Committer:Tanuj Nayak

[CHORE](sysdb): simplify finish API

Commit:615044f
Author:Tanuj Nayak
Committer:Tanuj Nayak

[BUG](fn-consumer): Read offsets from sysdb

Commit:486b99a
Author:Tanuj Nayak
Committer:Tanuj Nayak

[BUG](work-queue): expose compaction frontier Return the queue compaction frontier from GetWork while completion_offset is no longer persisted in the WQS state file. Keep the field optional on the wire so older consumers remain compatible during rollout.

Commit:76b70aa
Author:Tanuj Nayak
Committer:Tanuj Nayak

Revert "[BUG](fn-consumer): Read offsets from sysdb" This reverts commit f4ea0ff9c6e4c494a6ddf912cf985be639ade204.

Commit:0d4e9b5
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](work-queue): require compaction offset

Commit:566e9f4
Author:Tanuj Nayak
Committer:Tanuj Nayak

[CHORE](sysdb): simplify finish API

Commit:f4ea0ff
Author:Tanuj Nayak
Committer:Tanuj Nayak

[BUG](fn-consumer): Read offsets from sysdb

Commit:b0df5ac
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](work-queue): require compaction offset

Commit:0f14332
Author:Robert Escriva
Committer:Robert Escriva

document new fields

Commit:c384b77
Author:Tanuj Nayak
Committer:Tanuj Nayak

[CHORE](sysdb): simplify finish API

Commit:31bf79f
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](work-queue): require compaction offset

Commit:215e11d
Author:Tanuj Nayak
Committer:Tanuj Nayak

[CHORE](sysdb): simplify finish API

Commit:e9e877e
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](work-queue): require compaction offset

Commit:2502aa4
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](work-queue): add optional compaction offset

Commit:d2182db
Author:Tanuj Nayak
Committer:Tanuj Nayak

[CHORE](sysdb): simplify finish API

Commit:b78eb70
Author:Tanuj Nayak
Committer:Tanuj Nayak

[ENH](work-queue): require compaction offset

Commit:9d44034
Author:Tanuj Nayak

[ENH](work-queue): add optional compaction offset

Commit:8e5bda3
Author:Robert Escriva
Committer:Robert Escriva

[ENH](logservice): add conditional push and insert offset to proto Extend PushLogsRequest and PushLogsResponse to support conditional log appends: - Add optional PushLogsCondition with observed_log_offset and read_ids - Add optional first_inserted_record_offset to PushLogsResponse Update all PushLogsRequest construction sites and the log-service response to populate the new fields with None defaults. Co-authored-by: AI

Commit:3f8fa38
Author:Robert Escriva
Committer:Robert Escriva

[ENH](logservice): add conditional push and insert offset to proto Extend PushLogsRequest and PushLogsResponse to support conditional log appends: - Add optional PushLogsCondition with observed_log_offset and read_ids - Add optional first_inserted_record_offset to PushLogsResponse Update all PushLogsRequest construction sites and the log-service response to populate the new fields with None defaults. Co-authored-by: AI

Commit:54c196d
Author:Robert Escriva
Committer:Robert Escriva

[ENH](logservice): add conditional push and insert offset to proto Extend PushLogsRequest and PushLogsResponse to support conditional log appends: - Add optional PushLogsCondition with observed_log_offset and read_ids - Add optional first_inserted_record_offset to PushLogsResponse Update all PushLogsRequest construction sites and the log-service response to populate the new fields with None defaults. Co-authored-by: AI

Commit:dd0ea9a
Author:Robert Escriva
Committer:Robert Escriva

[ENH](logservice): add conditional push and insert offset to proto Extend PushLogsRequest and PushLogsResponse to support conditional log appends: - Add optional PushLogsCondition with observed_log_offset and read_ids - Add optional first_inserted_record_offset to PushLogsResponse Update all PushLogsRequest construction sites and the log-service response to populate the new fields with None defaults. Co-authored-by: AI

Commit:dfd4413
Author:Robert Escriva
Committer:Robert Escriva

[ENH](logservice): add conditional push and insert offset to proto Extend PushLogsRequest and PushLogsResponse to support conditional log appends: - Add optional PushLogsCondition with observed_log_offset and read_ids - Add optional first_inserted_record_offset to PushLogsResponse Update all PushLogsRequest construction sites and the log-service response to populate the new fields with None defaults. Co-authored-by: AI

Commit:7331b12
Author:Tanuj Nayak
Committer:tanujnay112

[ENH]: Finish_work and output collection flush are atomic

Commit:6bd423f
Author:tanujnay112
Committer:GitHub

[ENH]: Add scaffolding for functions with many inputs in fn consumer (#7143) ## Description of changes This PR upgrades the fn consumer pipeline to support functions with **multiple input collections**. Previously, `AttachedFunctionExecutor::execute` received a single `Chunk` of records; now it receives a `Vec<Chunk>` — one chunk per input collection. The orchestration layer is restructured to carry batches of `(collection_id, materialized_logs)` pairs through the pipeline, and the fn consumer manager groups work items by function before dispatching them as a single batch. The key structural changes are: 1. `AttachedFunctionExecutor::execute` signature: `Chunk<…>` → `Vec<Chunk<…>>` 2. New `FunctionExecutionBatch` / `FunctionExecutionProgress` / `FunctionContext` types extracted into a new `function_execution` module 3. `ExecuteAttachedFunctionInput` gains `input_batches: Vec<ExecuteAttachedFunctionBatchInput>` replacing the flat `materialized_logs` + `input_record_segment` 4. `FinishAsyncWorkInput` now carries `Vec<FinishAsyncWorkItem>` instead of a single offset 5. `FnConsumerManager` groups work items by function ID and dispatches them as a batch via `FunctionExecutionContext::run` - Improvements & Bug fixes - ... - New functionality - ... ## Test plan _How are these changes tested?_ - [ ] Tests pass locally with `pytest` for python, `yarn test` for js, `cargo test` for rust ## Migration plan _Are there any migrations, or any forwards/backwards compatibility changes needed in order to make sure this change deploys reliably?_ ## Observability plan _What is the plan to instrument and monitor this change?_ ## Documentation Changes _Are all docstrings for user-facing APIs updated if required? Do we need to make documentation changes in the_ [_docs section](https://github.com/chroma-core/chroma/tree/main/docs/docs.trychroma.com)?_

Commit:ed66c10
Author:Tanuj Nayak
Committer:Tanuj Nayak

[CLN](worker): Box function execution fetch future

Commit:c136e32
Author:tanujnay112
Committer:GitHub

[ENH]: Add route to allow multiple inputs on an async function (#7142) ## Description of changes This PR adds a new `add_input` route that lets callers associate additional input collections with an existing async attached function. The key data model insight: the `attached_functions` table uses a composite primary key of `(id, input_collection_id)`, so one logical attached function can have multiple rows — one per input collection. Only async functions support multiple inputs. The change spans the full stack: protobuf → Go sysdb coordinator → Rust sysdb client → Rust frontend server → Python API client. - Improvements & Bug fixes - ... - New functionality - ... ## Test plan _How are these changes tested?_ - [ ] Tests pass locally with `pytest` for python, `yarn test` for js, `cargo test` for rust ## Migration plan _Are there any migrations, or any forwards/backwards compatibility changes needed in order to make sure this change deploys reliably?_ ## Observability plan _What is the plan to instrument and monitor this change?_ ## Documentation Changes _Are all docstrings for user-facing APIs updated if required? Do we need to make documentation changes in the_ [_docs section](https://github.com/chroma-core/chroma/tree/main/docs/docs.trychroma.com)?_

Commit:78e4719
Author:Tanuj Nayak
Committer:Tanuj Nayak

[TEST](coordinator): Fix attach function collection mocks

Commit:2cf2a47
Author:tanujnay112
Committer:GitHub

[BUG]: WQS repairs on bootup (#7126) ## Description of changes This PR renames `AreInvocationsDone` (which returned `[]bool`) to `CheckInvocationStatus` (which returns a three-state enum). The key semantic change is that the old binary done/not-done is replaced with three states: - `NOT_DONE` (0) — default, work still in progress - `DONE` (1) — completed (soft/hard deleted, or offset advanced with no pending heap entry) - `NEEDS_REPAIR` (2) — offset advanced but `heap_entry_pending=true`, meaning the WQS crashed mid-repair The new `NEEDS_REPAIR` state enables the work queue manager to automatically repair stale entries on bootup, which is the core bug fix this PR addresses. In the proto we replaced the RPC for `AreInvocationsDone` with ``CheckInvocationStatus` this is not backwards compatible but that is fine as there is not production use for the former RPC yet. - Improvements & Bug fixes - ... - New functionality - ... ## Test plan _How are these changes tested?_ - [ ] Tests pass locally with `pytest` for python, `yarn test` for js, `cargo test` for rust ## Migration plan _Are there any migrations, or any forwards/backwards compatibility changes needed in order to make sure this change deploys reliably?_ ## Observability plan _What is the plan to instrument and monitor this change?_ ## Documentation Changes _Are all docstrings for user-facing APIs updated if required? Do we need to make documentation changes in the_ [_docs section](https://github.com/chroma-core/chroma/tree/main/docs/docs.trychroma.com)?_

Commit:30211ce
Author:Tanuj Nayak
Committer:Tanuj Nayak

[BUG]: WQS repairs on bootup

Commit:9db5d08
Author:tanujnay112
Committer:GitHub

[BUG]: Fn consumer manager keys its job scheduler on output collection id (#7117) ## Description of changes The `FnConsumerManager` was keying its "in-progress" job scheduler on `(fn_id, input_coll_id)` — a `FnJobKey`. The bug: multiple different attached functions could share the same output collection, and the old key didn't prevent two jobs from writing to the same output collection concurrently. The fix re-keys the in-progress map on `output_coll_id` (`CollectionUuid`), which is the actual resource that must be serialized. To support this, `GetAttachedFunctions` gains a batch `ids` parameter so the manager can look up all output collection IDs in a single sysdb call rather than one per work item. - Improvements & Bug fixes - ... - New functionality - ... ## Test plan _How are these changes tested?_ - [ ] Tests pass locally with `pytest` for python, `yarn test` for js, `cargo test` for rust ## Migration plan _Are there any migrations, or any forwards/backwards compatibility changes needed in order to make sure this change deploys reliably?_ ## Observability plan _What is the plan to instrument and monitor this change?_ ## Documentation Changes _Are all docstrings for user-facing APIs updated if required? Do we need to make documentation changes in the_ [_docs section](https://github.com/chroma-core/chroma/tree/main/docs/docs.trychroma.com)?_

Commit:be08eca
Author:Tanuj Nayak
Committer:Tanuj Nayak

[BUG]: Fn consumer manager keys its job scheduler on output collection id

Commit:d5090bb
Author:Robert Escriva
Committer:Robert Escriva

[ENH] implement stratified sampling endpoint for collections Add a new sample API that randomly selects records from a collection using stratified sampling over SPANN index strata. Records can be filtered by ID, metadata, or document predicates before sampling. The endpoint accepts a limit and optional seed for reproducibility. Full-stack implementation across all layers: - API types and request/response models (Rust, Python, JS/TS) - Protobuf IDL for distributed query execution - Frontend server routing, metrics, and executor integration - Sample operator with SPANN-aware stratified selection - Sample orchestrator with filter, projection, and shard fan-out - SPANN index reader support for head-count enumeration - Python and JS client bindings Co-authored-by: AI