These commits are when the Protocol Buffers files have changed: (only the last 100 relevant commits are shown)
| Commit: | 27a06ec | |
|---|---|---|
| Author: | Meg Stepp | |
| Committer: | Meg Stepp | |
feat(ctrl): add the portal domain service and verification workflow
| Commit: | 40b31f5 | |
|---|---|---|
| Author: | Flo | |
ctrl: remove connection revocations from discovery Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf Co-authored-by: Amp <amp@ampcode.com>
| Commit: | d720a27 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
ctrl: describe replica DNS without retired enrollment gates Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf Co-authored-by: Amp <amp@ampcode.com>
| Commit: | bb14ec5 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
ctrl: share versioned private network snapshots Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 9534fa8 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
ctrl: publish private connections and region topology Stream complete caller-scoped connection snapshots and cluster-region topology from Ctrl. Include stored deployment capabilities in the state Krane receives. Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 6d124cf | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
deploy: select and retain private connection targets Select connection targets within the requested environment and preserve pinned deployments through replacement, promotion, and rollback. Persist the private-network capability at creation and tolerate GitHub clock skew. Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 478ceba | |
|---|---|---|
| Author: | Flo | |
ctrl: describe replica DNS without retired enrollment gates Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 71ad8f5 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
ctrl: publish private connections and region topology Stream complete caller-scoped connection snapshots and cluster-region topology from Ctrl. Include stored deployment capabilities in the state Krane receives. Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf Co-authored-by: Amp <amp@ampcode.com>
| Commit: | a91187b | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
ctrl: share versioned private network snapshots Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 755452a | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
deploy: select and retain private connection targets Select connection targets within the requested environment and preserve pinned deployments through replacement, promotion, and rollback. Persist the private-network capability at creation and tolerate GitHub clock skew. Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 64b3b94 | |
|---|---|---|
| Author: | Flo | |
| Committer: | GitHub | |
fix(ctrl): keep newest deployment live (#7152) * fix(ctrl): keep newest deployment live Amp-Thread-ID: https://ampcode.com/threads/T-01a026f7-b218-7099-8b5e-752206b8a824 * refactor(ctrl): simplify live deployment swap handling Reuse transaction queries and construct the swap response in one place. Derive the environment from the deployment instead of passing both. Remove redundant comments and name the duplicate-request test result accurately; it does not exercise Restate journal replay. Amp-Thread-ID: https://ampcode.com/threads/T-01a07735-3f08-768b-a68d-96c1243ec8c6 Co-authored-by: Amp <amp@ampcode.com> * refactor(ctrl): separate swap transaction from Restate handler Keep journaling and response handling in the Restate handler. Move the atomic database operation behind a plain context.Context boundary without changing journal steps or promotion decisions. Scope deployment-step errors to their callbacks and use each callback's context consistently. Amp-Thread-ID: https://ampcode.com/threads/T-01a07735-3f08-768b-a68d-96c1243ec8c6 Co-authored-by: Amp <amp@ampcode.com> * test(ctrl): adapt routing fixture to current app schema Amp-Thread-ID: https://ampcode.com/threads/T-01a07735-3f08-768b-a68d-96c1243ec8c6 Co-authored-by: Amp <amp@ampcode.com> * fix(ctrl): separate Restate and promotion error scopes Amp-Thread-ID: https://ampcode.com/threads/T-01a07735-3f08-768b-a68d-96c1243ec8c6 Co-authored-by: Amp <amp@ampcode.com> --------- Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 9e86aa2 | |
|---|---|---|
| Author: | Flo | |
ctrl: publish private connections and region topology Stream complete caller-scoped connection snapshots and cluster-region topology from Ctrl. Include stored deployment capabilities in the state Krane receives. Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf Co-authored-by: Amp <amp@ampcode.com>
| Commit: | ab5bc1c | |
|---|---|---|
| Author: | Flo | |
deploy: select and retain private connection targets Select connection targets within the requested environment and preserve pinned deployments through replacement, promotion, and rollback. Persist the private-network capability at creation and tolerate GitHub clock skew. Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 605d5f9 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
private DNS: add opt-in region selection Filter authorized ready endpoints by region after revision activation. Keep ordinary and self-discovery names all-region, and expose the own hostname as UNKEY_PRIVATE_DOMAIN. Local-first falls back only when no local endpoints are ready; explicit regions never fall back. Publish topology after connection cleanup, with a separate bounded call outside the endpoint lock, so optional metadata cannot block revocation. Co-authored-by: Amp <amp@ampcode.com> Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf
| Commit: | 7e4f0eb | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
private network: rename bindings to connections Use connection terminology across schema, RPCs, policies, DNS, dashboard, and documentation. Keep resource IDs, DNS aliases and protobuf field numbers unchanged. Existing installations require the documented coordinated table rename and binary upgrade; old dashboard URLs redirect. Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 9d39cf5 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
ctrl: name private discovery entries as bindings Make caller and target ownership explicit in the discovery contract. Keep protobuf field numbers and complete-snapshot behavior unchanged. Co-authored-by: Amp <amp@ampcode.com> Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf
| Commit: | 9b157ac | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
ctrl: stream complete snapshots for enrolled workspaces Read every page in one transaction, then stream chunks with an explicit completion marker. Send the replica hostname only for enrolled workspaces and valid app slugs so Krane can opt each deployment into private DNS. Co-authored-by: Amp <amp@ampcode.com> Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf
| Commit: | 5793d89 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
ctrl: expose caller-scoped private discovery catalog Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf Co-authored-by: Amp <amp@ampcode.com>
| Commit: | de95620 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
deploy: preserve bound targets across replacement and rollback Snapshot binding hostnames into deployment secrets. Keep replaced targets available during overlap, defer automatic stops while pinned, and limit live-pointer updates to the built-in production environment. Co-authored-by: Amp <amp@ampcode.com> Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf
| Commit: | ce0fba5 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
private DNS: add opt-in region selection Filter authorized ready endpoints by region after revision activation. Keep ordinary and self-discovery names all-region, and expose the own hostname as UNKEY_PRIVATE_DOMAIN. Local-first falls back only when no local endpoints are ready; explicit regions never fall back. Publish topology after connection cleanup, with a separate bounded call outside the endpoint lock, so optional metadata cannot block revocation. Co-authored-by: Amp <amp@ampcode.com> Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf
| Commit: | 4f20d5f | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
private network: rename bindings to connections Use connection terminology across schema, RPCs, policies, DNS, dashboard, and documentation. Keep resource IDs, DNS aliases and protobuf field numbers unchanged. Existing installations require the documented coordinated table rename and binary upgrade; old dashboard URLs redirect. Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf Co-authored-by: Amp <amp@ampcode.com>
| Commit: | d96448d | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
ctrl: name private discovery entries as bindings Make caller and target ownership explicit in the discovery contract. Keep protobuf field numbers and complete-snapshot behavior unchanged. Co-authored-by: Amp <amp@ampcode.com> Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf
| Commit: | 4e86411 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
ctrl: stream complete snapshots for enrolled workspaces Read every page in one transaction, then stream chunks with an explicit completion marker. Send the replica hostname only for enrolled workspaces and valid app slugs so Krane can opt each deployment into private DNS. Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf Co-authored-by: Amp <amp@ampcode.com>
| Commit: | b2d43ff | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
ctrl: expose caller-scoped private discovery catalog Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf Co-authored-by: Amp <amp@ampcode.com>
| Commit: | e8e16ce | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
deploy: preserve bound targets across replacement and rollback Snapshot binding hostnames into deployment secrets. Keep replaced targets available during overlap, defer automatic stops while pinned, and limit live-pointer updates to the built-in production environment. Co-authored-by: Amp <amp@ampcode.com> Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf
| Commit: | 06bf697 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
fix(ctrl): retain removal work until regional confirmation Derive permanent removal from missing or deleting ancestors. Recover lost acknowledgments and notify Krane when an already-stopped topology needs cleanup, without reconciling settled snapshot rows. Deploy ctrl before the new Krane client, which requires deployment IDs on delete events. Amp-Thread-ID: https://ampcode.com/threads/T-01a0c960-4827-7091-bd71-bb8b7aa78f9c Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 1a97207 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
fix(ctrl): retain removal work until regional confirmation Derive permanent removal from missing or deleting ancestors. Recover lost acknowledgments and notify Krane when an already-stopped topology needs cleanup, without reconciling settled snapshot rows. Deploy ctrl before the new Krane client, which requires deployment IDs on delete events. Amp-Thread-ID: https://ampcode.com/threads/T-01a0c960-4827-7091-bd71-bb8b7aa78f9c Co-authored-by: Amp <amp@ampcode.com>
| Commit: | bd939f9 | |
|---|---|---|
| Author: | Flo | |
certificate: fix renewal results and health metrics Treat Restate Send results as invocation handles, not errors. Return real challenge failures and journal rate-limit delays before durable sleeps. Remove response status strings and reserve their protobuf field numbers. Count issuance attempts in handlers and expose current certificate state through an isolated scrape-driven collector. Database failures fail the scrape rather than retaining stale health values. Amp-Thread-ID: https://ampcode.com/threads/T-01a0f129-3914-75bb-b7b3-25c6d1adf08f Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 478faea | |
|---|---|---|
| Author: | Oz | |
| Committer: | GitHub | |
feat(frontline): add remoteIp match expression (#7638) * feat(frontline): match policies by client IP range **tldr;** policies can now match on the client IP. Frontline reads a new `remoteIp` match expression with `in` or `notIn` CIDR lists. Nothing writes it yet, the API change is the next PR. ```sh in: [203.0.113.0/24] + ACTION_DENY block that range notIn: [198.51.100.0/24] + ACTION_DENY block everyone but the office ``` The client IP is the one `X-Forwarded-For` already carries to the app: the TCP peer, or the signed metadata IP after a cross-region hop. Client headers are never read. Checked on prod, a spoofed `X-Forwarded-For` is replaced. ## Also changed - Invalid regex and invalid CIDR entries return `InvalidConfiguration` (422) instead of an uncoded 500 that pages us. - `zen.Session.ClientIP()` returns the `netip.Addr` behind `Location()`. ## Heads up - Ship this to frontline and the API before the API accepts `remoteIp`. Old frontlines drop the field and skip the policy, and old API pods strip it from sibling policies on `updatePolicy`. - No rollback past this commit once customers use `remoteIp`. * fix(frontline): align remoteIp comments and docs with repo terms * docs(frontline): say when an invalid remoteIp entry fails the request * docs(frontline): simplify the RemoteIpMatch proto comment * feat(frontline): accept bare IP addresses in remoteIp lists * fix(frontline): address pre-review findings on remoteIp * feat(frontline): accept IPv4 remoteIp entries only * feat(frontline): support ipv6 entries * chore: tidy up proto name and use remote_ip * fix: typo
| Commit: | 9a38280 | |
|---|---|---|
| Author: | Flo | |
deploy: reconcile container state after missed events Include container observations in deployment snapshots so periodic resync can clear stale waiting errors without relying on another watch event. Reconcile observations in the instance-upsert transaction and retain exit history independently of current-state ordering. Keep other containers out of the application status record. Cover delivery permutations, missing rows, counter resets, and failed resync delivery. Amp-Thread-ID: https://ampcode.com/threads/T-01a0e84e-32a5-72a9-8243-4ba869bde30c Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 1a3501b | |
|---|---|---|
| Author: | Flo | |
ctrl: name private discovery entries as bindings Make caller and target ownership explicit in the discovery contract. Keep protobuf field numbers and complete-snapshot behavior unchanged. Co-authored-by: Amp <amp@ampcode.com> Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf
| Commit: | c1a669f | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
ctrl: stream complete snapshots for enrolled workspaces Read every page in one transaction, then stream chunks with an explicit completion marker. Send the replica hostname only for enrolled workspaces and valid app slugs so Krane can opt each deployment into private DNS. Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 6272cd6 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
ctrl: expose caller-scoped private discovery catalog Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 01d3142 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
deploy: preserve bound targets across replacement and rollback Snapshot binding hostnames into deployment secrets. Keep replaced targets available during overlap, defer automatic stops while pinned, and limit live-pointer updates to the built-in production environment. Co-authored-by: Amp <amp@ampcode.com> Amp-Thread-ID: https://ampcode.com/threads/T-01a0867e-f00a-7718-961d-5d9cf2250eaf
| Commit: | 8887bda | |
|---|---|---|
| Author: | ogzhanolguncu | |
feat(frontline): match policies by client IP range **tldr;** policies can now match on the client IP. Frontline reads a new `remoteIp` match expression with `in` or `notIn` CIDR lists. Nothing writes it yet, the API change is the next PR. ```sh in: [203.0.113.0/24] + ACTION_DENY block that range notIn: [198.51.100.0/24] + ACTION_DENY block everyone but the office ``` The client IP is the one `X-Forwarded-For` already carries to the app: the TCP peer, or the signed metadata IP after a cross-region hop. Client headers are never read. Checked on prod, a spoofed `X-Forwarded-For` is replaced. ## Also changed - Invalid regex and invalid CIDR entries return `InvalidConfiguration` (422) instead of an uncoded 500 that pages us. - `zen.Session.ClientIP()` returns the `netip.Addr` behind `Location()`. ## Heads up - Ship this to frontline and the API before the API accepts `remoteIp`. Old frontlines drop the field and skip the policy, and old API pods strip it from sibling policies on `updatePolicy`. - No rollback past this commit once customers use `remoteIp`.
| Commit: | 099ddd7 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
fix(ctrl): keep newest deployment live Amp-Thread-ID: https://ampcode.com/threads/T-01a026f7-b218-7099-8b5e-752206b8a824
| Commit: | dbccb03 | |
|---|---|---|
| Author: | Andreas Thomas | |
| Committer: | GitHub | |
feat(logdrain): add HEC delivery format (#7605) * feat(logdrain): add HEC delivery format Add HEC configuration and dashboard selection, split format-specific delivery, and retain shared payload encoding with destination-owned timestamps. Reject truncated HEC acknowledgments and unspecified HTTP formats. Amp-Thread-ID: https://ampcode.com/threads/T-01a0d4c4-b92f-748e-a503-9137ec1be40f Co-authored-by: Amp <amp@ampcode.com> * fix(logdrain): isolate lease integration test database Keep global lease acquisition from stealing engine scheduling fixtures in shared MySQL. Add a private MySQL helper and verify row isolation. Amp-Thread-ID: https://ampcode.com/threads/T-01a0d4c4-b92f-748e-a503-9137ec1be40f Co-authored-by: Amp <amp@ampcode.com> --------- Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 2130c10 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
ctrl: process durable failure events for opt-in apps Poll the OOM and crash-loop inbox once per minute with bounded pages, fanout, and retries. Commit alert updates and handled identities in one MySQL transaction, then reconcile the existing per-group Restate object. Fresh evidence resets recovery without advancing five-minute ledgers. Keep collection workspace-gated and the local caller suspended. Retain processed identities for the beta so delayed replays cannot reopen alerts. Real MySQL/Restate tests cover killed invocations, delayed and duplicate events, topology suppression, recovery, capped sweeps, and 100 affected apps with 100,000 older processed identities. Amp-Thread-ID: https://ampcode.com/threads/T-01a065c7-a7fe-746a-a6f4-54a13271f6af Co-authored-by: Amp <amp@ampcode.com>
| Commit: | dd406fa | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
fix(ctrl): chase missed anomaly shard windows Amp-Thread-ID: https://ampcode.com/threads/T-01a065e3-7fa4-755e-9b34-f6dd5de842ef Co-authored-by: Amp <amp@ampcode.com>
| Commit: | ec5e1aa | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
fix(ctrl): serialize deploy anomaly shard handoff Amp-Thread-ID: https://ampcode.com/threads/T-01a065e3-7fa4-755e-9b34-f6dd5de842ef Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 6c4271f | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
fix(ctrl): derive anomaly lifetime from app creation Amp-Thread-ID: https://ampcode.com/threads/T-01a065e3-7fa4-755e-9b34-f6dd5de842ef Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 38b997c | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
fix(ctrl): delay anomaly evaluation for ingest settling Amp-Thread-ID: https://ampcode.com/threads/T-01a065e3-7fa4-755e-9b34-f6dd5de842ef Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 046c0e2 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
fix(ctrl): suppress request drops for stopped topology Amp-Thread-ID: https://ampcode.com/threads/T-01a065e3-7fa4-755e-9b34-f6dd5de842ef Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 96d5873 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
feat(ctrl): shard deploy anomaly evaluation Fan each closed window into stable workspace partitions, batch metadata reads, and keep incomplete telemetry from mutating alert state. Persist open, touch, and resolve transitions without sending notifications; a separate consumer will notify later. Amp-Thread-ID: https://ampcode.com/threads/T-01a065e3-7fa4-755e-9b34-f6dd5de842ef Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 042dd05 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
feat(ctrl): refine deploy anomaly signals and contracts Use robust traffic-drop detection, error-rate thresholds, telemetry watermarks, and instance-averaged memory utilization. Add the SQLC and protobuf contracts needed by the production Restate worker, including workspace notification muting. Amp-Thread-ID: https://ampcode.com/threads/T-01a065e3-7fa4-755e-9b34-f6dd5de842ef Co-authored-by: Amp <amp@ampcode.com>
| Commit: | ca853ed | |
|---|---|---|
| Author: | Andreas Thomas | |
feat(logdrain): add HEC delivery format Add HEC configuration and dashboard selection, split format-specific delivery, and retain shared payload encoding with destination-owned timestamps. Reject truncated HEC acknowledgments and unspecified HTTP formats. Amp-Thread-ID: https://ampcode.com/threads/T-01a0d4c4-b92f-748e-a503-9137ec1be40f Co-authored-by: Amp <amp@ampcode.com>
| Commit: | f52cb0c | |
|---|---|---|
| Author: | Andreas Thomas | |
fix(deploy): fail promptly on fatal startup events Amp-Thread-ID: https://ampcode.com/threads/T-01a0af6a-9266-7075-ad10-8e5c2f29fe8b Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 8bfd217 | |
|---|---|---|
| Author: | Oz | |
| Committer: | GitHub | |
ctrl(worker): drop buildslot and replace with restate flow control (#7453) * feat(worker): add new build handler * chore(worker): address review sweep findings * chore: docs * fix: wording * chore: improve wording * fix(worker): address advisory findings on the build handler * fix(cron): delete stale build rules at their listed version * docs: plain language for the build concurrency section * chore(worker): trim wording and keep the build retry policy with the binds * chore: doc * test(worker): pin that a waiting callee keeps its flow control slot A Restate concurrency slot is held for the whole running attempt, waits included, and is freed only when Restate suspends the invocation after the inactivity timeout. A rule on the deploying phase would therefore hold a slot through the wait for pods and then re-queue the resumed deployment behind newer work. Only builds are capped; the new case and the doc bullet say why. * fix(worker): serialise the topology quota check per workspace The build slot on main was held to the end of Deploy and so serialised the quota check in createTopologies by accident. With the slot narrowed to the build, two deployments of one workspace could both sum the running topologies and both insert. reserveTopologies now locks the workspace's limits row, sums the other deployments' topologies and inserts in one transaction. The sum excludes the deployment's own rows so a re-executed Run does not refuse a deployment that fits. The quota charges the clamped replica count that is actually inserted. * docs: drop the deploying-phase note from the deployments page * refactor(db): rename ListWorkspaceBuildConcurrencyAbove to ListLimitsWithBuildsConcurrentMaxAbove Table first like the other limits queries, with the column in the name * fix(worker): bound the build backend and correct the cancel docs Build holds the workspace's Restate build slot for as long as its restate.Run is running, and the Depot backend that production uses had no deadline of its own, so a backend that never answers stopped every other deployment in the workspace. withBuildkit now derives a 30 minute context that both backends and the build closure take. It is a child of the run context, so Restate cancellation still propagates, and it equals buildImageRetryCeiling so an attempt that hits it is not retried. The cancel comments, the package doc, and the product and engineering pages claimed a cancel frees the slot and that a CLI image deploy skips the queue. A running Build keeps its slot until its image build returns, and Deploy calls Build for every source type. Also move Build's limit key check above the terminal status return, journal the queued step's end timestamp so a retry reports the queue wait it measured rather than the time since the row was written, and drop EndActiveDeploymentStepsWithError, which lost its last caller with buildslot. * fix(worker): let a cancel abort a running build Restate suspends an invocation after about a minute of inactivity, and restate.Run's context is cancelled with the invocation's http2 stream rather than on suspend. A suspended Build therefore could not be told it was cancelled: it finished its image build and held the workspace's build slot until it did, while every other deployment in the workspace stayed pending. Cancelling in the first minute worked, which is why this was easy to miss. Bind Build with WithInactivityTimeout(BuildKeepAliveWindow) so it stays unsuspended for the whole build and a cancel reaches the image build through the Run context. The window sits above buildBackendDeadline so the build's own bound always fires first. TestCancelAbortsRunningBuild pins it: the caller suspends, the cancel lands, and the queued sibling starts. It waits past the one minute boundary on purpose, because a cancel before it passes either way. * docs(worker): say what the keep alive window constrains, not how The comments explained Restate's transport instead of the rule a reader acts on: keep the window above buildBackendDeadline or a cancel stops working for every build that outlives it. * docs: drop the CLI image note from the queued step It described an incidental consequence of Build being called for every source type as if it were intended behaviour. * docs(worker): explain the build bounds with a scenario * refactor(worker): drop the starting step It held one in-memory runtime settings check and no I/O, so a deployment sat at status=starting for the two writes it took to open and close the step and then went straight to building. The check now runs first inside the building step, which keeps its terminal error and the message the API classifier turns into InvalidRuntimeSettings. Nothing writes the starting step or status any more. Both enum values stay for rows that recorded them before this. * fix(worker): resolve a pre-built image without taking a build slot Build is concurrency limited because building is expensive. A pre-built image only needs its digest resolved, so queueing it behind someone else's build delays a deployment without limiting anything: a long build would hold up an image deploy that has no build at all. Deploy now calls Build for a git source only. A pre-built image ends its queued step and resolves in Deploy, unscoped, and records no building step it did not earn. The runtime settings backstop moves to Deploy as well, since it validates the deployment row rather than the build, which also keeps its InvalidRuntimeSettings message: an error returned from Build loses its public message crossing the Restate service boundary, and Deploy now only falls back to the generic one when there is none. * refactor(worker): give each deployment step its caller's context DeploymentStep served both Deploy, which has a WorkflowContext, and the shared Build handler, which has a WorkflowSharedContext, so its parameter had to narrow to restate.Context and the callback lost its own. Making it generic over the context type hands each callback exactly what its caller holds, and waitForDeployments needing a WorkflowContext is type checked through the parameter rather than through a captured variable. Go has no type parameters on methods, so it is a function now. * fix(worker): journal the superseded status write on its own UpdateDeploymentStatus is unconditional, and it shared a Run with the step write. A step write that failed retried both, so the retry could stamp superseded over a status that had changed in between. Separate Runs mean the status write is recorded once. * refactor(dashboard): drop the starting step from the deployment timeline Nothing writes the step any more, so resolveDeploymentStep saw a nullish step with no implicitlyComplete and resolved it to pending. Every new deployment showed a grey "Deployment Starting" card reading "Preparing deployment for building", and it stayed there after the deployment went ready. The skipped view hardcoded the same card without reading the step data, so it listed six steps where the other views listed five. The enum value stays in the drizzle schema and in DEPLOYMENT_STATUSES. The steps procedure validates its output keys against that enum, and statusGroupOf throws on a status it cannot group, so a row recorded before the step was dropped still renders. * refactor(worker): drop the unreachable runtime settings check Deploy re-validated port, cpu and memory off the deployment row. Create reads FindDeployTarget once, validates that struct, and writes the row from the same struct, so the values Deploy read were always ones Create had already approved. Nothing updates those three columns after the insert, and Create is Deploy's only caller. The approval path is the one place the row and the validation can disagree: the second Create validates current settings, its insert hits a duplicate key, and the swallow compares only app id and status, so the row keeps its original snapshot. That snapshot passed the same bounds when it was written. Since the check moved out of the step wrapper it wrote no step error, so deriveError already fell through to unknown rather than the InvalidRuntimeSettings its comment claimed. The classifier keeps the three runtime rules for rows written while the check still ran inside the starting step, which is the step the test now names. * docs(worker): Build never resolves an image buildImage fails terminally on a non-git source and Deploy only calls Build when the source needs building, which is what the handler's own doc comment says. * docs: say building where deployments used to say starting The three supersede comments named the step a deployment commits at, and that step is building now. The OpenAPI example offered starting as a step a failure can name, which no new deployment reaches, and the cancel docs listed it as a state a user can cancel from. The status enums keep starting: old rows carry it and the CLI filter still accepts it.
| Commit: | a3f8c63 | |
|---|---|---|
| Author: | Andreas Thomas | |
| Committer: | GitHub | |
feat(ctrl): stream deployment changes through Vitess (#7418) * feat(ctrl): stream deployment changes through Vitess Amp-Thread-ID: https://ampcode.com/threads/T-01a0a03f-4343-73ad-b2ee-b836285de946 Co-authored-by: Amp <amp@ampcode.com> * chore(dev): remove legacy MySQL volume preservation Amp-Thread-ID: https://ampcode.com/threads/T-01a0a03f-4343-73ad-b2ee-b836285de946 Co-authored-by: Amp <amp@ampcode.com> * fix: address VStream deployment sync review findings * refactor: extract reusable CDC client and deployment wrapper * docs: clarify CDC lifecycle and callback contracts * refactor(cdc): separate stream reads and checkpoint state * refactor: clarify deployment watch lifecycle boundaries * refactor(cdc): assert settings and simplify sync comments * use one callback for CDC changes and checkpoints Amp-Thread-ID: https://ampcode.com/threads/T-01a0a03f-4343-73ad-b2ee-b836285de946 Co-authored-by: Amp <amp@ampcode.com> * handle checkpoints inside the Krane stream watcher * make CDC clients own rules and resume progress * simplify CDC clients to in-memory watch progress * separate CDC structs and their methods by file * Give each CDC client its own Vitess connection Amp-Thread-ID: https://ampcode.com/threads/T-01a0a03f-4343-73ad-b2ee-b836285de946 Co-authored-by: Amp <amp@ampcode.com> * refactor(cdc): simplify watcher API and verify deletion events Amp-Thread-ID: https://ampcode.com/threads/T-01a0a03f-4343-73ad-b2ee-b836285de946 Co-authored-by: Amp <amp@ampcode.com> * fix(deps): update Vitess to v0.23.4 Amp-Thread-ID: https://ampcode.com/threads/T-01a0a03f-4343-73ad-b2ee-b836285de946 Co-authored-by: Amp <amp@ampcode.com> * fix(krane): restart snapshot after repeated stream failures Amp-Thread-ID: https://ampcode.com/threads/T-01a0a03f-4343-73ad-b2ee-b836285de946 Co-authored-by: Amp <amp@ampcode.com> --------- Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 66f1588 | |
|---|---|---|
| Author: | Andreas Thomas | |
fix(krane): order deployment revisions Amp-Thread-ID: https://ampcode.com/threads/T-01a0c865-b733-7580-98c6-ce63f00f7cbc Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 8b3676b | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | ogzhanolguncu | |
docs(worker): Build never resolves an image buildImage fails terminally on a non-git source and Deploy only calls Build when the source needs building, which is what the handler's own doc comment says.
| Commit: | 3385c04 | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | ogzhanolguncu | |
fix(worker): bound the build backend and correct the cancel docs Build holds the workspace's Restate build slot for as long as its restate.Run is running, and the Depot backend that production uses had no deadline of its own, so a backend that never answers stopped every other deployment in the workspace. withBuildkit now derives a 30 minute context that both backends and the build closure take. It is a child of the run context, so Restate cancellation still propagates, and it equals buildImageRetryCeiling so an attempt that hits it is not retried. The cancel comments, the package doc, and the product and engineering pages claimed a cancel frees the slot and that a CLI image deploy skips the queue. A running Build keeps its slot until its image build returns, and Deploy calls Build for every source type. Also move Build's limit key check above the terminal status return, journal the queued step's end timestamp so a retry reports the queue wait it measured rather than the time since the row was written, and drop EndActiveDeploymentStepsWithError, which lost its last caller with buildslot.
| Commit: | 92db351 | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | ogzhanolguncu | |
chore(worker): trim wording and keep the build retry policy with the binds
| Commit: | ac9870f | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | ogzhanolguncu | |
chore: improve wording
| Commit: | 788259d | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | ogzhanolguncu | |
fix: wording
| Commit: | 8865b48 | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | ogzhanolguncu | |
feat(worker): add new build handler
| Commit: | 0c7fc83 | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | GitHub | |
chore: improve wording
| Commit: | 198bfec | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | GitHub | |
feat(worker): add new build handler
| Commit: | c7b1495 | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | GitHub | |
fix(worker): bound the build backend and correct the cancel docs Build holds the workspace's Restate build slot for as long as its restate.Run is running, and the Depot backend that production uses had no deadline of its own, so a backend that never answers stopped every other deployment in the workspace. withBuildkit now derives a 30 minute context that both backends and the build closure take. It is a child of the run context, so Restate cancellation still propagates, and it equals buildImageRetryCeiling so an attempt that hits it is not retried. The cancel comments, the package doc, and the product and engineering pages claimed a cancel frees the slot and that a CLI image deploy skips the queue. A running Build keeps its slot until its image build returns, and Deploy calls Build for every source type. Also move Build's limit key check above the terminal status return, journal the queued step's end timestamp so a retry reports the queue wait it measured rather than the time since the row was written, and drop EndActiveDeploymentStepsWithError, which lost its last caller with buildslot.
| Commit: | 87456f1 | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | GitHub | |
chore(worker): trim wording and keep the build retry policy with the binds
| Commit: | b7052c0 | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | GitHub | |
fix: wording
| Commit: | b1da2fe | |
|---|---|---|
| Author: | Oz | |
| Committer: | GitHub | |
ctrl(worker): add cron to sync builds_concurrent_max to restate (#7452) * feat(ctrl): add new cron for syncing build limits to restate * chore: fmt * refactor: move scope key to shared * chore: change cron interval * chore: fix wording * chore: drop useless wording * chore: improve wording
| Commit: | 95afa85 | |
|---|---|---|
| Author: | Andreas Thomas | |
Merge remote-tracking branch 'origin/main' into chronark/vstream-deployment-sync Amp-Thread-ID: https://ampcode.com/threads/T-01a0a03f-4343-73ad-b2ee-b836285de946 Co-authored-by: Amp <amp@ampcode.com> # Conflicts: # docs/engineering/architecture/services/control-plane/worker/workflows/deployments.mdx # svc/ctrl/api/run.go # svc/ctrl/services/cluster/service.go
| Commit: | b6c2074 | |
|---|---|---|
| Author: | Andreas Thomas | |
refactor(cdc): simplify watcher API and verify deletion events Amp-Thread-ID: https://ampcode.com/threads/T-01a0a03f-4343-73ad-b2ee-b836285de946 Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 75b2c9c | |
|---|---|---|
| Author: | Andreas Thomas | |
refactor(cdc): assert settings and simplify sync comments
| Commit: | 1672968 | |
|---|---|---|
| Author: | Oz | |
| Committer: | GitHub | |
refactor: clean up trigger proto (#7347)
| Commit: | 15bffad | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | ogzhanolguncu | |
refactor: clean up trigger proto
| Commit: | dff3425 | |
|---|---|---|
| Author: | Oz | |
| Committer: | GitHub | |
refactor: remove deployVO for good (#7344) * refactor: remove deployVO for good * chore: fix stale comments * fix: build issue
| Commit: | 11160af | |
|---|---|---|
| Author: | Andreas Thomas | |
feat(ctrl): stream deployment changes through Vitess Amp-Thread-ID: https://ampcode.com/threads/T-01a0a03f-4343-73ad-b2ee-b836285de946 Co-authored-by: Amp <amp@ampcode.com>
| Commit: | c2b8cc4 | |
|---|---|---|
| Author: | Oz | |
| Committer: | GitHub | |
refactor(worker): deployVO to deployWorkflow (#7343) * refactor: move VO to workflow * refactor: deployVO to deployWorkflow * chore: tidy comments * chore: comment improvement
| Commit: | bba67d6 | |
|---|---|---|
| Author: | Oz | |
| Committer: | GitHub | |
refactor: remove dead deploy VO code (#7342)
| Commit: | 1dbd317 | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | ogzhanolguncu | |
refactor: clean up trigger proto
| Commit: | 6906090 | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | ogzhanolguncu | |
refactor: remove dead deploy VO code
| Commit: | 9f2ab4a | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | ogzhanolguncu | |
chore: fix stale comments
| Commit: | 48e7481 | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | ogzhanolguncu | |
chore: tidy comments
| Commit: | 68f6ccd | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | ogzhanolguncu | |
refactor: remove deployVO for good
| Commit: | a349cdc | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | ogzhanolguncu | |
refactor: deployVO to deployWorkflow
| Commit: | b3e3d7f | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | ogzhanolguncu | |
refactor: move VO to workflow
| Commit: | 159b227 | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | ogzhanolguncu | |
chore: fix stale comments
| Commit: | 91928b5 | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | ogzhanolguncu | |
chore: tidy comments
| Commit: | 9e0fc5c | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | ogzhanolguncu | |
refactor: remove deployVO for good
| Commit: | 6bf33cb | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | ogzhanolguncu | |
refactor: deployVO to deployWorkflow
| Commit: | 993fd5f | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | ogzhanolguncu | |
refactor: move VO to workflow
| Commit: | 3be3551 | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | ogzhanolguncu | |
refactor: remove dead deploy VO code
| Commit: | 474be17 | |
|---|---|---|
| Author: | Oz | |
| Committer: | GitHub | |
refactor: move wake and stop to deploymentVO (#7309) * refactor: move wake and stop to deploymentVO * chore: drop redundant comments * chore: move definiton to run
| Commit: | 4b05db0 | |
|---|---|---|
| Author: | Oz | |
| Committer: | GitHub | |
chore: cleanup old deployVO (#7300)
| Commit: | dcd7c5c | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
ctrl: process durable failure events for opt-in apps Poll the OOM and crash-loop inbox once per minute with bounded pages, fanout, and retries. Commit alert updates and handled identities in one MySQL transaction, then reconcile the existing per-group Restate object. Fresh evidence resets recovery without advancing five-minute ledgers. Keep collection workspace-gated and the local caller suspended. Retain processed identities for the beta so delayed replays cannot reopen alerts. Real MySQL/Restate tests cover killed invocations, delayed and duplicate events, topology suppression, recovery, capped sweeps, and 100 affected apps with 100,000 older processed identities. Amp-Thread-ID: https://ampcode.com/threads/T-01a065c7-a7fe-746a-a6f4-54a13271f6af Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 32f334c | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
fix(ctrl): chase missed anomaly shard windows Amp-Thread-ID: https://ampcode.com/threads/T-01a065e3-7fa4-755e-9b34-f6dd5de842ef Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 49f9c47 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
fix(ctrl): serialize deploy anomaly shard handoff Amp-Thread-ID: https://ampcode.com/threads/T-01a065e3-7fa4-755e-9b34-f6dd5de842ef Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 04f5256 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
fix(ctrl): derive anomaly lifetime from app creation Amp-Thread-ID: https://ampcode.com/threads/T-01a065e3-7fa4-755e-9b34-f6dd5de842ef Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 335f7d1 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
fix(ctrl): delay anomaly evaluation for ingest settling Amp-Thread-ID: https://ampcode.com/threads/T-01a065e3-7fa4-755e-9b34-f6dd5de842ef Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 2c35df3 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
fix(ctrl): suppress request drops for stopped topology Amp-Thread-ID: https://ampcode.com/threads/T-01a065e3-7fa4-755e-9b34-f6dd5de842ef Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 51b18c0 | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
feat(ctrl): shard deploy anomaly evaluation Fan each closed window into stable workspace partitions, batch metadata reads, and keep incomplete telemetry from mutating alert state. Persist open, touch, and resolve transitions without sending notifications; a separate consumer will notify later. Amp-Thread-ID: https://ampcode.com/threads/T-01a065e3-7fa4-755e-9b34-f6dd5de842ef Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 70367ba | |
|---|---|---|
| Author: | Flo | |
| Committer: | Flo | |
feat(ctrl): refine deploy anomaly signals and contracts Use robust traffic-drop detection, error-rate thresholds, telemetry watermarks, and instance-averaged memory utilization. Add the SQLC and protobuf contracts needed by the production Restate worker, including workspace notification muting. Amp-Thread-ID: https://ampcode.com/threads/T-01a065e3-7fa4-755e-9b34-f6dd5de842ef Co-authored-by: Amp <amp@ampcode.com>
| Commit: | 655edde | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | GitHub | |
chore: fix stale comments
| Commit: | d74edf3 | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | GitHub | |
refactor: remove deployVO for good
| Commit: | 325bdc0 | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | GitHub | |
chore: tidy comments
| Commit: | 4dc72d7 | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | GitHub | |
refactor: deployVO to deployWorkflow
| Commit: | a51f0a0 | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | GitHub | |
refactor: remove dead deploy VO code
| Commit: | addaf57 | |
|---|---|---|
| Author: | ogzhanolguncu | |
| Committer: | GitHub | |
refactor: move VO to workflow