These commits are when the Protocol Buffers files have changed: (only the last 100 relevant commits are shown)
| Commit: | f39485a | |
|---|---|---|
| Author: | buarki | |
chore(proto): format the three files buf reports on this branch Left logs.proto and service.proto alone: they are unformatted on main too and its CI is green, so the format check excludes them the same way buf.yaml excludes them from lint and breaking.
| Commit: | 5db9b87 | |
|---|---|---|
| Author: | buarki | |
fix(proto): carry the proto lint fixes from main buf lint fails on every PR against this branch: the runspace bridge protos miss field and oneof comments, and its two stream envelopes trip the RPC naming rule. All three fixes already exist on main, so they are taken from there rather than written again, generated code included since proto comments become Go doc comments.
| Commit: | 7b546c3 | |
|---|---|---|
| Author: | Vladislav Sukhin | |
| Committer: | GitHub | |
feat(proto): add the execution cache grants to the agent API (#8371) Two calls, both answering with a presigned URL: one asks whether an entry for a key exists and grants a read, the other grants a write for a new entry. The agent never computes an object name - it only follows a URL the control plane signed - which is what keeps the storage layout a server-side decision. A miss is a normal response rather than an error, so a cold cache costs the agent one call and no retry. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
| Commit: | 9ef3d7f | |
|---|---|---|
| Author: | Vladislav Sukhin | |
Merge remote-tracking branch 'origin/main' into vsukhin/feature/step-dependency-cache
| Commit: | 15ef92e | |
|---|---|---|
| Author: | Vladislav Sukhin | |
| Committer: | GitHub | |
feat(proto): add the execution cache grants to the agent API Two calls, both answering with a presigned URL: one asks whether an entry for a key exists and grants a read, the other grants a write for a new entry. The agent never computes an object name - it only follows a URL the control plane signed - which is what keeps the storage layout a server-side decision. A miss is a normal response rather than an error, so a cold cache costs the agent one call and no retry. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
| Commit: | 536c5f8 | |
|---|---|---|
| Author: | Vladislav Sukhin | |
| Committer: | Vladislav Sukhin | |
feat(proto): add the execution cache grants to the agent API Two calls, both answering with a presigned URL: one asks whether an entry for a key exists and grants a read, the other grants a write for a new entry. The agent never computes an object name - it only follows a URL the control plane signed - which is what keeps the storage layout a server-side decision. A miss is a normal response rather than an error, so a cold cache costs the agent one call and no retry. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
| Commit: | afc402d | |
|---|---|---|
| Author: | Vladislav Sukhin | |
| Committer: | GitHub | |
feat(proto): carry the rerun base and the derived lineage on the wire (#8296) Part of the execution-lineage stack, split for review. The layers below this one are its dependencies; the tip of the stack is identical to vsukhin/feature/execution-lineage. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
| Commit: | 58c3d36 | |
|---|---|---|
| Author: | Vladislav Sukhin | |
| Committer: | GitHub | |
feat(proto): add the execution cache grants to the agent API Two calls, both answering with a presigned URL: one asks whether an entry for a key exists and grants a read, the other grants a write for a new entry. The agent never computes an object name - it only follows a URL the control plane signed - which is what keeps the storage layout a server-side decision. A miss is a normal response rather than an error, so a cold cache costs the agent one call and no retry. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
| Commit: | a1f59d2 | |
|---|---|---|
| Author: | Vladislav Sukhin | |
Merge remote-tracking branch 'origin/main' into vsukhin/feature/step-dependency-cache # Conflicts: # go.mod # pkg/cloud/service.pb.go
| Commit: | acbea4b | |
|---|---|---|
| Author: | Vladislav Sukhin | |
feat(proto): carry the rerun base and the derived lineage on the wire Part of the execution-lineage stack, split for review. The layers below this one are its dependencies; the tip of the stack is identical to vsukhin/feature/execution-lineage. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
| Commit: | 3f249a4 | |
|---|---|---|
| Author: | Vladislav Sukhin | |
Merge branch 'main' into vsukhin/feature/execution-lineage pkg/cloud/service.pb.go conflicted inside the serialized descriptor bytes, where both sides had added a field. proto/service.proto merged cleanly, so the file was regenerated from it rather than hand-resolved - buf.gen.old.yaml pins its plugins to remote versions, so that output is the same one CI produces. Checked that both sides survived: our base_execution_id (field 13 on ScheduleRequest) and main's AGENT_CAPABILITY_EXECUTION are both in the regenerated descriptor. Regenerating also rewrote line endings in pkg/logs/pb, which has no content change, so those were restored. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
| Commit: | 57d4e6f | |
|---|---|---|
| Author: | Vladislav Sukhin | |
fix: make the index order legacy chain roots first, and say what lineage promises Three review findings, all places where a comment and the code it describes had drifted apart. The chain index ordered by the bare lineage_attempt. A legacy row carries NULL there and means attempt 1, so an ASC scan put the chain's original after every rerun of it - NULLS LAST - and an ORDER BY that corrected for that with COALESCE would no longer match the index, giving up the ordered scan the index exists for. Both index expressions now use COALESCE, so the index agrees with what EffectiveLineage() reports for the same row. No query had to change: nothing filters or orders on these columns yet, they only appear in select lists and the insert, so the index is still forward-looking. The DROP INDEX that heals an interrupted CONCURRENTLY build also rebuilds it for a database that ran the earlier version. The Postgres test still justified NULL-not-empty-string by a partial index, which this index has not been since it moved to COALESCE(lineage_root_id, id). The reason is better than the one it replaced: COALESCE substitutes only for NULL, so an empty string in lineage_root_id would survive it and drop the row out of its own chain, and an empty lineage_base_id would make every original run answer the "is a rerun" predicate. ExecutionStart.lineage was documented as present on every execution. It is not: lineageProtoOf returns nil for an execution recorded before lineage existed, and nothing backfills those. The comment now says it may be absent and that a consumer applies the original-run defaults - which is what lineageConfigFromProto and buildLineage already do, so this documents the behaviour rather than changing it. The generated comment was applied by hand after checking it reproduces buf's output byte for byte; running the generator here also reflows doc comments in six unrelated files, which is a local protoc-gen-go version difference and was left out. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
| Commit: | 0db4c1c | |
|---|---|---|
| Author: | Vladislav Sukhin | |
Merge branch 'main' into vsukhin/feature/step-dependency-cache
| Commit: | 528e09e | |
|---|---|---|
| Author: | Tristan Rasmussen | |
| Committer: | GitHub | |
refactor: add new `execution` capability to replace `runner` (#8227) * refactor: add execution capability to agents * chore: remove create/install runner commands * refactor: remove remaining references to runner capability
| Commit: | 7e4f1ec | |
|---|---|---|
| Author: | Vladislav Sukhin | |
Merge branch 'main' into vsukhin/feature/execution-lineage
| Commit: | 6b9e969 | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | Dejan Zele Pejchev | |
docs: call the stop and start reasons codes, not tokens
| Commit: | 0f5bf0e | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
docs: call the stop and start reasons codes, not tokens
| Commit: | cae2761 | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | Dejan Zele Pejchev | |
feat(runner): carry the stop actor on the execution state transition
| Commit: | 655f33f | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | Dejan Zele Pejchev | |
feat(executionworker): carry the stop reason as a token with words
| Commit: | eb19f79 | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
feat(executionworker): carry the stop reason as a token with words
| Commit: | ded2efa | |
|---|---|---|
| Author: | Vladislav Sukhin | |
feat: record what an execution is a rerun of pkg/executiondata lets a workflow read another execution's outputs, reports and artifacts, all addressed through a reference. It had no way to say "the execution I am a rerun of", so a rerun could not read the run it descends from without being handed the id as configuration - and the workflow is written once and rerun many times, so the author cannot know it in advance. The base id was not recorded anywhere usable either. Reruns carried it only as runningContext.actor.executionReference: free text on the user actor, overloaded across actor kinds, invisible to the pod, and carrying no chain root or attempt number. Add the reserved "rerun" reference, and the lineage that makes it resolve. TestWorkflowExecutionLineage{BaseId, RootId, Attempt} is recorded on every execution and threaded from ScheduleRequest.base_execution_id through the enqueuer, the execution record, ExecutionStart and ExecutionConfig to the pod, where RerunExecutionId() feeds the resolver. execution("rerun").outputs, .reports and read_artifact("rerun", ...) all follow from it. Three decisions worth stating: An original run is its own root at attempt 1 rather than having empty lineage, so "everything in chain R" is the single predicate lineage_root_id = R and includes the original. EffectiveLineage() synthesizes that for rows written before the columns existed, so nothing has to be backfilled and both repositories agree on what an unrecorded lineage means. Loading the base to derive the root and attempt is also the environment check service.proto asks for: PostgresRepository.Get is scoped to the organization and environment, so a base belonging to another one comes back not-found and the request is refused rather than silently honoured. Only the base travels on the wire - a caller able to assert a root or an attempt could forge a chain. Three scalar columns rather than JSONB, because these are queried: a composite btree gives the ordered range scan a GIN containment index cannot, and it sidesteps the nil-pointer-through-interface trap that storing an absent pointer as JSONB "null" would reintroduce. Two traps are pinned by tests: an original run has an empty BaseId and must resolve to no rerun rather than to its own RootId, and a lineage recorded on every execution must not be mistaken for a narrowing signal. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| Commit: | 150a525 | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | GitHub | |
feat(runner): apply the reason the control plane sends with an abort or cancel (#8225)
| Commit: | 8cdf4d8 | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | GitHub | |
feat(runner): send the reason and message with a declined execution (#8224)
| Commit: | 1920130 | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | Dejan Zele Pejchev | |
feat(runner): send the reason and message with a declined execution
| Commit: | c463abe | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | Dejan Zele Pejchev | |
feat(runner): apply the reason the control plane sends with an abort or cancel
| Commit: | 1374b90 | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | Dejan Zele Pejchev | |
feat(runner): apply the reason the control plane sends with an abort or cancel
| Commit: | 1c7d53a | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | Dejan Zele Pejchev | |
feat(runner): send the reason and message with a declined execution
| Commit: | c1d2530 | |
|---|---|---|
| Author: | Tristan Rasmussen | |
refactor: add execution capability to agents
| Commit: | 2b75360 | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | Dejan Zele Pejchev | |
feat(runner): apply the reason the control plane sends with an abort or cancel
| Commit: | 54e3d40 | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | Dejan Zele Pejchev | |
feat(runner): send the reason and message with a declined execution
| Commit: | da59e39 | |
|---|---|---|
| Author: | Vladislav Sukhin | |
feat: send the parse of a JUnit report, not just the XML The toolkit now parses a report before uploading it and sends counts and the non-passing test cases alongside the raw bytes, resolving the TODO that sat on AppendExecutionReportRequest.report. The agent already parses this report in the pod - muting decides the step's verdict and only the pod can still change it - so re-parsing server-side was duplicated work on a file that may be megabytes of XML. `report` is still sent, because a control plane predating the parsed fields reads it. Report.Digest() caps what travels: 2000 failures and 256KB of messages. An execution with more failures than that is not a "re-run the failures" scenario, and the report file it accompanies always holds the whole story. When a report describes more test cases than it names, the declared counters are believed over the named ones and the digest is marked truncated - the same reading the verdict takes, because a report naming one passing case while declaring four failures is not a passing report. The digest deliberately carries nothing about muting. A report says which test cases failed; whether a failure was *tolerated* is the step's verdict, decided in a different container from the one that uploads artifacts. The step result carries that, and the two join on step reference. So TestReportSummary.muted, .unexpected and .tolerated are left unset here - they need either that join or a sidecar, which is a decision rather than an oversight. A report that cannot be parsed is still uploaded unparsed rather than dropped: the artifact is the record, and an older control plane reads the bytes regardless. Also removes isJUnitReport, which duplicated testresults.Sniff, and reworded the `report` field's proto comment - it began "Deprecated:", which protoc turned into a Go deprecation marker, so populating a field every client MUST still send tripped staticcheck. The generated marker stays stale until the protobuf is regenerated, hence the nolint at the call site. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| Commit: | 99465be | |
|---|---|---|
| Author: | Vladislav Sukhin | |
feat: declare the wire contract for rerunning failed test cases Proto only; nothing consumes these yet. The generated code follows from `make generate-protobuf`, and the Go wiring after that, the same way the OpenAPI models landed. Three additions, all backwards compatible - new fields and new messages, so an older peer ignores what it does not know. RerunPolicy narrows an execution to specific test cases. It is defined twice on purpose: `cloud.RerunPolicy` on ScheduleRequest for the legacy path, and `testkube.testworkflow.execution.v1.RerunPolicy` on ExecutionStart for the typed one. The two start paths do not share types, and importing the legacy `cloud` package into the v1 API to save a duplicate would be the worse trade. The selection is deliberately *not* expanded on the wire. Only the reference travels; the pod resolves it against the report the referenced execution produced. A suite of ten thousand test cases would not survive being carried as a list through scheduling, and the pod is where the report already is. `test_cases` exists for small explicit selections and is capped. ExecutionRunningContext gains execution_reference, which fixes a gap that predates this work: runningContextFromProto can only carry across what the message declares, so on the typed start path execution.runningContext.actor.executionReference already resolves to empty inside the pod, while the legacy path fills it from the execution record. The same workflow behaves differently depending on which path started it. AppendExecutionReportRequest gains kind, summary, failures and failures_truncated, so the agent sends the parse rather than the raw XML - resolving the TODO that has sat on the `report` field. The agent parses in the pod regardless, because muting decides the step's verdict and only the pod can still change it, so re-parsing server-side is duplicated work. `report` stays for a control plane that predates the new fields. Two things the Control Plane must honour, both security-relevant: - RerunPolicy.execution_id must be confirmed to belong to the caller's environment. Otherwise it is a way to read another environment's artifacts. - A request carrying a RerunPolicy for a workflow with no step declaring `testCases.select` should be rejected, not silently run in full. Testkube does not know any runner's filter flag by design, so without such a step there is nothing for the policy to act on. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| Commit: | e8dfe68 | |
|---|---|---|
| Author: | Vladislav Sukhin | |
| Committer: | Vladislav Sukhin | |
feat: make a stored cache entry immutable with a conditional write Closes the race review found. Looking a key up and granting the URL are two steps, so the check alone could never stop two executions saving the same key from both being told it was absent and both overwriting. The write itself is now conditional, which does stop it. PresignCreateFileToBucket signs a PUT that the store applies only while the object is absent. MinIO's PresignHeader covers the condition header with the signature, so an upload that drops it is rejected rather than quietly becoming a plain overwrite - the condition cannot be negotiated away by the client. The headers travel back in the response rather than being assumed by the agent, because object stores spell the condition differently: If-None-Match on S3 and MinIO, a generation precondition on GCS. The commercial control plane can pick its own and the agent needs no knowledge of which. Losing the race is now an outcome rather than a corruption. The store answers the second upload with 412, or 409 for a concurrent write on one key, and the agent reports that another execution stored the key first and exits zero - which is correct, because both executions reached the same content-derived key, so the entry that won is the one this step would have written. It is deliberately not retried: a refused condition refuses again. The existence check is kept in front of the grant. It no longer carries the guarantee, but it still avoids packing and uploading an archive that would only be refused. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
| Commit: | 644ad86 | |
|---|---|---|
| Author: | Vladislav Sukhin | |
fix: stop claiming a saved cache entry is immutable Review is right and the claim was wrong. Stat and the presigned PUT are two steps, and a presigned PUT is an ordinary overwrite, so two executions saving the same scoped key can both find nothing, both be granted a URL, and the later upload replaces the earlier entry. "The first writer of a key wins" was not true, and it was asserted in six places here and in the commit message that introduced the handler. What the check actually provides is deduplication, and the comments now say so, along with why the cost is bounded: racing writers reached the same content-derived key, so they store equivalent trees, and a save only follows a miss, so once an entry exists later runs hit it and never write. The residual exposure is that a workflow able to write a scope can replace an entry there by racing, on top of being able to write one at all - which is the caveat scope: environment already carries. Immutability was an extra mitigation, not the boundary; the scope is the boundary. Making it a real guarantee needs the write itself to be conditional - a presigned PUT carrying If-None-Match, or a reservation the grant takes atomically - and neither storage interface can express that today. A third option avoids the protocol change by never overwriting: store each write under its own name in a per-key folder and serve the earliest. All three are design commitments rather than a comment fix, so they are written up in the plan instead of being chosen here. The generated trailing comment in pkg/cloud is updated to match the proto so verify-protobuf-generated stays green; it is worth confirming against a real regeneration. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
| Commit: | 3ed1262 | |
|---|---|---|
| Author: | Vladislav Sukhin | |
| Committer: | GitHub | |
Potential fix for pull request finding Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
| Commit: | 2e7a1cc | |
|---|---|---|
| Author: | Vladislav Sukhin | |
feat: add a dependency cache to test workflow steps Test suites re-download node_modules, wheels, Maven and Gradle artifacts and Go module caches on every run, because there is no cache concept in the CRD. The workarounds are a PVC with hand-rolled copy steps, which nothing invalidates when a lockfile changes, or an artifact fetched back by execution id, which cannot express "the most recent entry for this lockfile". A step now declares: cache: key: 'npm-{{ hash_files("package-lock.json") }}' restoreKeys: ['npm-'] paths: ['node_modules'] scope: workflow Entries are keyed, scoped per workflow by default, and evicted by a prefix-filtered bucket lifecycle rule. A cache is only an optimization, so every runtime failure - a miss, an unreachable control plane, a control plane too old to advertise the capability, a corrupt archive, a refused upload - degrades to an uncached run. Only authoring mistakes fail, and they fail at bundle time, before a pod exists. Three consequences of the surrounding code shaped the design: - The specification travels base64-encoded in one argument. testworkflow-init resolves every container argument with FinalizerFail and exits the step on failure, so a key holding hash_files() over an absent lockfile would kill the step rather than miss the cache. This is why the execute and parallel steps already encode their specs. - The two stages hand the resolved key over through a state file on the shared volume rather than each computing it. They are separate containers, and an install may rewrite the very lockfile the key hashes, so a recomputed key could store the entry where nothing will later search for it. - Each stage is its own container, and containers share volumes but not their root filesystems, so a cached path outside every volume is restored where the container running the install cannot see it. Paths are therefore mounted automatically, and mount: false on an uncovered path is refused instead of silently doing nothing. The save stage sets no condition, inheriting "passed", deliberately unlike the artifacts stage: publishing a failed install under a content-hash key would poison every later run with no way for a user to invalidate it. Keys are attacker-influenced text, so they are percent-encoded into a single path segment - prefix-preserving, which is what lets a restore key be answered by a plain prefix query - and the scope prefix is derived from the execution rather than the request, so one workflow cannot address another's entries. Also hardens the tar unpacker, which a cache makes reachable from outside the producing execution for the first time. It validated entry names but never symlink targets, so an archive of a symlink to / followed by a regular file beneath it wrote outside the destination and returned no error. Extraction now goes through os.Root, which refuses the whole class, and gained entry-count and decompressed-size limits and a permission-bit mask. Relative symlinks still work, which pnpm and yarn workspaces need. The two RPCs are declared but not yet implemented. Server embeds UnimplementedTestKubeCloudAPIServer, so they answer Unimplemented, which the degrade path already treats as a miss. Wiring the transport is a change to newCacheRepository plus the handlers, both built on the key derivation and match policy in pkg/executioncache. Generated CRDs, OpenAPI models and protobuf still need regenerating; the deepcopy and models here were written by hand in generator style so that the new field is not silently dropped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
| Commit: | 80c3482 | |
|---|---|---|
| Author: | Oscar Reyes | |
| Committer: | GitHub | |
chore(ci/cd): Move lint and tests to Testkube's Advanced Github Integration (#8183) * fix(proto): Comment fields so buf COMMENTS lint passes Quality Loop actually runs proto lint; GH never did because lint.yaml was push-only and buf-action defaults lint to pull_request. Ignore Connect RPC request/response names — they are bidi envelopes, not a wire rename. Co-authored-by: Cursor <cursoragent@cursor.com> * test: Fire inventory notifier burst without sleeps Sleeping 5ms between 10 events raced the 50ms debounce on a contended runner and produced a mid-burst extra push. Co-authored-by: Cursor <cursoragent@cursor.com> * ci: Remove GH Actions lint and test workflows Lint, unit/integration tests, and protobuf/CRD verify now run as Quality Loop TestWorkflows. Semantic PR titles and bake/release stay on GitHub Actions. Co-authored-by: Cursor <cursoragent@cursor.com> * ci: Remove GH Actions PR title lint Conventional Commits title checks move to Testkube with the rest of lint/test. Bake and release stay on GitHub Actions. Co-authored-by: Cursor <cursoragent@cursor.com> * chore: Regenerate protobuf after proto comment updates verify-protobuf-generated diffs these files after buf generate; the COMMENTS lint comments need to land in the generated Go. Co-authored-by: Cursor <cursoragent@cursor.com> * style(proto): Apply buf format Quality Loop runs buf format; GH lint.yaml never did (push-only vs buf-action PR default). Whitespace and a stray semicolon only. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: Drop deprecated Dial and ListWatch funcs Quality Loop golangci-lint runs the full tree; GH used only-new-issues so these SA1019s never failed CI. Switch Transport to DialContext and informers to ListWithContextFunc / WatchFuncWithContext. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: Use ListWithContext in toolkit parallel test informer Same SA1019 as the operator informers: ListFunc/WatchFunc are deprecated in client-go v0.36. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: Use ephemeral port in init integration tests Quality Loop already binds 60434 in the same network namespace, so these tests cannot reuse the production control port. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com>
| Commit: | 1c4eb42 | |
|---|---|---|
| Author: | Oscar Reyes | |
style(proto): Apply buf format Quality Loop runs buf format; GH lint.yaml never did (push-only vs buf-action PR default). Whitespace and a stray semicolon only. Co-authored-by: Cursor <cursoragent@cursor.com>
| Commit: | d6bc8fb | |
|---|---|---|
| Author: | Oscar Reyes | |
fix(proto): Comment fields so buf COMMENTS lint passes Quality Loop actually runs proto lint; GH never did because lint.yaml was push-only and buf-action defaults lint to pull_request. Ignore Connect RPC request/response names — they are bidi envelopes, not a wire rename. Co-authored-by: Cursor <cursoragent@cursor.com>
| Commit: | 3dd49e8 | |
|---|---|---|
| Author: | Vladislav Sukhin | |
| Committer: | GitHub | |
feat: [TKC-6403] exchange data between test workflows run as a suite (#8066) * feat: exchange data between test workflows run as a suite Workflows composed into a suite through `execute.workflows` could not pass anything back to the parent: the toolkit polled the child execution and threw away everything but the status. Values and files had to travel through external state. A child now publishes with the same mechanism steps already use - writing to /testkube/outputs - and the parent reads it back with the execution() expression: execute: workflows: - name: producer as: p fetch: - paths: ['results/**'] to: /data/from-producer ... shell: echo '{{ execution("p").outputs.token }}' The same function resolves "parent" from the execution ancestry, so a child can read what scheduled it, and sibling exchange falls out of feeding one child's output into the next child's config. - step outputs are promoted to the execution record, so they cross the pod - pkg/executiondata owns the registry, the expression functions and the artifact transfer, modelled on the existing credential() machine - `as` gives an entry a stable reference; two entries claiming the same one is an error rather than an ambiguous winner - execute specs are finalized when their operation starts, not up-front, so a later entry can read an earlier one - read_artifact() returns small files inline (1 MiB cap); `fetch` writes larger payloads to disk Files need read access to artifact storage, which the control plane did not grant: ListExecutionArtifactsPresigned is new in proto/service.proto and implemented for OSS. The Enterprise control plane needs the same RPC before read_artifact() and fetch work there; values work everywhere today. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test: read a parent's artifact from a consumer workflow The suite fixture showed read_artifact() only from the parent reading its children. The child side is the direction that needs the "parent" reference and a control plane round-trip, so cover it too: the suite uploads a fixture file before scheduling anything, and the consumer reads it back. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: read a sibling test workflow's values and artifacts A workflow only registers the test workflows it ran itself, so a child of a suite could not reach another child: execution("p") failed with an unknown reference, and there was no way to name the sibling at all. Both directions already fall through to the control plane when the reference is an execution id, so the parent passing that id down is enough - what was missing was permission to use it, which is lifted in kubeshop/testkube-cloud-api@212a5f831's follow-up. execute: workflows: - name: consumer config: producerId: '{{ execution("p").id }}' # inside the consumer {{ execution(config.producerId).outputs.token }} {{ read_artifact(config.producerId, "results/summary.json") }} The unknown-reference error now says that anything the workflow did not run must be addressed by id, rather than only suggesting an earlier step - advice that was misleading inside a child, which never runs anything. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: aborted selectors with an alias, and corrupted sensitive outputs Two defects in the suite data exchange, both found in review. An entry combining `as` with a selector aborted the whole step. The reference was claimed once per matched workflow, so the second match reported a duplicate and failed the command before anything was scheduled. An aliased entry names one group regardless of how many workflows it matched - the instances were already told apart by execution("<alias>", index) at runtime, only the validation disagreed. The claim now happens once per entry, and every matched name is still claimed individually when there is no alias. A step output holding a sensitive value reached the parent mangled. The instruction that publishes outputs rides the log stream, which the runner obfuscates with ShowLastCharacters, so a token arrived as ****ue - a value that looks real, compares unequal, and gives no hint why. Such outputs are now withheld from the execution record with a warning, and stay usable within their own workflow through step.<id>.outputs. Withholding rather than bypassing the obfuscator is deliberate: publishing raw would put any credential a workflow writes to its outputs directory into the pod log, and no output is worth trading that for. Credentials belong in credential(). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Add remote endpoint to MCP server registry entry (#7744) * fix: capability Signed-off-by: Vladislav Sukhin <vladislav@kubeshop.io> * fix: testing Signed-off-by: Vladislav Sukhin <vladislav@kubeshop.io> * fix: fail loudly when a workflow reads a withheld sensitive output An output whose value holds a sensitive word cannot travel through the obfuscated log stream without being corrupted, so it is not published outside the workflow that produced it. Omitting it left the reader with nothing: a missing key resolves to an empty value, so a parent workflow silently configured a child with "" instead of the token it asked for. Publish a marker in its place. The real value still never leaves the workflow, but a reader now resolves something self-describing, and the two steps that would consume it - handing configuration to another workflow, and running a command - refuse to run with it instead of passing the marker on. Also hold sensitive words added at runtime to the same minimum length as the ones read from the environment. A one-character word matches nearly every line, which masked unrelated logs and withheld unrelated step outputs wholesale. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: state the contract for exchanging sensitive outputs Outputs cross the boundary between executions through the execution record, which they reach by being printed to the obfuscated log stream. An output holding a sensitive word therefore cannot be exchanged: it stays inside the workflow that produced it, and what leaves is a marker naming what was withheld. That was implemented but never written down, so it read as an unfinished feature rather than a decision. Say it where the feature is documented, with both reasons it is a decision - publishing the value would either corrupt it or leak it into a record anyone who can read the execution can read - and name the channels that can carry such a value instead. Pin it from the consumer side too: the test now asserts that a withheld output does not resolve to an empty value, and that the error names both the output to stop relying on and a channel that can carry it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: stop a withheld output from reaching the environment The guard covered command arguments only, so a withheld output assigned to a computed environment variable slipped through: UseEnv resolved the marker like any other value and installed it, and the tool received it as the value of the variable with nothing looking at it again. Gate every install instead, plain values as well as computed ones. A step that spawns workers resolves their specification itself, so the marker reaches a worker baked in as a literal rather than as something to compute - guarding only the computed branch would have missed it. Guard the parallel step at its own resolution too. Its workers were receiving the marker inside their specification, and failing before anything is spawned names the missing output once instead of once per worker. Two sites that resolve the same expressions are deliberately left alone: tarball patterns, where a marker yields an empty glob rather than a wrong value, and the retry condition, which is consumed as a boolean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep short credentials in the sensitive-word set Holding runtime-added words to minSensitiveWordLength was wrong. Those words are the values resolved credential() calls hand over, and that set is the only thing deciding whether a step output may be published. Production sets the minimum to four, so a one-to-three-character credential was dropped from the set, then masked nowhere in the logs and published verbatim into the execution record instead of being withheld. Restore the earlier behaviour and say why the asymmetry with the environment-derived words is deliberate. A short word does match aggressively, and that was the reason for the change - but over-matching costs noisy logs and withheld outputs, which is the safe direction to fail and is now loud rather than silent. Publishing a credential is not. The tests now pin the safe direction, end to end: a three-character credential stays in the set, and a short credential written to the outputs directory is withheld from the record. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: classify step outputs against every sensitive value The sensitive-word set skips values shorter than four characters, and output publication was classified against that same set. A short value from secret(), a sensitive configuration parameter, or one of the common sensitive variables was therefore promoted into the execution record verbatim, where execution().outputs reads it back from another workflow. The length minimum exists for the log stream: masking two characters rewrites nearly every line they appear in and hides almost nothing. That is a rendering decision, and it was doing double duty as a security classifier. Split the two - GetSensitiveWords keeps the minimum for masking, GetSensitiveValues answers with everything marked sensitive, and output publication asks the latter. Masking behaviour is unchanged; only classification gets stricter, so nothing that was hidden before becomes visible now. The cost of a short value matching is a withheld output, which is announced; the cost of missing it is a published secret. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: report an ambiguous execution reference instead of guessing An aliased entry stays addressable by the workflow it ran, which the reference validation does not reserve - it claims the alias only. An aliased selector covering workflow "a" and a separate unaliased entry running "a" therefore both answered to execution("a", 0), and Lookup returned whichever had been inserted first. Configuration built from it could read the outputs of the wrong child, with nothing to indicate it. Report it. Lookup now fails when a reference and position address more than one execution, naming both so the author can see which runs collided and say which was meant. Validating it at claim time instead would have been smaller, but it would also refuse to schedule two aliased entries running the same workflow - told apart by their aliases, which is a topology worth keeping. The workflow name is the only reference that becomes ambiguous there, and it now fails when it is used rather than when it is created. Addressing an aliased execution by its workflow name stays supported; it is only a collision on that name that is refused. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: follow the storage certificate setting when downloading artifacts Downloading an artifact talks to object storage directly, and it was the only storage path in the worker that verified the certificate. Storage in a self-hosted deployment commonly presents one signed by a private CA that the worker image has no reason to trust, so read_artifact() and fetch failed against the very storage the workflow had just uploaded to - uploading an artifact, and uploading logs, already skip verification. Take the client from the caller instead of hardcoding one, and build it from the same constant that configures the control plane client, so the two cannot disagree about a host they both talk to. A caller that passes nothing keeps verifying: a deployment that needs otherwise says so. The transport is a clone of the default rather than the default itself. Assigning to the shared one - as the log upload does - makes every other client in the process skip verification too. Tested against a TLS server with an untrusted certificate, which is the failure this fixes, in both directions: skipping verification reads it, verifying refuses it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: move the execution registry instruction out of the testworkflow family The dashboard groups output instructions by name, reading anything matching ^testworkflow(-.*)?$ as the status of a single child execution. The parent's registry instruction was named testworkflow-execution.<alias>, which matches - and it carries a list of executions rather than one, so every field the dashboard looked for came back undefined. It invented a nameless child stuck at Queued for each step that ran a workflow, next to the real ones. Name it executiondata.<alias> instead. Renaming here fixes every dashboard already deployed, where narrowing the pattern would only fix the next one, and the name is free to change while the feature is unreleased. The test now asserts the name stays out of both instruction families, since the reason for it lives in another repository. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: resolve an execution id regardless of its fan-out position Lookup filtered by index before it looked at any reference, so an id was found only when its position happened to match the index the caller asked for. Since an unspecified index means 0, execution("<id>") reached the first member of a fan-out group and nothing else. An id names one execution outright, while an alias or workflow name names a group whose members are told apart by index, so match ids first and without consulting the index. The two paths failed differently. execution("<id>") did not error - it fell through to the control plane, which answers from the execution record, and that record carries neither the alias nor the index. So a shard read its own position as 0 and its alias as empty, at the cost of a network call. 'fetch: from:' has no such fallback and failed outright. The new expression test wires a repository that fails the test if it is called at all, then asserts the index and alias come back - neither is recoverable from the record, so they can only come from the registry. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: resolve a fetch reference the way an expression resolves one 'fetch: from:' looked only in the local registry, so it refused any reference the workflow had not scheduled itself - including the sibling execution id a suite hands down as configuration, which read_artifact() accepted from the same workflow with the same id. The two ways of reaching another execution's artifacts disagreed about what a reference means. Share the resolution instead of teaching fetch a second copy of it. The registry-then-control-plane walk moves out of the expression machine into a Resolver both use, so a reference resolves identically in either place and 'from: parent' now works for the same reason execution("parent") does. Resolve takes a context, which the expression path cannot supply - an expression function has none - but fetch can, so a cancelled fetch now cancels the read behind it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Potential fix for pull request finding Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> --------- Signed-off-by: Vladislav Sukhin <vladislav@kubeshop.io> Co-authored-by: Omoniyi Omotoso <omoniyiomotoso@gmail.com>
| Commit: | 87467b7 | |
|---|---|---|
| Author: | Tristan Rasmussen | |
| Committer: | GitHub | |
chore: merge main into release/2-13 (#8164) * chore(deps): update dependency turbo to v2.10.9 (#8082) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix: generate zsh completion under the actual kubectl-testkube binary name (#8081) * fix: generate zsh completion under the actual kubectl-testkube binary name Cobra's built-in "completion zsh" command hardcodes RootCmd.Name() (the first word of Use, "testkube") in the generated script's #compdef header and function names. Since this CLI always ships and is invoked as "kubectl-testkube" (a kubectl plugin), the installed completion script never matched the real command name and zsh completion silently did nothing. Add a custom "completion" command that delegates bash/fish/powershell to Cobra's standard generators unchanged, but generates zsh under the actual binary name by temporarily swapping RootCmd.Use for the duration of that one call. Fixes #964 * docs: update AGENTS.md and ARCHITECTURE.md to document custom zsh completion command * fix(deps): update module github.com/ohler55/ojg to v1.28.4 (#8083) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update dependency @esbuild/linux-x64 to v0.28.2 (#8084) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * .dev webhooks - back to normal noticications (#8072) * fix(deps): update module github.com/google/go-containerregistry to v0.21.9 (#8074) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update docker docker tag to v29.7.2 (#8088) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module github.com/onsi/ginkgo/v2 to v2.32.1 (#8087) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module google.golang.org/protobuf to v1.36.12 (#8086) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module github.com/mark3labs/mcp-go to v0.58.0 (#8092) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module github.com/nats-io/nats.go to v1.53.1 (#8093) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module golang.org/x/crypto to v0.55.0 (#8094) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module golang.org/x/text to v0.41.0 (#8095) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update dependency @changesets/cli to v3 (#8096) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * feat: Tag Testkube-provisioned hosted runners in agent telemetry (#8085) Hosted trial runners send the same testkube_api_start and testkube_api_heartbeat payloads as runners users deploy in their own infrastructure, so they cannot be told apart in analytics. The control plane's hosted-runner markers (the cloudrunner deployment label and the tkcagent_hr_ name) only travel over gRPC at registration and never reach telemetry. Derive a hosted-runner capability tag from RUNNER_NAME, which the control plane already sets to the tkcagent_hr_ prefixed agent name. This needs no Helm chart or control plane change. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(webhooks): skip silent-webhook executions in previous-state lookup (#8091) * fix(webhooks): skip silent-webhook executions in previous-state lookup Silent executions with SilentMode.Webhooks=true were being counted as part of a workflow's finished history when computing become-* state transitions. A silent failed run followed by a non-silent passed run was emitting a spurious become-testworkflow-up webhook. Filter silent-webhook executions out of GetPreviousFinishedState in both Postgres (sqlc) and Mongo backends, matching the SilentMode.Webhooks predicate the webhook listener already uses at dispatch time. * fix(webhooks): also skip legacy DisableWebhooks executions Executions using the deprecated DisableWebhooks bool (silent-flat) were still leaking into GetPreviousFinishedState and could contaminate the become-* transition decision the same way silent-webhooks executions did. Add the equivalent predicate for the legacy field in both Postgres and Mongo, plus test coverage. * refactor(webhooks): make silent-webhook filter opt-in on GetPreviousFinishedState * fix(deps): update module github.com/gofiber/fiber/v2 to v2.52.15 (#8099) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module github.com/nats-io/nats-server/v2 to v2.14.5 (#8100) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update nats docker tag to v2.14.5 (#8101) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update golang docker tag to v1.26.6 (#8103) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix: support multiple kubeconfig files in KUBECONFIG env var (#8102) * fix: support multiple kubeconfig files in KUBECONFIG env var GetK8sClientConfig passed the whole KUBECONFIG value to BuildConfigFromFlags as a single file path, so a colon-separated list (semicolon on Windows) failed with a stat error like 'stat /a/config:/b/config: no such file or directory'. Split the value with filepath.SplitList and load it through clientcmd's loading rules, which merge the files with kubectl's own precedence semantics: the first file's current-context wins, and entries defined in later files remain reachable. Fixes #657 * fix: treat empty KUBECONFIG env var as unset A set-but-empty KUBECONFIG previously fell through to the default loading behavior; splitting the empty string produced no loading rules and errored instead. Skip the env branch when the value is empty so it falls back to ~/.kube/config or in-cluster config as before. * fix: replace App Engine logger with testkube logger in k8sclient (#8106) pkg/k8sclient imported google.golang.org/appengine/v2/log for a single error log in the port-forward helper, pulling the App Engine SDK into the CLI and agent for no benefit and diverging from the zap-based logger used across the codebase. Use pkg/log's DefaultLogger instead. Fixes #8105 * chore(deps): update postgres docker tag to v16.15 (#8104) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * feat(TKC-6504): rework init demo for new architecture (#8090) * feat(TKC-6504): rework init demo for new architecture * fix: make init demo re-install idempotent (reuse agent key) * chore(deps): update dependency microsoft.net.test.sdk to 18.9.0 (#8107) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update dependency turbo to v2.10.10 (#8110) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update dependency xunit.runner.visualstudio to v4 (#8113) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module github.com/minio/minio-go/v7 to v7.3.0 (#8114) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore: bump cloud-ui-e2e Playwright image to v1.62.1 (#8115) * fix(deps): update module github.com/adhocore/gronx to v1.20.3 (#8116) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module github.com/stretchr/testify to v1.12.0 (#8117) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * feat: add ability to rerun with latest workflow (#8097) * fix: make install.sh POSIX sh compatible and verify release checksums (#7917) * fix: make install.sh POSIX sh compatible and verify release checksums The documented install command pipes the script to `sh`, which is dash on Debian/Ubuntu, but the script relied on bash-only constructs ([[ ]], =~, set -o pipefail) and failed there. Also removes dead code paths (i386 assets are no longer published, the Windows uname case never matched, beta tag-name grep matched nothing since the v1.17 era), fails clearly when the version cannot be resolved instead of silently substituting "1", verifies the tarball against the release's checksums.txt, and downloads into a mktemp workdir with cleanup instead of the caller's directory. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: address review comments — add -L to API curls, exact-match checksum lookup Adds -L to the GitHub API curl calls for consistency with the tarball and checksums downloads, and replaces the grep regex lookup in checksums.txt with an exact awk field match since the tarball name contains dots that grep treats as wildcards. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> * feat(TKC-6504): Add warn message (#8118) * feat(TKC-6504): Add warn message * fix: lint * feat(proto): add GITINTEGRATION actor type for git-provider-triggered flows (#8108) * feat(proto): add GITINTEGRATION actor type for git-provider-triggered flows Adds a new value to the RunningContextType proto enum and to the TestWorkflowRunningContextActorType OpenAPI enum, plus the round-trip mapper cases between them. Downstream consumers (telemetry bucketer, CDEvents mapper) route the new actor into git-integration/event buckets. Kicks off the change needed on the cloud-api side to introduce a provider-agnostic "Git Integration" actor for Quality Loop executions (GitHub today, GitLab and Bitbucket without another proto bump). * fix: allow gitintegration actor in CRD schema and CLI validation * refactor: rename GITINTEGRATION actor to QUALITYLOOP for internal naming * refactor: keep proto RunningContextType_QUALITYLOOP but expose gitintegration to users Proto stays QUALITYLOOP internally per Ole's suggestion. OpenAPI actor type, CRD kubebuilder enum, CLI flag, and telemetry bucket surface as gitintegration so every customer-facing surface (REST API, kubectl testkube, helm charts, generated CRDs) speaks the provider-agnostic name. The mapper cross-translates between the two. * refactor: alias QUALITYLOOP for internal Go references, wire stays gitintegration Adds regen-safe alias files so Go code can reference testkube.QUALITYLOOP_TestWorkflowRunningContextActorType while the underlying wire value stays "gitintegration" for CLI, REST, CRDs, and helm charts. Consumers refactored to use the internal name; customer surfaces unchanged. * chore: fix goimports alignment * feat(controlplaneclient): let chained children inherit the parent's actor for sticky families (#8119) * feat(controlplaneclient): let chained children inherit the parent's actor for sticky families Today every chained execution is stamped as RunningContextType_EXECUTION, which the server maps to actor.type = testworkflow. That is the right default for regular composites but it does not survive a filter by the Quality Loop actor: the QL parent carries actor.type = gitintegration and the children it schedules end up as testworkflow, so filtering the Executions page by the parent's actor returns an empty list. This introduces a small sticky-actor family: when the parent's actor belongs to it (only QUALITYLOOP for now), the child inherits the parent's actor type instead of falling back to the EXECUTION default. Parent-chain walkers keep working because we extend the QUALITYLOOP mapper branch to populate actor.executionId / actor.executionPath from ParentExecutionIds the same way the EXECUTION branch does. Everything outside the sticky set (user-authored composites, cron/testtrigger/CR-scheduled runs) is byte-identical to today: the default injection path still stamps EXECUTION and the mapper still produces testworkflow. Telemetry buckets, Mixpanel dimensions, and downstream webhooks for non-QL flows are unaffected. * chore: reword sticky-actor comments to reference gitintegration instead of Quality Loop * refactor: move child-context decision onto the actor type via ChildRunningContextType method Reads better at the call site: instead of an external helper that returns (value, ok) and an if-branch to conditionally overwrite the default, the actor type answers directly what its chained children should carry. Injection funnel collapses to: Type: c.opts.ParentActorType.ChildRunningContextType() The default (RunningContextType_EXECUTION, mapped server-side to actor.type = testworkflow / the "Workflow" chip on the Executions page) lives inside the method along with the special cases, so extending it means editing one file next to the type itself. Tests moved from the controlplane client into the same package as the method for the same reason. * chore: drop verbose doc block on ChildRunningContextType * feat(runner): propagate parent RunningContext to child pods via GetExecutionWorkflow The runner builds ExecutionConfig for a scheduled pod from the ExecutionStart proto, which does not carry RunningContext. That leaves cfg.Execution.RunningContext nil in the toolkit; parentActorTypeFromRunningContext returns empty; the sticky family cannot see it's chained under a gitintegration parent and children fall back to actor.type=testworkflow. The runner already makes a follow-up GetExecutionWorkflow call for every ExecutionStart to fetch the enriched workflow. That response is documented as the vehicle for "any additional information that may be related" - the perfect place to attach the parent's actor identity without touching the ExecutionStart proto (which is used by many non-sticky paths). Adds a minimal ExecutionRunningContext projection (actor type + actor name) to GetExecutionWorkflowResponse, populates it server-side from the stored TWE, and reconstructs a TestWorkflowRunningContext on the runner side that parentActorTypeFromRunningContext can read unchanged. When the field is nil (older control planes, non-actor executions), behaviour matches today. * chore(deps): update dependency turbo to v2.10.11 (#8122) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update mongodb docker tag to v8.3.8 (#8123) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update postgresql docker tag to v18.6 (#8124) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update docker.io/kubeshop/testkube-postgres docker tag to v18.6 (#8127) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update docker.io/kubeshop/bitnami-mongodb docker tag to v8.3.8 (#8126) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update natsio/nats-server-config-reloader docker tag to v0.24.0 (#8129) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * feat: extend git content spec with verbosity and retry configuration (#8089) * feat: model update for git content spec * chore: run auto generated for new testworkflows * feat: add verbosity and retry to git content types * chore: run auto generation for CRDs * chore: add git content updates to mappers + MCP schemas * feat: process git verbosity + retry logic * chore: restore extraneous change to test triggers * fix: cap retries and parsing around duration * feat: [TKC-6541] E2E tests workflow - UI build+serve with service, GH-integration-related changes (#8134) * tests - cloud-ui-e2e workflow for ui build using service * E2E tests - workflow updated to support GH integration * E2E tests - distributed reanabled * fix(quality-loop): mask parent workflow name on child running context (#8136) * fix: remediate CVE-2026-56865 (#8121) * fix: remediate CVE-2026-56865 * chore: upgrade helm database dependencies * fix: update chart.lock * chore(deps): update dependency @changesets/cli to v3.0.1 (#8130) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update module github.com/mikefarah/yq/v4 to v4.53.4 (#8131) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix: go mod Signed-off-by: Vladislav Sukhin <vladislav@kubeshop.io> * Revert "fix: go mod" This reverts commit 2a833c16d7c8cced5473d6fc0941d10c14521383. * Revert "chore(deps): update module github.com/mikefarah/yq/v4 to v4.53.4 (#8131)" This reverts commit 5f4adc9072efb6b5426dbd44d8f4790b9aecaad2. * chore(deps): update docker/setup-buildx-action action to v4.3.0 (#8133) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module google.golang.org/grpc to v1.83.1 (#8132) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module github.com/stretchr/testify to v1.12.1 (#8137) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update go toolchain directive to v1.27.0 (#8138) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update natsio/prometheus-nats-exporter docker tag to v0.20.2 (#8140) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update dependency inquirer to v14.1.0 (#8141) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update dependency golangci/golangci-lint to v2.13.0 (#8143) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update golang docker tag to v1.27.0 (#8144) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update dependency go to v1.27.0 (#8145) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * feat: add support for full git-ops ownership (#8120) * feat: add sync error handling with non-retyable errors * refactor: simplify logging around errors * doc: update documentation for new gitops ownership model * feat: add scheduler policy to targeted execution contracts (#8109) * feat: add targeted execution scheduler policy * fix(deps): update module github.com/gofiber/fiber/v2 to v2.52.15 (#8099) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module github.com/nats-io/nats-server/v2 to v2.14.5 (#8100) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update nats docker tag to v2.14.5 (#8101) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update golang docker tag to v1.26.6 (#8103) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix: support multiple kubeconfig files in KUBECONFIG env var (#8102) * fix: support multiple kubeconfig files in KUBECONFIG env var GetK8sClientConfig passed the whole KUBECONFIG value to BuildConfigFromFlags as a single file path, so a colon-separated list (semicolon on Windows) failed with a stat error like 'stat /a/config:/b/config: no such file or directory'. Split the value with filepath.SplitList and load it through clientcmd's loading rules, which merge the files with kubectl's own precedence semantics: the first file's current-context wins, and entries defined in later files remain reachable. Fixes #657 * fix: treat empty KUBECONFIG env var as unset A set-but-empty KUBECONFIG previously fell through to the default loading behavior; splitting the empty string produced no loading rules and errored instead. Skip the env branch when the value is empty so it falls back to ~/.kube/config or in-cluster config as before. * fix: replace App Engine logger with testkube logger in k8sclient (#8106) pkg/k8sclient imported google.golang.org/appengine/v2/log for a single error log in the port-forward helper, pulling the App Engine SDK into the CLI and agent for no benefit and diverging from the zap-based logger used across the codebase. Use pkg/log's DefaultLogger instead. Fixes #8105 * chore(deps): update postgres docker tag to v16.15 (#8104) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * feat(TKC-6504): rework init demo for new architecture (#8090) * feat(TKC-6504): rework init demo for new architecture * fix: make init demo re-install idempotent (reuse agent key) * chore(deps): update dependency microsoft.net.test.sdk to 18.9.0 (#8107) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update dependency turbo to v2.10.10 (#8110) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update dependency xunit.runner.visualstudio to v4 (#8113) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module github.com/minio/minio-go/v7 to v7.3.0 (#8114) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore: bump cloud-ui-e2e Playwright image to v1.62.1 (#8115) * fix(deps): update module github.com/adhocore/gronx to v1.20.3 (#8116) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module github.com/stretchr/testify to v1.12.0 (#8117) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * feat: add ability to rerun with latest workflow (#8097) * fix: make install.sh POSIX sh compatible and verify release checksums (#7917) * fix: make install.sh POSIX sh compatible and verify release checksums The documented install command pipes the script to `sh`, which is dash on Debian/Ubuntu, but the script relied on bash-only constructs ([[ ]], =~, set -o pipefail) and failed there. Also removes dead code paths (i386 assets are no longer published, the Windows uname case never matched, beta tag-name grep matched nothing since the v1.17 era), fails clearly when the version cannot be resolved instead of silently substituting "1", verifies the tarball against the release's checksums.txt, and downloads into a mktemp workdir with cleanup instead of the caller's directory. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: address review comments — add -L to API curls, exact-match checksum lookup Adds -L to the GitHub API curl calls for consistency with the tarball and checksums downloads, and replaces the grep regex lookup in checksums.txt with an exact awk field match since the tarball name contains dots that grep treats as wildcards. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> * feat(TKC-6504): Add warn message (#8118) * feat(TKC-6504): Add warn message * fix: lint * feat(proto): add GITINTEGRATION actor type for git-provider-triggered flows (#8108) * feat(proto): add GITINTEGRATION actor type for git-provider-triggered flows Adds a new value to the RunningContextType proto enum and to the TestWorkflowRunningContextActorType OpenAPI enum, plus the round-trip mapper cases between them. Downstream consumers (telemetry bucketer, CDEvents mapper) route the new actor into git-integration/event buckets. Kicks off the change needed on the cloud-api side to introduce a provider-agnostic "Git Integration" actor for Quality Loop executions (GitHub today, GitLab and Bitbucket without another proto bump). * fix: allow gitintegration actor in CRD schema and CLI validation * refactor: rename GITINTEGRATION actor to QUALITYLOOP for internal naming * refactor: keep proto RunningContextType_QUALITYLOOP but expose gitintegration to users Proto stays QUALITYLOOP internally per Ole's suggestion. OpenAPI actor type, CRD kubebuilder enum, CLI flag, and telemetry bucket surface as gitintegration so every customer-facing surface (REST API, kubectl testkube, helm charts, generated CRDs) speaks the provider-agnostic name. The mapper cross-translates between the two. * refactor: alias QUALITYLOOP for internal Go references, wire stays gitintegration Adds regen-safe alias files so Go code can reference testkube.QUALITYLOOP_TestWorkflowRunningContextActorType while the underlying wire value stays "gitintegration" for CLI, REST, CRDs, and helm charts. Consumers refactored to use the internal name; customer surfaces unchanged. * chore: fix goimports alignment * feat(controlplaneclient): let chained children inherit the parent's actor for sticky families (#8119) * feat(controlplaneclient): let chained children inherit the parent's actor for sticky families Today every chained execution is stamped as RunningContextType_EXECUTION, which the server maps to actor.type = testworkflow. That is the right default for regular composites but it does not survive a filter by the Quality Loop actor: the QL parent carries actor.type = gitintegration and the children it schedules end up as testworkflow, so filtering the Executions page by the parent's actor returns an empty list. This introduces a small sticky-actor family: when the parent's actor belongs to it (only QUALITYLOOP for now), the child inherits the parent's actor type instead of falling back to the EXECUTION default. Parent-chain walkers keep working because we extend the QUALITYLOOP mapper branch to populate actor.executionId / actor.executionPath from ParentExecutionIds the same way the EXECUTION branch does. Everything outside the sticky set (user-authored composites, cron/testtrigger/CR-scheduled runs) is byte-identical to today: the default injection path still stamps EXECUTION and the mapper still produces testworkflow. Telemetry buckets, Mixpanel dimensions, and downstream webhooks for non-QL flows are unaffected. * chore: reword sticky-actor comments to reference gitintegration instead of Quality Loop * refactor: move child-context decision onto the actor type via ChildRunningContextType method Reads better at the call site: instead of an external helper that returns (value, ok) and an if-branch to conditionally overwrite the default, the actor type answers directly what its chained children should carry. Injection funnel collapses to: Type: c.opts.ParentActorType.ChildRunningContextType() The default (RunningContextType_EXECUTION, mapped server-side to actor.type = testworkflow / the "Workflow" chip on the Executions page) lives inside the method along with the special cases, so extending it means editing one file next to the type itself. Tests moved from the controlplane client into the same package as the method for the same reason. * chore: drop verbose doc block on ChildRunningContextType * feat(runner): propagate parent RunningContext to child pods via GetExecutionWorkflow The runner builds ExecutionConfig for a scheduled pod from the ExecutionStart proto, which does not carry RunningContext. That leaves cfg.Execution.RunningContext nil in the toolkit; parentActorTypeFromRunningContext returns empty; the sticky family cannot see it's chained under a gitintegration parent and children fall back to actor.type=testworkflow. The runner already makes a follow-up GetExecutionWorkflow call for every ExecutionStart to fetch the enriched workflow. That response is documented as the vehicle for "any additional information that may be related" - the perfect place to attach the parent's actor identity without touching the ExecutionStart proto (which is used by many non-sticky paths). Adds a minimal ExecutionRunningContext projection (actor type + actor name) to GetExecutionWorkflowResponse, populates it server-side from the stored TWE, and reconstructs a TestWorkflowRunningContext on the runner side that parentActorTypeFromRunningContext can read unchanged. When the field is nil (older control planes, non-actor executions), behaviour matches today. * chore(deps): update dependency turbo to v2.10.11 (#8122) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update mongodb docker tag to v8.3.8 (#8123) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update postgresql docker tag to v18.6 (#8124) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update docker.io/kubeshop/testkube-postgres docker tag to v18.6 (#8127) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update docker.io/kubeshop/bitnami-mongodb docker tag to v8.3.8 (#8126) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update natsio/nats-server-config-reloader docker tag to v0.24.0 (#8129) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * feat: extend git content spec with verbosity and retry configuration (#8089) * feat: model update for git content spec * chore: run auto generated for new testworkflows * feat: add verbosity and retry to git content types * chore: run auto generation for CRDs * chore: add git content updates to mappers + MCP schemas * feat: process git verbosity + retry logic * chore: restore extraneous change to test triggers * fix: cap retries and parsing around duration * feat: [TKC-6541] E2E tests workflow - UI build+serve with service, GH-integration-related changes (#8134) * tests - cloud-ui-e2e workflow for ui build using service * E2E tests - workflow updated to support GH integration * E2E tests - distributed reanabled * fix(quality-loop): mask parent workflow name on child running context (#8136) * fix: remediate CVE-2026-56865 (#8121) * fix: remediate CVE-2026-56865 * chore: upgrade helm database dependencies * fix: update chart.lock * chore(deps): update dependency @changesets/cli to v3.0.1 (#8130) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update module github.com/mikefarah/yq/v4 to v4.53.4 (#8131) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix: go mod Signed-off-by: Vladislav Sukhin <vladislav@kubeshop.io> * Revert "fix: go mod" This reverts commit 2a833c16d7c8cced5473d6fc0941d10c14521383. * Revert "chore(deps): update module github.com/mikefarah/yq/v4 to v4.53.4 (#8131)" This reverts commit 5f4adc9072efb6b5426dbd44d8f4790b9aecaad2. * chore(deps): update docker/setup-buildx-action action to v4.3.0 (#8133) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module google.golang.org/grpc to v1.83.1 (#8132) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module github.com/stretchr/testify to v1.12.1 (#8137) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update go toolchain directive to v1.27.0 (#8138) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update natsio/prometheus-nats-exporter docker tag to v0.20.2 (#8140) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update dependency inquirer to v14.1.0 (#8141) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore: regenerate * chore: regenerate * chore(deps): update module github.com/mikefarah/yq/v4 to v4.53.4 (#8131) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix: go mod Signed-off-by: Vladislav Sukhin <vladislav@kubeshop.io> * Revert "fix: go mod" This reverts commit 2a833c16d7c8cced5473d6fc0941d10c14521383. * Revert "chore(deps): update module github.com/mikefarah/yq/v4 to v4.53.4 (#8131)" This reverts commit 5f4adc9072efb6b5426dbd44d8f4790b9aecaad2. * chore: fix golang version * fix: remove unrelated generated model changes --------- Signed-off-by: Vladislav Sukhin <vladislav@kubeshop.io> Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> Co-authored-by: gangadhar-res <gangadhar@resolve.ai> Co-authored-by: Valentin-Marko <69463262+Valentin-Marko@users.noreply.github.com> Co-authored-by: Razvan Topliceanu <47887589+topliceanurazvan@users.noreply.github.com> Co-authored-by: Tristan Rasmussen <tristan@testkube.io> Co-authored-by: Ole Lensmar <ole@lensmar.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Aurelio Buarque <aurelio@testkube.io> Co-authored-by: Tomasz Konieczny <tomasz.konieczny@kubeshop.io> Co-authored-by: Vladislav Sukhin <vladislav@kubeshop.io> * chore(deps): update go toolchain directive to v1.27.0 (#8146) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update depot/setup-action action to v1.7.2 (#8149) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update dependency golangci/golangci-lint to v2.13.1 (#8150) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update dependency @rollup/rollup-linux-x64-gnu to v4.62.5 (#8152) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update kubernetes monorepo to v0.36.4 (#8153) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix: dep update Signed-off-by: Vladislav Sukhin <vladislav@kubeshop.io> * chore(deps): update module github.com/mikefarah/yq/v4 to v4.53.6 (#8142) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore: run go mod tidy (#8156) * fix(deps): update k8s.io/kube-openapi digest to be32def (#8159) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module github.com/ohler55/ojg to v1.28.5 (#8155) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * Revert "fix: create builtin testworkflow templates via post-install hook job (#7988)" (#8158) This reverts commit d25992d855a4b7868861d5716970b142681fd101. * fix(deps): update module go.mongodb.org/mongo-driver/v2 to v2.8.1 (#8162) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update dependency nunit3testadapter to 6.3.0 (#8163) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> --------- Signed-off-by: Vladislav Sukhin <vladislav@kubeshop.io> Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> Co-authored-by: MsfPablo <129399053+MsfPablo@users.noreply.github.com> Co-authored-by: Tomasz Konieczny <tomasz.konieczny@kubeshop.io> Co-authored-by: Ole Lensmar <ole@lensmar.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Aurelio Buarque <aurelio@testkube.io> Co-authored-by: gangadhar-res <gangadhar@resolve.ai> Co-authored-by: Valentin-Marko <69463262+Valentin-Marko@users.noreply.github.com> Co-authored-by: Razvan Topliceanu <47887589+topliceanurazvan@users.noreply.github.com> Co-authored-by: Caio Medeiros Pinto <caio@testkube.io> Co-authored-by: Vladislav Sukhin <vladislav@kubeshop.io> Co-authored-by: ypoplavs <45286051+ypoplavs@users.noreply.github.com>
| Commit: | df570bb | |
|---|---|---|
| Author: | Aurelio Buarque | |
| Committer: | buarki | |
feat(controlplaneclient): let chained children inherit the parent's actor for sticky families (#8119) * feat(controlplaneclient): let chained children inherit the parent's actor for sticky families Today every chained execution is stamped as RunningContextType_EXECUTION, which the server maps to actor.type = testworkflow. That is the right default for regular composites but it does not survive a filter by the Quality Loop actor: the QL parent carries actor.type = gitintegration and the children it schedules end up as testworkflow, so filtering the Executions page by the parent's actor returns an empty list. This introduces a small sticky-actor family: when the parent's actor belongs to it (only QUALITYLOOP for now), the child inherits the parent's actor type instead of falling back to the EXECUTION default. Parent-chain walkers keep working because we extend the QUALITYLOOP mapper branch to populate actor.executionId / actor.executionPath from ParentExecutionIds the same way the EXECUTION branch does. Everything outside the sticky set (user-authored composites, cron/testtrigger/CR-scheduled runs) is byte-identical to today: the default injection path still stamps EXECUTION and the mapper still produces testworkflow. Telemetry buckets, Mixpanel dimensions, and downstream webhooks for non-QL flows are unaffected. * chore: reword sticky-actor comments to reference gitintegration instead of Quality Loop * refactor: move child-context decision onto the actor type via ChildRunningContextType method Reads better at the call site: instead of an external helper that returns (value, ok) and an if-branch to conditionally overwrite the default, the actor type answers directly what its chained children should carry. Injection funnel collapses to: Type: c.opts.ParentActorType.ChildRunningContextType() The default (RunningContextType_EXECUTION, mapped server-side to actor.type = testworkflow / the "Workflow" chip on the Executions page) lives inside the method along with the special cases, so extending it means editing one file next to the type itself. Tests moved from the controlplane client into the same package as the method for the same reason. * chore: drop verbose doc block on ChildRunningContextType * feat(runner): propagate parent RunningContext to child pods via GetExecutionWorkflow The runner builds ExecutionConfig for a scheduled pod from the ExecutionStart proto, which does not carry RunningContext. That leaves cfg.Execution.RunningContext nil in the toolkit; parentActorTypeFromRunningContext returns empty; the sticky family cannot see it's chained under a gitintegration parent and children fall back to actor.type=testworkflow. The runner already makes a follow-up GetExecutionWorkflow call for every ExecutionStart to fetch the enriched workflow. That response is documented as the vehicle for "any additional information that may be related" - the perfect place to attach the parent's actor identity without touching the ExecutionStart proto (which is used by many non-sticky paths). Adds a minimal ExecutionRunningContext projection (actor type + actor name) to GetExecutionWorkflowResponse, populates it server-side from the stored TWE, and reconstructs a TestWorkflowRunningContext on the runner side that parentActorTypeFromRunningContext can read unchanged. When the field is nil (older control planes, non-actor executions), behaviour matches today.
| Commit: | 29ea89d | |
|---|---|---|
| Author: | Aurelio Buarque | |
| Committer: | buarki | |
feat(proto): add GITINTEGRATION actor type for git-provider-triggered flows (#8108) * feat(proto): add GITINTEGRATION actor type for git-provider-triggered flows Adds a new value to the RunningContextType proto enum and to the TestWorkflowRunningContextActorType OpenAPI enum, plus the round-trip mapper cases between them. Downstream consumers (telemetry bucketer, CDEvents mapper) route the new actor into git-integration/event buckets. Kicks off the change needed on the cloud-api side to introduce a provider-agnostic "Git Integration" actor for Quality Loop executions (GitHub today, GitLab and Bitbucket without another proto bump). * fix: allow gitintegration actor in CRD schema and CLI validation * refactor: rename GITINTEGRATION actor to QUALITYLOOP for internal naming * refactor: keep proto RunningContextType_QUALITYLOOP but expose gitintegration to users Proto stays QUALITYLOOP internally per Ole's suggestion. OpenAPI actor type, CRD kubebuilder enum, CLI flag, and telemetry bucket surface as gitintegration so every customer-facing surface (REST API, kubectl testkube, helm charts, generated CRDs) speaks the provider-agnostic name. The mapper cross-translates between the two. * refactor: alias QUALITYLOOP for internal Go references, wire stays gitintegration Adds regen-safe alias files so Go code can reference testkube.QUALITYLOOP_TestWorkflowRunningContextActorType while the underlying wire value stays "gitintegration" for CLI, REST, CRDs, and helm charts. Consumers refactored to use the internal name; customer surfaces unchanged. * chore: fix goimports alignment
| Commit: | fd92a0e | |
|---|---|---|
| Author: | Vladislav Sukhin | |
Merge branch 'main' into vsukhin/feature/suite-step-exchange Conflict: pkg/cloud/service.pb.go. Both sides had regenerated it - main added scheduler_policy to ExecutionTarget, this branch added the artifact-read RPCs - and the conflicts fell inside the serialized rawDesc byte array, where merging by hand would leave a corrupt descriptor. proto/service.proto merged cleanly with both sides' additions, so the generated file was regenerated from it with the pinned toolchain (buf v1.68.1, proto/buf.gen.old.yaml, protoc-gen-go v1.32.0, grpc-go v1.2.0). Regeneration also rewrote pkg/logs/pb/* and pkg/cloud/service_grpc.pb.go identically to what was already committed, which is the check that the toolchain matches; those were restored to avoid line-ending-only noise. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
| Commit: | 4c2c672 | |
|---|---|---|
| Author: | Caio Medeiros Pinto | |
| Committer: | GitHub | |
chore: pre-release 2.12.2 (#8148) * fix: create builtin testworkflow templates via post-install hook job (#7988) * fix: create builtin testworkflow templates via post-install hook job * update docs * update kubectl image * fix docs * fix docs * feat: add support for matchexpressions (#8038) * add support for matchexpressions * print an error * fix a typo * correct ident * feat(TKC-6504): rework init demo for new architecture (#8090) * feat(TKC-6504): rework init demo for new architecture * fix: make init demo re-install idempotent (reuse agent key) * feat(TKC-6504): Add warn message (#8118) * feat(TKC-6504): Add warn message * fix: lint * fix: missing model and test * fix: remediate CVE-2026-56865 (#8121) * fix: remediate CVE-2026-56865 * chore: upgrade helm database dependencies * fix: update chart.lock * chore: go mod tidy * feat: add scheduler policy to targeted execution contracts (#8109) * feat: add targeted execution scheduler policy * fix(deps): update module github.com/gofiber/fiber/v2 to v2.52.15 (#8099) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module github.com/nats-io/nats-server/v2 to v2.14.5 (#8100) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update nats docker tag to v2.14.5 (#8101) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update golang docker tag to v1.26.6 (#8103) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix: support multiple kubeconfig files in KUBECONFIG env var (#8102) * fix: support multiple kubeconfig files in KUBECONFIG env var GetK8sClientConfig passed the whole KUBECONFIG value to BuildConfigFromFlags as a single file path, so a colon-separated list (semicolon on Windows) failed with a stat error like 'stat /a/config:/b/config: no such file or directory'. Split the value with filepath.SplitList and load it through clientcmd's loading rules, which merge the files with kubectl's own precedence semantics: the first file's current-context wins, and entries defined in later files remain reachable. Fixes #657 * fix: treat empty KUBECONFIG env var as unset A set-but-empty KUBECONFIG previously fell through to the default loading behavior; splitting the empty string produced no loading rules and errored instead. Skip the env branch when the value is empty so it falls back to ~/.kube/config or in-cluster config as before. * fix: replace App Engine logger with testkube logger in k8sclient (#8106) pkg/k8sclient imported google.golang.org/appengine/v2/log for a single error log in the port-forward helper, pulling the App Engine SDK into the CLI and agent for no benefit and diverging from the zap-based logger used across the codebase. Use pkg/log's DefaultLogger instead. Fixes #8105 * chore(deps): update postgres docker tag to v16.15 (#8104) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * feat(TKC-6504): rework init demo for new architecture (#8090) * feat(TKC-6504): rework init demo for new architecture * fix: make init demo re-install idempotent (reuse agent key) * chore(deps): update dependency microsoft.net.test.sdk to 18.9.0 (#8107) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update dependency turbo to v2.10.10 (#8110) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update dependency xunit.runner.visualstudio to v4 (#8113) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module github.com/minio/minio-go/v7 to v7.3.0 (#8114) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore: bump cloud-ui-e2e Playwright image to v1.62.1 (#8115) * fix(deps): update module github.com/adhocore/gronx to v1.20.3 (#8116) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module github.com/stretchr/testify to v1.12.0 (#8117) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * feat: add ability to rerun with latest workflow (#8097) * fix: make install.sh POSIX sh compatible and verify release checksums (#7917) * fix: make install.sh POSIX sh compatible and verify release checksums The documented install command pipes the script to `sh`, which is dash on Debian/Ubuntu, but the script relied on bash-only constructs ([[ ]], =~, set -o pipefail) and failed there. Also removes dead code paths (i386 assets are no longer published, the Windows uname case never matched, beta tag-name grep matched nothing since the v1.17 era), fails clearly when the version cannot be resolved instead of silently substituting "1", verifies the tarball against the release's checksums.txt, and downloads into a mktemp workdir with cleanup instead of the caller's directory. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: address review comments — add -L to API curls, exact-match checksum lookup Adds -L to the GitHub API curl calls for consistency with the tarball and checksums downloads, and replaces the grep regex lookup in checksums.txt with an exact awk field match since the tarball name contains dots that grep treats as wildcards. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> * feat(TKC-6504): Add warn message (#8118) * feat(TKC-6504): Add warn message * fix: lint * feat(proto): add GITINTEGRATION actor type for git-provider-triggered flows (#8108) * feat(proto): add GITINTEGRATION actor type for git-provider-triggered flows Adds a new value to the RunningContextType proto enum and to the TestWorkflowRunningContextActorType OpenAPI enum, plus the round-trip mapper cases between them. Downstream consumers (telemetry bucketer, CDEvents mapper) route the new actor into git-integration/event buckets. Kicks off the change needed on the cloud-api side to introduce a provider-agnostic "Git Integration" actor for Quality Loop executions (GitHub today, GitLab and Bitbucket without another proto bump). * fix: allow gitintegration actor in CRD schema and CLI validation * refactor: rename GITINTEGRATION actor to QUALITYLOOP for internal naming * refactor: keep proto RunningContextType_QUALITYLOOP but expose gitintegration to users Proto stays QUALITYLOOP internally per Ole's suggestion. OpenAPI actor type, CRD kubebuilder enum, CLI flag, and telemetry bucket surface as gitintegration so every customer-facing surface (REST API, kubectl testkube, helm charts, generated CRDs) speaks the provider-agnostic name. The mapper cross-translates between the two. * refactor: alias QUALITYLOOP for internal Go references, wire stays gitintegration Adds regen-safe alias files so Go code can reference testkube.QUALITYLOOP_TestWorkflowRunningContextActorType while the underlying wire value stays "gitintegration" for CLI, REST, CRDs, and helm charts. Consumers refactored to use the internal name; customer surfaces unchanged. * chore: fix goimports alignment * feat(controlplaneclient): let chained children inherit the parent's actor for sticky families (#8119) * feat(controlplaneclient): let chained children inherit the parent's actor for sticky families Today every chained execution is stamped as RunningContextType_EXECUTION, which the server maps to actor.type = testworkflow. That is the right default for regular composites but it does not survive a filter by the Quality Loop actor: the QL parent carries actor.type = gitintegration and the children it schedules end up as testworkflow, so filtering the Executions page by the parent's actor returns an empty list. This introduces a small sticky-actor family: when the parent's actor belongs to it (only QUALITYLOOP for now), the child inherits the parent's actor type instead of falling back to the EXECUTION default. Parent-chain walkers keep working because we extend the QUALITYLOOP mapper branch to populate actor.executionId / actor.executionPath from ParentExecutionIds the same way the EXECUTION branch does. Everything outside the sticky set (user-authored composites, cron/testtrigger/CR-scheduled runs) is byte-identical to today: the default injection path still stamps EXECUTION and the mapper still produces testworkflow. Telemetry buckets, Mixpanel dimensions, and downstream webhooks for non-QL flows are unaffected. * chore: reword sticky-actor comments to reference gitintegration instead of Quality Loop * refactor: move child-context decision onto the actor type via ChildRunningContextType method Reads better at the call site: instead of an external helper that returns (value, ok) and an if-branch to conditionally overwrite the default, the actor type answers directly what its chained children should carry. Injection funnel collapses to: Type: c.opts.ParentActorType.ChildRunningContextType() The default (RunningContextType_EXECUTION, mapped server-side to actor.type = testworkflow / the "Workflow" chip on the Executions page) lives inside the method along with the special cases, so extending it means editing one file next to the type itself. Tests moved from the controlplane client into the same package as the method for the same reason. * chore: drop verbose doc block on ChildRunningContextType * feat(runner): propagate parent RunningContext to child pods via GetExecutionWorkflow The runner builds ExecutionConfig for a scheduled pod from the ExecutionStart proto, which does not carry RunningContext. That leaves cfg.Execution.RunningContext nil in the toolkit; parentActorTypeFromRunningContext returns empty; the sticky family cannot see it's chained under a gitintegration parent and children fall back to actor.type=testworkflow. The runner already makes a follow-up GetExecutionWorkflow call for every ExecutionStart to fetch the enriched workflow. That response is documented as the vehicle for "any additional information that may be related" - the perfect place to attach the parent's actor identity without touching the ExecutionStart proto (which is used by many non-sticky paths). Adds a minimal ExecutionRunningContext projection (actor type + actor name) to GetExecutionWorkflowResponse, populates it server-side from the stored TWE, and reconstructs a TestWorkflowRunningContext on the runner side that parentActorTypeFromRunningContext can read unchanged. When the field is nil (older control planes, non-actor executions), behaviour matches today. * chore(deps): update dependency turbo to v2.10.11 (#8122) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update mongodb docker tag to v8.3.8 (#8123) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update postgresql docker tag to v18.6 (#8124) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update docker.io/kubeshop/testkube-postgres docker tag to v18.6 (#8127) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update docker.io/kubeshop/bitnami-mongodb docker tag to v8.3.8 (#8126) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update natsio/nats-server-config-reloader docker tag to v0.24.0 (#8129) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * feat: extend git content spec with verbosity and retry configuration (#8089) * feat: model update for git content spec * chore: run auto generated for new testworkflows * feat: add verbosity and retry to git content types * chore: run auto generation for CRDs * chore: add git content updates to mappers + MCP schemas * feat: process git verbosity + retry logic * chore: restore extraneous change to test triggers * fix: cap retries and parsing around duration * feat: [TKC-6541] E2E tests workflow - UI build+serve with service, GH-integration-related changes (#8134) * tests - cloud-ui-e2e workflow for ui build using service * E2E tests - workflow updated to support GH integration * E2E tests - distributed reanabled * fix(quality-loop): mask parent workflow name on child running context (#8136) * fix: remediate CVE-2026-56865 (#8121) * fix: remediate CVE-2026-56865 * chore: upgrade helm database dependencies * fix: update chart.lock * chore(deps): update dependency @changesets/cli to v3.0.1 (#8130) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update module github.com/mikefarah/yq/v4 to v4.53.4 (#8131) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix: go mod Signed-off-by: Vladislav Sukhin <vladislav@kubeshop.io> * Revert "fix: go mod" This reverts commit 2a833c16d7c8cced5473d6fc0941d10c14521383. * Revert "chore(deps): update module github.com/mikefarah/yq/v4 to v4.53.4 (#8131)" This reverts commit 5f4adc9072efb6b5426dbd44d8f4790b9aecaad2. * chore(deps): update docker/setup-buildx-action action to v4.3.0 (#8133) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module google.golang.org/grpc to v1.83.1 (#8132) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module github.com/stretchr/testify to v1.12.1 (#8137) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update go toolchain directive to v1.27.0 (#8138) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update natsio/prometheus-nats-exporter docker tag to v0.20.2 (#8140) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update dependency inquirer to v14.1.0 (#8141) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore: regenerate * chore: regenerate * chore(deps): update module github.com/mikefarah/yq/v4 to v4.53.4 (#8131) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix: go mod Signed-off-by: Vladislav Sukhin <vladislav@kubeshop.io> * Revert "fix: go mod" This reverts commit 2a833c16d7c8cced5473d6fc0941d10c14521383. * Revert "chore(deps): update module github.com/mikefarah/yq/v4 to v4.53.4 (#8131)" This reverts commit 5f4adc9072efb6b5426dbd44d8f4790b9aecaad2. * chore: fix golang version * fix: remove unrelated generated model changes --------- Signed-off-by: Vladislav Sukhin <vladislav@kubeshop.io> Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> Co-authored-by: gangadhar-res <gangadhar@resolve.ai> Co-authored-by: Valentin-Marko <69463262+Valentin-Marko@users.noreply.github.com> Co-authored-by: Razvan Topliceanu <47887589+topliceanurazvan@users.noreply.github.com> Co-authored-by: Tristan Rasmussen <tristan@testkube.io> Co-authored-by: Ole Lensmar <ole@lensmar.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Aurelio Buarque <aurelio@testkube.io> Co-authored-by: Tomasz Konieczny <tomasz.konieczny@kubeshop.io> Co-authored-by: Vladislav Sukhin <vladislav@kubeshop.io> * chore: go mod tidy --------- Signed-off-by: Vladislav Sukhin <vladislav@kubeshop.io> Co-authored-by: ypoplavs <45286051+ypoplavs@users.noreply.github.com> Co-authored-by: Valentin-Marko <69463262+Valentin-Marko@users.noreply.github.com> Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> Co-authored-by: gangadhar-res <gangadhar@resolve.ai> Co-authored-by: Razvan Topliceanu <47887589+topliceanurazvan@users.noreply.github.com> Co-authored-by: Tristan Rasmussen <tristan@testkube.io> Co-authored-by: Ole Lensmar <ole@lensmar.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Aurelio Buarque <aurelio@testkube.io> Co-authored-by: Tomasz Konieczny <tomasz.konieczny@kubeshop.io> Co-authored-by: Vladislav Sukhin <vladislav@kubeshop.io>
| Commit: | 679cd69 | |
|---|---|---|
| Author: | Caio Medeiros Pinto | |
| Committer: | GitHub | |
feat: add scheduler policy to targeted execution contracts (#8109) * feat: add targeted execution scheduler policy * fix(deps): update module github.com/gofiber/fiber/v2 to v2.52.15 (#8099) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module github.com/nats-io/nats-server/v2 to v2.14.5 (#8100) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update nats docker tag to v2.14.5 (#8101) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update golang docker tag to v1.26.6 (#8103) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix: support multiple kubeconfig files in KUBECONFIG env var (#8102) * fix: support multiple kubeconfig files in KUBECONFIG env var GetK8sClientConfig passed the whole KUBECONFIG value to BuildConfigFromFlags as a single file path, so a colon-separated list (semicolon on Windows) failed with a stat error like 'stat /a/config:/b/config: no such file or directory'. Split the value with filepath.SplitList and load it through clientcmd's loading rules, which merge the files with kubectl's own precedence semantics: the first file's current-context wins, and entries defined in later files remain reachable. Fixes #657 * fix: treat empty KUBECONFIG env var as unset A set-but-empty KUBECONFIG previously fell through to the default loading behavior; splitting the empty string produced no loading rules and errored instead. Skip the env branch when the value is empty so it falls back to ~/.kube/config or in-cluster config as before. * fix: replace App Engine logger with testkube logger in k8sclient (#8106) pkg/k8sclient imported google.golang.org/appengine/v2/log for a single error log in the port-forward helper, pulling the App Engine SDK into the CLI and agent for no benefit and diverging from the zap-based logger used across the codebase. Use pkg/log's DefaultLogger instead. Fixes #8105 * chore(deps): update postgres docker tag to v16.15 (#8104) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * feat(TKC-6504): rework init demo for new architecture (#8090) * feat(TKC-6504): rework init demo for new architecture * fix: make init demo re-install idempotent (reuse agent key) * chore(deps): update dependency microsoft.net.test.sdk to 18.9.0 (#8107) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update dependency turbo to v2.10.10 (#8110) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update dependency xunit.runner.visualstudio to v4 (#8113) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module github.com/minio/minio-go/v7 to v7.3.0 (#8114) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore: bump cloud-ui-e2e Playwright image to v1.62.1 (#8115) * fix(deps): update module github.com/adhocore/gronx to v1.20.3 (#8116) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module github.com/stretchr/testify to v1.12.0 (#8117) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * feat: add ability to rerun with latest workflow (#8097) * fix: make install.sh POSIX sh compatible and verify release checksums (#7917) * fix: make install.sh POSIX sh compatible and verify release checksums The documented install command pipes the script to `sh`, which is dash on Debian/Ubuntu, but the script relied on bash-only constructs ([[ ]], =~, set -o pipefail) and failed there. Also removes dead code paths (i386 assets are no longer published, the Windows uname case never matched, beta tag-name grep matched nothing since the v1.17 era), fails clearly when the version cannot be resolved instead of silently substituting "1", verifies the tarball against the release's checksums.txt, and downloads into a mktemp workdir with cleanup instead of the caller's directory. Co-authored-by: Cursor <cursoragent@cursor.com> * fix: address review comments — add -L to API curls, exact-match checksum lookup Adds -L to the GitHub API curl calls for consistency with the tarball and checksums downloads, and replaces the grep regex lookup in checksums.txt with an exact awk field match since the tarball name contains dots that grep treats as wildcards. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> * feat(TKC-6504): Add warn message (#8118) * feat(TKC-6504): Add warn message * fix: lint * feat(proto): add GITINTEGRATION actor type for git-provider-triggered flows (#8108) * feat(proto): add GITINTEGRATION actor type for git-provider-triggered flows Adds a new value to the RunningContextType proto enum and to the TestWorkflowRunningContextActorType OpenAPI enum, plus the round-trip mapper cases between them. Downstream consumers (telemetry bucketer, CDEvents mapper) route the new actor into git-integration/event buckets. Kicks off the change needed on the cloud-api side to introduce a provider-agnostic "Git Integration" actor for Quality Loop executions (GitHub today, GitLab and Bitbucket without another proto bump). * fix: allow gitintegration actor in CRD schema and CLI validation * refactor: rename GITINTEGRATION actor to QUALITYLOOP for internal naming * refactor: keep proto RunningContextType_QUALITYLOOP but expose gitintegration to users Proto stays QUALITYLOOP internally per Ole's suggestion. OpenAPI actor type, CRD kubebuilder enum, CLI flag, and telemetry bucket surface as gitintegration so every customer-facing surface (REST API, kubectl testkube, helm charts, generated CRDs) speaks the provider-agnostic name. The mapper cross-translates between the two. * refactor: alias QUALITYLOOP for internal Go references, wire stays gitintegration Adds regen-safe alias files so Go code can reference testkube.QUALITYLOOP_TestWorkflowRunningContextActorType while the underlying wire value stays "gitintegration" for CLI, REST, CRDs, and helm charts. Consumers refactored to use the internal name; customer surfaces unchanged. * chore: fix goimports alignment * feat(controlplaneclient): let chained children inherit the parent's actor for sticky families (#8119) * feat(controlplaneclient): let chained children inherit the parent's actor for sticky families Today every chained execution is stamped as RunningContextType_EXECUTION, which the server maps to actor.type = testworkflow. That is the right default for regular composites but it does not survive a filter by the Quality Loop actor: the QL parent carries actor.type = gitintegration and the children it schedules end up as testworkflow, so filtering the Executions page by the parent's actor returns an empty list. This introduces a small sticky-actor family: when the parent's actor belongs to it (only QUALITYLOOP for now), the child inherits the parent's actor type instead of falling back to the EXECUTION default. Parent-chain walkers keep working because we extend the QUALITYLOOP mapper branch to populate actor.executionId / actor.executionPath from ParentExecutionIds the same way the EXECUTION branch does. Everything outside the sticky set (user-authored composites, cron/testtrigger/CR-scheduled runs) is byte-identical to today: the default injection path still stamps EXECUTION and the mapper still produces testworkflow. Telemetry buckets, Mixpanel dimensions, and downstream webhooks for non-QL flows are unaffected. * chore: reword sticky-actor comments to reference gitintegration instead of Quality Loop * refactor: move child-context decision onto the actor type via ChildRunningContextType method Reads better at the call site: instead of an external helper that returns (value, ok) and an if-branch to conditionally overwrite the default, the actor type answers directly what its chained children should carry. Injection funnel collapses to: Type: c.opts.ParentActorType.ChildRunningContextType() The default (RunningContextType_EXECUTION, mapped server-side to actor.type = testworkflow / the "Workflow" chip on the Executions page) lives inside the method along with the special cases, so extending it means editing one file next to the type itself. Tests moved from the controlplane client into the same package as the method for the same reason. * chore: drop verbose doc block on ChildRunningContextType * feat(runner): propagate parent RunningContext to child pods via GetExecutionWorkflow The runner builds ExecutionConfig for a scheduled pod from the ExecutionStart proto, which does not carry RunningContext. That leaves cfg.Execution.RunningContext nil in the toolkit; parentActorTypeFromRunningContext returns empty; the sticky family cannot see it's chained under a gitintegration parent and children fall back to actor.type=testworkflow. The runner already makes a follow-up GetExecutionWorkflow call for every ExecutionStart to fetch the enriched workflow. That response is documented as the vehicle for "any additional information that may be related" - the perfect place to attach the parent's actor identity without touching the ExecutionStart proto (which is used by many non-sticky paths). Adds a minimal ExecutionRunningContext projection (actor type + actor name) to GetExecutionWorkflowResponse, populates it server-side from the stored TWE, and reconstructs a TestWorkflowRunningContext on the runner side that parentActorTypeFromRunningContext can read unchanged. When the field is nil (older control planes, non-actor executions), behaviour matches today. * chore(deps): update dependency turbo to v2.10.11 (#8122) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update mongodb docker tag to v8.3.8 (#8123) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update postgresql docker tag to v18.6 (#8124) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update docker.io/kubeshop/testkube-postgres docker tag to v18.6 (#8127) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update docker.io/kubeshop/bitnami-mongodb docker tag to v8.3.8 (#8126) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update natsio/nats-server-config-reloader docker tag to v0.24.0 (#8129) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * feat: extend git content spec with verbosity and retry configuration (#8089) * feat: model update for git content spec * chore: run auto generated for new testworkflows * feat: add verbosity and retry to git content types * chore: run auto generation for CRDs * chore: add git content updates to mappers + MCP schemas * feat: process git verbosity + retry logic * chore: restore extraneous change to test triggers * fix: cap retries and parsing around duration * feat: [TKC-6541] E2E tests workflow - UI build+serve with service, GH-integration-related changes (#8134) * tests - cloud-ui-e2e workflow for ui build using service * E2E tests - workflow updated to support GH integration * E2E tests - distributed reanabled * fix(quality-loop): mask parent workflow name on child running context (#8136) * fix: remediate CVE-2026-56865 (#8121) * fix: remediate CVE-2026-56865 * chore: upgrade helm database dependencies * fix: update chart.lock * chore(deps): update dependency @changesets/cli to v3.0.1 (#8130) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update module github.com/mikefarah/yq/v4 to v4.53.4 (#8131) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix: go mod Signed-off-by: Vladislav Sukhin <vladislav@kubeshop.io> * Revert "fix: go mod" This reverts commit 2a833c16d7c8cced5473d6fc0941d10c14521383. * Revert "chore(deps): update module github.com/mikefarah/yq/v4 to v4.53.4 (#8131)" This reverts commit 5f4adc9072efb6b5426dbd44d8f4790b9aecaad2. * chore(deps): update docker/setup-buildx-action action to v4.3.0 (#8133) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module google.golang.org/grpc to v1.83.1 (#8132) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update module github.com/stretchr/testify to v1.12.1 (#8137) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update go toolchain directive to v1.27.0 (#8138) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore(deps): update natsio/prometheus-nats-exporter docker tag to v0.20.2 (#8140) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix(deps): update dependency inquirer to v14.1.0 (#8141) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * chore: regenerate * chore: regenerate * chore(deps): update module github.com/mikefarah/yq/v4 to v4.53.4 (#8131) Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> * fix: go mod Signed-off-by: Vladislav Sukhin <vladislav@kubeshop.io> * Revert "fix: go mod" This reverts commit 2a833c16d7c8cced5473d6fc0941d10c14521383. * Revert "chore(deps): update module github.com/mikefarah/yq/v4 to v4.53.4 (#8131)" This reverts commit 5f4adc9072efb6b5426dbd44d8f4790b9aecaad2. * chore: fix golang version * fix: remove unrelated generated model changes --------- Signed-off-by: Vladislav Sukhin <vladislav@kubeshop.io> Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com> Co-authored-by: gangadhar-res <gangadhar@resolve.ai> Co-authored-by: Valentin-Marko <69463262+Valentin-Marko@users.noreply.github.com> Co-authored-by: Razvan Topliceanu <47887589+topliceanurazvan@users.noreply.github.com> Co-authored-by: Tristan Rasmussen <tristan@testkube.io> Co-authored-by: Ole Lensmar <ole@lensmar.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Aurelio Buarque <aurelio@testkube.io> Co-authored-by: Tomasz Konieczny <tomasz.konieczny@kubeshop.io> Co-authored-by: Vladislav Sukhin <vladislav@kubeshop.io>
| Commit: | c487c04 | |
|---|---|---|
| Author: | Vladislav Sukhin | |
Merge branch 'main' into vsukhin/feature/suite-step-exchange Signed-off-by: Vladislav Sukhin <vladislav@kubeshop.io> # Conflicts: # pkg/cloud/service.pb.go
| Commit: | 58eebbc | |
|---|---|---|
| Author: | Aurelio Buarque | |
| Committer: | Caio Medeiros Pinto | |
feat(controlplaneclient): let chained children inherit the parent's actor for sticky families (#8119) * feat(controlplaneclient): let chained children inherit the parent's actor for sticky families Today every chained execution is stamped as RunningContextType_EXECUTION, which the server maps to actor.type = testworkflow. That is the right default for regular composites but it does not survive a filter by the Quality Loop actor: the QL parent carries actor.type = gitintegration and the children it schedules end up as testworkflow, so filtering the Executions page by the parent's actor returns an empty list. This introduces a small sticky-actor family: when the parent's actor belongs to it (only QUALITYLOOP for now), the child inherits the parent's actor type instead of falling back to the EXECUTION default. Parent-chain walkers keep working because we extend the QUALITYLOOP mapper branch to populate actor.executionId / actor.executionPath from ParentExecutionIds the same way the EXECUTION branch does. Everything outside the sticky set (user-authored composites, cron/testtrigger/CR-scheduled runs) is byte-identical to today: the default injection path still stamps EXECUTION and the mapper still produces testworkflow. Telemetry buckets, Mixpanel dimensions, and downstream webhooks for non-QL flows are unaffected. * chore: reword sticky-actor comments to reference gitintegration instead of Quality Loop * refactor: move child-context decision onto the actor type via ChildRunningContextType method Reads better at the call site: instead of an external helper that returns (value, ok) and an if-branch to conditionally overwrite the default, the actor type answers directly what its chained children should carry. Injection funnel collapses to: Type: c.opts.ParentActorType.ChildRunningContextType() The default (RunningContextType_EXECUTION, mapped server-side to actor.type = testworkflow / the "Workflow" chip on the Executions page) lives inside the method along with the special cases, so extending it means editing one file next to the type itself. Tests moved from the controlplane client into the same package as the method for the same reason. * chore: drop verbose doc block on ChildRunningContextType * feat(runner): propagate parent RunningContext to child pods via GetExecutionWorkflow The runner builds ExecutionConfig for a scheduled pod from the ExecutionStart proto, which does not carry RunningContext. That leaves cfg.Execution.RunningContext nil in the toolkit; parentActorTypeFromRunningContext returns empty; the sticky family cannot see it's chained under a gitintegration parent and children fall back to actor.type=testworkflow. The runner already makes a follow-up GetExecutionWorkflow call for every ExecutionStart to fetch the enriched workflow. That response is documented as the vehicle for "any additional information that may be related" - the perfect place to attach the parent's actor identity without touching the ExecutionStart proto (which is used by many non-sticky paths). Adds a minimal ExecutionRunningContext projection (actor type + actor name) to GetExecutionWorkflowResponse, populates it server-side from the stored TWE, and reconstructs a TestWorkflowRunningContext on the runner side that parentActorTypeFromRunningContext can read unchanged. When the field is nil (older control planes, non-actor executions), behaviour matches today.
| Commit: | ec2df28 | |
|---|---|---|
| Author: | Aurelio Buarque | |
| Committer: | Caio Medeiros Pinto | |
feat(proto): add GITINTEGRATION actor type for git-provider-triggered flows (#8108) * feat(proto): add GITINTEGRATION actor type for git-provider-triggered flows Adds a new value to the RunningContextType proto enum and to the TestWorkflowRunningContextActorType OpenAPI enum, plus the round-trip mapper cases between them. Downstream consumers (telemetry bucketer, CDEvents mapper) route the new actor into git-integration/event buckets. Kicks off the change needed on the cloud-api side to introduce a provider-agnostic "Git Integration" actor for Quality Loop executions (GitHub today, GitLab and Bitbucket without another proto bump). * fix: allow gitintegration actor in CRD schema and CLI validation * refactor: rename GITINTEGRATION actor to QUALITYLOOP for internal naming * refactor: keep proto RunningContextType_QUALITYLOOP but expose gitintegration to users Proto stays QUALITYLOOP internally per Ole's suggestion. OpenAPI actor type, CRD kubebuilder enum, CLI flag, and telemetry bucket surface as gitintegration so every customer-facing surface (REST API, kubectl testkube, helm charts, generated CRDs) speaks the provider-agnostic name. The mapper cross-translates between the two. * refactor: alias QUALITYLOOP for internal Go references, wire stays gitintegration Adds regen-safe alias files so Go code can reference testkube.QUALITYLOOP_TestWorkflowRunningContextActorType while the underlying wire value stays "gitintegration" for CLI, REST, CRDs, and helm charts. Consumers refactored to use the internal name; customer surfaces unchanged. * chore: fix goimports alignment
| Commit: | 3b00607 | |
|---|---|---|
| Author: | Aurelio Buarque | |
| Committer: | GitHub | |
feat(controlplaneclient): let chained children inherit the parent's actor for sticky families (#8119) * feat(controlplaneclient): let chained children inherit the parent's actor for sticky families Today every chained execution is stamped as RunningContextType_EXECUTION, which the server maps to actor.type = testworkflow. That is the right default for regular composites but it does not survive a filter by the Quality Loop actor: the QL parent carries actor.type = gitintegration and the children it schedules end up as testworkflow, so filtering the Executions page by the parent's actor returns an empty list. This introduces a small sticky-actor family: when the parent's actor belongs to it (only QUALITYLOOP for now), the child inherits the parent's actor type instead of falling back to the EXECUTION default. Parent-chain walkers keep working because we extend the QUALITYLOOP mapper branch to populate actor.executionId / actor.executionPath from ParentExecutionIds the same way the EXECUTION branch does. Everything outside the sticky set (user-authored composites, cron/testtrigger/CR-scheduled runs) is byte-identical to today: the default injection path still stamps EXECUTION and the mapper still produces testworkflow. Telemetry buckets, Mixpanel dimensions, and downstream webhooks for non-QL flows are unaffected. * chore: reword sticky-actor comments to reference gitintegration instead of Quality Loop * refactor: move child-context decision onto the actor type via ChildRunningContextType method Reads better at the call site: instead of an external helper that returns (value, ok) and an if-branch to conditionally overwrite the default, the actor type answers directly what its chained children should carry. Injection funnel collapses to: Type: c.opts.ParentActorType.ChildRunningContextType() The default (RunningContextType_EXECUTION, mapped server-side to actor.type = testworkflow / the "Workflow" chip on the Executions page) lives inside the method along with the special cases, so extending it means editing one file next to the type itself. Tests moved from the controlplane client into the same package as the method for the same reason. * chore: drop verbose doc block on ChildRunningContextType * feat(runner): propagate parent RunningContext to child pods via GetExecutionWorkflow The runner builds ExecutionConfig for a scheduled pod from the ExecutionStart proto, which does not carry RunningContext. That leaves cfg.Execution.RunningContext nil in the toolkit; parentActorTypeFromRunningContext returns empty; the sticky family cannot see it's chained under a gitintegration parent and children fall back to actor.type=testworkflow. The runner already makes a follow-up GetExecutionWorkflow call for every ExecutionStart to fetch the enriched workflow. That response is documented as the vehicle for "any additional information that may be related" - the perfect place to attach the parent's actor identity without touching the ExecutionStart proto (which is used by many non-sticky paths). Adds a minimal ExecutionRunningContext projection (actor type + actor name) to GetExecutionWorkflowResponse, populates it server-side from the stored TWE, and reconstructs a TestWorkflowRunningContext on the runner side that parentActorTypeFromRunningContext can read unchanged. When the field is nil (older control planes, non-actor executions), behaviour matches today.
| Commit: | 0999968 | |
|---|---|---|
| Author: | Aurelio Buarque | |
| Committer: | GitHub | |
feat(proto): add GITINTEGRATION actor type for git-provider-triggered flows (#8108) * feat(proto): add GITINTEGRATION actor type for git-provider-triggered flows Adds a new value to the RunningContextType proto enum and to the TestWorkflowRunningContextActorType OpenAPI enum, plus the round-trip mapper cases between them. Downstream consumers (telemetry bucketer, CDEvents mapper) route the new actor into git-integration/event buckets. Kicks off the change needed on the cloud-api side to introduce a provider-agnostic "Git Integration" actor for Quality Loop executions (GitHub today, GitLab and Bitbucket without another proto bump). * fix: allow gitintegration actor in CRD schema and CLI validation * refactor: rename GITINTEGRATION actor to QUALITYLOOP for internal naming * refactor: keep proto RunningContextType_QUALITYLOOP but expose gitintegration to users Proto stays QUALITYLOOP internally per Ole's suggestion. OpenAPI actor type, CRD kubebuilder enum, CLI flag, and telemetry bucket surface as gitintegration so every customer-facing surface (REST API, kubectl testkube, helm charts, generated CRDs) speaks the provider-agnostic name. The mapper cross-translates between the two. * refactor: alias QUALITYLOOP for internal Go references, wire stays gitintegration Adds regen-safe alias files so Go code can reference testkube.QUALITYLOOP_TestWorkflowRunningContextActorType while the underlying wire value stays "gitintegration" for CLI, REST, CRDs, and helm charts. Consumers refactored to use the internal name; customer surfaces unchanged. * chore: fix goimports alignment
| Commit: | ca19e44 | |
|---|---|---|
| Author: | buarki | |
feat(runner): propagate parent RunningContext to child pods via GetExecutionWorkflow The runner builds ExecutionConfig for a scheduled pod from the ExecutionStart proto, which does not carry RunningContext. That leaves cfg.Execution.RunningContext nil in the toolkit; parentActorTypeFromRunningContext returns empty; the sticky family cannot see it's chained under a gitintegration parent and children fall back to actor.type=testworkflow. The runner already makes a follow-up GetExecutionWorkflow call for every ExecutionStart to fetch the enriched workflow. That response is documented as the vehicle for "any additional information that may be related" - the perfect place to attach the parent's actor identity without touching the ExecutionStart proto (which is used by many non-sticky paths). Adds a minimal ExecutionRunningContext projection (actor type + actor name) to GetExecutionWorkflowResponse, populates it server-side from the stored TWE, and reconstructs a TestWorkflowRunningContext on the runner side that parentActorTypeFromRunningContext can read unchanged. When the field is nil (older control planes, non-actor executions), behaviour matches today.
| Commit: | 07c820a | |
|---|---|---|
| Author: | buarki | |
refactor: rename GITINTEGRATION actor to QUALITYLOOP for internal naming
| Commit: | eab10df | |
|---|---|---|
| Author: | buarki | |
feat(proto): add GITINTEGRATION actor type for git-provider-triggered flows Adds a new value to the RunningContextType proto enum and to the TestWorkflowRunningContextActorType OpenAPI enum, plus the round-trip mapper cases between them. Downstream consumers (telemetry bucketer, CDEvents mapper) route the new actor into git-integration/event buckets. Kicks off the change needed on the cloud-api side to introduce a provider-agnostic "Git Integration" actor for Quality Loop executions (GitHub today, GitLab and Bitbucket without another proto bump).
| Commit: | 8c19375 | |
|---|---|---|
| Author: | Caio Medeiros Pinto | |
| Committer: | Caio Medeiros Pinto | |
feat: add targeted execution scheduler policy
| Commit: | 11b7054 | |
|---|---|---|
| Author: | Vladislav Sukhin | |
feat: exchange data between test workflows run as a suite Workflows composed into a suite through `execute.workflows` could not pass anything back to the parent: the toolkit polled the child execution and threw away everything but the status. Values and files had to travel through external state. A child now publishes with the same mechanism steps already use - writing to /testkube/outputs - and the parent reads it back with the execution() expression: execute: workflows: - name: producer as: p fetch: - paths: ['results/**'] to: /data/from-producer ... shell: echo '{{ execution("p").outputs.token }}' The same function resolves "parent" from the execution ancestry, so a child can read what scheduled it, and sibling exchange falls out of feeding one child's output into the next child's config. - step outputs are promoted to the execution record, so they cross the pod - pkg/executiondata owns the registry, the expression functions and the artifact transfer, modelled on the existing credential() machine - `as` gives an entry a stable reference; two entries claiming the same one is an error rather than an ambiguous winner - execute specs are finalized when their operation starts, not up-front, so a later entry can read an earlier one - read_artifact() returns small files inline (1 MiB cap); `fetch` writes larger payloads to disk Files need read access to artifact storage, which the control plane did not grant: ListExecutionArtifactsPresigned is new in proto/service.proto and implemented for OSS. The Enterprise control plane needs the same RPC before read_artifact() and fetch work there; values work everywhere today. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
| Commit: | ab67752 | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | Dejan Zele Pejchev | |
fix: start cluster-inventory CRD watcher only on listener-capable agents
| Commit: | a382d48 | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | Dejan Zele Pejchev | |
feat(triggers): schema-aware match conditions (#7566) * feat(agent): add GET /v1/cluster-resources for UI autodiscovery * feat(proto): AgentInventoryService with PutClusterResources * feat(triggers): schema-aware match conditions WIP * fix(triggers): propagate listenerAgentIds through mappers/scraper, release informers by old GVK on unpin, gzip inventory push, validate change-operators on empty event * feat(triggers): pin listeners via a Target selector (listener.match.id) instead of listenerAgentIds * refactor(triggers): extract ValidateMatchConditions with typed reason codes so cp-api can reuse match validation
| Commit: | 0ed76c5 | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | GitHub | |
feat(triggers): schema-aware match conditions (#7566) * feat(agent): add GET /v1/cluster-resources for UI autodiscovery * feat(proto): AgentInventoryService with PutClusterResources * feat(triggers): schema-aware match conditions WIP * fix(triggers): propagate listenerAgentIds through mappers/scraper, release informers by old GVK on unpin, gzip inventory push, validate change-operators on empty event * feat(triggers): pin listeners via a Target selector (listener.match.id) instead of listenerAgentIds * refactor(triggers): extract ValidateMatchConditions with typed reason codes so cp-api can reuse match validation
| Commit: | d18cda3 | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | Dejan Zele Pejchev | |
feat(proto): AgentInventoryService with PutClusterResources
| Commit: | c95189e | |
|---|---|---|
| Author: | Caio Medeiros Pinto | |
| Committer: | GitHub | |
fix(grpc): propagate execution tags in execution start updates (#7833) * fix(testworkflows): pass execution tags into expression machine * fix(grpc): propagate execution tags in start updates * chore(lint): apply goimports formatting
| Commit: | 0687d48 | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | Dejan Zele Pejchev | |
feat(proto): AgentInventoryService with PutClusterResources
| Commit: | afa2611 | |
|---|---|---|
| Author: | Caio Medeiros Pinto | |
| Committer: | Caio Medeiros Pinto | |
feat(agentserver): send registration metadata on reconnect (TKC-5876) (#7786) * feat(agentserver): send registration metadata on reconnect Carry labels, runner_group and is_global through the existing UpdateAgentCapabilitiesOnStartup RPC so runner-controlled fields stay fresh after the initial Register. Gated by a new opt-in flag so old servers stay untouched. - proto/service.proto: extend UpdateAgentCapabilitiesOnStartupRequest with labels, runner_group, is_global, update_registration_metadata. - cmd/api-server/main.go: extract collectRunnerRegistrationLabels and use it in both the Register and reconnect call sites. Refs: TKC-5876 * fix(agentserver): skip metadata refresh on failed deployment label read Address Greptile P1+P2: - collectRunnerRegistrationLabels and getDeploymentLabels now return an error so a Kubernetes Deployment lookup failure is observable. - In the reconnect path, when the lookup fails we skip the metadata update (UpdateRegistrationMetadata=false) so a transient K8s API or RBAC failure cannot replace the control plane's existing labels with a near-empty self-registration-only set. - Reconnect-path label fetch now uses updateCtx so a stalled K8s GET cannot block startup past the 5s control-plane RPC budget. - Initial Register path falls back to {registration:self} on lookup failure (preserves prior behavior; first registration has no prior labels to clobber). Refs: TKC-5876 * fix(agentserver): decouple labels from runner_policy in startup metadata refresh Split UpdateRegistrationMetadata into independent UpdateLabels and UpdateRunnerPolicy flags so a failed Deployment label lookup no longer suppresses runner_group / is_global propagation. Those values come from runner config and do not depend on the Kubernetes API. Addresses Greptile P1 review on PR #7786.
| Commit: | a203c04 | |
|---|---|---|
| Author: | Caio Medeiros Pinto | |
| Committer: | GitHub | |
feat(agentserver): send registration metadata on reconnect (TKC-5876) (#7786) * feat(agentserver): send registration metadata on reconnect Carry labels, runner_group and is_global through the existing UpdateAgentCapabilitiesOnStartup RPC so runner-controlled fields stay fresh after the initial Register. Gated by a new opt-in flag so old servers stay untouched. - proto/service.proto: extend UpdateAgentCapabilitiesOnStartupRequest with labels, runner_group, is_global, update_registration_metadata. - cmd/api-server/main.go: extract collectRunnerRegistrationLabels and use it in both the Register and reconnect call sites. Refs: TKC-5876 * fix(agentserver): skip metadata refresh on failed deployment label read Address Greptile P1+P2: - collectRunnerRegistrationLabels and getDeploymentLabels now return an error so a Kubernetes Deployment lookup failure is observable. - In the reconnect path, when the lookup fails we skip the metadata update (UpdateRegistrationMetadata=false) so a transient K8s API or RBAC failure cannot replace the control plane's existing labels with a near-empty self-registration-only set. - Reconnect-path label fetch now uses updateCtx so a stalled K8s GET cannot block startup past the 5s control-plane RPC budget. - Initial Register path falls back to {registration:self} on lookup failure (preserves prior behavior; first registration has no prior labels to clobber). Refs: TKC-5876 * fix(agentserver): decouple labels from runner_policy in startup metadata refresh Split UpdateRegistrationMetadata into independent UpdateLabels and UpdateRunnerPolicy flags so a failed Deployment label lookup no longer suppresses runner_group / is_global propagation. Those values come from runner config and do not depend on the Kubernetes API. Addresses Greptile P1 review on PR #7786.
| Commit: | b45ddfc | |
|---|---|---|
| Author: | Mark Gascoyne | |
| Committer: | GitHub | |
feat: [TKC-5638] add RunspaceBridgeService proto definition (#7647) Defines the bidirectional gRPC stream between the runspace bridge (Go, in cloud-api) and the AI service (Node.js), including session negotiation, workspace sync events, and file I/O RPC types.
| Commit: | 8fff975 | |
|---|---|---|
| Author: | Mark Gascoyne | |
feat: [TKC-5719] add FileChanged and Checkpoint to runspace bridge proto Adds FileChanged message (path, change_type enum, content bytes, seq) and Checkpoint marker to BridgeMessage oneof. Regenerated pb.go via make generate-protobuf.
| Commit: | a5a9190 | |
|---|---|---|
| Author: | Mark Gascoyne | |
feat: [TKC-5638] add RunspaceBridgeService proto definition Defines the bidirectional gRPC stream between the runspace bridge (Go, in cloud-api) and the AI service (Node.js), including session negotiation, workspace sync events, and file I/O RPC types.
| Commit: | 4ff5c98 | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | GitHub | |
feat: [TKC-5460] WorkflowTrigger CLI, OSS API, agent client (k8s + cloud), and cloud-watch (#7535) Signed-off-by: Dejan Zele Pejchev <pejcev.dejan@gmail.com>
| Commit: | f71bfe2 | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | GitHub | |
feat: [TKC-5460] WorkflowTrigger CRD types, field matching engine, internal trigger type (#7523) * feat: [TKC-5460] WorkflowTrigger v2 types, field matcher, and executor integration Signed-off-by: Dejan Zele Pejchev <pejcev.dejan@gmail.com> * feat: [TKC-5405] add dynamic informers + resourceRef for custom resource support (#7528) Signed-off-by: Dejan Zele Pejchev <pejcev.dejan@gmail.com> --------- Signed-off-by: Dejan Zele Pejchev <pejcev.dejan@gmail.com>
| Commit: | 9381b39 | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | Dejan Zele Pejchev | |
feat: [TKC-5460] cloud-watch + proto + controlplaneclient for WorkflowTrigger Signed-off-by: Dejan Zele Pejchev <pejcev.dejan@gmail.com>
| Commit: | d76ee5a | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
feat: [TKC-5460] WorkflowTrigger v2 types, field matcher, and executor integration Signed-off-by: Dejan Zele Pejchev <pejcev.dejan@gmail.com>
| Commit: | cfc81ab | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | Dejan Zele Pejchev | |
feat: [TKC-5460] add WorkflowTrigger CRD types, field matching engine, and internal trigger type Signed-off-by: Dejan Zele Pejchev <pejcev.dejan@gmail.com>
| Commit: | 1f3705a | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | Dejan Zele Pejchev | |
feat: [TKC-5460] add WorkflowTrigger CRD types, field matching engine, and internal trigger type Signed-off-by: Dejan Zele Pejchev <pejcev.dejan@gmail.com>
| Commit: | 9b7bef2 | |
|---|---|---|
| Author: | Povilas Versockas | |
| Committer: | GitHub | |
fix: [TKC-5412] fix workflow log streaming (#7488) * feat: [TKC-5412] add workflow stream resume protocol * chore: [TKC-5412] regenerate workflow stream protocol types * feat: [TKC-5412] make agent log streams resilient * feat: [TKC-5412] resume workflow logs in the CLI * fix: [TKC-5412] remove unused workflow stream test assignment * fix: [TKC-5412] harden notification stream manager * fix up * fix: [TKC-5412] preserve workflow log replay stream identity
| Commit: | 52efe1c | |
|---|---|---|
| Author: | Mark Gascoyne | |
| Committer: | GitHub | |
feat: [TKC-5159] add targetSelector to ListWebhooksV2 for webhook target filtering (#7462) * feat: [TKC-5159] add targetSelector field to ListWebhooksV2Request proto Add map<string, string> targetSelector (field 7) to ListWebhooksV2Request, allowing callers to opt-in to server-side webhook target filtering by passing their own labels. * feat: [TKC-5159] pass agent labels as targetSelector in ListWebhooksV2 Wire agent labels (including synthetic id/name) through CloudWebhookClient as a TargetSelector so the control plane can filter webhooks by target. Filtering is opt-in: agents that do not send a TargetSelector receive all webhooks (backwards compatible). * fix: [TKC-5159] guard against empty agent ID/name in target selector Only include synthetic "id" and "name" keys when non-empty, so a misconfigured agent omits them from the selector (less restrictive) rather than sending empty strings that match nothing.
| Commit: | 3f76026 | |
|---|---|---|
| Author: | Povilas Versockas | |
| Committer: | Povilas Versockas | |
feat: [TKC-5104] sync startup capabilities with control plane (#7175)
| Commit: | 920a45e | |
|---|---|---|
| Author: | Povilas Versockas | |
| Committer: | GitHub | |
feat: [TKC-5104] sync startup capabilities with control plane (#7175)
| Commit: | ea4a649 | |
|---|---|---|
| Author: | Copilot | |
| Committer: | Vladislav Sukhin | |
feat: Add runnerId field to TestWorkflowServiceNotificationsRequest and TestWorkflowParallelStepNotificationsRequest proto messages (#7088) * Initial plan * feat: Add runnerId field to TestWorkflowServiceNotificationsRequest and TestWorkflowParallelStepNotificationsRequest proto messages Co-authored-by: vsukhin <5984962+vsukhin@users.noreply.github.com> --------- Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> Co-authored-by: vsukhin <5984962+vsukhin@users.noreply.github.com>
| Commit: | 1147694 | |
|---|---|---|
| Author: | Copilot | |
| Committer: | GitHub | |
feat: Add runnerId field to TestWorkflowServiceNotificationsRequest and TestWorkflowParallelStepNotificationsRequest proto messages (#7088) * Initial plan * feat: Add runnerId field to TestWorkflowServiceNotificationsRequest and TestWorkflowParallelStepNotificationsRequest proto messages Co-authored-by: vsukhin <5984962+vsukhin@users.noreply.github.com> --------- Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> Co-authored-by: vsukhin <5984962+vsukhin@users.noreply.github.com>
| Commit: | 33ec7b3 | |
|---|---|---|
| Author: | Povilas Versockas | |
| Committer: | GitHub | |
feat: [TKC-4630] add support for webhook capability (#6906)
| Commit: | 367d756 | |
|---|---|---|
| Author: | Alisdair MacLeod | |
| Committer: | GitHub | |
feat: add super agent rollback (#6908) Allow the ability for agents to rollback to become super agents. This is going to cause any data that was added to the Control Plane to become inaccessible, at least until the agent migrates back, but this is a disaster recovery rollback mode anyway.
| Commit: | 93f77ad | |
|---|---|---|
| Author: | Alisdair MacLeod | |
| Committer: | GitHub | |
feat: super agent migration (#6895) * feat: add super agent migration endpoint and return current super-agent-ness with procontext * feat: handle procontext super agent status response field * feat: create an override option for super agent migration * feat: implement super agent migration on client side * feat: write to kubernetes termination log if unable to migrate * chore: move super agent migration to a separate file
| Commit: | c63fafb | |
|---|---|---|
| Author: | Wito | |
| Committer: | GitHub | |
chore: prepare source of truth migration (#6883) * chore: add source of truth capability * chore: deprecate agent.type * chore: generate protobuf
| Commit: | ff22081 | |
|---|---|---|
| Author: | Caio Medeiros Pinto | |
| Committer: | Caio Medeiros Pinto | |
feat: auto registering agent mode labels (#6841) * feat: add runner mode and labels when auto-registering through grpc method * fix: cloud interface new properties get methods * fix: linting alerts * chore: force global mode for super agent * fix: proto of agent request new fields * chore: add agent labels to agent pod as well * fix: debug agent labels to be configured * chore: remove forced debug mode for runner * fix: remove unneeded imports
| Commit: | bdd456b | |
|---|---|---|
| Author: | ed382 | |
| Committer: | GitHub | |
feat(TKC-4532): return runner name after registration (#6856) For legacy super agents the runner name is overriden to the environment ID.
| Commit: | 9fe76b1 | |
|---|---|---|
| Author: | Caio Medeiros Pinto | |
| Committer: | GitHub | |
feat: auto registering agent mode labels (#6841) * feat: add runner mode and labels when auto-registering through grpc method * fix: cloud interface new properties get methods * fix: linting alerts * chore: force global mode for super agent * fix: proto of agent request new fields * chore: add agent labels to agent pod as well * fix: debug agent labels to be configured * chore: remove forced debug mode for runner * fix: remove unneeded imports
| Commit: | ab28bd5 | |
|---|---|---|
| Author: | Alisdair MacLeod | |
| Committer: | GitHub | |
chore: deprecate unused RPCs and remove related code (#6833)
| Commit: | 83ffc1d | |
|---|---|---|
| Author: | ed382 | |
| Committer: | GitHub | |
feat: webhooks to control plane (#6757) * notes * notes * refactor: webhook loader with opts * refactor: webhook listener with opts * fix: check during become queries * notes * fix: handle not set metrics * fix: handle secret client not set * notes * fix: make webhooks repository optional * feat: make webhook template client optional * refactor: mark deprecated * refactor: rename repository for consistency * refactor: reduce the webhook client interface * refactor: unnecessary predicate * notes * notes * notes * chore: remove unused * notes * refactor: emitter reduce exported fields/methods * notes * notes * refactor: matching of events logic moved to listeners * refactor: single subscription per emitter * notes * refactor: make procontext optional * notes * refactor: make event emitter subject root configurable * notes * notes * feat: webhooks capabalities * feat: disable agent-based webhooks when cloud-based enabled * feat: registration updates and super agent registration * fix: pass procontext to deprecated system for auth * fix: resource and resource id in payloads * feat: event emitter leases * fix: switch to go tool * fix: unit tests * chore: fix lint * docs: clean up todos * chore: integration tests * fix: make lease name unique per release * chore: clean up * chore: clean up * chore: clean up * refactor: simplify mutex usage for listeners slice/map * fix: access to listeners and simplify notify call * fix: lease backend test
| Commit: | 2ccb7f0 | |
|---|---|---|
| Author: | Alisdair MacLeod | |
| Committer: | GitHub | |
feat: workflow retrieval rpc (#6828) This replaces passing workflows directly in the execution start message and provides a new rpc to retrieve a "resolved" workflow directly.
| Commit: | 0c7c402 | |
|---|---|---|
| Author: | Alisdair MacLeod | |
| Committer: | GitHub | |
feat: resolved workflow rpc (#6823) * feat: pass full workflow when starting execution * feat: use shared execution in sync rpc This is a breaking change but this is ok because the sync feature is unreleased currently. * chore: gen proto * fix: sync client use new proto definition * feat: runner client make use of full workflow
| Commit: | fb8c7e2 | |
|---|---|---|
| Author: | Alisdair MacLeod | |
| Committer: | GitHub | |
feat: execution start error reporting (#6816) * feat: add rpc for reporting executions startup errors * feat: implement startup error rpc
| Commit: | f25855a | |
|---|---|---|
| Author: | Alisdair MacLeod | |
| Committer: | Alisdair MacLeod | |
feat: add rpc for reporting executions startup errors
| Commit: | 48a930f | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | Dejan Zele Pejchev | |
fix: issues around credential expressions (#6799) Signed-off-by: Dejan Zele Pejchev <pejcev.dejan@gmail.com>
| Commit: | 7fed18c | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | GitHub | |
fix: issues around credential expressions (#6799) Signed-off-by: Dejan Zele Pejchev <pejcev.dejan@gmail.com>
| Commit: | 09a2578 | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
| Committer: | Dejan Zele Pejchev | |
fix: issues around credential expressions Signed-off-by: Dejan Zele Pejchev <pejcev.dejan@gmail.com>
| Commit: | 0fe1644 | |
|---|---|---|
| Author: | Kubeshop | |
Merge remote-tracking branch 'origin' into tkc-4250-execute-webhooks-within-control-plane
| Commit: | 095258b | |
|---|---|---|
| Author: | Dejan Zele Pejchev | |
fix: issues around credential expressions Signed-off-by: Dejan Zele Pejchev <pejcev.dejan@gmail.com>
| Commit: | 5417927 | |
|---|---|---|
| Author: | Kubeshop | |
| Committer: | Kubeshop | |
feat: registration updates and super agent registration
| Commit: | 6198d64 | |
|---|---|---|
| Author: | Alisdair MacLeod | |
| Committer: | GitHub | |
feat: Add RPC definitions for sync functionality (#6756) Whilst I hate that these RPCs use `bytes` to carry structured data, my attempts to generate or manually copy these data structures into proto definitions were fruitless so until we have some sort of CRD generation that can create protobuf definitions for us this is the best we can hope for. I did try `go-to-protobuf` which is what the Kubernetes project uses, and whilst it was mostly fine it would have required some wider changes to the repository to make it work nicely with the other protobuf changes that we have here, so maybe in the future for a v2 of these RPCs we can implement them, but for now this is the best we can hope for.
| Commit: | e2615c4 | |
|---|---|---|
| Author: | Kubeshop | |
feat: webhooks capabalities