Skip to content

Feature Request: trigger and read external CI/CD pipeline runs from a session #2786

Description

@gagarwal

Problem

An agent can change code, but it cannot close the loop on whether that change is actually good. Validation for most non-trivial repositories does not happen in the session's own environment: it happens in a CI/CD system that owns the build matrix, the signing identity, the device farm, the integration environment, and the historical baseline for what "passing" means.

A session today has no supported way to reach that system. It can produce a change and open a pull request, but it cannot start the validation run, watch it, read the result, or act on a failure. The outcome is a change that looks complete but is unverified, and a human has to pick up the loop manually.

This is distinct from running tests locally in the session environment. Some validation genuinely cannot be reproduced in a hosted session — multi-platform builds, signed artifacts, hardware-backed tests, or suites that need access to a provisioned integration environment. For those, delegating to the CI/CD system is not a workaround; it is the correct design.

What is missing

  • No way to trigger a pipeline or workflow run in an external CI/CD system from a session.
  • No way to poll or subscribe to the status of a run that the session started.
  • No way to retrieve logs, test results, or published artifacts from a completed run.
  • No documented pattern for an agent to act on a validation failure — read the failing step, correlate it to its own change, and iterate.
  • No credential model for a session to authenticate to a CI/CD system under appropriate authorization.

Proposed behavior

  • A session can trigger a run in a configured CI/CD system, supplying a ref and parameters.
  • The session can observe run status, and can suspend while waiting for a long run rather than holding execution open.
  • On completion, the session can read the outcome, per-step logs, structured test results, and artifact references.
  • Failures are actionable: the agent can identify the failing step and its output, and iterate on its change.
  • Authentication uses a caller-supplied or delegated credential, scoped to the specific pipelines involved, and never exposed in prompts or model context.
  • Systems that are not configured produce a clear capability error rather than an undifferentiated failure.

Example scenario

An agent fixes a defect in a library that ships to several platforms. The hosted session can compile and unit-test one of them, but the full matrix, the signed build, and the device tests only exist in the organization's CI/CD system. The agent pushes its branch, triggers the validation pipeline, suspends while it runs, then reads the result. One platform fails on a compatibility check; the agent reads that step's log, adjusts the change, and re-runs validation before requesting review.

Acceptance criteria

  • A session can trigger a run in a configured external CI/CD system.
  • Run status is observable, including across a suspension.
  • Logs, test results, and artifact references are retrievable on completion.
  • Credentials are scoped and not exposed to model context.
  • Unconfigured systems fail with a clear, documented capability error.

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions