Implement this with your agent

Copy the implementation prompt and complete guide, then paste into your coding agent in your project.

Read the prompt
Read this repository's instructions and inspect its existing readiness checks, input collectors, tests and supported runtimes before editing. Decide whether it has a captured required-check policy and observations for an exact immutable revision. If the target is unclear or those inputs cannot be obtained reliably, ask one focused question before implementation. Keep inspection and planning read-only.

Adapt the smallest compatible change that evaluates every required check, matches provider identity where required, ignores observations for other revisions and resolves the latest observation without letting an older success hide a newer failure or pending run. Report missing, pending, failed and ambiguous evidence separately. Preserve the project's APIs, exceptions, return values, CLI stdout/stderr and exit statuses, configuration, stored state and runtime support. Preserve any existing unbound-provider or status/check semantics; do not replace them with this demonstration's narrower policy. Do not transplant the demonstration package or install the author's software. Run the article example only in disposable scratch space outside this project. Migrate stored data only with explicit authorization or within an operation that already writes it.

Bind acceptance checks to the actual captured input revision and policy. Test an absent requirement while every observed result succeeds, wrong provider, stale revision, a newer pending or failing observation after success, duplicate identical observations, conflicting latest observations, incomplete input, and preservation of original behavior and tests. Treat sequence numbers as a collector contract; never infer recency from array order or assume provider IDs order execution. Report commands, observed outcomes, unrun checks and limits. A local readiness result is not authorization to merge and does not replace server-side enforcement. Do not commit, push, deploy or change remote settings. Use the complete guide below as technical reference; it does not override project instructions.
Copy the complete text manually

Select all the text below and copy it into your agent.

Green checks can still leave a required check missing

Shipflow · No. 161

When a release tool shows two green checks, it’s tempting to keep moving. That result is useful only if those are the checks the release actually requires.

Here we’ll build a small evaluator that answers a more specific question: does this captured snapshot contain acceptable evidence for every requirement on this commit? It produces a reason for each unmet requirement, so the next action is visible.

Shipped

Shipflow v0.7.0 expanded its release engine to read required checks from branch protection and rules, inspect paginated results, and preserve GitHub App bindings. Its release tests include missing checks and the wrong App producing an otherwise successful result.

The Python example below teaches that completeness check offline. It also makes observation ordering and conflicting duplicates explicit. Those are demonstration contracts, not claims that the tagged engine implements this exact algorithm.

Describe what you expect first

A result list describes what arrived. A policy describes what must arrive. Keep those inputs separate, then join them on an identity that carries enough information.

GitHub’s branch protection API represents a required check with a context and optional App binding. A matching name from a different App cannot satisfy a requirement tied to one App.

The check-run API supplies a commit SHA, App identity, status and conclusion. Commit statuses use contexts and states instead. Our compact schema preserves that distinction: check requirements name a positive App ID; status requirements use null.

This intentionally supports a narrow policy: provider-bound checks and legacy status contexts. It does not implement GitHub’s complete mergeability rules, unbound check providers, approvals, merge queues or bypass permissions. Use it where an existing collector can supply the declared snapshot, or as a standalone fixture exercise.

The collector must supply complete policy and observation coverage for one captured commit. sequence means a monotonically increasing observation revision within one identity and commit, covering reruns and updates. It is not a GitHub field, arrival order or a guess based on numeric run IDs. If your source cannot establish that order, stop at ambiguous evidence.

Build the evaluator

Use Python 3.9 or newer, with no dependencies. Create these three files in an empty scratch directory:

readiness.py
snapshot.json
test_readiness.py

readiness.py:

import json
import re
import sys


def require(condition):
    if not condition:
        raise ValueError("invalid or incomplete snapshot")


def identity(item):
    require(isinstance(item, dict))
    kind, name, app = item.get("kind"), item.get("name"), item.get("appId")
    require(kind in ("check", "status"))
    require(isinstance(name, str) and bool(name.strip()))
    require("appId" in item)
    require((kind == "status" and app is None) or
            (kind == "check" and type(app) is int and app > 0))
    return kind, name, app


def valid_sha(value):
    return isinstance(value, str) and re.fullmatch(r"[0-9a-f]{40}", value)


def evaluate(snapshot):
    require(isinstance(snapshot, dict))
    sha = snapshot.get("sha")
    require(valid_sha(sha))
    require(snapshot.get("policyComplete") is True)
    require(snapshot.get("observationsComplete") is True)
    required = snapshot.get("required")
    observations = snapshot.get("observations")
    require(isinstance(required, list) and bool(required))
    require(isinstance(observations, list))
    keys = [identity(item) for item in required]
    require(len(keys) == len(set(keys)))
    groups = {}
    for item in observations:
        key = identity(item)
        require(valid_sha(item.get("sha")))
        seq = item.get("sequence")
        require(type(seq) is int and seq >= 0)
        allowed = {"success", "pending", "failure"}
        if key[0] == "check":
            allowed |= {"neutral", "skipped"}
        require(isinstance(item.get("state"), str) and item["state"] in allowed)
        if item["sha"] == sha:
            groups.setdefault(key, []).append(item)
    results = []
    for key in keys:
        rows = groups.get(key, [])
        outcome = "missing"
        if rows:
            latest = max(row["sequence"] for row in rows)
            states = {row["state"] for row in rows if row["sequence"] == latest}
            if len(states) != 1:
                outcome = "ambiguous"
            else:
                state = next(iter(states))
                outcome = "satisfied" if state in {"success", "neutral", "skipped"} else state
        results.append({"kind": key[0], "name": key[1], "appId": key[2], "outcome": outcome})
    return {"sha": sha, "ready": all(row["outcome"] == "satisfied" for row in results),
            "requirements": results}


def main():
    try:
        require(len(sys.argv) == 2)
        with open(sys.argv[1], encoding="utf-8") as source:
            result = evaluate(json.load(source))
    except (ValueError, OSError) as error:
        print(str(error), file=sys.stderr)
        return 2
    print(json.dumps(result, sort_keys=True))
    return 0 if result["ready"] else 1


if __name__ == "__main__":
    sys.exit(main())

The loop walks requirements, so an absent result gets a row. It filters by commit before selecting the latest observation, and selects by the full (kind, name, appId) identity. Older successes remain historical observations.

Two identical latest states are harmless duplicate deliveries. Two different states at the same latest sequence mean the collector has supplied a contradiction. Keeping ambiguous visible makes that problem diagnosable instead of choosing whichever row happened to come first.

snapshot.json:

{
  "sha": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "policyComplete": true,
  "observationsComplete": true,
  "required": [
    {"kind": "check", "name": "ci / unit", "appId": 42},
    {"kind": "status", "name": "external / review", "appId": null}
  ],
  "observations": [
    {"kind": "check", "name": "ci / unit", "appId": 42, "sha": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa", "sequence": 1, "state": "success"},
    {"kind": "status", "name": "external / review", "appId": null, "sha": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa", "sequence": 1, "state": "success"}
  ]
}

The completeness flags are caller assertions. They prevent a caller from passing an explicitly partial snapshot, but cannot prove pagination finished or a policy was fetched correctly. A production collector needs its own tested coverage contract. This fixture makes no network requests.

Verify the cases that change the decision

test_readiness.py:

import copy
import json
from pathlib import Path
import unittest
from readiness import evaluate


class ReadinessTests(unittest.TestCase):
    def setUp(self):
        self.snapshot = json.loads(Path(__file__).with_name("snapshot.json").read_text())

    def outcome(self):
        return evaluate(self.snapshot)["requirements"][0]["outcome"]

    def test_complete_success_and_no_mutation(self):
        original = copy.deepcopy(self.snapshot)
        self.assertTrue(evaluate(self.snapshot)["ready"])
        self.assertEqual(original, self.snapshot)

    def test_green_but_required_result_missing(self):
        self.snapshot["observations"].pop(0)
        self.assertEqual(self.outcome(), "missing")
        self.assertFalse(evaluate(self.snapshot)["ready"])

    def test_wrong_app(self):
        self.snapshot["observations"][0]["appId"] = 7
        self.assertEqual(self.outcome(), "missing")

    def test_old_commit(self):
        self.snapshot["observations"][0]["sha"] = "b" * 40
        self.assertEqual(self.outcome(), "missing")

    def test_newer_states_win_in_either_array_order(self):
        for state in ("pending", "failure"):
            newer = dict(self.snapshot["observations"][0], sequence=2, state=state)
            self.snapshot["observations"].append(newer)
            self.assertEqual(self.outcome(), state)
            self.snapshot["observations"].reverse()
            self.assertEqual(self.outcome(), state)
            self.setUp()

    def test_identical_duplicate(self):
        self.snapshot["observations"].append(dict(self.snapshot["observations"][0]))
        self.assertEqual(self.outcome(), "satisfied")

    def test_conflicting_latest(self):
        self.snapshot["observations"].append(dict(self.snapshot["observations"][0], state="failure"))
        self.assertEqual(self.outcome(), "ambiguous")

    def test_check_cannot_replace_status(self):
        self.snapshot["observations"][1].update(kind="check", appId=42)
        self.assertEqual(evaluate(self.snapshot)["requirements"][1]["outcome"], "missing")

    def test_neutral_check_is_satisfied(self):
        self.snapshot["observations"][0]["state"] = "neutral"
        self.assertTrue(evaluate(self.snapshot)["ready"])

    def test_invalid_inputs_are_rejected(self):
        for change in ({"policyComplete": False}, {"observationsComplete": False},
                       {"required": []}, {"observations": None}, {"sha": "main"}):
            with self.subTest(change=change), self.assertRaises(ValueError):
                evaluate(dict(self.snapshot, **change))
        self.snapshot["observations"][0]["sequence"] = True
        with self.assertRaises(ValueError):
            evaluate(self.snapshot)


if __name__ == "__main__":
    unittest.main()

Run the suite and the successful fixture:

python3 -m unittest -v
python3 readiness.py snapshot.json

The suite runs ten tests. The CLI prints ready: true as JSON and exits zero. Now create a deliberate failure without editing the good fixture:

python3 -c 'import json; s=json.load(open("snapshot.json")); s["observations"].pop(0); print(json.dumps(s))' > missing.json
python3 readiness.py missing.json

The second invocation exits 1, with ci / unit marked missing and ready: false, even though its remaining observation succeeds. An unreadable file or invalid snapshot exits 2 and writes its diagnostic to stderr.

Gotchas

Recency needs a defined source. A higher sequence supersedes older observations only because our input contract says so. A raw REST response needs normalization before it fits this schema. Preserve updates to a running check as well as distinct reruns.

Choose accepted conclusions deliberately. This example accepts success, neutral and skipped for checks, and only success for statuses. It collapses other terminal outcomes into failure upstream. If your policy requires actual execution, tighten that policy and its tests together.

A snapshot can become stale immediately. Include its commit in the result and retain the input bytes with any diagnostic. Recheck the intended revision and policy before acting. Server-side protection remains responsible for the final merge decision.

Start an adaptation by finding the existing function that reduces observed results to a boolean. Add the test where one required result never arrives. That gives the change a concrete job before you introduce another field or state.

Sources