Skip to content

git-workflow

Recipe card from the charly-internals plugin (Development — contributor internals).

This card has additional detail pages:

git-workflow — branch-per-change, PR-only org-workflow-validated landing

Section titled “git-workflow — branch-per-change, PR-only org-workflow-validated landing”

Every change to an OpenCharly repo lands through ONE discipline: a pull request gated by the ORG-WIDE charly/pr-validator GitHub Actions workflow (opencharly/.github), which validates and, on a PASS verdict, enables GitHub’s OWN native auto-merge (squash) inline — there is no separate auto-merge workflow; the CalVer tag and the CHANGELOG are written afterwards by the independent tag-on-merge workflow, triggered by the merge. A direct push to main is FORBIDDEN and mechanically disabled — ONE organization branch RULESET blocks it. On GitHub Team that ONE ruleset carries both the branch rules (creation/deletion/non_fast_forward + a strict required validate / validate status check) and the workflows rule (“Require workflows to pass”) naming opencharly/.github/.github/workflows/org-wide-pr-validator-required.yml, so the validator is required ONCE org-wide with no per-repo stub to install. The legacy branch-protection API is NOT used and org-ruleset.sh removes it wherever it survives, because it has no bypass slot for the app that writes the CHANGELOG. The pre-push-gate adds a local backstop in every harness that wires it — including Claude Code, via .claude/settings.json’s PreToolUse hooks. The R10 pass authorizes OPENING the PR, never a self-merge: the two-step landing separates the author (who opens the PR) from the fresh validator (a check run on the same PR). This skill is the mechanics; the project rulebook “Post-Execution Policies” (AGENTS.md) carries the mandate, /charly-internals:cutover-policy the one-phase rule, /charly-build:migrate the schema-version/tag coupling, and the marketplace’s internals/agents/pr-validator.md the validator’s own spec.

THE BODY-BEFORE-PUSH RULE (the one that bites hardest)

Section titled “THE BODY-BEFORE-PUSH RULE (the one that bites hardest)”

Write the WHOLE PR body BEFORE the push. The validator validates the PR — its diff at the branch head, and its body — each time its run is triggered, and it fetches the PR LIVE by number. Getting the body final first is the cheapest order: a run that fires before the body is final reviews the old text and costs an extra cycle.

The org required workflow fires ONLY on the default push-driven types. MEASURED: the REQUIRED workflow (the org ruleset workflows rule) acts on opened / synchronize / reopened and IGNORES on.types for anything else — a body edit did NOT fire it even with edited listed, and a draft->ready transition did NOT fire it even with ready_for_review listed. A PLAIN per-repo workflow with the same types: DOES fire (verified live), but the org no longer ships one. So do NOT rely on a body edit to re-run the required gate.

The head SHA is known BEFORE the push, so this is always possible. A commit’s identity is content-addressed: git rev-parse HEAD (and git diff --stat origin/main...HEAD) give the exact SHA and diff-stats the push will publish, with nothing pushed yet. The correct order:

  1. Commit the SOURCE locally (do not push yet).
  2. Compute the head + diff-stats from the committed tree: git rev-parse HEAD, git log --oneline origin/main...HEAD, git diff --stat origin/main...HEAD.
  3. Write the WHOLE body (--body-file …) keyed to that real SHA — Summary, evidence with the real SHAs/diff-stats, rulebook section, attribution footer LAST.
  4. Push.

If you must move the head again (a real fix), the body is stale again — repeat 1–4 in ONE batch: edit the body, then push.

A body-only fix after a pushed head: re-run the gate MANUALLY with gh run rerun <run-id>; an empty re-freeze commit is the alternative. A body edit does NOT re-run the required workflow (the ruleset workflows rule fires only on opened/ synchronize/reopened and ignores types — MEASURED and confirmed by the GitHub docs, “Troubleshooting rules”), so the fix is applied by explicitly re-running the failed run: gh run rerun <run-id> (find it with gh run list --repo <r> --json databaseId,headSha,conclusion). A re-run reuses the SAME GITHUB_SHA and updates THAT run’s validate / validate check run IN PLACE — no duplicate same-name check run — so it clears the POISON state (a completed prior FAILURE on that head) and re-reads the corrected body. Because it preserves the head SHA it is the cheaper choice.

There is no automatic rerun-label channel any more. The org-wide rerun label plus scheduled sweep was RETIRED: a label can be added for ANY reason — including a comment — and the sweep re-ran the gate without a body change. gh run rerun <run-id> is the ONE body-fix path now. An empty re-freeze commit (git commit --allow-empty -m "docs: re-freeze …") ALSO re-fires the required gate: the ruleset workflows rule acts on the push-driven synchronize type, so the push mints a NEW head SHA and a fresh validate / validate run on it (the old head’s checks no longer apply). It is equally valid, just more expensive — the head moves, so re-key the body to the new SHA (repeat steps 1–4 above). Neither is ever REQUIRED: the cheapest correct move is to write the body before the push and need neither.

gh workflow run pr-validator.yml -f pr-number=<N> --ref <branch> is the manual alternative, and --ref is MANDATORY: a workflow_dispatch with no --ref runs on the DEFAULT branch, so its check registers on main and counts for nothing on the PR. Measured on opencharly/plugin-vm#39 (head c9457a9): the same dispatch without --ref produced no head check; carrying the branch ref – --ref feat/deploy-shape-override – it registered the green check on the branch head. But a gh workflow run mints a NEW run (a new same-name check run on the head) and can re-poison the rollup; prefer gh run rerun <run-id>.

There is NO self-heal. The reusable workflow sets a head_ref output but never consumes it (grep -c 'steps.pr.outputs' in opencharly/.github/.github/workflows/pr-validator.yml is 0); the org required workflow contains no re-dispatch step. A workflow_dispatch does NOT land its check on the PR head by itself — supply --ref.

The POISON state — green verdict, still BLOCKED

Section titled “The POISON state — green verdict, still BLOCKED”

A branch protection rule requires ALL same-name validate / validate check runs on the head to pass, so a COMPLETED earlier FAILURE keeps the PR BLOCKED even after a later SUCCESS of the same name — it reads like a verdict BLOCK but is not. Two runs on ONE head produce two such check runs; a gh workflow run re-dispatch on the same head mints another and can keep the PR stuck. The capability-free remedy is gh run rerun <run-id> on the failed run: a workflow re-run reuses the SAME GITHUB_SHA/GITHUB_REF and updates THAT run’s check run in place — no duplicate is minted (GitHub’s own “Re-running workflows and jobs” contract). Find the failed run with gh run list --repo <r> --json databaseId,headSha,conclusion,attempt and re-run the one whose headSha is the PR head. This is actions: write-free and needs no new SHA. The per-PR concurrency dedupe lives in the ONE org required workflow (org-wide-pr-validator-required.yml), not a per-repo dispatcher. Full mechanics and the dedupe YAML: references/validator-and-calver.md “The POISON state”.

The INCONCLUSIVE (verdict-less) class — the gate ran but produced no verdict

Section titled “The INCONCLUSIVE (verdict-less) class — the gate ran but produced no verdict”

A third terminal state sits beside PASS and BLOCK: ## validator INCONCLUSIVE — no review verdict was produced (not a BLOCK; no code finding). The required check stays RED on purpose (unreviewed code must never merge), but this is NOT a finding about your diff — it means the review engine produced no Verdict: line at all. Read the ## validator INCONCLUSIVE comment’s Diagnostics tail before touching your branch: it is the evidence that names the class. Do NOT “fix” your code for an INCONCLUSIVE — there is no code finding to fix — and do NOT re-dispatch the same head hoping for a different result; classify the cause from the diagnostics, then act on THAT.

The two root causes, and how to tell them apart from the diagnostics tail:

  • turn 1: N tool call(s) → turn 2: final content len=0 = the STALE-ENGINE class. The gate’s review engine is whatever plugin-review is welded into the charly binary it runs; an engine OLDER than the fix that produced those lines spins a tool loop and emits no verdict. The org pins a current engine via vars.CHARLY_VERSION; a self-hosted runner whose IMAGE bakes an older charly used to short-circuit that pin (if command -v charly; then exit 0) and silently ran the stale welded engine. The pin enforcement that closes this is opencharly/.github#115 (ensure-charly now verifies the on-PATH charly against the pin and downloads the pinned release when they differ). If you see this signature, the cause is the RUNNER ENVIRONMENT, not your diff — escalate to the operator (a stale runner image; the pin enforcement makes a stale on-PATH engine harmless once the workflow at main is the enforced one).
  • inconclusive: … provider did not respond / attempt timed out (or an HTTP status / stall marker) = the PROVIDER/ENDPOINT class. The evidence is in the message: the provider returned a status, or the stream stalled / the whole-request deadline elapsed. This is owned by the gate’s provider configuration (the org AI_REVIEW_* vars: provider/model, AI_REVIEW_ATTEMPT_TIMEOUT, AI_REVIEW_REASONING_EFFORT, AI_REVIEW_MAX_TOKENS) — escalate to the operator, who owns that configuration, rather than re-running blindly. Do not record it as a “flake”; a provider-class INCONCLUSIVE is a real signal that the provider bound or budget needs the operator’s attention.

The distinction that matters: a real BLOCK has a ## Review — BLOCK heading and a ### Blocks list you MUST fix; an INCONCLUSIVE has neither and must NOT be treated as a review finding (reporting it as one, or “fixing” code to satisfy it, is an R1 misdiagnosis). The gate’s own workflow classifies INCONCLUSIVE as exit 3 and keeps the check RED; pr_state_watch.sh reports it distinctly from a verdict BLOCK. Never merge around the red check — an INCONCLUSIVE is an environment/provider condition to fix or escalate, not a licence to bypass the gate. The only merge under a red required check is an explicit, operator-issued override decision, taken by the operator on evidence THEY accept — it is never an agent’s call, and never a documented standing procedure.

pr_state_watch.sh — STOP on a terminal state, never poll in a loop

Section titled “pr_state_watch.sh — STOP on a terminal state, never poll in a loop”

marketplace/scripts/pr_state_watch.sh <owner>/<repo> <pr-number> watches the required check across ALL its same-name runs on the head and exits the instant a terminal state is reached, distinguishing a real verdict BLOCK from POISON. Exit codes: 0 MERGED · 2 BLOCKED · 3 CLOSED · 4 TIMEOUT · 5 ERROR. It is the ONLY sanctioned poll — gh pr checks --watch cannot see POISON (it reads the collapsed rollup), and a hand-rolled while/sleep loop is the R4 band-aid this replaces. On exit 2: read the verdict, fix, re-finalize the body, push a NEW commit — never re-dispatch the same head. Detail: the reference + the script header.

  • MANDATORY PRECONDITION — read this skill (and every dispatched skill) BEFORE ANY code change. As soon as a task will edit a file, create a branch, commit, push, open or update a PR, touch a submodule, or run any git/gh action, this skill MUST be loaded FIRST. It is NOT optional: a git/PR action taken before loading it is an R0 violation and the change is not landable.

  • No direct push to main (the project rulebook’s PR-only landing mandate — see “Post-Execution Policies”). Enforced by ONE organization branch RULESET on refs/heads/main (creation + deletion + non_fast_forward + a strict required status check named exactly validate / validate), with the charly-auto-merge GitHub App as the only bypass actor (its scoped bypass is what lets tag-on-merge’s CHANGELOG commit land on a protected main). On GitHub Team that ONE ruleset also carries the workflows rule (“Require workflows to pass”) naming opencharly/.github/.github/workflows/org-wide-pr-validator-required.yml, so the validator is required ONCE org-wide with no per-repo dispatcher to install. The LEGACY branch-protection API is deliberately NOT used: it has no bypass slot for that app, so org-ruleset.sh deletes it wherever it survives, and enforce_admins plays no part. The pre-push-gate adds a local backstop in every harness that wires it — Claude Code included, via .claude/settings.json’s PreToolUse hooks. Organization-wide apply/verify is owned only by opencharly/.github/scripts/org-ruleset.sh.

  • Never force-push, on any branch, ever (mandate, same rulebook section). main only fast-forwards via native auto-merge’s squash; a feat/ branch, once pushed, advances only by ADDING commits (the author’s change plus any review-round fix commits), and the squash-merge collapses them. A stale feat/ catches up with gh pr update-branch (a merge, NOT a rebase-force); tags are add-only. Amending a feat/ branch is a normal authoring action — legal until the first push (amending a pushed branch would require a force-push, which is forbidden).

  • R10-gated; the merge requires the charly/pr-validator gate’s green check run. R10 PASS authorizes opening the PR (with pasted evidence); a rule violation or R10 FAIL means a red charly/pr-validator check and no merge — fix in the same tree, re-run R10, re-push; the check resets and the validator re-runs.

  • Zero warnings is part of R10 (project rulebook R1). A version-mismatch warning clears with charly box reconcile; any other warning gets /charly-internals:root-cause-analyzer then a real fix — “warning” is never an accepted end state.

  • Atomic on main, never on feat/. The org-wide charly/pr-validator workflow’s PASS enables GitHub native auto-merge (squash), which folds the author’s change and any review-round fix commits into one commit on main; the feat/ branch may freely accumulate fix commits across review rounds. The merge-time CalVer tag and the CHANGELOG/<CalVer>.md entry (written from the PR body — the PR body IS the changelog) are created after merge by the org-wide tag-on-merge workflow (see “CalVer” in references/validator-and-calver.md). Two separate cutovers must never share one PR.

  • Update the PR; never close-and-recreate (except for work that will not land at all — a disproven premise, an abandoned approach). When a review demands changes, append a commit and push it fast-forward — the check resets and the validator re-runs. This is what makes the no-force-push rule livable: because main gets a squash, a branch carrying five fix commits still lands as one. If the PR is AUTO-CLOSED after too many failed validation rounds (or the operator closes it) and the work is carried forward in a NEW PR, you MUST post a comment on the OLD (closed) PR that references the new one (its number/URL) and states what it supersedes — a closed PR is a durable public record that must point to where the work continued.

  • FINISH THE BODY BEFORE THE PUSH. The validator validates the PR — the diff at the branch head and the body — each time it runs, and fetches the PR live. The org REQUIRED workflow fires only on the default push-driven types (opened/synchronize/reopened) and IGNORES on.types (MEASURED); because the head SHA is content-addressed and therefore known BEFORE the push (git rev-parse HEAD), write the body first: commit → compute head/diff-stats → write the WHOLE body (footer last) → push. A body-only fix after the push needs no empty commit — re-run the gate MANUALLY with gh run rerun <run-id> on the failed run (find it with gh run list --repo <r> --json databaseId,headSha,conclusion; a re-run reuses the same GITHUB_SHA and updates the SAME check run in place — no duplicate, clears POISON). An empty re-freeze commit ALSO re-fires the gate (the ruleset workflows rule acts on the push-driven synchronize type) but mints a NEW head SHA — equally valid, just re-key the body to it. Full mechanics: “THE BODY-BEFORE-PUSH RULE” above. Corollary: a PR whose diff is EMPTY because the base already contains the change (you branched from a stale snapshot) is a no-op — close it rather than re-pushing (the validator flags it as body-truthfulness violation: body describes files the diff does not carry). Always git fetch origin main + diff against CURRENT main before opening or finalizing a PR.

  • BEFORE ANY UPDATE PUSH: ALWAYS read the PR’s current comments + validation AND ALWAYS write/update the PR body. Before ANY push that updates an existing PR — a fix commit, a body edit, or gh pr update-branch — do BOTH, in THIS order: (a) read the PR’s LIVE state — gh pr view <n> --repo <r> --json comments,reviews + gh pr checks <n> --repo <r> (or the sanctioned marketplace/scripts/pr_state_watch.sh: the latest charly/pr-validator verdict/validation result and EVERY new comment/review) AND the latest comments/state of every ISSUE this PR closes or relates to (gh issue view <N> --comments) — and ACT on each one: answer it in-thread, claim/hand-off on the issue, or change the pushed state to satisfy it); then (b) write the WHOLE updated body for the head you are about to publish (footer last). The validator re-reviews the diff + body + the FULL live thread on every run, so pushing with a stale read OR a stale body re-reviews the wrong state and can re-ship a defect an existing comment already named. A hook/classifier block is never reshaped-and-retried (B5).

  • Search existing issues/PRs before starting; file ONE proper issue if none; claim it before branching. Before any non-trivial work — and before filing anything — search the whole org (gh search issues <terms> / gh search prs <terms> / gh issue list -S <terms>) and ADD to the existing issue/PR thread rather than creating a duplicate. If none exists, file ONE proper issue (specific title, the problem, the evidence, the intended scope) and reference it from the PR (Closes #N / relates to #N). The issue is the coordination point: check its owner (assignee / claim comment / status label) and CLAIM it (comment + assign) BEFORE you create a branch; if another session already owns it, coordinate on the thread instead of opening a competing PR. Every issue resolved by a PR MUST be closed by an agent once that PR merges (with a comment linking the merge); never leave a resolved issue open. Full protocol: B2b.

  • Identity is in the footers; authority is in the comment verb. On a triggered scope — two or more agents on one issue/PR, OR a blocking dependency (any BLOCKS/UNBLOCKS in play) — EVERY agent-authored comment and PR body carries the two-line footer in ONE canonical order, Agent: FIRST and Assisted-by: LAST (a PR body thereby also satisfies the validator’s “Assisted-by FINAL line” requirement), and a coordination comment OPENS with a label from the CLOSED set — CLAIM · OWNING · HANDING OVER · TAKING OVER · BLOCKS · UNBLOCKS · STATUS · RESOLVED. Both are optional for a solo agent on an uncontended PR; the slug is NEVER appended to Assisted-by. Because same-account sessions are indistinguishable by comment author, the SLUG is authoritative for COORDINATION IDENTITY (who is doing the work), not the GitHub author: the LATEST OWNING (or TAKING OVER) for a scope wins, and you MUST NOT push to another slug’s claimed branch/PR without a HANDING OVER addressed to you, a TAKING OVER naming your authority:, or operator sign-off. The sign-off authority is a SEPARATE axis and is ACCOUNT-gated — a maintainer sign-off is valid ONLY when the comment is posted by a maintainer-set account (atrawog/aitrawog), verified by the author label (by @<login>), never the prose; the two rules do not conflict (the slug governs WHO, the account governs the sign-off’s VALIDITY). Progress is a COMPLETED charly/pr-validator run, never session activity; takeover is comment-FIRST over a 60-minute FLOOR window; the auto-close carry-forward touches FOUR surfaces (closed PR, successor body, the issue, ownership transfer); and there is no R10 class exemption for a library/schema change. An agent NEVER impersonates the operator. Full mechanics + rendered examples: B2b “agent identity and the coordination verb grammar”.

  • A dispatched workflow defaults to the DEFAULT BRANCH — --ref is MANDATORY. gh workflow run <wf> --ref <branch> is required to dispatch CI on a PR head; a dispatch with no --ref runs on main, so its check registers there and counts for nothing on the PR’s required checks. There is NO self-heal: the reusable workflow’s head_ref output is never consumed and no dispatcher re-dispatches itself. For run-on-head diagnostics, target the PR’s branch explicitly.

  • Tree-safety before destructive actions (R6). Check git status + git stash list before any destructive working-tree action — git stash discards in-progress work; rm on a tracked file is destructive. When the sandbox blocks an action, find a non-destructive alternative rather than working around it. The stash/pop cycle can itself silently un-stage a git rm: a stash taken while a deletion is staged restores the deletion as unstaged on pop, so a git status right after the cycle that shows the deleted file back as a plain unstaged change (rather than the staged deletion you left) has quietly lost the staging — re-stage it (git rm <path> again, or git add -u) before committing. A stash/pop round-trip is never a no-op on a mixed add+rm working tree.

  • Right worktree — pin one absolute path for the whole edit→commit→push sequence. Before branching, staging, or committing, confirm the worktree you are driving is the same one your edits landed in: git -C <path> rev-parse --show-toplevel must equal the path you edited, and git -C <path> status --short must list those edits. Under symlinked or near-twin sibling worktrees — a parent dir that is itself a symlink (~/projects → ~/Sync/projects), or look-alike names such as …/charly vs …/<other-worktree> — cd-ing to the wrong sibling makes git switch -c + git commit run against a clean tree and report “nothing to commit”, silently landing nothing (or landing in the wrong repo). Never change the path spelling mid-sequence. An unexpected “nothing to commit” right after editing a file is the signature of this mistake — stop and re-verify --show-toplevel before retrying (blind retry is an R1 violation).

  • The universal PR-gate — audit before any PR action, unconditionally. Before opening, updating, or merging any pull request, run the aggregate audit: gh pr list across every touched repo + git worktree list + the live teammate/agent roster. This is a standing preflight, run first every time — never reached for only once something already looks off — because skipping it risks a duplicate PR for scope already covered in flight, a branch update from a stale worktree, or a merge over a still-running validator’s verdict. Full operational detail: /charly-internals:agents “The universal PR-gate”.

  • Post-commit staging verification. After every commit, re-run git status --short (expect it empty, or only unrelated untracked paths) and git show --stat (confirm every intended file is actually listed) — a multi-path git add naming several paths where one is mistyped or stale can commit only the files that did resolve while git commit still succeeds, producing a commit that would not even compile. git show --stat catches that AFTER the fact, and it is the only post-commit check ON THE COMMITTING SIDE that sees a commit whose message describes an intent its diff does not carry — no gate that reads the tree can, because the mismatch is between the message and the tree, and the tree does not hold the message. It is NOT the only thing that can catch the class: any reader holding message and diff together — a review, a cross-repo reference sweep — catches it too, and catches the cases --stat cannot, such as a mismatch that spans two repositories.

  • Chain the edit to the commit so the failure cannot reach one. A script that edits and the git add / git commit that follow are SEPARATE statements: a guard that aborts the edit does not stop the commit, which then lands under a message describing an edit it does not contain. Joining them — edit && git add … && git commit … — makes a failed edit unable to produce a commit at all. The defect occurred three times in one session; the chain was exercised ONCE, and that occurrence left no artifact by construction — so three is the evidence for the PROBLEM and one is the evidence for the REMEDY. Do not report them under a single “measured”. The chain is necessary, not sufficient: && short-circuits on a non-zero exit, so an edit that fails LOUDLY cannot reach the commit, but a SILENT no-op edit — a pattern that matches nothing — exits 0 and lands the commit under a describing message exactly as before. That case still needs the post-commit read above; the chain closes only the aborting one. A rule held in a head is exactly as effective as a rule not held, unless something in the command line enforces it.

  • git add is all-or-nothing: never name a path that no longer exists. A single git add whose pathspec includes a vanished path fails the WHOLE add — fatal: pathspec '<old>' did not match any files, exit 128, nothing staged, including the paths that did resolve. Two shapes hit this, and both are silent because a PRIOR command already staged something, so the failed add leaves git status looking exactly as intended:

    • a deletion, where the removed file’s path is still named in the add (the deletion was already staged) — lands a partial commit carrying only the deletion;
    • a rename plus a content edit, where git mv old new is followed by git add old new (the mv already staged the rename) — lands the rename with the content edit dropped.

    Stage a rename+edit as git add <new> alone, and a deletion by DIRECTORY or with git add -u; never name the vanished path. This shape defeats the post-commit check above, which is why it needs stating separately: git show --stat DOES list the file, as old.md => new.md | 0, so “every intended file is listed” passes. Presence is not the signal — the tells are the magnitude (| 0 insertions on a commit meant to change content) and a leftover M <new> in git status afterwards. Read the numbers and the post-commit status, not just the filenames, and confirm content directly with git show HEAD:<new-path> | head -1 whenever the edit is the point of the commit.

  • Check-coverage is part of R10. The change must ship the test coverage that proves its functionality (check: checks for new/changed layers & images, Go tests for charly code) AND the live run must have exercised it. A change whose new functionality has no test that would fail without it is not landable.

  • Every repo is tagged at merge. The superproject, every box/<distro>, plugins, and docs all mint v<YYYY.DDD.HHMM> on their own merged HEAD — the tag marks the MERGE, decoupled from any charly.yml version: schema field, so a repo needs no charly.yml to be tagged. The sole exception is the sdk contract repo, which tags under its own Go-module scheme v0.<YYYYDDD>.<HHMM with all leading zeros stripped> (B2 step 0) — not an exemption but a hard Go-module requirement: v<YYYY.DDD.HHMM> is not a valid Go module version (semver forbids a leading-zero segment — 0733→733 — and a major ≥ 2 would force a /vN module-path suffix that breaks every import github.com/opencharly/sdk), so the stripped v0.<…> form is mandatory, not a choice. A skipped tag is a defect, not an exemption: the orchestrator verifies the tag landed after each merge (git ls-remote --tags origin v<VER> non-empty) and, if tag-on-merge skipped it, backfills it add-only on the merged HEAD (git tag -a v<VER> -m "<subject>" <merged-HEAD> + git push origin refs/tags/v<VER>) — tags are immutable, so a backfill only adds one, never moves an existing tag.

Umbrella mechanics (when you work in the ~400-submodule umbrella)

Section titled “Umbrella mechanics (when you work in the ~400-submodule umbrella)”

These are umbrella-REPO-maintenance mechanics, run inside a checkout of the opencharly/opencharly umbrella — never commands a charly user runs with only the charly binary. They are documented here because the umbrella AGENTS.md Skill Dispatcher routes its pinning/gitlink row to this skill, and the umbrella AGENTS.md names them as the sanctioned path for that maintenance work (“Umbrella-native mechanics are the sanctioned path for umbrella work”). A charly end-user never sees them; they are the umbrella’s own charly task entities.

Inside the umbrella checkout the umbrella AGENTS.md governs (rule 7: read the subrepo’s own rulebook before touching it; charly/AGENTS.md owns R0–R10 inside charly/). The umbrella’s own commands — never an ad-hoc substitute — are: charly task sync (policy-B pin bump), charly task verify (the full pinning gate — there is NO CI gate), charly task hooks (install the per-commit gate), charly task harness (config parity), charly task map. Hard rules there: never edit inside a submodule (change lands by PR to the owning repo; the umbrella only records gitlinks), run submodule git through git -C <absolute-path> from the umbrella root, no worktrees inside submodules, pin only MERGED refs, and bound every command’s output (SIGPIPE is ignored — grep floods on Broken pipe). Full detail: references/umbrella-mechanics.md.

Topic File
B1 (the two-step branch-per-change loop, concurrent landings, the cross-repo WIP landing sequence) and B4 (sync to upstream + prune) references/branch-and-pr-loop.md
B2 (multi-repo/multi-worktree coordination, per-module verification), B2b (cross-session coordination via PR comments when another session’s PR blocks you; the two-line identity footer + the CLAIM/OWNING/HANDING OVER/TAKING OVER/BLOCKS/UNBLOCKS/STATUS/RESOLVED verb grammar, mandatory on a contended or blocking scope), B3 (agent teams in per-teammate worktrees), B6 (cross-repo @github landing), and B7 (multi-worktree landing + refresh, the canonical end-to-end) references/multi-repo-coordination.md
B5 (the fresh evaluator + fork/PR path, the two-gate autonomous-landing model), CalVer generation, post-landing cleanliness + report format, and the validation-FAILS recovery sequence references/validator-and-calver.md
Evidence discipline — provenance vs plausibility of a pasted gate, the three freshness surfaces (head / body / pasted output), positive-vs-negative claim decay, sweeping for claims a fix invalidated, the merged-tree gate for a BEHIND PR, source-and-regeneration as one cross-repo cutover, submodule pointers reverted by a non-conflicting merge, and why status-absence on a known head proves nothing references/evidence-and-freshness.md
Umbrella mechanics — the ~400-submodule view, policy B, charly task sync/verify/hooks/harness, the no-edit-in-submodule rule, pin discipline references/umbrella-mechanics.md
Watch-and-wake — the self-sustaining watcher family (--auto-rearm, the single-instance lock, rate-limit backoff) and the arm → wake → act → re-arm runbook marketplace/scripts/pr_state_watch.sh, marketplace/scripts/pr_watch_many.sh, marketplace/scripts/gh_watch.sh (run one — never hand-roll a poll) + references/watch-and-wake.md
The pre-validator self-audit — the five BLOCK classes detectable before the first push, the nine-step preflight that removes them, and the delegating-parent duties references/pre-validator-self-audit.md

The pre-validator self-audit — one pass before the first push

Section titled “The pre-validator self-audit — one pass before the first push”

Before the first push, run the pre-validator self-audit: a short, five-class preflight that catches EVERY block cause a charly/pr-validator run would otherwise return. MEASURED (umbrella #286): across 9 docs PRs landed in one session, 6 first pushes blocked and 17 validator runs were spent; every block fell into one of five diff-derivable classes, and the one PR that passed first try had run this pass. The full checklist — the five classes, the nine steps, and the delegating-parent duties — lives in references/pre-validator-self-audit.md. A parent that spawns a worker MUST embed it in the brief (see /charly-internals:agents), and the worker MUST run it before its first push.

The watcher family — a self-sustaining loop

Section titled “The watcher family — a self-sustaining loop”

Three generic watcher scripts under scripts/ in the opencharly/marketplace repo (addressed as marketplace/scripts/<name> from the umbrella checkout), plus a shared helper library _watch_common.sh, turn a landing into a background command that EXITS — and, because the harness notifies the owning agent when a background command finishes, exiting IS the notification.

The harness constraint: an agent is woken ONLY when a background command COMPLETES. A watcher (or a while true loop) that never exits gives NO wake, and a one-shot watcher that exits leaves nothing watching until an agent re-arms — the dropped step this family exists to remove. Two patterns:

  • per-event notify (--no-rearm, the default) — one-shot; the agent re-arms on each wake;
  • durability (--auto-rearm) — on a non-terminal exit the watcher detaches a lock-guarded successor with the SAME args BEFORE it prints and exits, so a watch is ALWAYS alive independent of the agent. --auto-rearm keeps a WATCH alive; the agent’s re-arm keeps NOTIFICATIONS alive. A durable supervisor never exits, so it is NOT the answer.

THE RE-ARM INVARIANT. An agent that acts on a wake MUST leave a watcher armed — --auto-rearm makes that automatic; a per-event agent re-arms in the same turn. An agent that fires an event and does not re-arm has silently stopped watching.

Single instance; rate limits. A per-args lockfile (flock + a holder PID) means repeated arms never stack — a foreground arm takes over a live peer cleanly. Before each poll the remaining core quota is read from the FREE /rate_limit endpoint; below WATCH_RATE_MIN (default 200) the watcher backs off (4x, capped) rather than hammering into the observed HTTP-403 wall.

Arm this To watch Wake event
pr_state_watch.sh <owner>/<repo> <pr> ONE PR’s terminal state, across ALL same-name check-runs on the head 0 MERGED · 2 BLOCKED · 3 CLOSED · 4 TIMEOUT · 5 ERROR — and it names POISON (a green re-dispatch stuck behind an older same-name FAILURE) distinctly from a verdict BLOCK
pr_watch_many.sh [--repos R,…] [--interval S] [--stallmin M] [--validator NAME] [--auto-rearm] <owner>/<repo> <pr> … a CROSS-REPO PR batch the first of: a PR terminal (above, delegated); a NEW verdict — a new <validator> run COMPLETING on any watched repo; a STALL — no new verdict for the window while scopes stay open+unmerged
gh_watch.sh [--events L] [--interval S] [--stallmin M] [--workflow NAME] [--auto-rearm] <owner>/<repo>#<num> … a per-item list of PRs and/or issues a NEW comment; a NEW verdict; merged (unblocked); closed without merge (find the successor); or a stall (takeover candidate)

Both batch watchers poll every --interval and exit when a new --workflow run (default charly/pr-validator) COMPLETES; the STALL layer fires only when no new verdict lands within the window while the scope is still open+unmerged. That is the B2b.1 progress-signal rule applied to a monitor — progress is a completed validator run, never session activity (a looping agent never falls quiet; a peer WAITING on a running validator looks quiet but is working). Watch the scopes actually in flight — a stale watch list produces false stalls.

gh_watch.sh events are DELTA or STATE. comment and verdict are DELTA: arming seeds the current comment-count / newest run id, so a pre-existing comment or verdict never wakes you. merged, closed, and stall are STATE: they fire while the item IS in that state — arming merged on an already-merged PR wakes immediately (exactly how “has my blocker landed?” reads). STATE stall additionally requires the item to be open+unmerged, so a closed/merged item never emits a false (open, unmerged) alarm.

Three field-learned watcher requirements — mandatory, not tips:

  • Silence is an ALARM. A blocked item with no pushes emits no validator runs and no events, so an event-only watcher is blind to the worst case. Every watcher MUST carry a per-item stall/silence alarm — no progress (no new commit, comment, or completed charly/pr-validator run; i.e. updated_at unchanged) for the window while the item is open+unmerged. gh_watch.sh’s stall is therefore ON BY DEFAULT. The events tell you when something HAPPENED; the stall alarm tells you when something SHOULD have and did not. (Measured: a PR sat silently unchanged for ~2 hours with NO alert; the instant stall was armed it fired.)
  • Watch the OUTCOME, not every event. Waking on every comment of an actively iterating owner is noise. The DEFAULT event set is the terminal outcomes (merged,closed) plus the stall alarm; add comment/verdict ONLY for a wait that genuinely needs them (EVENTS=comment on an issue you asked a question on and want the moment anyone replies).
  • Liveness ≠ progress. Judge progress by artifacts — a pushed branch, a new commit, an opened PR, a merged tag — never by a session heartbeat or “still investigating”. The stall alarm is the mechanism: it keys on the item’s updated_at, so its silence is measured against artifacts, not activity.

On exit 2 of pr_state_watch.sh, read the verdict, fix, re-finalize the body, push a NEW commit — never re-dispatch the same head to “see if it clears”. It is preferred over gh pr checks --watch (which reads the collapsed rollup and cannot see POISON) and over any hand-rolled loop.

The wake feeds the coordination protocol (B2b.1): a stall wake makes the scope a TAKEOVER candidate — post a coordination comment FIRST, wait the window, then TAKING OVER — authority: window-expired BEFORE any push. A merged wake is the UNBLOCK signal. The full arm → wake → act → re-arm loop, the event catalog, and the takeover it feeds: references/watch-and-wake.md.

  • the project rulebook “Post-Execution Policies” — the mandate this skill operationalizes.
  • marketplace/internals/agents/pr-validator.md — the fresh evaluator’s full spec.
  • marketplace/scripts/pr_state_watch.sh — the per-PR terminal-state poll (stop, never loop).
  • marketplace/scripts/pr_watch_many.sh — the cross-repo PR-batch watcher (terminal / new verdict / stall).
  • marketplace/scripts/gh_watch.sh — the per-item PR/issue event watcher (comment / verdict / merged / closed / stall).
  • references/watch-and-wake.md — when to arm which watcher, the event semantics, and the re-arm-after-wake loop.
  • opencharly/.github/scripts/org-ruleset.sh — the sole organization-wide owner of the ONE branch ruleset (required workflow + branch rules) apply/verify.
  • /charly-internals:repo-setup — the org/dotgithub configuration, the landing automation (required workflow, native auto-merge, tag-on-merge CalVer), and the new-repo setup checklist.
  • /charly-internals:cutover-policy — one-phase, atomic-commit, R10-at-the-end.
  • /charly-build:migrate — version: ↔ tag coupling, per-merge tags, push order.
  • /charly-build:reconcile — cross-repo @github pin alignment used by B6.
  • /charly-check:check — the check-coverage gate (R10) every change must satisfy.
  • /charly-internals:root-cause-analyzer — run on any FAIL before re-trying.

Invoke before any git / gh action that commits, branches, pushes, opens a PR, or drives the pr-validator merge/tag — and whenever syncing to upstream, applying branch protection, or pruning branches/worktrees across the main repo and its submodules.