Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
40 changes: 38 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -301,6 +301,41 @@ Generation is identical to direct push. Only delivery changes: the same commit i

With the default `github.token`, the repository or organization must allow GitHub Actions to create pull requests. A GitHub App token or PAT can instead be passed as `github_token`. The same input is used for review comments and sync delivery.

### Save the diagram to a branch of its own

Set `sync_strategy: branch` to keep the analysis off your code branches entirely:

```yaml
- uses: CodeBoarding/CodeBoarding-action@v1
with:
mode: sync
llm: hosted
target_branch: main

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

target_branch, is this the target in tehms of where the sync will happen or in terms of which branch we will sync with, unsure that wording is clear again.

i think that most of these things will be read by ppl or even more by their agents so proly descriptive and somewhat clear names are worth investing in.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch, it was ambiguous, and under branch the old description ("Branch updated by sync mode") was just wrong. I kept the name since it's an existing v1 input and renaming would break every current workflow, but the description now says it's the code branch sync analyzes: with push / pull_request the analysis is also committed to it, with branch it's only read. Same wording in the README table and the new section.

sync_strategy: branch
```

Here `target_branch` is the code branch sync analyzes, and it is only read. Each sync adds one commit to `codeboarding/analysis` in the same repository (set `analysis_branch` to change the name), an orphan branch that shares no history with `main`. It holds the same `.codeboarding/` files sync would otherwise commit to `main`, plus `.codeboarding/source.json` naming the commit they describe and the configuration that made them; the commit message carries both as `CodeBoarding-Source:` and `CodeBoarding-Config:` trailers. Pushes only ever fast-forward, `main` is never written, and no pull request is opened. Reviews read their base from the branch, and the web platform reads the latest diagram from it. The [analysis branch section](docs/COMMIT_STRATEGY.md#the-analysis-branch) covers what happens if the branch is deleted, and a ruleset you should import to protect it: sync and review load a pickle from it.

**Moving an existing setup.** Nothing changes until you opt in: `push` and `pull_request` keep working as before. To switch, paste this into your coding agent:

```text
Move this repository's CodeBoarding sync to sync_strategy: branch.
1. In the workflow that runs CodeBoarding/CodeBoarding-action with mode: sync, set
`sync_strategy: branch` in its `with:` block (replace push or pull_request).
Keep every other input.
2. Delete the generated files under .codeboarding/ from the default branch, keeping
the user configuration: .codeboarding/.codeboardingignore,
.codeboarding/health/health_config.json and .codeboarding/health/.healthignore.
Reviews prefer a baseline committed on the branch, so a stale one left there
would keep being used.
3. Remove any .gitattributes lines that mark .codeboarding/ files as
linguist-generated, if nothing else is left under .codeboarding/ for them.
4. Open a pull request with these changes. After it merges, close any open
pull request from the codeboarding/sync branch and delete that branch.
```

The first sync after the merge creates `codeboarding/analysis`, catching up from a saved analysis when there is one.

## Inputs

| Input | Mode | Default | Description |
Expand All @@ -316,8 +351,9 @@ With the default `github.token`, the repository or organization must allow GitHu
| `parsing_model` | both | empty | Parsing-only override for `model`. |
| `depth_cap` | both | `2` | Positive integer maximum analysis depth, including full-analysis fallbacks. Changing it rebuilds incompatible state. |
| `github_token` | both | `${{ github.token }}` | Token for comments and sync delivery. |
| `sync_strategy` | sync | `push` | `push` or `pull_request`. |
| `target_branch` | sync | event branch | Branch receiving the baseline or rolling PR. |
| `sync_strategy` | sync | `push` | Where sync saves the analysis: `push` (a commit on `target_branch`), `pull_request` (a rolling PR into it), or `branch` (commits on `analysis_branch`). |
| `analysis_branch` | both | `codeboarding/analysis` | Branch in this repository that `sync_strategy: branch` saves the analysis to; reviews read their base from it when it exists. |
| `target_branch` | sync | event branch | Code branch sync analyzes. With `push` or `pull_request` it also receives the analysis commit or rolling PR; with `branch` it is only read. |
| `force_full` | sync | `false` | Ignore the committed baseline for this run. |
| `warmstart_retention_days` | review | `1` | Days to keep the reusable analysis. Only the next run reads it. |

Expand Down
16 changes: 14 additions & 2 deletions action.yml
Original file line number Diff line number Diff line change
Expand Up @@ -143,11 +143,15 @@ inputs:
required: false
default: ${{ github.token }}
sync_strategy:
description: 'Sync delivery method: push or pull_request.'
description: 'Where sync saves the analysis: push (a commit on target_branch), pull_request (a rolling PR into target_branch), or branch (commits on analysis_branch; target_branch is never written).'
required: false
default: 'push'
analysis_branch:
description: 'Branch in this repository that sync_strategy branch saves the analysis to, one commit per sync. Reviews read their base analysis from it when it exists.'
required: false
default: 'codeboarding/analysis'
target_branch:
description: 'Branch updated by sync mode. Defaults to the event branch.'
description: 'Code branch sync mode analyzes. With sync_strategy push or pull_request the analysis is also committed to it; with branch it is only read. Defaults to the event branch.'
required: false
default: ''
force_full:
Expand Down Expand Up @@ -223,6 +227,7 @@ runs:
HEAD_AUTHOR_EMAIL: ${{ github.event.head_commit.author.email }}
TARGET_BRANCH_INPUT: ${{ inputs.target_branch }}
SYNC_STRATEGY: ${{ inputs.sync_strategy }}
ANALYSIS_BRANCH: ${{ inputs.analysis_branch }}
COMMENT_BODY: ${{ github.event.comment.body }}
AUTHOR_ASSOCIATION: ${{ github.event.comment.author_association }}
ISSUE_PR_URL: ${{ github.event.issue.pull_request.url }}
Expand Down Expand Up @@ -451,6 +456,8 @@ runs:
CHECKOUT_DIR: ${{ github.workspace }}/.codeboarding-target
STAGE_DIR: ${{ runner.temp }}/cb-state/${{ github.action }}/out
FORCE_FULL: ${{ inputs.force_full }}
SYNC_STRATEGY: ${{ inputs.sync_strategy }}
ANALYSIS_BRANCH: ${{ inputs.analysis_branch }}
CFG_HASH: ${{ steps.state.outputs.cfg_hash }}
# Lets a branch without a usable committed baseline catch up from a saved analysis.
ANCESTOR_LOOKUP: ${{ github.server_url == 'https://github.com' && steps.state.outputs.cfg_hash != '' }}
Expand All @@ -475,6 +482,10 @@ runs:
TARGET_BRANCH: ${{ steps.guard.outputs.target_branch }}
SYNC_BRANCH_START_SHA: ${{ steps.guard.outputs.sync_branch_start_sha }}
SYNC_STRATEGY: ${{ inputs.sync_strategy }}
ANALYSIS_BRANCH: ${{ inputs.analysis_branch }}
ENGINE_VERSION: ${{ steps.state.outputs.engine_version }}
# Recorded with each analysis-branch commit, so a reader can tell which configuration made it.
CFG_HASH: ${{ steps.state.outputs.cfg_hash }}
GITHUB_TOKEN: ${{ inputs.github_token }}
GH_TOKEN: ${{ inputs.github_token }}
GH_ENTERPRISE_TOKEN: ${{ inputs.github_token }}
Expand Down Expand Up @@ -550,6 +561,7 @@ runs:
ANCESTOR_LOOKUP: ${{ github.server_url == 'https://github.com' && steps.state.outputs.cfg_hash != '' }}
# For rewriting the progress comment while a base is built from scratch.
PROGRESS_HEADER: ${{ steps.guard.outputs.comment_id }}
ANALYSIS_BRANCH: ${{ inputs.analysis_branch }}
BASE_REF: ${{ steps.guard.outputs.base_ref }}
REPOSITORY: ${{ github.repository }}
GH_HOST: ${{ github.server_url }}
Expand Down
67 changes: 67 additions & 0 deletions docs/COMMIT_STRATEGY.md
Original file line number Diff line number Diff line change
Expand Up @@ -107,6 +107,7 @@ them:
|---|---|
| the published `codeboarding-base-<cfg>-<merge_base>` artifact with a compatible depth cap | none |
| no usable artifact — check out the merge base, seed from a compatible baseline committed there, catch up | one incremental, full if Core requires it |
| the analysis branch (`sync_strategy: branch`): its commit for the merge base, else for the nearest of the merge base's last 100 first-parent ancestors, made under this configuration | none for the merge base itself (`reused`), one incremental otherwise (`incremental`) |
| no compatible committed baseline either: the nearest `codeboarding-base-<cfg>-<sha>` artifact among the merge base's last 100 first-parent ancestors | one incremental from that commit to the merge base |
| none within 100 commits either | full analysis directly, at the configured `depth_cap` |

Expand Down Expand Up @@ -145,6 +146,72 @@ diffs against, recorded as a digest in `origin.json`. Two runs of the engine ove
one commit need not name components identically, so a head descended from one
base and a diagram drawn against another would report changes nobody made.

## The analysis branch

`sync_strategy: branch` saves the analysis to a branch of its own in the same
repository, `codeboarding/analysis` unless `analysis_branch` names another.
`target_branch` is then the code branch sync analyzes; it is never written.

**What lives where.** The branch is an orphan: it shares no history with the code.
Each sync adds one commit holding the same `.codeboarding/` files the `push`
strategy would commit to the target branch, plus `.codeboarding/source.json`:

```json
{"schema": 1, "source_branch": "main", "source_sha": "<sha analysed>", "generated_at": "<iso>", "engine_version": "<v>", "config": "<cfg hash>"}
```

The commit is `chore(codeboarding): diagram of main @<sha7>` with two trailers:
`CodeBoarding-Source: <sha>` and `CodeBoarding-Config: <cfg hash>`, the same
configuration hash that names the base artifacts (engine version, provider, model,
depth cap). Engine output is never edited; which commit it describes and how it was
made live only in `source.json` and the trailers. The target branch is never
written, not even `.gitattributes`. The base artifacts are still published, named
for the analysed commit.

**How a sync writes it.** It seeds from the branch tip when the tip was made under
this configuration, and runs incrementally. The generated files are replaced
wholesale; only the checkout's own `.codeboardingignore` and health configuration
are kept. The push is a fast-forward onto the tip it fetched, never forced. Sync
refuses to write to an existing branch that is not an analysis branch (its tip has
no `CodeBoarding-Source` trailer, or holds anything besides `.codeboarding/`), so
pointing `analysis_branch` at a code branch fails instead of emptying it. If the
target branch moved during the analysis, the result is dropped, as with `push`. If
another sync moved the analysis branch, it builds on that tip once. A push the
remote refuses while the tip did not move is a branch rule, and the run fails
saying so.

The target branch is checked just before the push, not in the same transaction:
if it moves in that window, the branch can end on an analysis of the older commit.
Its trailer still names that commit, so no reader takes it for newer, and the run
queued for the newer commit replaces it.

**How a review reads it.** After an exact artifact and a baseline committed at the
merge base, a review lists the newest 100 commits of the branch (fetched without
file contents, so the listing costs commit messages only) and matches their
trailers against the merge base's first-parent history, up to 100 commits deep.
Only entries made under this run's configuration count: an entry for the merge
base itself is reused as is, so nothing else would catch a different engine or
model. An entry for an ancestor is caught up incrementally. Entries found only
under another configuration make a full run's reason `incompatible`. Only then
does it look for ancestor artifacts.

**If the branch is deleted**, the next sync creates it again as a new orphan,
seeding from a saved ancestor artifact when there is one and analyzing in full
otherwise. The history is lost; the current diagram is not.

**Protecting it.** Sync and review load `static_analysis.pkl` from this branch, and
a pickle runs code when loaded, so whoever can write the branch can run code in
the sync and review workflows. Import
[`analysis-branch-ruleset.json`](analysis-branch-ruleset.json) under Settings,
Rules, Rulesets, New ruleset, Import a ruleset. It blocks creating, updating,
deleting and force-pushing `codeboarding/analysis` for everyone except its bypass
actor, GitHub Actions (integration `15368`), which is what the default
`github.token` pushes as. If sync pushes with a GitHub App token instead, such as
the CodeBoarding Review app (`4021464`), make that app the only bypass actor:
any workflow can use `github.token`, while only the workflows you give the app's
key can push as the app. Rulesets on a private repository need a paid GitHub plan
(Pro, Team or Enterprise); on Free they apply to public repositories only.

## Trust boundary

`static_analysis.pkl` is a Python pickle, so state derived from code the
Expand Down
24 changes: 24 additions & 0 deletions docs/analysis-branch-ruleset.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,24 @@
{
"name": "CodeBoarding analysis branch",
"target": "branch",
"enforcement": "active",
"conditions": {
"ref_name": {
"include": ["refs/heads/codeboarding/analysis"],
"exclude": []
}
},
"rules": [
{ "type": "creation" },
{ "type": "update", "parameters": { "update_allows_fetch_and_merge": false } },
{ "type": "deletion" },
{ "type": "non_fast_forward" }
Comment on lines +11 to +15

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Restrict baseline writes to the bypass actor

In repositories where contributors may push branches but the main branch requires reviewed PRs, this ruleset still lets any contributor create or fast-forward codeboarding/baseline; these rules block only deletion and non-fast-forward updates. Review and sync runs later pass that branch's .codeboarding/static_analysis.pkl through restore_baseline to the engine, which loads pickle state in a privileged workflow, so a contributor can bypass code review with a crafted baseline commit. Add creation and update restrictions, leaving only the intended app/token as a bypass actor.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 7a65300 (now docs/analysis-branch-ruleset.json): it adds creation and update rules, so only the bypass actor can write. The bypass is GitHub Actions (15368), since that's what the default github.token pushes as, and the docs explain the pickle risk and say to make a GitHub App the only bypass actor when sync pushes with one, since any workflow can use github.token.

],
"bypass_actors": [
{
"actor_id": 15368,
"actor_type": "Integration",
"bypass_mode": "always"
}
]
}
104 changes: 103 additions & 1 deletion scripts/action/analyze.sh
Original file line number Diff line number Diff line change
Expand Up @@ -126,6 +126,20 @@ analyze_sync() {

REQUIRES_FULL=true
if [ "$(printf '%s' "${FORCE_FULL:-false}" | tr '[:upper:]' '[:lower:]')" != true ]; then
# The analysis branch's tip is this branch's last analysis. Without a usable
# one, the run below seeds from a saved ancestor or analyzes in full, and
# delivery creates the branch again.
if [ "${SYNC_STRATEGY:-}" = branch ]; then
local tip_entry
tip_entry="$(analysis_branch_index "${REPOSITORY:-}" | awk '{print $1, $3; exit}')"
if [ -z "$tip_entry" ]; then
echo "::notice::$ANALYSIS_BRANCH has no analysis to continue from; this sync creates it."
elif ! usable_config "${tip_entry#* }"; then
echo "::notice::The analysis on $ANALYSIS_BRANCH was made with another engine version or settings; not continuing from it."
elif ! restore_analysis_branch "${REPOSITORY:-}" "${tip_entry%% *}" "$state" "$CHECKOUT_DIR"; then
echo "::notice::Could not read the analysis on $ANALYSIS_BRANCH; not continuing from it."
fi
fi
if [ "$(depth_cap_from "$state/analysis.json")" = "$DEPTH_CAP" ]; then
incremental "$CHECKOUT_DIR" "$state"
fi
Expand Down Expand Up @@ -260,6 +274,73 @@ keep_user_config() {
done
}

# sync_strategy: branch keeps one commit per sync on ANALYSIS_BRANCH, each with
# trailers naming the commit it analysed (CodeBoarding-Source) and the
# configuration it ran under (CodeBoarding-Config). Lists them as
# "<branch commit> <source sha> <config>", newest first, at most
# ANALYSIS_BRANCH_DEPTH of them; an entry without a config has none to compare.
# Fetched without blobs into a scratch repository: the lookup needs messages, and
# a hundred pickles would cost more than it saves.
ANALYSIS_BRANCH_DEPTH="${ANALYSIS_BRANCH_DEPTH:-100}"
analysis_branch_index() {
local repository="$1" scratch="$RUNNER_TEMP/codeboarding-analysis-index.git" auth
[ -n "${ANALYSIS_BRANCH:-}" ] || return 0
rm -rf "$scratch"
git init -q --bare "$scratch"
auth="$(printf 'x-access-token:%s' "${GIT_TOKEN:-}" | base64 -w0)"
git -C "$scratch" -c "http.extraheader=AUTHORIZATION: basic $auth" fetch -q --filter=blob:none \
--depth="$ANALYSIS_BRANCH_DEPTH" "${GITHUB_SERVER_URL%/}/${repository}.git" "refs/heads/$ANALYSIS_BRANCH" 2>/dev/null ||
return 0
git -C "$scratch" log --format='%H %(trailers:key=CodeBoarding-Source,valueonly,separator=%x20) %(trailers:key=CodeBoarding-Config,valueonly,separator=%x20)' FETCH_HEAD |
awk 'NF >= 2 {print $1, $2, (NF >= 3 ? $3 : "-")}'
}
# An entry is only as good as the configuration it ran under: the artifact name
# pins it for saved analyses, the trailer does here. Without a configuration hash
# this run cannot tell, so it uses none, as it reuses no artifact either.
usable_config() {
[ -n "${CFG_HASH:-}" ] && [ "$1" = "$CFG_HASH" ]
}
# Replaces the generated state in $3 with an analysis-branch commit's, keeping the
# user configuration of checkout $4. source.json is provenance, not engine state.
restore_analysis_branch() {
local repository="$1" commit="$2" state="$3" config_from="$4" scratch="$RUNNER_TEMP/codeboarding-analysis-restore"
fetch_commit "$repository" "$commit" || return 1
rm -rf "$scratch"
mkdir -p "$scratch"
git -C "$CHECKOUT_DIR" archive "$commit" .codeboarding | tar -x -C "$scratch" || return 1
[ -f "$scratch/.codeboarding/analysis.json" ] || return 1
rm -f "$scratch/.codeboarding/source.json"
rm -rf "$state"
cp -a "$scratch/.codeboarding" "$state"
keep_user_config "$config_from" "$state"
}
# Seeds $3 from the analysis branch's entry for $2, or for its nearest first-parent
# ancestor that has one under this configuration, keeping checkout $4's user
# configuration. Sets BRANCH_SOURCE to the commit it describes and BRANCH_DISTANCE
# to how far below $2 that is; BRANCH_REASON=incompatible when the only entries
# found were made under another configuration.
BRANCH_SOURCE="" BRANCH_DISTANCE="" BRANCH_REASON=""
seed_from_analysis_branch() {
local repository="$1" tip="$2" state="$3" config_from="$4" index commit line entry="" distance=0
BRANCH_SOURCE="" BRANCH_DISTANCE="" BRANCH_REASON=""
index="$(analysis_branch_index "$repository")"
[ -n "$index" ] || return 1
fetch_commit "$repository" "$tip" "$(( CATCHUP_BOUND + 1 ))" || true
for commit in $(git -C "$CHECKOUT_DIR" rev-list --first-parent --max-count=$(( CATCHUP_BOUND + 1 )) "$tip" 2>/dev/null); do
while read -r line; do
if usable_config "${line#* }"; then
entry="${line%% *}"
break
fi
BRANCH_REASON=incompatible
done < <(awk -v source="$commit" '$2 == source {print $1, $3}' <<< "$index")
[ -z "$entry" ] || break
distance=$(( distance + 1 ))
done
[ -n "$entry" ] && restore_analysis_branch "$repository" "$entry" "$state" "$config_from" || return 1
BRANCH_SOURCE="$commit" BRANCH_DISTANCE="$distance" BRANCH_REASON=""
}

# Rewrites the sticky progress comment while the base is built from scratch. A
# fork's read-only token makes every call fail, which costs nothing.
PROGRESS_PID=""
Expand Down Expand Up @@ -342,6 +423,26 @@ analyze_review() {
elif [ -f "$base_state/analysis.json" ]; then
full_cause=incompatible
fi
# The analysis branch: its entry for the merge base is that commit's own
# analysis, and an entry for an ancestor is caught up like a committed one.
if [ "$REQUIRES_FULL" = true ] &&
seed_from_analysis_branch "$REVIEW_BASE_REPO" "$REVIEW_BASE_SHA" "$base_state" "$base_checkout"; then
if [ "$(depth_cap_from "$base_state/analysis.json")" != "$DEPTH_CAP" ]; then
full_cause=incompatible
elif [ "$BRANCH_DISTANCE" -eq 0 ]; then
REQUIRES_FULL=false base_method=reused
else
incremental "$base_checkout" "$base_state"
if [ "$REQUIRES_FULL" = true ]; then
full_cause=incompatible
else
base_method=incremental base_from_sha="$BRANCH_SOURCE"
catchup_commits="$(catchup_count "$BRANCH_SOURCE" "$REVIEW_BASE_SHA")"
fi
fi
elif [ "$BRANCH_REASON" = incompatible ]; then
full_cause=incompatible
fi
# Nothing at the merge base to grow from: catch up from the nearest saved
# ancestor, and publish the result under the merge base's own name below.
if [ "$REQUIRES_FULL" = true ] && seed_from_ancestor "$REVIEW_BASE_REPO" "$REVIEW_BASE_SHA" "$base_state" false "$base_checkout"; then
Expand Down Expand Up @@ -400,7 +501,8 @@ analyze_review() {
# under the same name every run, so normally only a run that produced one
# publishes it. The exception is lifetime: a review artifact references a base
# by id for its whole retention, so one about to expire is renewed rather than
# left dangling under a review that outlives it.
# left dangling under a review that outlives it. A base read from the baseline
# branch is published too: no artifact holds it yet.
local publish_base=false
if [ "$base_published" != true ] || [ "${RENEW_BASE:-false}" = true ]; then
stage "$base_state" base
Expand Down
Loading
Loading