Skip to content

ci: type-check tests/ and send a User-Agent on sim API requests - #57

Merged
scott-lowe-vapi merged 1 commit into
mainfrom
ci/test-workflow
Oct 3, 2026
Merged

scott-lowe-vapi merged 1 commit into
mainfrom
ci/test-workflow

Conversation

@scott-lowe-vapi

Copy link
Copy Markdown
Contributor

Value

V.A.L.U.E. tier: project — PR 1 of 10 for inline simulation PR checks (TEST-141); this PR is a small, behavior-preserving slice.

The CI workflow itself landed in #56, so this PR is smaller than PR 1 in the plan.

Evidence of value

Check main (69c7e83) This branch
tsc --noEmit with tests/ included 37 errors (6 test files) 0 errors
npm test 355 pass 357 pass (2 new, 1 rewritten)
User-Agent on POST /eval/simulation/run none vapi-gitops-sim/1.0.0 (asserted against a local HTTP server)

What the 37 errors were:

  • Fixture drift after the hash-store migration: state entries still carrying lastPulledHash/lastPushedHash, bare-string state values, and an untyped emptyLoaded().
  • One real gap: the reconcile-state-key harness never passed the required formatError, so any test reaching that error path would have thrown a TypeError instead of testing it.

Tests that used the removed hash fields as markers (state-merge, recanonicalize) now mark "which copy won" with distinct UUIDs or object identity, so they still check the same behavior.

Testing plan

  • npm run build (now covers src/ and tests/) and npm test locally on Node 22: green. CI on this PR runs both on Node 20 and 22.
  • New tests/user-agent.test.ts covers the header format against package.json's version, and the header actually sent on run create.
  • sim.test.ts now covers the legacy bare-string state value directly, replacing the old cast-based "forward-compat" test.
  • Not tested: a live run against the API (the header is asserted locally only), and Node 20 locally (left to CI). src/ behavior is unchanged apart from the added header.

Refs TEST-141

🤖 Generated with Claude Code

- tsconfig.json now includes tests/, so `npm run build` (already run by
  CI on every PR) type-checks the tests. tsx runs them without
  type-checking, which is how 37 type errors accumulated unnoticed.
- Fix those 37 errors. All were fixture drift after the hash-store
  migration (state entries carrying lastPulledHash/lastPushedHash or
  bare-string values, an untyped emptyLoaded fixture), plus one real
  gap: the reconcile-state-key harness never passed the required
  formatError, so any test reaching that error path would have thrown.
- New src/user-agent.ts; `npm run sim` sends
  `User-Agent: vapi-gitops-sim/<version>`, so gitops-started simulation
  runs can be counted in the platform's run-started analytics.

Refs TEST-141

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

scott-lowe-vapi commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor Author

Merge activity

  • Oct 3, 5:59 AM UTC: A user started a stack merge that includes this pull request via Graphite.
  • Oct 3, 5:59 AM UTC: @scott-lowe-vapi merged this pull request with Graphite.

@scott-lowe-vapi
scott-lowe-vapi merged commit d7b9a34 into main Oct 3, 2026
3 checks passed
scott-lowe-vapi added a commit that referenced this pull request Oct 3, 2026
…lse pass (#58)

## Value

**V.A.L.U.E. tier:** project — PR 2 of 10 for inline simulation PR checks ([TEST-141](https://linear.app/vapi/issue/TEST-141/gitops-run-simulation-suites-against-pr-changes-inline-as-ci-checks)). Stacked on #57.

- **Problem:** `npm run sim` reported every run as passed. It read a `results` field the simulation-run API doesn't return and counted `status === "pass"` (items are `passed`/`failed`), so it always summarised 0/0 and exited 0, failing runs included.
- **Who it affects:** anyone gating on `npm run sim`, locally or in CI. The PR check (later in this stack) reuses this verdict, so it has to be right.
- **What changes:**
  - A strict verdict (`src/sim-result.ts`).
  - Item fetching that handles both response shapes and late results.
  - The run link printed, plus each failing judge with expected vs extracted values.
  - `--timeout` and Ctrl-C both cancel the run.
  - Exit codes: 0 passed, 1 failed, 2 usage, 3 incomplete.
  - A config-free client that never retries run creation on a 5xx, because the run may already be queued.

## Evidence of value

**Same stub API, one passed and one failed item:**

| | Result |
|---|---|
| `main`'s `sim.ts` | `{"pass":0,"fail":0}` → exits 0 (false green) |
| This branch | `failed — 1 of 1 simulations failed` → exits 1 |

**Live, against a test org (chat transport):**

| Suite | Exit | Output |
|---|---|---|
| Designed to fail (judge: "open 24 hours?") | **1** | `✗ … open-24h (expected = true, got false)` — [run](https://dashboard.vapi.ai/simulations/run/9b933e3c-4ea6-4829-bedc-06bc795e6df3) |
| Designed to pass (judge: "open 8–5 on Fridays?") | **0** | `passed — 1 of 1 simulations passed` — [run](https://dashboard.vapi.ai/simulations/run/91bdb9d7-0bf5-4d07-87e3-bda9f5d1391b) |

The temporary resources were deleted afterwards.

## Testing plan

- `tests/sim-result.test.ts` is a verdict table covering:
  - the old false-green shape (no `results`);
  - 0 items, a short item list, and a count mismatch;
  - a failed item, with the failing judge listed;
  - canceled items;
  - all required evaluations skipped, and an optional skip alongside a scored required judge;
  - missing `itemCounts`, and a run that hasn't ended.
- `tests/sim-run.test.ts` runs `runSimulation` against a local HTTP stub:
  - pass, and fail using the bare-array item shape;
  - late item results;
  - timeout cancels the run, and an interrupt cancels it with the 400 "already ended" swallowed;
  - no retry of a 502 on create;
  - `--no-watch`;
  - pagination with overlapping pages deduped.
- `npm run build` and `npm test` pass (377 tests).
- **Not tested:** a live run that's still `running` when it's canceled (cancel was only exercised against the stub), and voice transport (the live runs used chat). Default behaviour change: unknown CLI arguments are now an error (exit 2) instead of being silently ignored.

Refs TEST-141

🤖 Generated with [Claude Code](https://claude.com/claude-code)
scott-lowe-vapi added a commit that referenced this pull request Oct 7, 2026
## Value

**V.A.L.U.E. tier:** project — PR 10 of 10 for inline simulation PR checks ([TEST-141](https://linear.app/vapi/issue/TEST-141/gitops-run-simulation-suites-against-pr-changes-inline-as-ci-checks)), the "check before deploy" step in promotion.

> Stacked on #65 (the promotion partial-failure fix), now that #57–#64 have merged. This PR is the single gate commit.

- **Problem:** promotion copies staging's reviewed files into production, but nothing checks that staging's agents still behave before they move on. Teams promoting dev → staging → prod need a behaviour gate between orgs, without new infrastructure.
- **Who it affects:** multi-org gitops users (the promotion pipeline), who get a "check before deploy" step with one line of `promotion.yml`. Single-org users are unaffected.
- **What changes:**
  - **`promotion.yml`** accepts `orgs.<slug>.check: <name>` (a slug), naming a `vapi-checks.yml` check.
  - **New `src/promotion-gate.ts`:**
    - **Validation before any transition:** the check must exist, and its `org` and `runOrg` must be the gated org, otherwise the run errors. `vapi-checks.yml` is required once any org is gated.
    - **Plan line:** what the gate would run, built offline.
    - **Live gate:** the check runs live and reduces to the worst target result.
  - **`src/promote-cmd.ts`:** in each transition, after the plan is built:
    - no changes skips the gate;
    - plan-only prints `check  would run <name> in <org> (<n> simulations × <t> targets)`;
    - `--apply` runs the check (after the bindings refresh, before `promotionPlanApply` writes anything). Any non-pass throws `Promotion out of <org> blocked: check <name> <outcome> (<run url>)`.
    - A pass is cached per source org and dropped once a transition applies into that org.
    - `promotionCommandRun(args, overrides)` now takes `Partial<PromotionDeps>` (`childRun`, `checkRun`).
  - **`.github/workflows/promotion.yml`:** `timeout-minutes: 90` on the "Reconcile configured promotions" **step**, not the job, so the `if: always()` commit step (fixed in #65) still runs after a blocked or slow gate.
  - **Docs:** `promotion.example.yml` (a commented `check:`), a README "Check before promoting" section, and a pointer from "PR Checks".

## Evidence of value

**The real gate, run live** in the owner's test org on the TEST-141 parity squad.
- **Setup:** a scratch repo whose `promotion.yml` gates `parity` on check `core`, with pipeline `parity → parity-prod`.
- **The run:** `promote --pipeline release --from parity --to parity-prod --apply`.
- **The fake:** the child runner was faked, so bindings pulls were no-ops and the downstream `apply.ts` was recorded but not run. No second org was needed or touched.

| Variant | Gate run | Result | Downstream apply | `resources/parity-prod/` |
|---|---|---|---|---|
| Degraded scheduler prompt | [7ed19587](https://dashboard.vapi.ai/simulations/run/7ed19587-d1ca-4d44-9232-7cdd15a50d67): 2 of 3 failed | `Promotion out of parity blocked: check core failed (https://dashboard.vapi.ai/simulations/run/7ed19587-…)` | **none** | **empty** (nothing written) |
| Fixture as-is | [95470670](https://dashboard.vapi.ai/simulations/run/95470670-2861-45d3-a483-7a220fc3591a): 3 of 3 passed | promoted | `["parity-prod"]` | written; 20 applied paths recorded |

The test org's resource counts were identical before and after both gate runs.

**Tests:** `npm test` goes from 484 (#65) to 492 passing, and #68's golden promotion test passes unchanged.

## Testing plan

- **`tests/promotion-gate.test.ts`** (6 tests, real git fixture, injected `childRun` / `checkRun`):
  - a pass applies;
  - failed and incomplete both block with the exact message, with no apply and the target untouched;
  - plan-only prints the line and runs nothing;
  - no changes skips the gate;
  - the three config errors (no `vapi-checks.yml`, unknown check, check in another org) stop before anything applies;
  - the pass cache: reused for two pipelines out of one org, and re-run after a transition applies into the gated org.
- **No gate configured, no change:** with no `check:` in `promotion.yml` and an **invalid** `vapi-checks.yml` present, plan and `--apply` both succeed, `checkRun` is never called, and the plan output equals a pinned string. That string is exactly what #65's code (before the gate existed) prints for the same fixture, which I confirmed by running #65's `promote-cmd` on it. So the gate is invisible unless someone opts in.
- **`tests/promotion.test.ts`:** `orgs.<slug>.check` is parsed, and a non-slug is rejected.
- **Not tested:**
  - **A real two-org promotion:** only one test org was available. The downstream apply was faked, so the blocked case shows nothing written, and the pass case shows the apply was called.
  - **A GitHub Actions promotion run with a gate**, including the step timeout firing.
- **Found while testing (pre-existing, out of scope):** promotion's dependency check rejects simulations that reference a **stock personality by UUID**, with "Referenced managed dependency is missing from source: personalities/a0000000-…". So a gated org's tests need local personality files until that's fixed.

Stacked on #65.

Refs TEST-141

## After review

The gate now refuses, when the config loads and before anything applies:

- an unknown key under an org in `promotion.yml`, so a misspelled `check:` can't silently drop the gate;
- `toolMocks: off` and `stripWebhooks: false`, because a gate runs in the real org, never a CI org;
- a check `baseUrl` that differs from the org's `baseUrl` in `promotion.yml`, so the org's key only goes to the host promotion uses (the gate always uses that host);
- a gate on an org that is last in every pipeline, where it would never run;
- gated checks whose combined budget is over 300 minutes.

It also fixes:

- **Deadline:** each batch of 3 targets gets a full `timeoutMinutes`, so a check with more than 3 targets is no longer falsely blocked as incomplete.
- **Step timeout:** raised from 90 to 330 minutes, as a safety net that no longer cuts short long ungated promotions.
- **Tests:** the deadline and the worst-target rule are pure helpers with their own tests.

The guide changes (blocks stop the whole run, simulation cost, the stock-personality limitation, accurate wording) are in #71. The `promote` User-Agent for gate runs is in #78. The block-report detail, fetch-stubbed gate test and deduplication are follow-ups.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants