Repository navigation
ci: type-check tests/ and send a User-Agent on sim API requests - #57
Merged
Merged
Conversation
- tsconfig.json now includes tests/, so `npm run build` (already run by CI on every PR) type-checks the tests. tsx runs them without type-checking, which is how 37 type errors accumulated unnoticed. - Fix those 37 errors. All were fixture drift after the hash-store migration (state entries carrying lastPulledHash/lastPushedHash or bare-string values, an untyped emptyLoaded fixture), plus one real gap: the reconcile-state-key harness never passed the required formatError, so any test reaching that error path would have thrown. - New src/user-agent.ts; `npm run sim` sends `User-Agent: vapi-gitops-sim/<version>`, so gitops-started simulation runs can be counted in the platform's run-started analytics. Refs TEST-141 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
Author
This was referenced Oct 1, 2026
scott-lowe-vapi
marked this pull request as ready for review
October 1, 2026 23:52
vtkovapi
approved these changes
Oct 3, 2026
Contributor
Author
Merge activity
|
scott-lowe-vapi
added a commit
that referenced
this pull request
Oct 3, 2026
…lse pass (#58) ## Value **V.A.L.U.E. tier:** project — PR 2 of 10 for inline simulation PR checks ([TEST-141](https://linear.app/vapi/issue/TEST-141/gitops-run-simulation-suites-against-pr-changes-inline-as-ci-checks)). Stacked on #57. - **Problem:** `npm run sim` reported every run as passed. It read a `results` field the simulation-run API doesn't return and counted `status === "pass"` (items are `passed`/`failed`), so it always summarised 0/0 and exited 0, failing runs included. - **Who it affects:** anyone gating on `npm run sim`, locally or in CI. The PR check (later in this stack) reuses this verdict, so it has to be right. - **What changes:** - A strict verdict (`src/sim-result.ts`). - Item fetching that handles both response shapes and late results. - The run link printed, plus each failing judge with expected vs extracted values. - `--timeout` and Ctrl-C both cancel the run. - Exit codes: 0 passed, 1 failed, 2 usage, 3 incomplete. - A config-free client that never retries run creation on a 5xx, because the run may already be queued. ## Evidence of value **Same stub API, one passed and one failed item:** | | Result | |---|---| | `main`'s `sim.ts` | `{"pass":0,"fail":0}` → exits 0 (false green) | | This branch | `failed — 1 of 1 simulations failed` → exits 1 | **Live, against a test org (chat transport):** | Suite | Exit | Output | |---|---|---| | Designed to fail (judge: "open 24 hours?") | **1** | `✗ … open-24h (expected = true, got false)` — [run](https://dashboard.vapi.ai/simulations/run/9b933e3c-4ea6-4829-bedc-06bc795e6df3) | | Designed to pass (judge: "open 8–5 on Fridays?") | **0** | `passed — 1 of 1 simulations passed` — [run](https://dashboard.vapi.ai/simulations/run/91bdb9d7-0bf5-4d07-87e3-bda9f5d1391b) | The temporary resources were deleted afterwards. ## Testing plan - `tests/sim-result.test.ts` is a verdict table covering: - the old false-green shape (no `results`); - 0 items, a short item list, and a count mismatch; - a failed item, with the failing judge listed; - canceled items; - all required evaluations skipped, and an optional skip alongside a scored required judge; - missing `itemCounts`, and a run that hasn't ended. - `tests/sim-run.test.ts` runs `runSimulation` against a local HTTP stub: - pass, and fail using the bare-array item shape; - late item results; - timeout cancels the run, and an interrupt cancels it with the 400 "already ended" swallowed; - no retry of a 502 on create; - `--no-watch`; - pagination with overlapping pages deduped. - `npm run build` and `npm test` pass (377 tests). - **Not tested:** a live run that's still `running` when it's canceled (cancel was only exercised against the stub), and voice transport (the live runs used chat). Default behaviour change: unknown CLI arguments are now an error (exit 2) instead of being silently ignored. Refs TEST-141 🤖 Generated with [Claude Code](https://claude.com/claude-code)
scott-lowe-vapi
added a commit
that referenced
this pull request
Oct 7, 2026
## Value **V.A.L.U.E. tier:** project — PR 10 of 10 for inline simulation PR checks ([TEST-141](https://linear.app/vapi/issue/TEST-141/gitops-run-simulation-suites-against-pr-changes-inline-as-ci-checks)), the "check before deploy" step in promotion. > Stacked on #65 (the promotion partial-failure fix), now that #57–#64 have merged. This PR is the single gate commit. - **Problem:** promotion copies staging's reviewed files into production, but nothing checks that staging's agents still behave before they move on. Teams promoting dev → staging → prod need a behaviour gate between orgs, without new infrastructure. - **Who it affects:** multi-org gitops users (the promotion pipeline), who get a "check before deploy" step with one line of `promotion.yml`. Single-org users are unaffected. - **What changes:** - **`promotion.yml`** accepts `orgs.<slug>.check: <name>` (a slug), naming a `vapi-checks.yml` check. - **New `src/promotion-gate.ts`:** - **Validation before any transition:** the check must exist, and its `org` and `runOrg` must be the gated org, otherwise the run errors. `vapi-checks.yml` is required once any org is gated. - **Plan line:** what the gate would run, built offline. - **Live gate:** the check runs live and reduces to the worst target result. - **`src/promote-cmd.ts`:** in each transition, after the plan is built: - no changes skips the gate; - plan-only prints `check would run <name> in <org> (<n> simulations × <t> targets)`; - `--apply` runs the check (after the bindings refresh, before `promotionPlanApply` writes anything). Any non-pass throws `Promotion out of <org> blocked: check <name> <outcome> (<run url>)`. - A pass is cached per source org and dropped once a transition applies into that org. - `promotionCommandRun(args, overrides)` now takes `Partial<PromotionDeps>` (`childRun`, `checkRun`). - **`.github/workflows/promotion.yml`:** `timeout-minutes: 90` on the "Reconcile configured promotions" **step**, not the job, so the `if: always()` commit step (fixed in #65) still runs after a blocked or slow gate. - **Docs:** `promotion.example.yml` (a commented `check:`), a README "Check before promoting" section, and a pointer from "PR Checks". ## Evidence of value **The real gate, run live** in the owner's test org on the TEST-141 parity squad. - **Setup:** a scratch repo whose `promotion.yml` gates `parity` on check `core`, with pipeline `parity → parity-prod`. - **The run:** `promote --pipeline release --from parity --to parity-prod --apply`. - **The fake:** the child runner was faked, so bindings pulls were no-ops and the downstream `apply.ts` was recorded but not run. No second org was needed or touched. | Variant | Gate run | Result | Downstream apply | `resources/parity-prod/` | |---|---|---|---|---| | Degraded scheduler prompt | [7ed19587](https://dashboard.vapi.ai/simulations/run/7ed19587-d1ca-4d44-9232-7cdd15a50d67): 2 of 3 failed | `Promotion out of parity blocked: check core failed (https://dashboard.vapi.ai/simulations/run/7ed19587-…)` | **none** | **empty** (nothing written) | | Fixture as-is | [95470670](https://dashboard.vapi.ai/simulations/run/95470670-2861-45d3-a483-7a220fc3591a): 3 of 3 passed | promoted | `["parity-prod"]` | written; 20 applied paths recorded | The test org's resource counts were identical before and after both gate runs. **Tests:** `npm test` goes from 484 (#65) to 492 passing, and #68's golden promotion test passes unchanged. ## Testing plan - **`tests/promotion-gate.test.ts`** (6 tests, real git fixture, injected `childRun` / `checkRun`): - a pass applies; - failed and incomplete both block with the exact message, with no apply and the target untouched; - plan-only prints the line and runs nothing; - no changes skips the gate; - the three config errors (no `vapi-checks.yml`, unknown check, check in another org) stop before anything applies; - the pass cache: reused for two pipelines out of one org, and re-run after a transition applies into the gated org. - **No gate configured, no change:** with no `check:` in `promotion.yml` and an **invalid** `vapi-checks.yml` present, plan and `--apply` both succeed, `checkRun` is never called, and the plan output equals a pinned string. That string is exactly what #65's code (before the gate existed) prints for the same fixture, which I confirmed by running #65's `promote-cmd` on it. So the gate is invisible unless someone opts in. - **`tests/promotion.test.ts`:** `orgs.<slug>.check` is parsed, and a non-slug is rejected. - **Not tested:** - **A real two-org promotion:** only one test org was available. The downstream apply was faked, so the blocked case shows nothing written, and the pass case shows the apply was called. - **A GitHub Actions promotion run with a gate**, including the step timeout firing. - **Found while testing (pre-existing, out of scope):** promotion's dependency check rejects simulations that reference a **stock personality by UUID**, with "Referenced managed dependency is missing from source: personalities/a0000000-…". So a gated org's tests need local personality files until that's fixed. Stacked on #65. Refs TEST-141 ## After review The gate now refuses, when the config loads and before anything applies: - an unknown key under an org in `promotion.yml`, so a misspelled `check:` can't silently drop the gate; - `toolMocks: off` and `stripWebhooks: false`, because a gate runs in the real org, never a CI org; - a check `baseUrl` that differs from the org's `baseUrl` in `promotion.yml`, so the org's key only goes to the host promotion uses (the gate always uses that host); - a gate on an org that is last in every pipeline, where it would never run; - gated checks whose combined budget is over 300 minutes. It also fixes: - **Deadline:** each batch of 3 targets gets a full `timeoutMinutes`, so a check with more than 3 targets is no longer falsely blocked as incomplete. - **Step timeout:** raised from 90 to 330 minutes, as a safety net that no longer cuts short long ungated promotions. - **Tests:** the deadline and the worst-target rule are pure helpers with their own tests. The guide changes (blocks stop the whole run, simulation cost, the stock-personality limitation, accurate wording) are in #71. The `promote` User-Agent for gate runs is in #78. The block-report detail, fetch-stubbed gate test and deduplication are follow-ups. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Value
V.A.L.U.E. tier: project — PR 1 of 10 for inline simulation PR checks (TEST-141); this PR is a small, behavior-preserving slice.
tests/was never type-checked (tsconfig.jsonincluded onlysrc/), so 37 type errors piled up silently. That's the gap test: restore the suite after the hash-store migration and run it in CI #56 named as follow-up inimprovements.mdfix(push,pull): recanonicalize stale UUID-suffixed state keys — root-cause duplicate generation #32. Separately, simulation runs started from gitops can't be told apart in the platform's analytics.npm run build, which CI already runs on every PR, now type-checkstests/.npm run simsendsUser-Agent: vapi-gitops-sim/<version>. The API records this asuser_agenton the[simulation] run startedevent.The CI workflow itself landed in #56, so this PR is smaller than PR 1 in the plan.
Evidence of value
main(69c7e83)tsc --noEmitwithtests/includednpm testUser-AgentonPOST /eval/simulation/runvapi-gitops-sim/1.0.0(asserted against a local HTTP server)What the 37 errors were:
lastPulledHash/lastPushedHash, bare-string state values, and an untypedemptyLoaded().reconcile-state-keyharness never passed the requiredformatError, so any test reaching that error path would have thrown aTypeErrorinstead of testing it.Tests that used the removed hash fields as markers (
state-merge,recanonicalize) now mark "which copy won" with distinct UUIDs or object identity, so they still check the same behavior.Testing plan
npm run build(now coverssrc/andtests/) andnpm testlocally on Node 22: green. CI on this PR runs both on Node 20 and 22.tests/user-agent.test.tscovers the header format againstpackage.json's version, and the header actually sent on run create.sim.test.tsnow covers the legacy bare-string state value directly, replacing the old cast-based "forward-compat" test.src/behavior is unchanged apart from the added header.Refs TEST-141
🤖 Generated with Claude Code