diff --git a/pstack/.cursor-plugin/plugin.json b/pstack/.cursor-plugin/plugin.json index 1fe039fcc..5dfdc6537 100644 --- a/pstack/.cursor-plugin/plugin.json +++ b/pstack/.cursor-plugin/plugin.json @@ -1,7 +1,7 @@ { "name": "pstack", "displayName": "pstack", - "version": "0.15.12", + "version": "0.15.13", "description": "if you want to go fast, go deep first. pstack helps you write less, but higher quality code. rigorous agent workflows you can parallelize with confidence.", "author": { "name": "Lauren Tan" diff --git a/pstack/docs/guide/01-setup.md b/pstack/docs/guide/01-setup.md index df9997272..054ddbdfe 100644 --- a/pstack/docs/guide/01-setup.md +++ b/pstack/docs/guide/01-setup.md @@ -30,10 +30,21 @@ You might be wondering what happens if you use Auto. Set a role to `inherit-pare At the end of setup, `/setup-pstack` looks for a way to prove app behavior in your project, either a `verify-*` skill or an existing harness. If it finds neither, it offers once to generate one with [`/create-verification-skill`](../../skills/create-verification-skill/SKILL.md). -Say yes and it writes `.cursor/skills/verify-/`, a project-local skill that teaches agents to drive your app the way a user does. It proves the skill works once before handing it over. Say no and setup moves on. You can run `/create-verification-skill` yourself any time. [Verify and ship](./06-verify-and-ship.md#create-a-project-verification-skill) covers when it earns its place. +Say yes and it writes `.cursor/skills/verify-/`, a project-local skill that teaches agents to drive your app the way a user does. It proves the skill works once before handing it over. Say no and setup moves on. You can run `/create-verification-skill` yourself any time. [Verify and ship](./06-verify-and-ship.md#create-a-project-verification-skill) covers it in depth. + +If you're new to pstack, say yes. An agent that can check its own work keeps going until the check passes. An agent that can't hands every result back to you to check by hand. Of everything in this guide, the verification skill pays off the most. After setup, start a new chat. The model rule applies to new sessions. +## Keep the cost in check + +pstack spends extra tokens on subagents and review panels. That's the price of the rigor. To spend fewer: + +- Rerun `/setup-pstack` and pick a smaller reasoning budget or cheaper models. A strong model in the main chat with cheaper, faster models in the code roles is a good split. +- Set a role to `auto` or `inherit-parent` so it runs on the chat's own model. +- Shorten a panel list. Each entry runs one subagent. +- Save `/poteto-mode` for work that needs rigor. A small, obvious edit doesn't. + ## Run your first task Pick something real but small, and describe it the way you'd describe it to a colleague: diff --git a/pstack/docs/guide/02-poteto-mode.md b/pstack/docs/guide/02-poteto-mode.md index 71e7ab1ad..8341604ae 100644 --- a/pstack/docs/guide/02-poteto-mode.md +++ b/pstack/docs/guide/02-poteto-mode.md @@ -37,6 +37,37 @@ You don't write a spec. You say what's wrong or what you want, plus anything you That's a Bug fix prompt. "repro first" is a real constraint, not politeness, and the playbook honors it. Watch the todo list fill with the Bug fix steps. A skipped step stays visible with `skip: `. +## What goes in a prompt + +A useful prompt carries up to five things, and each one fits in a sentence: + +- **The goal.** Say what's wrong, or what you want. +- **The done check.** It must be able to pass or fail. "Make it better" and "work on it for an hour" aren't checks. +- **The proof you want to see.** Ask for the real command output, a video of the flow, the stored value, or a before-and-after number. +- **What you already know.** A symptom, a repro step, a log line, or a link saves the agent a search. +- **The real constraints.** "repro first", "don't change any code yet", "zero behavior change", and "let me review before proceeding" each change what the agent does. + +Here's one prompt with all five: + +```text +/poteto-mode the csv export drops its last row since yesterday's deploy. failing job id is 4812. repro first, then fix. done means the 60k-row fixture exports every row. show me the row counts before and after. +``` + +Two things are worth leaving out: + +- **The how.** Say what to achieve, and leave the agent room to find a better path than the one you'd pick. The same goes for a list of skills, covered in the pitfall below. +- **Your theory of the cause, at first.** A stated guess narrows the search to wherever you pointed. Let the agent restate the problem before you share your hunch. + +For a noisy report, such as a long thread or a vague bug, make the restatement the first step: + +```text +/poteto-mode read this thread. restate the underlying issue in your own words, in plain english. don't change any code yet. +``` + +A misreading shows up in the restatement, before any code exists. Correct it there, and it costs you one message instead of one wrong fix. + +## Follow up short + When the conversation already carries the context, the prompt shrinks to almost nothing. All of these are enough: ```text @@ -63,15 +94,17 @@ A long chat accumulates context from the last task. When you change subjects, sa "new task" tells `/poteto-mode` to re-match rather than continue the prior playbook. "don't change any code yet" pins this one to Investigation. Without those two phrases, a mode mid-Feature tends to treat your question as the next feature step. -## Give parallel work its own worktree +## Give parallel work its own machine + +If you run several agents against one repository on one computer, they will fight over the working tree, the ports, and the build output. The cleanest isolation is a [cloud subagent](https://cursor.com/docs/subagents#cloud-subagents). Each one gets its own VM and branch, so it can install dependencies, run your app, and record video of the result without touching your machine. Type `/in-cloud` before the task, or ask the parent chat to hand work to cloud subagents. -If you run several agents against one repository, they will fight over the working tree. Ask for isolation up front: +When the work has to stay local, ask for a worktree up front: ```text /poteto-mode new task. branch off in a fresh worktree, then port the parser change there. ``` -Each task in its own branch and worktree means no agent stomps another's files. The [Opening a PR playbook](../../skills/poteto-mode/playbooks/opening-a-pr.md) already works from a worktree for code changes, so mostly you only say this when a specific base or location matters. +Each task in its own branch and worktree means no agent stomps another's files. Worktrees cost disk and machine resources, so a laptop runs only a handful at once. The [Opening a PR playbook](../../skills/poteto-mode/playbooks/opening-a-pr.md) already works from a worktree for code changes, so mostly you only say this when a specific base or location matters. Worktrees accumulate. When disk gets tight, ask: diff --git a/pstack/docs/guide/03-understand.md b/pstack/docs/guide/03-understand.md index af985988c..e15bea052 100644 --- a/pstack/docs/guide/03-understand.md +++ b/pstack/docs/guide/03-understand.md @@ -1,9 +1,21 @@ # Understand the code before changing it -Editing code you don't understand is how subtle regressions ship. pstack gives you four ways in. `/how` explains what the code does now. `/why` digs up the reasons it's shaped that way. `/teach` blends both into one explanation. `/recall` rebuilds your own recent context on a topic. +Editing code you don't understand is how subtle regressions ship, and that's as true for the agent as for you. Agents usually fail in one of two ways. They misread what you want, or they don't have the context to do the work right. [What goes in a prompt](./02-poteto-mode.md#what-goes-in-a-prompt) handles the first. This page handles the second. + +pstack gives you four ways in. `/how` explains what the code does now. `/why` digs up the reasons it's shaped that way. `/teach` blends both into one explanation. `/recall` rebuilds your own recent context on a topic. Each one also makes the agent explain itself in words you can check. That's how you supervise an agent that may know the code better than you do. ![A detective studies a machine blueprint with a magnifying glass while robots fetch case files; the evidence board behind her links clues under /how and /why.](./images/understanding.jpg) +## Start with a read-only investigation + +When the cause is unclear, ask for findings, not a fix: + +```text +/poteto-mode investigate why background jobs time out every few hours. give me what we know, what data you used, and your best hypotheses. don't change any code yet. +``` + +"don't change any code yet" routes this to the [Investigation playbook](../../skills/poteto-mode/playbooks/investigation.md). It runs `/how`, adds `/why` for questions about motivation, and returns a cited explanation. For a choice between options, it returns a recommendation with a trade-offs table. Asking "what data you used" makes the agent separate its evidence from its guesses. When the findings point at a fix, start the fix as a new task. + ## Trace behavior with `/how` ```text @@ -30,13 +42,27 @@ The two compose naturally. `do why first then how` is a perfectly good prompt wh [`/teach`](../../skills/teach/SKILL.md) is for when a summary isn't enough. It runs `/how` and `/why`, for a small change maybe just one of them, and weaves the findings into a plain explanation that builds up diagram by diagram. The "convince me" framing is worth stealing. It turns the explanation into an argument you can poke at instead of a tour. +It works on the agent's own choices too: + +```text +/teach me why you implemented it this way and not with a queue. what did you trade off, and why? +``` + +Teaching helps the agent as much as you. An agent that has to explain its work must read the code and back each claim with evidence, instead of stating it confidently and moving on. + ## Rebuild your own context with `/recall` ```text /recall catch me up on the export work from last week ``` -[`/recall`](../../skills/recall/SKILL.md) mines your own recent chats plus the shared record (issues, prior fixes, errors still firing) and hands back a brief on where things stand and what's next. Use it when you're returning to a topic cold. If you want to resume one specific chat, that's the Session pickup playbook below, not `/recall`. +[`/recall`](../../skills/recall/SKILL.md) mines your own recent chats plus the shared record (issues, prior fixes, errors still firing) and hands back a brief on where things stand and what's next. Your old chats hold context that a fresh agent lacks, so start new work on an old topic by loading it first, then hand over the new input: + +```text +/recall my work on the virtualized list from yesterday, then read this bug report. +``` + +If you want to resume one specific chat, that's the Session pickup playbook below, not `/recall`. ## Take over prior work with Session pickup diff --git a/pstack/docs/guide/04-design.md b/pstack/docs/guide/04-design.md index 681d39cff..b0cf10f13 100644 --- a/pstack/docs/guide/04-design.md +++ b/pstack/docs/guide/04-design.md @@ -2,6 +2,8 @@ One attempt at a hard design locks in the first shape the model thought of. `/architect` settles types and boundaries before implementation. `/arena` runs several attempts at the same brief and merges the best parts. `/interrogate` has other models try to break the result. When the job is coverage rather than design synthesis, `/swarm` fans out slices or races and aggregates their results. +The two most common design mistakes are taking the agent's first design and polishing a plan that no code has tested. This page fixes both. You plan through code: prototypes answer the open questions, a README or tutorial sets the target, and the written plan comes last. + ![Three robots draft competing bridge models at their own tables under /architect, /arena, and /interrogate panels, while a judge robot with a clipboard inspects skeptically.](./images/design.jpg) ## Settle the shape with `/architect` @@ -18,6 +20,8 @@ By default it proceeds straight from the synthesized design into implementation. /architect with checkpoint. stop and show me before implementing. ``` +The design isn't sacred once code starts. If implementation shows the same workaround in unrelated places, or types that only compile with `any` or forced casts, `/architect` treats that as proof the design is wrong. It scraps the sketch and starts over instead of patching around it. + ## Fan out attempts with `/arena` ```text @@ -66,6 +70,52 @@ Reach for it when parallelism buys coverage or lets independent checks race. `/a Read the dismissals too. The lead is a pragmatic senior engineer, not an oracle, and you can override it. +## Prototype instead of debating + +Never take the first design. Ask for a few, and pick from evidence you can see: + +```text +/poteto-mode prototype a few options for the new dropdown menu. take screenshots or videos for me to compare. +``` + +The [Prototype playbook](../../skills/poteto-mode/playbooks/prototype.md) builds throwaway sketches in a scratch directory, puts the variants behind one switcher, drives each one, and captures screenshots or timings. It also works for behavior and algorithms, not just UI. Prototypes are planning with code. They let the agent answer its own open questions by running something instead of asking you, and they leave room for an option you wouldn't have thought of. + +The same idea scales up to a real design. Pair `/architect` with prototypes and keep a review gate: + +```text +/poteto-mode we need rate limiting for external webhooks. /architect it first, and answer open questions with prototypes. let me review before proceeding. +``` + +Don't spend reviewers on an abstract plan. `/interrogate` belongs on a diff. Point adversarial review at a plan with no code behind it and the reviewers invent theoretical risks and edge cases that will never happen. Let prototypes settle the questions, then review what got built. + +## Write the README first for shared code + +For a package or API that other code will use, start with the doc a user would read: + +```text +/poteto-mode write a tutorial for how i would use the new config package first. then /teach me why it beats the current one. +``` + +Writing the tutorial first forces the caller's view. You describe the API to a hypothetical user and work back to the implementation. The doc also becomes a concrete target the agent checks its own work against. Name [`/technical-writing`](../../skills/technical-writing/SKILL.md) when the doc itself matters, so a tutorial stays a tutorial instead of drifting into reference and explanation at once. + +## Plan after the design settles + +pstack has no planning skill, on purpose. When you do want a written plan, ask for it once the design is settled: + +```text +/poteto-mode turn this design into a plan. small verifiable PRs, each with its own verification steps. +``` + +The [Multi-phase plan playbook](../../skills/poteto-mode/playbooks/multi-phase-plan.md) settles any remaining open questions by prototype, then writes one section per PR, each ending in proof that the change works. A passing test suite alone doesn't count as that proof. The plan is the deliverable. The playbook doesn't implement it, and it names which execution playbook should run it next. + +For a migration, state the bar in the prompt: + +```text +/poteto-mode plan the migration of our ui library to the new styling system. small verifiable PRs, each with visual regression checks. the result must match the original exactly, bugs included. +``` + +"bugs included" keeps the migration from quietly fixing things on the way, which would make the old and new output impossible to compare. For a project that spans many days, you can commit the plan to the repo for a while so other agents see the work in progress. Delete it when the work lands. + ## How much design work does a task deserve? You might be wondering whether every change needs this. No. Most changes need none of it. A rough ladder: @@ -74,7 +124,9 @@ You might be wondering whether every change needs this. No. Most changes need no - A change that crosses function boundaries or moves ownership earns `/architect`, which brings `/arena` with it. - A standalone decision where independent attempts would help, like naming, formats, or an algorithm, is `/arena` directly. - A coverage matrix, set of parallel checks, or race with declared arms is `/swarm`. +- An open question you could answer by running something, like a layout, a timing, or an approach, gets a prototype, not a debate. - A contested design that's expensive to reverse gets `/architect`, then `/interrogate` before shipping. +- Work that spans several PRs gets a plan, written after the design settles. `/poteto-mode` already applies this ladder. Boundary-crossing work triggers `/architect` on its own, so you reach for these directly mainly when you want more or less scrutiny than the default. diff --git a/pstack/docs/guide/05-build-and-clean.md b/pstack/docs/guide/05-build-and-clean.md index 57d22572b..fc48a46df 100644 --- a/pstack/docs/guide/05-build-and-clean.md +++ b/pstack/docs/guide/05-build-and-clean.md @@ -32,6 +32,14 @@ Each of these routes to its playbook ([Bug fix](../../skills/poteto-mode/playboo For sustained improvement of one number, there's the [Hillclimb playbook](../../skills/poteto-mode/playbooks/hillclimb.md). Give it the metric, a target, and a floor on attempts, and it loops one hypothesis at a time with a frozen measurement harness. It keeps wins and reverts everything else. +Both perf playbooks run [`/benchmark-checklist`](../../skills/benchmark-checklist/SKILL.md) on their numbers. Perf issue vets its baseline and every number after it, and Hillclimb vets its harness before freezing it. [Verify and ship](./06-verify-and-ship.md#vet-a-measured-number-with-benchmark-checklist) shows when to type it yourself. + +Sometimes you want the cause before any fix. For a live symptom, such as a leak, an idle CPU spin, or a visual glitch, the [Runtime forensics playbook](../../skills/poteto-mode/playbooks/runtime-forensics.md) instruments the running process. For a profile you already captured, the [Trace forensics playbook](../../skills/poteto-mode/playbooks/trace-forensics.md) reads the artifact and maps the hot frame to source. Both return a diagnosis, not a fix: + +```text +/poteto-mode here's a cpuprofile from the slow startup. tell me where the time goes and which source lines own it. no fix yet. +``` + ## Write the failing test first with `/tdd` When a bug has a cheap local test path, the whole prompt can be two words: diff --git a/pstack/docs/guide/06-verify-and-ship.md b/pstack/docs/guide/06-verify-and-ship.md index d96c76612..007ea2cab 100644 --- a/pstack/docs/guide/06-verify-and-ship.md +++ b/pstack/docs/guide/06-verify-and-ship.md @@ -1,6 +1,8 @@ # Verify the result and open a PR -"It compiles" is not evidence. The [Prove It Works principle](../../skills/principle-prove-it-works/SKILL.md) makes the agent check the real artifact before it reports success, and your job is to make "the real artifact" checkable. This page covers stating a finish condition, generating a verification skill for your app, opening the PR, and driving it to merged. +"It compiles" is not evidence. The [Prove It Works principle](../../skills/principle-prove-it-works/SKILL.md) makes the agent check the real artifact before it reports success, and your job is to make "the real artifact" checkable. This page covers stating a finish condition, vetting a measured number, generating a verification skill for your app, opening the PR, and driving it to merged. + +Verification is the slowest step in most agent work, because it's the step that usually waits on a human. Make the agent able to do it, and you stop being the bottleneck. Skip it, and running more agents only gets you more unchecked work to review. ![A prototype plane flies a real test course while she times it with a stopwatch and robots film and checklist the run; the terminal reads verify: pass, evidence: captured.](./images/verification.jpg) @@ -17,13 +19,35 @@ Now the agent has three checks it can run, not a mood to satisfy. When the reply Match the check to the change: - A CLI change runs the real command. -- A UI change walks the changed flow in the running app. +- A UI change walks the changed flow in the running app. When it must match a reference pixel for pixel, the [Visual parity playbook](../../skills/poteto-mode/playbooks/visual-parity.md) diffs screenshots against a frozen baseline instead of judging by eye. - A parser or migration replays a saved input. - A perf change compares before and after profiles. - A storage change reads back the written value. +Ask for the proof as an artifact you can inspect yourself: the failing test and then the passing one, a before-and-after video, the trace, the screenshot. If the fix already merged, ask for the same check again on main. An artifact beats a plausible explanation, because you can challenge it without replaying the whole run. + For a small diff you don't fully trust, [`/blast-radius`](../../skills/blast-radius/SKILL.md) finds what it could break elsewhere. It picks the one fact the change is safe because of and proves it by running code instead of writing an essay about it. +## Vet a measured number with `/benchmark-checklist` + +A before-and-after number is the easiest evidence to get wrong by accident. A warm cache, a debug build on one side, or work that never ran inside the timed region can each produce a convincing speedup. Before you report or act on a number, type: + +```text +/benchmark-checklist vet the export speedup before it goes in the pr +``` + +[`/benchmark-checklist`](../../skills/benchmark-checklist/SKILL.md) asks seven questions and wants evidence from a run for each: + +1. What limits the number, and why isn't it double? +2. Did every side run tuned the way production runs? +3. Does the result break a physical limit, like disk bandwidth or core count? +4. Did anything error or return wrong output? +5. Does it reproduce over alternating runs, with a median and a range? +6. Does it matter end to end, on the path a user waits on? +7. Did the work actually happen inside the timed region? + +The verdict comes back as faster, slower, no measurable difference, or inconclusive, with the run count, range, and limiter. It says inconclusive when it can't name the limiter or a side ran untuned. `/poteto-mode` already runs the checklist inside the Perf issue and Hillclimb playbooks, so you type it yourself when you measured something outside them, or when someone else's number looks too good. It's the working form of the [Explain the Number principle](../../skills/principle-explain-the-number/SKILL.md). + ## Create a project verification skill The UI bullet above hides a real requirement. The agent needs a scripted way to drive your app. If your project has one, great. If not, run: @@ -36,13 +60,32 @@ The UI bullet above hides a real requirement. The agent needs a scripted way to It writes `.cursor/skills/verify-/`, agent-facing instructions with exact Launch, Doctor, Drive, Evidence, and Cleanup sections, plus a feature map under `features/` that indexes what the app does and what result proves each feature works. The skill ships a [worked feature-map example](../../skills/create-verification-skill/references/feature-map-example/) with a README index and one file per feature using the four required H2s. Before handing it over, the generator proves the skill once end to end: launch, doctor check, drive one feature, capture evidence, clean up. If that proof fails, don't use the output. -From then on, "verify it in the app" is a step any agent can execute, in this repo, with no setup conversation. +From then on, "verify it in the app" is a step any agent can execute, in this repo, with no setup conversation. Name it in the prompt when you want the proof in a specific form: + +```text +/poteto-mode build the bulk-archive action. use /verify- to verify your changes and show me a video and screenshots as proof. +``` + +```text +/poteto-mode repro this with /verify-. if it repros on main, fix it and show me a video as proof. +``` + +Once the verify skill works, a [`/swarm`](../../skills/swarm/SKILL.md) can split a full pass by feature-map entry and aggregate the results. A swarm of verifiers also confirms a perf win over a big enough sample, or fuzzes the app for regressions before a PR ships. + +Treat the verification skill as infrastructure, not a one-off. Commit it, so every person and every agent on the team drives the app the same way. Then [build the lever](../../skills/principle-build-the-lever/SKILL.md). When agents keep writing throwaway scripts to click through the app, ask for a small control CLI that the skill calls instead. Agents spend fewer tokens, and every run becomes repeatable. A CLI that agents use well has these traits: + +- A few composable commands, each doing real work, rather than many thin ones. +- A `--dry-run` option on anything destructive. +- Subcommands that reveal features gradually instead of all at once. +- Error messages that say what to do instead. +- Rich `--help` text. +- Machine-readable output, such as JSON. -Once the verify skill works, a [`/swarm`](../../skills/swarm/SKILL.md) can split a full pass by feature-map entry and aggregate the results. +While you're there, make the dev setup repeatable too: seeded data, test users, and one command that brings the environment up the same way every time. ## Keep the verification skill honest -Apps change and feature maps rot. When yours drifts, run: +Apps change and feature maps rot. Run this at least once a day, ideally from a scheduled automation so nobody has to remember: ```text /maintain-verification-skill diff --git a/pstack/docs/guide/07-overnight.md b/pstack/docs/guide/07-overnight.md index 9ff66b770..cbad06723 100644 --- a/pstack/docs/guide/07-overnight.md +++ b/pstack/docs/guide/07-overnight.md @@ -1,9 +1,20 @@ # Run work while you sleep -This is the payoff for everything before it. An agent you can trust to verify its own work is an agent you can leave alone with a hard task. What makes that safe isn't hope. It's a checkable finish condition, an isolated worktree, and a decision log you audit in the morning. +This is the payoff for everything before it. An agent you can trust to verify its own work is an agent you can leave alone with a hard task. What makes that safe isn't hope. It's a checkable finish condition, an isolated worktree or cloud agent, and a decision log you audit in the morning. ![She waves goodnight from the door while robots keep the factory running, one updating a DECISION LOG wall board under a BUILD LOOP ACTIVE sign.](./images/overnight.jpg) +## Earn the trust before the loop + +A loop you don't trust just produces unchecked work faster, and the mess compounds with every iteration. Before you leave one running, check that it has earned it: + +- You've done the task once by hand, or watched an agent do it, so you know what good looks like. +- The agent has the tools and signals you'd use yourself: the verification skill, the profiler, the logs. +- Every stage proves its work and can stop the line when the work misses the bar. +- You've read a few transcripts and turned the repeated failures into tools, skills, or checks. + +Make the loop autonomous only after all four hold. Until then, run it while you watch. + ## The overnight contract A good handoff has the goal, the finish condition, permissions, and an escape hatch. It doesn't need to be long: @@ -26,6 +37,8 @@ Walk through what each line buys you: Because you'll review this work after stepping away, `/poteto-mode` routes it through [`/figure-it-out`](../../skills/figure-it-out/SKILL.md), which designs the run's phases before any code and wires in the decision log. +To stop a run on purpose, tell the agent to pause, or that you're about to go offline or restart Cursor. The [Pause safely playbook](../../skills/poteto-mode/playbooks/pause-safely.md) finishes or backs out of the current step, commits a work-in-progress checkpoint, and writes a resume note. A fresh chat picks the work up from that note through the Session pickup playbook. Saying "keep going" never triggers a pause. + ## What the loop does all night ```mermaid @@ -76,6 +89,32 @@ The contract above drives one task to one finish condition. Some nights hold mor /poteto-mode orchestrate the store migration. own it until every package is converted and merged. i'll check in twice a day. ``` +## Run many projects in parallel + +A [Cursor Project](https://cursor.com/blog/projects) gives one coordinator agent a persistent thread. The coordinator doesn't write code. It directs subagents, which run in the cloud by default, so the work continues when your laptop is closed. That's the shape the Orchestrate playbook expects. Start your prompts to the coordinator with `/poteto-mode`, and the subagents it spawns follow the playbooks. + +A few habits help: + +- Give each body of work its own Project, such as a feature, a migration, a perf push, or a tech-debt cleanup. Several can run side by side. +- Drag related chats into the Project, finished ones included. They become context for every agent in it. +- Give each PR a verification swarm before it merges, and let Autopilot-stack or Autopilot-full carry the queue. +- Ask the coordinator for a plan backed by data, and have it answer open questions with prototypes before it asks you. + +One prompt can carry a whole Project, from research through execution: + +```text +/poteto-mode refactor this repo so its architecture is more agent friendly. use /correct and /architect on past commits and review comments to find the mistakes agents make most here. use /recall for context from past chats. answer open questions with prototypes instead of asking me. come back with a plan backed by real data. once i approve it, run it with autopilot-stack or autopilot-full, and ask me which. +``` + +## Let loops start themselves + +Every loop above still waits for you to start it. A scheduled or event-driven automation removes that step. Software maintenance splits into stages that suit this well: triage a report, reproduce it, fix it, verify the fix. Two rules keep such a line trustworthy: + +- Every stage can stop the line. Triage can decide the report is expected behavior, repro can fail to reproduce it, and the fixer can judge the change too risky. Each of those outcomes is useful, because it keeps bad work from reaching the next stage, where it costs more to undo. +- Every stage hands over evidence. Repro attaches screenshots and video of the broken state, and the fix attaches before-and-after proof. A human can then check that the agent fixed the right thing before reading a line of code. + +pstack ships this as a dormant [automation pack](../../automations/benny/README.md) for Slack issue reports. One automation triages each report. The other reproduces confirmed bugs and may prepare a small draft fix. Point an agent at its [`FOR_AGENTS.md`](../../automations/benny/FOR_AGENTS.md) and name the target repository to set it up. + **Pitfall:** a duration is not a finish condition. "work on this for 4 hours" gives the agent nothing to check, and you'll wake up to four hours of motion instead of a result. Give `/loop` a predicate that can pass or fail. Next: [Steer with principle names](./08-principles.md). diff --git a/pstack/docs/guide/08-principles.md b/pstack/docs/guide/08-principles.md index 283a5e66b..0aaec930c 100644 --- a/pstack/docs/guide/08-principles.md +++ b/pstack/docs/guide/08-principles.md @@ -39,7 +39,7 @@ The core principles decide how much to build and when to rethink the design: - [Outcome-Oriented Execution](../../skills/principle-outcome-oriented-execution/SKILL.md) converges rewrites on the target design instead of preserving throwaway compatibility states. - [Experience First](../../skills/principle-experience-first/SKILL.md) chooses the user's result over implementation convenience. - [Exhaust the Design Space](../../skills/principle-exhaust-the-design-space/SKILL.md) builds two or three competing prototypes when there's no precedent. -- [Build the Lever](../../skills/principle-build-the-lever/SKILL.md) builds the script that does or proves the work, so a reviewer can rerun it. +- [Build the Lever](../../skills/principle-build-the-lever/SKILL.md) builds the script that does or proves the work, so a reviewer can rerun it. When an agent keeps doing the same thing by hand, have it write the tool or skill it wishes it had. If a script can do a step the same way every time, use the script, and save agents for the judgment calls. The architecture principles decide where state, validation, and compatibility live: @@ -56,7 +56,7 @@ The verification principles define what counts as proof: - [Fix Root Causes](../../skills/principle-fix-root-causes/SKILL.md) reproduces and traces to the cause before changing code. - [Sequence Work into Verifiable Units](../../skills/principle-sequence-verifiable-units/SKILL.md) ends each small unit in a check before starting the next. - [Test Behavior, Not Implementation](../../skills/principle-test-behavior-not-implementation/SKILL.md) calls the code the way its users do and asserts a literal expected value, and deletes a test that would still pass if every imported function returned `undefined`. -- [Explain the Number](../../skills/principle-explain-the-number/SKILL.md) names what limits a measured number and rules out that it measured something else, before anyone trusts or reports it. +- [Explain the Number](../../skills/principle-explain-the-number/SKILL.md) names what limits a measured number and rules out that it measured something else, before anyone trusts or reports it. [`/benchmark-checklist`](../../skills/benchmark-checklist/SKILL.md) turns it into seven questions you answer from real runs. The delegation principles keep parallel work sane: @@ -65,7 +65,7 @@ The delegation principles keep parallel work sane: And one meta principle: -- [Encode Lessons in Structure](../../skills/principle-encode-lessons-in-structure/SKILL.md) turns advice you've repeated twice into a lint, check, or script. +- [Encode Lessons in Structure](../../skills/principle-encode-lessons-in-structure/SKILL.md) turns advice you've repeated twice into a lint, check, or script. [`/correct`](../../skills/correct/SKILL.md) applies it to a whole repo, as shown in [Make it yours](./09-make-it-yours.md#fix-the-environment-with-correct). Don't memorize the list. Skim it now, then come back when you catch the agent doing something a name here would have prevented. That's how the vocabulary sticks. diff --git a/pstack/docs/guide/09-make-it-yours.md b/pstack/docs/guide/09-make-it-yours.md index 6723069b5..321caaf91 100644 --- a/pstack/docs/guide/09-make-it-yours.md +++ b/pstack/docs/guide/09-make-it-yours.md @@ -1,6 +1,8 @@ # Make it yours -poteto-mode is one person's style. The machinery underneath, playbooks, routing, model roles, works just as well wearing yours. This page covers generating a personal mode, capturing lessons from a session, authoring a focused skill, and testing a skill change before you trust it. +poteto-mode is one person's style. The machinery underneath, playbooks, routing, model roles, works just as well wearing yours. This page covers generating a personal mode, capturing lessons from a session, fixing the repo so agents stop repeating mistakes, authoring a focused skill, and testing a skill change before you trust it. + +Start smaller than you think. You don't need many skills on day one, or even this whole plugin. Prompt plainly, watch where agents fail, and add a skill or a check when the same failure shows up twice. ## Generate your own mode with `/automate-me` @@ -28,6 +30,25 @@ Right after a task that taught you something, run: [`/reflect`](../../skills/reflect/SKILL.md) sends the transcript to three parallel reviewers, then a synthesizer sorts the proposals into `Accepted`, `Rejected`, and `Backlog` and waits for your approval before any skill changes. Approve a proposal only if it would change a future decision. One weird session is an anecdote, not a rule. +## Fix the environment with `/correct` + +When you correct agents for the same mistake again and again, the fix belongs in the repo, not in your next prompt. Rank the options by how well they hold: + +1. Make the mistake impossible with architecture or a better data structure. +2. Block it with types, or with a lint or CI check whose error names the fix. +3. Catch it with a test. +4. Write it down as a doc or agent rule. Nothing fails when an agent skips a rule, so this comes last. + +Human review isn't on the list. A reviewer who must catch the same mistake on every PR is the problem this fixes. [`/correct`](../../skills/correct/SKILL.md) does the work: + +```text +/correct agents keep calling the database client directly instead of going through the repository layer +``` + +It reads recent commits, reverts, review comments, and comments that explain workarounds, then groups the mistakes into classes. A class counts once it has happened twice. It fixes the most frequent classes one commit each, at the highest level that works, and proves each new check fails on a real past mistake. It also keeps a table in the agent instruction file that pairs each rule with what enforces it, so a rule that nothing enforces shows up as a repeat. The reply lists each class with its evidence, the level chosen, and why a higher level didn't work. + +Run it with no argument and it finds the classes from history on its own. `/reflect` and `/correct` split the work. `/reflect` improves skills from one session. `/correct` changes the repo so a mistake class can't come back. Pair it with `/architect` when the fix is a new boundary. [Run many projects in parallel](./07-overnight.md#run-many-projects-in-parallel) has a prompt that does both across a whole repo. + ## Author a focused skill When you already know the workflow you want to capture: @@ -52,16 +73,28 @@ Skills aren't the only prose you ship. For docs, RFCs, readmes, PR descriptions, ## Test a skill change blind -A skill edit affects every future session, so test it like the experiment it is: +A skill edit affects every future session, so test it like the experiment it is. The same goes for adopting someone else's skill. Check that it makes the agent better on your work before you keep it. ```text /poteto-mode run the eval playbook on this skill change. same task for both variants, candidates stay blind. ``` +When a skill keeps missing and you know what it should do, change and test it in one task: + +```text +/poteto-mode update the review skill so it flags missing migrations, and eval the change. +``` + +Asking for the eval up front keeps the edit honest. A fix written from one bad session tends to overfit that session, and over many edits the skill drifts. The eval catches the drift before it ships. + The [Eval playbook](../../skills/poteto-mode/playbooks/eval.md) is built around one failure mode, the observer effect. An agent that knows it's being evaluated behaves differently. So candidate agents get an organic-looking task in sanitized directories, never the words "eval" or "candidate", and never each other's existence. One judge scores all outputs under neutral labels, and chain-following gets graded from which files each candidate actually read, not from what it claims. Read every output yourself before accepting the verdict. If you disagree with the judge, suspect the rubric before you suspect your judgment. +## Build a bot UI with `/make-bot-ui` + +One situational skill is for Grok Bot users. [`/make-bot-ui`](../../skills/make-bot-ui/SKILL.md) builds a small page whose buttons wake a bot over a webhook routine. For example, you could swipe through a review queue and have each swipe ask the bot to act on that item. A server on your machine holds the webhook's sender key, so the key never reaches the browser or the chat. The skill also covers exposing the page on Tailscale. + **Pitfall:** don't edit a skill mid-task because it's misbehaving. Fix it in its own PR and keep the task moving. A skill edit that ships tangled into feature work is invisible to review and impossible to evaluate. Next: [Recipes and pitfalls](./10-recipes-and-pitfalls.md). diff --git a/pstack/docs/guide/10-recipes-and-pitfalls.md b/pstack/docs/guide/10-recipes-and-pitfalls.md index da0af5e6f..68751e7c0 100644 --- a/pstack/docs/guide/10-recipes-and-pitfalls.md +++ b/pstack/docs/guide/10-recipes-and-pitfalls.md @@ -12,6 +12,30 @@ use /how first to understand how this initialization works. then use /why to fig Mechanics first, history second. Each skill's report tells you which sources it searched, so you know what the answer is grounded in. +## Restate a noisy report before touching code + +```text +/poteto-mode read this thread. restate the underlying issue in your own words, in plain english. don't change any code yet. +``` + +A misreading shows up in the restatement, where it costs one message to correct. Keep your own theory to yourself until the agent has stated its own. + +## Prototype before you pick + +```text +/poteto-mode prototype a few options for the settings layout. put them behind a switcher and send me screenshots of each. +``` + +You pick from things that run, not from descriptions. The agent answers its own layout and timing questions along the way. + +## Turn a settled design into a plan + +```text +/poteto-mode turn this design into a plan. small verifiable PRs, each with its own proof. +``` + +Ask only after the design settles. The plan is the deliverable, and it names the playbook that will execute it. + ## Get a second opinion on a design ```text @@ -44,6 +68,38 @@ The qualifiers do real work. "don't change anything yet" keeps it read-only, and "if there's a cheap test path" matters. Forcing a test through brittle mocks proves less than running the real command, and the playbook is allowed to say so. +## Repro and fix a report with proof + +```text +/poteto-mode repro this with /verify-. if it repros on main, fix it and show me a video as proof. +``` + +"if it repros on main" lets the run stop early when the bug is already gone. The video lets you check the fix before you read the diff. + +## Vet a number before you post it + +```text +/benchmark-checklist vet this 40% speedup before it goes in the pr description +``` + +You get faster, slower, no measurable difference, or inconclusive, with the run count, the range, and what limits the number. + +## Stop correcting the same mistake + +```text +/correct agents keep adding new config flags without registering them in the schema +``` + +The fix lands in the repo as architecture, a type, a lint, or a test, so the next agent can't make the mistake. + +## Ask how without starting the work + +```text +/poteto-help how do i get poteto-mode to stay on every turn? +``` + +You get an answer, a prompt to send, and a link to the source. Nothing runs until you send that prompt. + ## Keep a run honest while you're away ```text @@ -82,7 +138,13 @@ That's the whole prompt. [`/bro`](../../skills/bro/SKILL.md) restates the last m - **Enumerating skills in the prompt.** "use /how then /architect then /arena" reorders steps the playbook already sequences. State the goal and constraints. Name a skill only to override a default. - **A vague finish condition.** "make it better" gives `/loop` nothing to check. Give a command or artifact that can pass or fail. -- **Parallel agents in one worktree.** They overwrite each other and the diff becomes archaeology. Say "own worktree per attempt" and the isolation is free. +- **Leading with your theory of the cause.** The agent searches wherever you pointed. Ask it to restate the problem first, then share your hunch. +- **Taking the first design.** One attempt locks in the first shape the model thought of. Ask for prototypes or `/architect` and pick from evidence. +- **Polishing an abstract plan.** Adversarial review of a plan with no code behind it invents risks that will never happen. Settle the open questions with prototypes, then review what got built. +- **Parallel agents in one worktree.** They overwrite each other and the diff becomes archaeology. Run them as cloud agents, or say "own worktree per attempt". +- **Looping before you trust the loop.** A loop that can't verify its own work only makes unchecked work faster. Get the verification skill working first. +- **Trusting an unvetted number.** A warm cache or a skipped code path can fake a speedup. Run `/benchmark-checklist` before the number goes anywhere. +- **Correcting the same mistake by hand.** A correction in chat helps one run. `/correct` fixes the repo so no later run repeats it. - **Using `/arena` for coverage.** `/arena` repeats one design or code brief, then picks a base and grafts the best parts. `/swarm` partitions slices or declared race arms and aggregates one report. - **Accepting every review comment.** Bots and humans both file real catches and noise in one list. `/interrogate` sorts findings into act-on and dismissed buckets with reasons, and you can override either way. - **Treating `auto` as a model slug.** `auto` and `inherit-parent` mean "omit the model field so the subagent inherits the parent chat model." [Setup](./01-setup.md) covers the roles. diff --git a/pstack/docs/guide/README.md b/pstack/docs/guide/README.md index 144af4ab6..627bfb222 100644 --- a/pstack/docs/guide/README.md +++ b/pstack/docs/guide/README.md @@ -6,16 +6,24 @@ Here's what you'll learn: 1. [Set up pstack](./01-setup.md). Install the plugin and pick your models. 2. [Route work through `/poteto-mode`](./02-poteto-mode.md). Give it a goal and watch it pick a playbook. -3. [Understand the code](./03-understand.md). `/how`, `/why`, `/teach`, and `/recall` before you edit anything. -4. [Design the change](./04-design.md). `/architect`, `/arena`, `/swarm`, and `/interrogate` before code locks in a shape. +3. [Understand the code](./03-understand.md). A read-only investigation, then `/how`, `/why`, `/teach`, and `/recall` before you edit anything. +4. [Design the change](./04-design.md). `/architect`, `/arena`, `/swarm`, `/interrogate`, prototypes, and plans before code locks in a shape. 5. [Build and clean the change](./05-build-and-clean.md). The build playbooks, `/tdd`, `/unslop`, and `/no-comments`. -6. [Verify and ship](./06-verify-and-ship.md). Prove behavior on the real app, then open a focused PR and drive it to merged. -7. [Run work while you sleep](./07-overnight.md). An overnight contract, a decision log you can audit, and the playbooks that scale past one agent. +6. [Verify and ship](./06-verify-and-ship.md). Prove behavior on the real app, vet numbers with `/benchmark-checklist`, then open a focused PR and drive it to merged. +7. [Run work while you sleep](./07-overnight.md). Trust before loops, an overnight contract, a decision log you can audit, and Projects and automations that scale past one agent. 8. [Steer with principle names](./08-principles.md). The 24 names that redirect an agent mid-task. -9. [Make it yours](./09-make-it-yours.md). Your own mode, plus how to test a skill change. +9. [Make it yours](./09-make-it-yours.md). Your own mode, `/correct` for repeated mistakes, and how to test a skill change. 10. [Recipes and pitfalls](./10-recipes-and-pitfalls.md). Prompts to copy and mistakes to skip. -Read the pages in order the first time. After that, each page stands alone. When you're stuck, or can't tell which skill fits, ask [`/poteto-help`](../../skills/poteto-help/SKILL.md). It finds out what you're trying to do and points you at the right page or skill. +Read the pages in order the first time. After that, each page stands alone. + +When you're stuck, or can't tell which skill fits, type [`/poteto-help`](../../skills/poteto-help/SKILL.md) with your question: + +```text +/poteto-help which skill should i use to review this branch? +``` + +It answers, hands you a prompt to send, and links the skill or guide page the answer came from. It doesn't start the work, because a pstack run spends real tokens, so you send the prompt when you're ready. It runs only when you type it. ## If you only remember one thing