Repository navigation
讓一個功能最快最準:speed / accuracy 量測要內建在 IDD,還是另立 profiling plugin? #369
Description
Activity
- addedenhancementNew feature or requestNew feature or requestquestionFurther information is requestedFurther information is requested
on Oct 5, 2026 🎯 Decision
Original (2026-10-05):
「1. 我不確定 2. 可以分開來看,我剛開始想的是執行時間, 3. 要跑很多次的話,就要 FP 和 FN 分開看? 4. 可能要大量replicate, 5. 我覺得有些process的要求是要完全沒錯。例如你看apple-mail-mcp我最近就在做create_draft的方是。 5. 要看情況,例如create_draft的錢提就必須要是完全正確之下最快歐 6. 可能要避開。 我在想可能有耦合的問題」決定
The six Clarity Surface questions are answered and marked resolved in the body: speed starts as wall-clock execution time (tokens / cost and API calls are separate, later measures); accuracy is read as false positives and false negatives separately, estimated over many replicated runs; for processes like
create_draft, correctness is a hard constraint and speed is optimised only among fully correct candidates; the new capability should not be named "profiling". The level of "a feature" is undecided and moves to diagnose, withcreate_draftas the first concrete instance.Rationale
The answers turn "fastest and most accurate" from a trade-off into an ordering. For a process whose requirement is "never wrong" —
create_draftmust never wrap the body in a cite block or drop a recipient — no speed gain buys back a single error, so the objective is lexicographic: first exclude every candidate that is not fully correct, then pick the fastest of what remains. Other features may tolerate a known error rate and could use a genuine trade-off; that is why the answer to question 5 is "it depends", and the harness should let each feature declare which regime it is in rather than hard-code one.Reading false positives and false negatives separately follows from the same idea: #316 showed that the two error directions of one feature can have opposite costs (one silently routes a parked issue, the other halts real work), and a single accuracy number hides which one is happening. Replication is what turns "it worked when I tried it" into a rate, but it carries two conditions worth stating now. Each replicate still needs an oracle that judges right or wrong — for
create_draftthat can be a mechanical check of the saved draft. And zero failures does not mean a zero failure rate: after N independent clean runs, the one-sided 95 % upper bound on the failure rate is about 3/N, so a claim like "fails less than 1 % of the time" needs roughly 300 clean runs. A "completely correct" requirement therefore has to say which bound it means.The
create_draftexample also bears on placement. It lives inche-apple-mail-mcp, a Swift MCP server outside IDD, so whatever measures it has to work on code IDD does not own. The maintainer's remark about coupling (「我在想可能有耦合的問題」) is recorded as a decision input; my reading — unconfirmed — is that building the measurement into IDD would tie a general tool to IDD's lifecycle and to the existingidd-verify --profilevocabulary, which points toward a standalone tool that IDD consumes. That reading is for diagnose to confirm or reject, not settled here.Related
- bug: #298 的修正只落在 idd-list — 另三個 Complexity consumer 未動,且 idd-list 自身 Step 5 與 Step 3.7 互相矛盾 #316 — the FP / FN and frozen-corpus experience these answers build on
- Experimental create_draft path: write the draft directly into Mail's store, trigger upload by toggling its read status che-apple-mail-mcp#472 — the experimental
create_draftdirect-write path, a candidate second implementation to compare against the GUI path
Problem
The next goal is to measure a feature's speed and accuracy, and to drive it toward being as fast and as accurate as possible. Before building anything, one placement decision has to be made: does this live inside IDD (as part of the issue lifecycle), or in a separate profiling plugin?
Type
feature — with an open placement decision (to be settled in diagnose / discuss before any implementation)
Expected
Actual
idd-verifychecks fidelity — does the diff satisfy the issue — not performance.### Blockingsections) with hand-reviewed truth, false-positive / false-negative counts, and a rule that a reader-rule change must flip a measured corpus row. None of that machinery is reusable outside bug: #298 的修正只落在 idd-list — 另三個 Complexity consumer 未動,且 idd-list 自身 Step 5 與 Step 3.7 互相矛盾 #316.Context relevant to the placement decision
idd-verify --profile(code / prose / academic) already exists and means a reviewer-lens configuration, not performance profiling.idd-route recordwritesrouting-stats.jsonl— process metrics (round trips, findings per verify), not metrics of the feature itself.bestocr,bestasr,llm-context-benchmark,parallel-ai-agents:ensemble-eval.deep-integration-over-hardcodeapplies if the answer is "separate plugin": IDD would declare a dependency on it rather than carry its own copy.Impact
Every feature whose output can be right or wrong — gate verdicts, classifiers, OCR / ASR, parsers — currently re-invents its own measurement, and speed is not measured at all. Settling the placement first avoids building the same harness twice.
Priority
P2 — direction-setting; assumed scheduled after the open #316 / PR #318 lands (change on request).
Clarity Surface(idd-clarify run 2026-10-05T08:45:43Z)
create_draft(一個 MCP tool 操作,有多種實作方式))create_draft可機械檢查))create_draft這類 process 要求完全正確:正確是硬限制,在零錯誤的前提下求最快(不是兩者取捨))idd-verify --profile(指審查視角的組合),新的效能量測要避開 profiling 這個名字嗎?Current Status
Phase: created
Last updated: 2026-10-05 by /idd-comment → /idd-update
Key Decisions
create_draft, correctness is a hard constraint — fastest among fully correct candidates (lexicographic, not a trade-off)idd-verify --profile); coupling concern recorded as a placement inputche-apple-mail-mcpcreate_draft(GUI path vs Experimental create_draft path: write the draft directly into Mail's store, trigger upload by toggling its read status che-apple-mail-mcp#472 direct write)Scope Changes
Blocking
/idd-diagnose #369(placement decision: inside IDD vs standalone tool IDD consumes)Commits