Skip to content

讓一個功能最快最準:speed / accuracy 量測要內建在 IDD,還是另立 profiling plugin? #369

Description

@kiki830621

Problem

Original text:
「我發現接下來一個重要的東西是看一個功能的 speed and accuracy,要能夠讓一個功能最快最準是目標,我在想是要在idd還是要另外創立一個profiling 的plugin呢」
— Source: maintainer, /idd-issue invocation, 2026-10-05

The next goal is to measure a feature's speed and accuracy, and to drive it toward being as fast and as accurate as possible. Before building anything, one placement decision has to be made: does this live inside IDD (as part of the issue lifecycle), or in a separate profiling plugin?

Type

feature — with an open placement decision (to be settled in diagnose / discuss before any implementation)

Expected

  1. A repeatable way to measure, for one feature, how fast it is and how accurate it is, so that "fastest and most accurate" becomes a checkable target instead of an impression.
  2. A recorded decision on where this capability lives: an IDD lifecycle step or skill, a standalone plugin that IDD consumes, or a split between the two.

Actual

Context relevant to the placement decision

  • Name collision: idd-verify --profile (code / prose / academic) already exists and means a reviewer-lens configuration, not performance profiling.
  • idd-route record writes routing-stats.jsonl — process metrics (round trips, findings per verify), not metrics of the feature itself.
  • Measurement-first tools already exist outside IDD, each for one domain: bestocr, bestasr, llm-context-benchmark, parallel-ai-agents:ensemble-eval.
  • The dev rule deep-integration-over-hardcode applies if the answer is "separate plugin": IDD would declare a dependency on it rather than carry its own copy.

Impact

Every feature whose output can be right or wrong — gate verdicts, classifiers, OCR / ASR, parsers — currently re-invents its own measurement, and speed is not measured at all. Settling the placement first avoids building the same harness twice.

Priority

P2 — direction-setting; assumed scheduled after the open #316 / PR #318 lands (change on request).

Clarity Surface(idd-clarify run 2026-10-05T08:45:43Z)

Type Source Question for you Status
ambiguity 「看一個功能的 speed and accuracy」 這裡的「一個功能」是哪個層級:一支 helper 函式、一個 skill、一個 MCP tool,還是整個 plugin? resolved @ 2026-10-05T09:52:03Z (reason: 使用者未定 → 移交 diagnose 判定層級;首個實例是 che-apple-mail-mcp 的 create_draft(一個 MCP tool 操作,有多種實作方式))
ambiguity 「speed」 speed 要量的是執行時間、token 與費用,還是 API 呼叫次數? resolved @ 2026-10-05T09:52:03Z (reason: 分開看;先量執行時間(wall-clock),token/費用與 API 次數之後另看)
ambiguity 「accuracy」 accuracy 要用單一個準確率,還是像 #316 那樣把 false positive 和 false negative 分開看(兩個方向的代價不同)? resolved @ 2026-10-05T09:52:03Z (reason: FP 與 FN 分開看;要跑很多次(replicate),兩個方向各自估錯誤率)
missing-context 「accuracy」 accuracy 的真值從哪裡來:人工判讀的凍結語料、事先寫好的標準答案,還是跨模型一致? resolved @ 2026-10-05T09:52:03Z (reason: 大量 replicate;每次結果仍需一個判對錯的 oracle,依功能而定(create_draft 可機械檢查))
ambiguity 「讓一個功能最快最準」 快和準衝突的時候,哪一個優先?還是兩個都列出來讓人自己取捨? resolved @ 2026-10-05T09:52:03Z (reason: 看情況;像 create_draft 這類 process 要求完全正確:正確是硬限制,在零錯誤的前提下求最快(不是兩者取捨))
ambiguity 「profiling 的plugin」 IDD 已有 idd-verify --profile(指審查視角的組合),新的效能量測要避開 profiling 這個名字嗎? resolved @ 2026-10-05T09:52:03Z (reason: 避開 profiling 這個名字;使用者另提可能有耦合問題(記於 comment))

Current Status

Phase: created
Last updated: 2026-10-05 by /idd-comment → /idd-update

Key Decisions

Scope Changes

  • (none)

Blocking

  • (none) — next: /idd-diagnose #369 (placement decision: inside IDD vs standalone tool IDD consumes)

Commits

  • (none)

Activity

  1. kiki830621 commented on Oct 5, 2026

    @kiki830621
    MemberAuthor

    🎯 Decision

    Original (2026-10-05):
    「1. 我不確定 2. 可以分開來看,我剛開始想的是執行時間, 3. 要跑很多次的話,就要 FP 和 FN 分開看? 4. 可能要大量replicate, 5. 我覺得有些process的要求是要完全沒錯。例如你看apple-mail-mcp我最近就在做create_draft的方是。 5. 要看情況,例如create_draft的錢提就必須要是完全正確之下最快歐 6. 可能要避開。 我在想可能有耦合的問題」

    決定

    The six Clarity Surface questions are answered and marked resolved in the body: speed starts as wall-clock execution time (tokens / cost and API calls are separate, later measures); accuracy is read as false positives and false negatives separately, estimated over many replicated runs; for processes like create_draft, correctness is a hard constraint and speed is optimised only among fully correct candidates; the new capability should not be named "profiling". The level of "a feature" is undecided and moves to diagnose, with create_draft as the first concrete instance.

    Rationale

    The answers turn "fastest and most accurate" from a trade-off into an ordering. For a process whose requirement is "never wrong" — create_draft must never wrap the body in a cite block or drop a recipient — no speed gain buys back a single error, so the objective is lexicographic: first exclude every candidate that is not fully correct, then pick the fastest of what remains. Other features may tolerate a known error rate and could use a genuine trade-off; that is why the answer to question 5 is "it depends", and the harness should let each feature declare which regime it is in rather than hard-code one.

    Reading false positives and false negatives separately follows from the same idea: #316 showed that the two error directions of one feature can have opposite costs (one silently routes a parked issue, the other halts real work), and a single accuracy number hides which one is happening. Replication is what turns "it worked when I tried it" into a rate, but it carries two conditions worth stating now. Each replicate still needs an oracle that judges right or wrong — for create_draft that can be a mechanical check of the saved draft. And zero failures does not mean a zero failure rate: after N independent clean runs, the one-sided 95 % upper bound on the failure rate is about 3/N, so a claim like "fails less than 1 % of the time" needs roughly 300 clean runs. A "completely correct" requirement therefore has to say which bound it means.

    The create_draft example also bears on placement. It lives in che-apple-mail-mcp, a Swift MCP server outside IDD, so whatever measures it has to work on code IDD does not own. The maintainer's remark about coupling (「我在想可能有耦合的問題」) is recorded as a decision input; my reading — unconfirmed — is that building the measurement into IDD would tie a general tool to IDD's lifecycle and to the existing idd-verify --profile vocabulary, which points toward a standalone tool that IDD consumes. That reading is for diagnose to confirm or reject, not settled here.

    Related

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestquestionFurther information is requested

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions