fix(skills): summarize mass updates before semantic dispatch - #3978
azizur100389 wants to merge 1 commit into
Conversation
Co-Authored-By: OpenAI Codex <noreply@openai.com>
|
Thanks for the pull request, @azizur100389. A maintainer will review it soon. Want to talk it through while it is in review? Come join us on our Discord server. For longer-form discussion there is also GitHub Discussions. A couple of things that speed up review: make sure the test suite passes on Python 3.10 and 3.13, and that the change keeps extraction deterministic. |
There was a problem hiding this comment.
Graphify reviewed this change.
Worth a look — the grounded gate found no coupling regressions or blocking issues, but 1 advisory finding(s) below merit a look before merge.
Graphify review — findings
Adds summarize_incremental_changes, which returns the top-level directories by changed-file count and bytes when more than half the corpus has changed. It uses only the paths already detected and skips anything resolving outside the scan root. Missing files keep their count with zero bytes. Every agent's incremental-update skill now prints this summary after detect_incremental and asks the user whether the dominant directories belong to the corpus before semantic dispatch. Any exclusions go into .graphifyignore only when the user agrees, followed by a rerun of detection.
Worth a look
- Skill docs import
summarize_incremental_changes, which may not exist in graphify.detect —graphify/skills/copilot/references/update.md:17· Escalate · medium- agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
Review partial — this diff was larger than one review pass covers, so later files were not reviewed; some findings may be missing.
Analysis details — impact, health, verification
Impact & health
Graphify review
Impact — 3066 functions depend on the 569 functions this change touches.
Health — this change adds coupling hotspots:
- new:
extract()— 723 callers, 48 callees - new:
_rebuild_code()— 147 callers, 56 callees - new:
detect()— 112 callers, 15 callees - new:
_extract_generic()— 18 callers, 29 callees - new:
save_manifest()— 41 callers, 12 callees - new:
extract_js()— 87 callers, 4 callees - new:
extract_files_direct()— 17 callers, 20 callees - new:
extract_xaml()— 19 callers, 17 callees - …and 59 more — each is listed as a finding
Verification — 3066 functions in the blast radius were not formally verified this run (proofs are advisory here).
Gate & verification
graphify gate
PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.
Advisory (not blocking):
- verification_scope: 1343 function(s) in the blast radius were not formally verified this run
Test selection
Test selection
313 of 313 test file(s) selected (100%) via static blast radius.
Escalated to a full run for safety — the selection is not trustworthy on its own (see below). CI should run the whole suite.
tests/test_affected_cli.py— full-run-safetytests/test_affected_member_seed.py— full-run-safetytests/test_agents_platform.py— full-run-safetytests/test_analyze.py— full-run-safetytests/test_anthropic_custom_endpoint.py— full-run-safetytests/test_antigravity_install.py— full-run-safetytests/test_apm_fallback_version.py— full-run-safetytests/test_architecture_doc.py— full-run-safetytests/test_astro_extraction.py— impact, full-run-safetytests/test_astro_import_ids.py— full-run-safetytests/test_atomic_canvas_export.py— full-run-safetytests/test_atomic_version_stamp.py— full-run-safetytests/test_atomic_writes.py— impact, full-run-safetytests/test_backend_env_isolation.py— full-run-safetytests/test_backend_extras.py— full-run-safetytests/test_benchmark.py— full-run-safetytests/test_benchmark_raw_graph.py— full-run-safetytests/test_blade_extractor.py— full-run-safetytests/test_build.py— impact, full-run-safetytests/test_build_merge_dedup_scope.py— full-run-safetytests/test_build_merge_hyperedges_and_prune.py— full-run-safetytests/test_build_merge_shrink_guard.py— full-run-safetytests/test_builtin_global_type_refs.py— full-run-safetytests/test_cache.py— full-run-safetytests/test_callflow_html.py— full-run-safetytests/test_cargo_introspect.py— impact, full-run-safetytests/test_cargo_missing_manifest.py— full-run-safetytests/test_carried_hyperedge_remap.py— full-run-safetytests/test_case_sensitive_resolution.py— full-run-safetytests/test_charmap_encoding.py— impact, full-run-safetytests/test_chunking.py— impact, full-run-safetytests/test_cjs_module_extension.py— impact, full-run-safetytests/test_claude_cli_backend.py— impact, full-run-safetytests/test_claude_md.py— full-run-safetytests/test_cli_broken_pipe.py— full-run-safetytests/test_cli_export.py— full-run-safetytests/test_cli_help.py— full-run-safetytests/test_cluster.py— full-run-safetytests/test_cluster_exclude_hubs.py— full-run-safetytests/test_cobol_extractor.py— full-run-safetytests/test_codebuddy.py— full-run-safetytests/test_community_hub_labels.py— full-run-safetytests/test_community_labels_skill.py— full-run-safetytests/test_confidence.py— full-run-safetytests/test_corrupt_graph_json.py— full-run-safetytests/test_cpp_nested_and_cli.py— impact, full-run-safetytests/test_cpp_objc_cross_file_calls.py— full-run-safetytests/test_cpp_preprocess.py— full-run-safetytests/test_cross_extension_reexport_self_cycle.py— full-run-safetytests/test_cross_language_call_resolution.py— full-run-safety- … and 263 more
non-code file(s) changed (
graphify/skill-aider.md,graphify/skill-devin.md,graphify/skills/agents/references/update.md,graphify/skills/amp/references/update.md,graphify/skills/claude/references/update.md…) → running the full suite for safety (a code graph can't see config/fixture/data deps)
changed code file(s) with no mapped test (
graphify/skill-aider.md,graphify/skill-devin.md,graphify/skills/agents/references/update.md,graphify/skills/amp/references/update.md,graphify/skills/claude/references/update.md…) — a coverage gap or a missing link — running the full suite rather than only the selected tests
Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.
· 67 more finding(s) on lines outside this diff (see the check run).
Closes #3945.
When more than half of the detected corpus changes, incremental skills show the five largest top-level directory boundaries, with file counts and bytes, before semantic dispatch. Users confirm whether dominant directories belong in the corpus; agreed exclusions require rerunning detection.
The helper reuses the existing changed-file list, bounds paths to the scan root, deduplicates files, tolerates disappearing files and sorts deterministically. Source fragments, generated hosts and snapshots stay synchronized. Six regressions cover counts, bytes, thresholds, ordering, path boundaries and dispatch ordering.
Validation: Ruff, pre-commit, all five skillgen validators, fresh wheel installation, installed-artifact smoke checks, CLI help/install and the AST graph refresh pass. Regression tests reproduce the defects on unchanged upstream and pass with these changes.
Full CI passes on Python 3.10 (6224 passed, 13 repository-defined skips) and 3.12/3.13/3.14 (6223 passed, 14 repository-defined skips each). Frozen dependencies include all extras; CLI help/install passes on every version. No test filters or new skips were added.
The full Windows Python 3.12 suite finished with 6142 passed, 37 failed and 58 repository-defined skips. All 37 failing test IDs were rerun on unchanged upstream with the same Git Bash environment and reproduced there; none are introduced by this PR.
Pyright reports the same 586 baseline diagnostics with no introduced diagnostics. Bandit has the same 12 filtered baseline findings (4 high, 8 medium); pip-audit has the same 41 vulnerability records across 8 packages, matched by advisory aliases to upstream. Security CI is non-blocking in this repository, so its green status is not represented as a clean security audit. Dependencies and lockfile are unchanged.