fix(report): surface JavaScript files without structural symbols - #3979
azizur100389 wants to merge 2 commits into
Conversation
Co-Authored-By: OpenAI Codex <noreply@openai.com>
|
Thanks for the pull request, @azizur100389. A maintainer will review it soon. Want to talk it through while it is in review? Come join us on our Discord server. For longer-form discussion there is also GitHub Discussions. A couple of things that speed up review: make sure the test suite passes on Python 3.10 and 3.13, and that the change keeps extraction deterministic. |
Co-Authored-By: OpenAI Codex <noreply@openai.com>
There was a problem hiding this comment.
Graphify reviewed this change.
Looks safe to merge — no coupling regressions and no blocking issues, checked against the code graph (not a self-assessment).
Formal verification. PR-changed functions: 0/3 verified (0 proven, 0 may-equivalent, 0 distinguished) · 3 not verified (2 vacuous, 1 unsupported).
Not verified on this run: extract (vacuous: never exercised), \_extract\_generic (unsupported), generate (vacuous: never exercised).
Graphify review — findings
Flags JavaScript files whose parsed tree contains no functions, classes, imports, calls or new expressions by marking their file node with _no_structural_symbols and the source byte size; files with parser errors, non-contains edges or raw calls are excluded. extract prints a stderr note with the count, and generate adds a Files without structural symbols section to GRAPH_REPORT.md listing the five largest such files with sizes, using only the persisted metadata so it works without the original source. The section is only a coverage hint and does not trigger semantic extraction.
No blocking issues surfaced. 1 lower-confidence candidate did not survive cross-model review.
Analysis details — impact, health, verification
Impact & health
Graphify review
Impact — 2571 functions depend on the 546 functions this change touches.
Health — this change adds coupling hotspots:
- new:
extract()— 731 callers, 48 callees - new:
_rebuild_code()— 147 callers, 56 callees - new:
_extract_generic()— 18 callers, 29 callees - new:
extract_js()— 87 callers, 4 callees - new:
extract_xaml()— 19 callers, 17 callees - new:
generate()— 38 callers, 8 callees - new:
main()— 98 callers, 3 callees - new:
dispatch_command()— 2 callers, 126 callees - …and 49 more — each is listed as a finding
Verification — 2571 functions in the blast radius were not formally verified this run (proofs are advisory here).
Gate & verification
graphify gate
PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.
Advisory (not blocking):
- verification_scope: 2391 function(s) in the blast radius were not formally verified this run
Test selection
Test selection
313 of 313 test file(s) selected (100%) via static blast radius.
Escalated to a full run for safety — the selection is not trustworthy on its own (see below). CI should run the whole suite.
tests/test_affected_cli.py— full-run-safetytests/test_affected_member_seed.py— full-run-safetytests/test_agents_platform.py— full-run-safetytests/test_analyze.py— full-run-safetytests/test_anthropic_custom_endpoint.py— full-run-safetytests/test_antigravity_install.py— full-run-safetytests/test_apm_fallback_version.py— full-run-safetytests/test_architecture_doc.py— full-run-safetytests/test_astro_extraction.py— impact, full-run-safetytests/test_astro_import_ids.py— impact, full-run-safetytests/test_atomic_canvas_export.py— full-run-safetytests/test_atomic_version_stamp.py— full-run-safetytests/test_atomic_writes.py— full-run-safetytests/test_backend_env_isolation.py— full-run-safetytests/test_backend_extras.py— full-run-safetytests/test_benchmark.py— full-run-safetytests/test_benchmark_raw_graph.py— full-run-safetytests/test_blade_extractor.py— impact, full-run-safetytests/test_build.py— impact, full-run-safetytests/test_build_merge_dedup_scope.py— full-run-safetytests/test_build_merge_hyperedges_and_prune.py— full-run-safetytests/test_build_merge_shrink_guard.py— full-run-safetytests/test_builtin_global_type_refs.py— impact, full-run-safetytests/test_cache.py— full-run-safetytests/test_callflow_html.py— full-run-safetytests/test_cargo_introspect.py— full-run-safetytests/test_cargo_missing_manifest.py— full-run-safetytests/test_carried_hyperedge_remap.py— full-run-safetytests/test_case_sensitive_resolution.py— impact, full-run-safetytests/test_charmap_encoding.py— full-run-safetytests/test_chunking.py— full-run-safetytests/test_cjs_module_extension.py— impact, full-run-safetytests/test_claude_cli_backend.py— full-run-safetytests/test_claude_md.py— full-run-safetytests/test_cli_broken_pipe.py— full-run-safetytests/test_cli_export.py— full-run-safetytests/test_cli_help.py— full-run-safetytests/test_cluster.py— full-run-safetytests/test_cluster_exclude_hubs.py— full-run-safetytests/test_cobol_extractor.py— impact, full-run-safetytests/test_codebuddy.py— full-run-safetytests/test_community_hub_labels.py— full-run-safetytests/test_community_labels_skill.py— full-run-safetytests/test_confidence.py— impact, full-run-safetytests/test_corrupt_graph_json.py— full-run-safetytests/test_cpp_nested_and_cli.py— impact, full-run-safetytests/test_cpp_objc_cross_file_calls.py— impact, full-run-safetytests/test_cpp_preprocess.py— full-run-safetytests/test_cross_extension_reexport_self_cycle.py— impact, full-run-safetytests/test_cross_language_call_resolution.py— impact, full-run-safety- … and 263 more
non-code file(s) changed (
README.md) → running the full suite for safety (a code graph can't see config/fixture/data deps)
changed code file(s) with no mapped test (
README.md) — a coverage gap or a missing link — running the full suite rather than only the selected tests
Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.
Formal verification
Could not verify: Could not verify extract.
The verifier did not have enough to check extract, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 90 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly TypeError — names the real obstacle, not a sampling gap)
Could not verify: Could not verify \_extract\_generic.
The verifier did not have enough to check \_extract\_generic, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: parameter `config` is annotated `LanguageConfig` — outside the synthesizable primitive/collection set
Could not verify: Could not verify generate.
The verifier did not have enough to check generate, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 200 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly ValueError — names the real obstacle, not a sampling gap)
· 57 more finding(s) on lines outside this diff (see the check run).
Addresses #3946.
Data-shaped JavaScript can contribute only a file node or simple bindings while its contents stay invisible to AST extraction. The extractor now records that limitation on the file node, prints a bounded summary, and reports the five largest affected files with UTF-8 byte sizes.
Classification uses the existing parsed JavaScript tree, including unreported call initializers, and excludes functions, classes, imports, re-exports, calls and syntax errors. Metadata survives JSON persistence and report regeneration without reopening source files; incremental replacement removes stale notices. No automatic semantic fallback is added. Twenty-two regressions cover the seven reported forms, structural exclusions, extension variants, byte sizes, bounded ordering incremental replacement and filenames containing backticks.
Validation: Ruff, pre-commit, all five skillgen validators, fresh wheel installation, installed-artifact smoke checks, CLI help/install and the AST graph refresh pass. Regression tests reproduce the defects on unchanged upstream and pass with these changes.
Full CI passes on Python 3.10 (6240 passed, 13 repository-defined skips) and 3.12/3.13/3.14 (6239 passed, 14 repository-defined skips each). Frozen dependencies include all extras; CLI help/install passes on every version. No test filters or new skips were added.
The full Windows Python 3.12 suite finished with 6157 passed, 37 failed and 58 repository-defined skips. All 37 failing test IDs were rerun on unchanged upstream with the same Git Bash environment and reproduced there; none are introduced by this PR. This Windows run collected before the final filename-formatting regression was added; all 40 focused report tests pass at the final commit, and the final full CI matrix includes that regression. The final wheel also passes the installed-report check for a filename containing a backtick; a CommonMark rendering check covers HTML and newline characters.
Pyright reports the same 586 baseline diagnostics with no introduced diagnostics. Bandit has the same 12 filtered baseline findings (4 high, 8 medium); pip-audit has the same 41 vulnerability records across 8 packages, matched by advisory aliases to upstream. Security CI is non-blocking in this repository, so its green status is not represented as a clean security audit. Dependencies and lockfile are unchanged.