Skip to content

fix(report): surface JavaScript files without structural symbols - #3979

Open
azizur100389 wants to merge 2 commits into
Graphify-Labs:v8from
azizur100389:codex/data-only-code
Open

azizur100389 wants to merge 2 commits into
Graphify-Labs:v8from
azizur100389:codex/data-only-code

Conversation

@azizur100389

@azizur100389 azizur100389 commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Addresses #3946.

Data-shaped JavaScript can contribute only a file node or simple bindings while its contents stay invisible to AST extraction. The extractor now records that limitation on the file node, prints a bounded summary, and reports the five largest affected files with UTF-8 byte sizes.

Classification uses the existing parsed JavaScript tree, including unreported call initializers, and excludes functions, classes, imports, re-exports, calls and syntax errors. Metadata survives JSON persistence and report regeneration without reopening source files; incremental replacement removes stale notices. No automatic semantic fallback is added. Twenty-two regressions cover the seven reported forms, structural exclusions, extension variants, byte sizes, bounded ordering incremental replacement and filenames containing backticks.

Validation: Ruff, pre-commit, all five skillgen validators, fresh wheel installation, installed-artifact smoke checks, CLI help/install and the AST graph refresh pass. Regression tests reproduce the defects on unchanged upstream and pass with these changes.

Full CI passes on Python 3.10 (6240 passed, 13 repository-defined skips) and 3.12/3.13/3.14 (6239 passed, 14 repository-defined skips each). Frozen dependencies include all extras; CLI help/install passes on every version. No test filters or new skips were added.

The full Windows Python 3.12 suite finished with 6157 passed, 37 failed and 58 repository-defined skips. All 37 failing test IDs were rerun on unchanged upstream with the same Git Bash environment and reproduced there; none are introduced by this PR. This Windows run collected before the final filename-formatting regression was added; all 40 focused report tests pass at the final commit, and the final full CI matrix includes that regression. The final wheel also passes the installed-report check for a filename containing a backtick; a CommonMark rendering check covers HTML and newline characters.

Pyright reports the same 586 baseline diagnostics with no introduced diagnostics. Bandit has the same 12 filtered baseline findings (4 high, 8 medium); pip-audit has the same 41 vulnerability records across 8 packages, matched by advisory aliases to upstream. Security CI is non-blocking in this repository, so its green status is not represented as a clean security audit. Dependencies and lockfile are unchanged.

Co-Authored-By: OpenAI Codex <noreply@openai.com>
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown

Thanks for the pull request, @azizur100389. A maintainer will review it soon.

Want to talk it through while it is in review? Come join us on our Discord server. For longer-form discussion there is also GitHub Discussions.

A couple of things that speed up review: make sure the test suite passes on Python 3.10 and 3.13, and that the change keeps extraction deterministic.

Co-Authored-By: OpenAI Codex <noreply@openai.com>
@azizur100389
azizur100389 marked this pull request as ready for review October 1, 2026 19:28

@graphify-labs graphify-labs Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graphify reviewed this change.

Looks safe to merge — no coupling regressions and no blocking issues, checked against the code graph (not a self-assessment).

Formal verification. PR-changed functions: 0/3 verified (0 proven, 0 may-equivalent, 0 distinguished) · 3 not verified (2 vacuous, 1 unsupported).

Not verified on this run: extract (vacuous: never exercised), \_extract\_generic (unsupported), generate (vacuous: never exercised).


Graphify review — findings

Flags JavaScript files whose parsed tree contains no functions, classes, imports, calls or new expressions by marking their file node with _no_structural_symbols and the source byte size; files with parser errors, non-contains edges or raw calls are excluded. extract prints a stderr note with the count, and generate adds a Files without structural symbols section to GRAPH_REPORT.md listing the five largest such files with sizes, using only the persisted metadata so it works without the original source. The section is only a coverage hint and does not trigger semantic extraction.

No blocking issues surfaced. 1 lower-confidence candidate did not survive cross-model review.

Analysis details — impact, health, verification

Impact & health

Graphify review

Impact — 2571 functions depend on the 546 functions this change touches.

Health — this change adds coupling hotspots:

  • new: extract() — 731 callers, 48 callees
  • new: _rebuild_code() — 147 callers, 56 callees
  • new: _extract_generic() — 18 callers, 29 callees
  • new: extract_js() — 87 callers, 4 callees
  • new: extract_xaml() — 19 callers, 17 callees
  • new: generate() — 38 callers, 8 callees
  • new: main() — 98 callers, 3 callees
  • new: dispatch_command() — 2 callers, 126 callees
  • …and 49 more — each is listed as a finding

Verification — 2571 functions in the blast radius were not formally verified this run (proofs are advisory here).

Gate & verification

graphify gate

PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.

Advisory (not blocking):

  • verification_scope: 2391 function(s) in the blast radius were not formally verified this run

Test selection

Test selection

313 of 313 test file(s) selected (100%) via static blast radius.

Escalated to a full run for safety — the selection is not trustworthy on its own (see below). CI should run the whole suite.

  • tests/test_affected_cli.py — full-run-safety
  • tests/test_affected_member_seed.py — full-run-safety
  • tests/test_agents_platform.py — full-run-safety
  • tests/test_analyze.py — full-run-safety
  • tests/test_anthropic_custom_endpoint.py — full-run-safety
  • tests/test_antigravity_install.py — full-run-safety
  • tests/test_apm_fallback_version.py — full-run-safety
  • tests/test_architecture_doc.py — full-run-safety
  • tests/test_astro_extraction.py — impact, full-run-safety
  • tests/test_astro_import_ids.py — impact, full-run-safety
  • tests/test_atomic_canvas_export.py — full-run-safety
  • tests/test_atomic_version_stamp.py — full-run-safety
  • tests/test_atomic_writes.py — full-run-safety
  • tests/test_backend_env_isolation.py — full-run-safety
  • tests/test_backend_extras.py — full-run-safety
  • tests/test_benchmark.py — full-run-safety
  • tests/test_benchmark_raw_graph.py — full-run-safety
  • tests/test_blade_extractor.py — impact, full-run-safety
  • tests/test_build.py — impact, full-run-safety
  • tests/test_build_merge_dedup_scope.py — full-run-safety
  • tests/test_build_merge_hyperedges_and_prune.py — full-run-safety
  • tests/test_build_merge_shrink_guard.py — full-run-safety
  • tests/test_builtin_global_type_refs.py — impact, full-run-safety
  • tests/test_cache.py — full-run-safety
  • tests/test_callflow_html.py — full-run-safety
  • tests/test_cargo_introspect.py — full-run-safety
  • tests/test_cargo_missing_manifest.py — full-run-safety
  • tests/test_carried_hyperedge_remap.py — full-run-safety
  • tests/test_case_sensitive_resolution.py — impact, full-run-safety
  • tests/test_charmap_encoding.py — full-run-safety
  • tests/test_chunking.py — full-run-safety
  • tests/test_cjs_module_extension.py — impact, full-run-safety
  • tests/test_claude_cli_backend.py — full-run-safety
  • tests/test_claude_md.py — full-run-safety
  • tests/test_cli_broken_pipe.py — full-run-safety
  • tests/test_cli_export.py — full-run-safety
  • tests/test_cli_help.py — full-run-safety
  • tests/test_cluster.py — full-run-safety
  • tests/test_cluster_exclude_hubs.py — full-run-safety
  • tests/test_cobol_extractor.py — impact, full-run-safety
  • tests/test_codebuddy.py — full-run-safety
  • tests/test_community_hub_labels.py — full-run-safety
  • tests/test_community_labels_skill.py — full-run-safety
  • tests/test_confidence.py — impact, full-run-safety
  • tests/test_corrupt_graph_json.py — full-run-safety
  • tests/test_cpp_nested_and_cli.py — impact, full-run-safety
  • tests/test_cpp_objc_cross_file_calls.py — impact, full-run-safety
  • tests/test_cpp_preprocess.py — full-run-safety
  • tests/test_cross_extension_reexport_self_cycle.py — impact, full-run-safety
  • tests/test_cross_language_call_resolution.py — impact, full-run-safety
  • … and 263 more

non-code file(s) changed (README.md) → running the full suite for safety (a code graph can't see config/fixture/data deps)

changed code file(s) with no mapped test (README.md) — a coverage gap or a missing link — running the full suite rather than only the selected tests

Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.

Formal verification

Could not verify: Could not verify extract.

The verifier did not have enough to check extract, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 90 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly TypeError — names the real obstacle, not a sampling gap)

Could not verify: Could not verify \_extract\_generic.

The verifier did not have enough to check \_extract\_generic, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: parameter `config` is annotated `LanguageConfig` — outside the synthesizable primitive/collection set

Could not verify: Could not verify generate.

The verifier did not have enough to check generate, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 200 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly ValueError — names the real obstacle, not a sampling gap)

· 57 more finding(s) on lines outside this diff (see the check run).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant