Reuse the shared real node filter in the community listing loop - #3836
ayushcodes10 wants to merge 3 commits into
Conversation
The per community Nodes count and listing loop only excluded file nodes, so a rationale docstring fragment node or a concept node leaked into Nodes (N) as if it were a real code declaration, and inflated the reported count by including it. The Knowledge Gaps section a few lines below already excluded file, concept and rationale nodes together, the correct set an earlier issue established for this same function, the community listing loop was never given the same treatment. Consolidates every real node check this function makes, the summary thin and shown counts, the community listing, and the gaps section isolated and thin community counts, onto one shared predicate instead of four separate inline filters that had drifted apart. Fixing only the display line without also fixing the count computations that decide which communities are thin or shown would have reintroduced the exact header versus render disagreement an earlier issue fixed, since the header counts and the rendered content need to agree on what counts as a real node. Fixes issue 3794.
Covers issue 3794 directly: a community made of real code nodes plus a rationale docstring fragment node and a concept node must only count and list the real code nodes in its Nodes (N) line, matching the exact repro shape from the issue. Confirmed against pre fix code via git stash, the test fails there and reproduces the issue's own example output almost exactly.
There was a problem hiding this comment.
Graphify reviewed this change.
Looks safe to merge — no coupling regressions and no blocking issues, checked against the code graph (not a self-assessment).
Graphify review — findings
Excludes rationale (docstring-fragment) and concept nodes from GRAPH_REPORT.md's per-community Nodes (N): ... listing, so a class with a long docstring no longer prints raw prose fragments as node names or inflates the reported community size. Routes every self-reported node count in generate (summary thin/shown totals, community listing, and Knowledge Gaps isolated/thin counts) through a single _real_node predicate so they can't drift apart.
No blocking issues surfaced.
Analysis details — impact, health, verification
Impact & health
Graphify review
Impact — 598 functions depend on the 241 functions this change touches.
Health — this change adds coupling hotspots:
- new:
_rebuild_code()— 144 callers, 55 callees - new:
main()— 98 callers, 3 callees - new:
generate()— 36 callers, 8 callees - new:
dispatch_command()— 2 callers, 125 callees - new:
run_pipeline()— 8 callers, 13 callees - new:
watch()— 5 callers, 7 callees - new:
load_learning_for_report()— 4 callers, 3 callees - new:
test_report_shows_avg_confidence_for_inferred()— 0 callers, 7 callees - …and 2 more — each is listed as a finding
Verification — 598 functions in the blast radius were not formally verified this run (proofs are advisory here).
Gate & verification
graphify gate
PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.
Advisory (not blocking):
- verification_scope: 421 function(s) in the blast radius were not formally verified this run
Test selection
Test selection
301 of 301 test file(s) selected (100%) via static blast radius.
Escalated to a full run for safety — the selection is not trustworthy on its own (see below). CI should run the whole suite.
tests/test_affected_cli.py— full-run-safetytests/test_affected_member_seed.py— full-run-safetytests/test_agents_platform.py— full-run-safetytests/test_analyze.py— full-run-safetytests/test_anthropic_custom_endpoint.py— full-run-safetytests/test_antigravity_install.py— full-run-safetytests/test_apm_fallback_version.py— full-run-safetytests/test_architecture_doc.py— full-run-safetytests/test_astro_extraction.py— full-run-safetytests/test_astro_import_ids.py— full-run-safetytests/test_atomic_canvas_export.py— full-run-safetytests/test_atomic_version_stamp.py— full-run-safetytests/test_atomic_writes.py— full-run-safetytests/test_backend_env_isolation.py— full-run-safetytests/test_backend_extras.py— full-run-safetytests/test_benchmark.py— full-run-safetytests/test_benchmark_raw_graph.py— full-run-safetytests/test_build.py— full-run-safetytests/test_build_merge_dedup_scope.py— full-run-safetytests/test_build_merge_hyperedges_and_prune.py— full-run-safetytests/test_build_merge_shrink_guard.py— full-run-safetytests/test_builtin_global_type_refs.py— full-run-safetytests/test_cache.py— full-run-safetytests/test_callflow_html.py— full-run-safetytests/test_cargo_introspect.py— full-run-safetytests/test_cargo_missing_manifest.py— full-run-safetytests/test_carried_hyperedge_remap.py— full-run-safetytests/test_case_sensitive_resolution.py— full-run-safetytests/test_charmap_encoding.py— full-run-safetytests/test_chunking.py— full-run-safetytests/test_cjs_module_extension.py— full-run-safetytests/test_claude_cli_backend.py— full-run-safetytests/test_claude_md.py— full-run-safetytests/test_cli_broken_pipe.py— full-run-safetytests/test_cli_export.py— full-run-safetytests/test_cli_help.py— full-run-safetytests/test_cluster.py— full-run-safetytests/test_cobol_extractor.py— full-run-safetytests/test_codebuddy.py— full-run-safetytests/test_community_hub_labels.py— full-run-safetytests/test_community_labels_skill.py— full-run-safetytests/test_confidence.py— impact, full-run-safetytests/test_corrupt_graph_json.py— full-run-safetytests/test_cpp_nested_and_cli.py— full-run-safetytests/test_cpp_objc_cross_file_calls.py— full-run-safetytests/test_cpp_preprocess.py— full-run-safetytests/test_cross_extension_reexport_self_cycle.py— full-run-safetytests/test_cross_language_call_resolution.py— full-run-safetytests/test_cross_repo_external_call_guards.py— full-run-safetytests/test_cross_repo_member_calls.py— full-run-safety- … and 251 more
non-code file(s) changed (
CHANGELOG.md) → running the full suite for safety (a code graph can't see config/fixture/data deps)
changed code file(s) with no mapped test (
CHANGELOG.md) — a coverage gap or a missing link — running the full suite rather than only the selected tests
Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.
· 10 more finding(s) on lines outside this diff (see the check run).
Atomic label sidecars (#3853), docx table/all-text extraction (#3833), wiki index-collision (#3818) and escaped wikilink aliases (#3772), community-listing real-node filter (#3836), warn-once on missing pypdf extra (#3710), plus the --help logo banner + app.graphify.com line. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
|
Shipped in v0.9.72 (now on PyPI) via an authorship-preserving cherry-pick, so your commit keeps contributor-graph credit. Thanks @ayushcodes10! Community listing reuses the shared real-node filter. Release: https://github.com/Graphify-Labs/graphify/releases/tag/v0.9.72 |
Fixes #3794.
Summary
GRAPH_REPORT.md's per-communityNodes (N): ...listing only excluded file nodes from display, so arationale(docstring-fragment) node or aconceptnode leaked into the list as if it were a real code declaration — on a real project this printed raw docstring prose, including mid-sentence truncation and Sphinx:param:/:rtype:directives, as if they were node names, and inflated the reported community size since the count islen(real_nodes).The Knowledge Gaps section a few lines below in the same function already excludes file, concept, and rationale nodes — the correct set #1768 established for this function. The community-listing loop was never given the same treatment; this is the third spot with this exact missing-filter pattern (#1768, then a Knowledge Gaps spot, now this one).
Fix
Rather than patching only the display line, this consolidates every "is this a real node" check the function makes — the summary's thin/shown counts, the community listing, and the Gaps section's isolated/thin-community counts — onto one shared
_real_nodepredicate. Fixing only the display line while leaving the count computations on the narrower file-only filter would have reintroduced the exact header-vs-render disagreement #3148 fixed, since a community made up entirely of rationale/concept nodes would still count toward "shown" even though nothing real is rendered for it.Testing
test_community_listing_excludes_rationale_and_concept_nodestotests/test_report_gap_thresholds.py, reproducing the issue's own repro shape (code nodes + a rationale node + a concept node in one community) — confirmed it reproduces the issue's exact example output against pre-fix code viagit stash.test_report.py/test_report_gap_thresholds.pyplus every other test file touchingreport.py(test_semantic_similarity.py,test_hypergraph.py,test_labeling.py,test_pipeline.py,test_confidence.py) passes — 108 tests.tests/test_incremental.py::test_update_preserves_cross_file_imports_and_calls_to_unchanged_file) — confirmed via a separate cleanupstream/v8worktree that it fails identically without any of this PR's changes, so it's not introduced here; flagging separately.python3 -m tools.skillgen --checkpasses.