Skip to content

feat(mermaid): flow diagram thể hiện logic + call graph thay vì chuỗi thẳng - #12

Closed
hungpham10 wants to merge 65 commits into
Cleboost:mainfrom
hungpham10:feat/mermaid-flow-call-graph
Closed

hungpham10 wants to merge 65 commits into
Cleboost:mainfrom
hungpham10:feat/mermaid-flow-call-graph

Conversation

@hungpham10

Copy link
Copy Markdown

Vấn đề

Task yêu cầu: tạo Mermaid cho một hàm cụ thể để biết logic hàm làm gì và call đi đâu.

codegraph_mermaid (MCP) / mermaid (GraphQL) đã có kind: FLOW, nhưng control_flow chỉ:

  • nối tuần tự mọi node trong flow.chain bằng nodes.windows(2), nên nhánh if/else, switch, vòng lặp đều bị vẽ thành một đường thẳng;
  • bỏ qua hoàn toàn FlowResult.calls — không hiện line, condition, effect, callee;
  • gộp chung call nội bộ và call ngoài repo (không phân biệt · ext).

Trong khi đó flow.calls đã có sẵn mọi thứ cần thiết (position, to_name, to_id, line, condition, effect).

Tái hiện (hàm mermaid tại query.rs:167)

Trước — chuỗi thẳng, mất hết nhánh và call-site info:

flowchart TD
  c0["mermaid"]
  c1["ctx.data::<Arc<AppState>>"]
  c2{"IF_TRUE"}
  c3(["RETURN"])
  ...
  c0 --> c1
  c1 --> c2
  c2 --> c3   (nối tuyến tính, không phải logic)

Thay đổi

crates/codegraph-api/src/mermaid.rs — viết lại control_flow dùng cả flow.calls:

  • Điều kiện (IF_TRUE / IF_FALSE / SWITCH_CASE) → hình thoi, label kèm guard text lấy từ call record trong nhánh (IF_TRUE: !state.mermaid).
  • Vòng lặp / kết thúc (LOOP, LOOP_BACK, RETURN, THROW, …) → stadium.
  • Call → hộp chữ nhật: tên_callee · L<line> · if <condition> · <effect>.
  • Call không resolve (ngoài repo) → hộp subroutine [[...]] + · ext.
  • Cạnh suy từ marker, không nối tuyến tính:
    • IF_FALSE (else) rẽ từ chính node điều kiện if;
    • BRANCH_END hợp nhất nhánh (kể cả nhánh false khi không có else);
    • LOOP_BACK vẽ back edge về header loop;
    • SWITCH_CASE toả từ node entry của switch;
    • node thường nối tuần tự node liền trước; root không có cạnh vào.

Sau:

flowchart TD
  c0["mermaid"]
  c1[["ctx.data::<Arc<AppState>> · L174 · ext"]]
  c2{"IF_TRUE"}
  c3(["RETURN"])
  c4[["Err · L176 · if !state.mermaid · ext"]]
  c6(["BRANCH_END"])
  c7["parse_id · L180"]
  c10["api_for · L182"]
  c11{"SWITCH_CASE"}
  ...
  c1 --> c2
  c1 --> c6        (nhánh false)
  c2 --> c3 --> c4 --> c5 --> c6
  c10 --> c11
  c10 --> c18      (các case switch toả từ entry)
  c10 --> c24
  c10 --> c30

Ngoài ra cập nhật doc/mô tả:

  • crates/codegraph-mcp/src/tools.rs: mô tả tool codegraph_mermaid nêu rõ FLOW thể hiện nhánh + call-site.
  • crates/codegraph-graphql/src/types.rs: doc MermaidKind mô tả FLOW / CALLERS / CALLEES / IMPACT.

GraphQL

Không đổi schema, chỉ cải thiện nội dung trả về của field sẵn có:

{ mermaid(id: "156", kind: FLOW) }          # logic + call đi đâu
{ mermaid(id: "156", kind: CALLEES, depth: 2) }  # hàm gọi ai (graph LR)
{ mermaid(id: "156", kind: CALLERS, depth: 2) }  # ai gọi hàm
{ mermaid(id: "156", kind: IMPACT,  depth: 2) }  # bán kính ảnh hưởng

Server cần chạy serve --graphql --mermaid. MCP tương ứng serve --mcp --mermaid và codegraph_mermaid {node, kind, depth}.

Kiểm thử

  • 6 unit test mới cho control_flow: if/else không nối tuyến tính, loop back edge, switch fan-out, label call (line/condition/effect/ext), guard của điều kiện, root không có cạnh vào.
  • cargo test -p codegraph-api (10 integration + 6 unit) — pass.
  • cargo clippy -p codegraph-api -p codegraph-graphql -p codegraph-mcp --all-targets -- -D warnings — sạch.
  • Verify end-to-end qua GraphQL --mermaid trên chính repo này (hàm mermaid id 156).

hungpham7-tiki and others added 30 commits August 2, 2026 20:49
…-index-algorithm

Replace bfs with search index algorithm
* Remove old data structure

* Benchmark

* Remove temporary codegraph-viz and move benches into .github

* Fix issue codspeed

* Fix benchmark

* Implement sandbox and diff

* Add sandbox execution for diff and origin

* Remove unused unittest

* Fix lint

* Remove unused unittest

* Temporal remove another environments
* Implement new storage to improve performance

* style: apply rustfmt

* Fix lint

* Fix issue multiple reading in multiple streams
* Finetune to reduce LLM token

* style: apply rustfmt

* Fix lint
…es (#6)

* Implement MCP server for GA

* Setup unit-tests to new storages

* style: apply rustfmt

* Potential fix for pull request finding 'CodeQL / Workflow does not contain permissions'

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>

* Add tests

* Fix tests

* Fix tests

* Fix tests

* Fix lint

---------

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
* Implement new embed vector and similar search

* style: apply rustfmt

* Fix unit test

* Add badges

* Fix coverage
Add a security policy document outlining supported versions and vulnerability reporting.
* Truncate APIs and support graphql with mermaid

* style: apply rustfmt
* Truncate APIs and support graphql with mermaid

* Implement to support windows

* style: apply rustfmt

* Potential fix for pull request finding 'CodeQL / Workflow does not contain permissions'

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>

* Fix lint

---------

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
* Implement flow to install codegraph to MacOS and linux

* Add missing pipeline to support releasing

* Fix issue with doctor

* style: apply rustfmt
…rce code as possible (#13)

* Bump version to v2.0.3

* Update ci

* Update ci again

* Remove coverage

* Setup codecov.yml
hungpham10 and others added 26 commits September 6, 2026 23:03
- Add support to build graph from binary using radare2 for reverse binary
code
- Fix issue when working with lambda function which is a core feature to
analyze obfuscated code
* Fix issue missing function name and crash when picking to wrong function

* Bump version to v2.1.1

* style: apply rustfmt
* Improve by showing more meaningful object naming from r2

* Increase code-coverage

* style: apply rustfmt
* Implement tool to convert structure documents into graph

* Bump version to v2.1.3
* Integrate codegraph-docs into main flow

* style: apply rustfmt

* Fix lint

* style: apply rustfmt
* Persistent document graph into disk

* Inplement nginx parser

* style: apply rustfmt

* Fix lint

* style: apply rustfmt

* Bump version to v2.1.5
* Split storage of binary graph into different database

* Fix lint
* Support YAML multiple files

* style: apply rustfmt

* Bump version to v2.1.6
* Show progress bar while indexing documents

* Improve performance when starting

* Bump version to v2.1.7

* style: apply rustfmt

* Fix lint
* Fix document graph pipeline and wire real pattern/value search

The document tools were unusable end-to-end due to four latent bugs:

- parsers: parent->children links were never persisted (nodes cloned into
  Document.nodes before children wiring), so hydrate could never descend.
- graph: path token chains were built in reverse order (root ended up last),
  so full-path queries never matched; also node cache was materialized after
  trie insertion, so the first ingest interned no keys.
- graph: the four trie projections shared one storage without namespace —
  radix root pointers collided per shard, leaving only the last-written trie
  reachable. Add shard_bias to Radix/Search (default 0) and give each docs
  trie a distinct shard range.
- graph: the string interner was RAM-only while tries persisted, so interned
  token payloads dangled after restart. Persist the interner blob alongside
  docs and restore it in open(); doc ids move to their own id range (>=6e11)
  so they no longer collide with node ids.

MCP/CLI wiring:
- doc_search now parses dotted patterns into token chains via the interner,
  with a case-insensitive key-scan fallback for non full-path queries.
- new tools: doc_search_value (scalar substring scan), doc_ingest_dir
  (recursive bulk ingest with limit), doc_remove.
- doc_hydrate takes max_depth to keep LLM payloads small; doc_list returns
  per-doc metadata (doc_id, path, format, root_node_id, nodes) instead of
  bare counts; search results include key/index.
- CLI: fix `doc ingest` clap panic (positional `path` clashed with the
  global --path arg, renamed to `file`); doc search uses the same real
  pattern resolution.

* Bump version to v2.1.8

* style: apply rustfmt
#30)

* Add fuzzy key match, pattern mining (P#) and IDF ranking to document graph

Phase 3 of the structured document graph: fuzzy retrieval, cross-document
pattern mining and rarity-based ranking (rare structure = high information,
IDF-style scoring).

- Fuzzy key match: search_key_fuzzy scores distinct keys by exact/prefix/
  contains/Levenshtein similarity with an IDF bonus (rarer keys score
  higher). `doc_search` accepts `~segment` to force fuzzy and falls back
  exact-substring -> fuzzy when the full-path trie misses.
- Pattern mining: mine_patterns() counts kind chains (wildcard FIELD/IDX
  payloads) ending at scalar leaves over a max_depth window, assigns stable
  pattern ids (P#, registry persisted in storage and restored on open) and
  indexes chains into the previously dead pattern_trie. Results carry
  node_count / doc_count / doc_freq, sorted by document frequency ascending
  so characteristic patterns surface first and background (~1.0) sinks.
- Structural search: search_kind_chain() matches nodes whose ancestor kind
  window ends with the query chain (e.g. "MAP, FIELD, NUMBER"); results are
  ranked by pattern uniqueness IDF. search_path_scan() complements the radix
  trie whose leaf holds a single record per chain, returning all nodes with
  the same key path across documents (spec.replicas: 97 hits on the infra
  repo instead of 1).
- Depth semantics fixed in doc_search: depth counts extra levels BELOW the
  pattern, so the radix key-length filter gets pattern_len + depth.
- New MCP tools: doc_mine_patterns, doc_list_patterns, doc_search_struct.
  New CLI commands: `codegraph doc patterns`, `codegraph doc struct`.

* style: apply rustfmt

* Fix lint

* Bump to version v2.1.9
#31)

* Fix binary callees/callers/flow always empty: route id >= bin_base to BinaryGraph (#26 regression)

After the storage split, binary symbols live only in binary.sqlite (id range
bin_base = 2e9) while codegraph_callees/callers/flow still queried the main
GraphIndex, so every binary returned empty results. Now those tools route to
BinaryGraph when the node id falls in the binary range, with a fallback to the
old path if the binary DB is unavailable.

- BinaryGraph: add callees (call records resolved by name within the same
  binary), callers (BFS over a new caller_of:{name} reverse index built at
  ingest) and flow (chain render with CFG markers + call sites, same shape as
  GraphIndex::flow).
- Ingest: stop corrupting chain entries — CFG markers (id < SYMBOL_BASE) and
  unresolved-call placeholders (0) are kept as-is; only real symbol ids are
  remapped into the bin_base range.
- MCP: dispatch_binary_graph routes callees/callers/impact/flow for binary ids.
- Config template: document the [bingraph] section (enabled/bin_base/storage)
  with a warning against overlapping id ranges.
- Docs: binary-analysis.md updated for the split-storage architecture.
- Tests: bingraph unit tests extended + end-to-end test compiling a real
  shared library and running it through r2 extraction and queries.

* style: apply rustfmt
r2 sometimes fails to recover call ops for a shared library (entrypoint
detection fails on Mach-O dylibs), producing chains without call sites. Retry
extraction up to 3 times with the extract cache disabled and only accept a run
whose chains contain at least one resolved call.
- codegraph_graphcode_*: class, list_types, function_scope,
  search_by_annotation, files, dependencies, sandbox, diff,
  diff_simulate, origin_simulate + new codegraph_graphcode_stats
- codegraph_doc_* -> codegraph_graphdoc_* (all 11 tools)
- codegraph_binary_list/addr/stats -> codegraph_graphbin_*;
  binary_search merged into codegraph_search_symbol via new
  `source` arg (all|code|binary)
- Shared/session tools keep their names (symbol, search_symbol,
  callers, callees, impact, flow, context, references, mermaid,
  search_flow, init/deinit/index, query_usage_report)
- codegraph_status now aggregates all three datasets (code + doc +
  binary, null when absent)
- Minimize everywhere: doc/binary outputs routed through
  emit_value (omit_defaults); graphbin list/addr support
  `format` with fixed-order row arrays
- Docs updated: server-instructions.md, README, architecture,
  binary-analysis
* Update graphql to support new models

* Bump version to v2.2.1
* Add subcommand `clean` and remove subcommand `doc`

* Bump version to v2.2.2
* Add resume-able functions to avoid hanging

* Optimize cost by compressing response from MCP

* Fix lint

* Change minimize to minimal
* perf(lmdb): batch radix DFS reads + cache probe DBI handle

Search song song tren LMDB chạm gioi han reader-slot (`MDB_BAD_RSLOT`):
moi `get_node`/`get_children` la mot `begin_ro_txn` rieng, mot lan search
tao O(nodes) read-txn, nhieu request thi nhan lai gap gioi han.

Tang 1 - cache DBI handle trong probe_version:
`probe_env` giu them `Database` handle (Copy) cho `sg_meta` thay vi
`open_db` lai moi request. `SharedGraphIndex::ensure_fresh` chay
`current_version` truoc moi request nen moi lan deu ton mot `open_db`
(tha bang `Environment`) va mot `begin_ro_txn`; gio con 1 read-txn.

Tang 2 - batch read tren CategoryStorage:
them `get_nodes`/`get_childrens` doc nhieu id trong MOT read-txn. Default
impl goi lai ham don le nen 6 backend nen (in-memory/sqlite/redis/pg/mysql/
cached) khong doi. LMDB override de dung 1 txn cho ca danh sach.

`Radix::search_dfs` chuyen sang goy cac lan doc: doc node cua frame + tat ca
child chua duyet trong 1 lenh goi, thay vi mot lenh goi cho tung node. So
read-txn moi vong DFS giam tu ~2*children xuong 1. `base` giu nguyen thu tu
duyet nen hanh vi DFS khong doi.

Test:
- batch khop `get_node`/`get_children` tuang tung, thu tu va sort giu nguyen
- id khong ton tai van tra `BranchOutOfRange` nhu ban don le
- regression: 8 task doc song song khong loi
- `probe_version` qua cache handle 100 lan van dung version

Benchmark moi `lmdb_batch_read` (feature `lmdb`) do doc 1/16/64 node don
lue vs batch, ca voi `get_children`:
cargo bench -p codegraph-graph --bench lmdb_batch_read --features lmdb

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>

* style: apply rustfmt

* fix(lmdb): sửa 2 lỗi compile trong batch read

- `lmdb::Error::Other` nhận `c_int` (mã lỗi C), không phải `String`.
  `open_db` đã trả `lmdb::Result<Database>` nên chuyển thẳng với `?`.
- `children` trong `search_dfs` (Collect state) cần `mut` để gọi
  `.pop()` — tương tự `node` ngay dòng trên.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>

* fix(lmdb): sửa 3 lỗi compile trong test batch read

Test viết tay sai cú pháp:
- `storage::EMPTY` không resolve trong `mod tests` (chỉ có `EMPTY` từ
  `use super::*`) → E0433
- `.map(|i| s.new_node(..).await)` — `.await` trong closure không async
  → E0728. Vòng for thay thế.
- `b.new_node(..)` gọi method trên `usize` do nhầm biến `b` (id) với
  `s` (storage) → E0599
- `assert_eq!(batch[0], kids)` so `Vec` với array → so với bản sort

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>

* fix(radix): bỏ needless borrow khi đọc child node từ batch

clippy::needless_borrow: `Self::to_vec(&cp_bytes)` trong khi `cp_bytes`
đã là `&Vec<u8>` từ `&child_nodes[rel]` → bỏ dấu & thừa.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>

* fix(storage): allow clippy::double_must_use cho các trait async_trait

clippy 1.99 (rust-toolchain.toml dùng channel = stable) bật lint
double_must_use: async_tair tự sinh must_use cho future, trùng với
kiểu future boxed vốn đã must_use. Lỗi nằm trong macro của async_trait
nên allow tại chỗ sinh ra, không đổi hành vi.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>

---------

Co-authored-by: Hung Pham <[email protected]>
Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
* perf(bench): thêm profiler RAM + workflow đo memory trên CI

`GraphIndex` là in-memory-first: `rebuild()` nạp toàn bộ symbols/chains/
call_names/edges vào HashMap. CodSpeed chỉ đo thời gian nên cái giá RAM này
không ai thấy — thêm đường đo riêng trước khi tối ưu.

- `codegraph-graph/src/meminfo.rs`: đọc RSS std-only (Linux `/proc/self/statm`
  + `VmHWM`, macOS fallback `ps -o rss=`), kèm `MemTracker` lấy mẫu theo
  phase. Không thêm dependency.
- `codegraph-bench/src/bin/mem`: đo RSS theo extract → index → query trên
  repo thật, đồng thời dự đoán RAM theo cấu trúc
  (`count × size_of::<T>()`) để quy kết quả về đúng HashMap nào.
  Chế độ `--synthetic N` sinh index deterministic (không phụ thuộc network)
  với quy mô đủ để `edges`/`call_names` chiếm RAM đáng kể.
- `.github/workflows/memory.yml`: chạy profiler, in bảng vào job summary,
  upload artifact. Không chặn merge — GitHub runner là VM dùng chung nên
  RSS chỉ dùng so sánh tương đối, phần `predicted` mới là số ổn định.

Còn lại: tối ưu `edges`/`call_names`/`symbols` ra khỏi RAM (storage-backed
lazy). PR này mới phần đo, chưa đụng storage trait.

* style: apply rustfmt

* fix(bench): sửa 3 lỗi CI — needless_return, const chết, thiếu import

Báo qua CI của PR #38:

- `meminfo.rs`: tách `rss_linux()` riêng để `rss_bytes()` không có `return`
  (clippy `needless_return`, fail cả job clippy lẫn slim build).
- `meminfo.rs`: `FALLBACK_PAGE_SIZE` chỉ dùng trong nhánh linux → gate
  `#[cfg(target_os = "linux")]` (dead_code trên Windows).
- `bin/mem.rs`: bổ sung import `Annotation`, `CallRecord`, `EffectType`,
  `ScopeLevel`, `SymbolKind` — patch import đầu tiên không áp dụng nên các
  kiểu này chưa có tên trong scope.

* style: apply rustfmt

* fix(bench): mỗi repo một process riêng khi đo memory

Chạy `mem` với nhiều repo trong cùng process làm RSS cộng dồn: repo sau đo
trên nền của repo trước nên `rss_after_extract` không còn đọc được (synthetic
thì không dính vì chỉ đo 1 index).

Tách từng repo ra process riêng — chậm hơn vài giây nhưng số đo đúng.

* perf(graph): quy RSS về đúng cấu trúc dữ liệu (mem_breakdown)

Baseline đo được trên CI cho thấy 320 MiB `Δ index` với 50k symbol / 200k edge,
nhưng `symbols + edges` chỉ quy được 29.7 MiB — 91% nằm ở nơi khác. Không có
cách nào chọn đúng cấu trúc cần tối ưu khi chỉ có tổng RSS.

- `memtrack.rs`: deep size từng cấu trúc in-memory của `GraphIndex`
  (`symbols`, `chains_map`, `call_names`, `edges`, `name_index`, `scope_index`,
  `name_keys`, `files`). Heap của `String`/`Vec` đo exact bằng `capacity()`;
  bucket `HashMap` ước lượng theo công thức hashbrown. Ghi rõ trong doc phần
  nào exact / ước lượng / không tính.
- `LruCache::len` / `is_empty` (DashMap::len O(1)) + `Storage::cache_occupancy`
  (default `Vec::new()`, chỉ `CachedStorage` override) → biết LLR đang giữ bao
  nhiêu entry.
- `GraphIndex::mem_breakdown()` gom tất cả lại, `ranked()` sort giảm dần để
  thứ tự tối ưu là thứ tự xếp hạng.
- `bin/mem`: in bảng breakdown + JSON (`breakdown`, `accounted_total`, `caches`).

Chưa đổi hành vi production: các method mới đều chỉ đọc, không vào hot path.

* style: apply rustfmt

* fix(graph): memtrack — `&str`/`&[T]` không có `.capacity()`

CI báo 4 lỗi E0599 + 1 E0631: helper deep-size nhận `&str`/`&[T]` nhưng gọi
`.capacity()` — method đó chỉ có trên `String`/`Vec`.

Đổi helper sang `&String`/`&Vec<T>` và thêm `#[allow(clippy::ptr_arg)]` kèm lý
do: cần bytes **đã cấp** (`capacity`), slice không đọc được — dùng `len` sẽ
lệch tới ~2× do `Vec` over-allocate.

* fix(graph): sửa clippy — dead_code `is_empty`, 4 redundant_closure, sort_by

- `LruCache::is_empty` không nơi nào gọi → bỏ (chỉ cần `len` cho occupancy).
- `.map(|x| f(x))` → `.map(f)` ở 4 chỗ trong `memtrack`.
- `ranked()` dùng `sort_unstable_by` + tiebreaker theo tên: vừa tránh
  `stable_sort_primitive`, vừa giữ thứ tự **deterministic** khi hai cấu trúc
  trùng bytes (sort thuần tuỳ thứ tự ban đầu → báo cáo khác nhau giữa 2 lần chạy).

* style: apply rustfmt

* perf(bench): đo vi sai radix engine bằng 3 shape synthetic

Breakdown hiện tại chỉ tính HashMap của `GraphIndex` (97.5 MiB / 320 MiB Δ).
~223 MiB còn lại nằm ở hai radix engine — sống trong `InMemoryStorage` chứ không
phải HashMap nên `mem_breakdown` không thấy.

Cách đo: dựng cùng 50k symbol nhưng bỏ dần từng thành phần, lấy hiệu RSS.

- `symbols` — chỉ symbol → `symbols` + `name_index` + **name engine**
- `chains`  — + chain, không call record → **chain engine** (trie + rt_shortcuts)
- `full`    — + call record → `call_names` + `edges`

Không đụng `Storage` trait nên không có rủi ro 6 backend; đổi hoàn toàn ở tầng
benchmark. Kết quả quyết định có nên tối ưu `call_names` (30 MiB) hay không.

* style: apply rustfmt

* perf(bench): thêm `--mode open` để đo chi phí thường trực, không phải peak

Sửa lỗi phương pháp: `rss_after_index` đo **high-water của allocator**, không
phải live memory. Trong `ingest` tồn tại 3 bản sao call song song —
`results` (borrow, `lib.rs:983`) → `all_calls` (clone, `:1054`) →
`recs_by_caller` (clone, `:1389`) — rồi cả hai bản sao cuối bị drop nhưng
allocator không trả arena về OS. Nên "137 MiB do call records" ở lượt trước là
đỉnh lúc ingest, không phải chi phí giữ lại sau restart.

`--mode open`: ingest vào sqlite → `drop(index)` → **`drop(parsed)`** → `open`
lại. `open` chỉ chạy `rebuild()` (nạp blob + dựng HashMap) nên không có bản sao
tạm nào; `rss_after_open − rss_before_open` là chi phí thường trực thật.

Chạy cả `--mode ingest` và `--mode open` cho 3 shape trong CI để so trực tiếp:
hiệu giữa hai mode = phần transient, còn mode open = phần phải giữ.

* perf(graph): bỏ 2 bản clone `CallRecord` khi ingest

`ingest` chỉ cần sửa đúng một field của `CallRecord` là `caller_id` (sang
id global) nhưng lại clone toàn bộ struct (176 B + heap) hai lần:

- `all_calls: Vec<CallRecord>` → `Vec<CallRef>` (16 B, mượn từ `ParseResult`
  của caller vốn sống suốt `ingest` + `caller_id` đã remap).
- `recs_by_caller: HashMap<u64, Vec<CallRecord>>` → `Vec<u32>` (index vào
  `calls`), tra lại qua `calls[i].rec`.

200k call record: 2 × ~43.6 MiB → ~4 MiB, tiết kiệm ~83 MiB / peak 319.7 MiB
(~26%) theo số đo `--mode ingest` shape `full`.

Tiền lệ có sẵn: `resolve_calls` đã dùng `Vec<&CallRecord>` để group.

Ghi chú: `caller_id` trong blob persist giờ mang id **local** thay vì global.
Vô hại — `flow`, `rebuild_edges`, `bingraph` đều lấy caller id từ key của
blob, không đọc field đó; đã ghi chú tại chỗ ghi.

Test mới `call_records_grouped_by_global_caller_id`: hai file cùng dùng
`SYMBOL_BASE` làm caller local, chặn bug gom nhầm theo id local.

Kèm: `--mode open` trong CI hạ 50k → 5k (rebuild() còn nạp `all_call_records()`
rồi deserialize lần hai, 50k mất ~1h40m/shape ≈ 5h/job), và sửa 3 comment
mô tả "3 bản sao call" đã lỗi thời.

* Bump version to v2.2.5

---------

Co-authored-by: Hung Pham <[email protected]>
* refactor(source): gom mọi đường đọc I/O về một trait `Source`

Thêm crate `codegraph-source` — chiều **đọc** dữ liệu đầu vào, đối xứng với
`Storage` trait (chiều ghi) đã có sẵn 7 backend.

## Vấn đề

Trước đó phần đọc rải rác ở 4 crate, mỗi nơi tự gọi `std::fs`:

- 3 bản `WalkBuilder` **copy y hệt** cùng 5 option (hidden, git_ignore,
  git_exclude, parents, `.codegraphignore`) ở `extract/walker.rs`,
  `extract/config.rs` (`detect_project_header_hint`) và `binary/scan.rs`.
- `binary/scan.rs` đọc **cả file** chỉ để so 4 byte magic → cấp phát hàng trăm
  MB cho file nhị phân lớn mà không dùng tới phần đệm.
- `codegraph-context` đọc `std::fs` lúc **query** (`include_source`).

## Thiết kế

```rust
pub trait Source: Send + Sync {
    fn kind(&self) -> SourceKind;
    fn root(&self) -> &Utf8Path;
    fn config(&self) -> &SourceConfig;
    async fn list(&self) -> Result<Vec<SourceEntry>>;
    async fn read(&self, entry: &SourceEntry, limit: Option<usize>) -> Result<Vec<u8>>;
    async fn materialize(&self, entry: &SourceEntry) -> Result<Utf8PathBuf>;
}
```

Ba method vì ba lý do khác nhau:

- `list` — discovery, lọc diễn giải theo `SourceKind` (Code: extension,
  Document: glob, Binary: không lọc).
- `read(entry, limit)` — đọc in-memory tối đa N byte đầu.
- `materialize` — **bắt buộc cho binary**: `R2Pipe::spawn(path)` là tiến
  trình ngoài tự mở file bằng path hệ thống, không nhận bytes qua stdin.

`SourceConfig` cho phép cấu hình **theo dịch vụ** (`kind` quyết định cách diễn
giải `include`, mỗi dịch vụ một bộ default riêng).

## Quy ước giữ được schema persist

`SourceEntry.path` luôn **tương đối so với root**. `Symbol.file` lưu dạng
`<root>/<rel>` nên `entry_from_symbol_file(file, root)` tách prefix là dựng lại
được entry — đường query **không cần cột `source_id`, không cần migration**.

## Áp dụng

- `binary/scan.rs`: bỏ `WalkBuilder` thứ 3, đọc `Some(4)` byte đầu thay vì
  `std::fs::read` cả file.
- `docs/graph.rs`: tách `ingest_bytes` / `ingest_source` — `ingest_file` còn là
  wrapper mỏng nên hành vi cũ giữ nguyên.
- `context`: `build`/`build_response` nhận `Option<&dyn Source>`; MCP dựng
  `DiskSource::for_query(root)`, GraphQL dùng `session.root()`.
- Bỏ `if let Ok(...)` **nuốt lỗi im lặng** ở đường query — giờ `tracing::warn!`.

## Vì sao core của `DiskSource` là sync

Mọi đường ingest hiện tại đã block sẵn: `parse_files` dùng rayon (blocking
pool), `collect_binaries` spawn process `r2`. Bọc `spawn_blocking` chỉ tốn một
`SourceConfig` clone mỗi lần gọi mà không đổi gì quan sát được — nên core là
sync, method `async` của trait chỉ là bọc mỏng. Trait vẫn `#[async_trait]` để
provider từ xa (git host, object store) có chỗ cho I/O bất đồng bộ thật.

## Chưa làm (đợt sau)

`codegraph-extract` (`walker::walk`, `detect_project_header_hint`, `parse_one`,
`doc_config` glob) chưa đi qua trait — là phần rủi ro cao nhất vì `parse_files`
phải tách 2 pha async-fetch + rayon-parse.

* style: apply rustfmt

* fix(source): sửa 2 lỗi compile CI — camino extension() trả &str, Utf8PathBuf không có read_bytes

* fix(source): sửa 3 clippy lint — manual_contains, double_must_use, repeat_n (giữ MSRV 1.80)

* fix(binary): import trait Source để gọi được `root()` trên DiskSource

* fix(context): sửa assertion test — sym có end_line=2 nên lấy cả 2 dòng

* fix(source): test cần tạo .git — `ignore` chỉ áp .gitignore bên trong git repo (khớp hành vi walker.rs cũ)

* refactor(binary): stat đi qua SourceEntry thay vì tự gọi fs::metadata

`cache_path` gọi `fs::metadata` 2 lần (mtime + size) và được `collect_binaries`
gọi tới **3 lần** cho mỗi binary → 6 syscall chỉ để dựng cache key. Tệ hơn:
đó là I/O nằm ngoài `Source` trait, nên provider không có stat cục bộ sẽ vỡ ở
đúng chỗ này.

Thay vì thêm trait `SourceMetadata`, dùng thứ đã có sẵn:

- `SourceEntry` thêm field `mtime` (song song `size`).
- `DiskSource::list_blocking` lấp cả size và mtime từ **một** syscall `metadata`
  mà nó vốn đã gọi — chi phí bằng 0.
- `find_binaries*` trả `Vec<SourceEntry>` thay vì `Vec<Utf8PathBuf>`. Trước đó nó
  vứt mất stat khi chuyển sang path tuyệt đối.
- `cache_path` nhận `&SourceEntry`, bỏ `mtime()`/`size()`.
- `collect_binaries` tính cache path **một lần** thay vì 3 lần.

Net: 6 syscall → 1 cho mỗi binary, và `stat` nằm trong crate `codegraph-source`.
Hành vi cache **không đổi**: key vẫn là `(EXTRACT_VERSION, path, mtime, size)`.

* style: apply rustfmt

* refactor(extract): gom 2 `WalkBuilder` cuối về `Source::list`, sniff header qua `Source::read`

Bản `WalkBuilder` thứ 3 và thứ 4 biến mất — traversal giờ chỉ có một nguồn sự
thật trong `codegraph-source`.

## `header_looks_like_cpp` đọc cả file chỉ để sniff 8KB

```rust
let bytes = std::fs::read(path)?;           // cả file
let sample = &bytes[..bytes.len().min(8192)]; // lấy 8KB
```

Nay là `source.read_blocking(entry, Some(8192))` — đúng 8KB. Cùng bug đã sửa ở
`find_binaries`. Header C++ sinh máy hoặc header kéo theo header sẽ không còn bị
nạp toàn bộ vào RAM chỉ để xem 8KB.

## 2 pass → dùng chung 1 traversal

`orchestrator::parse_project` gọi `walker::walk`, mà `walk_options` lại gọi
`detect_project_header_hint` (dựng `WalkBuilder` thứ hai để đếm `.c`/`.cpp`). Giờ
cả hai đi qua cùng `DiskSource`.

Hành vi giữ nguyên: cùng policy ignore (hidden, git_ignore, git_exclude, parents,
`.codegraphignore`), cùng bộ đếm extension, cùng logic chọn parser.

## Thêm 2 test

- `hint_nhận_diện_project_c_thuần_và_cpp_thuần` — C thuần / C++ thuần / mixed.
- `sniff_chỉ_đọc_8kb_đầu_header_lớn` — fixture >8KB chứng minh chỉ đọc ngưỡng.

## Gỡ dep thừa

`ignore` không còn được dùng trong `codegraph-extract` (chuyển sang
`codegraph-source`).

* refactor(source): tách `Source` nhiều tầng + gom 2 đường đọc config.toml

`Source` trước là một trait 6 method. Tách thành 4 tầng để giảm áp lực khi
thêm provider: `SourceInfo` (nhận diện + policy) → `SourceListing`
(discovery) / `SourceReader` (đọc bytes) / `SourceMaterializer` (file thật
cho binary). `Source` gộp lại nên chỗ đã dùng `&dyn Source` không phải sửa.
Cùng cách `Storage` tách `CategoryStorage` / `NodeMetaStorage` trong
codegraph-graph.

`SourceReader` có thêm `read_optional()`: config là đường đọc *tuỳ chọn*
(thiếu `.codegraph/config.toml` là bình thường) — trước đây mỗi call site tự
`match Err(_)` nuốt lỗi. Chỉ `NotFound` mới thành `None`, lỗi I/O khác vẫn
là `Err` để không giấu sự cố thật.

Gom 2 chỗ đọc config còn sót lại về Source:
- `ExtractConfig::load_from` và `SboxConfig::load_from` nhận `&DiskSource`
  thay vì tự gọi `fs::read_to_string`.
- `CONFIG_REL_PATH` / `CONFIG_MAX_BYTES` nằm trong `codegraph-source` làm
  một nguồn sự thật duy nhất cho quy ước `.codegraph/config.toml`;
  `ensure_repo_id` dùng chung const (phần ghi vẫn gọi `fs` trực tiếp —
  thuộc chiều ghi, không gom).
- Đường đọc config **không qua discovery**: `.codegraph/` là hidden +
  gitignore nên `list()` không bao giờ trả về entry này.

Test dùng fixture dựng root đúng hình dạng thật (`.codegraph/config.toml`)
thay vì path tuỳ ý — không còn đường đọc config nào để bypass.

* style: apply rustfmt

* fix(source): tên test snake_case — `None`/`Some` viết hoa vi phạm non_snake_case

---------

Co-authored-by: Hung Pham <[email protected]>
… thẳng

`control_flow` trước đây chỉ nối tuần tự mọi node trong chain
(`windows(2)`), nên diagram của một hàm không cho biết nhánh if/else, vòng
lặp, switch, và call đi đâu — dù `FlowResult.calls` đã có đủ dữ liệu.

Thay bằng cách render dùng cả `flow.calls`:
- điều kiện (IF_TRUE/IF_FALSE/SWITCH_CASE) → hình thoi, label kèm guard text;
- LOOP/LOOP_BACK/RETURN/THROW/... → stadium;
- call → hộp chữ nhật có tên callee, `· L<line>`, `· if <condition>`, effect;
- call không resolve (ngoài repo) → hộp subroutine + `· ext`;
- cạnh suy từ marker: else rẽ từ if, LOOP_BACK vẽ back edge về header,
  SWITCH_CASE toả từ entry, BRANCH_END hợp nhất nhánh.

Cập nhật mô tả tool MCP `codegraph_mermaid` và doc MermaidKind (GraphQL).
Thêm 6 unit test cho control_flow.
@hungpham10

Copy link
Copy Markdown
Author

Đóng nhầm repo — PR này được tạo trên base sai (upstream đang ở v1.2.0). Branch phát triển thực tế dựa trên hungpham10/codegraph-rs main (v2.2.5). Sẽ mở lại PR đúng repo.

@hungpham10 hungpham10 closed this Oct 6, 2026
@Cleboost

Cleboost commented Oct 6, 2026

Copy link
Copy Markdown
Owner

Pls for your next PR make it in english

@Cleboost

Cleboost commented Oct 6, 2026

Copy link
Copy Markdown
Owner

And pls split your code in features with multiple pr, not big one. It's easier for me to review :)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants