Skip to content

feat: add scale and missingness bundling capacity experiments - #2

Merged
prrao87 merged 2 commits into
mainfrom
codex/testing-bundling-capacity-2
Oct 5, 2026
Merged

prrao87 merged 2 commits into
mainfrom
codex/testing-bundling-capacity-2

Conversation

@prrao87

@prrao87 prrao87 commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

Why

The questionnaire study tests how many facts can be recovered from one MAP bundle. Part 2 asks what happens when a small record must be retrieved among 920,000 candidates and some fields are missing. These experiments provide the measured evidence for that continuation and HYP-83.

What changed

  • Add an exhaustive synthetic PERSON fixture with age, job, region and three interests, separate calibration/evaluation panels, and a fixed 30% missingness mask.
  • Measure chance similarity across five dimensions, three seeds and candidate pools from 100 to 920,000; compare omission, shared null tokens and mixed null spellings; check retrieval with missing fields; and compare LanceDB IVF_PQ with exact search.
  • Separate finite-dimensional MAP error from similarity introduced by the missing-value strategy, using exact content and encoder baselines, frozen thresholds and query-level uncertainty.
  • Include a locked environment, reproduction commands, per-query evidence, summaries and PNG/SVG figures. Keep the generated database and reproducible intermediates out of Git.
  • Select index scores explicitly to avoid deprecated projection warnings, and document that fresh IVF_PQ builds can vary because training has no fixed seed.
  • Organize both studies under their own reports, update the repository overview, repair report and figure links, and apply Canadian English and routine formatting fixes.

The saved results show 99.1% content top-ten agreement with omission versus 67.9% with a shared null token at 30% missingness. Complete records produced no observed unrelated matches above calibrated thresholds. The recorded IVF_PQ run achieved 80.0% recall without refinement and up to 99.7% with refinement. These are results for the declared synthetic fixture and query panel, with no all-pairs or billion-record guarantee.

Validation

  • All 15 tests pass, including encoder invariants, exact float16 storage, streamed search against brute force, and threshold calibration.
  • Ruff lint and formatting checks pass; the installed environment matches the lock.
  • All 28 local documentation and figure links resolve; the saved figures were visually inspected.
  • Full reproduction completed in a clean temporary store. All five fixture inputs match byte for byte; all 42,000 per-query rows, 720,000 pairwise rows and 21 frozen thresholds reproduce exactly.
  • A fresh IVF_PQ build confirmed the refinement result: 79.75% to 79.80% unrefined recall, up to 99.85% with refinement, and no unrelated returned records above threshold. The small difference from the saved run reflects index training variation; saved evidence is preserved.
  • The explicit distance projection returned identical records and scores in 36 searches spanning all probe/refinement settings.

@prrao87
prrao87 marked this pull request as ready for review October 5, 2026 14:55
@prrao87
prrao87 merged commit 450c4e6 into main Oct 5, 2026
@prrao87
prrao87 deleted the codex/testing-bundling-capacity-2 branch October 5, 2026 14:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant