Skip to content

fix(server): send float16 multivectors to the sidecar as bytes - #416

Merged
fm1320 merged 3 commits into
mainfrom
fix/sidecar-f16-multivectors
Sep 30, 2026
Merged

fm1320 merged 3 commits into
mainfrom
fix/sidecar-f16-multivectors

Conversation

@fm1320

@fm1320 fm1320 commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

In SIE's cluster setup, which managed SIE uses, a request does not reach the Python server over HTTP. It goes from the gateway through the queue to a Rust sidecar that runs next to each GPU worker, and the sidecar talks to the Python worker over a local socket (IPC). When the worker hands encode results back, multi-vector outputs were converted to lists of Python floats, which msgpack packs as 9-byte doubles, and the whole batch's response is capped at 128 MiB.

Models with wide multi-vector outputs overflow that cap, and then the whole batch fails. TopK-Embed-V1, being added in #417, returns 2,048 numbers for every token: one ordinary full batch of it is about 130 MiB on this hop. Wide models SIE already ships can hit the same cap with large enough batches, for example nvidia/llama-nemoretriever-colembed-3b-v1 (3,072 numbers per token) and nvidia/nemotron-colembed-vl-4b-v2 (2,560).

The sidecar already accepts a float16 matrix as raw little-endian bytes (MultivectorOutput.values_f16, which publisher.rs turns into the wire payload) and says so on every encode batch (accepts_batched_f16_multivectors). The worker now uses that for float16 multivectors: 2 bytes a value instead of 9, so the same batch is about 29 MiB.

Sizes

Measured with SIE's own packing code (_maybe_multivector_raw_output, the response envelope, msgpack.packb), before and after:

Batch Before After
TopK-Embed-V1 small (2,048 numbers per token), one page 21.6 MiB 4.8 MiB
TopK-Embed-V1 small, one full 8,192-token server batch 129.7 MiB, rejected 28.8 MiB
TopK-Embed-V1 small, one 8,192-token document 144.0 MiB, rejected 32.0 MiB (sent in chunks)
TopK-Embed-V1 xsmall (1,024 numbers per token), one full 16,384-token server batch 140.5 MiB, rejected 31.2 MiB

Compatibility

Only float16 multivectors change, and only when the sidecar sets the flag; production sidecars always do (dispatcher.rs). float32 outputs, and sidecars without the flag, keep the list of floats. The single-server HTTP path does not use this hop and is unchanged.

Testing

  • New tests in test_queue_executor_stage1d.py: bytes are sent only with the flag and only for float16, the outcome through process_encode_batch carries them, and a packed matrix takes about 2 bytes a value instead of 9.
  • Full default suite: 9,432 passed.
  • Not run on a live cluster.

Summary by CodeRabbit

  • New Features
    • Float16 multivector results can use a compact binary format when the connected service supports it, reducing payload size. Float32 results and float16 results without that support continue to use the existing list format. The binary format carries contiguous little-endian float16 values, while the existing format remains available for other result types and unsupported connections.

On the cluster path the worker hands each encode result to the Rust sidecar
over IPC. Multivectors went as lists of Python floats, which msgpack packs as
9-byte doubles, in one response capped at 128 MiB. Wide models overflow it: a
full 8,192-token batch of TopK-Embed-small (2,048-dim tokens) packed to
130 MiB and one 8,192-token document to 144 MiB, so the whole batch failed.

The sidecar already takes a float16 matrix as little-endian bytes
(MultivectorOutput.values_f16) and says so on every encode batch
(accepts_batched_f16_multivectors). The worker now sends float16
multivectors that way when the flag is set: 2 bytes a value, so the same
batch is 29 MiB and the document 32 MiB. float32 outputs, and sidecars
without the flag, keep the list.

Tests: bytes only with the flag and only for float16, the outcome through
process_encode_batch, and the packed size per value.
@fm1320
fm1320 requested a review from a team as a code owner September 29, 2026 13:17
@coderabbitai

coderabbitai Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Team

Run ID: 945378ab-809c-4feb-a4b4-815c97e38dea

📥 Commits

Reviewing files that changed from the base of the PR and between 799a45e and 71e1f76.

📒 Files selected for processing (2)
  • packages/sie_server/src/sie_server/queue_executor.py
  • packages/sie_server/tests/test_queue_executor_stage1d.py

Included review availability: This review used your included allowance. 9 included reviews remain after this review. Your included PR review attempts over the past 7 days set your current allowance at 10 reviews per hour.


📝 Walkthrough

Walkthrough

The encode queue now supports little-endian byte output for float16 multivectors when the request enables the capability. Tests cover byte framing, list output, float32 output, packed size, and end-to-end encoding.

Changes

Float16 multivector encoding

Layer / File(s) Summary
Float16 byte framing and output tests
packages/sie_server/src/sie_server/queue_executor.py, packages/sie_server/tests/test_queue_executor_stage1d.py
The output helper can encode float16 values as contiguous little-endian bytes in values_f16, with values empty. Tests cover the byte and list forms, confirm float32 remains a list, and compare packed sizes.
Capability propagation through batch execution
packages/sie_server/src/sie_server/queue_executor.py, packages/sie_server/tests/test_queue_executor_stage1d.py
Encode-batch execution passes the request capability through encode groups and invalid-input isolation retries. An end-to-end test checks byte output when the capability is enabled.

Sequence Diagram(s)

sequenceDiagram
  participant process_encode_batch
  participant _run_encode_group
  participant _maybe_multivector_raw_output
  process_encode_batch->>_run_encode_group: Pass accepts_batched_f16_multivectors
  _run_encode_group->>_maybe_multivector_raw_output: Pass f16_bytes for multivector output
  _maybe_multivector_raw_output-->>_run_encode_group: Return values_f16 bytes or list values
Loading

Suggested reviewers: svonava

Priority: ➖ Normal

Merge Risk: ⚪ Minimal · up to 71e1f

The float16 byte output matches the sidecar’s supported format and preserves existing output for other requests. No actionable merge-blocking risk remains; normal test checks should complete before merging.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 41.67% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 12 functions across 2 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: sending float16 multivectors to the sidecar as bytes when supported.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

@fm1320
fm1320 merged commit 7ee9f75 into main Sep 30, 2026
20 checks passed
@fm1320
fm1320 deleted the fix/sidecar-f16-multivectors branch September 30, 2026 08:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant