Conversation
Merge standalone Live usage into compatible model events during evaluation conversion, retaining unmatched reports without mutating session events. Fixes google#7321
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Link to Issue or Description of Change
Closes #7321.
Live usage-only events were discarded during evaluation conversion, leaving
token_usage_v1unavailable. Following the maintainer's suggestion, retain those reports and merge them into an existing model event in the same invocation so they do not introduce an extra inference call.Match the author and, when supplied, model version; prefer preceding events and allow usage to arrive before content. Preserve existing usage and grounding metadata, keep unmatched reports standalone, and leave source session events and final response text unchanged.
Testing Plan
Added regression tests: standalone usage (including empty content), usage before/after content and in the same Live message, no usage, multiple chunks/calls/invocations, and different authors/models. Existing content+usage coverage remains intact.
Targeted evaluation tests: 97 passed. Before the fix, 12 new cases failed because usage was dropped.
Full suite run with
tox -p 2 -x 'testenv.commands=pytest tests/unittests -n 4'across Python 3.10–3.14 (3.10/3.11/3.13/3.14 rerun with fresh interpreters after an initial environment-creation failure):The full suite is not entirely green locally: three
test_load_web_page.pycases take the proxy path on this macOS host despite clearing proxy environment variables. Python 3.12 additionally reports two import-allowlist failures caused by Homebrew'ssitecustomize. Reproduced the same five failures on unmodified main (044a1ec3) with the same Python 3.12 environment (5 failed, 39 passed); also reproduced the three web failures on main with Python 3.10 (3 failed, 26 passed).Baseline failure names
pre-commit run --files src/google/adk/evaluation/evaluation_generator.py tests/unittests/evaluation/test_evaluation_generator.py: passed.Type-check comparison against
044a1ec3: no new errors in the changed module (same pre-existingStreamingModeexport error; Python 3.12 override required by installed NumPy stubs).uv build: passed; installed the resulting wheel with[eval]into a clean Python 3.12 environment.Manual end-to-end evidence (local transport): Ran the issue reproducer against the installed wheel, through
GeminiLlmConnection.receive()→Event→ evaluation conversion → both efficiency evaluators. This exercises the affected conversion path without a remote Gemini service.{"input_usage_events": 1, "retained_usage_events": 1, "expected_tokens": 15, "actual_tokens": 15.0, "inference_calls": 1.0, "final_response": "Hello"}Reproduce against the built wheel
Checklist
Additional context
This preserves the existing inference count for multi-chunk Live output; it does not change the evaluator's existing per-model-event counting behavior.