Skip to content

fix(eval): preserve live usage without extra inference calls - #7323

Open
yang0228 wants to merge 1 commit into
google:mainfrom
yang0228:codex/7321-live-eval-usage
Open

yang0228 wants to merge 1 commit into
google:mainfrom
yang0228:codex/7321-live-eval-usage

Conversation

@yang0228

Copy link
Copy Markdown
Contributor

Link to Issue or Description of Change

Closes #7321.

Live usage-only events were discarded during evaluation conversion, leaving token_usage_v1 unavailable. Following the maintainer's suggestion, retain those reports and merge them into an existing model event in the same invocation so they do not introduce an extra inference call.

Match the author and, when supplied, model version; prefer preceding events and allow usage to arrive before content. Preserve existing usage and grounding metadata, keep unmatched reports standalone, and leave source session events and final response text unchanged.

Testing Plan

  • Added regression tests: standalone usage (including empty content), usage before/after content and in the same Live message, no usage, multiple chunks/calls/invocations, and different authors/models. Existing content+usage coverage remains intact.

  • Targeted evaluation tests: 97 passed. Before the fix, 12 new cases failed because usage was dropped.

  • Full suite run with tox -p 2 -x 'testenv.commands=pytest tests/unittests -n 4' across Python 3.10–3.14 (3.10/3.11/3.13/3.14 rerun with fresh interpreters after an initial environment-creation failure):

    Python Passed Failed Skipped Xfailed Xpassed
    3.10 16,242 3 87 26 2
    3.11 16,251 3 86 26 2
    3.12 16,240 5 87 26 2
    3.13 16,242 3 87 26 2
    3.14 16,242 3 87 26 2

    The full suite is not entirely green locally: three test_load_web_page.py cases take the proxy path on this macOS host despite clearing proxy environment variables. Python 3.12 additionally reports two import-allowlist failures caused by Homebrew's sitecustomize. Reproduced the same five failures on unmodified main (044a1ec3) with the same Python 3.12 environment (5 failed, 39 passed); also reproduced the three web failures on main with Python 3.10 (3 failed, 26 passed).

    Baseline failure names
    test_load_web_page_allows_public_nat64_ip
    test_load_web_page_fetches_public_urls_by_pinning_the_resolved_ip
    test_load_web_page_tries_another_resolved_address_after_connect_error
    test_entry_point_loads_only_allowlisted_packages[agent]  # Python 3.12 only
    test_entry_point_loads_only_allowlisted_packages[runner] # Python 3.12 only
    
  • pre-commit run --files src/google/adk/evaluation/evaluation_generator.py tests/unittests/evaluation/test_evaluation_generator.py: passed.

  • Type-check comparison against 044a1ec3: no new errors in the changed module (same pre-existing StreamingMode export error; Python 3.12 override required by installed NumPy stubs).

  • uv build: passed; installed the resulting wheel with [eval] into a clean Python 3.12 environment.

Manual end-to-end evidence (local transport): Ran the issue reproducer against the installed wheel, through GeminiLlmConnection.receive() → Event → evaluation conversion → both efficiency evaluators. This exercises the affected conversion path without a remote Gemini service.

{"input_usage_events": 1, "retained_usage_events": 1, "expected_tokens": 15, "actual_tokens": 15.0, "inference_calls": 1.0, "final_response": "Hello"}
Reproduce against the built wheel
uv build
uv venv /tmp/adk-7321-smoke
uv pip install --python /tmp/adk-7321-smoke/bin/python 'dist/google_adk-2.10.0-py3-none-any.whl[eval]'
# Save the script below outside the checkout, then run it with:
/tmp/adk-7321-smoke/bin/python /tmp/live_usage_smoke.py
import asyncio
import json

from google.genai import types
from google.adk.events import Event
from google.adk.models.gemini_llm_connection import GeminiLlmConnection
from google.adk.evaluation.evaluation_generator import EvaluationGenerator
from google.adk.evaluation._efficiency_evaluators import _TokenUsageV1Evaluator, _InferenceCallCountV1Evaluator


class LocalTransport:
    session_id = "local-test"

    async def receive(self):
        yield types.LiveServerMessage(
            server_content=types.LiveServerContent(
                model_turn=types.Content(
                    role="model", parts=[types.Part(text="Hello")]
                )
            )
        )
        yield types.LiveServerMessage(
            usage_metadata=types.UsageMetadata(
                prompt_token_count=10,
                response_token_count=5,
                total_token_count=15,
            )
        )
        yield types.LiveServerMessage(
            server_content=types.LiveServerContent(turn_complete=True)
        )


async def main():
    connection = GeminiLlmConnection(
        LocalTransport(), model_version="gemini-live-test"
    )
    events = [Event(
        author="user", invocation_id="inv1",
        content=types.Content(role="user", parts=[types.Part(text="Hi")]),
    )]
    async for response in connection.receive():
        events.append(Event(
            author="agent", invocation_id="inv1",
            **response.model_dump(exclude_none=True),
        ))

    invocations = EvaluationGenerator.convert_events_to_eval_invocations(events)
    score = _TokenUsageV1Evaluator().evaluate_invocations(invocations).overall_score
    call_count = _InferenceCallCountV1Evaluator().evaluate_invocations(invocations).overall_score
    answer = invocations[0].final_response.parts[0].text
    print(json.dumps({
        "input_usage_events": sum(e.usage_metadata is not None for e in events),
        "retained_usage_events": sum(
            e.usage_metadata is not None
            for i in invocations
            for e in i.intermediate_data.invocation_events
        ),
        "expected_tokens": 15,
        "actual_tokens": score,
        "inference_calls": call_count,
        "final_response": answer,
    }))
    assert score == 15, f"Expected 15 tokens, got {score!r}"
    assert call_count == 1
    assert answer == "Hello"


asyncio.run(main())

Checklist

  • Read CONTRIBUTING and incorporated the maintainer's feedback.
  • Performed a self-review and added regression coverage.
  • Verified the installed wheel through the affected Live-to-evaluation path.
  • No public API or documentation changes.

Additional context

This preserves the existing inference count for multi-chunk Live output; it does not change the evaluator's existing per-model-event counting behavior.

Merge standalone Live usage into compatible model events during evaluation
conversion, retaining unmatched reports without mutating session events.

Fixes google#7321
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Live evaluation drops standalone usage metadata, so token_usage_v1 reports n/a

2 participants