Describe the Bug
For a single completed Gemini Live model response, inference_call_count_v1 changes with the number of text chunks delivered by the transport. The same final answer produces counts of 1, 2, or 3 when split into 1, 2, or 3 chunks. Transport chunking should not be interpreted as additional model calls or reasoning steps.
This reproducer emits no usage metadata, so it isolates the chunk-counting problem from the standalone usage loss in #7321 / PR #7323.
Steps to Reproduce
-
Check out the tested main commit and install the evaluation dependencies:
git clone https://github.com/google/adk-python.git
cd adk-python
git checkout d312c0eca3707dd162b21bbe00a2325758ea3d6a
uv sync --extra eval
-
Save the minimal reproduction below as repro_live_inference_count.py.
-
Run it against the checkout:
PYTHONPATH="$PWD/src" uv run --extra eval python repro_live_inference_count.py
-
Observe counts of [1.0, 2.0, 3.0] and the failed assertion expecting [1.0, 1.0, 1.0]. All three cases contain one turn_complete and produce the same final response.
Expected Behavior
One completed model response should contribute one inference call regardless of how its text is chunked. Genuine additional model calls, including multiple calls or sub-agents within the same invocation, should still be counted separately.
Observed Behavior / Logs
{"text_chunks": 1, "completed_model_turns": 1, "final_response": "Hello world", "inference_call_count_v1": 1.0}
{"text_chunks": 2, "completed_model_turns": 1, "final_response": "Hello world", "inference_call_count_v1": 2.0}
{"text_chunks": 3, "completed_model_turns": 1, "final_response": "Hello world", "inference_call_count_v1": 3.0}
AssertionError: ('Chunk boundaries should not change the number of model calls', [1.0, 2.0, 3.0])
Environment Details
- ADK: source version 2.10.0, main commit
d312c0ec, verified on 2026-09-30.
- OS: macOS (Darwin 25.6.0).
- Python: 3.12.13.
google-genai: 2.24.0.
- The local reproduction imported ADK from an unmodified snapshot of that commit; no source changes were applied.
Model Information
- LiteLLM: No.
- Model: N/A — deterministic local Live transport. The
gemini-live-test model-version label is supplied to the real GeminiLlmConnection adapter. No remote model, API key, network model request, or tool call is required.
Regression
Unknown; not claiming a regression from a particular release. Reproduced on current main.
Additional Context / Source Trace
The affected path is GeminiLlmConnection.receive() → Event → EvaluationGenerator.convert_events_to_eval_invocations() → _InferenceCallCountV1Evaluator.
PR #7323 preserves the existing multi-chunk inference count while fixing standalone token usage. This issue concerns identifying logical calls across chunks, so it needs a separate focused fix. Deduplicating all events by invocation ID or model version would undercount genuine repeated calls and should be avoided.
Minimal Reproduction Code
import asyncio
import json
from google.genai import types
from google.adk.events.event import Event
from google.adk.models.gemini_llm_connection import GeminiLlmConnection
from google.adk.evaluation.evaluation_generator import EvaluationGenerator
from google.adk.evaluation._efficiency_evaluators import _InferenceCallCountV1Evaluator
class FakeTransport:
session_id = 'local-test'
def __init__(self, chunks):
self.chunks = chunks
async def receive(self):
for chunk in self.chunks:
yield types.LiveServerMessage(
server_content=types.LiveServerContent(
model_turn=types.Content(
role='model', parts=[types.Part(text=chunk)]
)
)
)
yield types.LiveServerMessage(
server_content=types.LiveServerContent(turn_complete=True)
)
async def run(chunks):
connection = GeminiLlmConnection(
FakeTransport(chunks), model_version='gemini-live-test'
)
events = [Event(
author='user', invocation_id='inv1',
content=types.Content(role='user', parts=[types.Part(text='Hi')]),
)]
async for response in connection.receive():
events.append(Event(
author='agent', invocation_id='inv1',
**response.model_dump(exclude_none=True),
))
invocations = EvaluationGenerator.convert_events_to_eval_invocations(events)
call_count = _InferenceCallCountV1Evaluator().evaluate_invocations(
invocations
).overall_score
answer = invocations[0].final_response.parts[0].text
print(json.dumps({
"text_chunks": len(chunks),
"completed_model_turns": 1,
"final_response": answer,
"inference_call_count_v1": call_count,
}))
assert answer == "Hello world"
return call_count
async def main():
scores = [await run(chunks) for chunks in (
['Hello world'], ['Hello ', 'world'], ['Hel', 'lo ', 'world'],
)]
assert scores == [1.0, 1.0, 1.0], (
"Chunk boundaries should not change the number of model calls", scores
)
asyncio.run(main())
Frequency
Always (100% in the deterministic local reproduction; repeated with the same output).
Screenshots / Video
N/A — console reproduction.
Contribution / Assignment Request
I would like to work on this issue. Is anyone already addressing it? Following CONTRIBUTING's request to ask before contributing, could a maintainer confirm the intended counting semantics and assign this issue to @yang0228 if this contribution is welcome?
I propose a focused fix in the shared conversion/evaluation path, with regression coverage for:
- Chunk-invariant counts for a single model response (1, 2, and 3 text chunks).
- Separate genuine model calls and sub-agents within one invocation, plus unchanged non-Live evaluation behavior.
- Usage with content, standalone usage, missing usage, and preserved final text/tool events.
Before a PR, I will follow CONTRIBUTING's unit-test, Python-version matrix, pre-commit, wheel, and reproducible end-to-end evidence requirements.
Describe the Bug
For a single completed Gemini Live model response,
inference_call_count_v1changes with the number of text chunks delivered by the transport. The same final answer produces counts of 1, 2, or 3 when split into 1, 2, or 3 chunks. Transport chunking should not be interpreted as additional model calls or reasoning steps.This reproducer emits no usage metadata, so it isolates the chunk-counting problem from the standalone usage loss in #7321 / PR #7323.
Steps to Reproduce
Check out the tested main commit and install the evaluation dependencies:
Save the minimal reproduction below as
repro_live_inference_count.py.Run it against the checkout:
Observe counts of
[1.0, 2.0, 3.0]and the failed assertion expecting[1.0, 1.0, 1.0]. All three cases contain oneturn_completeand produce the same final response.Expected Behavior
One completed model response should contribute one inference call regardless of how its text is chunked. Genuine additional model calls, including multiple calls or sub-agents within the same invocation, should still be counted separately.
Observed Behavior / Logs
{"text_chunks": 1, "completed_model_turns": 1, "final_response": "Hello world", "inference_call_count_v1": 1.0} {"text_chunks": 2, "completed_model_turns": 1, "final_response": "Hello world", "inference_call_count_v1": 2.0} {"text_chunks": 3, "completed_model_turns": 1, "final_response": "Hello world", "inference_call_count_v1": 3.0}Environment Details
d312c0ec, verified on 2026-09-30.google-genai: 2.24.0.Model Information
gemini-live-testmodel-version label is supplied to the realGeminiLlmConnectionadapter. No remote model, API key, network model request, or tool call is required.Regression
Unknown; not claiming a regression from a particular release. Reproduced on current main.
Additional Context / Source Trace
The affected path is
GeminiLlmConnection.receive()→Event→EvaluationGenerator.convert_events_to_eval_invocations()→_InferenceCallCountV1Evaluator.model_version, including partial chunks.model_version._is_model_call_event()treats anymodel_versionas a model call; the evaluator counts all such events.PR #7323 preserves the existing multi-chunk inference count while fixing standalone token usage. This issue concerns identifying logical calls across chunks, so it needs a separate focused fix. Deduplicating all events by invocation ID or model version would undercount genuine repeated calls and should be avoided.
Minimal Reproduction Code
Frequency
Always (100% in the deterministic local reproduction; repeated with the same output).
Screenshots / Video
N/A — console reproduction.
Contribution / Assignment Request
I would like to work on this issue. Is anyone already addressing it? Following CONTRIBUTING's request to ask before contributing, could a maintainer confirm the intended counting semantics and assign this issue to @yang0228 if this contribution is welcome?
I propose a focused fix in the shared conversion/evaluation path, with regression coverage for:
Before a PR, I will follow CONTRIBUTING's unit-test, Python-version matrix, pre-commit, wheel, and reproducible end-to-end evidence requirements.