Skip to content

feat(eval): expose session state to custom evaluators - #7363

Draft
RainieLLM wants to merge 1 commit into
google:mainfrom
RainieLLM:codex/evaluator-session-state-4532
Draft

RainieLLM wants to merge 1 commit into
google:mainfrom
RainieLLM:codex/evaluator-session-state-4532

Conversation

@RainieLLM

@RainieLLM RainieLLM commented Sep 30, 2026 •

Copy link
Copy Markdown

Link to Issue or Description of Change

Related: #4532

Custom evaluators can inspect the conversation, but they cannot read session state. This adds EvaluationContext and an evaluate_with_context method for metrics that need it. The default method calls evaluate_invocations with its original arguments, so existing sync and async evaluators need no changes.

Local inference saves the actual state before the first turn and after the run ends. Each metric receives its own copy, plus the expected final state from the eval case. Later session updates or deletion do not change these snapshots. Old inference results without snapshots keep None for the actual state fields.

Testing Plan

Unit Tests:

  • I have added or updated unit tests for my change.
  • All unit tests pass locally.

pytest tests/unittests/evaluation -q: 960 passed on Python 3.13.15 at the final commit. The 22 new cases cover old and new sync/async evaluators, regular and live inference, reused sessions, state isolation, and saved results after a session update or deletion.

The full suite ran through tox on Python 3.10–3.14, with two pytest workers per version. A 60-second per-test timeout let the suite continue past stalled tests.

Python Passed Failed
3.10 16,608 5
3.11 16,619 3
3.12 16,611 2
3.13 16,610 3
3.14 16,607 6

The evaluation tests pass on all five versions: 960 per version. The full suite does not pass. Two LiveKit tests exceed the timeout, and the import allowlist check fails. All three failures also occur on unmodified main (9633ab9f) in the same Python 3.14 environment. Some CLI deploy tests fail in the full 3.10 and 3.14 runs, but all 68 tests in that file pass in the isolated baseline check. Their cause is not confirmed.

The checks for all changed files pass. The package build and a clean wheel import also pass. Mypy reports the same seven errors on this change and the original base revision, with no new errors.

Manual End-to-End (E2E) Tests:

A local BaseAgent adds a book to a cart through the real Runner. The test saves the inference result as JSON, deletes the session, then evaluates the saved result with a custom state metric. It needs no model API or credentials.

The metric receives an empty cart as the initial state and a cart with one book as the final state. The final state matches the expected state. The result is PASSED with a score of 1.0.

The check uses InMemorySessionService, LocalEvalService, and an isolated metric registry. The same setup and public inference/evaluation path are in test_evaluate_uses_snapshots_after_session_change. Run it from the repo after the normal development setup:

pytest tests/unittests/evaluation/test_local_eval_service.py -q -k uses_snapshots_after_session_change

Console output from the manual run:

Metric state: {"initial_session_state":{"cart":{"items":[]}},"final_session_state":{"cart":{"items":["book"]}},"expected_final_session_state":{"cart":{"items":["book"]}}}
{"status": "PASSED", "score": 1.0}

Checklist

  • I have read the CONTRIBUTING.md document.
  • I have performed a self-review of my own code.
  • I have commented my code, particularly in hard-to-understand areas.
  • I have added tests that prove my fix is effective or that my feature works.
  • New and existing unit tests pass locally with my changes.
  • I have manually tested my changes end-to-end.
  • Any dependent changes have been merged and published in downstream modules.

Additional context

Companion docs: google/adk-docs#2289. The strict docs build and local page check pass.

This is a draft for API feedback. Does the separate context method fit the evaluator interface, and are the snapshot boundaries appropriate? The full suite failures are listed above, so this is not ready to land yet.

Save actual session state before and after inference, and pass separate copies to each metric through an optional context method. Keep the original evaluator interface for existing sync and async metrics.

Related: google#4532
@google-cla

google-cla Bot commented Sep 30, 2026

Copy link
Copy Markdown

Thanks for your pull request! It looks like this may be your first contribution to a Google open source project. Before we can look at your pull request, you'll need to sign a Contributor License Agreement (CLA).

View this failed invocation of the CLA check for more information.

For the most up to date status, view the checks section at the bottom of the pull request.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants