feat(eval): surface Vertex rubric verdicts and explanations in eval results - #7366
Open
DalmPhilipe wants to merge 2 commits into
Open
DalmPhilipe wants to merge 2 commits into
DalmPhilipe wants to merge 2 commits into
Conversation
…esults Vertex-backed metrics (multi_turn_*_v1, safety_v1, response_evaluation_score) only read the mean score, so a failed case gave no reason. The Vertex response already carries rubric verdicts, the judge's explanation and an error message; map them into RubricScore the same way ADK-native judges do, so `adk eval --print_detailed_results` and evalset_result.json show them. Fixes google#7350
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Please ensure you have read the contribution guide before creating a pull request.
Link to Issue or Description of Change
1. Link to an existing issue (if applicable):
2. Or, if no issue exists, describe the change:
Problem:
Vertex-backed metrics (
multi_turn_task_success_v1,multi_turn_tool_use_quality_v1,multi_turn_trajectory_quality_v1,safety_v1,response_evaluation_score) go through_VertexAiEvalFacade._get_score(), which reads onlysummary_metrics[0].mean_score. The Vertex response also carries, per metric,rubric_verdicts,explanationanderror_message, but ADK discards them. A failed case gives no reason, and for the adaptive-rubricmulti_turn_*_v1metrics there is no way to tell which generated rubric failed.Solution:
Map what Vertex returns into the existing
RubricScoretype, the same way ADK-native judges do:rubric_verdictsbecomes aRubricScore:rubric_id: the rubric description, then itsrubric_id, thenrubric_{index}rationale:reasoningscore:1.0ifverdictelse0.0(Vertex omitsverdicton a failed rubric)explanationbecomesRubricScore(rubric_id="explanation", rationale=...)error_messagebecomesRubricScore(rubric_id="error", rationale=...)PerInvocationResult.rubric_scoresEvaluationResult.overall_rubric_scoressafety_v1,response_evaluation_scoremulti_turn_*_v1Behavior change worth reviewing: when a multi-turn metric gets no score but Vertex returns an error or explanation, the facade now returns
NOT_EVALUATEDwith those details, instead of an emptyEvaluationResult(). When Vertex returns neither a score nor details, the result is still empty, as before.LocalEvalServicealready allowsNOT_EVALUATEDresults.When Vertex returns no details,
rubric_scoresstaysNone, so the output does not change. No config change.Testing Plan
Unit Tests:
New tests in
tests/unittests/evaluation/test_vertex_ai_eval_facade.py:safety_v1explanation kept per invocation; error kept when unscored; no rubric scores when Vertex returns only a scoremulti_turn_*_v1metrics): verdicts on the last turn and overall, including an omittedverdict; error kept when unscoredManual End-to-End (E2E) Tests:
Ran on this branch against real Vertex (
GOOGLE_CLOUD_LOCATION=global), with no custom metrics. I used a throwaway agent that has aget_order_statustool but is told never to call it, and to ask for phone number, home address and date of birth. The evalset has 2 turns asking for the status of order A123.Output on this branch (excerpt)
In the Invocation Details table, both turns get a new
Rubric: explanationcolumn forsafety_v1:Reasoning: Violated policies: PII & Demographic Data.Same run on main (excerpt)
The Invocation Details table has no rubric or explanation columns.
What it shows:
multi_turn_tool_use_quality_v1rubrics come with their reasoning. Vertex omitsverdicton a failed rubric, and those map to0.0.safety_v1keeps the explanation on each invocation.multi_turn_task_success_v1passed because Vertex judged the agent against its own system instruction ("say you cannot look up orders"). Without the rubrics, a 1.0 here would be impossible to interpret.Checklist
Additional context
With the same mapping patched into google-adk 2.10.0,
response_evaluation_scoreturned out to fail on every call with400 INVALID_ARGUMENT(Vertex fails to parse its own judge output). Today it shows up only asN/A. That is a separate issue and is not changed here.