Skip to content

bug: hallucinations_v1 judging each claim #230

Description

@AdminTurnedDevOps

Description

Trace replay always stores the captured trajectory as ADK IntermediateData: tool_uses and tool_responses. hallucinations_v1 never reads that shape. ADK's HallucinationsV1Evaluator._get_steps_to_evaluate loads evidence only when intermediate_data is an InvocationEvents object:

all_events = []
  if isinstance(actual.intermediate_data, InvocationEvents):
   all_events = actual.intermediate_data.invocation_events or []

Anything else leaves all_events empty. The judge still scores the final response (intermediate natural-language responses are off by default), but the context it sees is only developer instructions, the user prompt, and tool definitions. Tool call arguments and tool outputs never appear, so a claim that is true only because it is in the tool output is marked unsupported. A mean of those labels is 0.00.

agentevals already has the converter for this, _to_invocation_events in builtin_metrics.py. It pairs each tool_uses entry with its tool_responses entry (by id, otherwise by position when both lack an id) and emits InvocationEvents. It runs only for the three multi-turn Vertex metrics:

_METRICS_NEEDING_INVOCATION_EVENTS = {
    "multi_turn_task_success_v1",
    "multi_turn_trajectory_quality_v1",
    "multi_turn_tool_use_quality_v1",
}

Putting the tool output on the golden eval set does not change the score. This metric does not read expected invocations. The evidence has to be on the actual invocation, in invocation_events, and the trace-replay path does not put it there. The code change that would fix it is adding hallucinations_v1 to _METRICS_NEEDING_INVOCATION_EVENTS so the existing adapter runs before the judge. That change is not in the tree.

Expected behavior

hallucinations_v1 should score the final response against the captured tool output, and a response whose every claim is in that output should score 1.0.

For each invocation the judge does two steps:

  1. It splits the final response into sentences. Intermediate agent text is left out unless evaluate_intermediate_nl_responses is turned on.
  2. It labels each sentence against a context that includes the developer instructions, the user prompt, the tool definitions, and the invocation's tool calls and tool outputs.
┌───────────────────────────────────────────────────────┬───────────┐
│ Label                                                 │ Counts as │
├───────────────────────────────────────────────────────┼───────────┤
│ supported                                             │ 1         │
├───────────────────────────────────────────────────────┼───────────┤
│ not_applicable (greeting, plan, question, disclaimer) │ 1         │
├───────────────────────────────────────────────────────┼───────────┤
│ unsupported                                           │ 0         │
├───────────────────────────────────────────────────────┼───────────┤
│ contradictory                                         │ 0         │
├───────────────────────────────────────────────────────┼───────────┤
│ disputed                                              │ 0         │
└───────────────────────────────────────────────────────┴───────────┘

The invocation score is the fraction of sentences that count as 1. The run score is the mean of those invocation scores. The default pass threshold is 0.5.

A claim that the tool output fully entails is supported. If every factual sentence in the final response is entailed that way, and any remaining sentences are not_applicable, the score is 1.0 and the run passes. A sentence that adds a fact the tool output does not contain is unsupported and pulls the score down. A sentence the tool output falsifies is contradictory.

Steps to reproduce

hallucinations_v1 scored 0.00 even though every response claim exists in the captured tool output. It appears trace replay stores evidence as tool_uses/tool_responses while this evaluator expects InvocationEvents.

How are you using agentevals?

Other

Config dump (ZIP from the web UI)

No response

Version information

Running in kagent enterprise v0.5.8.

Eval config (if applicable)

Trace and eval set files

No response

Relevant logs or error output


Additional context

No response

Human confirmation

  • I am a human (not a bot, agent, or AI) filing this issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions