The full judge prompt that was/would be sent
Whether this was evaluated by an actual LLM or simulated
Model used for judgment
Raw score on the rubric scale
The reasoning provided by the judge
Which rubric level was matched
Normalized score (0-1)
The full judge prompt that was/would be sent