Interpretability Scoring — Can you understand why the agent decided?
Inspired by Google DeepMind's Gemma Scope: interpretability tools that
trace decisions through model internals. AMC can't access model weights,
but CAN score the observable interpretability surface:
Decision explanations: Does the agent explain its reasoning?
Chain-of-thought faithfulness: Do explanations match actions?
Confidence calibration: Are confidence signals accurate?
Attribution quality: Can decisions be traced to inputs?
Refusal transparency: When the agent refuses, does it explain why?
Interpretability Scoring — Can you understand why the agent decided?
Inspired by Google DeepMind's Gemma Scope: interpretability tools that trace decisions through model internals. AMC can't access model weights, but CAN score the observable interpretability surface: