Agent-Initiated Pause Quality — Measures when agents stop to ask
Based on Anthropic's autonomy research: Claude Code pauses for
clarification more often than humans interrupt it. Agents that know
when to ask score higher on maturity.
Dimensions:
Pause frequency (too many = indecisive, too few = reckless)
Pause relevance (did the pause address a genuine ambiguity?)
Pause timing (early enough to prevent wasted work?)
Resolution quality (did the human response unblock effectively?)
Agent-Initiated Pause Quality — Measures when agents stop to ask
Based on Anthropic's autonomy research: Claude Code pauses for clarification more often than humans interrupt it. Agents that know when to ask score higher on maturity.
Dimensions: