Process Deception Detection Scoring Module
Scores AMC-7.1 to AMC-7.12: whether an AI system tests for and resists alignment faking, sandbagging, scheming, and goal-directed deception.
Part of the AI Safety Research Alignment evaluation lane.
Evidence artifacts found
Specific gaps and recommendations
Maturity level L0-L5
Per-question scores (0-1) keyed by question ID
Overall score 0-100
Process Deception Detection Scoring Module
Scores AMC-7.1 to AMC-7.12: whether an AI system tests for and resists alignment faking, sandbagging, scheming, and goal-directed deception.
Part of the AI Safety Research Alignment evaluation lane.