An AI evaluation is only useful if it is auditable
A score does not prove control. Boards, security teams and regulators need to reconstruct what was tested, against which version, with which traces, which judge, which evidence and which residual risk decision.