| Permissionscritical | Preregistered permission probes replayed through the real runtime policy path — instruction adherence, delegated-authority boundary, unsafe-tool refusal, approval escalation — with the resulting runtime decision evidence linked to the scenario. Identical probes are deduplicated by probe identity, so repetition cannot inflate the denominator. | A runtime policy decider on the evaluation path. | not_assessed “no runtime policy decider configured” |
| Groundingcritical | Source attribution and citation validity, abstention where support is absent, rejection of stale documents, escalation on conflicting sources. Scored deterministically — no model is ever the sole blocking signal. | An agent runner able to execute the agent-under-test and capture its outputs. | not_assessed “grounding cannot be measured without executing the agent-under-test” |
| Robustnesscritical | Repeat consistency, degradation under padded context, duplicate and conflicting input handling, schema conformance, behaviour under bounded concurrent load. | An agent runner able to execute the agent-under-test. | not_assessed “robustness cannot be measured without executing the agent-under-test” |
| Securitycritical | The same deterministic inline detector battery the gateway enforces with — prompt injection, jailbreak, indirect injection, prompt extraction, secret and regulated-identifier exfiltration, unsafe destination — run read-only over the execute seam, with no dispatch and no side effect. It resolves the account's effective outcome overrides fresh on every probe, so it scores the posture you actually run, not the default one. | A shared store the control plane can read the account's effective outcome overrides from. | not_assessed “no inline security inspector configured” |
| Privacycritical · composite | Isolation: a bounded, serialized, unconditionally rolled-back transaction per crown-jewel table that seeds a row as one synthetic tenant and asserts a second cannot read, update, delete, or write into it — exercising the real database isolation policies, not a re-implementation. PII: whether the agent's own outputs disclose sensitive data. | Isolation needs a database carrying the real isolation policies. PII needs an agent runner. | composite not_assessed The domain cannot pass on half its evidence. The isolation component's own counts and records are still preserved and surfaced, and the reason names the status of each half. |
| Efficiencynon-critical | Latency, provider-metered token and cost consumption, task-success-adjusted cost, bounded consumption. Only usage carrying explicit provenance is counted. | An agent runner supplying metered usage with provenance. | not_assessed Projected or unprovenanced usage is never promoted to a measurement. Any unassessable probe caps the domain at not_assessed. |