Equivalent cases should not produce inexplicably different decisions
Two compliance cases can look different and still be materially equivalent. A payment narrative rephrased, the same facts arriving in a different order, an analyst’s summary written more tersely: none of that changes what the institution’s policy should conclude. They should not receive different treatment simply because one was phrased differently.
At the same time, two cases that look similar should not receive the same treatment when a material risk factor has changed. A counterparty newly connected to a flagged network is not the same case as its lookalike from last quarter, however similar the paperwork appears.
That creates two important tests for any AI-assisted compliance system, and most evaluation practice today applies neither.
Invariance and sensitivity
Invariance: immaterial changes should not change the outcome. If the same underlying facts, expressed differently, produce different dispositions, the system is responding to phrasing rather than to policy, and every disposition it has ever produced becomes harder to defend.
Sensitivity: material changes should change the outcome, or at least the analysis. A system that treats a materially different case identically because it pattern-matches the surface is failing in the more dangerous direction: it is missing risk while appearing consistent.
Hallucination testing is not decision testing
Most evaluation of AI in compliance today asks whether the model invented a fact, cited a nonexistent source, or drifted from its prompt. Those checks matter, but they operate at the level of the model’s output. The institution’s exposure sits one level higher: at the decision.
The deeper question is whether the institution’s policy, evidence requirements and authority model are being applied consistently across cases. A model can be perfectly factual and still produce dispositions that vary with wording, evidence order, model version or time of day. Every one of those inconsistencies is a finding waiting to be written, by an internal auditor, a model-risk reviewer or a supervisor.
What decision-level testing looks like
Testing the decision rather than the output means constructing case families deliberately: equivalent cases with different wording; cases with irrelevant details added; cases where exactly one material factor changes; the same case executed repeatedly, and across model versions; cases with missing evidence; cases with conflicting evidence; cases evaluated under different policy versions; and cases that must route to a human regardless of what any model concludes.
For each family, the expected behavior is defined in advance, not by the model, but by the institution’s policy. The system passes when equivalence produces consistency, materiality produces change, missing evidence produces escalation rather than confident completion, and human-authority conditions produce a human every time.
Consistency is an architecture property
A system cannot be tested this way as an afterthought. It has to be built for it: policy expressed as something executable rather than prompt prose; evidence retrieval recorded, so what the decision saw is knowable; model and policy versions pinned to each decision; and the decision itself retained as a complete record: inputs, reasoning, both sides of the argument where the system argues both sides, and the human signer where one is required.
A consequential decision should be testable, challengeable and reconstructable. Not simply generated.
Fenori engineers compliance decision systems designed to be tested this way: inside the institution’s environment, under its control.
See Decision Infrastructure