AAIA Question of the Day: Shadow Testing Decision Criteria

AI audit training by expert Yazan Abu Ghosh for auditors with certifications.

AAIA exam practice question — AAIA Question of the Day: Shadow Testing Decision Criteria

AAIA exam practice question: daily practice for the ISACA Advanced in AI Audit (AAIA) exam — domain: AI Operations.

Question

A payments team has run a new fraud-detection model in shadow mode for six weeks, collecting predicted fraud labels alongside production decisions but not affecting transactions. As an auditor evaluating readiness for promotion to live traffic, which control gives the most reliable evidence that the candidate model is not introducing regressions?

  • A. Compare only aggregated accuracy and loss metrics between the shadow model and production model over the test window.
  • B. Apply pre-defined statistical significance tests to outcome-level metrics (false positives, false negatives, financial loss) before promoting the model.
  • C. Route a small percentage of live transactions to the new model and allow it to block or approve to observe real-world effects before switching fully.
  • D. Require weekly manual review of a random sample of shadow predictions by domain experts and sign-off before promotion.
Show the answer and explanation

Correct answer: B. Apply pre-defined statistical significance tests to outcome-level metrics (false positives, false negatives, financial loss) before promoting the model.

Pre-defined statistical significance testing on outcome-level metrics is best because it quantifies whether observed differences are unlikely to be due to chance and focuses on business-impact measures (false positives/negatives, monetary loss). This provides objective, replicable criteria for promotion. Comparing only aggregated accuracy and loss is insufficient because those aggregates can mask changes in critical subgroups or economic impact. Routing live traffic to let the model act changes its shadow purpose and exposes customers to untested risk; it is a deployment strategy, not an audit control. Weekly manual review is useful for qualitative insight but is subjective, labor-intensive, and unlikely to detect small but important statistical regressions reliably.

Want more practice?

Prepare for the ISACA Advanced in AI Audit (AAIA) exam with AI Audit & Compliance Framework: Practical Methods & Evaluation.