AI Safety

Study warns agentic AI in FinTech faces a dangerous 'Verifiability Gap'

Frontier AI models can't reproduce their own financial decisions—local 3B models do.

Deep Dive

A new research paper from Henry Han on arXiv ('Governing Agentic AI in FinTech', arXiv:2608.11344) delivers a stark warning: the biggest problem with AI in financial services isn't capability, it's verifiability. Han argues that as financial institutions hand consequential decisions to agentic AI systems—AI that can decompose goals, coordinate models and tools, and act with minimal human oversight—the industry is flying blind. He defines the 'Verifiability Gap' as the shortfall between what delegated authority demands and what explainability and reproducibility remain after a decision is made. The paper analyzes nine model versions, from a 3-billion-parameter local model to commercial frontier systems.

Study 1 found that provider updates rewrite historical financial actions, and frontier models outright reject temperature, top_p and top_k controls while exposing no random seed. Under the tightest allowed settings, the local model reproduced 320 of 320 executions, while hosted models managed 319 of 320 and 959 of 960. Study 2 reveals that orchestration itself is a hidden policy layer—no execution record repeated in any config at any scale, and the frontier model reproduces its own actions more often than local models yet still loses a comparable share of differentiation. Study 3 shows deterministic credit-model versions perfectly reproduce their current action but can't recover a historical one. The conclusion: capability buys a higher starting point, not auditability. Governance must become evidence-contingent, where authority is defensible only while retained evidence backs it—an idea that extends beyond finance to any high-stakes domain.

Key Points
  • Frontier models expose no random seed and reject temperature/top_p/top_k controls, making runs unrepeatable
  • A 3B-parameter local model reproduced 320/320 executions; hosted models only 319/320 and 959/960 even under tightest settings
  • No execution record repeated across any configuration, proving orchestration is an ungoverned policy layer

Why It Matters

If AI decisions can't be reproduced, no financial institution can legally defend them—regulators need verifiability standards now.

📬 Get the top 10 AI stories daily