← All insights

GenAI in Cash Management: Promise or Hype?

Published: 27 November 2025

GenAI in Cash Management: Promise or Hype?

In the last months, I have spoken with boards and family offices about one question: how far can we go with generative AI in mission-critical financial operations?

A recent working paper from the Bank for International Settlements (BIS) adds important insights to this discussion. The authors tested whether a GenAI agent, based on ChatGPT’s o3 reasoning model, could simulate intraday liquidity decisions within a payment system.

The results are impressive. Without domain-specific training, the AI was able to prioritise payments, manage liquidity buffers, and balance risks in a way that resembles the logic used by experienced treasury teams. I tested similar scenarios, and the results were comparable.

So far, so good. But as someone who has led organisations through digital transformation, cybersecurity events, regulatory audits, and operational restructuring, I look at the bigger picture and ask: is AI trustworthy in real financial infrastructures and platforms?

The distance between technical capability and responsible deployment remains significant. Bridging this gap is exactly where strategic leadership is needed.

The BIS working paper highlights the potential of GenAI in cash management; that much is clear. In my view, while GenAI holds promise for this use case, public and commercial LLMs, as applied in the BIS study, may not be the most suitable approach.

The promise: AI mirrors expert decision-making

The BIS experiment shows that GenAI can replicate certain aspects of cash-management behaviour:

  • Holding back liquidity when large payments are expected
  • Prioritising payments under uncertainty
  • Balancing operational risks against liquidity costs

This proves that AI can learn and imitate expert patterns, even in specialised financial domains. Imitation is not reliability, though. Technology can impress, but execution, governance, and resilience are what truly matter.

The reality: missing governance and control

In my different roles I have seen the consequences when complex systems fail, whether due to technology gaps, unclear ownership, or weak governance.

Payment and treasury systems are not playgrounds. These platforms move trillions worldwide every day. When something goes wrong, the consequences extend far beyond an IT incident. They can trigger a systemic risk event. This is why the BIS results must be interpreted through the lens of real-world accountability, not technological excitement.

Confidentiality and data integrity

GenAI models, such as the commercial LLM used in this working paper, are not designed to process highly sensitive, real-time transaction data within regulated infrastructures. Even in a secure setup (e.g. via API, information falsification, etc.), risks remain non-trivial.

  • A single leaked log entry can violate obligations such as GDPR or FINMA regulations.
  • A misconfigured interface can send data outside the security perimeter.
  • A prompt used by an operator can expose confidential details.

With some experience, I can say with certainty: these risks are not theoretical. They are existential and real.

Non-determinism conflicts with financial governance

One of the biggest challenges is that GenAI is non-deterministic. The same input does not always produce the same output. In a research lab, this is acceptable. In operational cash management, it is a governance nightmare.

Boards and regulators expect:

  • Consistency
  • Traceability
  • Reproducibility
  • Clear accountability

A system that behaves differently each time undermines all four pillars.

Bias risks and training gaps

Models trained on generalised or synthetic data may perform well on “normal days” in standard situations. Real stress events rarely follow normal patterns, though. Any leader who has navigated a crisis knows that assumptions collapse under pressure, and crises do not wait for model retraining.

Missing explainability

Even though generative models simulate reasoning, their decision-making logic remains difficult to verify. AI “explains” decisions with plausible text, not with verifiable calculations. Unlike rule-based or deterministic algorithms, LLMs cannot reconstruct why they prioritised a particular payment or withheld liquidity. In an audit or compliance case, this would be a serious problem.

New cyberattack surfaces

Agentic AI introduces new vulnerabilities: prompt injection, model manipulation, and input tampering are real and evolving risks. If many institutions use similar models, systemic exposure follows.

Board members and executives always need to think in terms of risk management under pressure and must anticipate failure modes before they appear. In AI-assisted financial systems, that level of maturity remains far off.

The human remains the ultimate safety layer

The BIS authors highlight an important point: generative AI can support, but cannot replace, professional judgement. I would go a step further: AI must be embedded into a governance structure where humans remain the decision authority, not the passive overseer.

This requires at least:

  • Human-in-the-loop mechanisms (HITL)
  • Clear accountability chains
  • Audit-ready decision logs
  • Regulatory frameworks before deployment

AI becomes valuable when it strengthens human decision-making, not when it substitutes it.

Agentic AI: limited to an assistant role today

There seems to be a misconception in the market. Many assume that because AI agents can “reason” and “take actions”, they are ready for autonomous operations. The current generation of agentic AI remains at an assistant level, though:

  • It can prepare scenarios.
  • It can simulate decisions.
  • It can provide recommendations.

It must not autonomously release liquidity, trigger payments, or make binding decisions.

True autonomy demands:

  • Deterministic control paths
  • Certified decision models
  • Secure, sovereign data infrastructures
  • Real-time monitoring and kill-switch mechanisms

Achieving this autonomy requires much more than technological progress. It requires board-level alignment, regulatory clarity, and cultural readiness.

The path forward for boards and executives

The potential of AI in cash management is undeniable. In leadership, speed must never replace responsibility. To move forward safely, organisations should focus on:

  • Establishing robust AI governance: Data sovereignty, explainability, and oversight must be foundational.
  • Testing AI in supervised sandboxes: Controlled environments allow learning without exposing real systems to risk.
  • Developing AI literacy at board level: Understanding capabilities and limitations is now a strategic imperative.
  • Building cyber-resilient AI architectures: Confidentiality by design and resilience must be non-negotiable.

Conclusion

Generative AI opens exciting new possibilities in cash management. The BIS study shows that AI can perform surprisingly well in structured test environments. A large gap remains, however, between impressive performance in a test lab and responsible deployment in a real-life scenario. The models are powerful, but not predictable. They are “intelligent” and efficient, yet not inherently trustworthy without strong governance.

For boards and executives, one point must be clear: governance must come before automation, in any digital transformation project, not only with AI. Only then can GenAI evolve from a promising experiment into a stable element of global financial infrastructure.

Source: BIS: AI Agents for Cash Management in Payment Systems