Sphere Partners
Senior executive gesturing during a strategic discussion in a corporate boardroom, representing EU AI Act compliance governance for banks deploying LLMs

The EU AI Act Compliance Checklist for Banks Deploying LLMs

The EU AI Act is now enforceable, and most bank LLM use cases fall into its high-risk category. Here's the practical compliance checklist — and what to ask a vendor when you can't answer it yourself.

6 min read
In this article

The EU AI Act is now a live compliance obligation, not a future one — and for a bank deploying a large language model, it is no longer optional homework. Prohibited practices became enforceable in February 2025, obligations for general-purpose AI models followed in August 2025, and the high-risk system requirements that cover most AI used in credit, fraud, and compliance decisioning phase in through August 2026 and into 2027. A private LLM platform touching any of those use cases needs a compliance answer today, not a roadmap.

This isn't a substitute for legal advice — the Act is long, the guidance is still being finalized by the EU AI Office, and every institution's exposure depends on its specific use cases and jurisdictions. What follows is a practical checklist: the questions a bank's compliance and procurement teams should be able to answer before signing off on an LLM deployment, and what to ask a vendor when you can't answer them yourself.

Why This Applies Even Outside the EU

The EU AI Act applies extraterritorially: it covers any AI system whose output is used within the EU, regardless of where the provider or deployer is based. A US or GCC bank with EU clients, EU subsidiaries, or EU-domiciled counterparties can be in scope even without a single employee in the bloc. That's part of why the Act shows up as a named requirement in RFPs issued well outside Europe — it has become a de facto baseline for AI governance maturity, the same way GDPR did for data privacy.

In short

Most bank-deployed LLM use cases — credit scoring inputs, fraud detection, creditworthiness assessment, internal risk scoring — fall into the Act's "high-risk" category, which carries the heaviest documentation, oversight, and testing obligations. A general-purpose internal chat or document-retrieval assistant with no role in an automated decision typically does not, but the line depends on exactly how the output is used downstream.

The Compliance Checklist

Eight areas a bank's compliance and IT teams should be able to speak to before an LLM platform goes live:

Requirement areaWhat it means in practiceWhat to ask a vendor
Risk classificationDetermine whether each specific LLM use case is minimal-risk, limited-risk, or high-risk under the Act's Annex III categories.Can you help us document a use-case-by-use-case risk classification, not just a platform-level statement?
Risk management systemHigh-risk systems require a documented, continuous risk management process covering the system's full lifecycle, not a one-time assessment.What does ongoing risk monitoring look like after deployment, and who owns it — us or you?
Data governanceTraining and retrieval data must be relevant, representative, and checked for errors and bias where relevant to the use case.What data governance controls apply to the retrieval corpus, not just the base model's training data?
Technical documentationHigh-risk systems require documentation sufficient for a regulator to assess compliance — architecture, intended purpose, and known limitations.Will you provide technical documentation in the format our regulator expects, and keep it current as the system changes?
Human oversightHigh-risk systems must be designed so a human can meaningfully understand, monitor, and override the system's output.Can a reviewer see why the model produced a given output, in plain language, before acting on it?
Accuracy and robustness testingHigh-risk systems must be tested for accuracy, robustness, and resilience against errors and adversarial manipulation before and during deployment.What adversarial and robustness testing has this specific deployment undergone, not just the base model?
Logging and traceabilityHigh-risk systems must automatically log events across their lifecycle to support post-hoc audit and incident investigation.Is every model call — including retrieval and tool use — captured in one audit trail we can hand to an examiner?
Transparency to end usersCustomers and employees interacting with certain AI systems must be informed they are doing so, in clear and understandable terms.Does the platform support the disclosure language our jurisdiction requires, built into the interface itself?

Why Private Deployment Makes This Checklist Easier, Not Harder

It's tempting to assume more regulation means avoiding AI deployment altogether, or defaulting to whatever a hosted model provider offers as a compliance add-on. In practice, a private or on-premise deployment tends to make several of these requirements easier to satisfy, not harder — because the bank controls the full stack rather than inheriting a vendor's shared, general-purpose configuration.

Logging and traceability are native to the deployment rather than dependent on what a third-party API exposes. Data governance applies to a retrieval corpus the bank actually curates. Human oversight can be built into the specific workflow rather than bolted onto a general-purpose chat interface. None of that makes the compliance work optional — it just means the architecture decision and the compliance decision are the same decision, which is why EU AI Act readiness increasingly shows up as a named requirement alongside MCP support and inference performance in the same RFPs.

faq

Frequently asked questions

It can. The Act applies to any AI system whose output is used within the EU, regardless of where the provider or deployer is based — so a non-EU bank with EU clients, subsidiaries, or counterparties can still be in scope for those specific use cases.

No. Risk classification is use-case specific. An LLM used for credit decisioning, fraud detection, or creditworthiness assessment is likely to be high-risk; an internal document-search or drafting assistant with no role in an automated decision about a person typically is not. The same platform can host both classifications for different use cases.

Prohibited AI practices became enforceable in February 2025. Obligations for general-purpose AI models followed in August 2025. High-risk system obligations phase in on a staggered timeline through August 2026 and into 2027 depending on the specific Annex III category. Institutions should confirm current deadlines with legal counsel, as implementing guidance from the EU AI Office is still being finalized.

No single deployment model automatically satisfies the Act — compliance depends on how the system is configured, documented, tested, and monitored, not just where it runs. A private deployment does make several requirements more directly achievable, since the bank controls logging, data governance, and oversight design rather than inheriting a shared vendor configuration.

Responsibility is typically shared and defined by role: the Act distinguishes between "providers" (who build the system) and "deployers" (who put it into use), with different obligations for each. A bank deploying a vendor's platform for a high-risk use case retains deployer obligations — like human oversight and monitoring — regardless of what the vendor contributes.

Mapping EU AI Act requirements against a specific LLM deployment plan? Talk to a Sphere AI Engineer.

We'd love to hear from you!

Please provide your contact details, and our team will get back to you promptly.