On 2 September 2026 the Financial Conduct Authority (FCA) published “Frontier AI and cyber resilience”, a collection of what firms have told it about using the most advanced AI models in their own cyber defence. The paper sets no new rules. It is still the most interesting document to appear on the subject this year, because it is the only one that turns the camera around. The attacker is not the subject. The firm picks up the same tool, and discovers that what it runs into is itself.

Everything else follows from the first of the five observations: “Vulnerability discovery is accelerating faster than firms' ability to respond.” Finding faster without fixing faster buys no security. It buys a backlog. And under European law that backlog is not an operational irritation. It is a register the firm is obliged to keep.

In brief

What: FCA multi-firm review on frontier AI and cyber resilience, published 2 September 2026

Legal status: explicitly not a new rulebook. It summarises observations that firms reported to the regulator themselves

Audience: according to the FCA, smaller and mid-sized firms in particular, with operational resilience, technology, risk and cyber leaders named

Central finding: the ability to find is outgrowing the ability to fix. Governance, validation capacity and change processes become the constraint

What binds European firms: not the FCA paper, but Article 10 of Commission Delegated Regulation (EU) 2024/1774, which already contains the chain of scan, record and deadline

Three papers in four months, three assurances that nothing new is being said

The background is short. On 15 May 2026 the Bank of England, the FCA and HM Treasury stated jointly that the cyber capabilities of current frontier models already exceed what a skilled practitioner can achieve, and do so “at a significantly higher speed, greater scale, and lower cost”. June brought guidance from the Cross Market Operational Resilience Group (CMORG), a body of British financial institutions and public authorities. Now the review.

What the three documents share is worth noting. The May statement says it is “not intended to introduce new expectations”. The CMORG guidance stresses that its use is voluntary and that it does “not constitute regulatory rules or supervisory expectations”. The September review puts it most plainly: “It does not introduce new rules, guidance or regulatory expectations.” Three times in four months, the same family of institutions insists it is saying nothing binding, while the pressure to act keeps rising. That is not a contradiction. It is how a regulator behaves when it knows its existing rulebook already covers the case and intends to leave it there.

Which is precisely the position. For European firms the decisive provision has been on the books for some time. It is rarely read in this context.

The deadlines a firm sets for itself

Vulnerability management in European financial firms is not spelled out in the text of the Digital Operational Resilience Act (DORA) itself, but in the regulatory technical standards (RTS) beneath it, Commission Delegated Regulation (EU) 2024/1774. The distinction matters and is not pedantry: saying “DORA requires a vulnerability register” cites one level too high. The requirement sits in Article 10 of the delegated regulation, where three duties lock together.

First, frequency. For ICT assets supporting critical or important functions, automated vulnerability assessments and scans must run at least weekly. That duty holds regardless of how good the scanning tool happens to be.

Second, the record. Article 10(2)(h) requires, in the words of the German text, a record of all detected vulnerabilities affecting ICT systems together with the monitoring of their remediation. Not the ones already fixed. All detected ones, with their remediation tracked.

Third, the deadline. There is no EU-wide patching deadline, and anyone citing one has invented it. What is mandatory is something else, namely that deadlines exist: Article 10(4)(d) requires firms to set deadlines for installing software and hardware patches and updates, and to define escalation procedures for cases where those deadlines cannot be met. The firm sets the clock itself, calibrated to the criticality of the asset. But it must set one, and it must have a defined route for what happens when the clock runs out.

Frontier AI changes none of these three duties. It multiplies the number of cases in which they bite. Christian Schablitzki, the agentic banker

That is the whole mechanism, and it stays unremarkable right up to the moment the input volume grows tenfold. The machinery does not change: scans run, findings are recorded, each one starts a self-imposed clock, and an escalation route opens when that clock expires. A firm that raises its detection capability without raising its remediation capacity is therefore not simply creating more work. It is creating entries in a register it is required to maintain, and a rising number of escalations it triggered itself.

The FCA describes the same dynamic without the legal citation, but with the practical ratios. Even where a substantial share of model output is discounted on expert review, the review notes, the remaining volume of genuine vulnerabilities can “still put considerable pressure on remediation teams, engineering resources, and change management processes”. The regulator even names the pinch points: validation capacity, engineering resources, patch testing, emergency change controls, evidence of closure, and the ability to keep important business services running while remediation is accelerated.

That last item deserves attention, because it is the awkward one. Fast patching is itself a source of operational risk. Push emergency changes into production under time pressure often enough and the firm raises the odds of causing an outage no attacker needed to arrange. The question is therefore not how fast a firm can patch, but how fast it can patch without switching itself off.

The legal definition sits in Article 3 of DORA

The bridge into European terminology holds. Where the FCA speaks of “important business services”, European law has a legal definition in Article 3(22) DORA: a critical or important function is one whose failure would materially impair the financial performance of a financial entity, or the soundness or continuity of its services and activities. That definition carries the weekly scanning duty, and it carries the prioritisation question that comes next.

Translating the FCA's questions therefore does not require building anything new. It requires applying them to the functions a firm has already classified as critical or important.

Why the severity ranking gives way

The second substantive finding is technical, and it bears directly on a method almost every firm relies on. Frontier models, participating firms report, can combine several individually low-rated weaknesses into what the review calls “vulnerability chaining”, creating “alternative routes to compromise”. They find more than individual issues. They find the relationships between them.

This undercuts a load-bearing assumption of current practice. Ranking by severity presumes that the risk of a vulnerability can be measured in isolation. That presumption fails once three low-rated findings combine into a route to a production system. The FCA accordingly reports that vulnerability management decisions are increasingly informed by the disruption that would follow if an attack path were exploited, “rather than the ratings of vulnerabilities in isolation”. The factors it names are exploitability, business service impact, prerequisites to exploit, risk-reduction controls, and dependency on the vulnerable system.

The CMORG guidance from June confirms that this is more than the impression of a few firms. It puts a question to firms aimed at the same point: whether prioritisation rests on something other “than inherent severity alone”. Two documents, drafted separately, reaching the same diagnosis.

In practice this imposes an uncomfortable sequence. Before a firm can change its prioritisation, it needs something many do not hold at sufficient quality: a reliable map of which systems carry which critical functions, and what those systems depend on. Without that map, “we prioritise by attack path” is an intention rather than a method.

In trading, an escalation route that fires constantly wears out

The mechanism of a self-set threshold backed by mandatory escalation is not new to financial services. It simply lives somewhere else. Trading limits follow the same construction: a firm sets its own position, loss and risk limits within the supervisory frame, and a breach triggers a defined procedure, namely notification to risk control, a decision to reduce or approve, and a record of what happened. No regulator prescribes the level of an individual limit. What is prescribed is that a limit exists and that breaching it has consequences.

Anyone who has run such systems for any length of time also knows how they wear. An escalation route walked once a quarter forces a genuine decision. One that fires every week becomes routine: it gets signed off rather than examined. The threshold loses its effect not because it was set wrongly, but because it is reached too often.

The difference to vulnerability management is precisely where this turns uncomfortable. In trading, the number of breaches depends on market moves and on the firm's own positioning, quantities a house watches and helps shape. With vulnerabilities it will depend on the capability of a tool the bank switches on itself, and the rate of improvement is set by the model provider rather than by the bank. A deadline regime calibrated for a known volume of findings therefore meets an input that grows from the outside.

The part that is not the model

The most revealing element is a term the FCA borrows from AI evaluation research and turns into a governance category. Harness engineering, by its own definition, means “the environment, controls and processes around an AI model that makes its outputs useful, safe and reliable”. The finding attached to it is uncomfortable for anyone about to choose a model: the value obtained is determined “less by the models themselves and more by the technical and operational environment they're deployed in”.

The review names four components: specialist tooling, robust validation processes, operational guardrails and human expertise. On guardrails the regulator gets specific, meaning limits on model permissions, human approval for higher-risk actions, and controls over access to sensitive systems and data. Without these, it reports, models generate large numbers of findings that are “technically possible but difficult to validate, prioritise or act upon”.

For practitioners the implication is direct. Asking which model to deploy is a fair question, but it is the second one. The first is whether the firm has the environment in which a good model produces any value at all. Several firms told the FCA they approach the subject through deliberately narrow use cases, testing their own readiness before scaling. That is the most practical recommendation in the paper, even though the paper does not present it as one.

Several firms reduce it to a phrase: frontier AI acts as a “stress test of their existing cyber-resilience capabilities”. What the models expose is not primarily technical gaps but, in the regulator's words, weaknesses in “the people, systems and processes that fix them”.

The paper names neither sector nor method

A qualification belongs here, and the review invites it. The FCA files it under “Multi-firm review”, yet the text describes itself as a summary of “observations reported by firms during our engagement”. That is self-reporting by participants rather than examination by the regulator. The text accordingly avoids verbs such as “reviewed” or “assessed” throughout, working instead with “engaging”, “told us” and “reported”.

Comparison with nine other FCA multi-firm publications from 2018 to 2026 makes the difference measurable. All nine name the sector examined, eight set out their method in detail, and five give an exact number of firms, from “a sample of 14 firms” in the 2023 liquidity review to “eleven wholesale banks” in the 2025 off-channel review. The frontier AI review gives neither sector nor method nor count. The absent count alone would be unremarkable, since four of the nine comparators also omit it. The combination is what stands out: this is the only one of the ten leaving both scope and method open at once.

No accusation follows from that, but a caveat does. Anyone citing these observations is citing the experience of firms already advanced enough to run frontier AI in cyber defence. That is a slice of the front runners, not a picture of the market. It fits a second peculiarity: the FCA directs its findings expressly at “particularly small to medium-sized firms”. The advanced were asked; the followers are the readership.

A firm's own AI use falls outside the high-risk regime

One clarification, because the market reflex runs the other way. A bank deploying a frontier model internally to probe its own systems for weaknesses is not operating a high-risk AI system within the meaning of the AI Act. Annex III lists eight areas, from biometrics through employment to creditworthiness assessment. Cybersecurity and vulnerability discovery appear in none of them. Classification follows the purpose of use, not the capability of the model.

What remains points elsewhere. A model of this scale may be classified as a general-purpose AI model with systemic risk, and that assessment expressly takes offensive cyber capabilities into account. Those obligations fall on the model provider, not on the bank as a deployer. The workload for the firm arises from the DORA regime, not from the AI Act.

Recommendations

1. Measure the bottleneck before detection capacity rises

Before the next tooling rollout: The number that matters is not how many vulnerabilities a firm finds, but how many it can validate, test and put into production per unit of time. That figure is usually unknown and can be reconstructed from the change records of the past twelve months. Once known, it shows how much additional finding the operation can absorb, and the rollout can be staged accordingly.

2. Test self-set patch deadlines against the expected increase

This month: Article 10(4)(d) of Commission Delegated Regulation (EU) 2024/1774 requires deadlines to be set and an escalation route defined for when they are missed. Both were typically calibrated against a finding volume produced by conventional scanning. Multiplying detection without revisiting that calibration produces escalations by design. Deadlines may be differentiated on a risk basis; they simply have to be reasoned and documented.

3. Put the dependency map ahead of the prioritisation reform

Next quarter: Prioritising by attack path rather than severity presumes it is known which systems carry which critical or important functions and what they depend on. Without that map the change stays an intention. The order is dependencies first, prioritisation logic second, not the other way round.

4. Ask suppliers now rather than after the first incident

At the next contract round: The DORA regime already requires firms to check whether ICT third-party providers address and report vulnerabilities promptly, and to track third-party software libraries including open source. The practice the FCA reports goes one step further, asking suppliers how they handle AI-enabled vulnerability discovery themselves and whether they can absorb rising patch volumes. That is the answer worth having before it is needed.

Glossary

Frontier AI: in the FCA's definition, the most advanced AI models available at any given time. The term is deliberately relative and denotes no fixed level of capability.

Harness engineering: the environment of tooling, controls and processes around a model that makes its output usable, safe and reliable. On the FCA's finding it matters more to the value obtained than the choice of model.

Vulnerability chaining: combining several individually low-rated weaknesses into a viable attack path. It undercuts prioritisation based on severity considered in isolation.

Critical or important function: defined in Article 3(22) DORA. The weekly scanning duty and the prioritisation of remediation both attach to it. It corresponds to what the FCA calls “important business services”.

CMORG: the Cross Market Operational Resilience Group, a body of British financial institutions and public authorities. Its June 2026 guidance on frontier AI is voluntary and expressly sets no supervisory expectations.