FCA warns frontier AI is finding cyber vulnerabilities faster than firms can fix them
Key Points
- The FCA said frontier AI is finding vulnerabilities faster than firms can validate and fix them
- Firms report AI chaining low-rated flaws into attack paths invisible to traditional scanning
- The value of frontier AI depends on the governance and tooling around the model, not the model itself
- Several firms called frontier AI a stress test of their existing cyber-resilience capabilities
- The publication introduces no new rules and is aimed at helping smaller firms prepare
The Financial Conduct Authority (FCA) has warned that financial firms using frontier AI are now discovering security flaws faster than their engineering, patching, and change teams can act on them.
The regulator published the finding in a new set of insights drawn from its engagement with firms on how they are using, testing and preparing for frontier AI models with cyber capabilities.
It follows a joint statement from the FCA, the Bank of England and HM Treasury in May 2026, which described frontier AI models as a step-change in capability with significant implications for cybersecurity and operational resilience.
The FCA said it wanted small and medium-sized firms in particular to learn from the experience of larger firms already running these models.
Vulnerability discovery is outpacing fixes
Firms told the regulator that frontier AI is increasingly supporting the identification, validation and prioritisation of vulnerabilities, which is increasing the pressure on their remediation processes.
Firms using the models said they can now identify weaknesses in software, systems and infrastructure rapidly, and that the challenge has shifted from finding vulnerabilities to responding to a continuous flow of findings.
The FCA said that even where expert review discounts a substantial proportion of model outputs, the remaining volume of genuine vulnerabilities still puts considerable pressure on remediation teams, engineering resources and change management processes.
It said firms may need to understand where bottlenecks will arise in validation capacity, engineering resource, patch testing, emergency change controls and evidence of closure, and whether existing processes can speed up without creating operational instability.
Models are chaining low-rated flaws into attack paths
Firms highlighted that frontier AI models can combine multiple lower-rated security flaws, a technique known as vulnerability chaining, to create alternative routes to compromise.
Several firms said the relationships between vulnerabilities, systems and dependencies were not visible through traditional scanning and testing.
With that visibility, firms said their vulnerability management decisions are increasingly driven by the disruption an exploited attack path would cause, rather than by the severity rating of each flaw in isolation.
The FCA said this is pushing firms towards a risk-based approach that weighs exploitability, business service impact, prerequisites to exploit, compensating controls and dependency on the vulnerable system.
Not model-specific
Firms reported that the value they get from frontier AI is determined less by which model they use and more by the governance, tooling, controls, human oversight and operational environment around it, which the FCA calls the harness.
The regulator said models are most effective when supported by specialist tooling, robust validation processes, operational guardrails and human expertise, alongside an understanding of how systems and assets underpin important business services.
Operational guardrails cited by firms include limits on model permissions, human approval for higher-risk actions and controls over access to sensitive systems and data.
Without these, firms said models can generate large numbers of findings that are technically possible but difficult to validate, prioritise or act upon.
A true stress test
Several firms characterised frontier AI as a stress test of their existing cyber-resilience capabilities, with organisational readiness emerging as the primary challenge.
The FCA said the models are exposing weaknesses in vulnerability management, access management controls, dependency mapping and remediation processes, as well as in the people, systems and processes that fix them.
Some firms are running targeted deployments rather than enterprise-wide rollouts as a way to test readiness before scaling.
Those firms said the approach gave them a clearer view of where post-discovery validation becomes constrained, how remediation ownership operates, whether change processes can absorb more findings, and how well cyber teams can separate theoretical weaknesses from credible exploitation routes.
The FCA said firms that do basic cyber-resilience practices well, with clear accountability and effective oversight, are likely to be better positioned for the challenges frontier AI brings.
Human judgement still a factor
Despite the increase in automation, firms told the regulator that human oversight remains critical, with specialist expertise still needed to validate findings, assess relevance, set priorities and make risk-based decisions.
Several firms observed that the benefits of autonomous discovery are limited where processes cannot keep pace with the volume of output.
The FCA said senior leaders, governance forums and risk committees may need clearer visibility of how frontier AI affects vulnerability registers, remediation capacity, supplier dependencies and operational resilience.
Firms also flagged supplier preparedness, cloud dependencies, software supply chain visibility and shared infrastructure as growing concerns.
Some firms said they are already asking suppliers whether they use AI-enabled vulnerability discovery, how they validate findings, how they notify customers and whether they can remediate quickly.
The FCA set out questions for firms to consider across each area, including who owns decisions about frontier AI in cyber-resilience work, whether firms can tell technically plausible outputs from genuinely exploitable findings, where remediation bottlenecks sit, and whether key suppliers are preparing for increased patch volumes.