Synthetic Data Gains Traction in Financial AI Development
Financial institutions are facing a significant challenge in developing and training artificial intelligence (AI) systems due to limited access to sensitive data. According to a joint survey by the Bank of England and Financial Conduct Authority, 75% of responding UK financial firms were already using AI in 2024, with another 10% planning to do so within three years.
The growth in AI adoption increases demand for models that can detect fraud, support underwriting, improve customer service, model liquidity, and automate internal processes. However, the data needed to train and test these systems is often highly sensitive and difficult to share or expose.
Synthetic data, artificially generated information designed to reproduce statistical characteristics of real datasets, is emerging as a potential solution. Used carefully, it can give financial institutions more room to train, test, and stress AI systems while reducing dependence on raw customer data.
However, using synthetic data carelessly can reproduce bias, create false confidence, and even leak information about the real data from which it was generated. The strategic value of synthetic data is not just in privacy but also in controllability, as it can be designed to reflect rare or extreme events that are difficult to recreate with real-world datasets.
The Bank of England has explicitly noted that generative AI can support its work through the production of synthetic data, and its AI Consortium highlighted synthetic data as a potential response to gaps in training-data adequacy. Synthetic data could change how financial AI is tested by allowing institutions to ask not only whether a model performs well on historical cases but also how it behaves across thousands of controlled combinations of inputs.
The UK Information Commissioner’s Office describes synthetic data as a privacy-enhancing technique that can be useful for AI training when organisations cannot access large real-world datasets. However, institutions should not treat synthetic data as a shortcut around privacy governance and still need to understand how the source data was obtained, how the generator was trained, whether individual records can be inferred, what privacy tests were applied, and whether the output remains sufficiently useful after safeguards are introduced.
Representativeness is another problem with synthetic data, as it can preserve or amplify historical bias if the source data contains weaknesses. The EU AI Act places emphasis on data governance for high-risk AI systems, requiring training and validation processes to ensure that AI systems do not perpetuate existing biases.