Right now, it seems everyone’s talking about AI. Policymakers want strategies, consultants are selling transformation, and the demos make it all look effortless. Meanwhile, investment is following demand accordingly: around a third of British businesses plan to invest in AI in 2026, while governments are racing to embed it into public services.

But when it comes to data – arguably AI’s most critical dependency – it seems far less attention is being paid to whether it’s accurate, consistent or even understood.
The average government organisation of today is now rich in data that’s prime for AI – think citizen records, service usage data, operational reports, and staff information. Yet most of this data sits in multiple systems, is duplicated and, in many cases, the knowledge of what that data represents is lacking. And let’s not forget the data that continues to come in as departments and services grow.
AI learns from patterns and, whatever you give it, it inherits. This means that while it reflects the strengths in your dataset, it will just as easily absorb the mistakes and biases. It also means that any AI trained on open data from the internet (a place largely shaped by the experiences of a relatively narrow group that’s far from representative), as opposed to internal data, is likely to reproduce the same demographic and systemic biases embedded in the web itself.
So, what happens when you throw messy, inconsistent or open data at an algorithm expecting good results? Usually a poorly-executed AI project that can’t scale, delivers unreliable outputs, and ultimately ends up costing more.
This is why the AI story starts with getting your data in order.
If you liked this content…
How to get data ready for an AI deployment
Getting data ready for AI is, unfortunately, not particularly glamorous. It means digging through systems to figure out what you actually have, cleaning anything corrupted, and connecting sources that have never spoken to each other. It also means moving your thinking from reactive to predictive, and making sure you have the right skills in place.
Most think this is a quick phase (it is not). In reality, it’s a lengthy process that requires sustained attention across the following areas:
Visibility, trust and governance (the foundations): To start with, you need visibility. By that I mean you must know what data exists, where it lives, and who owns it. Then, you need trust, which means fixing duplicates, standardising formats, and resolving inconsistencies. And finally, you need governance, from access rules to auditing to ethics. All of this means you must operate within a ‘walled garden’, and train only on internal, governed data. This will not only avoid producing skewed outputs that have little relevance to citizens, staff, or public services, but will also strengthen data protection, trust and accountability. As many will know, public confidence in AI is very much up and down in general, especially as staff experiment with open AI tools in the absence of clear guidelines. But this worry is particularly strong when it comes to the government. So in order to trust AI, people need a bias-mitigated environment that contains only government data, ensuring secure and efficient knowledge sharing that stays within government walls. All of this is the heavy-lifting phase, but skip it and your AI project will fail. I like to compare it to water infrastructure, whereby clean water is essential for public health. Clean data is the same for AI.
From foundations to scale: Once you’re set with a solid foundation – whereby quality, governed data is in place – that’s when it gets interesting and you can start to move your thinking from reactive to predictive. You can set up DataOps to automate cleaning and monitoring so your data stays reliable; MLOps handles deploying your models and retraining them as new data comes in; and alongside this, it’s important to have oversight, assurance teams, and clear rules around ethics and governance. Together this creates a sort of data fabric; a living system that keeps AI running properly. In my experience, those that treat it like this – an ecosystem rather than a one-off project – are the ones that actually scale their AI project successfully.
Skills, people and the next bottleneck: With the above in mind, it’s also important to think about your people and their skills. I believe this year we’ll start to see data readiness become as crucial as cybersecurity was ten years ago; a shift that in turn will see data engineers, not just prompt engineers, become organisations’ most valuable players. Every sector will need their skills, from government and banking to healthcare and any organisation automating anything. As such, my advice is to start upskilling your team in metadata, data lineage, ontologies and ethics immediately, as all these skills will soon be in high demand.
The bottom line here is that government organisations must start accepting AI as the brain, and data readiness as the circulatory system, and plan accordingly. Because when AI performs well, it’s usually down to clean and complete data. And when it fails, this foundation is likely missing.








