
Imagine inheriting a filing cabinet. Not a neat, organised one, but a vast, overflowing one that has been added to for fifteen years by dozens of different people, never sorted, never labelled, and now legally yours to manage. You need to know what's in it, what to keep, what to delete and what might be sensitive. And there are thousands more like it across the building.
That is not far from the reality facing Knowledge and Information Management (KIM) professionals working with government email today.
The problem hiding in plain sight
Email is everywhere in government, and it has been accumulating for decades. Unlike documents stored in formal systems, email tends to pile up quietly. It’s a mixture of business decisions, administrative threads, sensitive exchanges and personal messages, all sitting together with no obvious way to tell them apart. At some point, someone must go through it.
That someone is usually a KIM specialist, and the process is painstaking. Mailbox by mailbox, email by email, they work through content that can run into the hundreds of thousands of items. They need to make judgements about what is worth keeping for the long term and what meets legal and compliance requirements. They also assess what might be sensitive enough to need careful handling, including decisions about whether it can be shared or released at all. It is skilled, important work and at the scale government operates, doing it manually simply is not sustainable.
This is the problem Project Kestrel was built to explore. Led by the AI Data Readiness Incubator team within Government Digital Service and developed in close collaboration with the Government KIM team, Phase 1 set out to test whether modern data science techniques could offer a better way. Not to replace human judgement, but to support it, making the work faster, more consistent and easier to explain.
The data problem behind the data problem
Here is where it gets interesting. Before the team could even begin testing new approaches to email management, they ran into a challenge that will feel familiar to anyone working on data innovation in government: you cannot safely experiment with the data you need.
Real government emails contain personal information, sensitive material and content that is legally protected. Using them to develop and test new tools, however well-intentioned, creates real risks. So, the team made a deliberate choice, rather than finding ways to work around that constraint, they would solve it directly.
The result is Tiger Heron, a synthetic email generator. Using large language models, Tiger Heron produces realistic-looking emails that closely resemble the kind of content government teams encounter in practice, without containing any real or sensitive information. Teams can build with it, test with it, break things with it and start again, all without ever touching a real inbox.
This might sound like a workaround, but it is something more significant than that. Access to a safe, realistic data set is one of the most common blockers to data innovation in the public sector. Tiger Heron may offer a practical model for getting past it and one that could apply well beyond email.
Bringing order to the chaos
With a safe data set in place, the team were able to build and test Cuckoo, the capability at the heart of Project Kestrel. Where Tiger Heron solves the data problem, Cuckoo takes on the scale problem.
Within a controlled test environment, Cuckoo uses machine learning to read through large volumes of emails and group them into meaningful clusters based on their content. Instead of opening every message individually, a KIM professional can see their data organised into coherent groupings - similar topics, similar threads, similar types of content - presented through an interactive visual interface. That shift, from an undifferentiated pile to a structured view, changes what is possible. Reviewers can direct their attention where it matters most, spot patterns they might otherwise miss and make decisions that are easier to document and defend.
Together, Tiger Heron and Cuckoo form an end-to-end prototype, generate realistic test data, then analyse it in a way that supports faster, more transparent decision-making.
What careful experimentation looks like
Phase 1 was never going to produce a finished product, and it was not designed to. Alongside the two tools, the team carried out user research with KIM professionals, gathered ethics input to ensure responsible practices were built in from the start, engaged stakeholders across government and developed best-practice materials to support what comes next. Human oversight has been central throughout; the goal is always to keep people in control of the decisions that matter.
That is what responsible AI adoption actually looks like in practice, not bold claims about transformation, but careful groundwork, honest about what is working and what still needs testing.
Project Kestrel is an early example of that kind of experimentation. The filing cabinet is still full. But there is now a better sense of how to approach it.
If you have a government use case where synthetic data could help you work around the challenges of personal data, get in touch with the AI and Data Readiness team on dataAIReadiness_incubator@dsit.gov.uk
Leave a comment