Dutch Intelligence Got Caught Feeding Citizen Data Into Its Own AI, and It’s Not Even the First Time

The CyberSec Guru

Dutch Intelligence Got Caught Feeding Citizen Data

If you like this post, then please share it:

Buy me A Coffee!

Support The CyberSec Guru’s Mission

🔐 Fuel the cybersecurity crusade by buying me a coffee! Why your support matters: Zero paywalls: Keep the main content 100% free for learners worldwide.

“Your coffee keeps the servers running and the knowledge flowing in our fight against cybercrime.”☕ Support My Work

Buy Me a Coffee Button

Governments often ask citizens to trust them with sensitive personal information. That trust comes with an expectation that the data will be collected only when necessary, stored securely, and handled according to the law. A newly published Dutch oversight report suggests that, inside the Netherlands’ intelligence services, those expectations haven’t always matched reality.

On July 1, the Dutch intelligence watchdog, the CTIVD, released a supervisory report examining how the country’s civilian intelligence agency (AIVD) and military intelligence agency (MIVD) manage what are known as bulk datasets. These collections can contain millions of individual records, ranging from names and phone numbers to location history, social media activity, internet metadata, and even the contents of communications.

The report paints a worrying picture. It describes weak technical controls, inconsistent oversight, excessive data retention, and employees gaining access to information they should never have been able to see. None of those findings are good news on their own. But hidden among the technical details is something that raises a much larger question about where this data is ultimately ending up, especially as intelligence agencies increasingly adopt artificial intelligence tools.

Understanding bulk datasets

The term “bulk dataset” might sound abstract, but the concept is fairly straightforward. Rather than collecting information about one specific suspect, intelligence agencies sometimes acquire enormous collections of personal data containing information about large numbers of people, the overwhelming majority of whom are not under investigation.

Dutch law permits the AIVD and MIVD to use these datasets for national security purposes such as counter-terrorism, counter-espionage, and identifying foreign threats. That authority is not unlimited. The law places several conditions on how the data must be handled.

Retention periods are limited. Access is supposed to be restricted to specifically authorized personnel. When information is found to have no operational value, it should eventually be deleted instead of remaining in storage indefinitely. The entire system is intended to balance intelligence gathering with citizens’ privacy rights.

The CTIVD report suggests that this balance has not always been maintained. One particularly section discusses where these datasets originate. Some are supplied by other public authorities. Others come from commercial providers. More unusually, some originate from information that first appeared through criminal activity, including datasets stolen by hackers before being traded or sold online.

That isn’t speculation or a hypothetical scenario imagined by privacy advocates. The report itself lists illegally obtained datasets as one of the categories that intelligence services may encounter and process under existing legal frameworks. It illustrates just how complicated modern intelligence collection has become. Information stolen during cybercrime investigations can eventually become intelligence material, provided legal requirements are met.

For security professionals, that raises uncomfortable questions. Data breaches that expose millions of ordinary people may not simply disappear into criminal marketplaces. The same information can continue circulating for years, passing between brokers, researchers, law enforcement, and intelligence agencies long after the original victims believed the incident had faded into history.

What the watchdog found

Reading through the CTIVD’s findings feels surprisingly familiar if you’ve ever reviewed the results of an enterprise security audit. The report describes repeated failures in basic governance rather than a single catastrophic incident.

📬 Stay Ahead of Cyber Threats

Get the latest cybersecurity news, critical vulnerabilities, threat intelligence, tutorials, and exclusive giveaways delivered straight to your inbox. No spam. Unsubscribe anytime.

Subscribe to the Newsletter →

Employees from certain departments were able to access personal data beyond their authorized permissions. Large datasets remained stored beyond the legally permitted retention periods. More teams across both intelligence services are now working with bulk datasets, yet no central function consistently verifies whether those datasets are being managed according to legal requirements.

According to the CTIVD, these aren’t isolated administrative mistakes. The problems stem from both technical limitations within existing systems and operational procedures that fail to enforce compliance in practice. In other words, the policies may exist on paper, but the systems enforcing those policies are not consistently doing their job.

CTIVD chairman Hugo Hillenaar addressed the privacy implications directly. Processing bulk datasets, he noted, inevitably affects everyone whose information appears inside them. Most of those individuals have no connection whatsoever to terrorism, espionage, or any activity that originally justified collecting the data. They appear simply because large datasets, by their nature, contain information about vast numbers of ordinary people alongside any individuals who may actually interest intelligence investigators.

The watchdog ultimately issued thirteen recommendations aimed at improving the situation.

Some recommendations focus on legal clarity, including a more precise definition of what qualifies as a bulk dataset. Others deal with classification, such as ensuring data is treated as non-public whenever its original public availability cannot be verified. Several recommendations, however, focus squarely on technical controls.

Among the most significant is the call for substantially better access logging. The oversight body wants intelligence services to maintain records showing who accessed a dataset, whether that access was authorized, when it occurred, and what actions were performed afterward. That may sound like a basic cybersecurity requirement, because it is.

Most commercial organizations handling sensitive customer information already regard comprehensive audit logging as a foundational security control. Banks, healthcare providers, cloud platforms, and enterprise software vendors routinely rely on detailed access logs to investigate suspicious activity, satisfy regulatory requirements, and demonstrate compliance during audits.

When an external oversight body discovers that similar visibility is missing inside agencies responsible for safeguarding some of the country’s most sensitive intelligence holdings, it naturally raises questions about how effectively those internal controls have been operating.

From a purely security perspective, this is less about one isolated compliance failure and more about governance. If you cannot reliably determine who accessed sensitive data, when they accessed it, or whether that access was justified, then detecting misuse becomes significantly harder. That’s precisely the kind of weakness defenders spend years trying to eliminate inside corporate environments, and it’s the same category of weakness the CTIVD has now identified inside Dutch intelligence systems.

The AI question changes everything

The technical findings alone would have made this an uncomfortable report for the Dutch intelligence services. Weak access controls, inconsistent oversight, and data being kept longer than the law allows are serious issues regardless of who is responsible. What turned this from a routine compliance story into something much bigger was what happened after the report was published.

Data Handling and Privacy Oversight
Data Handling and Privacy Oversight

Dutch digital rights organization Bits of Freedom, which has been following intelligence oversight issues for years, argued that the CTIVD’s findings point toward a broader concern. Based on the report and accompanying government documents, the group believes the AIVD and MIVD are increasingly using artificial intelligence to analyse bulk datasets and that these datasets may include information originating from previously leaked or illegally obtained sources.

The CTIVD report itself has discussed the use of AI-assisted analysis and bulk data processing. However, it does not explicitly state that citizens’ personal data is being used to train foundation models or large language models. That stronger interpretation comes from Bits of Freedom and several security researchers who have examined the findings. Even with that distinction, the privacy implications are difficult to ignore.

If personal data is collected unlawfully or retained beyond the legal deadline, there is at least a theoretical remedy. Oversight bodies can order it deleted. Courts can intervene. Citizens may eventually exercise their rights under European privacy law. Once that same information becomes part of an AI workflow, things become considerably more complicated.

Traditional databases are relatively easy to modify. Delete the records, update the backups, and the information is gone. AI systems don’t necessarily work that way. If data contributes to model development, feature extraction, or other machine learning processes, removing every trace afterward may be technically difficult or, in some cases, practically impossible without retraining systems from scratch.

Researchers have debated this issue for years. The challenge is no longer hypothetical. Courts, regulators, and AI developers across Europe are already wrestling with how the “right to erasure” under the GDPR should apply when personal information has influenced machine learning systems.

For ordinary Dutch citizens, the concern is obvious. They have no visibility into which datasets intelligence agencies hold, how long those datasets remain available, or whether AI systems are analysing information that should have been deleted years earlier.

Unlike a commercial AI company operating inside the European Union, intelligence agencies operate under a very different legal framework. There is no privacy dashboard. No opt-out button. No account settings page explaining how your data is being processed. Oversight largely happens behind closed doors, with watchdog organizations reporting their findings months or years later.

That lack of transparency is exactly why reports like this receive so much attention.

This isn’t the first time

Perhaps the most striking part of the entire story is how familiar it sounds.

The Netherlands has already been through a remarkably similar controversy.

History of Dutch Intelligence Agency Dataset Acquisition
History of Dutch Intelligence Agency Dataset Acquisition

Back in 2020, it emerged that the AIVD and MIVD had retained personal data belonging to millions of citizens well beyond the legal eighteen-month retention limit established under Dutch intelligence law.

Bits of Freedom filed a formal complaint, arguing that the agencies had continued storing information that should already have been destroyed.

The complaint eventually reached the CTIVD’s complaints division.

In 2022, the watchdog issued a binding decision ordering five major citizen datasets to be deleted. According to publicly available findings, the responsible ministers had initially decided to retain those datasets despite earlier advice from the CTIVD’s supervisory department recommending their destruction.

That sequence of events is difficult to overlook. First, the oversight body warned that the data should be deleted. The advice wasn’t followed. A complaint was soon filed. Only after a legally binding ruling were the agencies required to destroy the datasets.

Fast forward several years and another CTIVD report is again describing many of the same underlying problems. Unauthorized access. Retention periods exceeding legal limits. Weak technical safeguards. Insufficient procedural controls.

The technology has evolved since 2020, but many of the governance failures appear familiar.

For privacy advocates, that’s arguably the most troubling part of the report. A single compliance failure can sometimes be explained by technical debt or organizational complexity. Seeing similar weaknesses identified years apart raises harder questions about whether earlier recommendations produced lasting change.

A warning from the security community

Dutch security researcher Bert Hubert, who has written extensively about intelligence oversight and digital rights, argues that the report should also be viewed in the broader context of today’s data ecosystem.

Massive data breaches have become routine. Every year, billions of records are stolen from companies, government agencies, healthcare providers, financial institutions, and online platforms. Some datasets circulate freely on hacking forums. Others are sold through criminal marketplaces or commercial brokers before eventually appearing elsewhere. Taken individually, many of those leaks seem limited. Combined over years, they become something else entirely.

Hubert argues that intelligence agencies no longer need extraordinary new surveillance powers to assemble highly detailed profiles of large portions of the population. Much of the raw material already exists in commercially traded datasets, historic breaches, publicly available information, and other sources that can legally or operationally find their way into intelligence systems.

That observation doesn’t suggest intelligence agencies are building complete dossiers on every citizen. The available evidence doesn’t support such a claim. It does, however, highlight how dramatically the data landscape has changed. Twenty years ago, assembling detailed information about millions of people would have required enormous dedicated surveillance efforts. Today, unprecedented quantities of personal information already exist across countless breached databases, commercial data markets, public records, and online platforms.

The challenge has shifted from collecting data to managing it responsibly once it has been acquired. That’s precisely why the CTIVD’s findings matter beyond the Netherlands. The report isn’t simply about one country’s intelligence services failing compliance checks. It reflects a much broader problem facing governments around the world: intelligence agencies now have access to more personal information than ever before, while the legal, technical, and ethical safeguards governing that information are still struggling to keep pace.

How the Dutch government responded

The Dutch government does not dispute that processing bulk datasets affects the privacy of ordinary citizens. In its response to parliament, it acknowledged that handling such large collections of personal information inevitably interferes with individual privacy. At the same time, it argues that these capabilities remain essential for countering espionage, terrorism, cyber threats, and other national security risks.

Justice and Security Minister Dilan Yeşilgöz-Zegerius and Defence Minister Jan Willem Uitermark told parliament that they broadly accept the CTIVD’s recommendations and have already begun implementing several improvements. According to the government’s response, work is underway to strengthen internal procedures, improve technical controls, and clarify how bulk datasets should be managed throughout their lifecycle. That sounds reassuring on paper, but the details remain thin.

The government has not publicly committed to major changes regarding the use of artificial intelligence with bulk datasets. Nor has it directly addressed concerns raised by Bits of Freedom about whether personal information that should have been deleted could still be used in AI-assisted analytical systems. That silence has become one of the more closely watched aspects of the debate.

Privacy campaigners are no longer asking only whether intelligence agencies are collecting too much information. Increasingly, they’re asking what happens after the collection phase ends. Is the information archived? Shared? Used to train analytical systems? Deleted when required? Or does it continue influencing future intelligence work in ways that oversight bodies cannot easily verify?

Those questions remain largely unanswered.

A new intelligence law could reshape the debate

The timing of the CTIVD report is significant for another reason.

The Dutch Intelligence and Security Services Act, commonly known as the Wiv 2017, is currently undergoing revision. A draft of the updated legislation is expected to enter public consultation, giving citizens, academics, civil society groups, and technology experts an opportunity to comment before the proposals move further through the legislative process.

That consultation is likely to become an important test of how the Netherlands intends to regulate intelligence work in an era increasingly shaped by artificial intelligence and large-scale data analytics.

Supporters of stronger intelligence powers argue that agencies need access to massive datasets to identify sophisticated threats, especially as hostile states, ransomware groups, and terrorist organizations become more technologically advanced.

Privacy advocates see the situation differently. For organizations like Bits of Freedom, the question isn’t whether intelligence agencies should have powerful investigative tools. It’s whether meaningful safeguards exist once those tools are deployed.

If oversight reports continue identifying the same weaknesses every few years, critics argue, expanding surveillance capabilities before fixing governance problems risks repeating the same cycle with even larger quantities of data.

Whether lawmakers tighten oversight requirements or grant intelligence services greater operational flexibility is likely to become one of the most closely watched parts of the legislative process.

Why this matters beyond the Netherlands

It’s easy to dismiss this story as an internal Dutch oversight dispute, but the underlying issues extend far beyond one country.

Across Europe, North America, and much of the developed world, intelligence agencies are rapidly incorporating machine learning into investigations. Governments are exploring AI-assisted systems for language translation, document classification, anomaly detection, metadata analysis, facial recognition, and identifying patterns across enormous datasets that would be impossible for human analysts to process manually.

None of that is inherently controversial. Used carefully, these technologies can help analysts process information faster while allowing investigators to focus on genuine threats instead of drowning in raw data.

The controversy begins when oversight mechanisms fail to keep pace with those technological capabilities. Rules that were written years before generative AI became mainstream were largely designed around traditional databases and human analysts. Modern analytical systems introduce entirely new questions about data retention, traceability, accountability, explainability, and deletion. Legislators, regulators, and courts across Europe are still trying to determine where existing privacy laws end and AI governance begins.

The Dutch case offers a glimpse of what that transition looks like in practice. It shows how traditional compliance failures, such as weak access controls or excessive data retention, can become far more significant once artificial intelligence enters the picture.

The bottom line

Strip away the national security context for a moment and imagine the same findings applied to a private company.

Employees accessed personal information they weren’t supposed to see. Sensitive datasets remained stored after legal retention deadlines had expired. Audit logging wasn’t robust enough to establish who accessed what, or why. Some of the data being processed originated from information that had previously surfaced through criminal data leaks. Few regulators would consider that an acceptable state of affairs.

The difference, of course, is that intelligence agencies don’t operate like ordinary organizations. Much of their work necessarily takes place behind closed doors. Public scrutiny is limited, external auditing happens only periodically, and citizens often learn about problems years after the underlying practices began. That’s precisely why independent oversight bodies such as the CTIVD exist.

This latest report demonstrates that oversight can identify serious shortcomings. Whether it can consistently drive meaningful change is a different question. Perhaps the most troubling aspect of the entire episode isn’t any single finding. It’s the pattern.

Six years after earlier investigations uncovered excessive data retention involving millions of citizens, another watchdog report has identified familiar weaknesses: data kept longer than permitted, insufficient technical safeguards, incomplete access controls, and governance problems that remain unresolved.

The technology has evolved. Artificial intelligence has entered the intelligence workflow. The volume of available data has grown dramatically. Yet many of the same governance questions remain unanswered. For Dutch citizens, the immediate concern isn’t whether intelligence agencies should exist or whether they should investigate national security threats. Those powers are unlikely to disappear.

The more pressing question is whether the safeguards designed to protect ordinary people are evolving as quickly as the technologies those agencies now have at their disposal. The CTIVD has delivered another warning. Whether this one produces lasting reforms or simply joins the growing archive of oversight reports documenting the same recurring failures, is something that will only become clear over the next few years.

FAQs

What is the CTIVD and what does it oversee?

The CTIVD is the independent Dutch body that oversees the AIVD and MIVD, the country’s civilian and military intelligence services. It checks whether both agencies comply with the law when collecting, storing, and processing data, and it can issue binding rulings after complaints.

What are AIVD and MIVD bulk datasets?

Bulk datasets are large collections of personal data, sometimes covering millions of records, that can include names, phone numbers, location data, social media activity, and communication content. They come from other government bodies, commercial data brokers, and in some cases data obtained by hackers and resold online.

Are Dutch intelligence services really training AI on citizen data?

Bits of Freedom, a Dutch digital rights organization, says the pattern in the CTIVD’s July 2026 report strongly suggests AIVD and MIVD are training their own AI systems on citizen data, including data that appears to have been purchased from breach markets. This is an inference from the report rather than a confirmed admission by the agencies themselves.

Did this happen before 2026?

Yes. In 2020, the same watchdog found that AIVD and MIVD had retained citizen data on millions of people well beyond the legal retention limit. After a complaint from Bits of Freedom, the CTIVD’s complaints department issued a binding ruling in 2022 ordering the destruction of five bulk datasets.

What happens next with the Wiv 2017 revision?

The law governing the Dutch intelligence services, the Wiv 2017, is being revised and is expected to go to public consultation around August 2026, with a roughly six-week window for citizens and organizations to submit responses.

Buy me A Coffee!

Support The CyberSec Guru’s Mission

🔐 Fuel the cybersecurity crusade by buying me a coffee! Your contribution powers free tutorials, hands-on labs, and security resources.

Why your support matters:
  • Writeup Access: Get complete writeup access within 12 hours
  • Zero paywalls: Keep the main content 100% free for learners worldwide

Perks for one-time supporters:
☕️ $5: Shoutout in Buy Me a Coffee
🛡️ $8: Fast-track Access to Live Webinars
💻 $10: Vote on future tutorial topics + exclusive AMA access

“Your coffee keeps the servers running and the knowledge flowing in our fight against cybercrime.”☕ Support My Work

Buy Me a Coffee Button

If you like this post, then please share it:

News

Discover more from The CyberSec Guru

Subscribe to get the latest posts sent to your email!

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from The CyberSec Guru

Subscribe now to keep reading and get access to the full archive.

Continue reading