Skip to content
NLEN
Illustration: Why Public Services Choose Their Own Models

Why Public Services Choose Their Own Dutch Models

By Ivo Donker — compiled with AI assistance (Claude & Gemini)

The adoption of artificial intelligence within public administration is at a decisive tipping point. Where government institutions in the experimental early phase frequently relied on public, commercial cloud interfaces for pilots and exploratory projects, that approach runs into hard legal, ethical, and technical limits in operational practice. The overview article on the status of AI in government already made clear how government bodies must balance the drive for innovation against strict frameworks; the structural shift toward sovereign, open-source, and domestic language models is the direct answer to that challenge.

Public organizations — from major implementing agencies such as the Dienst Uitvoering Onderwijs (DUO), the Sociale Verzekeringsbank (SVB), and the Belastingdienst to regional environmental services and individual municipalities — manage exceptionally sensitive personal data and make automated or semi-automated decisions with far-reaching legal consequences for citizens. Routing official case files through external API services from commercial parties outside the European Union introduces fundamental risks in the areas of information security, data breaches, and legal violations. In addition, global mega-models turn out to fail remarkably often on the subtle terminology of Dutch administrative law, local regulations, and the specific tone that citizens may expect from a trustworthy government.

The choice for a government's own, locally hosted language models is therefore by no means a protectionist reflex or a technical dogma, but a deliberate, risk-averse management strategy. In this article we analyze the strategic, technical, legal, and linguistic motives that lead public services to set up their own language technology, including the associated measurement methods, cost considerations, and unavoidable operational limitations.

1. Digital Sovereignty and Control Over the Technology Chain

Digital sovereignty within the public sector has evolved from an abstract policy ideal into an operational minimum level for continuity management. When a government body grafts its primary work processes — such as permit granting, policy evaluation, or citizen correspondence — onto closed external cloud models, an acute technological dependency arises. A unilateral change in service levels, abrupt price increases, or the sudden phasing out of a model version can directly lead to the failure of vital administrative chains.

By opting for open weights and open-source model architectures, the government retains full control over the lifecycle of its software. IT departments can download model weights, archive them locally, apply version control, and enforce deterministic outcomes. This prevents an unannounced update at an external supplier from suddenly undermining the consistency of official summaries. Those curious about the technical implementation of such architectures can consult the guide on running European open models yourself to see how local inference is architecturally set up.

Moreover, technological autonomy safeguards mandatory archiving and reconstructability. Under the Archiefwet (Archives Act) and the Wet open overheid (Woo, Open Government Act), government bodies must be able to account for decision-making processes down to the level of detail for years. If an external model is no longer available via a commercial provider's API after twelve months, a formal reconstruction of a historical administrative decision has effectively become impossible.

2. Privacy, Confidentiality, and the GDPR in Administrative Practice

The processing of personal data by the government is subject to strict principles of purpose limitation, data minimization, and storage limitation under the General Data Protection Regulation (GDPR). As soon as confidential documents — such as reports from youth protection services, medical indications, criminal sanction orders, or tax assessments — are sent to an external cloud environment, an extensive legal risk process is triggered, including mandatory Data Protection Impact Assessments (DPIAs) and complex transfer assessments.

Even if commercial cloud providers contractually stipulate that submitted data will not be used for further development of models, structural vulnerabilities remain around telemetry, metadata storage, and necessary error logging. In addition, the legal threat of foreign extraterritorial legislation, such as the US Cloud Act, remains a lurking danger for governments that must guarantee confidentiality. For a civil servant bound by a statutory duty of confidentiality, entering case details into a foreign cloud environment is legally barely defensible.

Running your own open language model within a shielded government cloud (such as the Rijksoverheid Cloud) or on on-premise hardware within your own data center structurally rules out data leaks to third parties. No token crosses the controlled network perimeter. This not only drastically reduces the complexity of privacy audits, but also protects the fundamental societal trust between citizen and state.

3. Meeting the Requirements of the EU AI Act

The entry into force of the European AI Regulation is forcing public organizations toward far-reaching professionalization of their model governance. The analysis on the content of the AI Regulation clearly explains the strict requirements that apply to risk assessment, transparency, human oversight, and technical robustness for AI systems.

A significant portion of government tasks falls directly under the category of high-risk systems within the meaning of the AI Act. Think of systems for risk profiling in social benefits, automated prioritization in inspections, or advanced personnel selection procedures. For such applications, the legislator requires that training datasets be representative, error-free, and free of impermissible bias, and that the model's operation be technically auditable for supervisory authorities such as the Autoriteit Persoonsgegevens.

With commercial closed models, it is impossible to meet this legal burden of proof, since the exact training composition and filtering process are shielded as trade secrets. With its own open model base, a government body can account exactly for which training corpora were used, which synthetic fine-tuning data was added, and how the bias evaluations were carried out. This satisfies the documentation obligation and prevents costly legal enforcement proceedings.

4. Transparency, Auditability, and the Algorithm Register

Openness of governance forms the cornerstone of the democratic rule of law. Dutch government bodies are required to include the algorithms they use in the national algorithm register to give citizens insight into decision-making structures. The publication on disclosures in the algorithm register demonstrates that registrations featuring proprietary black-box models almost always run into incompleteness due to a lack of system information from the supplier.

In an administrative context, transparency means that a regulator, judge, or citizen must be able to trace the logic and parameters by which a model generates certain summaries or categorizes policy texts. If a government body uses a closed API, the responsible minister or municipal council cannot account for any hidden system prompts or unknown weights in the network.

Property Commercial Closed API Own / Sovereign Model
Training data insight Not accessible (trade secret) Fully auditable and reproducible
Control over model drift Opaque (changes with provider updates) Fully static and deterministically manageable
Algorithm register conformity Often insufficient due to missing specifications Meets all openness requirements and metadata standards
Data processing and perimeter External clouds (often multi-tenant) Within own sovereign government server
Archiving and reconstruction Uncertain in case of model deprecation by supplier Permanently reconstructable in accordance with the Archives Act

The use of sovereign models enables institutions to transparently publish the full technical passport — including architecture choices, tokenizer parameters, and validation results. This makes automated support verifiable for society once again.

5. Linguistic Accuracy and Administrative Jargon

Globally, the Dutch language accounts for only a fraction of the training data on which large commercial base models are pretrained. As explained in the analysis on the position of the Dutch language in AI, this skewed ratio leads in practice to creeping style errors, anglicisms, and a fundamental lack of understanding of institutional contexts.

In an administrative context, word choice matters exceptionally precisely. Legal terms such as 'last onder dwangsom' (order subject to a penalty payment), 'bestuurlijke boete' (administrative fine), 'beginselplicht tot handhaving' (principle obligation to enforce), 'voorlopige voorziening' (interim measure), or 'passend onderwijs' (appropriate education) have a clearly defined meaning under the General Administrative Law Act (Awb) or sector-specific laws. Generic commercial models regularly confuse these concepts with Anglo-Saxon legal terms (such as 'injunction' or 'contempt of court'), which can cause letters to citizens to contain legally incorrect commitments or intimidating phrasing.

Through targeted pre-training or fine-tuning of open-source base models on representative Dutch-language government corpora — such as public parliamentary papers (Kamerstukken), official Staatscourant publications, anonymized rulings from the Council of State (Raad van State), and municipal ordinances — a language model emerges that does understand administrative nuance. Such a model formulates letters in accordance with the applicable standards for plain government language (such as language level B1) without sacrificing legal precision.

6. Procurement Conditions, Tenders, and Cost Structures

The procurement of software and IT services by public bodies is bound by European and national tendering rules. Purchasing commercial cloud AI proves in practice to be exceptionally difficult due to strict requirements around interoperability and vendor independence. The guide on requirements in AI tenders emphasizes how contracting authorities must prevent vendor lock-in by explicitly steering toward open standards and data portability.

Commercial API models generally use a pricing model based on token consumption. For small-scale pilots, the initial costs are low, but when scaling up to thousands of employees who search extensive case files daily, variable costs become volatile and budgetarily unmanageable. In contrast, sovereign infrastructure — although it comes with initial investments in hardware and implementation — ensures predictable, fixed operational costs that do not explode under intensive use.

In addition, an open architecture creates a healthy separation between the model supplier and the hosting party. If a particular open-source language model is eventually surpassed by a more efficient or more accurate alternative, the government can replace the model weights within the existing application without having to re-tender the entire integration layer.

7. Typical Architecture of a Sovereign Government Pipeline

In practice, public institutions build so-called Retrieval-Augmented Generation (RAG) systems around local open language models. In such a configuration, the model does not generate facts from its own memory, but reasons exclusively over verified policy documents that are supplied in real time from a secure source.

# Architectuur van een lokale en soevereine overheids-RAG-omgeving
[Beveiligde Opslag: Dossiers & Wetgeving]
             |
             v
[Lokaal Embeddingmodel (bv. Gecureerd NL-model)]
             |
             v
[Lokale Vectordatabase binnen Overheidsperimeter]
             |
[Ambtelijke Vraag / Burgerdossier] --> [Context Selectie & Anonimisering]
                                                  |
                                                  v
                                     [Lokaal Open Taalmodel]
                                     (Gehost op overheidsserver)
                                                  |
                                                  v
                                     [Geverifieerd Advies / Conceptbesluit]

Within this architecture, the entire data flow remains hermetically sealed off from the public internet. The embedding model converts official texts into vector representations within the organization's own network, after which the local language model formulates only context-driven answers with direct source references to the underlying laws and regulations.

8. Measurement Methods: Quantifying Performance and Safety

The decision to put a proprietary model into production requires objective measurement methods. Public institutions cannot rely on subjective impressions or general marketing benchmarks from model makers. In administrative practice, three specific evaluation domains are therefore quantified:

1. Language-specific legal benchmarks: Models are tested on representative datasets of administrative cases. This measures how accurately the model cites legislation, whether it correctly assigns powers to government bodies, and whether it handles formal legal terminology flawlessly. The metric focuses on factual precision and ruling out hallucinations in legal qualifications.

2. Comprehensibility and language level (B1 testing): For public communication, the generated text is tested automatically and manually against readability indices (such as the Flesch-Kincaid score and the Cito readability index for Dutch). The percentage of sentences that meet the guidelines for plain government language (short sentences, active constructions, avoiding unnecessary administrative jargon) is measured.

3. Safety and leak resistance (Prompt Injection & Jailbreaking): A model deployed in a front-desk function must be resistant to manipulation by bad actors. Automated 'red teaming' scripts are used to measure how resistant the system is to indirect prompt injection (for example, malicious instructions hidden in uploaded PDF applications) and whether the model consistently refuses to disclose confidential system prompts or unauthorized source data.

9. Limitations, Edge Cases, and Real Costs

Despite the compelling advantages, deploying proprietary language models brings significant operational and financial challenges. Ignoring these limitations inevitably leads to failed IT projects and write-offs. Government bodies must take the following constraints into account:

The costs of computing power and management: Running local models requires specialized graphics processors (GPUs). The purchase, energy consumption, cooling, and daily management of such clusters entail significant fixed costs. For a small rural municipality with 25,000 residents, it is not financially viable to set up its own dedicated GPU cluster. This creates an urgent need for shared services, in which provinces, umbrella organizations such as the Vereniging van Nederlandse Gemeenten (VNG), or nationwide shared service centers collectively host the infrastructure.

The quality gap in abstract reasoning tasks: Although specialized and compact open models excel at domain-specific tasks such as summarizing, classifying, and extraction, the very largest commercial cloud models remain superior at complex, multi-step reasoning tasks and broad general knowledge. A local model with 8 to 70 billion parameters can struggle with extremely complex, non-standard cases in which dozens of policy areas interfere with one another. In such situations, strict human supervision is indispensable.

Maintenance and data currency: A language model is a snapshot in time. When legislation changes (such as with the introduction of a new Omgevingswet, Environment and Planning Act), the static model knows nothing about it on its own. Keeping a sovereign system up to date therefore requires a robust and continuous document management and RAG pipeline. Building this internal technical expertise is, in today's tight labor market, one of the biggest bottlenecks for the public sector.

Conclusion

The movement of Dutch public services toward their own, sovereign language models marks a coming of age for digital government. The realization has set in that public values such as privacy, equal treatment, transparency, and democratic accountability cannot be outsourced to external, commercial black boxes outside European jurisdiction.

By investing in open-source architectures, high-quality Dutch-language government corpora, and a secure, shared computing infrastructure, the government is laying a sustainable foundation for reliable digital services. In doing so, it not only ensures that administrative processes become more efficient, but above all that citizens can continue to trust a government that keeps control over its own decision-making and data firmly in its own hands at all times.