Skip to content
NLEN
Illustration: Accessibility and AI Technology: Opportunities and Barriers

Accessibility and AI Technology: Two Sides of the Same Coin

By Ivo Donker — compiled with AI support (Claude & Gemini) · Last updated: 6 August 2026

The rise of artificial intelligence often provokes predominantly optimistic or, conversely, pessimistic reactions in public debate. Within the domain of digital accessibility, this divide is strikingly visible, yet the two sides are rarely discussed in close connection with each other. On the one hand, technology based on advanced algorithms offers powerful tools to make information accessible to a broad group of users. On the other hand, these same technologies create entirely new, unforeseen barriers in digital interfaces and information flows. Anyone wanting to take stock must therefore pay attention to both the problem-solving capabilities and the exclusionary mechanisms of modern systems.

This tension touches the foundations of how we shape, consume, and share information. The influence of these systems extends across various sectors, as is also visible in broader societal developments surrounding AI in education, where the accessibility of learning materials directly affects students' equal opportunities. To understand how this dynamic works, it is necessary to systematically break down the concrete applications and the inherent limitations of the underlying techniques.

Concrete Applications and Their Functional Value

To see what technology concretely contributes to the accessibility of digital media, it helps to look at the specific needs per target group. The functionalities offered by modern systems range from visual support to textual simplification and auditory processing.

For people who are visually impaired or blind, automatic image descriptions offer a way to understand visual content. Systems generate textual alternatives for images, charts, and photos. This allows screen readers to read the content aloud, meaning visual information is no longer necessarily lost for users who depend on auditory or tactile output. In addition, multimodal models play an increasingly important role in processing text, image, and audio simultaneously, which facilitates the translation between different media forms.

For people who are hard of hearing or deaf, automatic subtitling and summaries of spoken audio or video are crucial tools. These applications convert spoken language into written text, often in real time. This significantly lowers the barrier to participating in online meetings, lectures, or video content. It enables users to quickly grasp the core of an argument without depending on the quality of the original audio track.

A third important category concerns simplifying complex texts for people who struggle with language level, such as those with low literacy or a cognitive disability. This involves shortening long sentences, converting difficult concepts into everyday language, and making the structure clearer. This closely relates to the broader discussion about how complex language use in official documents can be reduced to understandable proportions, a theme that also plays a role within the AI in Dutch government: the current state of affairs.

Speech Recognition, Speech Synthesis, and the Dutch Language Area

Speech recognition and speech synthesis are among the oldest applications within the field of artificial intelligence. A computer's ability to convert spoken language into text, and vice versa, has for decades been an important tool for hands-free operation and auditory support.

The quality of these systems in Dutch, however, has a story of its own. Because the Dutch-language area is relatively small compared to English, the amount of available training data for models is often limited. In practice, this leads to specific challenges:

Improving these aspects requires targeted attention to language diversity and the collection of representative datasets, as further explained in analyses on speech models and the specific characteristics of language diversity in LLMs for Dutch.

The Limits of Automatic Image Descriptions

The generation of image descriptions by algorithms shows how technological progress and fundamental limitations go hand in hand. Automatically generated descriptions are almost always an improvement over a complete absence of alternative text. Still, they cannot replace a human description written with knowledge of the context.

The fundamental problem lies in the difference between what is factually visible in an image and what that image means in its specific context. An automatic system can determine that 'two people at a table' are visible. However, it lacks the context of whether this is an informal meeting, a job interview, or an artistic work. Such nuances are precisely what determine the informational value of the image.

In addition, an incorrect image description carries a specific risk: it can be more harmful than no description at all. When a system misinterprets a crucial element in a chart or photo and the user has no other way to verify the content, a false sense of certainty arises. The user builds on incorrect information in the belief that digital accessibility has been correctly ensured.

Subtitling and the Vulnerability of the Margins

In automatic subtitling and transcription, the focus is often on the average accuracy score across an entire document or video. Although a high average percentage looks impressive, this figure says nothing about where the errors occur.

In practice, errors in automatic speech conversion tend to concentrate precisely on the most critical parts of a message: personal names, specific technical terms, unique locations, and unusual accents or speaking speeds. This means the errors pile up at exactly the moments when the listener or reader needs the information most to understand the core of the story. A high average accuracy therefore offers no guarantee of an accessible user experience at the places where it truly matters.

Text Simplification Versus Substantive Accuracy

Automatically simplifying complex texts helps make information more accessible to a broader audience. However, this brings a subtle tension in terms of legal and factual accuracy.

In official communications, such as regulations, legislation, or policy documents, meaning often depends on precisely formulated nuances. When an algorithm rewrites a text to make it simpler, there is a risk that subtle conditions, exceptions, or legal implications get lost or shift in meaning. What appears clear in simple language may take on an entirely different legal weight. This requires a constant balancing act between comprehensibility and the absolute accuracy of the source information.

New Barriers Through Generative Interfaces

While existing applications try to remove old barriers, modern generative systems introduce entirely new types of obstacles. The shift toward interfaces that rely entirely on free-text input or voice control changes the way users interact with systems.

An important aspect of this is the dynamic nature of the output. Where traditional websites work with fixed structures, predictable navigation menus, and fixed headings that a screen reader can step through, modern chat interfaces generate unique responses each time. This makes it difficult for users who depend on keyboard navigation or fixed operating patterns to predict and control the interaction.

Time limits also play a role. Conversational interfaces often expect a quick response within a certain context, which increases the pressure on users who need more time to process information or formulate input. Moreover, if a chat window is the sole entry point to a service, this turns from a convenient preferred option into an insurmountable accessibility problem. The absence of a traditional, structured menu excludes users who cannot cope with open-ended dialogues.

The Relationship to Traditional Guidelines and Checks

Existing accessibility guidelines rest on a number of fixed pillars: information must be perceivable, interfaces must be operable, and the operation must be understandable and robust. Generative interfaces affect all three of these pillars in a fundamental way:

Automated accessibility checks, which are often used to scan websites for errors, can only identify part of this problem. They check technical characteristics such as color contrast and element structure, but miss the semantic depth of dynamic AI interactions. This is why practical testing with the people concerned — users who live with these barriers in daily practice — remains completely irreplaceable.

Practical Implications for Organizations

Organizations responsible for digital services and content face the task of deploying this technology thoughtfully. This requires a sober approach in which the strengths of automation are used without losing sight of the risks:

Phase Role of AI Technology Human Responsibility
Starting Point Generating an initial draft for transcripts, alternative texts, or summaries Checking for factual accuracy and context
Application Offering tools for translation and simplification Validating whether the core meaning is preserved
Design Providing flexible input options in interfaces Always keeping an alternative route open without AI compulsion

Applying these principles means that automated tools are seen as a valuable starting point, and never as a self-sufficient endpoint. By always keeping an accessible alternative — such as a traditional form, a fixed menu structure, or a human contact route — open, technological innovation is prevented from leading to exclusion.

Further reading