Publication date 23/07/2026
 Imagen abstracta de la recopilación de datos procedentes de múltiples fuentes a través de la red.
Description

Let's think for a moment about how any public administration works. Every day, files are processed, technical reports are written, minutes of meetings and collegiate bodies are drawn up, thousands of emails are exchanged, content is published on the electronic headquarters, paper documents are digitized and images, audio recordings and videos of all kinds are generated. All this documentary production contains valuable information about how the organization works, what decisions it makes and why it makes them. And yet, most of that information remains off the radar of traditional information systems.

When we talk about the data of an administration we usually think of databases, spreadsheets, registers, budgets or indicators. It makes sense: for decades, data management has focused on this type of structured information. However, that vision only shows a small part of reality.

Much of the knowledge generated by administrations does not reside in their databases, but in files, reports, minutes, resolutions, emails and multimedia content that are part of their daily activity. Information that exists, that contains enormous value and that, traditionally, has remained on the margins of data governance strategies.

The emergence of artificial intelligence has turned this documentary heritage into an unprecedented opportunity, but it has also shown that technology, alone, is not enough. In this article, we will look at why unstructured data has become one of the main assets of public administrations, what obstacles prevent it from being exploited to its full potential, and how a data governance strategy can transform this enormous volume of information into a reliable, reusable resource that is ready to generate value.

The data we don't see

The best way to understand this situation is to imagine an iceberg. The visible part represents the structured data on which most corporate applications, dashboards, and official statistics work. Below the much more extensive  waterline is the enormous volume of unstructured information describing decisions, procedures, technical knowledge, and administrative context.

According to estimates by analyst firms such as IDG or Gartner, this type of content represents around 80% of the information handled by an organization, and everything indicates that this percentage will continue to grow.

 

Visual explanation of the iceberg of public administration data. Structured data accounts for 20 per cent of the total data, whilst unstructured data accounts for 80 per cent.

Figure 1. Visual explanation of the type of data of a public administration. Source: own elaboration – datos.gob.es

The paradox is obvious: most of the knowledge of an administration is not in its databases, but submerged in its documents. For many years that part of the iceberg could hardly be used. Today, however, the situation has changed radically.

The latent value that is being missed

Ignoring the submerged part of the iceberg means wasting one of the largest information assets of the public sector. These contents document the experience accumulated over the years, the knowledge of public employees, the relationships between files and a good part of the context that is never stored in a database. Its use has, therefore, a direct and transversal impact on public management.

When that information can be located, understood and related, a file ceases to be merely a set of documents and becomes a source of reusable knowledge. The journey is always the same: documents contain institutional knowledge in the form of context, decisions, criteria and evidence, and it is the governance of data that allows this knowledge to be transformed into public value. This translates into better services for citizens, greater traceability of administrative activity, more transparency, greater internal efficiency and better reuse of institutional knowledge.

This enhancement also connects with a very specific legal obligation: the "one-time" principle, set out in Article 28.2 of Law 39/2015, which recognises the right of citizens not to provide documents that are already in the possession of any administration. Giving effect to this right requires that the information that the administration already possesses, often in the form of documents, can be located, interpreted and exchanged between agencies: every time a citizen is asked for data that already appears in a file, the problem is not normative, but one of management and interoperability of the information.

But there is one factor that has made this need a strategic priority: artificial intelligence. These contents are today the main raw material for AI applications, especially those based on natural language processing (NLP) and Large Language Models (LLMs).

Harnessing unstructured information, however, did not begin with generative artificial intelligence. For years, techniques such as optical character recognition, ETL (extract, transform, load) processes, web scraping, regular expressions, taxonomies, rule engines, and classic natural language processing techniques have been used to extract, classify, standardize and transfer documentary information to exploitable structures. These approaches continue to be especially effective when sources are relatively homogeneous, patterns are stable, and extraction rules can be precisely defined. Artificial intelligence does not necessarily replace these techniques, but rather expands their scope and allows more variable, ambiguous or difficult-to-process content to be addressed through previously defined rules.

Until a few years ago, much of this information could be exploited using traditional techniques, but doing so required very specific processes, dependent on the format and difficult to scale or maintain when document diversity and complexity increased. Documents were generated and stored in a wide variety of systems and formats, their localization often depended on keywords or folder paths, and understanding content that could not be processed by rules required time-consuming and costly manual reading. The emergence of generative artificial intelligence has completely changed this scenario: today it is possible to summarise documents, classify files, extract relevant entities, detect relationships between documents or answer questions about regulations automatically. AI has not created the unstructured data; it has simply made it possible to tap into a documentary heritage that has been waiting for decades for its opportunity.

The iceberg, therefore, is no longer just a metaphor for what we don't see: it's a fairly accurate description of where the value we are not yet capturing is.

However, there is a common misconception: thinking that having artificial intelligence models is enough to take advantage of all that knowledge. However, it is not.

The obstacle is not technological, it is data governance

The technology is already here: tools capable of processing documents and applying artificial intelligence are becoming more and more accessible and evolving at a dizzying pace. However, the real challenge is not to incorporate new algorithms, but to properly govern the information on which they work.

A language model can summarize thousands of files in minutes, but it cannot determine what the valid version of a document is, whether its contents are still valid, or who is responsible for keeping it up to date. You can only work with the information you receive. If that information is incomplete, inconsistent, or lacks context, your answers will inherit those same limitations.

An infographic explaining artificial intelligence with and without data governance


 Figure 2. Visual explanation about artificial intelligence without data governance vs. data governance. Source: own elaboration – datos.gob.es

Without metadata, cataloguing, classification, quality criteria, and clear responsibilities, unstructured data ceases to be an asset and becomes an organizational liability: abundant, expensive to maintain, and difficult to locate, interpret, and reuse. Applying artificial intelligence to poorly governed documentation does not generate knowledge; it generates apparently plausible answers built on unreliable information, probably the worst possible scenario for a public administration.

This situation leads us directly to a concept that we already discussed in the article From the Swamp to the Lake: how to prevent your data from becoming a swamp. The accumulation of information without government ends up producing a data swamp, because accumulating information is not equivalent to generating knowledge. If this is already true for structured data, it is even more true for document repositories, where volume grows faster and context is lost sooner. Without data governance, the organization acts as a simple digital storage room; With data governance, that content is transformed into a useful, reliable asset that is ready for exploitation, both by people and by artificial intelligence.

The natural evolution of data

The governance of data fulfills another less obvious and yet fundamental function here: to help identify when information that was born as unstructured should cease to be so.

Many administrations continue to store certain data in text documents, PDF forms, or comment fields simply because that is how the original process was designed. With the passage of time, this information begins to be repeated in thousands of files and ceases to be an exception and becomes a stable pattern.

An infographic explaining how to move from documents to data: a decision by the data-driven government.

Figure 3. Explanatory visual about the passage from document to data. Source: own elaboration – datos.gob.es

Not all unstructured data should remain unstructured. One of the functions of data governance is precisely to identify when the time has come to structure it. Governing data means deciding which information should retain its documentary richness and which should be transformed into structured data to facilitate its validation, interoperability, exploitation and reuse. It is a decision of information design, not an inertial consequence of how things began to be done decades ago.

How to activate that value: a governance roadmap

Taking advantage of this documentary heritage does not require starting by implementing artificial intelligence. It requires first building a solid basis for data governance. The principles are the same as those we already applied to structured data, now extended to all the organization's information.

Pillar

Objective

Policies and Responsibilities Define who is responsible for each type of content and under what rules it is created, modified, shared, and deleted.
Metadata Describe documents so that they can be automatically located, understood, and related.
Classification Organize information by taxonomies, document typologies, and levels of sensitivity.
Quality Ensure that information is complete, up-to-date, free of duplication, and ready for reuse.
Interoperability To facilitate the exchange of documents, files and systems through common standards.
Life cycle Manage information from its creation to its archiving or elimination, applying homogeneous criteria throughout the process.

Figure 4. Table showing the pillars and objectives of the governance roadmap. Source: own elaboration - datos.gob.es

As a methodological reference to follow this path, Spain has the ecosystem of UNE standards on data governance, management and quality (UNE 0077, UNE 0078, UNE 0079, UNE 0080 and UNE 0081). This framework allows for a homogeneous approach to the management of both structured and unstructured data, relying on processes, responsibilities and continuous improvement to turn information into a governed and measurable asset.

Beyond storage

For years, public administrations have made an enormous effort to digitize documents and files. This process has made it possible to replace paper with electronic files, but digitizing does not always mean better management of information.

The real challenge in the coming years is not to store more documents, but to turn that immense documentary heritage into a governed, reusable asset ready to generate value: improving public services, strengthening transparency and providing a reliable basis for the artificial intelligence applications that are already transforming public management.

And it is important not to lose sight of the fact that the iceberg will continue to grow. The documentary production of the administrations will continue to increase, and with it the volume of knowledge that remains under the surface. The difference between organizations that turn this hidden mass into an advantage and those that suffer it as a burden will not be in the technology they use, but in how they govern their information.

Because, as in any iceberg, the greatest value is not in the visible part. It is under the surface, waiting for the administrations to develop the necessary capacities to discover it, understand it and put it at the service of citizens.

Content produced by Dr Fernando Gualo, Lecturer at the University of Castilla-La Mancha (UCLM) and Consultant on Governance and Data Quality. The content and views expressed in this publication are the sole responsibility of the author.

Comments