Making an organization more secure: collect less data
For years, organizations have been accumulating data. Customer, employee, and supplier data, transactions, communications, digital behaviors, equipment, movements, system access, videos, technical logs, internal documents: digital transformation has made it possible to retain a quantity of information that was once unimaginable.
The arrival of artificial intelligence reinforces this trend further. Since data can feed analyses, automate processes, train models, and reveal new knowledge, it can seem logical to keep as much of it as possible.
“It might be useful one day” then becomes an almost natural justification.
Yet retained data has another characteristic: it has to be protected.
Every new piece of information collected potentially increases the organization’s responsibilities, the systems to secure, the access to control, the regulatory obligations to meet, and the possible consequences of a cyberattack.
The strategic question is therefore no longer only: “What data can we collect?” It becomes: “What data is truly worth taking on the responsibility of holding?” This distinction could become one of the important principles of modern cybersecurity.
Data is an asset, but also a liability
We frequently hear that data is one of an organization’s most valuable assets. That is true. It makes it possible to understand customers, improve operations, optimize decisions, automate certain tasks, and develop new products and services. However, not all information assets have the same value.
An organization may retain millions of files, years of emails, historical databases, multiple backups, system logs, copies of documents, and old customer data without really knowing what remains useful.
Over the years, this accumulation creates a form of information debt. Gradually, no one has a perfectly clear view of all the information the organization holds:
- data copied
- systems replaced
- employees gone
- suppliers changed
- applications abandoned
- backups persist
- access sometimes stays active longer than necessary
This situation is a governance challenge. It is also a cybersecurity challenge.
You cannot properly protect what you do not know about
To protect a piece of information, you first have to know that it exists. You then have to know where it is located, its value, its level of sensitivity, the people and systems that can access it, the applications that use it, the places where it is copied, and how long it must be retained.
In a modern technology environment, this mapping can become extremely complex. Data circulates between internal systems, cloud platforms, SaaS software, mobile devices, APIs, business partners, development environments, collaboration tools, and now artificial intelligence systems. The same piece of information can therefore exist simultaneously in several environments. This multiplication mechanically increases the number of places an organization has to secure.
The first step in a data protection strategy is therefore not necessarily to add a new technology layer. It can begin with a far more fundamental question: Do we really need to keep everything we have?
Every unnecessary piece of data increases the risk surface
Data that no longer exists in an organization’s systems can no longer be exfiltrated during a cyberattack. This obvious fact has important consequences.
Cybersecurity strategies naturally focus on firewalls, detection systems, encryption, identity management, segmentation, backups, and incident response mechanisms. These technologies remain essential. However, reducing the quantity of information exposed is also a form of risk reduction.
When an organization eliminates data that has become useless, limits copies, reduces access privileges, and shortens certain retention periods, it simultaneously reduces what an attacker could obtain. Data minimization then becomes a cybersecurity mechanism.
This logic can be applied to personal information, business information, internal documents, financial data, trade secrets, and many other categories of digital assets.
Artificial intelligence changes the value of historical data
AI does, however, introduce a paradox. At the very moment when organizations should be exercising better control over their data, that data becomes potentially more valuable. Archives considered of little use a few years ago can today feed artificial intelligence systems, help uncover trends, or contribute to building a particularly rich body of organizational knowledge. Systematically deleting old data would therefore be just as imperfect an approach as keeping everything.
The real question concerns value. A mature organization must be able to distinguish data that constitutes a strategic information heritage from data that primarily represents a liability. This distinction requires knowledge governance. You have to understand what the organization holds, why it holds it, what must be retained, what can be anonymized, what can be archived, and what should be eliminated. In the age of AI, this discipline becomes even more important.
Artificial intelligence does not automatically turn an accumulation of files into usable knowledge. The quality, structure, provenance, usage rights, reliability, and context of data largely determine its value. Accumulating information and building a body of knowledge are two very different things.
The proliferation of copies deserves particular attention
A piece of data can be perfectly protected in its system of origin and become vulnerable when it is copied elsewhere:
- A file extracted to produce a report.
- A database copied into a test environment.
- A customer list exported into a spreadsheet.
- A document downloaded onto a personal computer.
- A backup kept in a different environment.
- Information transferred to a SaaS application.
- A document submitted to an artificial intelligence tool.
Each of these operations can create a new instance of the data. As copies multiply, control becomes more difficult. That is why modern governance must concern itself not only with official systems, but also with the movement of information. Knowing where data resides is important. Understanding how it circulates is even more so.
Zero Trust should also apply to data
The Zero Trust principle rests in particular on the idea that access should never be granted simply because a user or a system is inside a perimeter considered secure. Every access must be justified according to context, identity, privileges, and genuine need. This logic can be extended all the way to the data itself:
- Why does this person need access to this information?
- Why must this application receive the entire record when it only uses a few fields?
- Why must this information be copied?
- Why must it be retained for ten years?
- Why must an artificial intelligence system access an entire document repository when a fraction of it is enough to complete its task?
Reducing privileges and minimizing data ultimately pursue a similar objective: limiting exposure to what is genuinely necessary.
Zero-Knowledge takes this logic even further
Some architectures even make it possible to demonstrate that a piece of information is true without revealing the information itself. This approach is particularly interesting for digital identity, authentication, and environments where sensitive information has to be validated.
Depending on the use case, an organization might need to know that a person holds an authorization without necessarily knowing all the information used to establish it. It might need to confirm an identity without multiplying copies of data that would allow that identity to be reconstructed. It might have to verify a condition without retaining all the information used for that verification. This evolution is gradually changing how we think about security.
For a long time, holding more information could seem to offer more control. Modern architectures sometimes make it possible to achieve more trust while holding less sensitive information.
Less data can also mean lower costs
The question in fact goes beyond cybersecurity. Retaining data entails costs. Storage, backups, replication, availability, encryption, monitoring, administration, compliance, audits, classification, and recovery all have to be considered. Individually, the cost of a few extra gigabytes may seem insignificant.
At the scale of an organization and over several years, however, the accumulation can become considerable. Above all, the real cost does not lie solely in storage. It lies in complexity. The more data an organization holds scattered across different environments, the more resources it must devote to understanding, administering, and securing its digital estate. Reducing this complexity can therefore improve security, governance, and operational efficiency at the same time.
Knowing what you hold becomes a strategic advantage
An organization’s digital maturity could ultimately be measured less by the quantity of data it holds than by its ability to understand that data’s value.
- Which data genuinely supports operations?
- Which constitutes strategic knowledge?
- Which must be protected as a priority?
- Which is necessary for artificial intelligence?
- Which is subject to specific obligations?
- Which can be anonymized?
- Which should simply disappear?
An organization able to answer these questions holds a considerable advantage. It can better protect its critical assets, prepare its data for artificial intelligence, reduce its risks, and invest its cybersecurity resources where they produce the most value.
For several decades, the digital growth of organizations has been accompanied by an almost natural accumulation of data. Artificial intelligence now obliges us to reconsider that approach.
Data can become a tremendous source of knowledge and value. It can also represent a significant liability when it is poorly understood, needlessly copied, insufficiently governed, or retained for no specific reason.
The challenge is therefore to develop genuine organizational intelligence about data: knowing what you hold, understanding its value, knowing its dependencies, and determining the level of protection it deserves.
This approach draws on several disciplines: information governance, cybersecurity, identity and access management, technology architecture, Zero Trust, Zero-Knowledge, digital sovereignty, AI governance, and risk management.
At Quantum Beyond, our experts work with the teams already responsible for these environments in order to bring them complementary capabilities and expertise. The goal is to help them better understand their information estate, identify unnecessary exposures, strengthen their architectures, and prepare their data for new digital uses and artificial intelligence.
Internal teams remain at the heart of this transformation. They know the operations, the systems, and the reality of their organization. By giving them better visibility, suitable methods, and specialized expertise, they can make better decisions and strengthen their environment for the long term.
The most secure data is not systematically the data that benefits from the most sophisticated encryption. Sometimes it is the data the organization had the intelligence not to keep.
