When law enforcement and intelligence agencies discuss critical and often sensitive data, the conversation usually begins with familiar questions: Is it secure? Where is it stored? Does it meet applicable compliance requirements? Who can access it, and can those access controls support the mission?
Those are essential questions. But another question is too often overlooked. It usually surfaces only after an agency has experienced a cyber incident, major outage, difficult vendor relationship, or urgent need to move data to another system: Who really controls the data?
The issue is not new. More than a decade ago, while serving in government, I encountered a situation in which there needed to be answers to difficult and highly specific questions around data retention and deletion.
The challenge was not simply determining whether information had been deleted from a primary system. It was understanding where else it existed: in backups, replicated environments, logs, archives, partner systems, vendor-managed infrastructure, and other parts of the broader data lifecycle.
Even then, it was clear data rarely exists in only one place, and questions of who has accessed it, who administers it, who can retain it, and who can truly delete it are often more complicated than expected.
Fifteen years later, cloud computing, software-as-a-service platforms, advanced analytics, and artificial intelligence (AI) have made those questions significantly more complex at both the customer and vendor levels.
This complexity does not eliminate accountability; it makes clear, shared understanding more important. Agencies must understand how their information moves through the full technology ecosystem. Vendors must explain those processes clearly, identify the limits of their control, and support customers’ legal, operational, and mission requirements.
Data sovereignty cannot be treated as a customer concern or a vendor obligation alone. It must be understood, designed for, and managed by both.
What Is the Difference Between Data Residency and Data Sovereignty?
Data residency is the physical location where data is stored, such as whether a cloud service keeps information in the United States, the European Union, or another region.
Data sovereignty is an agency’s ability to control its information throughout its lifecycle: where that information is processed, who can access or administer it, which legal jurisdictions apply, how long it is retained, whether it can be reused, how it is deleted, and whether the agency keeps control when conditions change. Residency is only one input to sovereignty.
A modern data environment is not a single repository. Information can move among applications, cloud platforms, identity systems, networks, backup environments, cybersecurity tools, support organizations, analytics platforms, and AI services.
Even if a primary database resides in a single country, the personnel, administrative systems, encryption keys, telemetry, software-update processes, subcontractors, or legal authorities with access to the environment may fall outside the agency’s intended trust boundary.
Effective data governance also extends beyond source files to metadata, audit logs, permissions, configurations, backup copies, support records, security telemetry, usage history, and records generated through workflow automation.
The scope expands further when AI is involved, making AI data sovereignty a distinct consideration. In addition to source records, agencies must account for prompts, outputs, retrieval activity, logs, embeddings, indexes, configurations, and other artifacts created as AI systems process agency information.
Those artifacts may remain in the environment even after an original document has been deleted. More on this later.
How Does Vendor Lock-In Affect Data Sovereignty?
Dependence on outside technology providers is not necessarily a problem. Agencies cannot and should not build every technology capability themselves. Commercial providers bring scale, security investment, and specialized expertise that agencies may not be able to replicate internally.
Even so, agencies should treat vendor lock-in as an enterprise-level risk, not simply a procurement concern. Four risks are particularly important:
- Mission Continuity
A provider may experience a major outage, cyber incident, acquisition, financial failure, loss of authorization, product discontinuation, or a change in service that no longer meets agency requirements. If the agency depends heavily on that provider for mission-critical operations, even a short disruption can affect employees, partners, and mission outcomes. - Financial Control
Once a platform becomes central to operations, leaving can cost more than staying. A provider may increase subscription costs, alter licensing terms, move to a usage-based pricing model, bundle formerly included features into paid products, or require the purchase of additional components to retain the same level of service. The agency may have little negotiating leverage if it has no usable alternative. - Data Portability
Data may be difficult to move because it exists in proprietary formats or is tied to the vendor’s workflow logic, permissions, audit history, records classifications, integrations, and reporting tools. An agency that can export raw documents but cannot recreate the business process, verify the records, or preserve associated metadata does not have true portability. - Agency Flexibility
Agencies must be able to respond to changing mission requirements, emerging cybersecurity threats, legal obligations, and technology. Lock-in can limit their ability to improve security, adopt better tools, or change direction.
The goal is not to avoid single-vendor relationships. In many cases, a single provider may be the most efficient and secure option. The goal is to ensure that the agency enters those relationships with its eyes open, retains meaningful control, and has a credible path to continue operations if the provider no longer meets agency needs.
What Should Agencies Require from Technology Vendors?
The contract is where data sovereignty is secured or lost. It should clearly establish agency ownership and control of source data, records, metadata, audit logs, configurations, workflow artifacts, prompts, outputs, embeddings, and other information generated through use of the platform.
Vendors should provide usable exports in documented, nonproprietary, or widely interoperable formats whenever feasible. The export should include not only raw files, but also the metadata, classifications, retention labels, permissions, audit history, configurations, and workflow information needed to make the material usable in another environment.
Contracts should define an exit process, including transition assistance, technical documentation, data export support, knowledgeable personnel, and cooperation with a successor provider. That process should be established before the service becomes deeply embedded.
Deletion and retention obligations are equally important. At the conclusion of a contract, the provider should certify that agency data has been deleted from active systems, development environments, backups, and applicable subcontractor systems, subject only to legally required retention obligations. NIST’s Guidelines for Media Sanitization provides a useful technical reference for what secure deletion actually requires.
The agency should be able to place records under legal hold when needed and retain the ability to preserve data for audit, investigation, discovery, and records-management purposes.
Agencies need visibility into privileged access by provider personnel and subcontractors. Access should be authorized, time-bound, logged, and monitored. For sensitive environments, the agency should require approval or notification before provider personnel access agency data or systems.
Agencies should also require prompt notification, to the extent permitted by law, if a provider receives a subpoena, court order, regulatory request, law-enforcement inquiry, or other governmental demand for agency data. The provider should disclose only what is legally required, notify the agency whenever possible, preserve relevant records, and cooperate with the agency’s legal and records-management responsibilities.
Finally, encryption remains an essential safeguard. Wherever possible, data should be encrypted both in transit and at rest. For sensitive workloads, agencies should evaluate customer-managed keys or stronger control over key generation, storage, rotation, revocation, and emergency access. Key control reduces sovereignty risks but does not eliminate them completely.
How Does AI Change Data Sovereignty?
AI can move from experiment to an operational dependency quickly. An agency might begin by using AI to improve writing or summarize public material. It may soon expand AI use into internal research, knowledge management, document review, case support, workflow automation, and decision support.
As AI becomes embedded in operations, staff develop prompt libraries, specialized assistants, evaluation datasets, retrieval systems, and workflows around one provider’s model, APIs, and tools. Over time, the agency’s knowledge and operating processes may become tied to a specific vendor environment.
Agencies should therefore avoid allowing authoritative source information to become permanently embedded in a proprietary AI platform. Records and authoritative knowledge bases should remain in agency-controlled repositories whenever possible.
The AI system should retrieve only the information necessary for a defined task and apply the same access controls that govern the source records. An AI assistant should not provide a user with information the user would not be authorized to access directly.
AI also creates a specific data-provenance concern. Agencies need to know what source materials informed an AI response, whether those materials were current and authoritative, which model and configuration were used, and who reviewed the result.
Without that information, an agency may have difficulty explaining or defending an AI-supported recommendation, analysis, or decision. NIST’s Generative AI Profile identifies information integrity, privacy, cybersecurity, intellectual-property, and third-party value-chain risks as key areas to manage throughout the generative-AI lifecycle.
Unless expressly approved in writing by the agency, providers should be prohibited from using agency prompts, uploaded content, outputs, feedback, telemetry, or derivative artifacts for generalized model training, fine-tuning, benchmarking, or product improvement.
Any authorization should define exactly what data may be used, for what purpose, where it may be processed, how long it may be retained, and what rights the agency retains in the resulting customized model, configuration, or artifact.
How Can Agencies Adopt AI Without Losing Data Control?
A closed-system-first model uses controlled environments for sensitive work while building portability into the design. Approved external services can still serve lower-risk uses under defined controls.
Before approving a technology or AI use case, agency leaders should ask three questions:
- What type of information will be processed?
- What are the consequences if that information is exposed, altered, unavailable, or retained improperly?
- How dependent will the agency become on the service to perform its mission?
The answers determine the appropriate level of containment.
| Information type | Environment | Level of Agency Control |
|---|---|---|
| Public, non-sensitive, low-impact | Approved external AI service, through a managed enterprise account | Identity and access, audit logs, retention settings, data residency, and a contractual prohibition on model training |
| Internal, sensitive, operationally important | Private AI environment connected to agency repositories through permission-aware access | Source documents, access rules, logs, retention settings, security monitoring, and encryption keys where appropriate |
| Highly sensitive or high-impact | Segregated environment with limited connectivity | Administrative access, ongoing security testing, mandatory human review, and a documented fallback process |
At the highest sensitivity level, the fallback process matters as much as the controls. Agencies should be able to continue essential work if the AI platform, the underlying cloud provider, or a key integration becomes unavailable.
Approval should not rest with any single function. The agency should establish a cross-functional governance process that includes mission leadership, information security, privacy, legal, procurement, records management, data governance, and program owners.
This group should evaluate proposed AI uses, approve data categories and security requirements, review vendor commitments, maintain an inventory of material dependencies, and oversee periodic reassessments.
Agencies should test the exit plan at defined intervals: export representative data, metadata, configurations, audit records, prompts, and workflows. Validate usability and, where practical, move a representative workflow to another environment. An exit plan that has never been tested is an assumption, not a capability.
Conclusion
Data sovereignty is not achieved simply by locating information in a particular cloud region or inserting a data-ownership clause into a contract.
It is achieved when the agency can govern its information throughout its lifecycle: where it resides, who can access it, what laws apply, how it is used, how it is retained, and whether it can be retrieved and moved without losing essential context, control, or function.
Vendor lock-in is the practical test of that control. An agency that cannot readily access its data, preserve its records, continue its operations, or transition to another qualified provider does not control its technology environment, no matter what the contract says.
AI makes that test more urgent. Dependencies that once formed over the life of a multi-year contract can now form in a matter of months, through everyday adoption rather than a procurement decision.
A closed-system-first approach can support innovation without sacrificing agency control, mission continuity, public trust, or future choice.