# Sovereign intelligence

Thijs Verreck · 30 September 2026

https://wavyr.com/sovereign-intelligence

If your organisation changed AI provider today, which parts of its work would carry over? Documents may be easy to move. Accumulated context, working integrations and reliable behaviour may be harder to reproduce.

This briefing uses sovereign intelligence to mean the ability to choose and replace AI systems while retaining useful company knowledge and control over access. That ability depends on where information lives, how it is accessed and whether another system can use it effectively.

The sections below review research on cost and capability, explain the parts of an agent and offer a checklist for assessing switching costs. The research findings and the proposed assessment framework are separate; this is not a new empirical study.

## The economics keep changing

Epoch AI estimates that the cost of reaching a given benchmark performance fell roughly 47% per quarter over 2023–26. Its historical comparison puts that pace ahead of DNA sequencing, compute, lithium batteries and electricity. These are different technologies measured over different periods, not competing products observed in a controlled experiment.

For a CIO, the implication is architectural: a fixed choice of model can become economically outdated while the surrounding implementation is still being rolled out. The relevant question is how much surrounding work must change when the model changes.

Lower inference prices do not automatically lower the total cost of a workflow. Tool usage, retries, integration and human review also count. Measure cost per successfully completed task, alongside quality, before changing the system.

## The frontier is moving in two directions

The cost of a particular capability can fall while the best available capability improves. These are separate developments. A cheaper substitute for yesterday’s model is not the same thing as a system that can solve a harder problem today.

The Epoch Capabilities Index brings benchmark results onto a common scale. The snapshot below shows both individual models and the highest score available over time. It provides context for periodic evaluation, but does not establish which model is suitable for a particular business workflow.

A benchmark score is not a percentage of business tasks completed or a promise of productivity. For routine work, a smaller model may be sufficient. For difficult work, assess what the strongest eligible models can do before designing a process around a weaker model’s limitations.

## Four terms, four different decisions

Conversations about ‘the AI’ often bundle several decisions together. Separating the layers makes it easier to ask what you are buying, what you are configuring and what you need to control.

The agent is the combined model and harness at work, not an extra piece of software wrapped around them. The environment is what that system can access: documents, applications, databases and other services. Our diagram uses nested boundaries to explain these relationships, not to prescribe a deployment topology.

In an integrated frontier-lab platform, the provider controls both the model and the harness that runs it. You configure and use those layers; you do not necessarily own or control them. That arrangement is a product choice, not a requirement of the architecture: a separate harness can connect to models from different providers.

Access also has a time dimension. Reading a file gives an agent a snapshot. If the file changes, the agent needs to read it again. Connecting a system is only the beginning; useful environments also make current information discoverable and clearly distinguish it from stale material.

## Better know-how can improve the same model

The WikiSkill preprint studies how agent experience can become persistent knowledge and reusable skills. In its Gemini-3.5-Flash comparison, average accuracy across five benchmarks rises from 49.5% without skills to 68.1% with WikiSkill: an increase of 18.6 percentage points, with the same inference model.

That is evidence for the value of structured, reusable know-how. It is not a measured uplift from adding proprietary company data, and it should not be presented as a forecast of enterprise performance. The reported results average three independent runs of the skill-evolution process.

The business hypothesis is worth testing: your definitions, worked examples and validated procedures may make an agent more effective on your tasks. A comparison on held-out work can test that hypothesis. More documents do not necessarily produce more accurate results.

## Build the environment the company keeps

Owning the model and harness gives a frontier lab two reasons to keep work inside its platform: more model usage and a deeper role in the customer’s operations. Memory, procedures and connections can make the product more useful while also increasing the work required to leave. This creates a commercial incentive for company knowledge to accumulate inside the platform.

The alternative is to make the company’s environment the permanent home for the services agents rely on: MCP tools, long-term memory, knowledge bases, reusable procedures and messaging between agents. In this design, the model and harness access those services; they are not the only place those services or their records exist.

The five areas below turn that principle into enterprise responsibilities. These services can be self-hosted or managed: what matters is control over access, usable records and the ability to change providers. The suggested owners should fit your organisation’s existing accountabilities.

### Knowledge and memory

What belongs here: Knowledge bases, source documents, customer and case history, agreed definitions, and durable records of decisions and preferences. Keep provenance and retention rules with the information.

Why it matters for independence: A replacement agent can recover the organisation’s context without people having to teach it everything again. Its immediate working context still lives in the harness; durable memory needs an explicit read and write mechanism.

Suggested owner: Business knowledge owners and the data team

### Tools and business systems

What belongs here: Company-controlled APIs and MCP servers that expose approved actions in CRM, ERP, document stores and other operational systems. Credentials and access rules belong with these services.

Why it matters for independence: Different compatible harnesses can use the same business capabilities. You adapt the agent’s connection instead of recreating every integration. The MCP client remains in the harness; the server exposes the company’s tools and information.

Suggested owner: Enterprise IT and integration teams

### Skills and ways of working

What belongs here: Versioned procedures, reusable skills, instructions, worked examples and escalation rules. Store the authoritative versions in repositories or document systems the company controls.

Why it matters for independence: Your operational know-how survives a change of AI product. A new harness may need different packaging or prompts, but the underlying procedure and its history remain available to reuse and test.

Suggested owner: Process owners and domain specialists

### Messaging and coordination between agents

What belongs here: Shared inboxes, message history, task queues, assignments, handoff records and persistent workflow status. Record who owns each task and what has already happened.

Why it matters for independence: Agents from different providers can coordinate through shared services. A replacement can pick up a recorded handoff instead of losing the work inside a platform’s private conversation. Active sessions do not transfer automatically; messaging requires its own service, not just MCP.

Suggested owner: Workflow owners and the automation team

### Permissions, audit and evaluation

What belongs here: Identity and access policies, approval requirements, durable action logs, evaluation cases and acceptance criteria. Enforce access at the services agents use, as well as in the harness.

Why it matters for independence: You retain the evidence needed to compare a replacement and reconstruct past actions. Verify that the new harness enforces approvals and produces the required records before letting it take over.

Suggested owner: Security, risk and the accountable business owner

Consider replacing the agent handling a customer request. The replacement can retrieve the same case history, use the same business-system tools and pick up a recorded handoff from another agent. The company retains the shared records and services; the replacement still needs compatible connections, permissions and an evaluation on the work. Active sessions and model behaviour do not automatically transfer. That is the practical independence this architecture aims for: changing the intelligence without recreating the company’s operating context.

## What would you lose if you switched provider?

Choose one workflow and name both the current system and a possible replacement. Specify whether the change covers just the model, the harness, or the entire platform. Switching a model within an existing harness is a different exercise from moving the whole workflow.

For each item below, record what would happen in that specific move: it carries over, needs rebuilding, is lost, or remains unknown. Assign an owner and estimate the effort and disruption. An export option is useful evidence only after the exported material has been used successfully in the replacement.

Treat unknowns as unanswered questions, not evidence of portability. Before deciding to switch, run the same held-out work in both systems and compare result quality, review effort and total cost. A dependency may be acceptable when its benefits and replacement costs are understood.

## Figure data and definitions

Cost comparison: LLM inference 47.0%, DNA sequencing 14.2%, compute 9.9%, lithium batteries 3.6%, electricity 1.2% compounded decline per quarter over the different historical periods in Figure 1.csv. Lines use relative cost = (1 − quarterly decline)^(4 × elapsed years), restricted to observed period lengths. Logarithmic cost axis; not a forecast.

Model: the capability to reason. Harness: software that turns capability into action. Agent: model and harness working toward a goal. Environment: the world the agent can observe and change.

WikiSkill: no skills 49.5%; with WikiSkill 68.1%; difference +18.6 percentage points. Average across five benchmarks for Gemini-3.5-Flash, not an enterprise data uplift.

Portability check: retain source knowledge, permission rules, evaluation cases and change history; compare agents on quality, cost per successful task and review required.

# AI provider switching assessment

Proposed CIO discussion framework. Assess one specific workflow and replacement; this is not a vendor rating or certification.

Workflow: ____________________
Current model / harness / platform: ____________________
Replacement: ____________________
Scope of change: ____________________
Assessment date: ____________________

Outcomes: Carries over / Needs rebuilding / Lost / Unknown. Confirm carry-over in the replacement, rather than relying on an export promise.

## 1. Company records and source material

Can you take the original documents, metadata and source references with you?

Potential loss: Missing material, provenance or usable access rules.

Evidence to request: Import a representative export into the replacement and check completeness, links and permissions.

Outcome: Unknown
Owner: ____________________
Evidence / gaps: ____________________
Effort / disruption: ____________________

## 2. Accumulated context and memory

What has the system learned about your work that exists only in its memory or conversation history?

Potential loss: Continuity: people may have to explain the same work again.

Evidence to request: Recover decisions, preferences and relevant history, then check whether the replacement can retrieve and use them.

Outcome: Unknown
Owner: ____________________
Evidence / gaps: ____________________
Effort / disruption: ____________________

## 3. Instructions, skills and workflows

Which procedures can you reuse, and which depend on this harness?

Potential loss: Working routines and the effort already spent refining them.

Evidence to request: Run a documented procedure in the alternative. Record changes to prompts, tools and approval steps.

Outcome: Unknown
Owner: ____________________
Evidence / gaps: ____________________
Effort / disruption: ____________________

## 4. Business-system connections

Which integrations and scheduled jobs would stop working?

Potential loss: Access to operational systems or uninterrupted automation.

Evidence to request: Exercise required read and write operations, authentication and error handling in a test environment.

Outcome: Unknown
Owner: ____________________
Evidence / gaps: ____________________
Effort / disruption: ____________________

## 5. Permissions, approvals and audit history

Can you retain the controls and evidence required for this workflow?

Potential loss: Control coverage or the ability to reconstruct past actions.

Evidence to request: Check role mappings, approval paths and accessible historical logs. Verify the replacement enforces the same boundaries.

Outcome: Unknown
Owner: ____________________
Evidence / gaps: ____________________
Effort / disruption: ____________________

## 6. Evaluations and feedback

Do you own the examples and feedback that establish whether the system works?

Potential loss: A reliable quality baseline and accumulated feedback.

Evidence to request: Use the same held-out cases and acceptance criteria in both systems. Keep results outside either platform.

Outcome: Unknown
Owner: ____________________
Evidence / gaps: ____________________
Effort / disruption: ____________________

## 7. Model-specific behaviour

Which results depend on this model, a fine-tune or a proprietary feature?

Potential loss: Task performance, even when the same documents and instructions transfer.

Evidence to request: Compare difficult cases, output formats and failure modes. Verify whether any custom model assets are actually transferable.

Outcome: Unknown
Owner: ____________________
Evidence / gaps: ____________________
Effort / disruption: ____________________

## 8. Day-to-day operation

How much interruption, retraining and reconfiguration would the move create?

Potential loss: Time, service continuity and the team’s familiarity with the system.

Evidence to request: Run a limited parallel trial. Record staff time, review load, service dependencies and a rollback plan.

Outcome: Unknown
Owner: ____________________
Evidence / gaps: ____________________
Effort / disruption: ____________________

## Decision

Blocking dependencies: ____________________
Acceptable dependencies and why: ____________________
Quality / cost comparison: ____________________
Next action and owner: ____________________

Source: https://wavyr.com/sovereign-intelligence#practice



A practical test of independence is whether an alternative works with what you can take with you.

## Sources

- [Model Context Protocol architecture](https://modelcontextprotocol.io/docs/learn/architecture). MCP documentation. Hosts and clients connect to servers that expose tools, resources and prompts. MCP itself does not establish company ownership or portability.

- [The plunging price of thought](https://epoch.ai/publications/the-plunging-price-of-thought). Epoch AI · Luke Emberson & David Roodman · 22 September 2026. Cost at a fixed level of benchmark performance; estimates and limitations.

- [Inference-cost: code and data](https://github.com/droodman/inference-cost/tree/9163c17ee7217b9f09ac51bcc71ed64c18c04bc8). David Roodman · repository snapshot. Figure 1.csv supplies the historical quarterly decline rates used in our comparison.

- [Epoch Capabilities Index](https://epoch.ai/eci). Epoch AI · data snapshot 28 September 2026. ECI point estimates by release date. The displayed frontier is a running maximum, not a forecast.

- [WikiSkill](https://arxiv.org/html/2608.27454). Tang et al. · August 2026 preprint · Table 1. Compiling Agent Experience into Persistent Knowledge for Skill Evolution.

- [AI Coding Dictionary](https://github.com/mattpocock/dictionary-of-ai-coding). Matt Pocock. Model, harness, agent and environment. Definitions adapted for an executive audience.

[Cost data](https://wavyr.com/research/sovereign-intelligence/cost-comparison.csv)

[ECI data snapshot](https://wavyr.com/research/sovereign-intelligence/epoch-eci-2026-09-28.csv)

