For the past two years I have been talking almost every week with leaders who want to use artificial intelligence but cannot afford to send their data to someone else’s cloud. A hospital with medical records, a bank with transactions, a law firm with case files, a grid operator with substation diagrams. At SystemOne we work with cloud, DevOps and security every day, and I can see how the question itself has changed. It used to be “should we use AI?” Now it is “where exactly does it run, and who has access to everything it sees?”
For critical data my answer is more and more often the same: an open-weight model running on your own hardware, inside a perimeter isolated from the internet. That is sovereign AI in a practical rather than declarative sense.
What I mean by sovereign AI
Sovereignty is not a slogan on a slide. It comes down to three concrete conditions:
- Open weights. The model can be downloaded and run independently, without a vendor’s API. It is important to distinguish: “open weights” does not mean “open code and data”. Licenses differ significantly, and only a handful of developers fully disclose their training data.
- Your own hardware. Inference happens on servers owned by the organization or the country — in your own data center or in a private or national cloud under your jurisdiction.
- Isolation. In the strictest version, the environment is completely cut off from the internet (air-gapped): prompts, documents, answers and logs never leave the perimeter.
Why it matters: independence, control, security
- Independence. You do not depend on a vendor’s pricing, changing terms of use, sanctions, geopolitics or a decision to retire a model. A model you have downloaded and verified will keep working for as long as you need it.
- Full control and governance. You decide which version to run, when to update, which tools the model may call and what gets logged. For audits this is a fundamental difference: you show your own configuration, not a vendor’s promise.
- Data security. The model sees the same things as your most sensitive systems: personal data, trade secrets, technical documentation. When processing happens only locally, a whole class of risks disappears — from leaks through a third party to foreign authorities gaining access to data.
- Security of the model itself. Weights are an asset too, and they can be swapped or “poisoned”. Inside your own perimeter you control the model’s integrity just as you control the integrity of any other critical software.
- Resilience. For Ukraine this is not theory: a system that does not depend on an external link keeps working during communication outages or attacks on infrastructure.
From the state down to a single department
The national level. Ukraine is building a national language model called “Siaivo”. It is being developed by the WINWIN AI Center of Excellence under the Ministry of Digital Transformation together with Kyivstar, on top of Google’s open Gemma 3 model. According to DOU, a smaller version entered closed testing in June 2026, a full test model is expected by the end of the year, and the plan is to open it after validation. Another telling example is Switzerland’s Apertus from EPFL, ETH Zurich and CSCS: a fully open model (weights, data, training recipes) under the Apache 2.0 license, explicitly positioned by its authors as a foundation for sovereign AI.
The alliance level. The EU is building shared compute infrastructure: 19 AI Factories and 13 associated “antennas” are currently being set up, and on 30 July 2026 EuroHPC launched a tender for up to seven AI Gigafactories (proposals are due by 12 November 2026). In parallel, the OpenEuroLLM consortium is working on open multilingual models. The logic is simple: shared compute under European jurisdiction plus open models that can be deployed anywhere within that space.
The organization or department level. This is where I believe the practical impact is greatest. A few typical scenarios:
- a hospital where a model on a local server helps structure discharge summaries and search clinical protocols, while medical data never leaves the hospital network;
- a bank or insurer analyzing contracts and customer requests without sending personal data to third-party APIs;
- a law firm building search over its own case archive while preserving attorney-client privilege;
- an energy or water utility where an assistant works with diagrams, procedures and incident logs inside an isolated segment of the operational network;
- a city council processing citizens’ requests and internal documents on its own infrastructure;
- defense organizations, for which isolation is a baseline requirement rather than an option.
How to build it: key approaches
1. Air-gapped inference on your own hardware
The baseline design is GPU servers inside your own perimeter with no outbound internet access. The model, the inference server, the vector database, the applications and the logs all live inside. Where full isolation is not required, the alternative is a private sovereign cloud or national compute centers with clear jurisdiction and contractual control.
2. Right-sizing and quantization
The largest model is not always the best one. Quantization (FP8, FP4, 4–8-bit formats) cuts memory requirements several times over with a moderate loss of quality, and Mixture-of-Experts architectures offer hundreds of billions of parameters of which only a small fraction is active for each token. Start from the task: document search and classification often need just a single-GPU model, while complex analysis and agentic scenarios call for a server with several GPUs.
3. The inference stack
For server workloads, vLLM and SGLang have become the standard, with TensorRT-LLM as an option on NVIDIA infrastructure. For workstations, laptops and edge devices there are llama.cpp and Ollama. All of them can expose an API compatible with common clients, so applications are not tied to a specific model.
4. RAG over local data
Most of the value usually comes not from a “smarter model” but from access to your documents. Retrieval-augmented generation (RAG) with a local vector database and a local embedding model produces answers grounded in internal sources. It is critical that retrieval respects access rights: the model must never show a user a document they are not entitled to see.
5. Fine-tuning on your own data
When you need specific terminology or formats, fine-tuning (for example, LoRA) on your own data helps — again, inside the perimeter. But I recommend exhausting RAG and good instructions first: they are cheaper and easier to maintain.
6. Supply chain of the weights
Model weights are software, with all the supply-chain risks that implies. Minimum checks: download only from the developer’s official repositories (among popular models on Hugging Face there are plenty of modified copies with safeguards removed), verify checksums and, where available, digital signatures, use the safe safetensors format instead of formats that can execute code, review the license and provenance, and pin versions in an internal registry.
7. Access control, logging, red teaming
A model is one more privileged service. You need sign-in through corporate identities, role separation, logging of prompts and answers in line with data protection requirements, and limits on which tools the model can call. Plus regular resilience testing: prompt injection via documents, data extraction attempts, policy bypasses. Useful references are the OWASP Top 10 for LLM Applications and MITRE ATLAS.
8. Updates in an offline environment
Isolation does not mean “install and forget”. You need a process: a separate internet-connected zone to download and vet new models, containers and dependencies; scanning and testing against your own task set; controlled transfer into the closed environment through an approved channel; and the ability to roll back quickly. In other words, the same disciplined change management you apply to any critical system.
The strongest open-weight models as of early October 2026
Below are the models that top the open-model rankings on LMArena (Text Arena, open-source filter) and Artificial Analysis (Intelligence Index, open weights) in early October, plus notable compact models from the same rankings. Sizes are from Artificial Analysis and model cards; licenses are from LMArena and model cards. Treat this as a reference point, not a recommendation: before choosing, check the current state, the full license text and results on your own tasks.
Frontier tier: a GPU cluster or several top-end servers
- Kimi K3 (Moonshot AI) — 2.8T parameters, 104B active; custom license with separate terms for large commercial services.
- MiMo-V2.6-Pro (Xiaomi) — about 1T / 42B active; MIT.
- GLM-5.3 (Z.ai) — 753B / 40B active; MIT.
- DeepSeek V4 Pro — 1.6T / 49B active; MIT.
- Qwen3.8-2.4T-A95B (Alibaba) — 2.4T / 95B active; custom license, not Apache 2.0.
Mid tier: a single server with several GPUs
- GLM-5.3 Flash (Z.ai) — 320B / 18B active; MIT.
- DeepSeek V4.1 Flash — 552B / 16B active; MIT.
- MiniMax-M3 — 428B / 23B active; MiniMax Community License.
- Qwen3.5-397B-A17B (Alibaba) — Apache 2.0.
- NVIDIA Nemotron 3 Ultra — 550B / 55B active; OpenMDW-1.1.
- Mistral Medium 3.5 (128B; modified MIT) and Mistral Small 4 (119B / 6.5B active; Apache 2.0) — European options.
- gpt-oss-120b (OpenAI) — 117B / 5.1B active; Apache 2.0. Not a ranking leader, but very resource-efficient.
Compact tier: a single GPU, workstation, laptop, edge
- Qwen3.8 27B (Alibaba) — 27B dense; Apache 2.0. The strongest compact model on Artificial Analysis.
- Gemma 4 (Google) — 31B, 26B A4B, 12B, plus E4B and E2B for devices; Apache 2.0. Gemma 4 31B ranks highest among compact models on LMArena.
- gpt-oss-20b (OpenAI) — Apache 2.0.
- NVIDIA Nemotron 3 Nano (30B A3B) — NVIDIA Open Model License.
- IBM Granite 4.2 (30B, 8B, 3B) — Apache 2.0.
- Apertus (8B and 70B) and Olmo 3.1 (Ai2, 32B) — Apache 2.0 with open training data and process. Not quality leaders, but valuable where transparency of provenance matters.
One more observation: a large share of the open-ranking leaders come from Chinese labs. In an air-gapped environment the model sends no data to its developer, but questions of origin, built-in biases and your sector’s regulatory constraints deserve a separate, documented assessment.
The rankings change every month
According to Epoch AI, since the start of 2026 the best open models have lagged the best closed ones by about four months on average. The leaders of the open rankings change literally every month: some of the models listed above did not exist six months ago. That is why the architecture should let you swap the model without rewriting applications: a standardized API, your own test task set, and a well-rehearsed process for vetting and transfer. I recommend reviewing the rankings and reassessing your choice at least once a quarter.
I have written about the broader risk context in my posts on collective cyberdefense and AI inside companies.
Further reading
- LMArena — Text Arena, open-model leaderboard
- Artificial Analysis — open-source model comparison and overall leaderboard
- Hugging Face — trending text-generation models
- Epoch AI — how far open models lag closed ones
- DOU — status of Ukraine’s national LLM “Siaivo” (in Ukrainian)
- Apertus — Switzerland’s fully open model
- European Commission — AI Factories and EuroHPC AI Gigafactories call
- OpenEuroLLM
- OWASP Top 10 for LLM Applications and MITRE ATLAS
- Inference stack documentation: vLLM, SGLang, llama.cpp, Ollama, TensorRT-LLM
In lieu of a conclusion
If you are responsible for a critical organization — a hospital, a bank, an energy company, a government body or a municipality — pay attention to this architecture now. Not because cloud AI is bad, but because for certain data the cost of a mistake is too high to depend on someone else’s infrastructure. You can start small: one task, one model on one server, your own test set and clear access rules. If you need help with an assessment or the architecture, I will be glad to help — get in touch or message me on LinkedIn.