Running AI locally means executing an artificial intelligence model on a machine you control. That machine can be a laptop, a server in your own building or a graphics processor (GPU) rented from a European host. Your prompts and documents no longer travel to OpenAI, Anthropic or Google. In 2026 this has become realistic, thanks to the open models released by Mistral, Google, OpenAI and Alibaba.
Three motives keep coming up with the executives I work with. They want to protect client data and strategic documents, reduce their exposure to US law, and stay compliant with GDPR and the European AI regulation, the AI Act. The right method has two stages. You first run a proof of concept (pilot) on a workstation, in a few days and with almost no budget. You then move to a server, but only if the pilot proves that the open model does the job.
This guide covers which models to pick, the hardware, the costs in euros, the state of the law in October 2026 and the method to go from test to deployment. The security of autonomous agents is covered in our article on AI agent security in companies, and the split of roles in the one on AI governance.

Local AI, sovereign AI, on-premise, what are we talking about?
Local AI runs on a machine the company controls. The term covers two situations. The first is the workstation, where the model is downloaded and then runs offline. The second is the dedicated server, installed on your premises (on-premise) or rented exclusively from a hosting provider.
Sovereign AI means a chain that is entirely subject to European law, from the model to the hosting. Local AI on a server you control is sovereign by design. A vendor API hosted in Europe can be sovereign too, provided the vendor is outside the reach of US law.
The vocabulary around models needs a clarification. Most so-called open source models are in fact open-weight. You download the model file and run it wherever you want, but the training data stays private. Licences also vary from one model to the next. Gemma 4 and gpt-oss ship under Apache 2.0, while Mistral Medium 3.5 requires a commercial licence above roughly 17 million euros of monthly revenue.
Why companies want to run AI locally
Four reasons come up, and the first one outweighs the other three combined. According to France's national statistics office, Insee, 43% of companies using AI say concerns about data protection hold them back. The figure comes from a study published in July 2026.
- Confidentiality. Client contracts, HR files, health data or a planned sale of the business have to stay in the company. A local model processes them without any third party seeing them go by.
- Sovereignty. The US Cloud Act lets American authorities demand data held by a US provider, even when it is stored in Europe. Before the French Senate in June 2025, Microsoft France's head of legal admitted he could not guarantee otherwise.
- Control over the model. A vendor can pull a model from its API overnight. Mistral removed Devstral 2 from its own in July 2026, while the weights remain downloadable. A company running it in-house keeps using it.
- Offline use. A local model answers with Wi-Fi switched off. That matters on a building site, on the road or on an isolated industrial network.
The figure comes from Insee, and I have been making the same point since June 2024. In my newsletter on the threat to your data, I summed it up in one line. If it is free, you are the product. Paid plans and APIs exclude your data from training, but your data still passes through the vendor. My video on data protection in ChatGPT reached the same conclusion. The real guarantee comes from running the model on your own server.
What large French organisations already do
Large organisations have already made the move, often with Mistral, and almost always for the same reasons.
- France's armed forces ministry signed a framework agreement in January 2026 that lets it deploy Mistral models on its own infrastructure, according to Acteurs publics.
- BNP Paribas renewed its partnership with Mistral for three years in May 2026. The bank favours deployment in its own data centres for sensitive data, reports Le Monde Informatique.
- Crédit Agricole plans 500 million euros of AI investment over 2026-2028. Between 100 and 200 million will fund two computer rooms for inference on its critical data, according to Le Monde Informatique in June 2026.
- The French state rolled out its AI Assistant in June 2026, built on Mistral models served through its Albert platform and hosted by Outscale. The tool grew from 10,000 to 43,000 users over the summer, according to Acteurs publics.
I pointed this out in my comparison of AI ecosystems published on LinkedIn in November 2025. Mistral came out as the vendor that best combined open weights with possible on-premise hosting. That is exactly what these organisations are buying.

What local AI settles for GDPR and the AI Act, and what it leaves to you
Local AI settles the question of transfers and subcontracting. It leaves your obligations as data controller and as deployer fully in place. Where the model is hosted matters for GDPR, whereas the AI regulation applies according to use.
France's data protection authority, the CNIL, leans towards local deployment when data is sensitive. In its FAQ on generative AI, published in July 2024, it considers on-premise solutions more appropriate and more secure for personal data or strategic documents. It also acknowledges that remote hosting under a proper contract is often simpler, and that nothing prevents a small company from using shared infrastructure.
The AI Act works differently. It applies according to your role and the use of the system, whatever the server. A company using a model in a professional setting is its deployer. If it builds its own chatbot on an open model and puts it into service under its own name, it becomes its provider. It must then tell users they are talking to an AI, under Article 50, which has applied since August 2026.
The timeline moved this summer. The Omnibus regulation adopted by the Council, published in July 2026, postpones the obligations for high-risk systems. They now arrive in December 2027 for recruitment, education or credit, and in August 2028 for regulated products. The text also softens Article 4, which now asks companies to support AI literacy rather than guarantee it. In France, the CNIL is preparing to become the supervisory authority for the regulation, but the law designating it was still waiting for the National Assembly in October 2026.
| Question | With local AI |
|---|---|
| Data transfers outside the European Union | Removed, the data stays on your machine |
| Exposure to the Cloud Act | Removed if the server and host fall under European law |
| Reuse of your data for training | Ruled out, no vendor sees your prompts |
| Legal basis and purpose under GDPR | Unchanged, to be documented like any processing |
| Impact assessment on sensitive data | Still required in the cases GDPR sets out |
| AI Act obligations | Unchanged, they depend on use and on your role |
| Security of the server and access | Entirely your responsibility |
Which models can you run locally in 2026?
The right model depends first on the memory available. Open models fall into three families, and size decides the hardware far more than the brand does.

Small models, between 1 and 15 billion parameters, run on most recent laptops. They summarise, rephrase, classify and extract fields well. Medium models, from 20 to 120 billion, need a Mac with 64 to 128 GB or an 80 GB GPU. Qwen3.8 27B, Gemma 4 31B, gpt-oss-120b and Mistral Small 4 belong here. Large models such as GLM-5.3, Kimi K3 or DeepSeek V4 exceed 400 GB and need an eight-GPU server.
A simple rule gives you the memory you need. With 4-bit quantisation, the standard compression format for local use, a model takes about 0.6 GB per billion parameters. Add a 20 to 50% margin for the conversation context. That is how OpenAI can say gpt-oss-120b fits on a single 80 GB GPU.
The quality gap with closed models
Open models still trail the best closed models. On the Artificial Analysis index checked in October 2026, the best open-weight model scores 46 points, against 58 for Claude Opus 5.5. Gemma 4 31B scores 15, on a par with Claude Haiku 4.5 at 17. That level is already enough for many business tasks, such as extraction, sorting or summarising documents.
The case of Chinese models
The best open-weight models today are Chinese, such as Alibaba's Qwen, DeepSeek, Moonshot's Kimi or Z.ai's GLM. European measures, like the Italian authority's in January 2025, target the DeepSeek app and its servers in China. A downloaded model running on your machine without a connection avoids that transfer risk. The risk lies elsewhere. The US standards institute, NIST, measured in September 2025 a high vulnerability of DeepSeek models to guardrail bypass. For sensitive use, Mistral, Gemma or gpt-oss are easier to defend in front of a committee.
What hardware do you need to run AI locally?
Memory decides everything. On a Mac, unified memory serves both the processor and the GPU, which makes it the simplest pilot machine. On a PC, only the graphics card's memory really counts.
| Machine | Memory | Price | What it can run |
|---|---|---|---|
| 14-inch MacBook Pro M5 Pro | 24 to 64 GB | From 2,899 euros incl. VAT | Small models, medium up to 30 billion with 64 GB |
| Mac mini M5 Pro | 64 GB | 3,319 euros incl. VAT | Medium models up to 30 billion, team pilot |
| Mac Studio M5 Max | 128 GB | 6,189 euros incl. VAT | gpt-oss-120b and the heaviest medium models |
| RTX 5090 graphics card | 32 GB | 5,500 to 7,000 euros incl. VAT | Small and medium models up to 27 billion, very fast |
| Rented H100 GPU, OVHcloud or Scaleway | 80 GB | About 2.80 euros per hour excl. VAT | gpt-oss-120b and medium models, for a whole team |
| Eight rented H100 or H200 | 640 GB to 1.1 TB | 18,500 to 30,700 euros per month excl. VAT | Large models such as GLM-5.3 or DeepSeek V4 |
Speed follows memory bandwidth. On the same llama.cpp community test, an M5 Pro produces about 66 tokens per second, an M5 Max 120 and an H100 close to 270. A token is a fragment of a word. These figures apply to a small test model, and a 27-billion-parameter model answers noticeably more slowly.
Setting up a local AI pilot, step by step
A local AI pilot takes half a day to install and a few days to judge. It answers a single question. Does the open model do the job on your real cases?
- Pick one use case and twenty real examples. Anonymise them for the first run. A meeting summary, field extraction from contracts or a standard reply to a client all work well.
- Install an inference engine. Ollama installs with one command on Mac, Windows or Linux. LM Studio does the same with a graphical interface, free for business use since July 2025.
- Download a model sized to your memory. On a 16 GB laptop, the command
ollama run gpt-oss:20bis enough. With 32 to 64 GB, tryollama run gemma4:31b. - Switch off the Wi-Fi and test. The model must answer offline. That proves nothing leaves the machine.
- Compare with your current cloud tool. Run the twenty cases through both tools. Score every answer on a simple three-level grid, correct, usable or to redo.
- Measure speed and memory. A model that saturates the machine or takes a minute to answer will not survive the pilot stage.
I presented Ollama as the most powerful tool in my selection in a May 2024 video on six AI tools. The models have changed since, the logic has stayed the same. The engine downloads the model, manages its versions and answers your applications in the background.
Desktop app, terminal or browser, where should you plug in the local model?
The engine runs in the background. Your teams need an interface, and three families exist depending on the number of users.

Desktop apps such as LM Studio, AnythingLLM or Goose are enough for a single user. They feel close to ChatGPT or Claude, with conversations, documents and sometimes connectors.
The terminal suits technical profiles. Since January 2026, Ollama accepts Anthropic's API format, which lets you plug Claude Code into a local model. You keep the tool's skills and connectors, with a model running on your side. Anthropic states, however, that it does not support this setup. Also switch off telemetry with the CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC variable, otherwise usage metrics keep going out. Our Claude Code guide for non-developers covers getting started with the tool itself.
The browser serves teams. Open WebUI deploys on a server, handles accounts, roles and single sign-on, and is shared with a simple link. Its licence requires you to keep the Open WebUI branding beyond 50 users, unless you buy an enterprise licence.
Ben AI walked through these three families in a video published in August 2026, with Claude Code plugged into gpt-oss-120b and a shared Open WebUI interface. It shows the setup well in real conditions. His GPU cost estimates are higher than the European prices listed below.
From pilot to server, three ways to deploy local AI
Once the pilot succeeds, three options open up. They differ in level of isolation, cost and operating effort.
A server on your premises
You buy the machine and install it in your server room. Control is total and the cost is a capital expense. The trade-off is skills. The French Directorate General for Enterprise guide estimated in 2024 that an on-premise assistant querying your documents took about six months to deploy, against two in the cloud. This kind of assistant relies on the RAG technique, explained in our RAG guide.
A dedicated GPU rented from a European host
You rent a machine reserved for you alone, from OVHcloud or Scaleway, in France. An 80 GB H100 costs about 2,000 euros per month excluding VAT, running around the clock. It is enough to serve gpt-oss-120b to a whole team. For health data or sensitive public data, ANSSI recommends a SecNumCloud-qualified offer. OVHcloud, for instance, offers servers with L4 and L40S GPUs in its qualified Bare Metal Pod range.
A sovereign API, paid per use
You call an open model hosted in Europe and pay by volume, with no server to manage. OVHcloud AI Endpoints serves gpt-oss, Qwen or Mistral from France, and the host states that your data stays out of model training. Your prompts still share the infrastructure with other clients, and this service is absent from the SecNumCloud catalogue. This option protects better than a consumer service, without the isolation of a dedicated server.
Tandem used a hybrid setup of this kind at Santiane, a health insurance broker. The automation workflow runs on the group's own infrastructure, with no third-party API outside the European Union. The Santiane case study details the outcome, eight weeks of implementation and a target cost under 0.10 euro per file processed. The orchestration layer, n8n, can be self-hosted too, as our n8n guide explains.

How much does local AI cost compared with cloud licences?
Local AI costs almost nothing for a pilot and pays off at high usage volumes. In between, everything depends on model size and the number of users.

An existing workstation is enough for the pilot. A 64 GB Mac mini works out at under 80 euros per month excluding VAT over three years. A rented H100 running continuously costs about 2,000 euros per month, and eight GPUs for a large model between 18,500 and 30,700 euros.
By comparison, 50 cloud assistant seats cost around 1,250 euros per month, at about 25 euros per user according to our comparison of AI assistants. The comparison is misleading, because leading closed models keep a clear lead. Local AI pays off in three cases. Volumes are massive, processing runs day and night, or confidentiality simply rules out the cloud.
The same government guide gave a telling order of magnitude in 2024 for an internal assistant. Electricity for a local model came to about 2 euros per employee per year, against about 106 euros for the same use in the cloud. The figures are dated, but the structural gap still holds.
The limits of local AI to know before you start
- Quality. The best open model trails the best closed model by about a dozen points. Long reasoning and fine writing suffer the most.
- Concurrent users. A Mac serves one person well. As soon as ten colleagues query the model at once, you need a GPU server and an engine like vLLM.
- Security. Ollama has no built-in authentication. In September 2025, Cisco Talos found 1,139 Ollama servers exposed on the internet, reports The Register. A local server belongs behind a VPN and authenticated access.
- Maintenance. Models change every quarter, and so do their licences. GLM-5.2 was MIT-licensed, GLM-5.3 moved to a custom licence in August 2026. Someone has to track these versions, test updates and document the choices.
Should your company move to local AI?
Yes for the pilot, without hesitation. It costs a few days and next to nothing. Above all, it answers the only question that matters, whether an open model does the job on your real cases. For production, the answer depends on your data.
Most SMEs and mid-sized companies gain from a hybrid architecture. A cloud assistant under an enterprise contract covers everyday use. A local or sovereign model handles the sensitive flows, such as HR files, health data or plans to sell the business. Local AI becomes the rule when confidentiality, regulation or volume demand it.
Tandem helps make that call, use case by use case. That is the purpose of our AI audit, which ranks your uses by data sensitivity and the level of model they need. When an internal assistant on your documents is the right answer, our internal AI tools are deployed on whichever model and hosting your constraints require.



