AI Applied

Running AI locally in your company: from a laptop pilot to a sovereign server

Open models, hardware, costs in euros, GDPR and the AI Act. The method to test on a computer before investing in a server.

Louis Graffeuil
Louis Graffeuil
Founder Tandem
October 3, 2026Published
14 minread
Tandem illustration of local AI, a laptop linked by a cable to a small dark server locked with a coral padlock

Running AI locally means executing an artificial intelligence model on a machine you control. That machine can be a laptop, a server in your own building or a graphics processor (GPU) rented from a European host. Your prompts and documents no longer travel to OpenAI, Anthropic or Google. In 2026 this has become realistic, thanks to the open models released by Mistral, Google, OpenAI and Alibaba.

Three motives keep coming up with the executives I work with. They want to protect client data and strategic documents, reduce their exposure to US law, and stay compliant with GDPR and the European AI regulation, the AI Act. The right method has two stages. You first run a proof of concept (pilot) on a workstation, in a few days and with almost no budget. You then move to a server, but only if the pilot proves that the open model does the job.

This guide covers which models to pick, the hardware, the costs in euros, the state of the law in October 2026 and the method to go from test to deployment. The security of autonomous agents is covered in our article on AI agent security in companies, and the split of roles in the one on AI governance.

Tandem diagram of the three ways to run generative AI, in the vendor's cloud, locally on a workstation or on a dedicated server, with the company perimeter drawn as a dashed line
The model can stay the same from one option to the next. The boundary of your data is what moves.

Local AI, sovereign AI, on-premise, what are we talking about?

Local AI runs on a machine the company controls. The term covers two situations. The first is the workstation, where the model is downloaded and then runs offline. The second is the dedicated server, installed on your premises (on-premise) or rented exclusively from a hosting provider.

Sovereign AI means a chain that is entirely subject to European law, from the model to the hosting. Local AI on a server you control is sovereign by design. A vendor API hosted in Europe can be sovereign too, provided the vendor is outside the reach of US law.

The vocabulary around models needs a clarification. Most so-called open source models are in fact open-weight. You download the model file and run it wherever you want, but the training data stays private. Licences also vary from one model to the next. Gemma 4 and gpt-oss ship under Apache 2.0, while Mistral Medium 3.5 requires a commercial licence above roughly 17 million euros of monthly revenue.

Why companies want to run AI locally

Four reasons come up, and the first one outweighs the other three combined. According to France's national statistics office, Insee, 43% of companies using AI say concerns about data protection hold them back. The figure comes from a study published in July 2026.

  • Confidentiality. Client contracts, HR files, health data or a planned sale of the business have to stay in the company. A local model processes them without any third party seeing them go by.
  • Sovereignty. The US Cloud Act lets American authorities demand data held by a US provider, even when it is stored in Europe. Before the French Senate in June 2025, Microsoft France's head of legal admitted he could not guarantee otherwise.
  • Control over the model. A vendor can pull a model from its API overnight. Mistral removed Devstral 2 from its own in July 2026, while the weights remain downloadable. A company running it in-house keeps using it.
  • Offline use. A local model answers with Wi-Fi switched off. That matters on a building site, on the road or on an isolated industrial network.

The figure comes from Insee, and I have been making the same point since June 2024. In my newsletter on the threat to your data, I summed it up in one line. If it is free, you are the product. Paid plans and APIs exclude your data from training, but your data still passes through the vendor. My video on data protection in ChatGPT reached the same conclusion. The real guarantee comes from running the model on your own server.

What large French organisations already do

Large organisations have already made the move, often with Mistral, and almost always for the same reasons.

  • France's armed forces ministry signed a framework agreement in January 2026 that lets it deploy Mistral models on its own infrastructure, according to Acteurs publics.
  • BNP Paribas renewed its partnership with Mistral for three years in May 2026. The bank favours deployment in its own data centres for sensitive data, reports Le Monde Informatique.
  • Crédit Agricole plans 500 million euros of AI investment over 2026-2028. Between 100 and 200 million will fund two computer rooms for inference on its critical data, according to Le Monde Informatique in June 2026.
  • The French state rolled out its AI Assistant in June 2026, built on Mistral models served through its Albert platform and hosted by Outscale. The tool grew from 10,000 to 43,000 users over the summer, according to Acteurs publics.

I pointed this out in my comparison of AI ecosystems published on LinkedIn in November 2025. Mistral came out as the vendor that best combined open weights with possible on-premise hosting. That is exactly what these organisations are buying.

Tandem comparison table from November 2025 between OpenAI, Anthropic, Gemini, Mistral and Grok, covering open weights, privacy and on-premise hosting
Mistral already ticked open weights and on-premise hosting. The OpenAI row needs a correction, since gpt-oss has shipped with open weights since August 2025.

What local AI settles for GDPR and the AI Act, and what it leaves to you

Local AI settles the question of transfers and subcontracting. It leaves your obligations as data controller and as deployer fully in place. Where the model is hosted matters for GDPR, whereas the AI regulation applies according to use.

France's data protection authority, the CNIL, leans towards local deployment when data is sensitive. In its FAQ on generative AI, published in July 2024, it considers on-premise solutions more appropriate and more secure for personal data or strategic documents. It also acknowledges that remote hosting under a proper contract is often simpler, and that nothing prevents a small company from using shared infrastructure.

The AI Act works differently. It applies according to your role and the use of the system, whatever the server. A company using a model in a professional setting is its deployer. If it builds its own chatbot on an open model and puts it into service under its own name, it becomes its provider. It must then tell users they are talking to an AI, under Article 50, which has applied since August 2026.

The timeline moved this summer. The Omnibus regulation adopted by the Council, published in July 2026, postpones the obligations for high-risk systems. They now arrive in December 2027 for recruitment, education or credit, and in August 2028 for regulated products. The text also softens Article 4, which now asks companies to support AI literacy rather than guarantee it. In France, the CNIL is preparing to become the supervisory authority for the regulation, but the law designating it was still waiting for the National Assembly in October 2026.

QuestionWith local AI
Data transfers outside the European UnionRemoved, the data stays on your machine
Exposure to the Cloud ActRemoved if the server and host fall under European law
Reuse of your data for trainingRuled out, no vendor sees your prompts
Legal basis and purpose under GDPRUnchanged, to be documented like any processing
Impact assessment on sensitive dataStill required in the cases GDPR sets out
AI Act obligationsUnchanged, they depend on use and on your role
Security of the server and accessEntirely your responsibility

Which models can you run locally in 2026?

The right model depends first on the memory available. Open models fall into three families, and size decides the hardware far more than the brand does.

Tandem diagram of the three families of open models in October 2026, small, medium and large, with examples, memory needed, hardware and realistic uses
The middle family is where serious pilots happen. It fits on a well-specced Mac or a single GPU.

Small models, between 1 and 15 billion parameters, run on most recent laptops. They summarise, rephrase, classify and extract fields well. Medium models, from 20 to 120 billion, need a Mac with 64 to 128 GB or an 80 GB GPU. Qwen3.8 27B, Gemma 4 31B, gpt-oss-120b and Mistral Small 4 belong here. Large models such as GLM-5.3, Kimi K3 or DeepSeek V4 exceed 400 GB and need an eight-GPU server.

A simple rule gives you the memory you need. With 4-bit quantisation, the standard compression format for local use, a model takes about 0.6 GB per billion parameters. Add a 20 to 50% margin for the conversation context. That is how OpenAI can say gpt-oss-120b fits on a single 80 GB GPU.

The quality gap with closed models

Open models still trail the best closed models. On the Artificial Analysis index checked in October 2026, the best open-weight model scores 46 points, against 58 for Claude Opus 5.5. Gemma 4 31B scores 15, on a par with Claude Haiku 4.5 at 17. That level is already enough for many business tasks, such as extraction, sorting or summarising documents.

The case of Chinese models

The best open-weight models today are Chinese, such as Alibaba's Qwen, DeepSeek, Moonshot's Kimi or Z.ai's GLM. European measures, like the Italian authority's in January 2025, target the DeepSeek app and its servers in China. A downloaded model running on your machine without a connection avoids that transfer risk. The risk lies elsewhere. The US standards institute, NIST, measured in September 2025 a high vulnerability of DeepSeek models to guardrail bypass. For sensitive use, Mistral, Gemma or gpt-oss are easier to defend in front of a committee.

What hardware do you need to run AI locally?

Memory decides everything. On a Mac, unified memory serves both the processor and the GPU, which makes it the simplest pilot machine. On a PC, only the graphics card's memory really counts.

MachineMemoryPriceWhat it can run
14-inch MacBook Pro M5 Pro24 to 64 GBFrom 2,899 euros incl. VATSmall models, medium up to 30 billion with 64 GB
Mac mini M5 Pro64 GB3,319 euros incl. VATMedium models up to 30 billion, team pilot
Mac Studio M5 Max128 GB6,189 euros incl. VATgpt-oss-120b and the heaviest medium models
RTX 5090 graphics card32 GB5,500 to 7,000 euros incl. VATSmall and medium models up to 27 billion, very fast
Rented H100 GPU, OVHcloud or Scaleway80 GBAbout 2.80 euros per hour excl. VATgpt-oss-120b and medium models, for a whole team
Eight rented H100 or H200640 GB to 1.1 TB18,500 to 30,700 euros per month excl. VATLarge models such as GLM-5.3 or DeepSeek V4

Speed follows memory bandwidth. On the same llama.cpp community test, an M5 Pro produces about 66 tokens per second, an M5 Max 120 and an H100 close to 270. A token is a fragment of a word. These figures apply to a small test model, and a 27-billion-parameter model answers noticeably more slowly.

Setting up a local AI pilot, step by step

A local AI pilot takes half a day to install and a few days to judge. It answers a single question. Does the open model do the job on your real cases?

  1. Pick one use case and twenty real examples. Anonymise them for the first run. A meeting summary, field extraction from contracts or a standard reply to a client all work well.
  2. Install an inference engine. Ollama installs with one command on Mac, Windows or Linux. LM Studio does the same with a graphical interface, free for business use since July 2025.
  3. Download a model sized to your memory. On a 16 GB laptop, the command ollama run gpt-oss:20b is enough. With 32 to 64 GB, try ollama run gemma4:31b.
  4. Switch off the Wi-Fi and test. The model must answer offline. That proves nothing leaves the machine.
  5. Compare with your current cloud tool. Run the twenty cases through both tools. Score every answer on a simple three-level grid, correct, usable or to redo.
  6. Measure speed and memory. A model that saturates the machine or takes a minute to answer will not survive the pilot stage.

I presented Ollama as the most powerful tool in my selection in a May 2024 video on six AI tools. The models have changed since, the logic has stayed the same. The engine downloads the model, manages its versions and answers your applications in the background.

The Ollama demo starts around the fourteenth minute. The models have changed, the steps are still the same. The video is in French.

Desktop app, terminal or browser, where should you plug in the local model?

The engine runs in the background. Your teams need an interface, and three families exist depending on the number of users.

Tandem diagram of the three families of interfaces for a local model, desktop app, terminal and shared browser, sitting on top of the inference engine
The same engine can serve all three. The choice depends on who uses it, and how many of them there are.

Desktop apps such as LM Studio, AnythingLLM or Goose are enough for a single user. They feel close to ChatGPT or Claude, with conversations, documents and sometimes connectors.

The terminal suits technical profiles. Since January 2026, Ollama accepts Anthropic's API format, which lets you plug Claude Code into a local model. You keep the tool's skills and connectors, with a model running on your side. Anthropic states, however, that it does not support this setup. Also switch off telemetry with the CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC variable, otherwise usage metrics keep going out. Our Claude Code guide for non-developers covers getting started with the tool itself.

The browser serves teams. Open WebUI deploys on a server, handles accounts, roles and single sign-on, and is shared with a simple link. Its licence requires you to keep the Open WebUI branding beyond 50 users, unless you buy an enterprise licence.

Ben AI walked through these three families in a video published in August 2026, with Claude Code plugged into gpt-oss-120b and a shared Open WebUI interface. It shows the setup well in real conditions. His GPU cost estimates are higher than the European prices listed below.

Ben AI's video. The diagrams in this article follow its logic, with prices rechecked in France.

From pilot to server, three ways to deploy local AI

Once the pilot succeeds, three options open up. They differ in level of isolation, cost and operating effort.

A server on your premises

You buy the machine and install it in your server room. Control is total and the cost is a capital expense. The trade-off is skills. The French Directorate General for Enterprise guide estimated in 2024 that an on-premise assistant querying your documents took about six months to deploy, against two in the cloud. This kind of assistant relies on the RAG technique, explained in our RAG guide.

A dedicated GPU rented from a European host

You rent a machine reserved for you alone, from OVHcloud or Scaleway, in France. An 80 GB H100 costs about 2,000 euros per month excluding VAT, running around the clock. It is enough to serve gpt-oss-120b to a whole team. For health data or sensitive public data, ANSSI recommends a SecNumCloud-qualified offer. OVHcloud, for instance, offers servers with L4 and L40S GPUs in its qualified Bare Metal Pod range.

A sovereign API, paid per use

You call an open model hosted in Europe and pay by volume, with no server to manage. OVHcloud AI Endpoints serves gpt-oss, Qwen or Mistral from France, and the host states that your data stays out of model training. Your prompts still share the infrastructure with other clients, and this service is absent from the SecNumCloud catalogue. This option protects better than a consumer service, without the isolation of a dedicated server.

Tandem used a hybrid setup of this kind at Santiane, a health insurance broker. The automation workflow runs on the group's own infrastructure, with no third-party API outside the European Union. The Santiane case study details the outcome, eight weeks of implementation and a target cost under 0.10 euro per file processed. The orchestration layer, n8n, can be self-hosted too, as our n8n guide explains.

Tandem four-step roadmap to move from a local AI pilot on a workstation to governed production on a server
The pilot answers a quality question. The server answers a volume question.

How much does local AI cost compared with cloud licences?

Local AI costs almost nothing for a pilot and pays off at high usage volumes. In between, everything depends on model size and the number of users.

Tandem chart of the monthly cost of local AI in October 2026, from a Mac mini depreciated over three years to eight rented H200 GPUs, compared with 50 cloud assistant seats
Between the Mac mini and eight rented GPUs, the gap is more than 300-fold. Model size drives almost the entire price.

An existing workstation is enough for the pilot. A 64 GB Mac mini works out at under 80 euros per month excluding VAT over three years. A rented H100 running continuously costs about 2,000 euros per month, and eight GPUs for a large model between 18,500 and 30,700 euros.

By comparison, 50 cloud assistant seats cost around 1,250 euros per month, at about 25 euros per user according to our comparison of AI assistants. The comparison is misleading, because leading closed models keep a clear lead. Local AI pays off in three cases. Volumes are massive, processing runs day and night, or confidentiality simply rules out the cloud.

The same government guide gave a telling order of magnitude in 2024 for an internal assistant. Electricity for a local model came to about 2 euros per employee per year, against about 106 euros for the same use in the cloud. The figures are dated, but the structural gap still holds.

The limits of local AI to know before you start

  • Quality. The best open model trails the best closed model by about a dozen points. Long reasoning and fine writing suffer the most.
  • Concurrent users. A Mac serves one person well. As soon as ten colleagues query the model at once, you need a GPU server and an engine like vLLM.
  • Security. Ollama has no built-in authentication. In September 2025, Cisco Talos found 1,139 Ollama servers exposed on the internet, reports The Register. A local server belongs behind a VPN and authenticated access.
  • Maintenance. Models change every quarter, and so do their licences. GLM-5.2 was MIT-licensed, GLM-5.3 moved to a custom licence in August 2026. Someone has to track these versions, test updates and document the choices.

Should your company move to local AI?

Yes for the pilot, without hesitation. It costs a few days and next to nothing. Above all, it answers the only question that matters, whether an open model does the job on your real cases. For production, the answer depends on your data.

Most SMEs and mid-sized companies gain from a hybrid architecture. A cloud assistant under an enterprise contract covers everyday use. A local or sovereign model handles the sensitive flows, such as HR files, health data or plans to sell the business. Local AI becomes the rule when confidentiality, regulation or volume demand it.

Tandem helps make that call, use case by use case. That is the purpose of our AI audit, which ranks your uses by data sensitivity and the level of model they need. When an internal assistant on your documents is the right answer, our internal AI tools are deployed on whichever model and hosting your constraints require.

Frequently asked questions

Can you install ChatGPT or Claude locally?

No. The models behind ChatGPT and Claude are closed and only run at OpenAI and Anthropic. OpenAI does, however, publish gpt-oss, a family of open-weight models under the Apache 2.0 licence. The 20b version runs on a 16 GB computer, the 120b version on an 80 GB GPU or a 128 GB Mac. Claude Code can also be plugged into a local model through Ollama, without official support from Anthropic.

Does local AI work without an internet connection?

Yes. Once the model is downloaded, Ollama or LM Studio run it entirely on the machine, with Wi-Fi switched off. That is in fact the best test to check nothing leaves. Two exceptions call for care. Models tagged cloud in the Ollama library run on remote servers, and an agent installed locally may call an online model for every task.

Are Chinese models such as Qwen or DeepSeek safe to run locally?

The transfer risk disappears when the model runs on your premises without a connection, since European measures target the DeepSeek app and its servers in China. A model-specific risk remains. The US standards institute, NIST, measured in September 2025 a high vulnerability of DeepSeek models to guardrail bypass. For sensitive uses, Mistral, Gemma or gpt-oss are easier choices to justify.

How long does it take to deploy local AI?

A pilot on a workstation installs in half a day and is judged within a few days. A team pilot on a shared interface takes two to four weeks. A dedicated production server takes one to two months with a rented GPU. France's Directorate General for Enterprise estimated in 2024 that an on-premise assistant querying your documents took about six months, against two in the cloud.

Do you need a developer to install local AI?

Not for the pilot. LM Studio installs like any other application and downloads models from its interface, and Ollama takes a single command. Yes for production. A shared server needs a suitable engine such as vLLM, authentication, backups, monitoring and update tracking. That operating work should be planned from the pilot stage.

Read next

All articles →
AI AppliedAI governance in a company: illustration of a committee table with one seat highlighted

AI governance in a company: who decides, who builds, who spreads it

By Louis Graffeuil
AI AppliedTwo blocks, one pale and one dark, joined by a coral bridge, with a laptop resting on top

Forward Deployed Engineer: the job of deploying AI

By Louis Graffeuil
AI AppliedIllustration of Claude Code training, a laptop wearing a graduation cap

Claude Code training: the path for people who do not code

By Louis Graffeuil