07 Local intelligence

Local AI, run properly

Open-weight models on hardware you own, wired into the tools you already use — and the serving, versioning and measurement that turn a working prototype into something a company can depend on.

Lead time
1+ weeks
Hosting
EU / EEA only
Licence
Open source throughout
Handover
Your repository

The interesting question about AI in a European company is rarely capability. It is jurisdiction. A prompt containing an unreleased product plan, a patient record or a client’s source code is a data transfer, and calling it a “feature” does not change that.

Ollama runs open-weight models on your own hardware, behind your own firewall. A single well-chosen GPU handles a team’s day-to-day drafting, classification, summarisation and code assistance with no request ever crossing your perimeter.

What we build first

We size the machine to the models you actually need, install Ollama declaratively on NixOS, and expose an OpenAI-compatible endpoint on your WireGuard network so existing tooling works unchanged. Then we connect it to where the work happens: editors, Redmine, internal search, and retrieval over your own document store.

Everything stays on your side of the firewall, which means the compliance conversation is short, and the monthly cost is electricity rather than tokens.

And then it has to be dependable

Getting a local model running takes a fortnight. Making it something forty engineers rely on during a release week is a different problem, and it is the one people underestimate. It used to be a separate engagement on this site; it is the same engagement, further along.

The prototype works, so adoption spreads. Then a second team wants a different model and the GPU has one. Someone pulls a newer tag and last month’s prompts quietly get worse, with no way to prove it. Retrieval indexes drift and nobody notices until an answer cites a policy replaced in March. The finance conversation arrives and there is no per-team usage figure to give.

None of those are model problems. They are the ordinary operational problems of a shared service — capacity, versioning, measurement, rollback — applied to a component that happens to be a language model. Which is why the second half of this engagement looks like infrastructure work, because it is.

Why it starts at a week

Because the first useful version genuinely is that quick: a GPU, Ollama on NixOS, an endpoint on your own network. Everything after that is scoped against what adoption actually does, and we quote it when we can see it rather than guessing at the outset.

An honest limit

Local open-weight models are not frontier models. For a large class of daily work — summarising, classifying, drafting, refactoring, answering questions about internal documents — the gap does not matter. For the rest, we help you draw a clear line about which work may leave, rather than pretending the line does not exist.

The long-form case

Why we reach for these in particular — the comparison tables, the numbers, and the situations where we would tell you to pick something else.

Nästa steg

Tell us what you are running.
We will tell you what it should be.

A first conversation costs nothing and takes forty minutes. You will leave it with an honest opinion about your stack — including, occasionally, that you should change nothing at all.