S. Ostermann

Private AI

Open models on your own hardware, so the data stays with you

  • Open models on your own hardware, in your own server room or with a German operator.
  • Audit, hardware, model selection and stack from one place.
  • At the end your own people run the system.

Frontier AI from OpenAI or Anthropic means everything your staff type leaves the building on the way to an answer. With contracts, personnel files, customer data and sometimes even source code, that is often not what you want. I run open models on my own hardware and set up the same for you: audit, hardware, models, software, until your people can operate it themselves.

01 — Why in your own house

Sovereignty first. But cost can matter too.

The data stays in your network, so the data protection question is about your own server room or a trustworthy German operator. Nobody changes your terms overnight, and nobody retires the model you built a process around. Per-seat subscriptions are cheap for as long as the providers do not have to earn money. I would not assume that lasts.

What I will not claim is that it is cheaper on day one. A capable machine for local AI is expensive, and whether it pays off depends on how hard it is used. That is what the audit is for.

02 — The audit

Before you buy any hardware.

We go through the work you want to hand to AI and sort it: what a local model handles well, what is worth sending to a large cloud model, and what needs no AI at all, just decent search. You get a written recommendation: which machine is enough for it, and which work it takes off your hands. And if renting is the better choice at your volume, it says so.

03 — The hardware

Sized from the workload.

The model decides how much memory you need, and that is where the significant cost sits. I specify the hardware, tell you what it costs, and go through procurement with you. Whether it stands on a desk, in your own server room or in a rented German data centre is mostly a question of who looks after power, cooling and physical access.

04 — The models

Chosen for your material.

Gemma, Qwen and GLM differ more in practice than their benchmark scores suggest. Especially on German, and on documents written the way yours are. I test the candidates on your own material and show you where it gets rough. You get a shortlist with numbers behind it, and a test you can repeat in six months when the next model appears.

05 — The stack

What turns a model into something people use.

Most of the work is not the model: the inference server, the interface your colleagues open, access to your documents, user accounts, the agents that handle recurring jobs, and the integration with what you already run. Established open source throughout, with logging and access rights built in from the start, so your team can take it over.

06 — Handover

A running system your people can operate.

Done means it runs, the setup is documented, and whoever is responsible knows how to update a model and what to do when a card fails. If you want me to keep an eye on it afterwards, ask me.