RB2SHosting · IT · Advisory

Scherwiller, near Sélestat — the model runs where you are

Alsatian AI: the model runs among the vineyards, not across the Atlantic.

For many organisations the question is not which model is cleverest, but where the sentences you hand it end up. A litigation file, an operative report, an industrial production plan, a council resolution — those texts have no business sitting with a third party. We install an open language model on a machine dedicated to you, in our technical room in Centre Alsace, and we run it.

Let us start with candour

This site's assistant does not run in Alsace yet. It will say so itself the day it does.

We would sell sovereign AI badly by lying about our own. The assistant you can try a little further down this page currently goes through a model provider's programming interface: your questions are passed to it. The gateway is already written to query our own machine first and fall back to the API only when it is unavailable — all that is missing is the machine, due when hosting goes live.

Today
The rb2s.com assistant queries an external programming interface. No question is stored on our servers; only aggregate counters are kept.
When hosting goes live
A dedicated Mac in our technical room runs an open model. The API becomes nothing more than a safety net for when the machine is down.
What you will see
Every answer produced by our machine carries the line “Answer produced by a model hosted in our own technical room”. Its absence means the answer came from the API. You do not have to take our word for it.
For your machines
What this page describes is waiting for nothing: a dedicated machine we install for you runs your model on site from the day it is commissioned.

Try it, right now

Talk to it. Same assistant, same instructions.

It answers from the information published on this site, makes no commitment on behalf of RB2S, and tells you when it does not know — including about the machine running it right now. No sign-up, nothing is kept.

What “Alsatian” actually means

Three different things, routinely conflated.

Plenty of offerings call themselves sovereign because the company is French, while inference runs on an American hyperscaler. The three questions are separate, and we answer them separately.

The machine

The computing happens here

The model is loaded into memory and executed on physical hardware at 6 rue du Sommerberg. Your requests cross no border, because they never leave the building.

The operator

The hands are here

We rack, power, monitor and restart it ourselves. Nothing is subcontracted to an operator whose name you would not know, and your contact is less than an hour's drive away.

The law

The contract is European

A GDPR-compliant processing agreement, French law, French jurisdiction. No extraterritorial legislation such as the Cloud Act layers itself on top.

Who this is for

When the text itself is the secret.

Local AI is pointless for drafting a job advert. It matters a great deal once the material is confidential by nature, or once handing it to a third party raises a professional, contractual or regulatory problem.

Law and accountancy

Professional secrecy cannot be delegated

Summarising a file, comparing contract versions, building a chronology: immediately useful work, on documents a solicitor or an accountant cannot pass to a third-party service without thinking twice.

Health and social care

Health data, therefore a regime of its own

Reports, letters, transcriptions. As soon as health data is involved, hosting and transfers obey strict rules: a dedicated machine on site makes the demonstration considerably simpler.

Industry and engineering

Know-how does not get sent away

Manufacturing procedures, test reports, technical documentation, internal code. These are precisely the texts that constitute the competitive advantage — and the ones people hesitate most to paste into a public chat window.

Public bodies

An expectation to set the example

Resolutions, correspondence from residents, public procurement. Public buyers look closely at where processing takes place, and regional infrastructure is easier to justify before a council than a distant framework agreement.

Software and studios

A cost that stops moving

At high volume, per-token billing becomes unpredictable. A dedicated machine has a fixed monthly price, whether you query it ten times or ten thousand times a day.

Those for whom this is not the answer

And we will say so

If you need the strongest models on the market, very long reasoning, or a massive and irregular load, a public API remains the right tool. We would rather lose a rental than install a machine that will disappoint you.

How it is set up

Four steps, with no hardware bought blind.

01

The use case

What you want to do, on which documents, for how many people at once. That conversation determines the size of the model, and therefore the machine — never the other way round.

First conversation, no commitment
02

The trial

We run the candidate model on your own documents, and you judge it on your texts rather than on a demonstration we picked. This is also where you find out whether a smaller model would do.

Over a few days
03

The machine

A dedicated Mac, a GPU server or a Jetson module, depending on what the trial showed. It is ordered, installed in the technical room, and you get root access and a fixed IP address.

Within seven to fifteen days
04

The connection

The model is exposed to your applications through an interface compatible with the market standard: anything that already talks to a public API talks to your machine by changing one address. Monitoring and backups included.

Then monthly, 30 days' notice

What it runs on

Memory decides everything.

The size of model you can run depends first on available memory, and its speed on that memory's bandwidth. These are the orders of magnitude we use when advising — and we redo the arithmetic with your figures.

Mac mini 16 GB
A quantised model of roughly 8 billion parameters. An internal assistant for a small team, a few concurrent requests, around ten watts at idle.
Mac mini 32 GB
Up to some fifteen billion parameters, or the same model with far more context: whole documents rather than extracts.
Mac Studio 64 to 256 GB
The large open models, without multiplying graphics cards. Apple's unified memory is a decisive advantage here, at a power draw nowhere near that of a GPU server.
NVIDIA RTX GPU server
When several dozen people query the model at once, or when it has to be fine-tuned on your data. Full CUDA ecosystem: vLLM, Ollama, image generation, transcription.
NVIDIA Jetson module
Inference running permanently at a few dozen watts: stream analysis, sensors, small specialised models. The most frugal option when throughput is not the point.
Models
Open models, executed on the machine: the Qwen, Gemma, Mistral and Llama families. We train nothing on your data and nobody sees it.

Prices for dedicated Macs are on the hosting page. Prices for GPU servers and Jetson modules are quoted within 48 hours: they depend on the hardware price at the time of order and on actual electricity consumption.

What we will not tell you

Three received ideas worth clearing up before signing.

On price

It is not a saving

At low volume, a machine left switched on costs more in electricity than moderate API use costs in tokens. The benefit lies elsewhere: confidentiality, and a cost that stops moving as volume grows.

On performance

It is not the best model in the world

The open models that fit on a company machine are good, sometimes excellent at bounded tasks — summarising, extracting, rewriting, classifying. On long and difficult reasoning, the large proprietary models keep an edge.

On solar power

It is only a share

The photovoltaic roof covers part of daytime consumption. A machine running day and night consumes well beyond that. We give the actual coverage rate to anyone who asks.

Frequently asked

Ask the assistant too — it has the same instructions.

Does this site's assistant already run on a machine in Alsace?
Not yet, and it is written above without evasion. It currently goes through a model provider's programming interface; the move is planned for the start of hosting operations, in November 2026. The gateway already queries our machine first once it is declared, and every answer coming from it is flagged on screen. You will be able to see the change yourself, without any announcement from us.
Is a locally hosted model cheaper than an API?
No, not at low volume, and we would rather say so straight away. A machine running continuously draws electricity every day whether you query it or not; a few hundred questions a month on a public API cost less than a coffee. The arithmetic reverses at high volume, or when the cost of your data leaving for a third party cannot be priced at all: that is usually where the decision is made.
Which models can be run?
Open models, downloaded and executed on the machine: the Qwen, Gemma, Mistral and Llama families, quantised to fit in memory. The choice is made on your documents during the trial. We impose no model: it is your machine, and you can change it.
Is my data used to train anything?
No. A model running on your machine sends nothing back to its publisher: the weights are downloaded once, then the machine works without re-emitting anything. We neither train nor fine-tune any model on your data, unless you explicitly ask us to as part of an engagement.
How do my applications talk to the model?
Through an interface compatible with the one that has become the market standard. In practice, a tool that already knows how to query a public API can query your machine by changing the address and the key. We expose the service encrypted, filtered and behind dedicated authentication — never the engine's administration interface, which has no business being on the internet.
What if the machine fails?
As with any hosting with us: intervention within four working hours, same-day return to service on hardware failure, with replacement equipment available on site. If your application must answer under all circumstances, we can configure it to fall back to an external API during the incident — which is exactly what this site's assistant does.
Can we come and see the machine?
The room is closed to the public and we do not run visits. Work on your hardware is carried out by our technicians at your written request. Everything concerning the building, power and connectivity is described in detail on the hosting page.
Do we need advisory first, or can we simply rent?
You can rent directly if you know what you want to run. If the use case is still to be defined — which is common — that is the work of RB2S Advisory: framing, model selection, governance and the AI Act. The two are ordered separately, and advisory is not a condition of renting.

First conversation

Tell us what you want the machine to read.

We will tell you which model is enough for it, which machine it needs, and whether a public API would serve you better. Reply within twenty-four working hours, first meeting with no commitment.

contact@rb2s.com Reply within 24 working hours