Aleph Alpha Kolibri: Self-Hosting a Sovereign AI Model

Aleph Alpha Kolibri: Self-Hosting a Sovereign AI Model

On 3 October 2026, Aleph Alpha, the AI company from Heidelberg, published the weights of a new model called Kolibri on Hugging Face under the Apache 2.0 licence. It is a 78-billion-parameter English and German model you can run on your own GPUs, with every prompt staying inside your network. If you run your own stack, this is the most interesting model release of the year so far.

Hitesh Saini has deployed self-hosted infrastructure for clients across Europe, the US and Australia since 2020.

Most of my European clients ask first where the data goes. With Nextcloud or Jitsi I have a clean answer. With AI, the honest answer has usually been "to someone else's servers". Kolibri is one of the few serious models where I can give a better answer, if the client will pay for the hardware.

What Kolibri actually is

Kolibri is an open-weight mixture-of-experts language model with 78.1 billion total parameters, about 3.46 billion of them active per token. It understands English and German and handles 262,144 tokens of context natively. Full details are on the Kolibri-1 model card on Hugging Face and in Aleph Alpha's "Kolibri has landed" launch post.

How it works, in plain words

A mixture-of-experts (MoE) model is one large model split into many specialist parts, called experts. A small router picks a few of them for each token, which is a word or part of a word.

Kolibri has 50 layers. Each has 384 experts plus 1 shared expert that is always used, and the router sends each token to 6 of the 384. That gives 78,103,074,560 total parameters but only 3,457,573,120 active per token, about 4.4%. Each token costs roughly what a 3.5B model costs in compute, which is why it can be fast.

The catch many people will miss: the router can pick any expert for the next token, so all 78 billion parameters must stay in GPU memory. MoE saves compute, not memory.

The weights ship in FP8, about 78 GB. FP8 is a form of quantisation: each number is stored in 8 bits instead of 16, so the model takes roughly half the memory.

  • Attention: a 4:1 mix of sliding-window and full attention. Four layers see only nearby text and one sees everything, which keeps long inputs cheaper.
  • Context: 262,144 tokens natively, which Aleph Alpha recommends for efficient serving. Validated up to 1,048,576 tokens (1 million).
  • Tokenizer: 128,000 tokens, using a custom method called UniBPE. It needs about 11.2% fewer tokens for German than GPT-5's tokenizer, and more savings on legal German, so more text fits in context at lower compute.
  • Reasoning effort: none, low, medium or high. Tool calling is supported.
  • Saying "I don't know": trained with Aleph Alpha's Merlin-Arthur protocol to abstain when the answer is not in the provided context. For document search over company files, this matters more than any benchmark.
  • Knowledge cutoff: 18 June 2026.
  • Training: 20 trillion tokens (about 62.5% English, 23.9% German, 13.6% code) on 768 NVIDIA B200 GPUs over 21 days, about 392,000 GPU-hours, in Germany and Finland.

What "sovereign" means here, and why it matters

Here, sovereign means you control and operate the model. You download the weights, run them on hardware you choose, and no request leaves your network. Nobody can change the model under you, raise the price or switch it off.

Kolibri was built in Germany, under European and German law. Aleph Alpha has signed the EU's General-Purpose AI Code of Practice and says it built Kolibri with the EU AI Act and GDPR in mind. That gives an EU client's data protection officer something concrete to read instead of a US cloud provider's terms of service.

One honest note: the model card says some synthetic training data was rephrased using non-European models. So "sovereign" is about who controls the model, not a promise that every input was European. I respect that they wrote it down.

Read the licence carefully. Apache 2.0 covers the weights and configuration files, so you can use, modify and ship them commercially. Aleph Alpha keeps the rights to its training code, model architecture, parameter settings and training methods.

Benchmarks: how it compares

These are Aleph Alpha's own numbers, not mine. Kolibri Origin is a second Kolibri variant in their comparison; the model card explains how it differs. In names like "35B-A3B", the first number is total parameters and the second is active parameters.

BenchmarkKolibriKolibri OriginQwen3.6 35B-A3BNemotron 3 Super 120B-A12BMistral Small 4 119B-A6B
AIME 2025 (EN)96.981.984.691.779.8
AIME 2025 (DE)87.573.582.985.672.3
GPQA diamond84.368.183.478.074.7
LiveCodeBench v685.959.282.582.071.2
SWE-Bench Verified66.4-51.073.860.8

AIME is competition maths, GPQA diamond tests hard science knowledge, LiveCodeBench tests writing code, and SWE-Bench Verified tests fixing real bugs in real repositories. Kolibri leads on maths, knowledge and code generation, even against models with far more active parameters. Nemotron 3 Super wins clearly on SWE-Bench Verified, the closest of these to real coding-agent work.

Aleph Alpha also reports overall scores of 75.5 in English and 70.8 in German, 61.4 on BFCL v4 for tool calling and 64.5 on LongBench Pro for long documents.

Honest limits

  • GPU memory: about 78 GB for the weights alone, plus room for the conversation cache. That rules out almost every homelab, including mine.
  • Languages: English and German only. If your users write in French, Hindi or Spanish, look elsewhere.
  • Coding agents: it writes good code, but the SWE-Bench gap shows it is not the best pick for an autonomous coding agent.
  • Serving stack: it needs vLLM on data-centre GPUs, not a desktop app on a gaming card.

How to self-host Kolibri

Hardware first. The minimum is 2x A100 80 GB, 2x H100 SXM5, 1x H200, 1x B200 or 1x B300. Aleph Alpha recommends 2x H100 SXM5, 2x H200, 1x B200 or 1x B300. If you do not own these, renting a single H200 or B200 node from a European GPU host is the realistic route.

You serve it with vLLM, an open-source engine that loads a model onto GPUs and exposes an OpenAI-compatible API. First install Aleph Alpha's inference package, whose source is in the aleph-alpha-inference repository on GitHub:

pip install 'aleph-alpha-inference>=1'

Then start the server:

vllm serve Aleph-Alpha/Kolibri-1 \
  --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice
  • --kv-cache-dtype fp8 stores the KV cache, the model's working memory of the current conversation, in FP8 to save GPU memory.
  • --reasoning-parser kolibri1 separates the model's reasoning from its final answer in the API response.
  • --tool-call-parser kolibri1 turns Kolibri's tool call format into standard tool call objects.
  • --enable-auto-tool-choice lets the model decide when to call a tool.

If you prefer containers, as I do for client work, there is a prebuilt image at ghcr.io/aleph-alpha/aleph-alpha-inference.

For contexts beyond 262,144 tokens, add two flags to the serve command:

--max-model-len 1048576 --hf-overrides '{"max_position_embeddings": 1048576}'

The first raises the maximum context vLLM accepts; the second lets the model config allow the longer positions. Memory use climbs sharply at that length, which is why Aleph Alpha recommends staying at 262,144 or below.

For sampling, use Aleph Alpha's recommended temperature=1.0, top_p=0.97 and top_k=128 in your client requests.

Should you run it?

Run it if you have, or can rent, an H200 or B200 class machine, your users work in English or German, and your data cannot leave your own infrastructure. It fits best for document search, internal assistants and German legal or compliance work. Skip it if you are on homelab hardware, need other languages, or mainly want an autonomous coding agent; a smaller model or a different MoE will serve you better.

Frequently asked questions

What is Kolibri?

An open-weight English and German language model from Aleph Alpha, released on 3 October 2026. It has 78.1 billion total parameters, about 3.46 billion active per token, and runs on your own hardware.

What is a sovereign AI model?

One you control and operate yourself, with no dependency on an outside provider's servers. Kolibri was also built in Germany under European and German law, with the EU AI Act and GDPR in mind.

What hardware does Kolibri need?

About 78 GB of GPU memory for the FP8 weights. Minimum: 2x A100 80 GB, 2x H100 SXM5, 1x H200, 1x B200 or 1x B300. Recommended: 2x H100 SXM5, 2x H200, 1x B200 or 1x B300.

Is Kolibri open source?

It is open-weight, not fully open source. The weights and configuration files are under Apache 2.0, so you can use them commercially, but Aleph Alpha keeps the rights to its training code, model architecture, parameter settings and training methods.

Final thoughts

Kolibri is the first European model I would seriously put in front of a client who needs AI and strict data control at the same time. The hardware bill is real, so plan for it before you promise anything. For the right team, it closes a gap that has been open for a long time.

This article was written with AI assistance and checked line by line against Aleph Alpha's launch post, the Hugging Face model card and the technical report. I did not have the hardware to run the 78B weights myself, so all performance numbers here are Aleph Alpha's own.

If you would like a setup like this planned and deployed for your team, you can see what I offer on my services page.