AI/News

Aleph Alpha releases Kolibri, a 78-billion-parameter open-weight German-English model under Apache 2.0, trained in Germany and Finland

Mixture-of-experts with about 3 billion parameters active per token and a context window of up to one million tokens, released on German Unity Day for public administration, industrials and aerospace. Every benchmark on the page is the company's own.

By Daily Aletheia · Checked against the primary source · 3 October 2026 · 2 min read


Image: Aleph Alpha - the cover image on "Kolibri Has Landed: A Sovereign Open-Weight Model", aleph-alpha.com, 3 October 2026

Aleph Alpha released Kolibri on 3 October - "the Day of German Reunification", the post notes - with full weights on Hugging Face under the Apache 2.0 licence. It is an English-German mixture-of-experts transformer with 78 billion total parameters and about 3 billion active per token (78.1 billion and 3.46 billion in the company's own table), supporting context lengths of up to one million tokens.

Who it is for

The company describes Kolibri as "a specialized language model built for sovereign mission-critical work in regulated areas including public administration, industrials and aerospace". Sovereignty is the sales pitch: the model was built by teams in Germany, trained on infrastructure in Germany and Finland "under European and German law, with no foreign control", with the EU AI Act, the General-Purpose AI Code of Practice and the GDPR "in mind from the ground up". Its size is meant to let customers run it on their own premises rather than send data to a third-party API.

German is a fifth of the training mix rather than a translation layer: 21.3 per cent of pre-training tokens, about 4.3 trillion of 20 trillion, with translated text held to 6 per cent overall because, the post says, "a model trained on it speaks German about a world that looks American". The model is trained to abstain - to say "I don't know" when the answer is not in the documents it is given - and to reason at four effort levels, none to high.

How it was built

Pre-training ran on 768 Nvidia B200 GPUs: 20 trillion tokens over 21 days, then 3.44 trillion of mid-training and 200 billion of long-context adaptation, nearly 24 trillion in all. Over those 21 days the run hit 38 unplanned interruptions, "roughly one per 10,000 GPU-hours", each recovered automatically. A predecessor, Kolibri Origin, finished pre-training on 11 June and was never released; Kolibri finished on 11 September. Knowledge cutoff is 18 June 2026.

The claims

Aleph Alpha says Kolibri "matches models with up to four times its active parameter count, such as Nemotron 3 Super" across maths, coding, grounding and long-context tasks, and "sits on the Pareto frontier for quality versus serving cost" in English and German, against Qwen3.6-35B-A3B, Nemotron 3 Super 120B-A12B and Mistral Small 4 119B-A6B. Those comparisons were run by the company releasing the model, on its own chart; its sector scores come from internal customer-proxy benchmarks that are not public. A tech report is linked from the post.

Sources & further reading

  1. 01Aleph Alpha - Kolibri Has Landed: A Sovereign Open-Weight Model, 3 Oct 2026 ↗Primary