Aleph Alpha has released Kolibri, its new sovereign open-weight language model, built for mission-critical workloads in public administration, industrials and aerospace. Kolibri was trained on Verda, in European data centers we own and operate, through to the final run. Congratulations to the team in Heidelberg. We're proud to have been their infrastructure partner.
Kolibri
The model is a 78B-parameter mixture of experts with 3.46B active parameters, controllable reasoning effort, native tool calling and a context window of up to 1M tokens. It's German-English native, and it's trained to abstain when context is insufficient rather than guess.
Aleph Alpha documented measures to address copyright, data protection and EU AI Act requirements, going beyond the legally required baseline. 21.3% of the pre-training tokens is German, and a large share of it was prepared in-house. The tokenizer is designed for efficiency in both languages: it compresses German better than other models without sacrificing English compression, leading to more efficient inference, lower costs and shorter response times. Aleph Alpha explores both the training data and the tokenizer in greater detail in their release post.
Our partnership
Training a model like Kolibri requires over a thousand GPUs running around the clock for months. We supplied the infrastructure and kept it running, while our engineers contributed to Aleph Alpha's training efficiency work.
Infrastructure & reliability
Verda supplied the GPU compute infrastructure behind Kolibri. We provisioned and operated the bare-metal fleet, custom-built to Aleph Alpha's requirements with NVIDIA Blackwell GPUs and NVIDIA InfiniBand interconnect. We maintain the cluster services on their behalf.
While hardware failures are inevitable, we worked together with Aleph Alpha to predict them early and mitigate them through maintenance. This included passive monitoring to identify and isolate nodes with degraded performance, automated recovery across multiple levels of the stack when failures occurred, and predictive maintenance to minimize hardware failures.
Software stack
To keep utilization high, we built the Kubernetes and observability stack that serves as a single control plane and observability suite, and we adapted the platform to Aleph Alpha's needs as they evolved across the training runs.
Contributing to training efficiency
“Their responsiveness in adapting the platform to our requirements throughout the trainings of Kolibri Origin and Kolibri was invaluable.” – Aleph Alpha, technical report
Our engineers also contributed to Aleph Alpha's training efficiency work. As acknowledged in the technical report: “We are especially grateful to Paul Chang, Riccardo Mereu, and Daniel Obolensky for their significant contributions to improving pre- and post-training efficiency. Their efforts were key in achieving a performant model.”
What's next
Kolibri can be downloaded with the full weights on Hugging Face and used under the Apache 2.0 license terms. You can learn more in the technical report.
We'll be sharing more about our work with Aleph Alpha soon. If you're attending NVIDIA GTC Berlin, you can learn more from us at booth #3059.
In the meantime, you can spin up your own InfiniBand cluster on Verda:
- Log onto our console
- Or explore our CLI, API, and other provisioning methods in our docs