Kolibri AI is a free, open-weight reasoning model from the German company Aleph Alpha, and my verdict is that it's an excellent model for German and English work that most people can't run at home.

That's the whole verdict in one sentence.

People also search for it as "Colibri AI", but the official name is Kolibri, the German word for hummingbird.

It launched on 3 October 2026 under the Apache 2.0 licence, with 78.1 billion total parameters and about 3.46 billion active per token.

In this review I'll score every feature, show you the real hardware bill, flag which benchmark numbers are vendor-reported, and tell you who should and shouldn't bother.

The Verdict: Kolibri AI Scores 7 Out Of 10

Kolibri AI earns a 7 out of 10 from me.

It's genuinely strong for German and English reasoning, it's truly open under Apache 2.0, and it has the best compliance story of any open model I've covered.

It loses points because the official version needs data-centre memory, and because the easy desktop apps can't load it yet.

If you have the hardware and you work in German, it's close to a must-test.

If you're on a normal laptop, it's a model to read about rather than run.

One more honest note before the scores.

I run a Mac Studio and I don't run much local AI myself, so these scores lean on the official model card, Aleph Alpha's vendor-reported benchmarks and the community's published numbers.

Kolibri AI Review Scorecard

Area Score out of 10 Why
German language quality 9 It has the best German overall score among the MoE models in Aleph Alpha's own table
English reasoning and maths 8 It scores 96.5 on the English maths average and 84.3 on GPQA Diamond, both vendor-reported
Agentic tool calling 7 It ties the best MoE agentic average at 63.4, but trails on some function-calling tests
Coding 7 It scores 89.3 on the code average and 66.4 on SWE-Bench Verified, behind the Qwen MoE models on SWE-Bench
Long context 8 It reads 262,144 tokens natively and Aleph Alpha validated it up to about 1 million
Licence and openness 10 It's Apache 2.0, ungated, with a detailed model card and technical report
Compliance positioning 9 Aleph Alpha signed the EU GPAI Code of Practice and designed it around the EU AI Act and GDPR
Hardware accessibility 3 The official build needs about 78 GB, and even 2-bit community builds need a 36 GB Mac
Ease of setup 4 Stock LM Studio, Ollama and llama.cpp can't load it yet
Factual reliability 5 Its own model card shows weak closed-book and fact-check results
Overall 7 A strong, honest open model held back by its memory bill

What Kolibri AI Actually Is

Kolibri AI is Kolibri 1, Aleph Alpha's mixture-of-experts reasoning model with a focus on German and English.

Aleph Alpha is a German company based in Heidelberg, and the model card names Aleph Alpha Research GmbH as the developer.

Here are the official specs.

Feature Official detail
Name Kolibri 1, often searched as Kolibri AI or Colibri AI
Repo Aleph-Alpha/Kolibri-1 on Hugging Face, plus a BF16 version
Released 3 October 2026
Licence Apache 2.0
Total parameters 78.1 billion
Active per token About 3.46 billion
Experts 384 per layer, with 6 routed and 1 shared
Layers 50
Attention Hybrid, with four sliding-window layers for every one full-attention layer
Context 262,144 tokens native, validated to 1,048,576
Training data About 20 trillion pre-training tokens, plus 3.44 trillion mid-training and 201 billion long-context tokens
Data mix About 62.5% English, 23.9% German and 13.6% code in pre-training
Knowledge cutoff 18 June 2026
Output Text only

In my video I quoted roughly 24 trillion training tokens, and the official breakdown above adds up to about 23.6 trillion.

I also said "3 billion active", and the exact official figure is 3.46 billion.

Kolibri AI Features, Reviewed One By One

Reasoning mode

Kolibri has an explicit reasoning mode with low, medium and high effort levels.

You can also set it to none, and the model answers straight away with no thinking step.

This is a great feature because you can trade speed for quality on every single request.

My only complaint is that some community builds default to high effort, which feels slow until you change it.

Tool calling

Kolibri supports tool calling, and the official vLLM setup uses a Hermes-style tool-call format.

That makes it a natural fit for agent frameworks like Hermes Agent.

Aleph Alpha's agentic average puts it level with Qwen3.5 35B-A3B at 63.4, and ahead of Qwen 3.6, Nemotron 3 Super and Mistral Small 4.

On the BFCL function-calling tests it trails several rivals, so it's strong but not the clear leader.

Long context

Kolibri was trained on 262,144-token sequences in its final phase, and that's its native window.

Because only the sliding-window layers use positional encoding, Aleph Alpha says it extends beyond that without extra tricks, and they validated it up to about 1 million tokens.

They still recommend staying at or under 262,144 tokens for speed and for complex tasks.

German-first tokenizer

Aleph Alpha built a tokenizer tailored to German word structure.

That means long German compound words get processed efficiently instead of being chopped into lots of tiny pieces.

It's a small detail that makes a real difference to cost and quality on German text.

Compliance and sovereignty

This is the feature that makes Kolibri different from almost every other open model.

Aleph Alpha says it was built with the EU AI Act, the General-Purpose AI Code of Practice and GDPR in mind from the ground up.

The model card states the company is a signatory of the EU GPAI Code of Practice, and it publishes training data sources, energy use and limitations in detail.

Their launch blog describes it as built for "mission-critical work in regulated areas", naming public administration, industrials and aerospace.

Human-in-the-loop design

Aleph Alpha is unusually direct about how it should be used.

The model card says it's meant for systems where a person reviews the output before it's acted on, not for unsupervised autonomous agents.

I respect that honesty, and it's good advice for any model.

๐Ÿ”ฅ Want to test open models like this without wasting a weekend? Inside the AI Profit Boardroom, I share the local model setups I actually use, with step-by-step tutorials and four coaching calls a week alongside 3,400+ members. โ†’ Get access here

Kolibri AI Benchmarks: The Vendor's Own Numbers

Every number in this section comes from Aleph Alpha's model card, so it's vendor-reported and run on their own evaluation framework at high reasoning effort.

Vendor-reported Kolibri Qwen3.5 35B-A3B Qwen3.6 35B-A3B Gemma 4 26B-A4B Nemotron 3 Super Mistral Small 4 GPT-OSS 120B Qwen 3.8 27B (dense)
Overall English 75.5 74.7 71.4 71.9 73.0 63.1 72.3 80.2
Overall German 70.8 69.8 67.3 66.3 67.9 61.4 70.2 79.9
Agentic average 63.4 63.4 62.1 54.6 54.9 40.7 54.0 66.7
Maths average 96.5 90.1 87.8 87.4 91.1 81.4 90.7 97.8
Code average 89.3 85.0 87.7 89.0 88.3 82.0 90.8 94.2
SWE-Bench Verified 66.4 71.6 73.8 57.8 60.2 60.8 not reported 72.6

Here's how I read that table.

Kolibri is the best all-rounder among the mixture-of-experts models Aleph Alpha tested, especially in German and maths.

A tweet I mentioned in my video said it beats Qwen, Nemotron and Mistral in maths, science and coding, and the vendor table broadly backs that up against those MoE models.

It isn't the best model in the table, though.

The dense Qwen 3.8 27B beats it on almost every average, which is exactly what one viewer told me in the comments.

The difference is cost per token, because Qwen 3.8 27B uses all 27 billion parameters on every token and Kolibri uses about 3.46 billion.

There are also weak spots in Aleph Alpha's own table that I don't see anyone talking about.

Kolibri scores 34.0 on the RGB fact-check test, where most rivals score between 53 and 90.

It scores 51.0 on RGB closed-book questions, which is the lowest score in the whole table.

So use it with documents and search, not as a know-it-all.

If you're choosing a model mainly for code, my roundup of the best AI models for coding is the better place to start.

Kolibri AI Hardware Requirements: The Real Bill

This is where the review score drops, so I'll be precise.

Build Size What it needs Status
Official FP8 About 78 GB 2ร— A100 80 GB, 2ร— H100, 1ร— H200, 1ร— B200 or 1ร— B300 minimum Official, via vLLM and Aleph Alpha's plugin
Official BF16 Larger than FP8 Even more GPU memory Official
Community MLX 4-bit About 41 GiB A 64 GB-plus Mac Unofficial, ships its own launcher
Community MLX 2-bit About 24 GiB A 36 GB-plus Mac Unofficial, with a bigger quality hit
Community GGUF Q4_K_M About 47.5 GB Plenty of RAM and a patched llama.cpp Unofficial

A viewer said a 78B model needs roughly 100 GB of memory to run properly, and once you add room for context on top of the 78 GB of weights, that's a fair estimate for the official build.

Another asked about a 2 GB graphics card, and that's a firm no.

On speed, I haven't measured Kolibri myself.

The community converters report around 52 to 56 tokens per second on an M1 Max for their MLX builds, and around 13 to 15 tokens per second on a CPU-only desktop for their 4-bit GGUF.

Treat those as their numbers, not mine.

Kolibri AI Setup: How Easy Is It?

In my video I said you can run Kolibri in LM Studio, and I have to mark that down for today.

As of 7 October 2026, stock LM Studio, Ollama and llama.cpp can't load Kolibri, because its architecture is new.

Support requests are open on llama.cpp and mlx-lm, and there's an open pull request on Ollama.

The routes that work right now are these three.

  • The official route uses vLLM with pip install 'aleph-alpha-inference>=1', then vllm serve Aleph-Alpha/Kolibri-1 with the Kolibri reasoning and tool-call parsers.
  • The Mac route uses a community MLX build, which ships its own model file and launcher script.
  • The tinkerer route uses a community GGUF with a llama.cpp patch you compile yourself.

All three give you an OpenAI-compatible server, which is how you plug it into Hermes Agent or any other agent tool.

Kolibri AI With Hermes Agent: Worth It?

Pairing Kolibri with Hermes Agent gives you a free local agent brain with strong German and solid tool calling.

That's the pitch I made in the video, and it holds up on paper.

My honest caveat is that the best local Hermes model I've tested so far is still LFM 2.5 at 2.6 billion parameters, because it's crazy fast and was trained with Hermes Agent in mind.

Kolibri is far smarter, but it's also far heavier.

So my verdict is that Kolibri is the better Hermes brain if you have the memory and need quality, and LFM 2.5 is the better one if you need speed on modest hardware.

Kolibri AI Pros And Cons

These are the pros.

  • It's free and Apache 2.0, so you can use it commercially and run it anywhere.
  • It has the best German overall score among the MoE models in its vendor table.
  • It has a reasoning switch with four levels, including none.
  • It supports tool calling that drops straight into agent frameworks.
  • It has a long context window that Aleph Alpha validated up to about 1 million tokens.
  • It comes with the most detailed compliance and transparency documentation I've seen on an open model.

These are the cons.

  • It needs about 78 GB of memory in its official form.
  • It doesn't load in stock LM Studio or Ollama yet.
  • It's text only.
  • It scores poorly on some fact-check and closed-book tests in its own model card.
  • It's beaten by the dense Qwen 3.8 27B on Aleph Alpha's own overall averages.

Who Should Use Kolibri AI?

Use Kolibri AI if you work in German and English and want an open model with a strong compliance story.

Use it if you have a GPU server, a rented cloud GPU or a 64 GB-plus Mac.

Use it if you're building agents for European clients who won't approve American cloud tools.

Skip it if you're on a 16 GB laptop, because a smaller local model will serve you much better.

Skip it if you need image input, or if you want a one-click install today.

If you want to see how I judge local models against hosted ones, my verdict on running Jev locally uses the same honest approach.

๐Ÿ”ฅ Want my full local AI and agent setups? The AI Profit Boardroom has the training, the Agent OS and weekly coaching where you can ask me which model fits your hardware. โ†’ Join the AI Profit Boardroom

Related Reading

Also On Our Network

Kolibri AI Review FAQ

Is Kolibri AI any good?

Yes, it's one of the strongest open mixture-of-experts models for German and English on Aleph Alpha's own tests.

Its weak points are the memory it needs and its fact-checking scores.

Is Kolibri AI better than Qwen 3.8 27B?

Not on Aleph Alpha's overall scores, where Qwen 3.8 27B leads with 80.2 in English and 79.9 in German.

Kolibri uses far fewer active parameters per token, so it can be cheaper to serve at scale.

What hardware does Kolibri AI need?

The official build needs about 78 GB, with two 80 GB A100s or one H200 as the minimum.

Community 4-bit MLX builds need a 64 GB Mac, and it won't run on a 16 GB laptop or a 2 GB graphics card.

Does Kolibri AI work in LM Studio?

Not in stock LM Studio as of 7 October 2026.

Use vLLM with Aleph Alpha's plugin, or a community MLX or GGUF build, until support lands.

Is it Kolibri or Colibri?

The official spelling is Kolibri, the German word for hummingbird.

Colibri AI is the common alternative spelling people search for.

Is Kolibri AI free for commercial use?

Yes, it's released under Apache 2.0, which allows commercial use.

Final Verdict

Kolibri AI is a serious open model with a clear purpose, which is strong German and English work for people who need to own their AI.

It's held back by memory, not by quality.

If you have the hardware, test it against Qwen on your own tasks this week.

๐Ÿ“บ Video notes + links to the tools ๐Ÿ‘‰

๐ŸŽฅ Learn how I make these videos ๐Ÿ‘‰

๐Ÿ†“ Get a FREE AI Course + Community + 1,000 AI Agents ๐Ÿ‘‰

About Julian

I'm Julian Goldie, an SEO entrepreneur, author and founder of the AI Profit Boardroom, which has 3,400+ members.

I help business owners scale with AI agents, automation and SEO.

  • I've built a 7-figure agency, Goldie Agency, with a team of around 50 people.
  • I've grown a YouTube channel to 400,000+ subscribers.
  • I wrote the Amazon best-sellers "SEO Link Building Mastery" and "Agency Marketing Mastery".
  • My Udemy courses have taught over 50,000 students.

โ†’ Get my best AI training inside the AI Profit Boardroom

My final score stands at 7 out of 10, and Kolibri AI is the open model I'd test first for German-language agent work.