Model Lab · 2 October 2026
We installed 22 open language models from 17 makers on one ordinary MacBook, ran the ones that fit, and asked every one of them the same 24 questions — in English, German, French, Polish, Portuguese, Spanish and Ukrainian. Then we asked the cloud. Every answer is graded and published.
The test
Every model got each knowledge question in each language — and an answer only counts in full if it is right and in the language it was asked in. A model that explains the sky perfectly, but in English when asked in Polish, gets half.
Then three checks with one right answer each. Each question was asked once, with temperature 0, so the answer is the model's best guess, not a lucky draw.
Local · MacBook Pro, Apple M1 Pro, 16 GB unified memory
Across all local models the easiest language was English (2.8 of 3 on average), the hardest Polish (1.8 of 3). Each language column scores three answers (0–3). Hover or tap a cell to read the model's own words. Click Score, tok/s or GB to sort.
| # | Model | EN | DE | FR | PL | PT | ES | UK | Bus | Words | Wall | Score | tok/s | GB |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | AMALIA 9BPortuguese public consortium | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | right | right | right | 24/24 | 22 | 5.6 |
| 2 | Gemma 3 12BGoogle | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | right | right | right | 24/24 | 16 | 8.1 |
| 3 | MamayLM 12BINSAIT, on Google Gemma 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | right | right | right | 24/24 | 15 | 7.3 |
| 4 | EuroLLM 9BEU research consortium | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | right | right | wrong | 23/24 | 21 | 5.6 |
| 5 | Tower+ 9BUnbabel, Lisbon | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | right | right | wrong | 23/24 | 18 | 5.8 |
| 6 | Aya Expanse 8BCohere | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | right | half right | wrong | 22.5/24 | 25 | 5.1 |
| 7 | Qwen3 14BAlibaba | 3 of 3 | 3 of 3 | 3 of 3 | 2 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | right | half right | right | 22.5/24 | 14 | 9.3 |
| 8 | Qwen3 8BAlibaba | 3 of 3 | 3 of 3 | 2.5 of 3 | 2 of 3 | 2.5 of 3 | 3 of 3 | 3 of 3 | right | half right | right | 21.5/24 | 25 | 5.2 |
| 9 | Mistral Nemo 12BMistral AI and NVIDIA | 3 of 3 | 3 of 3 | 3 of 3 | 2 of 3 | 3 of 3 | 3 of 3 | 2.5 of 3 | right | right | wrong | 21.5/24 | 23 | 7.1 |
| 10 | T-lite 2.1 8BT-Bank | 3 of 3 | 3 of 3 | 3 of 3 | 1 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | right | half right | right | 21.5/24 | 22 | 5.0 |
| 11 | Gemma 3 4BGoogle | 3 of 3 | 2.5 of 3 | 2 of 3 | 2.5 of 3 | 2.5 of 3 | 3 of 3 | 2.5 of 3 | half right | right | right | 20.5/24 | 41 | 3.3 |
| 12 | GigaChat 3.1 Lightning 10BSber | 3 of 3 | 3 of 3 | 3 of 3 | 1 of 3 | 2 of 3 | 3 of 3 | 3 of 3 | right | half right | wrong | 19.5/24 | 37 | 6.5 |
| 13 | Phi-4 14BMicrosoft | 3 of 3 | 2 of 3 | 2 of 3 | 2 of 3 | 3 of 3 | 2 of 3 | 1.5 of 3 | right | right | right | 18.5/24 | 15 | 9.1 |
| 14 | DeepSeek-R1 8BDeepSeek | 3 of 3 | 2.5 of 3 | 2 of 3 | 1 of 3 | 2 of 3 | 2.5 of 3 | 2 of 3 | right | right | right | 18/24 | 22 | 5.2 |
| 15 | Nemotron 3 Nano 4BNVIDIA | 3 of 3 | 3 of 3 | 2 of 3 | 1 of 3 | 1.5 of 3 | 3 of 3 | 2.5 of 3 | wrong | wrong | right | 17/24 | 32 | 2.8 |
| 16 | Granite 3.3 8BIBM | 2.5 of 3 | 3 of 3 | 2 of 3 | 1.5 of 3 | 1.5 of 3 | 3 of 3 | 1 of 3 | wrong | right | right | 16.5/24 | 24 | 4.9 |
| 17 | YandexGPT 5 Lite 8BYandex | 3 of 3 | 3 of 3 | 1.5 of 3 | 0 of 3 | 1 of 3 | 3 of 3 | 3 of 3 | wrong | wrong | right | 15.5/24 | 23 | 4.9 |
| 18 | Gervásio 8B PTPTUniversity of Lisbon | 2 of 3 | 0.5 of 3 | 1 of 3 | 1 of 3 | 2 of 3 | 3 of 3 | 2 of 3 | right | half right | wrong | 13/24 | 21 | 4.9 |
| 19 | Llama 3.1 8BMeta | 1 of 3 | 0.5 of 3 | 0.5 of 3 | 0.5 of 3 | 2 of 3 | 1.5 of 3 | 1 of 3 | right | half right | wrong | 8.5/24 | 25 | 4.9 |
| 20 | Llama 3.2 3BMeta | 2.5 of 3 | 1.5 of 3 | 0.5 of 3 | 0 of 3 | 1 of 3 | 1 of 3 | 1 of 3 | wrong | wrong | wrong | 7.5/24 | 60 | 2.0 |
Speed
Tokens per second while answering, averaged over the six questions. Comfortable reading is about 10–15.
Memory
macOS lets the graphics side use about 12 GB of the 16. Anything bigger spills onto the processor and into swap.
We tried anyway. gpt-oss 20B with half its layers on the GPU answered correctly — at about one token a second, with free memory down to 14%. Mistral Small 24B, a dense model, had not finished one short answer after fifteen minutes.
Cloud · Cloudflare Workers AI
Bigger versions of the same families, run on Cloudflare's GPUs, on the free plan. DeepSeek V4, Kimi K2.6 and GLM 5.3 sit behind the paid plan and are not here.
| Model | EN | DE | FR | PL | PT | ES | UK | Bus | Words | Wall | Score |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Llama 3.3 70BMeta | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | right | right | right | 24/24 |
| Gemma 4 26B (MoE)Google | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | right | right | right | 24/24 |
| Qwen3.8 27BAlibaba | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | right | right | right | 24/24 |
| Mistral Small 3.1 24BMistral AI | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | right | right | wrong | 23/24 |
| gpt-oss 120BOpenAI | 3 of 3 | 3 of 3 | 3 of 3 | 2 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | right | right | right | 23/24 |
| Qwen3 30B (MoE)Alibaba | 3 of 3 | 3 of 3 | 3 of 3 | 3 of 3 | 2.5 of 3 | 3 of 3 | 3 of 3 | wrong | right | right | 22.5/24 |
Verdict
AMALIA 9B (Portugal)
24 out of 24 — every language, every check — in 5.6 GB at 22 tokens a second. Portugal's public model, and the most even one here.
Gemma 3 12B · MamayLM 12B
Both 24 out of 24. About 15 tokens a second, and they want most of the memory: close other apps first.
Gemma 3 4B (Google)
20.5 out of 24 at 39 tokens a second in 3.3 GB. The one to leave loaded for quick questions.
AMALIA 9B · Tower+ 9B
Both knew every Portugal-only word; Tower+, from Lisbon's Unbabel, writes the most natural translations.
Aya Expanse 8B (Cohere)
Full marks in all seven languages; only the Portuguese words and the English sum let it down.
Gemma 4 26B · Llama 3.3 70B (cloud)
24 out of 24 on Workers AI's free plan; Llama 3.3 answers in seconds.
Reading the labels
8b, 14bBillions of parameters — the model's size. In the usual 4-bit form a model needs about 0.6 GB per billion, so 14b is roughly 9 GB.
30b-a3bMixture of experts: 30 billion in total, 3 billion working on each word. Fast like a small model, but all 30 billion still have to sit in memory.
q4_K_MCompression: about 4 bits per parameter instead of 16. Four times lighter for a small loss. q8 is near-lossless, q2–q3 noticeably weaker.
Decides speed. Each word means reading the whole model once, so tokens per second ≈ bandwidth ÷ model size. Memory size decides what runs at all.
Method
Local models ran through Ollama on MacBook Pro, Apple M1 Pro, 16 GB unified memory, one at a time, with a watchdog that unloads a model if free memory falls below 6%. Every model got the same short system prompt and temperature 0. Cloud models ran through Cloudflare Workers AI with the same prompt. Answers were graded by fixed rules, including a check of which language each reply is in; every raw answer is public: local.json, cloud.json.
24 questions are a probe, not an exam. They reward knowing European Portuguese and answering in the reader's language, and they penalise a single slip; a model can be excellent at things not asked here. Speeds are for this one laptop.