azenha.aiModel Lab

Model Lab · 2 October 2026

Open AI models on a 16 GB laptop, asked the same things in seven languages

We installed 22 open language models from 17 makers on one ordinary MacBook, ran the ones that fit, and asked every one of them the same 24 questions — in English, German, French, Polish, Portuguese, Spanish and Ukrainian. Then we asked the cloud. Every answer is graded and published.

20
models that run on 16 GB
7
languages, every question in each
24/24
best local score — AMALIA 9B · MamayLM 12B · Gemma 3 12B
6
cloud models for comparison

The test

Three questions in seven languages, three specialist checks

Every model got each knowledge question in each language — and an answer only counts in full if it is right and in the language it was asked in. A model that explains the sky perfectly, but in English when asked in Polish, gets half.

EnglishGermanFrenchPolishPortugueseSpanishUkrainian
  1. Why is the sky blue? Two sentences.Rayleigh scattering: short, blue wavelengths scatter most.
  2. A train leaves at 9:40 and travels 2 h 35 min. When does it arrive?12:15
  3. Who wrote “The Master and Margarita”, and when was it first published?Mikhail Bulgakov; the magazine Moskva, 1966–67.

Then three checks with one right answer each. Each question was asked once, with temperature 0, so the answer is the model's best guess, not a lucky draw.

  1. BusTranslate into European Portuguese: “Ich komme zu spät zum Bus, warte im Café auf mich.”asked in German → PortugueseAutocarro, not the Brazilian ônibus, and the lateness kept.
  2. WordsHow do you say “cell phone”, “bus”, “breakfast” and “train” in Portugal?asked in PortugueseTelemóvel, autocarro, pequeno-almoço, comboio.
  3. WallIf 3 workers build a wall in 6 hours, how many hours do 2 workers need?asked in English9

Local · MacBook Pro, Apple M1 Pro, 16 GB unified memory

The leaderboard

Across all local models the easiest language was English (2.8 of 3 on average), the hardest Polish (1.8 of 3). Each language column scores three answers (0–3). Hover or tap a cell to read the model's own words. Click Score, tok/s or GB to sort.

Local models, sorted by score then speed
#ModelENDEFRPLPTESUKBusWordsWallScoretok/sGB
1AMALIA 9BPortuguese public consortium3 of 33 of 33 of 33 of 33 of 33 of 33 of 3rightrightright24/24225.6
2Gemma 3 12BGoogle3 of 33 of 33 of 33 of 33 of 33 of 33 of 3rightrightright24/24168.1
3MamayLM 12BINSAIT, on Google Gemma 33 of 33 of 33 of 33 of 33 of 33 of 33 of 3rightrightright24/24157.3
4EuroLLM 9BEU research consortium3 of 33 of 33 of 33 of 33 of 33 of 33 of 3rightrightwrong23/24215.6
5Tower+ 9BUnbabel, Lisbon3 of 33 of 33 of 33 of 33 of 33 of 33 of 3rightrightwrong23/24185.8
6Aya Expanse 8BCohere3 of 33 of 33 of 33 of 33 of 33 of 33 of 3righthalf rightwrong22.5/24255.1
7Qwen3 14BAlibaba3 of 33 of 33 of 32 of 33 of 33 of 33 of 3righthalf rightright22.5/24149.3
8Qwen3 8BAlibaba3 of 33 of 32.5 of 32 of 32.5 of 33 of 33 of 3righthalf rightright21.5/24255.2
9Mistral Nemo 12BMistral AI and NVIDIA3 of 33 of 33 of 32 of 33 of 33 of 32.5 of 3rightrightwrong21.5/24237.1
10T-lite 2.1 8BT-Bank3 of 33 of 33 of 31 of 33 of 33 of 33 of 3righthalf rightright21.5/24225.0
11Gemma 3 4BGoogle3 of 32.5 of 32 of 32.5 of 32.5 of 33 of 32.5 of 3half rightrightright20.5/24413.3
12GigaChat 3.1 Lightning 10BSber3 of 33 of 33 of 31 of 32 of 33 of 33 of 3righthalf rightwrong19.5/24376.5
13Phi-4 14BMicrosoft3 of 32 of 32 of 32 of 33 of 32 of 31.5 of 3rightrightright18.5/24159.1
14DeepSeek-R1 8BDeepSeek3 of 32.5 of 32 of 31 of 32 of 32.5 of 32 of 3rightrightright18/24225.2
15Nemotron 3 Nano 4BNVIDIA3 of 33 of 32 of 31 of 31.5 of 33 of 32.5 of 3wrongwrongright17/24322.8
16Granite 3.3 8BIBM2.5 of 33 of 32 of 31.5 of 31.5 of 33 of 31 of 3wrongrightright16.5/24244.9
17YandexGPT 5 Lite 8BYandex3 of 33 of 31.5 of 30 of 31 of 33 of 33 of 3wrongwrongright15.5/24234.9
18Gervásio 8B PTPTUniversity of Lisbon2 of 30.5 of 31 of 31 of 32 of 33 of 32 of 3righthalf rightwrong13/24214.9
19Llama 3.1 8BMeta1 of 30.5 of 30.5 of 30.5 of 32 of 31.5 of 31 of 3righthalf rightwrong8.5/24254.9
20Llama 3.2 3BMeta2.5 of 31.5 of 30.5 of 30 of 31 of 31 of 31 of 3wrongwrongwrong7.5/24602.0
3 all right in that language1–2½ some wrong or in another language0 none✓ ◐ ✗ specialist checkstok/s — tokens per second on this Mac · GB — size on disk

Speed

How fast they write

Tokens per second while answering, averaged over the six questions. Comfortable reading is about 10–15.

Llama 3.2 3B59.6
Gemma 3 4B40.8
GigaChat 3.1 Lightning 10B37.4
Nemotron 3 Nano 4B31.7
Llama 3.1 8B24.9
Qwen3 8B24.8
Aya Expanse 8B24.8
Granite 3.3 8B24.3
Mistral Nemo 12B23
YandexGPT 5 Lite 8B22.9
T-lite 2.1 8B22.1
AMALIA 9B21.9
DeepSeek-R1 8B21.5
Gervásio 8B PTPT21.4
EuroLLM 9B20.9
Tower+ 9B17.7
Gemma 3 12B16.4
Phi-4 14B15.4
MamayLM 12B14.9
Qwen3 14B14.2

Memory

What fits in 16 GB

macOS lets the graphics side use about 12 GB of the 16. Anything bigger spills onto the processor and into swap.

~11.8 GB the GPU may use
Mistral Small 24B14 too big
gpt-oss 20B13 too big
Qwen3 14B9.3
Phi-4 14B9.1
Gemma 3 12B8.1
MamayLM 12B7.3
Mistral Nemo 12B7.1
GigaChat 3.1 Lightning 10B6.5
Tower+ 9B5.8
AMALIA 9B5.6
EuroLLM 9B5.6
Qwen3 8B5.2
DeepSeek-R1 8B5.2
Aya Expanse 8B5.1
T-lite 2.1 8B5
YandexGPT 5 Lite 8B4.9
Gervásio 8B PTPT4.9
Llama 3.1 8B4.9
Granite 3.3 8B4.9
Gemma 3 4B3.3
Nemotron 3 Nano 4B2.8
Llama 3.2 3B2

We tried anyway. gpt-oss 20B with half its layers on the GPU answered correctly — at about one token a second, with free memory down to 14%. Mistral Small 24B, a dense model, had not finished one short answer after fifteen minutes.

Cloud · Cloudflare Workers AI

The same questions, asked of the cloud

Bigger versions of the same families, run on Cloudflare's GPUs, on the free plan. DeepSeek V4, Kimi K2.6 and GLM 5.3 sit behind the paid plan and are not here.

Cloud models on Workers AI, same questions
ModelENDEFRPLPTESUKBusWordsWallScore
Llama 3.3 70BMeta3 of 33 of 33 of 33 of 33 of 33 of 33 of 3rightrightright24/24
Gemma 4 26B (MoE)Google3 of 33 of 33 of 33 of 33 of 33 of 33 of 3rightrightright24/24
Qwen3.8 27BAlibaba3 of 33 of 33 of 33 of 33 of 33 of 33 of 3rightrightright24/24
Mistral Small 3.1 24BMistral AI3 of 33 of 33 of 33 of 33 of 33 of 33 of 3rightrightwrong23/24
gpt-oss 120BOpenAI3 of 33 of 33 of 32 of 33 of 33 of 33 of 3rightrightright23/24
Qwen3 30B (MoE)Alibaba3 of 33 of 33 of 33 of 32.5 of 33 of 33 of 3wrongrightright22.5/24

Verdict

Which one for what

Best value that fits 16 GB

AMALIA 9B (Portugal)

24 out of 24 — every language, every check — in 5.6 GB at 22 tokens a second. Portugal's public model, and the most even one here.

Also perfect, a little heavier

Gemma 3 12B · MamayLM 12B

Both 24 out of 24. About 15 tokens a second, and they want most of the memory: close other apps first.

Fast and light

Gemma 3 4B (Google)

20.5 out of 24 at 39 tokens a second in 3.3 GB. The one to leave loaded for quick questions.

European Portuguese

AMALIA 9B · Tower+ 9B

Both knew every Portugal-only word; Tower+, from Lisbon's Unbabel, writes the most natural translations.

Many languages, small footprint

Aya Expanse 8B (Cohere)

Full marks in all seven languages; only the Portuguese words and the English sum let it down.

When the laptop is not enough

Gemma 4 26B · Llama 3.3 70B (cloud)

24 out of 24 on Workers AI's free plan; Llama 3.3 answers in seconds.

Reading the labels

What the names mean

8b, 14b

Billions of parameters — the model's size. In the usual 4-bit form a model needs about 0.6 GB per billion, so 14b is roughly 9 GB.

30b-a3b

Mixture of experts: 30 billion in total, 3 billion working on each word. Fast like a small model, but all 30 billion still have to sit in memory.

q4_K_M

Compression: about 4 bits per parameter instead of 16. Four times lighter for a small loss. q8 is near-lossless, q2–q3 noticeably weaker.

Memory bandwidth

Decides speed. Each word means reading the whole model once, so tokens per second ≈ bandwidth ÷ model size. Memory size decides what runs at all.

Method

How this was done, and what it does not prove

Local models ran through Ollama on MacBook Pro, Apple M1 Pro, 16 GB unified memory, one at a time, with a watchdog that unloads a model if free memory falls below 6%. Every model got the same short system prompt and temperature 0. Cloud models ran through Cloudflare Workers AI with the same prompt. Answers were graded by fixed rules, including a check of which language each reply is in; every raw answer is public: local.json, cloud.json.

24 questions are a probe, not an exam. They reward knowing European Portuguese and answering in the reader's language, and they penalise a single slip; a model can be excellent at things not asked here. Speeds are for this one laptop.