language · proved by recall of meaning English 26 of 29 models asked a question written entirely in English, in its own script, answered it correctly. Answering in English is a failure here, not a pass — a model with no coverage cannot parse the question in the first place. Every verdict below is signed by the machine that took it.
26 of 29 models answered correctly · 3 asked, and could notAnswered correctly Best recall-proved context first — the order in which you would try them.
hf.co/bartowski/Llama-3.2-1B-Instruct-GGUF:Q4_K_M 16,384 ctx · 103.9 tok/s · 2 machines hf.co/unsloth/gemma-4-E2B-it-GGUF:IQ4_NL 8,192 ctx · 54.9 tok/s · 2 machines llama3.2:1b 8,192 ctx · 49.9 tok/s · 5 machines hf.co/bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF:IQ3_XS 8,192 ctx · 46.8 tok/s · 1 machine gemma4:e2b 8,192 ctx · 45.1 tok/s · 5 machines hf.co/unsloth/gemma-4-E2B-it-GGUF:Q3_K_M 8,192 ctx · 38.4 tok/s · 3 machines hf.co/bartowski/microsoft_Phi-4-mini-instruct-GGUF:Q3_K_L 8,192 ctx · 32.1 tok/s · 1 machine hf.co/unsloth/gemma-4-E2B-it-GGUF:Q4_K_M 8,192 ctx · 11.3 tok/s · 2 machines hf.co/MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF:IQ3_XS 4,096 ctx · 107.5 tok/s · 2 machines hf.co/unsloth/gemma-4-E2B-it-GGUF:Q4_0 4,096 ctx · 57.6 tok/s · 2 machines hf.co/unsloth/Qwen3-VL-4B-Instruct-GGUF:Q4_0 4,096 ctx · 40.9 tok/s · 3 machines hf.co/MaziyarPanahi/Qwen3-0.6B-GGUF:Q4_K_M 4,096 ctx · 37.4 tok/s · 1 machine hf.co/bartowski/microsoft_Phi-4-mini-instruct-GGUF:IQ2_M 4,096 ctx · 37.3 tok/s · 2 machines hf.co/unsloth/Qwen3-VL-4B-Instruct-GGUF:Q4_1 4,096 ctx · 37.0 tok/s · 2 machines hf.co/MaziyarPanahi/Phi-4-mini-instruct-GGUF:Q3_K_S 4,096 ctx · 27.9 tok/s · 2 machines hf.co/mradermacher/Qwen3-VL-8B-Instruct-abliterated-GGUF:Q2_K 4,096 ctx · 22.4 tok/s · 3 machines hf.co/unsloth/gemma-4-E4B-it-GGUF:Q3_K_M 4,096 ctx · 20.4 tok/s · 1 machine hf.co/unsloth/Qwen3-VL-4B-Instruct-GGUF:IQ4_XS 4,096 ctx · 6.4 tok/s · 1 machine hf.co/unsloth/Qwen3-4B-GGUF:IQ3_XXS 4,096 ctx · 0.9 tok/s · 1 machine hf.co/MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF:Q4_K_M context not probed · 113.3 tok/s · 2 machines hf.co/MaziyarPanahi/Qwen3-0.6B-GGUF:Q3_K_M context not probed · 44.6 tok/s · 1 machine hf.co/unsloth/Phi-4-mini-reasoning-GGUF:Q2_K context not probed · 37.7 tok/s · 2 machines hf.co/unsloth/Phi-4-mini-reasoning-GGUF:Q3_K_M context not probed · 31.9 tok/s · 3 machines hf.co/unsloth/Phi-4-mini-reasoning-GGUF:IQ1_S context not probed · 26.4 tok/s · 4 machines hf.co/unsloth/Qwen3-VL-4B-Instruct-GGUF:Q2_K context not probed · 11.9 tok/s · 1 machine hf.co/bartowski/Llama-3.2-1B-Instruct-GGUF:F16 context not probed · 8.0 tok/s · 1 machine Asked, and could not On the page on purpose. These are the models a “multilingual” label would have sold you, and the reason the list above is worth anything.
A model absent from both lists was never asked in English, and that is not a claim either way. The probe set covers a fixed list of languages; a script missing from it is missing from the evidence, not from the model.
← Every model this network has measured