9 of 25 models asked a question written entirely in Malayalam, in its own script, answered it correctly. Answering in English is a failure here, not a pass — a model with no coverage cannot parse the question in the first place. Every verdict below is signed by the machine that took it.
Best recall-proved context first — the order in which you would try them.
On the page on purpose. These are the models a “multilingual” label would have sold you, and the reason the list above is worth anything.
A model absent from both lists was never asked in Malayalam, and that is not a claim either way. The probe set covers a fixed list of languages; a script missing from it is missing from the evidence, not from the model.
← Every model this network has measured