8 of 27 models asked a question written entirely in Bengali, in its own script, answered it correctly. Answering in English is a failure here, not a pass — a model with no coverage cannot parse the question in the first place. Every verdict below is signed by the machine that took it.
8 of 27 models answered correctly · 19 asked, and could not
Answered correctly
Best recall-proved context first — the order in which you would try them.
A model absent from both lists was never asked in Bengali, and that is not a claim either way. The probe set covers a fixed list of languages; a script missing from it is missing from the evidence, not from the model.