/inference/hf.co-maziyarpanahi-llama-3.2-1b-instruct-gguf-iq3_xs

hf.co/MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF:IQ3_XS

Measured by one node on hardware we do not own, and signed by each of them. What machines here actually proved — never what the card claims.

matrix built 21 Aug 2026 · every field verified against the signature of the node that produced it, and re-verifiable by you: each signed figure carries the exact bytes the node signed (signing_payload_b64) and the signature over them, and the ed25519 public key is embedded in node_did. effective_ctx is the context length a node PROVED by recall, never the length a model card advertises
context declared1,31,072what the runtime reported. A node attests it was told this; nobody attests it is true.
context proved by recallnot probedthe largest size at which a machine could still find tokens planted across the whole prompt — start, middle and end, all of which had to come back.
Both numbers are needed to state a gap, and one of them is missing.

Languages, proved

English

Each question was put entirely in its own language and script, and the answer had to be right. A model with no coverage cannot parse the question, so answering in English is a failure here, not a pass. 1 proved of 10 asked — a language absent from the probe set was never asked, and is not a claim either way.

Capabilities, measured

Reasons about codeprobed, and failed
Asked to predict exactly what a short program prints. This is the half of coding that matters for agent work and that writing valid syntax does not demonstrate: following state through control flow.
Writes codeprobed, and failed
Asked for a specific function and checked with the Go compiler’s own parser. Nothing is executed — running model-written code to score a model would be a security hole opened for a metric — but syntactic validity is a fact rather than an opinion.
Follows instructionspassed
Given a plain constraint on its output format and checked for exact obedience. This separates an instruct-tuned model from a base one — a base model answers fluently and ignores every instruction, which then looks like a dozen unrelated weaknesses instead of one.
Sustained outputpassed
Asked to produce a long structured answer and checked for length. A model that stops after a couple of hundred tokens cannot write a report, however capable it is otherwise.
Structured outputpassed
The ENGINE was asked to constrain decoding to a schema, and the answer conformed. This is stronger than politely asking for JSON and hoping: the grammar makes invalid output unreachable, which is what an operator actually builds on.
Extended thinkingprobed, and failed
The model emitted reasoning on the engine’s dedicated reasoning channel before answering. Read from that channel rather than by scanning the answer for tags, which is why models that reason are no longer reported as models that do not.
Synthesises sourcesprobed, and failed
Given three documents with one fact split across them, and required to combine rather than quote. No single document contains the answer, so a model that retrieves without reasoning cannot pass.
Tool callingprobed, and failed
The node asked the model to call a function and checked that it called the RIGHT one with a well-formed argument — not merely that it emitted a tool name.
Tool loop finishedprobed, and failed
Whether the conversation ENDED after the tool call, or the model kept calling until the turn budget ran out. Passing “tool calling” and failing this means the model will call a tool in production and never come back. A model chosen on the first flag alone is the reason an agent silently stops delivering work.
Visionprobed, and failed
Shown an image and asked what was in it. The node checked the answer.
Audioprobed, and failed
Given audio and asked to transcribe it.
faithful_tool_resultprobed, and failed

The llama family

siblings measured on the same network — compare what each actually proved
signed by this network

What machines here proved

best throughput21.74 tok/s
machines that measured it1
samples behind those numbers1
regions it was proved in1
Every line traces to a node signature.
from the public record · unsigned

What the runtime says about it

context declared1,31,072
parameters1.24B
parameter count1,23,58,14,432
quantisationIQ3_XS
familyllama
formatgguf
ollama /api/show; the node attests that the runtime reported these, not that they are true. Attested by did:epn:002408…9e3c79.The only unsigned block on this page. We attest we read these from the runtime — not that they are true. Nobody signs a model card.

Proven across regions

The machine that measured it

nothing is averaged — the machine is what you are choosing
did:epn:002408…9e3c79Hyderabad · 500050 · 0.31.2 · 21 Aug 2026
ctx provedprobe JsbXA/PfTW
21.74tok/s1 samples · 2s
Reasons about code Writes code Follows instructions Sustained output Structured output Extended thinking Synthesises sources Tool calling Tool loop finished Vision Audio faithful_tool_result1 languages

Context ladders attempted: did:epn:002408…9e3c79 8192,4096