/inference/hf.co-maziyarpanahi-qwen3-0.6b-gguf-q3_k_m

hf.co/MaziyarPanahi/Qwen3-0.6B-GGUF:Q3_K_M

Measured by one node on hardware we do not own, and signed by each of them. What machines here actually proved — never what the card claims.

matrix built 21 Aug 2026 · every field verified against the signature of the node that produced it, and re-verifiable by you: each signed figure carries the exact bytes the node signed (signing_payload_b64) and the signature over them, and the ed25519 public key is embedded in node_did. effective_ctx is the context length a node PROVED by recall, never the length a model card advertises
context declared40,960what the runtime reported. A node attests it was told this; nobody attests it is true.
context proved by recallnot probedthe largest size at which a machine could still find tokens planted across the whole prompt — start, middle and end, all of which had to come back.
Both numbers are needed to state a gap, and one of them is missing.

Languages, proved

English

Each question was put entirely in its own language and script, and the answer had to be right. A model with no coverage cannot parse the question, so answering in English is a failure here, not a pass. 1 proved of 10 asked — a language absent from the probe set was never asked, and is not a claim either way.

Capabilities, measured

Reasons about codeprobed, and failed
Asked to predict exactly what a short program prints. This is the half of coding that matters for agent work and that writing valid syntax does not demonstrate: following state through control flow.
Writes codepassed
Asked for a specific function and checked with the Go compiler’s own parser. Nothing is executed — running model-written code to score a model would be a security hole opened for a metric — but syntactic validity is a fact rather than an opinion.
Follows instructionsprobed, and failed
Given a plain constraint on its output format and checked for exact obedience. This separates an instruct-tuned model from a base one — a base model answers fluently and ignores every instruction, which then looks like a dozen unrelated weaknesses instead of one.
Sustained outputpassed
Asked to produce a long structured answer and checked for length. A model that stops after a couple of hundred tokens cannot write a report, however capable it is otherwise.
Structured outputpassed
The ENGINE was asked to constrain decoding to a schema, and the answer conformed. This is stronger than politely asking for JSON and hoping: the grammar makes invalid output unreachable, which is what an operator actually builds on.
Extended thinkingpassed
The model emitted reasoning on the engine’s dedicated reasoning channel before answering. Read from that channel rather than by scanning the answer for tags, which is why models that reason are no longer reported as models that do not.
Synthesises sourcespassed
Given three documents with one fact split across them, and required to combine rather than quote. No single document contains the answer, so a model that retrieves without reasoning cannot pass.
Tool callingpassed
The node asked the model to call a function and checked that it called the RIGHT one with a well-formed argument — not merely that it emitted a tool name.
Faithful to tool resultspassed
The model was handed a list through a tool and had to report every item back. Calling a tool and understanding its answer are different abilities: a model can produce a perfect call, receive nine items, and confidently report three. This is the flag that separates them.
Tool loop finishedpassed
Whether the conversation ENDED after the tool call, or the model kept calling until the turn budget ran out. A model that passes tool calling and fails this will call a tool in production and never come back.
Chained tool callsprobed, and failed
The model had to call one tool, read its answer, and call a second with a value derived from it. One call is not agentic work — real specialist jobs are sequences, and a model that never takes the second step looks, from outside, exactly like a specialist silently doing nothing.
Visionprobed, and failed
Shown an image and asked what was in it. The node checked the answer.
Audioprobed, and failed
Given audio and asked to transcribe it.
faithful_tool_resultpassed

The qwen3 family

siblings measured on the same network — compare what each actually proved
signed by this network

What machines here proved

best throughput44.56 tok/s
machines that measured it1
samples behind those numbers1
regions it was proved in1
Every line traces to a node signature.
from the public record · unsigned

What the runtime says about it

context declared40,960
parameters752M
parameter count75,16,32,384
quantisationQ3_K_M
familyqwen3
formatgguf
ollama /api/show; the node attests that the runtime reported these, not that they are true. Attested by did:epn:002408…9e3c79.The only unsigned block on this page. We attest we read these from the runtime — not that they are true. Nobody signs a model card.

Proven across regions

The machine that measured it

nothing is averaged — the machine is what you are choosing
did:epn:002408…9e3c79Hyderabad · 500050 · ollama · 21 Aug 2026
ctx provedprobe kADgzPcfCz
44.56tok/s11 samples · 38s
Reasons about code Writes code Follows instructions Sustained output Structured output Extended thinking Synthesises sources Tool calling Faithful to tool results Tool loop finished Chained tool calls Vision Audio faithful_tool_result1 languages

Context ladders attempted: did:epn:002408…9e3c79 8192,4096