/inference/hf.co-lmstudio-community-phi-4-reasoning-plus-gguf-q4_k_m

hf.co/lmstudio-community/Phi-4-reasoning-plus-GGUF:Q4_K_M

Measured by one node on hardware we do not own, and signed by each of them. What machines here actually proved — never what the card claims.

matrix built 16 Sept 2026 · every field verified against the signature of the node that produced it, and re-verifiable by you: each signed figure carries the exact bytes the node signed (signing_payload_b64) and the signature over them, and the ed25519 public key is embedded in node_did. effective_ctx is the context length a node PROVED by recall, never the length a model card advertises
context declarednot declaredwhat the runtime reported. A node attests it was told this; nobody attests it is true.
context proved by recallnot probedthe largest size at which a machine could still find tokens planted across the whole prompt — start, middle and end, all of which had to come back.
Both numbers are needed to state a gap, and one of them is missing.

Capabilities, measured

Tool callingprobed, and failed
The node asked the model to call a function and checked that it called the RIGHT one with a well-formed argument — not merely that it emitted a tool name.
Tool loop finishedprobed, and failed
Whether the conversation ENDED after the tool call, or the model kept calling until the turn budget ran out. Passing “tool calling” and failing this means the model will call a tool in production and never come back. A model chosen on the first flag alone is the reason an agent silently stops delivering work.
Structured outputprobed, and failed
Asked for JSON matching a schema, and got it. This is the difference between a model you can build on and one you must parse defensively.
Extended thinkingprobed, and failed
The model emits reasoning before its answer. Measured where the probe could observe it.
Visionprobed, and failed
Shown an image and asked what was in it. The node checked the answer.
Audioprobed, and failed
Given audio and asked to transcribe it.
Faithful to tool resultsprobed, and failed
The model was handed a list through a tool and had to report every item back. Calling a tool and understanding its answer are different abilities: a model can produce a perfect call, receive nine items, and confidently report three.
signed by this network

What machines here proved

machines that measured it1
regions it was proved in1
Every line traces to a node signature.
from the public record · unsigned

What the runtime says about it

No runtime metadata was attested for this model.The only unsigned block on this page. We attest we read these from the runtime — not that they are true. Nobody signs a model card.

Proven across regions

The machine that measured it

nothing is averaged — the machine is what you are choosing
did:epn:002408…afbb8aLalganj · 834002 · ollama-host · 15 Sept 2026Apple M4 · 10 cores · 16.0 GiB RAM
0ctx provedno recall at 16,384 · 8,192 · 4,096
tok/s now
Tool calling Tool loop finished Structured output Extended thinking Vision Audio Faithful to tool results

Context ladders attempted: did:epn:002408…afbb8a 16384,8192,4096