ModelHubby scout badgeModelHubbyFIELD MANUAL · EST. 2026
Measured 2026-08-19

llama3.2:latest

Running locally on Apple M4 · 24 GB unified memory, llama3.2:latest passed 21 of 22 capability tasks (95%), with a median latency of 1,521 ms and 16,052 tokens of context actually read.

Harness 1.1.0 · deterministic graders, no LLM judge · how this is measured

By category

categorypassedrateflaky tasksmedian latencywhat it means
code3/3100%1,504 msWrote code that executes and returns the expected value.
longcontext4/4100%16,326 msActually read a fact buried in a long prompt, rather than truncating it.
instruction5/5100%492 msFollowed a literal instruction exactly (casing, word choice, length).
json3/475%888 msEmitted parseable, schema-correct JSON with no prose wrapper.
reasoning4/4100%2,608 msGot a deterministic answer right — one with a single checkable result.
refusal2/2100%8,990 msAnswered a benign question instead of refusing it spuriously.

Failed every attempt

  • json-extract (json) price should be 24.5, got 8.17

What this review does not say

Nothing here measures prose quality, creativity, or taste — the graders are deterministic by design, which is what makes the numbers reproducible. A high instruction score does not mean this model writes well. It means it did what it was told, on Apple M4, on 2026-08-19, and you can re-run the suite and check.