Roster refreshed 21 August 2026

Current Ollama models.
One comparable KEMPO protocol.

This curated, non-exhaustive roster names one exact local artifact for each listed selected size through 120B. New target requests omit the think field for every model, allowing each model to use its own default behavior while isolation, requested context, temperature, system message, and question set remain fixed. The recorded campaign runtime is Ollama 0.32.15.

36
explicit local targets through 120B
100
isolated questions each
131K
requested context tokens

Evaluation status

This page shows what is being tested now.

Models is the source of truth for exact local artifacts, completed response sets, active protocol status, retest requirements, preflight needs, and retained historical evidence.

Standardized local campaign

Seven complete response sets await human scoring.

Six cards preserve answer-complete August 2026 evidence produced with an explicit thinking: low request and therefore require model-default reruns. Granite 4.1 3B is the first answer-complete run under the active model-default protocol. Every set contains 100 responses; independent human scoring is still pending.

Answers complete2.1 GB · 131K advertised · model-default

IBM · active protocol

Granite 4.1 · 3B

granite4.1:3b-q4_K_M

Next stageAwaiting human scoring

The first completed active model-default run, retained separately from the six historical explicit-low response sets.

Official Granite 4.1 Ollama entry
Answers complete18 GB · 256K advertised · Ollama ≥0.32.12

Alibaba Qwen

Qwen 3.8 · 27B

qwen3.8:27b-q4_K_M

Next stageAwaiting human scoring

Newest Qwen local release in the roster, selected as a current reasoning and multimodal representative.

Official Qwen 3.8 Ollama entry
Answers complete18 GB · 128K advertised · Ollama ≥0.32.8

Meta

Muse Glimmer · 30B

muse-glimmer:30b-q4_K_M

Next stageAwaiting human scoring

Meta's current local-agent model, included as a new safety-and-judgment comparison at workstation scale.

Official Muse Glimmer Ollama entry
Answers complete7.6 GB · 256K advertised

Google DeepMind

Gemma 4 · 12B

gemma4:12b-it-q4_K_M

Next stageAwaiting human scoring

A current mid-size Gemma target that fits fully within the local hardware envelope while retaining the campaign context.

Official Gemma 4 Ollama entry
Answers complete3.4 GB · 256K advertised

Alibaba Qwen

Qwen 3.5 · 4B

qwen3.5:4b-q4_K_M

Next stageAwaiting human scoring

Small-model control for the normalized lane and a direct generational comparison with Qwen 3.8.

Official Qwen 3.5 Ollama entry
Answers complete14 GB · 128K advertised

OpenAI · August explicit-low run

GPT-OSS · 20B

gpt-oss:20b

Next stageAwaiting human scoring

Retested because an earlier GPT-OSS 20B response set exists, allowing the new isolated protocol to stand beside historical evidence.

Official GPT-OSS 20B Ollama entry
Answers complete14 GB · 128K advertised

OpenAI · August explicit-low run

GPT-OSS Safeguard · 20B

gpt-oss-safeguard:20b

Next stageAwaiting human scoring

Retested as the safety-specialized counterpart to GPT-OSS 20B, without carrying forward its earlier total as a new score.

Official GPT-OSS Safeguard 20B Ollama entry

Every model above has 100 stored responses and a complete deterministic sentence-boundary audit. None has a human KEMPO score yet. The six explicit-low runs require model-default reruns before active-protocol comparison; Granite 4.1 3B does not. No critical fail recorded: only a literal human FAIL creates one.

Comparison controls

Fresh context is part of the result.

  1. Pin identity

    Each execution records the exact Ollama tag, full installed digest, byte size, quantization metadata, and Ollama version.

  2. Prove isolation

    Unload and verify absence before every prompt, send exactly one system and one user message, then unload and verify again.

  3. Use model-default thinking

    Request 131,072 context tokens and temperature 0.8, omit the think field for every model, run one sequential target, and use the exact 31-byte system message.

  4. Separate judgment

    Derive one versioned sentence-boundary measurement from every stored answer, while retaining earlier AI-checker output only as diagnostics. Human KEMPO scoring remains a later, accountable evaluation.

Local family size ladder

Thirty-six explicit artifacts. Every selected size through 120B.

Each row is a local Ollama target, not a cloud alias. Physical parameter count is recorded separately from the model’s public size label, which matters for effective-size and mixture-of-experts names. “Preflight required” means admission is scheduled for a low-contention system window; it does not deny this machine’s capability.

Active requests use 131,072 context tokens and omit think. Historical explicit-low answers are disclosed separately from active run status.
Exact local Ollama tagLabel · physical parametersArtifactAdvertised native context · inputActive model-default status
Granite 4.1 · IBM
granite4.1:3b-q4_K_M3B · 3.4B2.1 GB · Q4_K_M131,072 · textModel-default answers complete · 100 / 100 · awaiting human scoring
granite4.1:8b-q4_K_M8B · 8.79B5.3 GB · Q4_K_M131,072 · textInstalled · admission preflight pending · not started; old explicit-low admission failure retained
granite4.1:30b-q4_K_M30B · 28.9B17 GB · Q4_K_M131,072 · textInstalled · admission preflight pending · not started
Qwen 3.5 · Alibaba Qwen
qwen3.5:0.8b-q8_00.8B · 0.873B1.0 GB · Q8_0262,144 · text, imageNot started
qwen3.5:2b-q4_K_M2B · 2.27B1.9 GB · Q4_K_M262,144 · text, imageNot started
qwen3.5:4b-q4_K_M4B · 4.66B3.4 GB · Q4_K_M262,144 · text, imageRetest required · historical explicit-low answers complete
qwen3.5:9b-q4_K_M9B · 9.65B6.6 GB · Q4_K_M262,144 · text, imageNot started
qwen3.5:27b-q4_K_M27B · 27.8B17 GB · Q4_K_M262,144 · text, imageNot started
qwen3.5:35b-a3b-q4_K_M35B-A3B · 36B24 GB · Q4_K_M262,144 · text, imagePreflight required · not started
Qwen 3.8 · Alibaba Qwen
qwen3.8:27b-q4_K_M27B · 27.3B18 GB · Q4_K_M262,144 · text, imageRetest required · historical explicit-low answers complete
Gemma 4 · Google DeepMind
gemma4:e2b-it-q4_K_ME2B effective · 5.12B physical7.2 GB · Q4_K_M131,072 · text, imageNot started
gemma4:e4b-it-q4_K_ME4B effective · 8B physical9.6 GB · Q4_K_M131,072 · text, imageNot started
gemma4:12b-it-q4_K_M12B · 11.9B7.6 GB · Q4_K_M262,144 · text, imageRetest required · historical explicit-low answers complete
gemma4:26b-a4b-it-q4_K_M26B-A4B · 25.8B18 GB · Q4_K_M262,144 · text, imageNot started
gemma4:31b-it-q4_K_M31B · 31.3B20 GB · Q4_K_M262,144 · text, imageNot started
GPT-OSS · OpenAI
gpt-oss:20b20B · 20.9B14 GB · MXFP4131,072 · textRetest required · historical explicit-low answers complete
gpt-oss:120b120B · 117B65 GB · MXFP4131,072 · textOperator-confirmed previously run · low-load preflight · not started
GPT-OSS Safeguard · OpenAI
gpt-oss-safeguard:20b20B · 20.9B14 GB · MXFP4131,072 · textRetest required · historical explicit-low answers complete
gpt-oss-safeguard:120b120B · 117B65 GB · MXFP4131,072 · textPreflight required · not started
DeepSeek-R1 · DeepSeek
deepseek-r1:1.5b-qwen-distill-q4_K_M1.5B · 1.78B1.1 GB · Q4_K_M131,072 · textNot started
deepseek-r1:7b-qwen-distill-q4_K_M7B · 7.62B4.7 GB · Q4_K_M131,072 · textNot started
deepseek-r1:8b-0528-qwen3-q4_K_M8B · 8.19B5.2 GB · Q4_K_M131,072 · textNot started · 0528 / Qwen3 variant
deepseek-r1:14b-qwen-distill-q4_K_M14B · 14.8B9.0 GB · Q4_K_M131,072 · textNot started
deepseek-r1:32b-qwen-distill-q4_K_M32B · 32.8B20 GB · Q4_K_M131,072 · textNot started
deepseek-r1:70b-llama-distill-q4_K_M70B · 70.6B43 GB · Q4_K_M131,072 · textPreflight required · not started
LFM 2.5 · Liquid AI
lfm2.5-thinking:1.2b-q4_K_M1.2B · 1.17B731 MB · Q4_K_M128,000 · textBlocked: native context below 131,072 request
lfm2.5:8b-a1b-q4_K_M8B-A1B · 8.47B5.2 GB · Q4_K_M128,000 · textBlocked: native context below 131,072 request
North Mini Code · NVIDIA
north-mini-code-1.0:q4_K_M30B · 30.5B19 GB · Q4_K_M500,000 · textNot started
Muse Glimmer · Meta
muse-glimmer:30b-q4_K_M30B · 27.9B18 GB · Q4_K_M131,072 · text, imageRetest required · historical explicit-low answers complete
Nemotron 3 Nano · NVIDIA
nemotron-3-nano:4b4B · 3.97B2.8 GB · Q4_K_M262,144 · textNot started
nemotron-3-nano:30b-a3b-q4_K_M30B-A3B · 31.6B24 GB · Q4_K_M1,048,576 · textPreflight required · not started
Nemotron 3.5 Lightning · NVIDIA
nemotron-3.5-lightning:30b-a3b-q4_K_M30B-A3B · 32.9B25 GB · Q4_K_M1,048,576 · textPreflight required · not started
Ministral 3 · Mistral AI
ministral-3:3b-instruct-2512-q4_K_M3B · 3.85B3.0 GB · Q4_K_M262,144 · text, imageNot started · Ollama ≥0.13.1
ministral-3:8b-instruct-2512-q4_K_M8B · 8.92B6.0 GB · Q4_K_M262,144 · text, imageNot started · Ollama ≥0.13.1
ministral-3:14b-instruct-2512-q4_K_M14B · 13.9B9.1 GB · Q4_K_M262,144 · text, imageNot started · Ollama ≥0.13.1
Mistral Small 3.2 · Mistral AI
mistral-small3.2:24b-instruct-2506-q4_K_M24B · 24B15 GB · Q4_K_M131,072 · text, imageLow-load preflight · not started

Separate lanes

Protocol, capacity, and evidence are separate facts.

3B model-default answers complete

Granite 4.1 · 3B, 8B, and 30B

All three selected Granite artifacts are installed and protocol-suitable under the revised model-default request. Granite 3B now has a canonical 100-answer model-default run with 100 deterministic sentence-limit records. The 8B and 30B artifacts remain installed with preflight pending and active status not started. The earlier 8B zero-answer failure remains historical evidence that explicit thinking: low was unsupported; it is not a present blocker.

Capacity-sensitive local

Low-contention preflight

Seven larger or not-yet-characterized targets require a low-system-load admission window. This lane distinguishes everyday-load operation from capacity-sensitive operation; it does not label the models unsupported or deny machine capability. GPT-OSS 120B also has a separately disclosed operator-confirmed previous local run, while its older KEMPO answer set remains historical evidence.

Context-insufficient

LFM 2.5 · two local artifacts

Both selected LFM artifacts advertise 128,000 native context tokens, below the active 131,072-token request. Their active status is blocked by that exact protocol mismatch, not by thinking support or a general hardware finding.

Outside the selected cap

Mistral Medium 3.5 · 128B reference

mistral-medium-3.5:128b-q4_K_M

The official 128B Q4_K_M artifact is retained as an outside-cap reference. It is excluded because this ladder stops at 120B, not because KEMPO has made a machine-capability finding.

Official Mistral Medium 3.5 Ollama entry

Retained evidence

Earlier runs remain historical; they are never silently mixed into the new cohort.

PRECRISIS-BH:3b PRECRISIS-BH:20b PRECRISIS-BH:120b GPT-OSS:20B GPT-OSS:120B GPT-OSS-SAFEGUARD:20B GPT-OSS-SAFEGUARD:120B Granite4:32B-A9B-H

These eight source-backed response sets predate the current local campaign. Their provenance and any recoverable historical totals remain labeled separately on the Scores page.

Follow the evidence

A queue is not a score.
A completed answer run is not yet human judgment.