Compare models on the Artificial Analysis benchmark that matters to you. Every local family stays on the graph: runnable models use their highest-fidelity comfortable build, while models beyond your Mac are muted at their smallest available build. The cloud rail traces Claude, Gemini, and GPT intelligence over time.
Best local fit · adapts to RAMIntelligence
Qwen3.8 27B
Q8_0 · 29.1 GB · 33.9 Intelligence
Highest-scoring model, using its highest nominal-precision build that fits. Working estimate uses 45% of your 64 GB unified memory.
8 GB512 GB
Artificial Analysis Intelligence Index v4.3: ten evaluations of reasoning, knowledge, coding, and agentic work. Refreshed September 18, 2026, with models released since read from the same index version on September 22; AA estimates are marked.
28/50model families runnable on this MacAll other families stay visible in gray.
THE BENCHMARK MAP
Intelligence Index vs. local size.
Open weights Beyond this Mac Anthropic Google OpenAI Pareto frontier♥ Repo activity on hover1× One point per model family
Intelligence score ↑
0
12
24
36
48
60
2G
4G
8G
16G
32G
64G
128G
256G
512G
1T
2T
64 GB Mac comfort line
Published model artifact size → log scaleOne point per family · gray marks need more memory than your Mac.
COMFORTABLE FIT · PRIORITY
Qwen3.8 27B
A dense vision-language Qwen model with 262K native context and MTP-trained decoding. Scores shown here come from Artificial Analysis.
Qwen3.6 was trained with extra prediction heads. A compatible runtime can use them for speculative decoding: propose several future tokens, verify them together, and accept the ones the main model agrees with.
HOW GENERATION CHANGES
01Draft
The MTP heads predict the next few tokens cheaply.
→
02Verify
The full model checks those candidates in parallel.
→
03Accept
Correct candidates advance generation several tokens.
This changes decoding, not the Artificial Analysis capability score. Rejected drafts are discarded, so Unsloth reports no accuracy change from enabling MTP.
UNSLOTH REPORTED RANGE1.4–2.2×
faster token generation across Qwen3.6 MTP configurations
Standard decode1.0×
MTP · 2 draft tokens1.4× avg.
Typical dense-model gain1.4×
Extra headroom~1 GB
Recommended drafts2
SAME MODEL, TWO DECODING PATHS
Standard weights or MTP heads?
♥ Loading Hugging Face likes…
Both repositories contain Qwen3.6 35B-A3B. They share the same capability score; the MTP build adds prediction heads that a compatible runtime can use to accelerate generation.
+0.49 GBextra on disk for UD-Q4_K_XL
STANDARD GGUFFits 64 GB Mac
22.4 GB
Baseline decoding. No MTP-specific single-slot or projector restriction.
File sizes are summed from matching GGUFs in both Hugging Face repositories, checked July 26, 2026. The disk delta is exact for the selected artifact; speed and extra runtime memory are Unsloth-reported ranges. Like counts refresh from the Hugging Face Hub API and are summed across repositories, not deduplicated users.
Speed is hardware-dependent. Unsloth recommends trying draft counts 1–6, although two performed best on average. Their measured acceptance fell from 83% at two drafts to 50% at four, making extra drafts less useful.
Current llama.cpp support requires MTP-aware settings. Parallel slots above one and the multimodal projector are not yet supported with this path.
Actual published artifact sizes from unsloth/Qwen3.8-27B-GGUF. Fit reserves 20% of unified memory for macOS and inference overhead.
Q4Fits
17.1 GB
Q4_K_M
Q6Fits
22.9 GB
Q6_K
Q8Fits
29.1 GB
Q8_0
BF16Too large
55.6 GB
BF16
File size is not peak runtime memory. Context length, KV cache, multimodal projectors, and the inference engine add overhead. The benchmark score belongs to the base model—not to each quantized file.
THE FULL MTP REPOSITORY
22 ways to store 27B.
Filter by nominal bit depth. “Working estimate” adds Unsloth’s recommended ~1 GB MTP headroom to the actual file size.
Unsloth Dynamic: model-specific mixed precision, assigning more bits to sensitive layers.
IQ
Importance-aware quantization, designed to retain more useful signal at very low bit counts.
K
llama.cpp superblock quantization. S, M, and XL distinguish recipes that spend progressively more storage.
Q4_0 / Q4_1
Older uniform 4-bit formats. Widely compatible, but less selectively allocated than newer recipes.
Do not rank these files by tokens/second from size alone. Unsloth warns its per-quant throughput measurements are noisy; MTP draft acceptance, hardware bandwidth, and runtime settings can matter more than a small size difference.
CURRENT VIEW
Ranked for your Mac.
Open models with both a Intelligence Index score and a published Q4_K_M artifact, sorted by the selected benchmark.
ModelIntelligenceArtifact sizeFit
READ THE FINE PRINT
Two sources, two different claims.
01
Scores come from Artificial Analysis
The selector changes the independent model evaluation shown on the vertical axis. Missing evaluations are omitted, never treated as zero. Intelligence scores use the v4.3 index, refreshed September 18, 2026, with models released since added from the same index version on September 22. Values estimated by Artificial Analysis are marked “AA estimate”; inherited quantized-model scores are marked “base score.” Coding and Agentic are archived September 1 scores, and remain separate from the current index.
02
Sizes come from Hugging Face
Artifact sizes are the summed byte sizes of the named files in the linked repositories, including multi-part builds and native quantized checkpoints.
03
Quant scores are not implied
Artificial Analysis evaluates a model endpoint or base model. The app does not pretend every Q4, Q6, or Q8 artifact preserves that score exactly.
04
Fit leaves working room
“Comfortable” keeps 20% of unified memory free. Actual runtime needs still vary with context and cache precision. Each family appears once: the highest-fidelity comfortable artifact when one fits, otherwise its smallest available artifact in gray.