M MAPLE PREVIEW
ON-DEVICE MODEL ↗
Featured in Hugging Face Spaces of the Week · Aug 10

20B—A1B TERNARY REASONING MODEL

Intelligence,
grown locally.

Maple-Preview runs entirely on your GPU. Your prompts stay on your device, while a 2-bit official checkpoint delivers fast, private reasoning.

403 MB of KV cache — 5.71 GB of GPU memory in total.

20B
PARAMETERS
1B
ACTIVE / TOKEN
2-bit
OFFICIAL WEIGHTS

Made with Opus 5 and GPT 5.6 sol

DEVICE READINESS

Checking your GPU…

Maple's official browser checkpoint is 5.31 GB. A device with at least 8 GB of unified or dedicated GPU memory is required; 12 GB or more is recommended.

LOCAL RUNTIME

Inside this session

CONTEXT WINDOW

Changing this clears the cached conversation prefix.

SAMPLING

Greedy decoding is exactly reproducible but can get stuck repeating itself on open-ended prompts; the sampled modes avoid that. The first load streams the official 2-bit checkpoint from Hugging Face and keeps a copy in this browser's storage, so later loads skip the download. Clearing site data removes cached weights and chat state.