VRAM-aware autofit
Reads the model’s config.json and your VRAM, chooses the largest
context that fits, prints the arithmetic, and warns when the margin is thin.
How it works
CLI · local model server · EXL3
An Ollama-style CLI and local OpenAI-compatible server for ExLlamaV3, built for 4 GB NVIDIA cards. One static Go binary, zero dependencies.
curl -fsSL https://raw.githubusercontent.com/rizperdana/stone-llama/main/scripts/install.sh | sh
14 commands: doctor, setup, fit,
pull, rank, list, serve, ps,
stop, run, version, rm,
import, login.
MIT · Go ≥ 1.24 · static binary ≈ 7.4 MiB ·
NVIDIA CUDA · Linux/amd64 only ·
v0.1.0-rc1 ·
install docs.
01 — what’s different
Both run before you spend bandwidth or VRAM.
Reads the model’s config.json and your VRAM, chooses the largest
context that fits, prints the arithmetic, and warns when the margin is thin.
How it works
fit and pull check architecture → quant format
→ fit from metadata alone, then refuse with the numbers.
≤ ~16 MiB of metadata read per repo, worst case — it once refused a 1.68 GB model before a single weight byte moved.
02 — terminal
2026-09-26 (fit/gate re-captured 2026-09-29), RTX 3050 Laptop 4096 MiB, Linux/amd64. Raw files in docs/screenshots/.
$ stone-llama doctor
GPU NVIDIA GeForce RTX 3050 Laptop GPU
VRAM 4096 MiB
Driver 580.178.04
Runtime cu13 extra
$ stone-llama list
NAME QUANT SIZE VERDICT SOURCE
SmolLM3-3B-exl3_4.0bpw - 1.84 GiB - imported
$ stone-llama fit async0x42/Qwen3-1.7B-exl3_4.0bpw
gate: arch ✓ Qwen3ForCausalLM
gate: quant ✓ exl3
gate: fit ⚠ weights 1491 + KV 980 + overhead 128 = 2599 MiB (headroom 1344, budget 2752)
[…]
$ stone-llama fit async0x42/Qwen3-8B-exl3_4.0bpw
gate: arch ✓ Qwen3ForCausalLM
gate: quant ✓ exl3
gate: fit ✗ no config fits: weights 4950 + KV (Q4 @ 4096) 162 + overhead 128 = 5240 MiB > 3040 MiB budget (4096 MiB VRAM − 1056 headroom: 512 base + 512 prefill workspace [est] + 32 ctx margin [est])
largest ctx that would fit: none — no context fits; pull a smaller quant or use a bigger GPU
fit: refused — no context fits this model in VRAM (see the gate report above)
EXIT=3
$ stone-llama pull ggml-org/SmolLM3-3B-GGUF
stone-llama pull: repo ggml-org/SmolLM3-3B-GGUF has no config.json — not an EXL3 model layout
$ printf 'n\n' | stone-llama pull async0x42/Qwen3-1.7B-exl3_4.0bpw # declined: non-interactive, --yes required, 'n' never read
gate: arch ✓ Qwen3ForCausalLM
gate: quant ✓ exl3
gate: fit ⚠ weights 1491 + KV 980 + overhead 128 = 2599 MiB (headroom 1344, budget 2752)
[…]
pulling async0x42/Qwen3-1.7B-exl3_4.0bpw → Qwen3-1.7B-exl3_4.0bpw: 11 files, 1.47 GiB (sha256-verified), 76.05 GiB free
stone-llama pull: non-interactive pull requires --yes to confirm the download
$ stone-llama serve --attach 127.0.0.1:5002 --key-file <key-file>
stone-llama listening on 127.0.0.1:5111 (OpenAI-compatible)
upstream: http://127.0.0.1:5002 (attach)
[…]
$ stone-llama ps
stone-llama: running (pid 419204)
address: 127.0.0.1:5111
mode: attach (http://127.0.0.1:5002)
stone-llama does not own this process — it proxies only; load/unload is the upstream's
ready: yes
model: SmolLM3-3B-exl3
uptime: 3m45s
$ stone-llama stop
stone-llama: stopped
exit=0
03 — platforms
exllamav3 — the engine stone-llama wraps — has no CPU, AMD or Apple path. Not configurable.
| Machine | Works |
|---|---|
| NVIDIA GPU, Linux x86_64, driver ≥ 570 | ✓ cu12 |
| NVIDIA GPU, Linux x86_64, driver ≥ 580 | ✓ cu13 |
| NVIDIA GPU, driver < 570 | ✗ upgrade |
| Windows | ⚠ untested |
| macOS | ✗ never serves |
| AMD / Intel GPU | ✗ |
| CPU only | ✗ |
| Linux arm64 | ✗ not v1 |
Linux/amd64 is the only supported platform in v1 — on CPU, AMD or
Apple use Ollama with GGUF.
stone-llama doctor prints this verdict before anything is installed.
On macOS doctor/list/fit still run.
04 — comparison
Same-category local serving runtimes, on the axes that matter for a 4 GB machine. Compiled 2026-09-26 from each project’s own documentation — nothing was installed or run. This table includes the rows where stone-llama loses.
| Capability | stone-llama | ollama | llama.cpp | LM Studio | TabbyAPI |
|---|---|---|---|---|---|
| Checks fit before download | ✓ | ✗ | ✗ | ✓ | ✗ |
| VRAM-aware context | ✓ | ~ | ✓ | ✓ | ✗ |
| Runs without NVIDIA GPU | ✗ | ✓ | ✓ | ✓ | ✗ |
| Windows / macOS | ✗ | ✓ | ✓ | ? | ? |
| Several models at once | ✗ | ✓ | ✗ | ✓ | ✗ |
| Idle auto-unload | ✗ | ✓ | ~ | ✓ | ✗ |
ps / stop daemon |
✓ | ✓ | ✗ | ✓ | ✗ |
| Tool calling | ✓ | ✓ | ✓ | ? | ✗ |
| GGUF models | ✗ | ✓ | ✓ | ✓ | ✗ |
| EXL3 models | ✓ | ✗ | ✗ | ✗ | ✓ |
| Environment preflight | ✓ | ✗ | ✗ | ~ | ✗ |
✓ yes · ✗ no · ~ partial · ? not covered by our sources.
Sources: docs.ollama.com · llama.cpp server README · lmstudio.ai/docs · TabbyAPI wiki · this repo’s README. LM Studio’s pre-download states are third-party sourced.
05 — models
Metadata-only ranking, 2026-09-29 — zero weight bytes downloaded. RTX 3050 Laptop, 4096 MiB.
| Model | Size | Max ctx | Tool calls | Verdict |
|---|---|---|---|---|
| Qwen3-4B-Instruct-2507 ArtusDev/Qwen_Qwen3-4B-Instruct-2507-EXL3 @ 2.5bpw_H6 | 2.06 GiB | 35,415 [est] | ✓ | fits-tight |
| LFM2.5-8B-A1B turboderp/LFM2.5-8B-A1B-exl3 @ 2.10bpw_mul1 | 2.70 GiB | 128,000 | ✓ | fits |
| Ministral-3B UnstableLlama/Ministral-3-3B-Instruct-2512-exl3 @ 2.10bpw | 2.55 GiB | 31,301 [est] | ✓ | fits-tight |
| SmolLM3-3B turboderp/SmolLM3-3B-exl3 @ 3.5bpw — baseline | 1.82 GiB | 65,536 | ✓ | fits |
| Qwen3.5-2B UnstableLlama/Qwen3.5-2B-exl3-2.10bpw | 2.26 GiB | 262,144 | ✓ | fits |
| Qwen2.5-7B-Instruct turboderp/Qwen2.5-7B-Instruct-exl3 @ 2.0bpw | 2.92 GiB | 32,768 | ✓ | fits |
[est] = computed from metadata + your VRAM, not
measured. Tool calls = the runtime’s server-side parser.
Max ctx here uses the budget rule —
weights + KV ≤ 3508 MiB (4096 − 88 overhead − 500 headroom allowance, the
research note’s config rule) — which prices cache only.
stone-llama fit applies the stricter headroom rule
(weights + KV + overhead ≤ 4096 − headroom(ctx); headroom grows with ctx), so
its picks on this card are smaller: Qwen3-4B Q4 16384, LFM2.5 Q4 15872, Ministral
Q4 8192, SmolLM3 Q4 32768, Qwen3.5-2B Q8 32768; Qwen2.5-7B @2.0bpw is
refused there (3182 MiB > 3040 MiB budget at ctx 4096). Both figures
are true of their own rule — the table above is the budget rule, fit is the
headroom rule. fit speed estimates anchor to one measured decode point:
42.7 tok/s (SmolLM3-3B, this card); prefill estimates are unfitted.