stone-llama

CLI · local model server · EXL3

stone-llama

An Ollama-style CLI and local OpenAI-compatible server for ExLlamaV3, built for 4 GB NVIDIA cards. One static Go binary, zero dependencies.

  • Fits context to your VRAM — picks the largest safe context and prints the arithmetic.
  • Gates every download — architecture, quant format and fit checked from metadata, before any weight byte moves.
curl -fsSL https://raw.githubusercontent.com/rizperdana/stone-llama/main/scripts/install.sh | sh

14 commands: doctor, setup, fit, pull, rank, list, serve, ps, stop, run, version, rm, import, login.

MIT · Go ≥ 1.24 · static binary ≈ 7.4 MiB · NVIDIA CUDA · Linux/amd64 only · v0.1.0-rc1 · install docs.

01 — what’s different

Two checks before you commit

Both run before you spend bandwidth or VRAM.

VRAM-aware autofit

Reads the model’s config.json and your VRAM, chooses the largest context that fits, prints the arithmetic, and warns when the margin is thin. How it works

Pre-download fit gate

fit and pull check architecture → quant format → fit from metadata alone, then refuse with the numbers.

≤ ~16 MiB of metadata read per repo, worst case — it once refused a 1.68 GB model before a single weight byte moved.

02 — terminal

Four captures from a live run

2026-09-26 (fit/gate re-captured 2026-09-29), RTX 3050 Laptop 4096 MiB, Linux/amd64. Raw files in docs/screenshots/.

doctor-list.txt — doctor + list
$ stone-llama doctor
GPU        NVIDIA GeForce RTX 3050 Laptop GPU
VRAM       4096 MiB
Driver     580.178.04
Runtime    cu13 extra

$ stone-llama list
NAME                    QUANT  SIZE      VERDICT  SOURCE
SmolLM3-3B-exl3_4.0bpw  -      1.84 GiB  -        imported
fit.txt — fit: gate, arithmetic, refusal
$ stone-llama fit async0x42/Qwen3-1.7B-exl3_4.0bpw
gate: arch  ✓ Qwen3ForCausalLM
gate: quant ✓ exl3
gate: fit   ⚠ weights 1491 + KV 980 + overhead 128 = 2599 MiB (headroom 1344, budget 2752)
[…]
$ stone-llama fit async0x42/Qwen3-8B-exl3_4.0bpw
gate: arch  ✓ Qwen3ForCausalLM
gate: quant ✓ exl3
gate: fit   ✗ no config fits: weights 4950 + KV (Q4 @ 4096) 162 + overhead 128 = 5240 MiB > 3040 MiB budget (4096 MiB VRAM − 1056 headroom: 512 base + 512 prefill workspace [est] + 32 ctx margin [est])
      largest ctx that would fit: none — no context fits; pull a smaller quant or use a bigger GPU
fit: refused — no context fits this model in VRAM (see the gate report above)
EXIT=3
gate.txt — pull: GGUF refused before download
$ stone-llama pull ggml-org/SmolLM3-3B-GGUF
stone-llama pull: repo ggml-org/SmolLM3-3B-GGUF has no config.json — not an EXL3 model layout

$ printf 'n\n' | stone-llama pull async0x42/Qwen3-1.7B-exl3_4.0bpw   # declined: non-interactive, --yes required, 'n' never read
gate: arch  ✓ Qwen3ForCausalLM
gate: quant ✓ exl3
gate: fit   ⚠ weights 1491 + KV 980 + overhead 128 = 2599 MiB (headroom 1344, budget 2752)
[…]
pulling async0x42/Qwen3-1.7B-exl3_4.0bpw → Qwen3-1.7B-exl3_4.0bpw: 11 files, 1.47 GiB (sha256-verified), 76.05 GiB free
stone-llama pull: non-interactive pull requires --yes to confirm the download
serve.txt — serve + ps + stop (attach mode)
$ stone-llama serve --attach 127.0.0.1:5002 --key-file <key-file>
stone-llama listening on 127.0.0.1:5111 (OpenAI-compatible)
  upstream: http://127.0.0.1:5002 (attach)
[…]
$ stone-llama ps
stone-llama: running (pid 419204)
  address:  127.0.0.1:5111
  mode:     attach (http://127.0.0.1:5002)
            stone-llama does not own this process — it proxies only; load/unload is the upstream's
  ready:    yes
  model:    SmolLM3-3B-exl3
  uptime:   3m45s

$ stone-llama stop
stone-llama: stopped
exit=0

03 — platforms

NVIDIA-CUDA only. Read this first.

exllamav3 — the engine stone-llama wraps — has no CPU, AMD or Apple path. Not configurable.

MachineWorks
NVIDIA GPU, Linux x86_64, driver ≥ 570✓ cu12
NVIDIA GPU, Linux x86_64, driver ≥ 580✓ cu13
NVIDIA GPU, driver < 570✗ upgrade
Windows⚠ untested
macOS✗ never serves
AMD / Intel GPU✗
CPU only✗
Linux arm64✗ not v1

Linux/amd64 is the only supported platform in v1 — on CPU, AMD or Apple use Ollama with GGUF. stone-llama doctor prints this verdict before anything is installed. On macOS doctor/list/fit still run.

04 — comparison

How it compares

Same-category local serving runtimes, on the axes that matter for a 4 GB machine. Compiled 2026-09-26 from each project’s own documentation — nothing was installed or run. This table includes the rows where stone-llama loses.

Capability stone-llama ollama llama.cpp LM Studio TabbyAPI
Checks fit before download ✓ ✗ ✗ ✓ ✗
VRAM-aware context ✓ ~ ✓ ✓ ✗
Runs without NVIDIA GPU ✗ ✓ ✓ ✓ ✗
Windows / macOS ✗ ✓ ✓ ? ?
Several models at once ✗ ✓ ✗ ✓ ✗
Idle auto-unload ✗ ✓ ~ ✓ ✗
ps / stop daemon ✓ ✓ ✗ ✓ ✗
Tool calling ✓ ✓ ✓ ? ✗
GGUF models ✗ ✓ ✓ ✓ ✗
EXL3 models ✓ ✗ ✗ ✗ ✓
Environment preflight ✓ ✗ ✗ ~ ✗

✓ yes · ✗ no · ~ partial · ? not covered by our sources.

Sources: docs.ollama.com · llama.cpp server README · lmstudio.ai/docs · TabbyAPI wiki · this repo’s README. LM Studio’s pre-download states are third-party sourced.

05 — models

What fits a 4 GB card

Metadata-only ranking, 2026-09-29 — zero weight bytes downloaded. RTX 3050 Laptop, 4096 MiB.

Model Size Max ctx Tool calls Verdict
Qwen3-4B-Instruct-2507 ArtusDev/Qwen_Qwen3-4B-Instruct-2507-EXL3 @ 2.5bpw_H6 2.06 GiB 35,415 [est] ✓ fits-tight
LFM2.5-8B-A1B turboderp/LFM2.5-8B-A1B-exl3 @ 2.10bpw_mul1 2.70 GiB 128,000 ✓ fits
Ministral-3B UnstableLlama/Ministral-3-3B-Instruct-2512-exl3 @ 2.10bpw 2.55 GiB 31,301 [est] ✓ fits-tight
SmolLM3-3B turboderp/SmolLM3-3B-exl3 @ 3.5bpw — baseline 1.82 GiB 65,536 ✓ fits
Qwen3.5-2B UnstableLlama/Qwen3.5-2B-exl3-2.10bpw 2.26 GiB 262,144 ✓ fits
Qwen2.5-7B-Instruct turboderp/Qwen2.5-7B-Instruct-exl3 @ 2.0bpw 2.92 GiB 32,768 ✓ fits

[est] = computed from metadata + your VRAM, not measured. Tool calls = the runtime’s server-side parser. Max ctx here uses the budget rule — weights + KV ≤ 3508 MiB (4096 − 88 overhead − 500 headroom allowance, the research note’s config rule) — which prices cache only. stone-llama fit applies the stricter headroom rule (weights + KV + overhead ≤ 4096 − headroom(ctx); headroom grows with ctx), so its picks on this card are smaller: Qwen3-4B Q4 16384, LFM2.5 Q4 15872, Ministral Q4 8192, SmolLM3 Q4 32768, Qwen3.5-2B Q8 32768; Qwen2.5-7B @2.0bpw is refused there (3182 MiB > 3040 MiB budget at ctx 4096). Both figures are true of their own rule — the table above is the budget rule, fit is the headroom rule. fit speed estimates anchor to one measured decode point: 42.7 tok/s (SmolLM3-3B, this card); prefill estimates are unfitted.