A small API proxy for self-hosted llama.cpp setups. Routes across multiple llama-server instances, per-user API keys, token usage tracking. In-memory rate limiting (RPM/TPM) with async SQLite flush. ~0.13 ms is proxy logic (per-request), the 2–4 ms end-to-end figure includes the FastAPI/uvicorn/httpx baseline and asyncio event-loop contention at high concurrency. <=1100 code lines, ~53 MB RAM.
Built for the case where you run multiple llama-server instances (different models, different GPUs) and want to share them across users with token tracking.
smol-llm-proxy is built for one specific case: multiple llama-server instances, multiple users, per-user token accounting.
- LiteLLM — much broader scope: 100+ cloud providers, virtual keys, budgets, admin UI, fallbacks. Requires Postgres + Redis for full features. Use it if you need a production gateway across cloud LLMs.
- llama-swap — solves a different problem: hot-swapping models on one llama.cpp instance. No users, no accounting. Use it if you run many models on one machine and want them loaded on demand.
- llama.cpp router mode — built into llama-server itself. Same scope as llama-swap, no auth layer.
If you self-host several llama-server instances on one or more machines and want to share them with a small group while tracking usage, smol-llm-proxy is the smallest thing that does that. Otherwise, one of the above is probably a better fit.
[users] ──HTTP──> [proxy :port] ──HTTP──> [llama-server 1 :port]
│ [llama-server 2 :port]
│ [llama-server N :port]
│
├── in-memory cache (keys, routes) — TTL 30s
├── validate API key + resolve routing (SQLite on first call, then cache)
├── rate limiter: read-only DB check + in-memory reservation + reconcile on response
├── async batch flush to SQLite (1s interval)
├── forward request via connection-pooled httpx client
├── async usage logger (batch queue, 50 items / 1s timeout)
└── retention cleanup (90-day purge, daily)
TLS termination is handled externally — Cloudflare, Caddy, nginx, or any reverse proxy in front of the proxy.
- Per-user API keys (create / delete / toggle active)
- Rate limiting — per-key RPM/TPM, sliding 60-second window, 429 +
Retry-Afteron exceed - Multi-server routing by model name with in-memory cache (TTL 30s)
- Model aliases (
alias->model-name.gguf) - Token usage logging — prompt/completion tokens, timings, 90-day retention with daily purge
- Streaming and non-streaming proxy support for chat, completions, and embeddings
- Connection-pooled httpx client (keepalive connections to backends)
- SQLite backend (zero external DB dependencies)
Clone the repo, then:
cp .env.example .env # set ADMIN_KEY
cp config.example.yaml config.yaml # fill in your servers
docker compose up -d --buildThe proxy listens on 0.0.0.0:8000 by default.
docker build -t smol-llm-proxy .
docker run -p 8000:8000 \
-e ADMIN_KEY=secret \
-v db-data:/app/data \
-v $(pwd)/config.yaml:/app/config.yaml:ro \
smol-llm-proxypip install smol-llm-proxyExample configs ship with the package. Copy them:
python -c "import smol_llm_proxy, shutil, os; d=os.path.dirname(smol_llm_proxy.__file__); shutil.copy2(f'{d}/config.example.yaml','config.yaml'); shutil.copy2(f'{d}/.env.example','.env')"
# Edit .env and set ADMIN_KEY, then uncomment and fill in config.yaml with your servers
ADMIN_KEY=secret smol-llm-proxy
# or: ADMIN_KEY=secret python -m smol_llm_proxyOr download from GitHub:
curl -sO https://raw.githubusercontent.com/robolamp/smol-llm-proxy/main/.env.example && cp .env.example .env
curl -sO https://raw.githubusercontent.com/robolamp/smol-llm-proxy/main/config.example.yaml && cp config.example.yaml config.yamlgit clone https://github.com/robolamp/smol-llm-proxy.git
cd smol-llm-proxy
pip install -e .Copy example configs and run:
cp .env.example .env # set ADMIN_KEY
cp config.example.yaml config.yaml # fill in your servers
ADMIN_KEY=secret python -m smol_llm_proxy- Create a user key:
curl -X POST http://localhost:8000/admin/keys \
-H "Authorization: Bearer $ADMIN_KEY" \
-d '{"name": "my-user"}'The response contains a JSON object with a key field — that's the user's Bearer token. Save it; you'll need it for proxy requests.
- Send a chat completion with the user key:
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-<full-key-from-step-1>" \
-d '{
"model": "Qwen3.5-2B",
"messages": [{"role": "user", "content": "Hello!"}],
"stream": true
}'The proxy reads two files: config.yaml for routing and .env for runtime settings.
Loaded into SQLite at startup, persisted across restarts:
servers:
- name: my-server
url: http://host:port
api_key: "" # optional, if llama-server requires auth
models:
- model-name.gguf
aliases:
alias: model-name.gguf # short name -> real model nameOn startup, config.yaml is merged into the database:
- Servers/aliases present in
config.yamlare upserted into the database (new ones created, existing ones updated). - Servers/aliases present in the database but not in
config.yamlare kept — YAML is additive, not destructive. - Models listed under a server in YAML are synced; extra models on that server are removed.
- API keys, usage logs, and rate limits are not affected by config sync.
- The in-memory route cache is cleared after sync.
Servers and aliases managed via the Admin API persist across restarts. To remove them, use the Admin API (DELETE /admin/servers/{id}, DELETE /admin/aliases/{name}).
The Compose setup mounts two volumes:
${DATA_DIR:-./data}:/app/data— SQLite database (DB_PATH=/app/data/proxy.db), persists across container restarts. Override withDATA_DIRin.env../config.yaml:/config/config.yaml:ro— config file, read-only. SetCONFIG_PATH=/config/config.yamlin Compose (hardcoded, not from.env).
Reference: Admin API
All admin endpoints require Authorization: Bearer <ADMIN_KEY> header.
| Endpoint | Method | Description |
|---|---|---|
/admin/keys |
GET / POST |
List API keys / Create user key |
/admin/keys/{key_id} |
DELETE |
Revoke key (integer id, not key string) |
/admin/keys/{key_id}/toggle |
PATCH |
Activate/deactivate key |
/admin/keys/{key_id}/limits |
PUT |
Set per-key RPM/TPM rate limits |
/admin/servers |
GET / POST |
List servers / Register a llama-server |
/admin/servers/{server_id} |
DELETE / PATCH |
Remove or update server |
/admin/servers/{server_id}/models |
POST |
Assign model to server (use ?reassign=true to move from another server) |
/admin/servers/{server_id}/models/{model_name} |
DELETE |
Unassign model from server |
/admin/aliases |
GET / POST |
List aliases / Create alias |
/admin/aliases/{alias_name} |
PATCH / DELETE |
Update alias target / Delete alias |
/admin/usage |
GET |
Token usage logs + summary (accepts key_id, server_id, start_date, end_date, limit, offset) |
/admin/usage/summary/real |
GET |
Usage summary grouped by real model name + server |
Note: Key operations (DELETE, PATCH /toggle, PUT /limits) use integer key_id from the database, not the API key string itself.
Default limits: 100 RPM, 50,000 TPM. Set custom limits per key:
curl -X PUT http://localhost:8000/admin/keys/1/limits \
-H "Authorization: Bearer $ADMIN_KEY" \
-d '{"rpm_limit": 200, "tpm_limit": 100000}'Exceeding limits returns 429 Too Many Requests with a Retry-After header. The rate limiter reads from SQLite + in-memory counters on every check and commits in-memory with async batch flush — one read-only SELECT per request, no commit overhead. Note: rate counters and caches are process-local; under --workers N each worker maintains independent counters (limits multiply by N).
Each model can be assigned to exactly one server at a time. Assigning a model that is already on another server returns 409 Conflict:
curl -X POST http://localhost:8000/admin/servers/2/models \
-H "Authorization: Bearer $ADMIN_KEY" \
-d '{"model_name": "qwen3-30b-a3b-instruct.gguf"}'
# 409: Model 'qwen3-30b-a3b-instruct.gguf' already assigned to server 'server-a' (id=1). Use ?reassign=true to move.Move a model to a different server with ?reassign=true:
curl -X POST "http://localhost:8000/admin/servers/2/models?reassign=true" \
-H "Authorization: Bearer $ADMIN_KEY" \
-d '{"model_name": "qwen3-30b-a3b-instruct.gguf"}'Multiple aliases can point to the same model. Aliases can be switched to a different model via PATCH:
curl -X PATCH http://localhost:8000/admin/aliases/fast \
-H "Authorization: Bearer $ADMIN_KEY" \
-H "Content-Type: application/json" \
-d '{"real_model_name": "new-model.gguf"}'View raw usage logs:
curl "http://localhost:8000/admin/usage?key_id=1&start_date=2024-01-01T00:00:00&end_date=2024-01-31T23:59:59" \
-H "Authorization: Bearer $ADMIN_KEY"Usage summary grouped by real model name + server (aggregates across all aliases):
curl "http://localhost:8000/admin/usage/summary/real?key_id=1&start_date=2024-01-01T00:00:00&end_date=2024-01-31T23:59:59" \
-H "Authorization: Bearer $ADMIN_KEY"All endpoints accept: key_id, server_id, start_date (ISO 8601), end_date (ISO 8601), limit, offset. Usage logs are preserved even after a server or key is deleted.
Reference: Proxy Endpoints
These forward to llama-server backends based on model name routing.
| Endpoint | Method | Description |
|---|---|---|
/v1/chat/completions |
POST |
Chat completions (streaming + non-streaming) |
/v1/completions |
POST |
Legacy completions |
/v1/embeddings |
POST |
Embeddings |
/v1/models |
GET |
List available models (requires user API key — returns 401 without Authorization, 403 for invalid/inactive key, fans out to all backends) |
/health |
GET |
Health check (no auth required) |
Each request logs: user, server, model name, prompt/completion tokens, timings (ms), and total tokens. No conversation content is stored.
Reference: Environment Variables
| Variable | Default | Description |
|---|---|---|
ADMIN_KEY |
required | Bearer token for /admin/* endpoints. Startup fails without it. |
PROXY_HOST |
0.0.0.0 |
Listen address |
PROXY_PORT |
8000 |
Listen port |
DB_PATH |
data/proxy.db |
SQLite database location |
CONFIG_PATH |
config.yaml |
Path to YAML config file |
HTTPX_TIMEOUT |
120 |
Upstream HTTP client timeout (seconds) |
PROXY_MAX_CONNECTIONS |
50 |
httpx connection pool max size |
PROXY_MAX_KEEPALIVE |
20 |
httpx keepalive pool max size |
PROXY_KEEPALIVE_EXPIRY |
30.0 |
httpx keepalive expiry (seconds) |
BENCH_COLD_CACHE |
(unset) | Set to "1" to disable in-memory caches (for benchmarking) |
For pip install, set them in shell or via a .env-like loader (the proxy reads .env from the current working directory at startup). For Docker Compose, put them in .env:
ADMIN_KEY=secret
PROXY_PORT=8000Reference: Dependencies
| Package | Version | Notes |
|---|---|---|
fastapi |
>= 0.115.0 | Web framework |
uvicorn[standard] |
>= 0.30.0 | ASGI server |
httpx |
>= 0.27.0 | Async HTTP client |
pydantic |
>= 2.0.0 | Data validation |
pyyaml |
>= 6.0 | Config parsing |
orjson |
>= 3.9.0 | Fast JSON serialization |
uvloop |
>= 0.19.0 | Linux/macOS only (sys_platform != 'win32') |
Optional dev dependencies: pytest, pytest-rerunfailures, ruff (see pyproject.toml).
Benchmarking
Proxy overhead measured with Locust using parallel concurrent execution — both benchmarks (direct and proxy) hit the same backend simultaneously for a fair comparison.
All requests authenticated, routed, and logged via SQLite on every call (cold cache mode). In-memory cache disabled to measure worst-case per-request overhead.
Hardware: i9-14900K, RTX 4090 | Model: Qwen3.5-2B-GGUF | Backend: llama.cpp server
| Benchmark | Mean Overhead | P99 Overhead | RPS Delta |
|---|---|---|---|
| Mock (low load) | +2 ms | +10 ms | -1.0 |
| Mock (medium load) | +2 ms | +10 ms | -3.8 |
| Mock (high load) | +4 ms | +10 ms | -29.0 |
| Real (low load) | +1 ms | +20 ms | 0.0 |
Mock backend (fixed 100ms delay, clean conditions)
| Mode | Users | Direct P50 | Proxy P50 | Overhead P50 | Direct P95 | Proxy P95 | Overhead P95 | Direct P99 | Proxy P99 | Overhead P99 | Direct Mean | Through proxy | Overhead Mean | Direct RPS | Through proxy | RPS overhead |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Low | 5+5 | 100ms | 100ms | +0ms | 100ms | 100ms | +0ms | 100ms | 110ms | +10ms | 101ms | 104ms | +2ms | 49.1 | 48.1 | -1.0 |
| Medium | 20+20 | 100ms | 100ms | +0ms | 100ms | 110ms | +10ms | 100ms | 110ms | +10ms | 102ms | 104ms | +2ms | 194.5 | 190.7 | -3.8 |
| High | 100+100 | 100ms | 110ms | +10ms | 110ms | 110ms | +0ms | 110ms | 120ms | +10ms | 103ms | 106ms | +4ms | 893.7 | 864.7 | -29.0 |
Real llama-server backend
| Mode | Users | Direct P50 | Proxy P50 | Overhead P50 | Direct P95 | Proxy P95 | Overhead P95 | Direct P99 | Proxy P99 | Overhead P99 | Direct Mean | Through proxy | Overhead Mean | Direct RPS | Through proxy | RPS overhead |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Low | 5+5 | 430ms | 430ms | +0ms | 760ms | 750ms | -10ms | 850ms | 870ms | +20ms | 460ms | 461ms | +1ms | 10.8 | 10.8 | 0.0 |
Proxy overhead on clean conditions (mock): ~2–4 ms mean, +0–10 ms P99 across all load levels with DB-backed rate limiter. Against real backend: negligible at low load (+1 ms mean), higher load limited by SQLite I/O under concurrent bench runs.
Run your own benchmarks: python tests/benchmark/run.py [low|medium|high] (add --mock for fixed-delay backend)
| Workers | Idle | Under load | Growth |
|---|---|---|---|
| 1 | 53 MB | 62 MB | +9 MB |
| 4 | 252 MB | 273 MB | +21 MB |
Per-worker baseline: ~53 MB, load growth: +4–6 MB per worker.
Identical footprint against mock and real backends — the proxy forwards without buffering responses. No memory growth observed across extended runs.
The aggregate overhead numbers (+2–4 ms mean, +0–10 ms P99 on mock) include asyncio event loop contention at high concurrency (100+ concurrent users per worker). Per-request proxy logic itself is ~0.13 ms with uvloop — the difference comes from how asyncio handles many simultaneous awaits on a single thread. With 4 uvicorn workers, each worker handles ~25 requests, keeping contention minimal.
Real backend P99 spikes (Medium mode, ~14 s) are caused by llama.cpp single-thread inference bottleneck under 20 concurrent users — not proxy overhead. Proxy adds negligible latency regardless of backend saturation.