Configuration Reference
miLLM is configured entirely through environment variables (with .env file support, case-sensitive names). In Docker Compose these go in the api service's environment: block; in Kubernetes, in the deployment's env:.
Core
| Variable | Default | Description |
|---|---|---|
DATABASE_URL | postgresql+asyncpg://postgres:postgres@localhost:5432/millm | PostgreSQL DSN (async driver required) |
MODEL_CACHE_DIR | /app/model_cache | Where downloaded model weights live |
SAE_CACHE_DIR | /app/sae_cache | Where downloaded SAEs live |
HF_TOKEN | — | Server-wide HuggingFace token for gated repos; per-request tokens in API calls take precedence and are never persisted |
HOST / PORT | 0.0.0.0 / 8000 | Bind address and port |
DEBUG | false | Enables debug behavior; never use in production |
REDIS_URL | — | Optional Redis for cache/real-time state |
AUTO_LOAD_MODEL | — | Model ID or name to load automatically at startup — recommended for unattended deployments |
HTTP & logging
| Variable | Default | Description |
|---|---|---|
CORS_ORIGINS | * | Comma-separated allowed origins. Set explicitly (e.g. http://localhost:3000,https://webui.example.com) when browsers call the API cross-origin |
LOG_LEVEL | INFO | DEBUG / INFO / WARNING / ERROR |
LOG_FORMAT | console | console for humans, json for log pipelines |
Concurrency & timeouts
| Variable | Default | Description |
|---|---|---|
MAX_CONCURRENT_REQUESTS | 1 | Must stay 1. This is a correctness constraint, not a throughput knob: values above 1 race on steering apply/restore and on the shared sensing buffer. Leave it at 1 |
MAX_PENDING_REQUESTS | 10 | Queue depth beyond the concurrent slots; overflow returns 503 (QUEUE_FULL, backpressure) |
MAX_DOWNLOAD_WORKERS | 2 | Parallel model/SAE downloads |
MAX_LOAD_WORKERS | 1 | Parallel model loads (keep at 1) |
GRACEFUL_UNLOAD_TIMEOUT | 30.0 | Seconds an unload waits for the requests already running on the model before it moves any weight. Requests that arrive once the unload has begun are refused with 503 model_busy rather than waited for. If the running requests outlast this, the unload goes ahead and they fail. (Until 2026-09-14 this setting was read by nothing and the wait was a fixed 5 s.) |
DOWNLOAD_TIMEOUT | 3600.0 | Max seconds for a single download |
Model lease and backpressure
| Variable | Default | Description |
|---|---|---|
LEASE_DEFAULT_TTL_SECONDS | 7200 | TTL of a lease taken without ttl_seconds. Must not exceed the maximum (checked at startup) |
LEASE_MAX_TTL_SECONDS | 7200 | Largest ttl_seconds accepted; larger values are refused with 400, never clamped |
LEASE_HOLDER_MAX_CHARS | 128 | Longest holder label, after stripping |
LEASE_REASON_MAX_CHARS | 512 | Longest reason, after stripping |
LEASE_ENDED_MEMORY | 64 | Ended leases remembered, so renew/release of a seen ID answers 409 LEASE_EXPIRED with its reason instead of 404. Lost at restart |
QUEUE_DURATION_WINDOW | 50 | Recent slot-holding durations kept for estimated_wait_seconds |
RETRY_AFTER_QUEUE_DEFAULT_S | 5 | QUEUE_FULL with no estimate |
RETRY_AFTER_MAX_S | 60 | Ceiling on the QUEUE_FULL estimate |
RETRY_AFTER_LOAD_S | 15 | MODEL_BUSY / MODEL_LOADING while a load runs |
RETRY_AFTER_UNLOAD_S | 5 | MODEL_BUSY while an unload runs |
RETRY_AFTER_NOT_LOADED_S | 30 | MODEL_NOT_LOADED |
RETRY_AFTER_MEMORY_S | 30 | INSUFFICIENT_MEMORY on /v1 |
RETRY_AFTER_READINESS_S | 5 | The readiness probe's 503 |
RETRY_AFTER_FALLBACK_S | 10 | Any other 503 that reached the response without Retry-After; logged as retry_after_defaulted. Deliberately distinct from every other value |
Leases live in process memory: a restart ends every lease. miLLM runs one worker process; several workers would each hold their own lease registry.
Performance
| Variable | Default | Description |
|---|---|---|
TORCH_COMPILE | (auto) | true/false/unset. Unset = auto: compile on CUDA for non-bitsandbytes models. Compilation speeds up decoding significantly; first request after model load (and after SAE attach/detach) pays a one-time recompile (~20 s). The SAE hook is fully honored under compilation — see Architecture |
TORCH_COMPILE_MODE | reduce-overhead | default / reduce-overhead / max-autotune |
KV_CACHE_MODE | dynamic | static enables compiled static KV cache (needs a C compiler for triton in the image) |
SPECULATIVE_MODEL | — | HF model ID of a small draft model for speculative decoding. Works with steering: draft proposes, the steered main model verifies — output correctness preserved, acceptance rate lower |
SPECULATIVE_NUM_TOKENS | 5 | Tokens the draft proposes per step |
Continuous batching (opt-in)
High-throughput batched inference via HuggingFace ContinuousBatchingManager. Trade-offs in Architecture.
| Variable | Default | Description |
|---|---|---|
ENABLE_CONTINUOUS_BATCHING | false | Start CBM at model load |
CBM_MAX_QUEUE_SIZE | 256 | Manager queue size |
CBM_DEFAULT_TEMPERATURE | 0.7 | Fixed sampling temperature for batched requests; requests with a different value fall back to serial |
CBM_DEFAULT_TOP_P | 0.95 | Fixed top-p, same fallback rule |
CBM_DEFAULT_MAX_TOKENS | 512 | Default generation length |
CBM_FORCE_SERIAL_MONITORING | false | Route monitored requests to the serial path for exact per-request activation attribution (trades throughput for fidelity) |
Transformers models
Settings that decide whether a transformers (non-GGUF) model fits the GPUs. See Hardware Requirements for how they are used.
| Variable | Default | Notes |
|---|---|---|
TRANSFORMERS_MIN_CONTEXT | 4096 | The context, in tokens, whose KV cache every card a model uses must have room for, or the model's own maximum context if that is shorter. A floor for loading, not a limit on requests. SAE attachment keeps the same room free. Raise it to refuse a model that would only fit with a short context. |
TRANSFORMERS_CUDA_CONTEXT_MB | 500 | Memory each card keeps for its CUDA context, in MB. Measured on the node at 250–400 MiB a card. It covers the context only: a request's working memory (its prefill activations and what PyTorch's allocator keeps reserved) is sized per model and per card, and a bitsandbytes load's staging is counted too (see Hardware Requirements). |
TRANSFORMERS_IDLE_CACHE_RELEASE_S | 5.0 | Seconds a transformers model's cards stay idle before PyTorch's unused cached memory is returned to them. Without this, memory a finished request freed stays reserved, and nvidia-smi, which miStudio and other tenants of the node read, shows it as used until the model is unloaded. The release waits for the request queue to be idle and takes its slot, so it never runs during a request, and a request that arrives first cancels it. The next request then reserves that memory again, which costs tens of driver calls. 0 releases as soon as the queue is idle; a negative value never releases. Nothing is released while continuous batching is running, because its requests do not pass through the request queue. GGUF models are not affected. |
GGUF models
Settings that apply only when the loaded model is a GGUF file. See Hardware Requirements for the measurements behind the defaults.
| Variable | Default | Notes |
|---|---|---|
GGUF_CONTEXT_LENGTH | 32768 | A ceiling, not a target. miLLM reads what the file declares and loads the largest window that fits under this. 0 means "whatever the file declares", bounded only by what fits. |
GGUF_TENSOR_SPLIT | (empty) | How a GGUF model that needs more than one GPU divides its layers. Empty divides them in proportion to each used card's free memory. Otherwise give one proportion per card the split uses, in index order (1,3 puts a quarter on the lower-index card). A value with the wrong number of entries for a load is refused, not stretched. It has no effect on a model that fits one card. |
GGUF_KV_CACHE_TYPE | q8_0 | KV-cache precision — the biggest lever on how much context fits. q8_0 is near-lossless and roughly triples the window over f16. q4_0 buys another third at a real accuracy cost. f16 restores llama.cpp's default behaviour. |
GGUF_FLASH_ATTENTION | true | Required by a quantized KV cache, not merely an optimisation. The loader refuses to pair the two with this off. |
GGUF_ENABLE_EMBEDDINGS | true | Loads with embedding output so /v1/embeddings works without a second load. Costs ~6.7% generation throughput, measured. Must be decided at load time — it cannot be switched on later. |
llama.cpp's own default is 512 tokens, which truncates almost any real conversation. Unbounded is the opposite trap: a model declaring 262,144 tokens needs tens of gigabytes of KV cache, which either fails the load outright or reserves a GPU that other work shares.
Example configurations
AUTO_LOAD_MODEL=gemma-2-2b
MAX_CONCURRENT_REQUESTS=1
CORS_ORIGINS=http://localhost:3000
LOG_FORMAT=console
ENABLE_CONTINUOUS_BATCHING=true
CBM_FORCE_SERIAL_MONITORING=true
MAX_PENDING_REQUESTS=32
LOG_FORMAT=json
LOG_LEVEL=INFO