Skip to main content

Configuration Reference

miLLM is configured entirely through environment variables (with .env file support, case-sensitive names). In Docker Compose these go in the api service's environment: block; in Kubernetes, in the deployment's env:.

Core​

VariableDefaultDescription
DATABASE_URLpostgresql+asyncpg://postgres:postgres@localhost:5432/millmPostgreSQL DSN (async driver required)
MODEL_CACHE_DIR/app/model_cacheWhere downloaded model weights live
SAE_CACHE_DIR/app/sae_cacheWhere downloaded SAEs live
HF_TOKEN—Server-wide HuggingFace token for gated repos; per-request tokens in API calls take precedence and are never persisted
HOST / PORT0.0.0.0 / 8000Bind address and port
DEBUGfalseEnables debug behavior; never use in production
REDIS_URL—Optional Redis for cache/real-time state
AUTO_LOAD_MODEL—Model ID or name to load automatically at startup — recommended for unattended deployments

HTTP & logging​

VariableDefaultDescription
CORS_ORIGINS*Comma-separated allowed origins. Set explicitly (e.g. http://localhost:3000,https://webui.example.com) when browsers call the API cross-origin
LOG_LEVELINFODEBUG / INFO / WARNING / ERROR
LOG_FORMATconsoleconsole for humans, json for log pipelines

Concurrency & timeouts​

VariableDefaultDescription
MAX_CONCURRENT_REQUESTS1Must stay 1. This is a correctness constraint, not a throughput knob: values above 1 race on steering apply/restore and on the shared sensing buffer. Leave it at 1
MAX_PENDING_REQUESTS10Queue depth beyond the concurrent slots; overflow returns 503 (QUEUE_FULL, backpressure)
MAX_DOWNLOAD_WORKERS2Parallel model/SAE downloads
MAX_LOAD_WORKERS1Parallel model loads (keep at 1)
GRACEFUL_UNLOAD_TIMEOUT30.0Seconds an unload waits for the requests already running on the model before it moves any weight. Requests that arrive once the unload has begun are refused with 503 model_busy rather than waited for. If the running requests outlast this, the unload goes ahead and they fail. (Until 2026-09-14 this setting was read by nothing and the wait was a fixed 5 s.)
DOWNLOAD_TIMEOUT3600.0Max seconds for a single download

Model lease and backpressure​

VariableDefaultDescription
LEASE_DEFAULT_TTL_SECONDS7200TTL of a lease taken without ttl_seconds. Must not exceed the maximum (checked at startup)
LEASE_MAX_TTL_SECONDS7200Largest ttl_seconds accepted; larger values are refused with 400, never clamped
LEASE_HOLDER_MAX_CHARS128Longest holder label, after stripping
LEASE_REASON_MAX_CHARS512Longest reason, after stripping
LEASE_ENDED_MEMORY64Ended leases remembered, so renew/release of a seen ID answers 409 LEASE_EXPIRED with its reason instead of 404. Lost at restart
QUEUE_DURATION_WINDOW50Recent slot-holding durations kept for estimated_wait_seconds
RETRY_AFTER_QUEUE_DEFAULT_S5QUEUE_FULL with no estimate
RETRY_AFTER_MAX_S60Ceiling on the QUEUE_FULL estimate
RETRY_AFTER_LOAD_S15MODEL_BUSY / MODEL_LOADING while a load runs
RETRY_AFTER_UNLOAD_S5MODEL_BUSY while an unload runs
RETRY_AFTER_NOT_LOADED_S30MODEL_NOT_LOADED
RETRY_AFTER_MEMORY_S30INSUFFICIENT_MEMORY on /v1
RETRY_AFTER_READINESS_S5The readiness probe's 503
RETRY_AFTER_FALLBACK_S10Any other 503 that reached the response without Retry-After; logged as retry_after_defaulted. Deliberately distinct from every other value

Leases live in process memory: a restart ends every lease. miLLM runs one worker process; several workers would each hold their own lease registry.

Performance​

VariableDefaultDescription
TORCH_COMPILE(auto)true/false/unset. Unset = auto: compile on CUDA for non-bitsandbytes models. Compilation speeds up decoding significantly; first request after model load (and after SAE attach/detach) pays a one-time recompile (~20 s). The SAE hook is fully honored under compilation — see Architecture
TORCH_COMPILE_MODEreduce-overheaddefault / reduce-overhead / max-autotune
KV_CACHE_MODEdynamicstatic enables compiled static KV cache (needs a C compiler for triton in the image)
SPECULATIVE_MODEL—HF model ID of a small draft model for speculative decoding. Works with steering: draft proposes, the steered main model verifies — output correctness preserved, acceptance rate lower
SPECULATIVE_NUM_TOKENS5Tokens the draft proposes per step

Continuous batching (opt-in)​

High-throughput batched inference via HuggingFace ContinuousBatchingManager. Trade-offs in Architecture.

VariableDefaultDescription
ENABLE_CONTINUOUS_BATCHINGfalseStart CBM at model load
CBM_MAX_QUEUE_SIZE256Manager queue size
CBM_DEFAULT_TEMPERATURE0.7Fixed sampling temperature for batched requests; requests with a different value fall back to serial
CBM_DEFAULT_TOP_P0.95Fixed top-p, same fallback rule
CBM_DEFAULT_MAX_TOKENS512Default generation length
CBM_FORCE_SERIAL_MONITORINGfalseRoute monitored requests to the serial path for exact per-request activation attribution (trades throughput for fidelity)

Transformers models​

Settings that decide whether a transformers (non-GGUF) model fits the GPUs. See Hardware Requirements for how they are used.

VariableDefaultNotes
TRANSFORMERS_MIN_CONTEXT4096The context, in tokens, whose KV cache every card a model uses must have room for, or the model's own maximum context if that is shorter. A floor for loading, not a limit on requests. SAE attachment keeps the same room free. Raise it to refuse a model that would only fit with a short context.
TRANSFORMERS_CUDA_CONTEXT_MB500Memory each card keeps for its CUDA context, in MB. Measured on the node at 250–400 MiB a card. It covers the context only: a request's working memory (its prefill activations and what PyTorch's allocator keeps reserved) is sized per model and per card, and a bitsandbytes load's staging is counted too (see Hardware Requirements).
TRANSFORMERS_IDLE_CACHE_RELEASE_S5.0Seconds a transformers model's cards stay idle before PyTorch's unused cached memory is returned to them. Without this, memory a finished request freed stays reserved, and nvidia-smi, which miStudio and other tenants of the node read, shows it as used until the model is unloaded. The release waits for the request queue to be idle and takes its slot, so it never runs during a request, and a request that arrives first cancels it. The next request then reserves that memory again, which costs tens of driver calls. 0 releases as soon as the queue is idle; a negative value never releases. Nothing is released while continuous batching is running, because its requests do not pass through the request queue. GGUF models are not affected.

GGUF models​

Settings that apply only when the loaded model is a GGUF file. See Hardware Requirements for the measurements behind the defaults.

VariableDefaultNotes
GGUF_CONTEXT_LENGTH32768A ceiling, not a target. miLLM reads what the file declares and loads the largest window that fits under this. 0 means "whatever the file declares", bounded only by what fits.
GGUF_TENSOR_SPLIT(empty)How a GGUF model that needs more than one GPU divides its layers. Empty divides them in proportion to each used card's free memory. Otherwise give one proportion per card the split uses, in index order (1,3 puts a quarter on the lower-index card). A value with the wrong number of entries for a load is refused, not stretched. It has no effect on a model that fits one card.
GGUF_KV_CACHE_TYPEq8_0KV-cache precision — the biggest lever on how much context fits. q8_0 is near-lossless and roughly triples the window over f16. q4_0 buys another third at a real accuracy cost. f16 restores llama.cpp's default behaviour.
GGUF_FLASH_ATTENTIONtrueRequired by a quantized KV cache, not merely an optimisation. The loader refuses to pair the two with this off.
GGUF_ENABLE_EMBEDDINGStrueLoads with embedding output so /v1/embeddings works without a second load. Costs ~6.7% generation throughput, measured. Must be decided at load time — it cannot be switched on later.
Why the ceiling is needed in both directions

llama.cpp's own default is 512 tokens, which truncates almost any real conversation. Unbounded is the opposite trap: a model declaring 262,144 tokens needs tens of gigabytes of KV cache, which either fails the load outright or reserves a GPU that other work shares.

Example configurations​

.env — single-GPU research box
AUTO_LOAD_MODEL=gemma-2-2b
MAX_CONCURRENT_REQUESTS=1
CORS_ORIGINS=http://localhost:3000
LOG_FORMAT=console
.env — shared inference server
ENABLE_CONTINUOUS_BATCHING=true
CBM_FORCE_SERIAL_MONITORING=true
MAX_PENDING_REQUESTS=32
LOG_FORMAT=json
LOG_LEVEL=INFO