Skip to main content

Troubleshooting

Organized by symptom. For error codes returned by the API, see the Error Codes reference.

Steering has no effect​

The most important failure mode in miLLM — an intervention that silently isn't happening invalidates your experiment. Check in this order:

  1. Is steering actually enabled with non-zero values?
    curl -s http://localhost:8000/api/saes/steering | jq '.data'
  2. Is the hook firing? Generate a completion, then:
    curl -s http://localhost:8000/api/saes/attachment | jq '.data.steering_apply_count'
    The counter must increase across a steered generation. If it doesn't move, the hook isn't executing — re-attach the SAE and check server logs. (miLLM guards the known torch.compile-swallows-hooks failure automatically, and attach/detach forces a recompile; a non-moving counter after that indicates a genuinely unsupported architecture.)
  3. Right module? The attach response's layer_module_path should name a decoder layer (e.g. model.layers.12). On multimodal/exotic architectures the layer-resolution heuristics can land elsewhere — try attaching by a different layer index and compare.
  4. Strength too low or feature not what you think? Sweep to ±100 on a feature with an obvious signature (see the tutorial); check the feature's meaning on Neuronpedia.
  5. Wrong model variant? SAEs trained on gemma-2-2b are much weaker on gemma-2-2b-it. The attach warnings array tells you about this mismatch.

Out-of-memory (OOM)​

SymptomCauseFix
OOM during model loadModel too large for GPUMore aggressive quantization (Q8/Q4), or let device_map=auto offload to CPU
OOM during inferencemodel + KV cache + SAE exceeds VRAMReduce max_tokens, use a narrower SAE (16k vs 65k+), or reduce MAX_PENDING_REQUESTS to bound queued work (MAX_CONCURRENT_REQUESTS must stay 1 — see Configuration)
OOM only with hybrid/Mamba modelsmamba-ssm not installed; naive fallback allocates 20 GB+ intermediatesInstall mamba-ssm, or set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
VRAM used but no model loadedLeaked memory after a crashRestart the backend process/pod, then reload

On OOM with an SAE attached, miLLM degrades gracefully: the SAE is disabled and the base model continues serving — check the attachment status if steering suddenly stops.

nvidia-smi --query-gpu=memory.used,memory.free,memory.total --format=csv,noheader
nvidia-smi --query-compute-apps=pid,name,used_memory --format=csv,noheader

Sizing guidance: Hardware Requirements.

Downloads & model loading​

SymptomCauseFix
401 GATED_MODEL_NO_TOKENGemma/Llama are gatedAccept the license on the HF model page, supply hf_token
404 REPO_NOT_FOUNDTypo, or private repo without tokenCheck the repo ID; add a token
Load fails: "TokenizersBackend"Model needs a custom tokenizer packageRetry with trust_remote_code: true
Load fails: BitNet/GPTQ conflictsPre-quantized weights + bitsandbytes don't mixDownload with FP16 (miLLM auto-detects most cases and skips bitsandbytes)
First request after load is very slow (~20 s)One-time torch.compile warmupExpected; subsequent requests are fast. Also occurs once after SAE attach/detach. Disable with TORCH_COMPILE=false if warmup matters more than throughput
Download stuckNetwork or HF rate limitingCancel (POST /api/models/{id}/cancel) and retry; check GET /api/health/circuits for a tripped HuggingFace breaker

GGUF-specific​

SymptomCauseFix
Could not create a llama.cpp context at any context length (tried …)Usually not the context size. Some architectures refuse the pooling mode embeddings require and fail at every lengthmiLLM already retries without embeddings and records supports_embeddings: false. If it still fails at every rung with the card nearly empty, the file itself is the problem, not the window
400 AMBIGUOUS_MODEL_NAMEA bare repository name now matches two or more downloaded quantizationsName the tag exactly, e.g. …-GGUF:IQ4_XS. The error lists the tags that exist
400 context_length_exceeded on a prompt that used to workThe window shrank — usually GGUF_KV_CACHE_TYPE set back to f16, or GGUF_FLASH_ATTENTION turned offRestore q8_0 + flash attention, which roughly triples the window. See Hardware
Context window smaller than expectedGGUF_CONTEXT_LENGTH is a ceiling, and the loader also stops at what VRAM can holdCheck the load log for the predicted and achieved window; free VRAM, or lower the KV-cache precision
/v1/embeddings returns an error for one GGUF modelThat model loaded without embedding supportCheck supports_embeddings on the model record; use a different model for embeddings
A Continue action restarts the answer from the beginningFixed — GGUF chat templates seal the final turn, and miLLM now completes a trailing assistant message insteadUpdate miLLM; no client change needed. See OpenAI-compatible API

SAE attachment​

SymptomCauseFix
400 SAE_INCOMPATIBLE (dimension mismatch)SAE d_in ≠ model hidden sizeUse the SAE built for this exact model size (2b vs 9b vs 27b)
Attach warning: trained on different layer/modelDeliberate mismatch or wrong fileHeed it — features are only meaningful at the trained layer on the trained model
409 SAE_ALREADY_ATTACHEDOne SAE at a timeDetach first
Detach appears to hangWaiting (up to 30 s) for in-flight requests to finishExpected — the hook is never removed mid-generation

Inference & API​

SymptomCauseFix
503 model_not_loaded on /v1/*No model loadedLoad one; consider AUTO_LOAD_MODEL for restarts
503 QUEUE_FULL errorsRequest queue full (MAX_PENDING_REQUESTS) — backpressureRaise the limit, or slow the client (Open WebUI's parallel title-generation is a common culprit)
400 context_length_exceededprompt + max_tokens > model contextShorten the prompt or reduce max_tokens
Streaming stops with an in-stream error eventGeneration failed after headers were sent (HTTP is already 200)Check server logs; SSE consumers should watch for error objects before [DONE]
Responses look like completions, not chatBase (non--it) modelExpected — base models aren't instruction-tuned (why you might want that)
CORS errors in browserOrigin not allowedSet CORS_ORIGINS (Configuration)
500s on /v1/chat/completions after a crashLeaked GPU stateRestart the backend, reload the model

Monitoring​

SymptomCauseFix
No activation recordsMonitoring disabled, or no completions since enablingGET /api/monitoring to check state; records are written once per completed generation
Wrong/garbage feature indices in history(Fixed in current versions) watched-feature list desyncUpgrade; configure via /api/monitoring/configure
History entry per request, not per tokenBy design — records capture the final forward passSee monitoring semantics
WebSocket disconnects during long generationsEvent loop busyNormal; the client auto-reconnects — resync state via REST after reconnect

Kubernetes-specific​

kubectl get pods -n millm # pod states
kubectl logs -n millm deploy/millm-backend # backend logs
kubectl delete pod -n millm <pod-name> # restart to clear leaked VRAM

GPU scheduling requires the NVIDIA device plugin and a node with the GPU visible; see the Kubernetes install guide.

Still stuck?​

  • Set LOG_LEVEL=DEBUG and reproduce — steering/hook activity is logged
  • GET /api/health/detailed gives a per-component health breakdown
  • Open an issue: github.com/Onegaishimas/miLLM with logs and your model/SAE combination