#

inference

inference related cheatsheet collection with 5 quick references covering common commands, flags and real usage for developers and SRE.

llamacpp (llama.cpp) CLI Cheatsheet - llama-cli Reference

llamacpp (llama.cpp) llama-cli complete command reference: model load -m, interactive -i, GPU offload -ngl, context -c, threads -t, --temp sampling and -cnv chat, all matched to official docs.

2026-09-1020 commands
llama-cppllama-cligguf

LocalAI CLI Cheatsheet - local-ai Self-hosted Inference Server Full Reference

Full command reference for LocalAI local-ai CLI: run and serve to start, models list/install/pull for management, api and finetune commands, all matched to official docs.

2026-09-1020 commands
localailocal-aiself-hosted

LM Studio CLI Cheatsheet - lms Model Management Full Reference

Full command reference for LM Studio lms CLI: status, ls and ps to list models, get to download, load and unload, server start, and terminal chat, all matched to official docs.

2026-09-1020 commands
lmstudiolmslocal-llm

SGLang CLI Cheatsheet - sglang.launch_server Inference Full Reference

Full flag reference for SGLang launch_server: --model-path, --host and --port, --tp and --dp parallelism, --mem-fraction-static VRAM, --quantization, and --trust-remote-code, all matched to official docs.

2026-09-1020 commands
sglanginferenceserver

vLLM Cheatsheet - High-Performance LLM Inference & Serving

Command reference for vLLM, the high-performance LLM inference & serving engine: install, OpenAI-compatible server launch, host/port/API-key configuration, and quantization, covering pip install vllm, vllm serve, and --host/--port/--api-key.

2026-08-2325 commands
vllmllm-servingai