Ollama CLI Cheatsheet - Local LLM Runtime Full Reference
Ollama is the de-facto runtime for running LLMs locally, for developers and privacy-sensitive teams who want models offline on their own machine plus one-command serving. It is far easier than hand-rolling llama.cpp, with unified model management and an OpenAI-compatible API. This sheet covers pulling/running models, tag management, Modelfile temperature/weight tuning, and serving.
Run & Chat 5
ollama run llama3.1ollama run llama3.1 "summarize in one line"ollama run -mollama serveollama psModel Management 5
ollama pull <model>ollama push <model>ollama listollama rm <model>ollama cp <src> <dst>Modelfile 5
FROM llama3.1PARAMETER temperature 0.7SYSTEM "You are a rigorous assistant"TEMPLATE """{{ .Prompt }}"""ollama create mymodel -f ModelfileServe & Port 5
ollama serveOLLAMA_HOST=0.0.0.0 ollama servecurl http://localhost:11434/api/generateollama show <model>ollama embeddingsTips 5
ollama run llama3.1 -c 4096ollama list | headollama run --verboseollama ps to watch VRAMollama run <m> "translate..."Tips
- Ollama is the de-facto local runtime; a "local model" entry covers privacy/offline/zero-API-cost cases that complement cloud CLIs.
- Bake behavior into a Modelfile SYSTEM/PARAMETER — more stable and versionable than repeating system instructions in prompts.
- Tight on VRAM? Cap context with `-c` or watch parallel usage with `ollama ps` to avoid OOM.
- For programmatic use run `ollama serve` + the REST API instead of interactive mode.
FAQ
What are the most common Ollama commands?
The daily drivers are ollama run <model> to pull and chat interactively, ollama pull <model> to download without running, ollama list to see local models, ollama ps for running models, ollama create to build a custom model from a Modelfile, and ollama serve to expose models as an OpenAI-compatible API on localhost:11434.
How do I download and run a model with Ollama?
Run something like ollama run llama3.2 — if the model is not local it is pulled from the library automatically, then you drop into a chat. To download only, use ollama pull <model>.
How do I run Ollama as an API server?
ollama serve starts a local server on port 11434 with an OpenAI-compatible /v1 API. You often do not need to start it manually because ollama run brings the backend up on demand. Point your OpenAI client base URL at http://localhost:11434 to use it.
How do I list my local Ollama models and see what is running?
ollama list shows every pulled model with size and modified time; ollama ps shows only models currently loaded in memory. Remove one with ollama rm <model>.
Official References
Each command links to its official documentation below, so you can verify the latest usage and read deeper.
Maintained by LaoHand
Publicly updated on Aug 21, 2026, continuously proofread against official docs.
Contact Us
Wrong command or description? Send us corrections, business inquiries or product feedback by email.
Contact Us