explainer / Sep 27, 2026

Ollama Memory Has Three Budgets, Not One

A model download is not a memory budget. Understand loaded models, context length, and parallel requests before sizing local AI for a shared home server.

By Stackarr EditorialOllama · Local AI · Memory planning · Homelab
A compact workstation on a timber desk beside a closed notebook and ceramic mug, lit by a window and one violet power light.
Plan local AI around the workload sharing the home server, not only the model download.

Ollama capacity planning needs three separate questions: which models stay loaded, how much context each request needs, and how many requests run together. A model that works for one short conversation is not proof that the same home server can support a document-heavy agent and several simultaneous users.

This is a sizing explainer for locally executed models, not a deployment guide or a hardware benchmark. It applies to an existing self-hosted Ollama installation. Cloud-hosted inference has a different resource boundary. The observations below require permission to inspect the running instance; changing its startup configuration requires the operator's authority. No firewall, Docker port, or storage changes are needed to understand the budget.

Loaded models consume capacity between conversations

Ollama can leave a model resident after it finishes answering. That trades memory occupancy for less loading work on the next request. An idle chat window therefore does not imply an empty model runtime. According to the Ollama FAQ, models normally remain loaded for five minutes, but request-level and server-level retention settings can change that behavior.

The distinction matters on a shared homelab machine. A photo service, media server, or backup process still needs resources even when local AI is waiting. Keeping a frequently used assistant warm may be reasonable; keeping every occasional model resident is a different capacity decision. Neither policy can be judged from the download size alone.

The API's keep_alive setting controls residency, while OLLAMA_KEEP_ALIVE supplies a server setting. The FAQ specifies that a request-level value takes precedence. A front end can therefore influence model retention even when the server operator expects its default to apply. Residency is not the same control as the number of simultaneous models.

Context is a separate memory commitment

Context length determines the token window available to a model. Ollama's context documentation explicitly links larger windows to greater memory requirements. A long document or an agent's accumulated conversation can create a different demand from a brief chat using the same model.

Do not assume one universal default across installations. The dedicated context page describes hardware-dependent defaults, while the FAQ also contains a fixed-default description. That inconsistency is a reason to inspect the running allocation rather than build a purchasing decision around a remembered number. Explicit configuration and the deployed version matter.

For planning, separate the model's supported maximum from the context actually needed by the task. A larger advertised window is not a promise that the whole workload fits comfortably on the available accelerator. Conversely, reducing context just to fit hardware can undermine a task that genuinely needs the missing material.

Parallel work magnifies the context cost

Multiple requests are not automatically just more waiting time. When Ollama processes them in parallel for a loaded model, context-related memory demand increases. The FAQ describes the relationship between OLLAMA_NUM_PARALLEL and OLLAMA_CONTEXT_LENGTH; it also distinguishes parallel requests from loading several different models.

That produces two independent questions: how many model runtimes fit, and how many requests each runtime can serve together. A single successful prompt answers neither question for peak household use. Test conditions should resemble the intended workload, including context and overlapping demand, before treating a configuration as sufficient.

A queue is another boundary, not extra memory or extra inference capacity. OLLAMA_MAX_QUEUE determines how much pending work can wait before additional requests are rejected. Raising it may allow a longer backlog; it does not make a memory-constrained model run faster.

Read the allocation, not just the model label

On an existing installation, ollama ps is a read-only view of loaded models. The context documentation identifies its PROCESSOR field as the place to inspect CPU versus GPU placement. Its context and expiry information help distinguish an actively loaded runtime from a model merely stored on disk.

For software consuming the same information, the documented running-model API is GET /api/ps. Its response includes size_vram, context_length, and expires_at. Use the existing trusted connection to the instance; do not publish the API to the internet to collect these fields.

These are runtime observations, not a guarantee of application speed. CPU offloading, context demands, and competing services still need interpretation. An expiry timestamp explains intended residency, not whether a later request will extend it.

Choose the trade-off before buying hardware

A useful capacity note records decisions rather than one headline memory figure:

  • Which exact models must be available together?
  • Which tasks genuinely need long context?
  • How many requests need concurrent execution rather than queueing?
  • Which models benefit from remaining warm between requests?
  • What resources must remain available for the rest of the home server?

For an occasional assistant, accepting model loading time may be preferable to permanent residency. For several active users, concurrency deserves more attention than idle retention. For document-heavy work, the required context may be the dominant constraint. These are workload choices, not universal performance prescriptions.

What this budget cannot promise

There is no reliable single conversion from a model's downloaded bytes to the total RAM or VRAM needed by every deployment. The official documentation establishes the controls and their relationships, but it does not supply a benchmark for a particular home server. Hardware backends, model selection, and request patterns limit how far another installation's result transfers.

The practical outcome is a better question: can the chosen model, required context, and expected concurrency coexist with the other self-hosted services? Record those conditions alongside runtime observations. That makes a later hardware purchase or configuration change a response to a specific constraint rather than a guess based on model size.

Verification ledger

Sources and further reading

  1. Ollama FAQ: model residency and concurrent requestsOllama · Primary source
  2. Ollama context length and processor offloadingOllama · Primary source
  3. Ollama API: list running modelsOllama · Primary source