# Small Quantized Models: Getting Useful AI Work from a 96 GB Server

Use a small local model to prepare information before calling a larger AI. See our 96 GB server test results, memory tradeoffs and a downloadable starter kit.

Author: J.A. Watte
Published: October 4, 2026
Source: https://jwatte.com/blog/small-quantized-models-96gb-ai-server/

---


The most useful job for a small language model is often the least glamorous one. Read a short note. Extract the reference number. Mark the missing date. Return a small record that another program can check.

That can be a better use of a rented server than asking the largest model that fits in memory to handle every request. It can also reduce the amount of material you send to a paid model. The condition is that the preparation must preserve the information the next step needs. A shorter prompt is not a saving if it quietly loses the important exception.

We tested that narrow idea on the 96 GB machine described in the [Sovereign AI Playbook](/blog/sovereign-ai-playbook-remote-desktop-commander/). The results below are actual local measurements from October 4, 2026, with a deliberately small synthetic test. They are not a benchmark of all small models or proof of production accuracy.

<p><strong>Try it yourself:</strong> <a href="/downloads/quantized-local-model-starter.md" download>Download the local model starter kit (.md)</a>, <a href="/downloads/quantized-mini-benchmark.py" download>the Python test</a>, and <a href="/downloads/quantized-benchmark-results.json" download>our recorded results</a>. You can also <a href="/downloads/small-quantized-models-96gb-ai-server-article.md" download>save this article as Markdown</a>.</p>

## Quantization saves space, not judgment

A model contains numerical weights. Quantization stores an approximation of those values using fewer bits. The file gets smaller and the runtime may need less memory or move less data. The trade is that some numerical precision is lost. How much that affects the job depends on the model, quantization method, runtime, and input.

For a rough example, eight billion weights stored at 16 bits take about 16 billion bytes before other data. Four bits per weight would take about four billion bytes. Real files also need scales, metadata, and other structures, so that arithmetic is a planning estimate, not an exact download size. [llama.cpp quantization tooling](https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md).

The Hermes 3 8B file installed on our server uses **Q4_0** and occupies **4,661,227,243 bytes**, about 4.66 GB. The installed Llama 3 70B file also uses Q4_0 and occupies **39,969,745,349 bytes**, about 39.97 GB. Those are actual file sizes reported by Ollama, not the memory required for every possible request.

Do not read Q4_0 as a universal recommendation. We measured the model already installed in this build. Comparing another four bit format or an eight bit copy requires another controlled test. It would be misleading to claim that this run proved which quantization is best.

## Small models are good candidates for bounded work

I would first try classification, straightforward field extraction, short draft summaries, and converting repetitive prose into records. A repair note with one job reference and a clear date is a better starting point than a complicated contract with scattered exceptions.

Before using a model, check whether code can do the job. Adding quantities, checking a reference against a known list, matching a date format, and detecting an exact duplicate are ordinary programming tasks. There is no advantage in making an LLM guess an answer that a short function can calculate.

A practical sequence is:

1. Parse and clean the input with ordinary code where possible.
2. Ask the local model for a small, explicit record.
3. Validate that record and retain the source evidence.
4. Escalate ambiguous cases to a person or an approved larger model.

That is a routing decision, not a promise that a smaller model can replace the larger one. In particular, do not let a preliminary summary become the only surviving record of a customer's instructions.

## What we actually measured

The test used Ollama **0.35.1**, Hermes 3 8B in Q4_0, a **4,096 token context setting**, **24 requested CPU threads**, temperature zero, seed 42, and a maximum of 128 generated tokens. Ollama reported no GPU allocation for the loaded model.

There were three invented inputs: a bicycle repair with complete fields, a catering inquiry missing its quantity and date, and a cleaning job with an old deadline replaced by a confirmed new deadline. Each input ran three times, making nine calls. Every expected record had four fields. These were short notes, not invoices, scanned documents, long conversations, or hostile prompt injection tests.

| Measurement | Observed result |
|---|---|
| Exact expected records | 9 of 9 calls |
| Expected fields matched | 36 of 36 fields |
| Parseable JSON | 9 of 9 calls |
| Median total request time | 3.011 seconds |
| Fastest total request time | 2.716 seconds |
| First request time | 12.462 seconds |
| Model loading within the first request | 7.489 seconds |
| Median generation rate | 13.738 tokens per second |
| External model API calls | None |

The generation rate is calculated from Ollama's output token count and evaluation duration. It is not total request throughput. Total time also includes loading and processing the prompt. These calls went directly to local Ollama, so they are not a benchmark of the browser, Cloudflare, LiteLLM, or the full application chain. [Ollama usage metrics](https://docs.ollama.com/api/usage).

After the test, Ollama reported a loaded model allocation of **5,316,594,891 bytes**, about 5.32 GB, at the selected context length. That is the runtime's model allocation report, not the entire server's resident memory. The operating system, file cache, applications, and other processes still use memory.

The missing fields stayed null. The replacement deadline was selected instead of the old one. Those are useful acceptance checks. But nine successful calls are only three distinct cases repeated. They do not establish a general accuracy percentage, a confidence interval, or protection against a cleverly written malicious input.

The larger Llama model was downloaded, but it was **not part of this comparison**. Its presence on disk is not a measured performance result. The [test script and results](/downloads/quantized-benchmark-results.json) make the scope inspectable.

## A second test: how little model does a routing job need?

We also ran three short category messages through **Qwen 2.5 1.5B in Q4_K_M** and **Hermes 3 8B in Q4_0**, each at 12 and 24 requested threads. The messages concerned maintenance, scheduling and billing. Each configuration received the same three cases once, after a warmup. Context was 2,048 tokens and output was capped at 48 tokens.

| Model | Requested threads | Median total request time after warmup | Correct categories |
|---|---:|---:|---:|
| Qwen 2.5 1.5B | 12 | 0.416 seconds | 3 of 3 |
| Qwen 2.5 1.5B | 24 | 0.365 seconds | 3 of 3 |
| Hermes 3 8B | 12 | 1.394 seconds | 3 of 3 |
| Hermes 3 8B | 24 | 1.046 seconds | 3 of 3 |

The Qwen file occupies 986,061,892 bytes, about 0.99 GB. That is a concrete reason to test a genuinely small model for an easy routing job. It returned the expected categories faster in these twelve calls. It does not follow that it is more accurate on difficult messages, extraction, coding, or reasoning. We did not measure sustained concurrent load, repeat the routing cases enough to estimate variability, or compare equivalent quantization formats.

The 24 thread setting was faster in this small test, but that does not prove it is the best setting for every workload on a dual socket server. Warmup requests took roughly 2.4 to 2.6 seconds for Qwen and 7.9 to 8.1 seconds for Hermes. Keep loading time separate from an already resident model's response time.

[Download the routing results](/downloads/quantized-routing-results.json) and the [exact routing script](/downloads/quantized-routing-benchmark.py). These results are separate from the nine extraction requests above. They use different prompts and output lengths.

## Keep the context shorter than the advertisement

A model may advertise a large maximum context. That is not a suggestion to put every document into every request. More context and parallel requests can consume additional memory. Ollama documents how its context and concurrency settings affect operation. [Context length](https://docs.ollama.com/context-length) and [concurrency](https://docs.ollama.com/faq).

We started at 4,096 tokens for short extraction. That is a test setting, not a universal limit. A longer document needs deliberate handling. Break it at meaningful boundaries, keep document and section references, and carry necessary neighboring text where a sentence depends on what came before.

Suppose a booking note says the event date changed in a later paragraph. Extracting only the first paragraph produces a tidy but wrong record. A useful pipeline must preserve the correction or mark the result for review. Lowering context until a model runs quickly is not an optimization if it removes the evidence.

Use a budget for the system prompt, source text, and answer together. Reject or split input that exceeds the tested budget instead of silently truncating it. Measure the actual prompt token count when the runtime provides it.

## Give the 96 GB machine room to do its other jobs

The server is not just a model loader. It also runs a database, a gateway, a chat interface, development tools, and optional browser or automation services. A large model should not make those services unreliable.

The reference deployment starts with one model loaded at a time and one parallel model request. Its Ollama container has a 48 GiB memory limit and a CPU quota. Those limits are ceilings, not reserved resources. Other containers need their own limits, and the real test is what happens under combined load.

For another 96 GB installation, begin with the small model, modest context, and serial requests. Inspect available host memory before increasing any of them. Leave room for builds, database work, backups, and a browser rather than assigning all advertised RAM to model weights.

A requested thread count is not CPU affinity. A Docker CPU quota is not pinning to selected physical cores. We requested 24 threads in this test; we did not establish that NUMA placement or every vector instruction was optimal. On a machine with two CPU sockets, testing another thread count may be useful, but the answer should come from measurements on that host.

This is why I would not make the 70B model the default merely because the file fits. Check its latency, peak memory, context needs, and output quality on the actual task. Keep the smaller route when it meets the acceptance criteria.

## Request a schema, then check the meaning

Ollama accepts a JSON schema for structured output. That makes a parser's job easier, but valid JSON can still contain an invented date or the wrong quantity. The test asks for four fields and checks the returned values against known answers. [Structured output support](https://docs.ollama.com/capabilities/structured-outputs).

For real work, validate reference IDs, types, dates, allowed categories, and required evidence. Prefer an explicit missing value to an invented replacement. Preserve the source document, a content hash, and the exact excerpt supporting each important field.

Do not trust a model's self reported confidence as the gate for automatic action. A more useful gate is whether the necessary source evidence exists and the result passes rules you can inspect. Unsupported fields should go to a review queue.

The same applies to redaction. A small model may help identify information for removal, but that does not certify that an export contains no personal data. Use deterministic rules and a human review where the consequence warrants it. An outside model should receive only material you have deliberately approved.

## The saving comes from sending less unnecessary material

Consider an illustrative batch of 1,000 documents. Sending 12,000 input tokens from every document would mean 12 million input tokens. Sending a verified 750 token brief from each would mean 750,000. That is a 93.75 percent reduction in the material sent onward.

It is not automatically a 93.75 percent reduction in the bill. Provider prices, cached input, output length, retries, and the number of cases escalated all matter. Local processing and review also consume time and server capacity. This is arithmetic for a planning example, not an observed saving from our nine requests.

Compare the entire workflow. Record the local processing time, the outside model's usage, the human correction time, and the rate at which the smaller brief omits necessary evidence. When the original text is important, send the relevant original passages along with the record rather than substituting an unsupported summary.

Our [two lane server setup](/downloads/sovereign-ai-setup-guide.md) keeps supported subscription logins separate from API billing. A local preprocessing step does not turn a subscription into an API allowance or remove its limits. Nor should a failed local request silently become a paid cloud request.

## Other model families are candidates, not assumed equivalents

Hermes handled the extraction test. We also ran the small routing comparison below. A production selection still needs a representative test set for its own task.

Google's **Gemma** is an open model family to evaluate locally; **Gemini** is not simply another name for those downloadable weights. Alibaba's **Qwen**, suitable **Mistral** releases, and OpenAI's **gpt-oss** are other documented families to investigate. Check the exact model, license, supported runtime, available quantization, context behavior, and memory requirements before downloading. [Gemma](https://ai.google.dev/gemma/docs), [Qwen](https://github.com/QwenLM/Qwen3), [Mistral deployment](https://docs.mistral.ai/deployment/self-deployment/overview), and [gpt-oss](https://openai.com/open-models/).

Kimi Code or another agent running on your Linux machine does not necessarily mean its model runs locally. A client may call a cloud model. With very large or mixture of experts models, the number of active parameters per token is also not a sufficient estimate of the memory needed to hold the weights. Do not infer that a downloadable flagship will fit because a terminal client installed successfully.

Qwen has one narrow measured routing result here. We did not benchmark Gemma, Mistral, or gpt-oss in this exercise. The [companion comparison](/blog/desktop-commander-cowork-browser-agents-small-business/) covers their task and browser products separately. A product without a verified local model or supported browser equivalent should be marked that way rather than given an invented one.

## Where the rest of the software helps

Ollama loads the local model. LiteLLM can expose it through a controlled gateway. PostgreSQL keeps the gateway's persistent state, and Open WebUI provides the chat interface. Aider and native coding clients handle editing workflows. None of these removes the need to review a command before it changes something important.

The optional components in the Sovereign AI Playbook belong around that pipeline only when they solve a specific problem:

| Optional component | Useful role in this setup |
|---|---|
| Qdrant | Store searchable document vectors when the collection is too large for simple lookup. Retrieval supplies candidate evidence, not a guarantee of truth. |
| n8n | Coordinate ingestion, validation, review queues, and approved delivery. Keep retries from producing duplicate actions. |
| Cloudflare Tunnel and Access | Give approved users an HTTPS route to private applications. Test identity enforcement before exposing a terminal. |
| Tailscale | Provide a separate private connection for enrolled devices. It is not the same experience as opening a public browser link. |
| Browser desktop | Keep a persistent work browser near the server tools. Account sessions and downloaded files still require protection. |
| Deno | Run a preprocessing or deployment component that was written for Deno. It is not required just because an LLM is present. |
| Vercel CLI | Deploy a related interface when Vercel is its actual host. |
| PM2 | Supervise an existing Node worker where that is the chosen supervisor. Do not add a competing restart loop to a Docker service. |
| Gitleaks | Review possible secret exposure before committing the workflow or its example data. |
| Trivy | Review vulnerabilities in the runtime and container images, then act on relevant findings. |
| Figma integration | Bring approved interface or report designs into an editing task. It does not make inference faster. |
| Ideogram | Create conceptual artwork from a generic brief. It is an outside service, not local model processing. |

An embedding model is a separate choice from the chat model. Adding Qdrant does not automatically produce useful document vectors. Preserve document IDs, chunk boundaries, versions, and retrieval tests so a search result can lead back to the right source. [Qdrant documentation](https://qdrant.tech/documentation/).

Restic and rclone belong in the recovery plan. Back up the source documents and configuration as well as the database exports. A directory containing a Compose file does not contain all the data in its named volumes. Keep an independent encrypted backup and test a restore before removing the old workstation. [Restic](https://restic.readthedocs.io/en/stable/) and [Docker volumes](https://docs.docker.com/engine/storage/volumes/).

## Test the cases that could make the business regret trusting it

The next test should be harder than our three notes. Include missing fields, contradictory dates, unfamiliar wording, duplicate references, and instructions embedded in the input that try to change the task. Add scans only after the text extraction stage has its own checks.

Set aside examples that you do not use while adjusting the prompt. Compare candidates on those held out cases. Record failures instead of changing the expected answers to make the report look better. Repeat the run after a model, quantization, runtime, or prompt change.

Test under combined load as well. A model that responds quickly on an idle server may be inconvenient while a build, backup, and browser are competing for resources. Watch total latency, memory pressure, failures, and user correction time. More tokens per second is not the only useful outcome.

Finally, keep a manual route. A small model should make an ordinary job easier to review, not become the only way to understand the original material. The best result may be a short record, a source quote, and an honest request for clarification.

<p><a href="/downloads/quantized-local-model-starter.md" download><strong>Download the starter kit</strong></a> to run the synthetic test, record your own results, and build an evaluation set for your business. <a href="/downloads/small-quantized-models-96gb-ai-server-article.md" download>Download the full article</a>.</p>

*Measurements are from one server and three synthetic inputs repeated three times. They do not establish general accuracy, a production response time guarantee, or comparative superiority. The hero is a conceptual illustration generated with Ideogram for this article. No real customer documents or outside model APIs were used in the reported benchmark.*


---

Canonical HTML: https://jwatte.com/blog/small-quantized-models-96gb-ai-server/
RSS: https://jwatte.com/feed.xml
JSON Feed: https://jwatte.com/feed.json
Hero image: https://jwatte.com/images/small-quantized-models-96gb-ai-server.webp
