# Small Quantized Model Starter for a 96 GB Server Version: October 4, 2026 Article: https://jwatte.com/blog/small-quantized-models-96gb-ai-server/ Base server guide: https://jwatte.com/downloads/sovereign-ai-setup-guide.md Use this kit to test a bounded local task before relying on it for business work. It is not an unattended OS installer. It does not assume that the largest downloadable model is the best choice. No passwords, API keys or customer information belong in the kit. ## 1. Check the intended machine ```bash hostname uname -r free -h df -h / ss -lnt curl --fail --silent http://127.0.0.1:11434/api/version curl --fail --silent http://127.0.0.1:11434/api/tags ``` Stop if this is not the expected host or Ollama is not your intended local service. Keep port 11434 private. These checks do not install or restart anything. The reference host used Ubuntu 26.04.1, two Intel Xeon Silver 4214R processors, 24 physical cores in total, 48 logical threads, 96 GB nominal RAM and mirrored NVMe storage. Its LLM tests were CPU-only. Treat these as the reference machine's facts, not a hardware requirement for every reader. ## 2. Start with explicit limits The reference Compose service starts with one loaded model and one parallel request. The Ollama container has a 48 GiB memory ceiling and a CPU quota. A quota is not physical-core pinning and a memory ceiling is not reserved memory. For a new deployment, review this example against your actual Compose file before merging it. Never replace a working configuration blindly. ```yaml services: ollama: environment: OLLAMA_NUM_PARALLEL: "1" OLLAMA_MAX_LOADED_MODELS: "1" OLLAMA_KEEP_ALIVE: 5m ports: - "127.0.0.1:11434:11434" mem_limit: 48g cpus: 24 ``` Leave memory for the operating system, builds, databases, browser desktop, backups and other services. Use lower limits on a smaller machine. Inspect actual host pressure before increasing context or concurrency. Official configuration: https://docs.ollama.com/faq Context guidance: https://docs.ollama.com/context-length ## 3. Select the model deliberately Our extraction test used hermes3:8b in Q4_0. A separate short routing experiment used qwen2.5:1.5b in Q4_K_M and hermes3:8b. Neither test establishes the best model for other workloads. Check the exact license and model metadata before downloading. Pulling a model uses bandwidth and storage. This request asks your local Ollama service to download the selected model; it does not call a paid inference API. ```bash curl --fail --show-error http://127.0.0.1:11434/api/pull \ -H 'Content-Type: application/json' \ -d '{"model":"hermes3:8b","stream":false}' ``` Inspect metadata rather than guessing the quantization from a model name: ```bash curl --fail --silent http://127.0.0.1:11434/api/show \ -H 'Content-Type: application/json' \ -d '{"model":"hermes3:8b"}' ``` Do not download a 70B model just to fill RAM. Measure its actual latency and memory needs first. Model weights, KV cache, context, parallelism and runtime overhead are different parts of the memory budget. ## 4. Reproduce the published extraction experiment Create a new directory, save the Python script there, inspect it, then run it. It uses Python's standard library and localhost. The script refuses to overwrite an existing results file. ```bash set -eu run_dir="$HOME/quantized-pilot-$(date +%Y%m%d-%H%M%S)" mkdir "$run_dir" cd "$run_dir" curl --fail --show-error --location \ https://jwatte.com/downloads/quantized-mini-benchmark.py \ --output quantized-mini-benchmark.py python3 -m py_compile quantized-mini-benchmark.py ``` Read the script before the next command. It makes nine synthetic local requests, which will consume CPU time. It writes results beside the script. ```bash python3 quantized-mini-benchmark.py ``` Keep the published results in a different directory for comparison, because placing them beside the script would correctly trigger its overwrite protection: ```bash mkdir reference curl --fail --show-error --location \ https://jwatte.com/downloads/quantized-benchmark-results.json \ --output reference/published-results.json ``` Recorded reference result: nine exact records from three cases repeated three times; 36 fields matched; median total request time 3.011 seconds; first request 12.462 seconds including 7.489 seconds loading; median output generation rate 13.738 tokens per second. This is a narrow synthetic result, not a production accuracy estimate. The original script is supplied unchanged for reproducibility. It is a demonstration, not a complete input-validation framework. In particular, your production validator must verify required-field presence and types as well as values. ## 5. Inspect the second, smaller-model experiment The article also links to https://jwatte.com/downloads/quantized-routing-results.json and https://jwatte.com/downloads/quantized-routing-benchmark.py. That separate experiment uses three simple category messages, two models and two requested thread counts. It makes twelve routing requests, plus warmups. Do not mix its timing with the extraction test: the prompts, schemas, context and output lengths differ. The routing script downloads qwen2.5:1.5b when it is absent. Review storage capacity and the script before running it. Both experiments use synthetic material only. ## 6. Build a real acceptance set Choose a bounded output schema. For example: ```json { "reference": "JOB_001", "quantity": 2, "due_date": "2026-10-09", "needs_review": false, "source_excerpt": "Job JOB_001 needs two tires by 2026-10-09." } ``` Require null or an explicit unknown for missing information. Verify each reference against the input. Parse dates with code. Check arithmetic with code. Preserve the excerpt and original document instead of retaining only a summary. Include examples with missing quantities, conflicting dates, corrections in later paragraphs, duplicate references, unrelated material and instructions embedded in the source. The model must treat the source as data, not authority to expand the task. Keep a held-out set separate from examples used to adjust the prompt. Do not change expected answers after seeing an inconvenient model result. Record every failed case. ## 7. Compare the entire workflow For each candidate, record model digest, quantization, runtime, prompt version, context, thread count, request concurrency, warm/cold state, total latency, load time, output tokens, peak observed memory and reviewer corrections. Use the same input and task when comparing models. Ollama reports timing fields in nanoseconds. Output generation rate is eval_count divided by eval_duration in seconds. It is not end-to-end requests per second. See https://docs.ollama.com/api/usage. A smaller brief saves external input tokens only when it preserves what the next step needs. Track outside-model usage, retries and human correction time. Never silently fall back to a paid provider. Unsupported cases can go to a human. Suggested CSV header: ```text run_id,model,digest,quantization,context,threads,parallel_requests,cold_start,total_seconds,load_seconds,input_tokens,output_tokens,valid_schema,exact_match,review_minutes,external_cost,notes ``` ## 8. Optional components and their purpose | Component | Add it for | |---|---| | Qdrant | Searchable document vectors with provenance and retrieval tests | | n8n | A tested ingestion and review workflow with safe retries | | Cloudflare Tunnel and Access | Identity-protected HTTPS access to private services | | Tailscale | An independent private route for enrolled devices | | Browser desktop | A persistent approved browser session | | Deno | A project written for that runtime | | Vercel CLI | A project actually hosted on Vercel | | PM2 | An existing Node service whose chosen supervisor is PM2 | | Gitleaks | Review of possible secrets before committing code | | Trivy | Review of runtime and container vulnerabilities | | Figma integration | Approved report or interface design context | | Ideogram | Conceptual artwork from a nonconfidential brief | Ollama serves models; LiteLLM routes configured requests; PostgreSQL stores gateway state; Open WebUI is the chat interface. Restic and rclone serve different recovery and transport roles. These components are not interchangeable and installing one does not prove its workflow is tested. ## 9. Promotion checklist - [ ] Local endpoint is private and authenticated gateways reject unauthenticated requests. - [ ] Required fields, types, dates and evidence are checked outside the model. - [ ] Missing, contradictory and hostile inputs are in the acceptance set. - [ ] Held-out results and reviewer corrections are recorded. - [ ] Latency and memory are measured during combined server load. - [ ] Cloud escalation is explicit and budgeted. - [ ] Source material and database exports have an independent tested backup. - [ ] The workflow has a manual fallback and named owner. Passing the included toy test is the start of evaluation, not the end.