# Sovereign AI Workstation Playbook A build and operations guide for a Linux machine that keeps doing your work when nobody is logged in. Prepared October 8, 2026 UTC. Companion article: https://jwatte.com/blog/sovereign-ai-workstation-best-practices/ ## What this file is This is a reusable plan, not an installer. Every command is a small example you should read before running. Nothing here asks you to paste a script into a production machine unreviewed, and nothing here needs a graphics card. It is written for any Linux host: a rented bare metal box, a home server, a spare desktop, a small cloud instance. Where a number is given, it is the number that worked on one machine, not a requirement. ## 0. Decide what you are actually buying Write one sentence answering this before you spend anything: **what do I want to still exist tomorrow morning that does not exist today?** Good answers look like: a repository already checked out, a browser already signed in, a schedule that fired at 4am, a directory of notes from last month, a report waiting in my inbox. Answers that do not need this build: "faster responses", "a bigger model", "cheaper tokens". Those are model and billing questions, and a rented API usually wins them. If your answer is about continuity, keep reading. If it is about inference speed, buy a graphics card instead and stop here. ### Sizing | Workload | Reasonable floor | |---|---| | Agents, schedulers, repos, browser, no local models | 4 cores, 8 GB RAM, 100 GB disk | | The above plus containers and a desktop session | 8 cores, 32 GB RAM, 250 GB disk | | The above plus small quantized local models on CPU | 16 cores, 64 GB RAM, 500 GB disk | | Fast local inference on large models | A GPU, which is a different document | The reference machine for the companion article sits in the third row: 48 threads, 94 GB RAM, 877 GB mirrored disk, no GPU. It is comfortable, not minimal. ## 1. Base the machine on something boring Use a current long term support release of a mainstream distribution. Enable unattended security updates. Put the disk on a mirror if the hardware allows it. ```bash # confirm what you actually have before planning anything lscpu | grep -E '^Model name|^CPU\(s\)|^Core|^Socket' free -h df -h / lspci | grep -iE 'vga|nvidia|3d' ``` Record the output. You will want it later when something is slow and you are trying to remember what the machine is. ### One account owns the work Create one unprivileged account that owns the repositories, the keys, the browser profile and the scheduled jobs. Do not run any of this as root. Do not spread it across several accounts, because every file permission question afterwards becomes a puzzle. ## 2. Turn on lingering before anything else This is the single most common reason a setup like this works for a day and then stops. ```bash sudo loginctl enable-linger "$USER" loginctl show-user "$USER" | grep Linger # expect Linger=yes ``` Without it, every user service and timer stops when your last SSH session closes. ## 3. Put durable work in user units Not in a terminal multiplexer that someone left open, and not in a system unit that needs root. A minimal service: ```ini # ~/.config/systemd/user/example-agent.service [Unit] Description=Example long running agent After=network-online.target [Service] Type=simple ExecStart=%h/.local/bin/example-agent Restart=always RestartSec=10 # keep a crash loop from filling the disk StandardOutput=append:%h/logs/example-agent.log StandardError=inherit [Install] WantedBy=default.target ``` ```bash systemctl --user daemon-reload systemctl --user enable --now example-agent.service systemctl --user status example-agent.service ``` Rules that have held up: - `Restart=always` with a `RestartSec` of at least 5 seconds. Without the delay, a service that fails at startup will spin. - Log to a file you rotate, or to the journal with a size cap. An agent service can be chatty. - One service, one job. A service that does three things fails in three ways and tells you about one. ## 4. Watchdog the agent, because the agent is what fails The expensive, long lived process is the thing most likely to die. Do not ask it to notice its own death. ```ini # ~/.config/systemd/user/agent-watchdog.timer [Unit] Description=Check the agent session every five minutes [Timer] OnBootSec=2min OnUnitActiveSec=5min Persistent=true [Install] WantedBy=timers.target ``` ```bash #!/bin/bash # ~/.local/bin/agent-watchdog # Dumb on purpose. If it needs debugging, it is too clever. set -u if ! systemctl --user is-active --quiet example-agent.service; then logger -t agent-watchdog "agent not active; restarting" systemctl --user restart example-agent.service fi ``` The watchdog must be simpler than the thing it watches. If the watchdog can fail in an interesting way, you now have two problems. ## 5. Every scheduled run writes a status file This is the practice that pays for itself fastest. A scheduled agent that exits zero and writes nothing is indistinguishable from one that did its job. Make the absence of a report a failure. ```bash #!/bin/bash set -u OPS="$HOME/ops/myjob" STAMP=$(date -u +%Y%m%dT%H%M%SZ) MARK="$OPS/runs/.started-$STAMP" LOG="$OPS/runs/run-$STAMP.log" mkdir -p "$OPS/runs" : > "$MARK" # one run at a time, no overlap if a run goes long exec 9>"$OPS/.run.lock" flock -n 9 || { logger -t myjob "another run holds the lock"; exit 75; } timeout 2h my-agent-command > "$LOG" 2>&1 rc=$? # the agent is supposed to write status.json itself before exiting. # if it did not, the run failed even though the process may have exited zero. if [ ! "$OPS/runs/status.json" -nt "$MARK" ]; then printf '{"state":"no_report","rc":%d,"log":"%s"}\n' "$rc" "$LOG" > "$OPS/runs/status.json" fi rm -f "$MARK" state=$(python3 -c 'import json,sys;print(json.load(open(sys.argv[1]))["state"])' "$OPS/runs/status.json") logger -t myjob "finished rc=$rc state=$state" case "$state" in ok*) exit 0 ;; *) exit 1 ;; esac ``` The status file should carry a **state name**, not a boolean. Useful states: `ok`, `ok_nothing_to_do`, `blocked_dependency`, `partial`, `failed`, `no_report`. "Nothing happened and that is fine" and "nothing happened and that is wrong" must not look the same. If the job consumes a quota, record how much it spent. A failed run should not cost tomorrow's budget, and you can only prove that if the number is written down. ## 6. Give an unattended agent a small, explicit tool list Automatic approval is what makes unattended work possible, and it is the part that deserves the most care. The shape to aim for: ```bash my-agent \ --permission-mode default \ --allowedTools 'Bash(myjob-tool:*)' 'Read(//home/me/ops/myjob/**)' \ --max-turns 200 \ "$PROMPT" ``` Two ideas are doing the work there. **Write a wrapper command.** Instead of granting general shell access, write one small script that exposes exactly the verbs the job needs, and grant only that. The list of things the agent can do becomes a file you can read in one sitting. ```bash # ~/.local/bin/myjob-tool case "${1:-}" in plan) ... ;; check) ... ;; record) ... ;; status) ... ;; *) echo "usage: myjob-tool {plan|check|record|status}"; exit 2 ;; esac ``` **Scope the reads and writes to one directory.** An unattended agent should be able to do the specific job and nothing adjacent to it. Write the boundary into the prompt as well as the flags. State plainly what the agent must never do, and what it should do instead when it is blocked. An agent told "if X is unavailable, record this state and stop" behaves far better than one left to improvise. ## 7. Treat a browser dependency as an outage risk Some services have no API for the thing you need. The action exists only as a button in a signed in session. When that is true, your schedule now depends on a browser, which has none of the properties you want in a dependency: it holds credentials, its session expires, it is tied to one machine's desktop, and when it is missing the job cannot degrade, only stop. If you cannot avoid it: - Have the agent **verify the right browser** before it touches anything, and stop with a named state if it is the wrong one. - Never let a job fall back to a different browser or a different profile than the one it was scoped to. - Put the browser connection in your daily check, next to disk and services. - Expect to reconnect it by hand sometimes. Budget for that rather than engineering around it. ```bash # daily check, one screen systemctl --user --failed systemctl --user list-timers --no-pager | head df -h / | tail -1 for d in ~/work/*/; do s=$(git -C "$d" status --porcelain 2>/dev/null | wc -l) b=$(git -C "$d" branch --show-current 2>/dev/null) [ "$s" != "0" ] && echo "DIRTY $d ($b, $s files)" done cat ~/ops/*/runs/status.json 2>/dev/null ``` ## 8. Derive anything that appears in more than three places A value hand copied into many files will drift, and it will drift quietly. The failure to picture: a constant with a comment saying "bump this on every change", copied by hand into fifty four places. The constant was bumped. The copies were not. The mechanism it existed to provide had silently stopped working, and the comment describing that exact failure was sitting in the file where it was happening. A comment is not a mechanism. For anything that must agree across files: 1. Pick one source of truth. 2. Write the script that rewrites the copies from it. 3. Write the check that exits non zero when they disagree. 4. Run the check in whatever gate you already run. The same applies to operational lists: repositories to back up, services to check, domains to renew. Derive them from the filesystem or from one declared file. Every hand maintained second copy is a future surprise. ## 9. Test the populated path If your checks run against an empty environment, which is normal for anything scheduled or containerized, then every loop over an empty list runs zero times and the code inside it is never executed. This is not theoretical. A sixty nine check suite reported everything clean while four live pages returned server errors, because the thing that broke was inside a loop the tests never entered. ``` # the suite said OK. production said: TypeError: Cannot read properties of undefined (reading 'slug') ``` Two rules: - Feed at least one **synthetic fixture with real shape** into every renderer, report builder and transform. One record with data proves more than a thousand empty runs. - **Break it on purpose** and confirm the check goes red. A check you have never seen fail is a check you do not have. ## 10. Back up on one timer, verify on another ``` backup.timer daily backup-verify.timer weekly ``` A backup job reporting success is reporting that it finished, not that the archive restores. Put the verification on its own schedule so "we have been writing corrupt archives for a month" has a bounded lifetime. Verification means actually extracting something and comparing it, not reading the archive's own index. ```bash # weekly, pick one archive and prove a known file comes back intact latest=$(ls -t "$HOME/backups"/*.tar.zst | head -1) tmp=$(mktemp -d) tar --zstd -xf "$latest" -C "$tmp" ./canary.txt sha256sum -c <<<"$(cat "$HOME/backups/canary.sha256") $tmp/canary.txt" \ && echo "verify ok $latest" || echo "VERIFY FAILED $latest" rm -rf "$tmp" ``` ## 11. Keep memory in files the next session can read A directory of small Markdown notes, one fact each, with an index that loads at the start of every session. Not a database. ``` memory/ INDEX.md one line per note, loaded every session deploy-path-for-site-x.md owner-prefers-no-dashes.md why-we-dropped-approach-y.md ``` Write down what the repository **cannot** tell you. Code structure, commit history and configuration are already recoverable. What is not recoverable is why a decision was made, what was tried and rejected, and what the person you work for actually wants. Delete a note when it turns out to be wrong. A stale note is worse than a missing one, because it will be trusted. ## 12. Acceptance checklist Do not consider the build done until every line passes. - [ ] `loginctl show-user "$USER"` reports `Linger=yes` - [ ] Reboot the machine. Every service and timer comes back without a login. - [ ] `systemctl --user --failed` is empty. - [ ] Kill the main agent process by hand. The watchdog restarts it within its interval. - [ ] Every scheduled job has written a `status.json` with a state name in the last 24 hours. - [ ] Make a scheduled job fail on purpose. Confirm the status file says so and the exit code is non zero. - [ ] The unattended agent's tool list is a file you can read in under a minute. - [ ] Ask the agent to do something outside its scope. Confirm it refuses rather than improvising. - [ ] Restore one real file from the newest backup, on a different path, and compare hashes. - [ ] Disconnect any browser the schedule depends on. Confirm the job records a blocked state instead of acting wrongly. - [ ] Your daily check fits on one screen and you have run it by hand at least twice. ## 13. What not to build - **More desktops and portals than you need.** Every convenience service is another authentication surface. A plain SSH session does most of it. Add the rest only when a real requirement, like phone access, makes you. - **A long lived test container.** A throwaway container started per task, with no network, has been better in every case where isolation actually mattered. - **Agents with broad permission "for now".** The narrow ones are the ones that still work unattended. Broad grants do not get narrowed later; they get forgotten. - **A local model you have not measured.** Run the comparison on your own workload before buying hardware for it. On a machine with no GPU, small quantized models are useful for narrow, tolerant tasks and frustrating for everything else. ## The honest limit None of this makes a machine autonomous. The most reliable jobs are the ones with the narrowest definition of done. The least reliable are the ones that have to act like a person in someone else's interface, and those will need a human sometimes no matter how well you build. Build the narrow ones first. --- Licence: do what you like with this file. No attribution required.