# Sovereign AI Workstation Best Practices: What Three Months of Unattended Operation Taught Us

A dedicated AI workstation with no GPU, 94 GB of RAM and seven scheduled agents. The practices that survived contact with real failures, and the ones that did not.

Author: J.A. Watte
Published: October 8, 2026
Source: https://jwatte.com/blog/sovereign-ai-workstation-best-practices/

---

Our workstation has no graphics card. It has two Xeon Silver 4214 processors, 94 GB of memory, 877 GB of mirrored disk and an ASPEED chip that exists only so the motherboard can draw a console. If you came looking for a local model benchmark, this is the wrong article.

What it does have is continuity. Seven scheduled jobs, five long running services, forty repository checkouts and a browser that stays signed in. The practices below came out of running that for three months and watching which parts broke.

<p><strong>Build the same thing:</strong> <a href="/downloads/sovereign-ai-workstation-playbook.md" download>Download the full playbook as Markdown</a>. It is written for any Linux box, not for ours.</p>

## The sovereignty is in the state, not the model

The common version of this idea is that you buy a big graphics card, run a model locally, and stop paying for tokens. That is a real thing people do, and our earlier article compared the economics honestly. It is not what made the difference here.

What made the difference is that the work persists between sessions. A repository stays checked out. A browser stays logged in to a service that has no API. A memory directory accumulates what was learned. A schedule fires whether or not anyone is at a keyboard. None of that needs a graphics card, and all of it is lost the moment your agent runs in a container that evaporates when the conversation ends.

So the first practice is a reframing. Ask what you want to still exist tomorrow morning, and build for that. The model is rented either way.

## Put the long running pieces in user units, not in a terminal

Every durable piece here is a systemd user unit. Not a system unit, and not a terminal someone left open.

```bash
systemctl --user list-unit-files --state=enabled
```

That returns five on this box: a session spawner, a desktop control service, a private operations portal, a password protected browser terminal, and a tunnel. User units matter for a specific reason. They run as the same account that owns the files, the SSH keys and the browser profile, so nothing needs to be readable by root, and a mistake cannot take down the host.

The cost is that user units stop when the user's login session ends, unless you turn on lingering:

```bash
sudo loginctl enable-linger "$USER"
```

Forgetting that line is the single most common reason a setup like this works for a day and then quietly stops.

## A scheduled agent needs a watchdog, because the agent is what fails

Our session spawner has a watchdog timer that fires every five minutes. That looked like belt and braces when it was written. It is not.

An agent session is a long lived process holding a network connection to a model provider, a terminal multiplexer and sometimes a browser. It can die from a provider timeout, an out of memory event, a container restart or a bug in its own tooling. When it dies, nothing else notices, because the thing that was supposed to notice was the agent.

```bash
systemctl --user list-timers
```

The useful pattern is a cheap external timer whose only job is to check that the expensive thing is alive and restart it if not. Five minutes is a reasonable interval. The watchdog should be dumb enough that it cannot itself be the thing that breaks.

## Make every scheduled run write a status file, including the runs that do nothing

This is the practice that earned its place most clearly this week.

Our nightly search console job is driven by an agent in a real browser, because the button it has to press has no API behind it. On two consecutive nights it failed. Not silently: it wrote a status file.

```json
{
  "date": "2026-10-08",
  "state": "browser_missing",
  "accepted": 0,
  "refused": 0,
  "summary": "No Linux Chrome connected (only a Windows browser listed; switch_browser found none) - no tab opened, no page touched, zero spends"
}
```

Three things in that file are worth copying. It names a state rather than a success flag, so "nothing happened and that is wrong" is distinguishable from "nothing happened and that is fine". It records the quantities that matter for a job with a daily quota, so a failed night does not consume tomorrow's budget. And it was written by the agent itself before exiting, which means the runner script can detect a session that died without reporting:

```bash
if [ ! "$OPS/runs/status.json" -nt "$MARK" ]; then
  gsc-tool status failed 0 0 "session exited (rc=$rc) without reporting; see $LOG"
fi
```

A scheduled agent that exits zero and writes nothing is indistinguishable from one that did its job. Make the absence of a report a failure.

## Give an unattended agent a small, explicit tool list

Automatic approval is what makes unattended work possible. It is also the part that deserves the most care.

Our nightly job runs with a named list and nothing else:

```bash
claude -p --chrome --permission-mode default \
  --allowedTools 'mcp__claude-in-chrome__*' 'ToolSearch' 'Bash(gsc-tool:*)' \
                 'Read(//home/ubuntu/enclave/workspaces/gsc-operations/**)' \
  --max-turns 400 "$PROMPT"
```

It can drive a browser, call one purpose built command, and read files inside one directory. It cannot run arbitrary shell, cannot write outside that directory, and cannot reach the rest of the machine. The command it is allowed to run is a small wrapper we wrote, not a general tool, which means the list of things it can do is a file we can read in one sitting.

The general principle: an unattended agent should be allowed to do the specific job and nothing adjacent to it. Write the wrapper. It is an hour of work and it converts a broad grant into a narrow one.

## The browser is the weakest component, and it is the one you cannot replace

Three of our scheduled jobs touch a browser, because the services involved have no API for the action required. Search console indexing requests are the clearest case. The indexing API accepts only job postings and broadcast events, and the inspection API is read only, so there is no programmatic path to the button. A person or an agent has to click it.

That makes the browser session a dependency with none of the properties you want in a dependency. It holds credentials. It expires. It is tied to one machine's desktop. And when it is not there, the job cannot degrade gracefully, it can only stop.

Our failure this week was exactly that. The extension was connected to a browser on a different machine, the job correctly refused to use it, and two nights of work did not happen. The correct response is not to make the agent more permissive. It is to monitor the dependency directly:

```bash
systemctl --user --failed
```

Put the browser connection in your morning check, treat a disconnection as an outage rather than a nuisance, and accept that this part of the system has a human in it.

## Stop hand typing any value that appears in more than three places

This one came from a web project on the same machine, and it generalizes.

A cache busting version string lived in one constant and was hand copied into fifty four places across eighteen files. The constant's own comment said to bump it on every change. It was bumped. The fifty four copies were not, so the bust silently stopped working, and the comment describing that failure was sitting in the file where the failure was happening.

A comment is not a mechanism. If a value must agree across files, write the script that rewrites them and the check that fails when they disagree. Then the next person to bump it cannot get it wrong.

The same logic applies to the lists that accumulate in an operations setup: the repositories you back up, the services you check, the domains you renew. Derive them from the filesystem or from one declared source. Every hand maintained second copy drifts, and it drifts quietly.

## Test the populated path, because the empty one always passes

Our strictest check on a sixty nine gate test suite reported everything clean while four production pages returned server errors.

The reason is worth sitting with. The test rendered every page with no data store attached, so every list came back empty, so every loop over that list ran zero times. The code inside the loop was never executed. When a function's return shape changed, the one place still expecting the old shape was inside a loop that the test never entered.

```
# the test said OK, production said:
TypeError: Cannot read properties of undefined (reading 'slug')
```

If your checks run against an empty environment, which is normal for anything scheduled or containerized, feed them a small synthetic fixture with real shape. One record with data proves more than a thousand empty renders. Then break the code on purpose and confirm the check goes red, because a check you have never seen fail is a check you do not have.

## Back up on a timer, and verify on a different timer

Two units, not one:

```
enclave-optional-backup.timer         daily 09:30
enclave-optional-backup-verify.timer  weekly
```

A backup job that reports success is reporting that it finished, not that the archive can be restored. Separating the verification onto its own schedule means the failure mode "we have been writing corrupt archives for a month" has a bounded lifetime. Weekly is enough for most small operations. Never is what most people actually run.

## Keep the memory in files the next session can read

The agent that runs here keeps a directory of small Markdown notes, one fact each, with an index file that loads at the start of every session. It is not a database and it does not need to be.

The value is specific and easy to underrate. When a session three weeks from now asks "where does this site deploy from", the answer is a file rather than a re-derivation. When a practice turns out to be wrong, you delete the file. When the owner states a preference once, it survives.

The rule that keeps it useful is to write down what the repository cannot tell you. Code structure, commit history and configuration are already recoverable. What is not recoverable is why a decision was made, what was tried and rejected, and what the person you work for actually wants.

## What we would not build again

Three things on this machine are more than the job needed.

Forty repository checkouts on one box is more than anyone can hold in their head, and the only thing that makes it workable is that a scheduled check reports which ones are dirty. If we were starting again we would keep the checkouts and add the check on day one rather than month two.

A browser terminal and a portal are convenient and they are also two more services with their own authentication to get right. A plain SSH session does most of what they do. We kept them because mobile access matters here, which is a real requirement and not a general one.

And the second desktop container we built for testing has been worth less than the hour it takes to maintain. A throwaway container started per task, with no network, has been better in every case where we actually needed isolation.

## The honest limit

None of this makes the machine autonomous. Today it needed a person to reconnect a browser, and no amount of scheduling would have fixed that. Our most reliable jobs are the ones with the narrowest definition of done, and our least reliable are the ones that have to act like a human in someone else's interface.

Build the narrow ones first. They are the ones that still work when you are not watching.

## Fact-check notes and sources

- **Indexing API scope**: Google documents that the Indexing API can only be used to crawl pages with JobPosting or BroadcastEvent embedded in a VideoObject, which is why a news or local business page cannot be submitted through it. [Google Indexing API quickstart](https://developers.google.com/search/apis/indexing-api/v3/quickstart)
- **URL Inspection API is read only**: the API exposes a single method, urlInspection.index.inspect, and both of its OAuth scopes are read scopes, so there is no submit method behind the Request Indexing button. [URL Inspection API reference](https://developers.google.com/webmaster-tools/v1/urlInspection.index/inspect)
- **systemd lingering**: user units stop when the user's last session ends unless lingering is enabled for the account. [loginctl manual](https://www.freedesktop.org/software/systemd/man/latest/loginctl.html)
- **Hardware figures** in this article were read from the machine described: `lscpu` reports 2 sockets of Intel Xeon Silver 4214 at 2.20 GHz for 48 threads, `free -h` reports 94 GB total memory, and `df -h /` reports 877 GB on a mirrored array. There is no discrete graphics card; `lspci` shows only an ASPEED display controller.

## Related reading

- [The Sovereign Advantage: What a Dedicated AI Workstation Actually Changes](/blog/sovereign-advantage-dedicated-ai-workstation/) is the economics piece this one builds on.
- [Mobile AI Workspace Lessons](/blog/mobile-ai-workspace-lessons/) covers reaching the machine from a phone, which is where most of the browser trouble started.
- The full build playbook, written for any Linux host: [sovereign-ai-workstation-playbook.md](/downloads/sovereign-ai-workstation-playbook.md)


---

Canonical HTML: https://jwatte.com/blog/sovereign-ai-workstation-best-practices/
RSS: https://jwatte.com/feed.xml
JSON Feed: https://jwatte.com/feed.json
Hero image: https://jwatte.com/images/sovereign-ai-workstation-best-practices.webp
