A Claude usage limit should pause a Claude session, not strand your website files, deployment tools, research, and entire working environment. The practical answer is to own the workstation where the work happens, then choose which assistant gets access to it. That is what I mean by a sovereign AI workstation: a persistent server, portable projects, explicit data boundaries, and more than one way to get work done.
This is not a way around subscription limits. It is a way to keep your tools and files available when one service is unavailable. The reference build combines an OVHcloud dedicated server, Ubuntu, Remote Desktop Commander, Claude Code, OpenAI Codex, Kimi Code, and a private dashboard that can use local models or separately billed APIs.
Download the Sovereign AI Setup Guide in Markdown. It includes the software map, a private Docker Compose baseline, authentication boundaries, migration checks, backup requirements, and a decommission checklist. Read it before running commands on a server containing data.
What you own, and what you still depend on
The server holds the projects, Git checkouts, build tools, local model weights, and application data. A browser or laptop becomes a way to reach that environment rather than the only place it exists. You can close an office computer without making that computer the permanent home of every scheduled job.
But a rented server is not independence from every outside party. OVHcloud still provides the hardware and network. Remote Desktop Commander relays tool calls through its service. A prompt sent to Claude, OpenAI, Kimi, or Ideogram goes to that provider. Cloudflare can sit in the access path. Self-hosting the dashboard does not make every model request local. Desktop Commander's documentation describes both local execution and its remote relay, including temporary handling of tool arguments and results. Read its Remote MCP architecture.
The useful distinction is control, not an absolute privacy label. Decide which data stays on the machine, which snippets may leave, who can authorize changes, and how you recover without the original workstation.
Why the OVHcloud Eco line is worth a look
As checked on October 4, 2026, OVHcloud's US Eco catalog lists the SYS-5 from $125 per month, with a listed $125 installation fee. The configuration family starts with dual Intel Xeon Silver 4214R processors, 96 GB RAM, two 960 GB NVMe drives, and public bandwidth from 1 Gbps. Options, location, stock, taxes, and checkout terms can change the total. OVHcloud US Eco catalog.
That is a concrete example inside a $120 to $130 monthly server budget, not a promise that every 96 GB server costs that amount. It is also an Intel Xeon example, not an AMD EPYC configuration. EPYC is AMD's server family. Compare an available EPYC offer on its own CPU generation, memory, storage, network terms, and price rather than treating the names as interchangeable.
The appeal is predictable capacity. Enough RAM can hold several services, build processes, source trees, and a modest local language model without renting a separate instance for each job. Mirroring the two drives gives roughly one drive's usable capacity before filesystem overhead. It improves resilience to a drive failure; it does not protect against accidental deletion, compromised credentials, or a bad upgrade.
The advertised public network rate is not a guaranteed upload speed from your office. Your access connection, routing, provider policies, and destination affect throughput. Private-network bandwidth is a separate specification. Buy the machine for the workload you can demonstrate, not a single impressive number.
The software stack, without the mystery
There are three layers: the operating system and recovery tools, the assistants that work on files, and the model-serving applications. These components do different jobs.
| Component | What it does here | Why it earns its place |
|---|---|---|
| Ubuntu LTS | Runs the server and packages | A conventional Linux base with documented maintenance and recovery procedures. |
| OpenSSH, tmux, systemd | Provide secure access, reconnectable terminal sessions, and service startup | An SSH disconnect should not abandon an upgrade or permanently stop the workstation. |
| Node.js 24 and nvm | Run JavaScript command-line tools in a selected user runtime | Keep the tested runtime separate from whichever Node package the OS happens to provide. |
| Python and uv | Run scripts and isolated Python tools | Aider's interpreter should not disappear when the OS changes its default Python. |
| Remote Desktop Commander | Gives an authorized AI client terminal and filesystem tools on a selected device | The assistant can work on the actual server instead of only describing commands. |
| Claude Code, Codex, Kimi Code | Provide native coding-agent sessions | Different providers can use the same checked-out project through their supported clients. |
| Aider | Edits a Git working tree using a configured model | Useful for a second editing workflow, including a local-model route. |
| Docker Engine and Compose | Run and describe the application services | Make volumes, networks, limits, and restart behavior explicit and repeatable. |
| Ollama | Loads and serves local models | Provides a local inference option with no external model API on that route. |
| LiteLLM and PostgreSQL | Provide a model gateway, separate access keys, and persistent gateway state | Give applications limited credentials rather than distributing an administrator key. |
| Open WebUI | Provides the browser chat interface | One dashboard can use the models deliberately exposed by the gateway. |
| Git, GitHub CLI, Netlify CLI, Wrangler | Version, manage, build, and deploy website projects | Deployment should work from a replacement machine, not depend on one Windows profile. |
| Restic and rclone | Provide encrypted backup and controlled file transport | Recovery and synchronization are different jobs and need different checks. |
These are recommended roles, not a claim that every optional integration is installed or necessary. Docker installation, uv tool isolation, Ollama configuration, LiteLLM virtual keys, and Open WebUI setup describe the underlying mechanisms.
Optional tools should follow an actual need. Deno belongs where a project uses its runtime. Vercel CLI belongs where a site is hosted on Vercel. PM2 can supervise an existing Node application, but do not make PM2 and systemd both fight to restart the same process. Qdrant is useful when you have a retrieval workload that needs a vector database; a chat dashboard alone does not justify it. An automation engine such as n8n needs its own credentials, backups, and execution limits.
For design work, Figma integrations can expose approved design context, and Ideogram can generate editorial artwork. Those are external services, not local inference. Send a generic creative brief, not customer records, secrets, or a private server inventory. The image API produces expiring download links, so save the approved asset into your website's managed source tree. Ideogram image API.
Remote Desktop Commander is a bridge, not a blanket permission slip
The remote launcher runs on the machine you want to control. Launch it on AWS and the commands run on AWS. Launch it on the dedicated server and the commands run there. An authorized device and a connected AI-side integration are both required. Official remote setup.
I would begin with hostname, operating system, disk, and project-location checks. Only then grant the specific privileges the task requires. Keep destructive disk operations, secret extraction, and unrelated directories outside the normal workflow. Authentication prompts belong in the provider's browser or local terminal, not in a chat transcript.
One practical failure in this build was a dependency mismatch: the Ubuntu-provided Node 18 environment could not load a dependency chain used by the launcher. Moving the user runtime to Node 24 resolved it. The lesson is not to rewrite installed dependency files until the error disappears. Record a working version combination, pin it, and test the actual launcher. A package's declared minimum is not an end-to-end integration test.
The connector also needs a tested restart strategy. A user-level systemd service with the correct absolute Node path is more durable than a forgotten SSH window. Test a reconnect after a controlled restart. Keep a separate SSH and provider-console recovery path; never make the AI connection the only way into its own server.
Separate subscription sessions from API spending
There are two billing lanes, plus the local route. Mixing them is where an attractive architecture becomes an expensive misunderstanding.
Subscription lane: use the provider's supported native client and account login. Claude Code uses its supported Anthropic sign-in. Codex can use ChatGPT authentication. Kimi Code has its own account flow. Availability, allowances, and applicable terms depend on the provider and account. A working login is not an unlimited allowance, and a ChatGPT subscription is not a general OpenAI API key. Codex authentication, Claude Code authentication, Kimi Code login.
API lane: applications and approved automation call LiteLLM using a restricted gateway key. LiteLLM routes only the configured models to the intended external API or local server. Provider API credentials stay in protected configuration or the gateway's authenticated administration interface. Model identifiers and regional endpoints must be checked against the actual account; do not copy an old Moonshot endpoint into every deployment and assume it is universal.
Local route: a request goes through the gateway to Ollama on the private Docker network. There is no outside model API on that route. That does not mean the whole request is automatically private: a remote AI operator may already have received the prompt or tool output, and any logging or external retrieval you enable changes the data path.
A wrapper can unset conflicting API environment variables for a subscription session. It should never replace the user's entire Claude settings file to do so. Preserve permissions, hooks, MCP settings, and project configuration. Never put subscription OAuth tokens into LiteLLM as a substitute for provider-supported API authentication.
Start with a small local model and measure it
A 96 GB server is useful, but RAM is not a GPU. Model weights are only part of the memory budget. Context, concurrent requests, the operating system, databases, and browser processes also need room. Ollama documents how concurrency and context affect memory use. Ollama FAQ.
Start with an 8B-class model for controlled extraction, summaries, and small coding tasks. In this build, a brief direct Hermes 3 8B test completed in 7.2 seconds. A later authenticated request through LiteLLM returned the expected test response. Those are smoke tests on one machine, not a benchmark for long-context reasoning or a service-level guarantee.
A quantized 70B model may be a useful experiment on some 96 GB configurations, but fitting it into RAM does not make its response time acceptable. Measure your own documents, context lengths, concurrency, and accuracy. An AVX-512-capable processor also does not prove the chosen model runtime is using the implementation or thread placement you expect. A CPU quota is not CPU affinity.
The economical pattern is selective: extract and redact locally, retain provenance, and send only the material an external model needs. Verify that the reduction did not discard evidence important to the task. Smaller prompts are not automatically better prompts.
Moving the workstation is not moving every website's hosting
Netlify sites can keep running on Netlify while their source checkouts and deployment tools move from AWS to the dedicated server. Cloudflare Workers and Pages can likewise stay on their existing platforms. Do not point every production domain at the new server just because the workstation moved.
Inventory all accessible Netlify teams and Cloudflare accounts, then record each site's platform ID, domain, repository, source directory, build command, publish directory, functions, environment-variable names, and external data dependencies. Verify the authenticated account from the destination, not merely that the CLI prints a version. Netlify CLI, Wrangler commands.
One private repository per independent site is a sensible default. A motel portfolio can use one repository with a separate directory and deployment configuration for each property. Shared code belongs in an intentional shared package, not in an accidental copy of every other site's content. Existing Git history and functioning deployment links should be preserved until their replacements are tested.
DNS deserves its own map. Some domains may use Netlify's authoritative nameservers while others use Cloudflare. A domain registered or visible in one dashboard is not necessarily managed there. Check the actual nameserver delegation before changing a record. Do not disturb mail, domain-verification, or existing application records to make a new dashboard URL work.
A private dashboard comes before a public hostname
The conservative first version binds Open WebUI, LiteLLM, and Ollama to loopback on the server. PostgreSQL has no published host port. An SSH tunnel gives the operator access without making a new public login surface. Inside Docker, WebUI uses http://litellm:4000/v1, not the WebUI container's own 127.0.0.1 address.
Add Cloudflare Tunnel only after the local path works. Configure an explicit Access policy for the intended identities before exposing the hostname, and verify that an unauthorized visitor cannot reach the application. A healthy tunnel proves the connector reached Cloudflare; it does not prove that its route reaches the right local service or that Access is enforcing the intended rule. Tunnel setup, origin troubleshooting.
A browser desktop is optional. Its persistent profile is sensitive, and a browser debugging endpoint is effectively an administration interface. Do not publish that port. Copying Windows browser preferences or extension folders into Linux does not establish portable, working account sessions. Reauthenticate through supported flows and separately test extensions, MFA, and downloads.
The migration is finished when the old machine is no longer a dependency
A successful file upload is the beginning of acceptance testing. Preserve directory structure, including hidden project files and Git history. Hash the source files, transfer over an authenticated encrypted connection, and compare every selected destination file against the manifest. Record exclusions explicitly. A skipped configuration file is an unresolved dependency until someone classifies it.
Then check what a file listing misses: Windows Scheduled Tasks, services, Python environments, global npm packages, local databases, mapped cloud drives, refresh credentials, webhooks, build-time secrets, worker bindings, and background jobs. Windows-only scripts need a replacement or an explicitly retained Windows environment. An empty GitHub repository is not a backup of a website.
Build from the destination, deploy a preview, verify the application, and then test the public deployment by its content or identifier. An old homepage returning HTTP 200 does not prove a new deployment arrived. Keep a rollback path and a separate encrypted backup whose restore you have exercised. A copy stored only on the AWS VM is not independent recovery after that VM is destroyed.
Do not decommission AWS until these checks pass. The setup guide includes a sign-off checklist. This article documents an operating pattern and observed server tests; it does not certify that every application in a particular portfolio has completed migration.
Patch in stages, and keep honest evidence
The reference host progressed from Ubuntu 24.04 to Ubuntu 26.04.1 LTS, with a verified boot into the new kernel, healthy mirrored arrays, and all four application containers healthy. The provider-specific boot configuration was retained, and the updated EFI files were compared across both drives.
That result does not justify appending an unattended release upgrade to every installer. Back up first, test recovery access, review package removals, preserve a known working kernel, and validate after the reboot. Keep Python tool environments independent where appropriate. Pin container images by digest, then review and test updates deliberately rather than leaving them frozen forever. Ubuntu release-upgrade guidance, managed Python.
Firmware is another separate operation. fwupdmgr get-upgrades lists available updates; it does not install them. In this build, the live UEFI revocation-database version showed the requested update while the historical updater record retained an earlier detection failure. The honest report includes both. Do not erase warnings or repeatedly flash firmware just to manufacture a green status.
The real budget is larger than the server invoice
The $125 example pays for a server configuration, not all AI usage. Budget separately for subscriptions, external API calls, generated artwork, independent backup storage, any paid access or relay plans, domain services, and your maintenance time. Local inference avoids a per-request external model charge, but consumes the machine you rent.
That is still an attractive trade for a hands-on operator managing several projects. One persistent environment can hold the source, the build tools, the tests, and a local fallback model. The benefit is not an allegedly free 70B assistant. It is being able to replace a laptop, switch an AI provider, or rebuild a server without losing the operating knowledge of the business.
Get the downloadable setup guide and start with the smallest private system you can test. Add the next component only when its purpose, cost, credentials, and recovery procedure are clear.
Related reading
- Your Claude Code CLI is installed. Now what?
- AI data-flow security habits
- Plan for an AI vendor lockout
- APU versus GPU for on-premises AI
Conceptual hero illustration generated with Ideogram. Prices and product documentation checked October 4, 2026. This is a technical operating guide, not a guarantee of availability, security, model performance, or future pricing. Product names identify the tools discussed; no vendor affiliation is implied.