Corv: my Hermes agent.

There is a container on my Proxmox cluster whose entire job is to be asked things. It holds no data, serves no traffic, and backs up like any other guest. What it does have is credentials for nearly everything else in the lab, a Discord channel, and standing permission to act without checking with me first.

This post is a tour of it.

The box

It is LXC 209, hostname hermes, a Debian 13 unprivileged container with nesting=1, 4 cores, 8 GiB RAM, and a 30 GB disk, sitting on corvus, the primary node of the two-node corvidae cluster. Its existence and shape are declared in OpenTofu’s hosts.yaml like every other guest, so it is reproducible from a reviewed diff rather than a memory of what I typed.

Inside it runs Hermes Agent v0.21.5 as a systemd gateway, hermes-gateway.service. The gateway is the always-on half: it holds the Discord and Telegram connections, runs the cron scheduler, and spawns an agent conversation for each incoming message. I reach it from a Discord channel that opens a thread per top-level message, so each topic starts with a fresh context instead of inheriting the whole accumulated history. Telegram works too, but Discord is the better interface for exactly that reason.

The design decision that shapes everything else:

approvals:
  mode: 'off'          # manual | smart | off
  cron_mode: approve   # deny | approve

Approvals are off. An agent that asks permission before every command is a chatbot with extra steps; the point of this box is one that acts and then reports what it did. (One YAML detail worth passing on: bare off is a YAML 1.1 boolean, parses as False, and fails validation silently, leaving the prompts on. The quotes are load-bearing. hermes config get approvals.mode is the honest check, not reading the file back.) There is still a hardline blocklist underneath for the genuinely unforgivable things, rm -rf / and fork bombs and dd to raw devices, which no approval setting can override.

The model chain

The primary model is gpt-6-luna, reached through the ChatGPT subscription’s Codex path. Login there is a device-code flow, so it cannot be scripted: hermes auth add openai-codex --type oauth --no-browser prints a URL and a code, and that is the whole ceremony. Nothing on the vendor side documents which ChatGPT tiers are eligible, so Plus working here is a fact discovered by trying it rather than promised anywhere.

Behind the primary sits a four-deep fallback chain, and the interesting property is that it spans three separate billing paths:

PositionModelProviderBilling
Primarygpt-6-lunaopenai-codexChatGPT subscription
Fallback 1kimi-k3opencode-goZen Go subscription
Fallback 2glm-5.3opencode-goZen Go subscription
Fallback 3qwen3.8-maxopencode-goZen Go subscription
Fallback 4deepseek-v4-proopenrouterpay per token

The chain is walked on rate limits, 5xx errors, and connection failures, and resets to the primary on each new message. The last position is the safety valve: deepseek-v4-pro through OpenRouter, paid per token, no subscription involved. It costs money when it runs, which is exactly why it works as the final fallback. The three subscriptions above it can all be quota-capped or down at once, and a metered API key with credit still answers.

Switching models is done through a small script called switch-model, because the built-in hermes model is an interactive picker that wants a TTY, which is painful over SSH. It sets only the provider and default fields, it can probe candidates with real calls before trusting them, and it refuses to set models known to fail on this account. That last part matters: one model on the Zen Go subscription lists fine in /models and then rejects every real request with “This Go model requires Global regions”, a workspace privacy setting. A listing is not an entitlement check. Two models are recorded as known-broken and the script will not select them.

Coding is delegated

Hermes orchestrates; it does not grind through big code changes itself. Anything substantial gets handed to OpenCode, which runs on the same box under the opencode.ai Zen Go subscription, on kimi-k3.

The reason is quota separation. Both subscriptions are flat rate, so this saves nothing, but coding is the token-heavy half of any real task: reading whole files, iterating on diffs, re-reading after the fix. Keeping it on a different plan means a long refactoring session does not eat the same weekly limit as the conversations. Small, obvious edits stay inline, because a handoff round trip costs more than the edit and the delegate only receives a prompt, not the conversation.

The split looks like this:

%%{init: {"flowchart": {"useMaxWidth": false, "nodeSpacing": 16, "rankSpacing": 32, "diagramPadding": 4}}}%%
flowchart TB
    chat["Discord thread\nor Telegram"] --> gw["hermes-gateway\nLXC 209"]
    gw --> orch["Hermes\norchestration, memory, cron"]
    orch -->|"primary"| codex["gpt-6-luna\nChatGPT subscription"]
    orch -->|"fallback"| zen["kimi-k3 / glm-5.3 / qwen3.8-max\nZen Go subscription"]
    zen -->|"all three subscriptions down"| orr["deepseek-v4-pro\nOpenRouter, pay per token"]
    orch -->|"substantial code"| oc["OpenCode\nkimi-k3 on Zen Go"]

What it knows: wiki, skills, memory

The agent’s knowledge of the lab comes from three layers, and they are all in git.

The first is the homelab wiki, cloned to ~/projects/homelab-wiki and pulled to tip every hour by a cron job. It is the authoritative source for IPs, ports, procedures, and why things are the way they are. The agent’s homelab skill opens with a mandatory instruction: before changing any infrastructure, read the wiki first. The fleet table, the DNS layout, the Traefik route procedure, all of it is a file read away, and the skill’s stance is that guessing an IP wastes the user’s time.

The second layer is skills: markdown instruction packages in ~/.hermes/skills/, one per kind of work. Beyond the homelab skill there are around twenty others, covering things like browser automation, research workflows, document generation, and writing style. They load on demand: the agent sees each skill’s one-line description and pulls the full text only when a task matches. Some carry scripts and reference files, so a skill can amount to a small toolkit rather than a note.

The third is memory, a pair of files (MEMORY.md, USER.md) that persist facts across sessions: the fleet layout, where repos live, which tools are installed, how the user likes things done. The distinction from skills is that memory applies to every conversation regardless of task, so it stays small on purpose.

Skills and memory both live in a private git repo, hermes-config, symlinked into place. Editing a file there changes what the running agent uses, and the agent’s own memory writes show up as git changes. A second cron job commits and pushes that repo to main every half hour, so the agent’s brain is versioned and backed up like everything else in the lab.

MCP servers

For capabilities outside its own toolset, the gateway runs two MCP servers:

ServerWhat it adds
Playwrighta real browser: navigation, clicks, screenshots, DOM inspection
Context7current documentation for libraries and APIs, on demand

The browser one changes what the agent can honestly report. Anything from “does the new Glance dashboard actually render” to “log into the Technitium console and check the DHCP scope” becomes a question it can answer from pixels rather than from an API guess. Context7 exists because model training data goes stale: before writing code against a library, the agent can pull the current docs instead of trusting what it remembers.

Both are listed as enabled with all tools exposed, and they are just two entries in a config table. Adding a third MCP server is a one-line change and a gateway restart.

A self-managed toolchain

Since 0.21, Hermes provisions its own runtime rather than borrowing the system Python. On this box that means:

ToolVersionNotes
Python3.14.7Hermes-managed, in ~/.hermes/tools/; the system Python is 3.11
Node / npm26.7.0 / 12.0.2for the MCP servers
Chromium1208the Playwright backend
ffmpeg9.0.1audio and media skills
ripgrep15.2.0fast content search

hermes pm doctor reports what is missing and installs it, and hermes pm install <tool> targets one. On top of these, the box carries the homelab CLI set, jq, fd, ansible-core, OpenTofu, and fj for the Forgejo API, so the agent can validate a playbook or read a plan without leaving home.

What it can reach

The agent’s outbound access is deliberately narrow in shape.

SSH goes to the two Proxmox nodes with sudo rights for pct list/status/exec/pull/push and read-only pvesh get. That is the interesting grant: LXC guests do not accept SSH from peer boxes at all, by design, so the route into any of the fourteen containers is ssh corvus 'sudo pct exec 208 -- ...' through the node. One key, one grant, every guest. Sylvanus, the Docker VM, accepts SSH from this box for host-level inspection, and Forgejo is reachable over SSH for git.

The API side has three tokens: a Proxmox API token for inventory and status, a Forgejo token for issues, repos, and CI, and Komodo credentials for the Docker stacks. With those, the agent can move from an alert to a runbook, to the live service response, to the config that owns it, and then propose the change as a pull request, which is how changes are supposed to travel in this lab anyway. It does not merge its own PRs; CI runs the checks and I read the diff.

Cron

The scheduler runs three jobs, all as plain scripts with their stdout delivered directly:

  • Every 60 minutes, pull the homelab wiki clone to tip. This is what keeps the first knowledge layer honest.
  • Every 30 minutes, commit and push hermes-config if anything is dirty. This is what keeps the second and third layers backed up.
  • A paused event watcher, kept around because its delivery target needs re-pointing.

Failures surface in hermes cron list and the gateway notifies, since the useful notifications are the ones that tell me something went wrong; success stays silent.

Why this shape

The guiding idea is that the agent is infrastructure, so it gets the same treatment as infrastructure: declared in OpenTofu, configured by Ansible, its knowledge in git, its work shipped as reviewed pull requests. The box is boring on purpose. The interesting part is the boundary around it: it can act freely inside the lab’s own change process, the credentials define what “inside” means, and the process catches the rest.

Discussion