00Orientationready

The finished result, the build order, and the test that ends each phase.

Start here: what you are building and in what order

Guide TL;DR: Set up local recovery, work records, separate accounts, and access limits before installing an LLM. Then build one Linux server, reliable network and power paths, storage with a proved backup, and household services. Each configuration or dependency change comes as one tested script and one command; it reports success only after the real user test passes. After a failure, the agent usually recommends returning to the last working state, explains the risk, and asks before doing it. By the end, you can get back in, rebuild the Linux host and selected services, restore the data, and find the original reason for a change. The guide does not promise zero downtime, enterprise testing facilities, or fully automatic repair.

What you will have when you finish#

TL;DR: Finish with independent recovery paths, constrained agents, searchable original records, reproducible hosts, proved data restore, deliberately exposed services, and tested operations—without pretending every failure is automatic.

This guide assumes a very ordinary beginning: a Pi-hole on a Raspberry Pi, a consumer mesh Wi-Fi system, a recycled x86 desktop, several drives of uncertain suitability, and an always-on Mac mini. It does not assume Kubernetes, VLANs, a rack, business-grade hardware, a paid API budget, or Linux administration experience.

At the end you will have:

  • a Tier 0 local recovery path from a keyboarded device to the control Mac and server that survives WAN, DNS, cloud-account, and model-provider failure;
  • a separately tested Tier 1 remote path plus an honest map of everything it still depends on;
  • a standard agent account, narrow observer identities, target-side denials, and current low-friction Claude/Codex approval settings that let ordinary authorized work continue without discarding containment;
  • a human-controlled searchable archive of complete bring-up and future systems-administration transcripts, terminal evidence, and single-command action scripts, alongside a sanitized Git-friendly index that preserves original intent without treating the transcript as current truth;
  • a reproducible Linux host, deliberate update policy, direct-cable recovery, and one complete test-bearing human packet for any privilege the agent does not possess;
  • stable naming and identities, a storage layout chosen from the failure and growth model, deletion-resistant backup history, and a restored file/application/ACL proved before valuable migration;
  • deliberately exposed Compose services with versioned definitions, secrets outside Git/chat, user and denial tests, client-visible health, a prepared recovery path, and explicit household disruption;
  • sparse monitoring, incident state, risk-based drills, and recorded triggers for optional VLANs, directory services, GPUs, Kubernetes, remote power, or more hardware.

You will also know what remains manual: value judgments, secret entry, physical work, destructive authorization, and the decision to roll back after a failure. The agent should do the observation, comparison, drafting, validation, evidence capture, and routine execution it is actually authorized to perform. Every response ends by telling the human exactly what to do on which host/account/path—or explicitly says no action is needed and what event will wake the agent.

Tasks use two separate relative scales: operational work and expected LLM quota draw, each tiny, small, medium, large, or XL. A work profile also names human effort, agent effort, waiting, outage, affected tiers, and service/network/storage/household disruption. It gives a clock duration only when measured or sourced; otherwise it says unknown or labels a guess with its basis and uncertainty. For this interactive use, a predictable subscription allowance is the default; repeated quota stalls or forced model downgrades are a reason to discuss a larger plan, not to hide open-ended API billing inside automation.

The objective is not the largest service catalog. It is a home system that:

  • keeps ordinary household Internet and file access reliable;
  • can rebuild the recorded Linux host and selected services when their hardware changes or fails;
  • gives an LLM operator enough access to do useful work without giving it ownership of the house;
  • can continue a bounded job after the owner closes a laptop or phone;
  • lets many failures be diagnosed remotely, with software recovery over tested human paths when no physical work or new authority is required;
  • loses experiments before it loses family data, DNS, or the ability to get back in.

The shortest route is: control host, plain Linux host, recoverable network, identity and storage, then services.

An agent is not the foundation. It is a fast junior operator attached to a foundation. Build the attachment points and the safety boundary first.

The five service tiers#

TL;DR: Rank owner-visible outcomes from local control through optional experiments so higher-level services never become dependencies of recovery foundations.

These tiers are a decision tool, not an industry standard. A higher number may depend on a lower number; the reverse should not be true.

TierPromiseExamplesDefault recovery priority
0 — local controlFrom the only device with a keyboard, you can reach the control host and servers at home without Internet, public DNS, cloud SSO, or an LLM vendor.Direct IPs, local SSH, Screen Sharing, console, recovery accountsFirst
1 — household edgePeople at home can get out to the Internet, and the owner can get in securely from outside to perform Tier 0 work.ONT/modem, router, switch, primary Wi-Fi, DHCP, DNS, VPN, control hostFirst
2 — data safetyImportant data has integrity checks, snapshots where useful, and restorable copies outside the machine and outside the house.Storage pool, identity map, backup repository, restore testsBefore irreplaceable data
3 — household servicesDaily services can fail without taking Tiers 0–2 down.File shares, home automation, photo tools, internal appsAfter foundations
4 — media and experimentsLarge, noisy, optional, or fast-changing workloads are allowed to fail first.Plex, transcodes, download automation, test clustersLast

This ordering answers most arguments about what belongs on which machine, UPS outlet, account, or network rule. If Plex can exhaust the storage server that supplies family files, the tiers have been mixed. If the only VPN endpoint depends on the Kubernetes cluster you are trying to repair, the tiers have been inverted.

The build sequence#

TL;DR: Build in dependency order, let the agent batch safe work, and require each phase’s exit evidence before starting the next.

Follow the phases in order. Within a phase, the agent should batch safe work and ask for human help only at the listed gates. Desired state and evidence begin in Phase 1 and accumulate throughout; reproducibility is not a later cleanup phase.

Phase 0 — interview, inventory, and stop rules#

TL;DR: Record the initial control path, do-not-touch list, unknowns, and read-only authority without forcing decisions that belong to later phases.

Work profile: operational size small; quota small; human effort small; agent effort medium; wait/outage none; service/network/storage/household disruption none; clock duration unknown until the available records and targets are observed.

Ask only enough to establish the control surface safely: which Mac can stay on, which device has the keyboard, what router/Pi-hole/server hardware exists, which disks must not be touched, physical/noise/depth limits, and what read-only discovery is authorized. Defer the full storage, exposure, service, and power interviews until the corresponding phase, when the agent can prefill observable facts.

Exit test: there is a dated bootstrap inventory, an explicit Tier 0/1 sketch, a do-not-touch list, a list of unknowns, and a definition of which actions require a human.

Phase 1 — establish the control host#

TL;DR: Make the always-on Mac locally recoverable, durable, and able to continue agent work before relying on it for broader operations.

Work profile: operational size medium; quota medium; human effort medium; agent effort large; wait medium; planned outage limited to control-Mac reboot/cold-boot tests; household network and storage disruption none; duration unknown because hardware, FileVault, client permissions, transcript export, and current subscription state vary.

Dedicate the always-on Mac mini to long-running operator sessions. First prove local and independent human recovery, separate administrator/break-glass and standard accounts, FileVault/power behavior, a raw searchable session/evidence archive, model-safe workspace, backup, and access to every existing server. Then choose subscription/quota expectations, install one suitable agent client, test containment, and select the least interruptive current approval mode inside those handcuffs. Keep remote VPN entry as a Phase 3 completion item after the edge has been discovered.

Exit test: the keyboarded client can reach it by local IP with the WAN disconnected; a clean session can find an original bring-up record; the first agent can run there after the laptop closes without crossing tested denials; and both model-safe records and raw evidence survive outside the Mac.

Phase 2 — establish the first Linux host#

TL;DR: Install a minimal Linux system on an isolated OS disk, establish human and constrained agent access, and prove direct recovery before adding Docker or data.

Work profile: operational size large; quota medium; human effort large for hardware/installer boundaries; agent effort large; wait large for burn-in and reboots; outage limited to the new empty server; household network/storage disruption none; duration unknown until hardware health is observed.

Install a minimal supported LTS release on its own OS disk with all data disks disconnected. Establish human SSH, a separate unprivileged agent identity, automatic security patching, a source-constrained host firewall, persistent logs, and a direct-cable recovery path. Do not install Docker or build the data pool yet.

Exit test: the host cold-boots, survives a reboot, and is reachable from the Mac over both the normal LAN and a tested direct cable while router, Wi-Fi, DHCP, and DNS are unavailable.

Phase 3 — stabilize the household edge#

TL;DR: Map, stabilize, and battery-protect the complete network/control path, then prove local recovery, remote entry, and clean shutdown behavior.

Work profile: operational size XL; quota large; human effort medium; agent effort large; wait large for power drills; planned Tier 1 network and household disruption during attended maintenance; no valuable-storage mutation; duration unknown until the topology, UPS load, and management surfaces are discovered.

Map and battery-protect the complete path: utility power → UPS → ONT or modem → broadband router → switch → required mesh access point → DHCP/DNS → VPN endpoint/control host. Assign stable addresses to infrastructure. Keep a written direct-IP recovery card.

Exit test: ordinary clients resolve names and reach the Internet; remote entry works from a fresh cellular connection; a short power cut does not make the house immediately disappear; and a provisional extended-outage drill cleanly shuts down the empty server while networking and control remain available with measured battery margin.

Phase 4 — choose storage and identity before copying data#

TL;DR: Decide disk topology, growth, backup, and numeric identity from evidence, then prove access and restoration before migrating valuable data.

Work profile: operational size XL; quota XL; human effort large at disk identity/erasure/migration decisions; agent effort XL; wait XL for disk tests, backup, restore, and soak; planned storage and file-service disruption with unique-data risk; network disruption normally none; duration unknown until disk health, capacity, and data size are measured.

Inventory every disk by exact model, size, health, interface, and failure history. Decide what failures the pool must tolerate, how capacity will grow, and what will be backed up. Reserve a stable human and service UID/GID namespace even if centralized login comes later.

Exit test: the selected pool and datasets import after a cold reboot; the intended SMB allow/deny tests pass; an independent off-site backup exists; and a representative file plus its ACL has been restored from that backup into staging and opened. Valuable source data remains intact until the migration soak completes.

Phase 5 — add household services#

TL;DR: Introduce Docker, Compose, ingress, and selected services only after storage and recovery exist, retaining one consolidated human privilege boundary.

Work profile: operational size large; quota large; human effort small to medium; agent effort large; wait medium; planned per-service disruption but no lower-tier outage; storage changes limited to preapproved service paths; duration unknown per selected service.

Storage already established the first SMB share. Keep NFS just in time for a real server-to-server consumer. Install Docker Engine and Compose on one host, then run one disposable container before deploying a real Compose service. Put applications behind one ingress pattern, keep management surfaces on LAN/VPN, and separate application configuration from durable data. Until a privileged mechanism has a real implementation and independent security review, the agent stages and validates changes and the human executes one consolidated privileged command packet.

Exit test: each service has an owner, data path, health check, update method, rollback, backup classification, and explicit exposure.

Phase 6 — add media, automation, and orchestration when evidence demands them#

TL;DR: Add GPUs, download automation, and k3s only when measured workload or operational pain justifies their extra failure modes.

Work profile: operational size large to XL; quota large to XL; human effort medium; agent effort large to XL; wait depends on measured media and selected workloads; Tier 3/4 service and storage disruption is planned; Tier 0/1 disruption is prohibited; duration is unknown until scope is selected.

Measure Plex clients and media before buying a GPU. Prefer direct play. Treat download automation and its VPN as a distinct failure and security domain. Move from Compose to k3s only when multi-host scheduling, dependency races, restart behavior, service discovery, or deployment toil have become real recurring problems.

Exit test: for each selected optional workload, an attended stop-and-rebuild test leaves the Tier 0–2 checks passing.

Phase 7 — prove lights-out operation#

TL;DR: Validate sparse monitoring, loaded shutdown, bounded remote controls, and real recovery drills before claiming unattended operation.

Work profile: operational size large; quota medium to large; human effort medium for attended drills; agent effort large; wait large for representative observation; planned network/service/power disruption in declared windows; storage risk must be separately cleared; duration follows recorded drill evidence, not a generic calendar.

Add sparse external monitoring, revalidate graceful UPS shutdown with the real storage and accepted application load, add recovery notifications, and introduce only those narrowly mediated power controls whose enforcement actually exists. Run failure drills. An alert is not useful merely because it is technically true.

Exit test: the owner has successfully recovered from representative failures using the same remote paths, runbooks, and credentials that will exist during a real incident.

Decisions that become expensive to reverse#

TL;DR: Decide tiers, data value, storage shape, numeric identity, naming, exposure, and agent boundaries early while deferring complexity that lacks a demonstrated trigger.

Make these early:

  1. What is Tier 0 and Tier 1? This determines power, network, account, and placement dependencies.
  2. What data is irreplaceable? This determines backup before it determines RAID.
  3. What is the storage failure model and growth model? Filesystem and vdev/array shape follow from this.
  4. What is the UID/GID namespace? Stable numeric identities prevent years of NFS and permission repair.
  5. Which namespace will private DNS use? Use the reserved home.arpa namespace unless you have a specific reason to use split DNS under an owned domain.
  6. What is publicly reachable? Default to one intentional ingress for public applications; keep every management plane on LAN/VPN.
  7. Where is the security boundary for agents? Put it in infrastructure identities, ACLs, forced commands, network policy, and rollback—not prose. Until an enforcement mechanism exists and has been reviewed, retain the human privilege boundary.

These can usually wait:

  • Kubernetes;
  • VLANs that the router, Wi-Fi system, and switch cannot all enforce;
  • centralized LDAP or Active Directory;
  • a private certificate authority before a service actually requires TLS;
  • a rack-mount UPS;
  • a disk shelf;
  • a GPU;
  • high-cardinality metrics and elaborate dashboards.

Waiting is not neglect. It is preserving the option to make the decision using evidence.

The short first interview#

TL;DR: Ask only what is needed to establish safe control and discovery now, then defer phase-specific value judgments until the relevant evidence exists.

Use the following prompt with the operator agent before it touches anything:

Interview me only for the facts needed to establish a safe control surface: the always-on Mac, my keyboarded client, router/Pi-hole, candidate server and OS disk, every device or disk you must not touch, physical constraints, and the scope of read-only discovery. Ask in short batches. Separate preferences from facts you can discover. Start a decision register with: decision, why it matters, default, alternatives, evidence still needed, reversibility, and deadline. Be terse. After I authorize read-only discovery, continue through it without asking after every command. Never request or assume general sudo. Defer detailed storage, exposure, service, and power questions until their chapter.

The human answers value judgments; the agent discovers facts. At each later phase the agent first prefills its interview from the inventory, then asks only what remains: data value and loss tolerance before storage, users and privacy before sharing, audiences before exposure, and acceptable downtime before power design.

Who may do what on day one#

TL;DR: Begin with read-only agent access and one human-run privileged packet, adding fixed actions or a general control plane only after they truly exist and are independently reviewed.

Do not make a hypothetical privileged program part of the foundation. Use the first rung that actually exists:

  1. Read-only agent: the agent uses its own unprivileged SSH/API identities and prepares evidence and candidates.
  2. Consolidated human execution: the agent presents one complete, reviewed command packet for the approved phase; the human invokes it with privilege and returns the result.
  3. Fixed native action: a specific, root-owned service or helper exposes one enumerated operation with fixed targets, validation, limits, audit, and revocation. It is reviewed as privileged software.
  4. General typed broker or pull-based control plane: adopt only when a real implementation, threat model, update path, and independent security review exist.

The guide's default is rung 2 for privileged mutation and direct scoped application APIs where they already exist. Rung 3 is earned one operation at a time. Do not describe rung 4 in a handoff and then treat the description as authority.

One command means one complete change#

TL;DR: Give the owner one reviewed script and one exact command; the script installs or changes everything in scope, runs every relevant test, records the result, and reports success only after the user-visible outcome works.

This rule covers every configuration change and every package, server, or binary install or update. The script or fixed action includes the full set of packages, binaries, scripts, configuration files, includes, and generated files needed for that change. It checks its own Bash or Python where used, tests regular expressions against good and bad examples where used, runs the product's native configuration checks, seals the reviewed inputs, applies the change, checks live state, tests the real consumer and an expected denial, and saves the evidence. Exit code zero and the word SUCCESS are reserved for the end.

On failure, the same invocation stops, records what happened, and leaves a known state or completes its already-approved cleanup. It never tells the owner to paste another cleanup command. If returning to the previous state could lose data or worsen the failure, the agent explains that risk and asks before doing it. Physical work or a UI-only change arrives as one complete screen-by-screen set of instructions followed by one exact verifier command. The phase work order and change contract carry this rule into each job.

When to move to the next phase#

TL;DR: Continue through an authorized phase when recovery, lower-tier safety, durable records, and the human boundary are sound instead of pausing after every successful step.

Do not ask “is this perfect?” Ask:

  • Is the next phase reversible?
  • Does its rollback rely on the thing being changed?
  • Is the rollback prepared and tested, with a clear point at which the owner will be asked whether to use it?
  • Can it harm a lower tier?
  • Have we recorded the live facts rather than remembering them from chat?
  • Is the human being asked for a decision, a credential boundary, or a physical action that only a human can perform?

If the answers are sound, proceed through the authorized phase. Do not stop after every successful command to ask whether to run the next obvious one.