# lab note
Don’t give the homelab agent root: build a CI/CD control plane instead
An AI agent can draft infrastructure changes, but it should not hold a root SSH key to the hypervisor. This pattern puts proposed changes through protected CI/CD, least-privilege credentials, telemetry, and tested rollback paths.
Sean · wrote this at the bench
Giving an AI agent direct root SSH access to Proxmox sounds efficient right up until it efficiently deletes the wrong thing. The CyberSec Tuesday presentation behind this topic describes exactly that dead end: direct root access was not workable, so the experiment moved toward GitLab CI/CD, Terraform, Ansible, Vault, Wazuh, Argo CD, and Prometheus/Grafana instead youtube.com. That is the useful lesson here. The agent is not the control plane. It is an untrusted change author.
This does not make an agent harmless. It makes its authority narrow, reviewable, and reversible. An agent can open a merge request, generate Terraform, update an Ansible role, and explain a failing pipeline. It should not be able to bypass the merge request, mint permanent credentials, or improvise directly on a hypervisor. GitLab’s own CI root-cause-analysis material is a decent reminder of the boundary: a tool may identify a likely fix, such as a missing Redis service, but diagnosis is not authorization about.gitlab.com.
Severity and who should care
Treat an agent with root access to your virtualization host as a high-severity design problem. The immediate blast radius is every VM, container, backup target, bridge, firewall rule, and secret reachable from that host. If the agent reads tickets, chat, web pages, or repository issues, prompt injection is another path to bad instructions. You do not need a dramatic model failure. A plausible-looking command in the wrong context is enough.
- Homelabbers are affected if an agent can SSH to Proxmox, run privileged Ansible, invoke
terraform apply, or access a broad API token. - Teams are affected if their CI runner can deploy from arbitrary branches, if deploy jobs use long-lived secrets, or if pipeline logs expose credentials.
- The risk drops substantially when the agent can only propose a change and a protected pipeline applies an approved, immutable artifact. It does not drop to zero. Nothing with
rm -rfin its vocabulary gets a gold star.
The safer pattern: agent proposes, pipeline disposes
Put the desired state in Git. Give the agent a branch and a merge-request workflow, not shell access to the host. Use protected branches so only approved identities can merge production-bound changes, and protected environments so deployment jobs have a separate authorization boundary. GitLab documents both controls in its guidance for protected branches and protected environments.
- The agent creates a branch, commits an IaC or configuration change, and opens a merge request with an explanation of scope and rollback.
- CI runs formatting, validation, policy checks, and a non-destructive plan. Failed checks stop there.
- A reviewer approves the merge request and a separately authorized deploy job. Do not let “merge” silently mean “run against the hypervisor.”
- The deploy job obtains narrowly scoped, short-lived credentials from Vault. Vault policies should permit only the required secret path or backend action, not a general-purpose root key developer.hashicorp.com.
- Terraform and Ansible apply only the reviewed repository revision. For Proxmox, prefer a scoped service/API identity or a constrained automation account over root SSH. Proxmox has role-based access controls for exactly this sort of separation pve.proxmox.com.
Make CI prove as much as it can before apply
A plan is not a guarantee, but it is a useful choke point. HashiCorp documents that terraform plan -out creates a saved plan suitable for a later apply, rather than recalculating changes at the moment of deployment developer.hashicorp.com. Store the plan only as a protected, short-retention artifact. It can contain sensitive values.
# Run in a protected CI job, with remote state locking enabled.
terraform init -lockfile=readonly
terraform fmt -check -recursive
terraform validate
terraform plan -out=tfplan
terraform show -json tfplan > tfplan.json
# After policy checks and an explicit approval in a separate deploy job:
terraform apply tfplanFor configuration management, run a dry run and show the diff before the real play. Ansible warns that check mode is only a simulation: tasks may not support it, and later tasks can lack facts created by earlier real changes docs.ansible.com. It is still far better than discovering a typo halfway through a privileged run.
ansible-playbook -i inventory/production site.yml --check --diff
ansible-lint site.yml
# Only in the approval-gated deployment job:
ansible-playbook -i inventory/production site.ymlDetect bypasses, drift, and noisy automation
Watch both the infrastructure and the automation account. Wazuh file-integrity monitoring can alert when selected files change documentation.wazuh.com. On Proxmox nodes, monitor at least /root/.ssh, automation users’ authorized_keys, sudo policy, and relevant /etc/pve paths. Tune exclusions carefully. Configuration churn is not an intrusion, but unexplained churn deserves a look.
# Run on a Proxmox node as an administrator.
sudo pveum user list
sudo pveum acl list
sudo find /root/.ssh /home -type f -name authorized_keys -print
sudo journalctl -u ssh --since "24 hours ago" | tail -n 200
# In the infrastructure repository: fail the pipeline if drift is found.
terraform plan -detailed-exitcodeFor terraform plan -detailed-exitcode, exit code 0 means no changes, 2 means changes are proposed, and 1 is an error. In a scheduled drift job, alert on 2; do not automatically apply it. Record the pipeline ID, commit SHA, actor, target environment, and outcome with every deployment. Prometheus alerting guidance recommends alerts that are actionable and routed to a real response path prometheus.io. A dashboard nobody checks is just expensive wallpaper.
Rollback is a design requirement, not a button-shaped wish
Argo CD is useful for Kubernetes workloads because it can sync a declared Git state and supports application rollback argo-cd.readthedocs.io. Keep that scope straight: Argo CD does not roll back a Proxmox VM change made by Terraform. Terraform and Ansible need their own recovery runbooks, backups, and tested inverse changes.
# For an Argo CD-managed application. Confirm the target history ID first.
argocd app history homelab
argocd app rollback homelab <history-id>After an Argo rollback, also revert or fix the Git commit. Otherwise automated sync may faithfully reapply the broken desired state. That is not Argo being difficult. It is Argo doing exactly what it was told.
Start with a boring boundary
My practical starting point would be modest: let an agent open merge requests for a non-production service; require CI validation and one human approval; issue short-lived deployment credentials; collect logs and metrics; and rehearse rollback once. Expand privileges only after you can answer three questions: who approved this change, exactly what ran, and how do we undo it? If any answer is “the agent probably knows,” the factory is still missing a guardrail.
$ subscribe --email
Automation patterns and lab updates when something ships or breaks. No hype, no spam.