Ansible role for updating a Promxox Cluster
| Filename | Latest commit message | Latest commit date |
|---|---|---|
|
All checks were successful
ci/woodpecker/push/linting Pipeline was successful
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> |
||
| .woodpecker | ||
| meta | ||
| playbooks | ||
| .ansible-lint | ||
| .editorconfig | ||
| .gitattributes | ||
| .gitignore | ||
| .markdownlint-cli2.jsonc | ||
| .sops.yaml | ||
| .yamllint | ||
| AGENTS.md | ||
| ansible.cfg | ||
| readme.md | ||
| renovate.json | ||
Ansible Playbook: Proxmox
Playbooks to maintain a Proxmox VE cluster: put a node into maintenance, update it, take it out of maintenance and
set or unset the Ceph maintenance flags. This repository only contains playbooks (no role tasks, defaults or
requirements.yml) and defines no variables. All playbooks target hosts: all with become: true, so limit them
to the node you want to work on.
Requirements
- Proxmox VE nodes (Debian based) reachable over SSH, with a user that can become root.
ha-managerandpvesh(Proxmox VE) on the nodes,cephfor the Ceph playbooks andfwupdmgrfor firmware updates.- The inventory hostname must start with the Proxmox node name, because
inventory_hostname_shortis used as node name.
Playbooks
| Playbook | Description |
|---|---|
playbooks/pve-enter-maintenance.yaml |
Enable HA maintenance mode on the node, then wait until it is empty |
playbooks/pve-update-node.yaml |
Run a dist-upgrade, install firmware updates and reboot the node |
playbooks/pve-exit-maintenance.yaml |
Disable HA maintenance mode on the node, then wait 60 seconds |
playbooks/ceph-enter-maintenance.yaml |
Set the Ceph OSD flags nodeep-scrub, noout, norebalance, noscrub |
playbooks/ceph-exit-maintenance.yaml |
Unset the same Ceph OSD flags |
Details
- Enter maintenance runs
ha-manager crm-command node-maintenance enable <node>and then pollspvesh get /cluster/resources(every 10 seconds, up to 60 retries) until no guests remain on the node. Resources matchingnode/pve,storage/pve,sdn/pve,template,network,ocportestare ignored. - Update runs
aptwithdistupgrade (keeping existing config files), thenfwupdmgr upgrade(failures are ignored) and finally reboots the node, waiting up to 600 seconds for it to return. - Exit maintenance runs
ha-manager crm-command node-maintenance disable <node>(3 retries) and waits 60 seconds so the cluster can settle.
Tags
None of the playbooks define tags.
Usage
Run the playbooks through Semaphore, one task template per playbook, limited to a single node. A typical node update runs the playbooks in this order:
ceph-enter-maintenance.yaml(only when the cluster uses Ceph)pve-enter-maintenance.yamlpve-update-node.yamlpve-exit-maintenance.yamlceph-exit-maintenance.yaml(only when the cluster uses Ceph)
Run it manually with, for example:
ansible-playbook -i inventory playbooks/pve-update-node.yaml --limit pve01