added install script for agent

changed backend to agent
This commit is contained in:
Philipp
2026-06-04 14:05:50 +02:00
parent 7f2785fb05
commit 0b059aec1d
669 changed files with 767 additions and 70582 deletions
+108 -27
View File
@@ -1,8 +1,24 @@
# Issue Backlog
## Status snapshot
Last reviewed against code: 2026-06-04.
Several earlier backlog items have already landed in the codebase:
- Management persists accepted agent jobs in `job_history`.
- Management polls agent jobs and updates final status, error, finish time, duration, current step, and parsed backup snapshot ID.
- Operations page shows job history and audit log.
- Node-agent health checks cover required commands, ZFS pool state, `/dev/zvol`, and Restic repository access.
- Backup jobs verify the stored Restic file size and remove failed snapshots on verification errors.
Known caveat: agent-side jobs are persisted, but subprocesses cannot survive an agent restart. Active jobs are marked `failed` on startup; deeper cleanup recovery for partially changed host resources is still future work.
## 1. Persist final agent job status in management
Management currently records when a backup or restore was accepted by an agent, but it does not persist the final agent job result.
Status: mostly done.
Management records when a backup or restore was accepted by an agent and now persists the final agent job result when the node-agent remains reachable long enough to be polled.
### Goal
@@ -10,21 +26,27 @@ Persist reliable end-to-end job status in the management database.
### Tasks
- Add polling for accepted agent jobs from management.
- Store final `success` or `failed` status in `job_history`.
- Store duration, finished timestamp, error message, and agent job logs summary.
- Store created Restic snapshot ID for successful backup jobs when available.
- Surface final status in the Operations page.
- [x] Add polling for accepted agent jobs from management.
- [x] Store final `success` or `failed` status in `job_history`.
- [x] Store duration, finished timestamp, error message, and current step.
- [ ] Store agent job logs summary.
- [x] Store created Restic snapshot ID for successful backup jobs when available.
- [x] Surface final status in the Operations page.
- [x] Persist recent agent-side jobs.
- [x] Mark active agent jobs as `failed` after an agent restart so management can poll a final state.
### Acceptance Criteria
- A backup started through management eventually shows `success` or `failed`.
- A restore started through management eventually shows `success` or `failed`.
- Agent restart or management restart does not lose already persisted history.
- [x] A backup started through management eventually shows `success` or `failed` while the agent remains reachable.
- [x] A restore started through management eventually shows `success` or `failed` while the agent remains reachable.
- [x] Management restart does not lose already persisted history.
- [x] Agent restart does not lose the active job record; active work is marked `failed` because the subprocess cannot survive restart.
## 2. Expand node-agent health checks
The current health endpoint should provide deeper operational checks for backup readiness.
Status: mostly done.
The health endpoint now provides deeper operational checks for backup readiness.
### Goal
@@ -32,23 +54,52 @@ Make `/api/health` useful for diagnosing whether a node can actually run backup
### Tasks
- Check that required commands exist: `incus`, `zfs`, `zpool`, `restic`, `udevadm`, `dd`.
- Check that configured ZFS pool exists.
- Check that `/dev/zvol` is accessible.
- Check Restic repository access.
- Check S3/Restic credentials by running a safe Restic command.
- Include ZFS pool capacity and free space.
- Return structured check names and messages.
- [x] Check that required commands exist: `incus`, `zfs`, `zpool`, `restic`, `udevadm`, `dd`.
- [x] Check that configured ZFS pool exists.
- [x] Check that `/dev/zvol` is accessible.
- [x] Check Restic repository access.
- [x] Check S3/Restic credentials by running a safe Restic command.
- [x] Include ZFS pool capacity and free space.
- [x] Return structured check names and messages.
- [ ] Surface detailed per-node health output in the management UI, not just a compact aggregate.
### Acceptance Criteria
- Management UI shows degraded node health with actionable check names.
- A missing command, wrong pool, or wrong Restic credentials is visible in health output.
- [~] Management UI shows degraded node health with actionable check names. Detailed output exists through node health endpoints; UI can still be improved.
- [x] A missing command, wrong pool, or wrong Restic credentials is visible in health output.
## 3. Add per-VM backup policy
## 3. Harden backup verification
Status: partially done.
Backup jobs now verify the backed-up Restic file size after `restic backup` and attempt to remove failed snapshots. This reduces the risk of accepting a truncated backup, but the implementation still needs stronger stream failure handling and broader verification semantics.
### Goal
Ensure a backup is marked `success` only when the expected source data was fully stored and verified.
### Tasks
- [x] Parse the created Restic snapshot ID after backup.
- [x] Verify stored Restic file size against streamed source bytes or expected ZFS size.
- [x] Remove failed Restic snapshots with `forget` and `prune` on verification errors.
- [ ] Make stream/pipeline failure handling explicit for all sources.
- [ ] Avoid marking success if source stream closes early but Restic exits successfully.
- [ ] Add tests with mocked command failures and short reads.
### Acceptance Criteria
- [x] Successful backup jobs include a verifiable Restic snapshot ID in management history.
- [x] Size mismatches fail the job.
- [ ] Simulated source stream errors cannot produce a successful job.
- [ ] Verification behavior is covered by automated tests.
## 4. Add per-VM backup policy
Retention and schedule behavior is currently broad. Per-VM policies would make production usage more flexible.
Current state: schedules can be enabled per VM with interval and time of day. Retention is still global per agent.
### Goal
Allow each VM to define its own backup policy.
@@ -57,7 +108,7 @@ Allow each VM to define its own backup policy.
- Add management database table for VM backup policies.
- Support per-VM retention values: hourly, daily, weekly, monthly.
- Support per-VM schedule enablement, interval, and time.
- [x] Support per-VM schedule enablement, interval, and time.
- Add optional policy flag: backup only when VM is running.
- Update Scheduler UI to edit policies per node and VM.
- Send policy retention to the agent backup request or apply it in management scheduling.
@@ -68,10 +119,12 @@ Allow each VM to define its own backup policy.
- Two VMs on the same node can have different retention policies.
- Disabled policies do not trigger backups.
## 4. Improve restore safety workflow
## 5. Improve restore safety workflow
Restore is destructive and should be guarded with a clearer preflight and confirmation flow.
Current state: VM restore requires typed confirmation, validates the snapshot against the VM, checks source/target size, writes the Restic dump to a staged ZFS volume, stops the VM, and swaps the staged volume into place with `zfs rename`. The previous production volume is kept as a rollback copy.
### Goal
Reduce the risk of accidental or unsafe restores.
@@ -82,17 +135,20 @@ Reduce the risk of accidental or unsafe restores.
- Validate node health before restore.
- Show selected snapshot metadata before restore.
- Show target VM status and disk size before restore.
- Require explicit typed confirmation including VM name.
- [x] Require explicit typed confirmation including VM name.
- Optionally offer “create backup before restore” when the VM is accessible.
- Log restore intent in audit log before dispatch.
- [x] Log restore intent in audit log before dispatch.
- [x] Replace direct `dd` to production ZVOL with a staged restore workflow.
- [ ] Define a safe container restore workflow with `zfs receive` or disable container restore surfaces entirely.
### Acceptance Criteria
- UI displays a restore plan before the final confirmation.
- Restore is blocked when required preflight checks fail.
- Audit log records restore attempts and results.
- [x] Restore stream failures happen on the staged volume, not the production disk.
## 5. Add failure notifications
## 6. Add failure notifications
Operators need to know when scheduled backups or restores fail.
@@ -115,7 +171,7 @@ Send notifications for failed or degraded operations.
- A test notification can be triggered from the UI.
- Notification failures are visible in Operations or audit logs.
## 6. Implement real snapshot file browsing
## 7. Implement real snapshot file browsing
Current snapshot browsing shows Restic contents, which for block-level backups is usually only `/vm.raw`.
@@ -138,7 +194,7 @@ Allow browsing files inside a backed-up VM disk image.
- Mounted/temporary resources are cleaned up after use.
- The workflow refuses unsupported or unsafe disk images with a clear error.
## 7. Add agent and management version reporting
## 8. Add agent and management version reporting
Multi-node setups need version visibility.
@@ -159,3 +215,28 @@ Show software version and compatibility state for management and each agent.
- Nodes page displays agent version.
- Management can identify incompatible agents.
- Health output includes version information.
## 9. Containerized management and UI deployment
The node-agent must remain a host-level systemd service because it needs direct Incus, ZFS, `/dev/zvol`, Restic, and device access. Management and the frontend do not need those host privileges and are good candidates for container deployment.
### Goal
Provide a production-ready Docker/Compose deployment for the management API and frontend while keeping node-agents installed as systemd services on Incus hosts.
### Tasks
- Add a `management` container image.
- Add a frontend image that serves the Vite build through a small static server or Nginx.
- Provide `compose.yaml` with persistent SQLite volume for management.
- Mount the agent CA certificate into the management container as read-only.
- Document required `CORS_ORIGINS`, `SESSION_COOKIE_SECURE`, `DATABASE_PATH`, and reverse-proxy assumptions.
- Decide whether the frontend calls the management API through the same origin reverse proxy or a separate API origin.
- Add health checks for both containers.
### Acceptance Criteria
- Management API and frontend can be started with Compose without installing Node.js on the management host.
- Management database survives container recreation.
- Cookie login works behind HTTPS.
- Management can connect to HTTPS node-agents using the configured CA file.