9.7 KiB
Issue Backlog
Status snapshot
Last reviewed against code: 2026-06-04.
Several earlier backlog items have already landed in the codebase:
- Management persists accepted agent jobs in
job_history. - Management polls agent jobs and updates final status, error, finish time, duration, current step, and parsed backup snapshot ID.
- Operations page shows job history and audit log.
- Node-agent health checks cover required commands, ZFS pool state,
/dev/zvol, and Restic repository access. - Backup jobs verify the stored Restic file size and remove failed snapshots on verification errors.
Known caveat: agent-side jobs are persisted, but subprocesses cannot survive an agent restart. Active jobs are marked failed on startup; deeper cleanup recovery for partially changed host resources is still future work.
1. Persist final agent job status in management
Status: mostly done.
Management records when a backup or restore was accepted by an agent and now persists the final agent job result when the node-agent remains reachable long enough to be polled.
Goal
Persist reliable end-to-end job status in the management database.
Tasks
- Add polling for accepted agent jobs from management.
- Store final
successorfailedstatus injob_history. - Store duration, finished timestamp, error message, and current step.
- Store agent job logs summary.
- Store created Restic snapshot ID for successful backup jobs when available.
- Surface final status in the Operations page.
- Persist recent agent-side jobs.
- Mark active agent jobs as
failedafter an agent restart so management can poll a final state.
Acceptance Criteria
- A backup started through management eventually shows
successorfailedwhile the agent remains reachable. - A restore started through management eventually shows
successorfailedwhile the agent remains reachable. - Management restart does not lose already persisted history.
- Agent restart does not lose the active job record; active work is marked
failedbecause the subprocess cannot survive restart.
2. Expand node-agent health checks
Status: mostly done.
The health endpoint now provides deeper operational checks for backup readiness.
Goal
Make /api/health useful for diagnosing whether a node can actually run backup and restore operations.
Tasks
- Check that required commands exist:
incus,zfs,zpool,restic,udevadm,dd. - Check that configured ZFS pool exists.
- Check that
/dev/zvolis accessible. - Check Restic repository access.
- Check S3/Restic credentials by running a safe Restic command.
- Include ZFS pool capacity and free space.
- Return structured check names and messages.
- Surface detailed per-node health output in the management UI, not just a compact aggregate.
Acceptance Criteria
- [~] Management UI shows degraded node health with actionable check names. Detailed output exists through node health endpoints; UI can still be improved.
- A missing command, wrong pool, or wrong Restic credentials is visible in health output.
3. Harden backup verification
Status: partially done.
Backup jobs now verify the backed-up Restic file size after restic backup and attempt to remove failed snapshots. This reduces the risk of accepting a truncated backup, but the implementation still needs stronger stream failure handling and broader verification semantics.
Goal
Ensure a backup is marked success only when the expected source data was fully stored and verified.
Tasks
- Parse the created Restic snapshot ID after backup.
- Verify stored Restic file size against streamed source bytes or expected ZFS size.
- Remove failed Restic snapshots with
forgetandpruneon verification errors. - Make stream/pipeline failure handling explicit for all sources.
- Avoid marking success if source stream closes early but Restic exits successfully.
- Add tests with mocked command failures and short reads.
Acceptance Criteria
- Successful backup jobs include a verifiable Restic snapshot ID in management history.
- Size mismatches fail the job.
- Simulated source stream errors cannot produce a successful job.
- Verification behavior is covered by automated tests.
4. Add per-VM backup policy
Retention and schedule behavior is currently broad. Per-VM policies would make production usage more flexible.
Current state: schedules can be enabled per VM with interval and time of day. Retention is still global per agent.
Goal
Allow each VM to define its own backup policy.
Tasks
- Add management database table for VM backup policies.
- Support per-VM retention values: hourly, daily, weekly, monthly.
- Support per-VM schedule enablement, interval, and time.
- Add optional policy flag: backup only when VM is running.
- Update Scheduler UI to edit policies per node and VM.
- Send policy retention to the agent backup request or apply it in management scheduling.
Acceptance Criteria
- Two VMs on the same node can have different schedules.
- Two VMs on the same node can have different retention policies.
- Disabled policies do not trigger backups.
5. Improve restore safety workflow
Restore is destructive and should be guarded with a clearer preflight and confirmation flow.
Current state: VM restore requires typed confirmation, validates the snapshot against the VM, checks source/target size, writes the Restic dump to a staged ZFS volume, stops the VM, and swaps the staged volume into place with zfs rename. The previous production volume is kept as a rollback copy.
Goal
Reduce the risk of accidental or unsafe restores.
Tasks
- Add restore preflight endpoint.
- Validate node health before restore.
- Show selected snapshot metadata before restore.
- Show target VM status and disk size before restore.
- Require explicit typed confirmation including VM name.
- Optionally offer “create backup before restore” when the VM is accessible.
- Log restore intent in audit log before dispatch.
- Replace direct
ddto production ZVOL with a staged restore workflow. - Define a safe container restore workflow with
zfs receiveor disable container restore surfaces entirely.
Acceptance Criteria
- UI displays a restore plan before the final confirmation.
- Restore is blocked when required preflight checks fail.
- Audit log records restore attempts and results.
- Restore stream failures happen on the staged volume, not the production disk.
6. Add failure notifications
Operators need to know when scheduled backups or restores fail.
Goal
Send notifications for failed or degraded operations.
Tasks
- Add notification settings in management.
- Support webhook notifications first.
- Include node, VM, job type, error, and timestamp.
- Trigger notifications for failed scheduled backups.
- Trigger notifications for failed manual backups/restores.
- Add a “send test notification” action.
Acceptance Criteria
- A failed scheduled backup sends one notification.
- A test notification can be triggered from the UI.
- Notification failures are visible in Operations or audit logs.
7. Implement real snapshot file browsing
Current snapshot browsing shows Restic contents, which for block-level backups is usually only /vm.raw.
Goal
Allow browsing files inside a backed-up VM disk image.
Tasks
- Design a safe read-only raw image inspection workflow.
- Restore or mount raw image read-only in a temporary workspace.
- Detect partitions and filesystems.
- Browse directories through management UI.
- Allow downloading a single file.
- Ensure cleanup of mounts and temporary files.
Acceptance Criteria
- User can browse a Linux VM filesystem from a snapshot without restoring the VM.
- Mounted/temporary resources are cleaned up after use.
- The workflow refuses unsupported or unsafe disk images with a clear error.
8. Add agent and management version reporting
Multi-node setups need version visibility.
Goal
Show software version and compatibility state for management and each agent.
Tasks
- Add version field to management API.
- Add version field to agent health response.
- Show agent version in Nodes page.
- Flag unsupported or outdated agents.
- Document compatibility expectations.
Acceptance Criteria
- Nodes page displays agent version.
- Management can identify incompatible agents.
- Health output includes version information.
9. Containerized management and UI deployment
The node-agent must remain a host-level systemd service because it needs direct Incus, ZFS, /dev/zvol, Restic, and device access. Management and the frontend do not need those host privileges and are good candidates for container deployment.
Goal
Provide a production-ready Docker/Compose deployment for the management API and frontend while keeping node-agents installed as systemd services on Incus hosts.
Tasks
- Add a
managementcontainer image. - Add a frontend image that serves the Vite build through a small static server or Nginx.
- Provide
compose.yamlwith persistent SQLite volume for management. - Mount the agent CA certificate into the management container as read-only.
- Document required
CORS_ORIGINS,SESSION_COOKIE_SECURE,DATABASE_PATH, and reverse-proxy assumptions. - Decide whether the frontend calls the management API through the same origin reverse proxy or a separate API origin.
- Add health checks for both containers.
Acceptance Criteria
- Management API and frontend can be started with Compose without installing Node.js on the management host.
- Management database survives container recreation.
- Cookie login works behind HTTPS.
- Management can connect to HTTPS node-agents using the configured CA file.