Snapshots
Updated
A snapshot is a rollback point for one classic VM, held inside the disk files it owns as a qcow2 internal snapshot. You can take it while the machine is running or while it is completely stopped. On a running machine you choose what kind of capture you want; the host does not choose it for you.
What a snapshot is
Section titled “What a snapshot is”One snapshot writes the same-named internal snapshot into the root disk and into every writable attached data volume, in a single transaction, and records the machine’s shape alongside them. Snapshots form a tree: each one records the snapshot it was taken from as its parent, and the machine records the node its disks are on right now, shown on the snapshots page as a Now row. Take a snapshot after a restore and it becomes a child of the node you restored rather than a new root.
It also records the machine’s shape: vCPU count, memory, growth ceilings, memory reclaim target, nested virtualization, host CPU affinity, and CPU topology. A restore puts the configuration back along with the data.
A snapshot is not a backup
Section titled “A snapshot is not a backup”Snapshots live inside the machine’s own disk files, in the same storage pool as the machine they protect. If that pool or the host is lost, the snapshots are lost with it. They exist to undo a change you are about to make, not to survive a disaster. Taking the snapshot from a running machine does not change this: the snapshot data lands in the same files either way.
For copies that leave the host, see Backup and restore.
Taking one
Section titled “Taking one”Open the machine’s snapshots and choose Take snapshot. The action is available in two states and no others: Running, and fully Stopped. In any other state the console says that a snapshot captures a running machine or a fully stopped one, and asks you to start or stop the machine first.
Name the snapshot up to 64 characters. Names may repeat; the destructive steps ask you to type the snapshot’s id, which is what actually identifies it.
A stopped machine needs no consistency choice: the level follows from how the machine came to rest. A guest that shut itself down gives a Clean shutdown snapshot; a machine that was hard-stopped, or that crashed, gives a Crash-consistent one. The dialog says which level this machine will produce before you take it.
A running machine asks for Capture consistency, and the dialog opens on Crash-consistent. Change it to Filesystem-quiesced if you need the quieter option; the machine stays online and keeps serving its workload while the snapshot is taken.
Either way a snapshot covers every writable disk of the machine as one set, never part of it. One transaction carries an action for each disk: if any of them fails the transaction is rejected, so a cancelled or failed attempt leaves nothing half-registered behind.
| Captured | Not captured |
|---|---|
| The root disk | Live memory |
| Every attached writable data volume | Running state and open connections |
| vCPU, memory, and their growth ceilings | PCI device assignment |
| Memory reclaim target, nested virtualization | Cached images |
| Host CPU affinity and CPU topology | Metrics and logs |
A read-only volume is recorded, not captured: no internal snapshot is written into one, because nothing ever writes to a read-only volume, and the snapshot still records that it was attached so a restore puts the mount relationship back. There is no option to exclude a writable volume. Rolling some disks back and not others would leave the machine with a split view of its own history, which is the failure a rollback point exists to prevent.
Choosing a consistency level
Section titled “Choosing a consistency level”Every snapshot carries one of three badges, and the badge reports what was proven during the capture, not what was requested.
| Badge | How it was earned | What you are getting |
|---|---|---|
| Clean shutdown | The guest shut itself down and the hypervisor exited cleanly | The cleanest state available: no writer, no ambiguity |
| Crash-consistent | Captured while the machine was running, or after it was hard-stopped or crashed | The disk state after a power cut. Most guests come back from it after a filesystem check |
| Quiesced | Captured while the machine was running, with the guest’s filesystems frozen and thawed around that instant | Filesystems that had a chance to settle first. Needs the guest agent |
The API names the same three levels: stopped_clean, crash_consistent, and
filesystem_quiesced. Of those, stopped_clean is never something a request
can ask for, because it is the result of stopping rather than an option.
Choose with the workload in mind:
- Quiesced is fail-closed. It pauses guest writes through the guest agent
while the snapshot is committed, and it needs the
qemu-guest-agentpackage, with fsfreeze support, inside the machine. If the agent is unreachable, if the freeze is refused, or if the thaw cannot be confirmed, the snapshot fails. It does not come back labelled crash-consistent, and it does not come back. - There is no downgrade path in either direction. A running machine cannot be asked for the Clean shutdown label, because that would claim a proof the capture could not have: the request is rejected, and the answer names the two levels a running machine can actually earn. The same rejection applies the other way round, so a knob that cannot be honoured is never quietly ignored.
- The freeze is brief by design. The guest is held still only long enough to commit the transaction on every disk at the same instant. The guest is writing again immediately afterwards.
- Filesystem-quiet is not application-quiet. A quiesce stops the filesystem from changing underneath the snapshot. It does not ask a database to flush its cache or commit what it has in flight, and no level in Virtainer Free claims that it does.
- The evidence is kept. A quiesced snapshot records how many guest filesystems the capture window froze and thawed, and exposes those counts on the snapshot view. The level is a conclusion with proof behind it.
What the guest feels
Section titled “What the guest feels”A running capture copies no data: it writes the snapshot into each disk’s qcow2 metadata in one transaction, while the machine keeps running. The guest is held still for that transaction and no longer, about 150 ms for every 8 GiB of data the disks already hold. The pause comes from the transaction itself, not from the quiesce: a crash-consistent capture freezes nothing and stalls the same way.
That is the trade to weigh when a workload has a latency limit. Guests whose services are sensitive to a stalled disk are the ones worth scheduling around. A stopped machine feels none of it: its capture runs against disks nothing is writing to.
Two interruptions are worth knowing about in advance:
- The storage daemon that serves a machine’s disks is a machine-level dependency. If it dies during a capture, the capture fails and the machine goes Failed, which is what happens to that machine whenever that process is lost.
- If Virtainer Free restarts while a capture is running, the operation is left Uncertain and keeps the machine’s changes and boot blocked until you recheck it. The recheck looks for the snapshot on the disks: if every disk carries it the snapshot is registered, and if none does the capture fails and you take a new one.
Restoring one
Section titled “Restoring one”The machine must be stopped to restore, whatever level the snapshot was taken at,
and it stays stopped afterwards. Each disk is reverted in place with qemu-img snapshot -a, which is idempotent: an interrupted operation recheck runs the same
step again rather than leaving half the disks in the future.
Volumes and NICs travel with it. A volume or NIC attached since the snapshot is detached and kept, a NIC going back to the NIC pool; one that was detached since is attached again. A volume or NIC the snapshot needs that is gone, or that now belongs to another machine, refuses the restore before any disk is touched.
If the snapshot’s configuration differs from the machine’s current settings, the confirmation lists exactly what will change before you agree to it.
PCI devices do not travel with a snapshot. A passed-through device belongs to the host, so restoring re-applies the machine’s configuration and then validates it against the host’s current CPU, firmware, and PCI state. Anything invalid is refused before a single disk is touched.
When Virtainer Free refuses
Section titled “When Virtainer Free refuses”Snapshots fail loudly and early rather than producing something that cannot be restored safely.
| Refused when | Why |
|---|---|
| The machine is neither running nor fully stopped | A capture needs one or the other, and the host will not guess which |
| A request for a running machine names no level | The level is the whole point of the request, so there is nothing to fall back on |
| A running machine is asked for Clean shutdown, or a stopped one is handed a level | The state and the level have to agree; neither is quietly resolved |
| A quiesced capture cannot prove its freeze and thaw | It fails instead of publishing a weaker snapshot under a stronger label |
| Another operation on the machine is still open | One operation at a time, and a capture holds the machine for as long as it runs |
| A running capture cannot identify the machine’s storage daemon | The identity has to hold for the whole capture, so the host refuses rather than guess |
| The pool is being created, changed, or recovered | A capture runs only while the pool is settled; it refuses rather than queue, and can be retried |
| The root disk is not managed by the host, or a disk is not qcow2 | A snapshot is written into the disks themselves, so every disk has to be one the host owns and keeps as qcow2 |
| A volume or NIC the snapshot needs is gone, or attached elsewhere | A restore puts the mount relationships back and will not take another machine’s resource or invent one |
| The saved configuration is invalid here | Checked against the host’s current CPU, firmware, and PCI state before any disk changes |
| The pool is below its 10% free-space floor | See below |
| It is an AppVM | AppVMs roll back by redeploying instead. See Logs, exec, and redeploy |
What it costs
Section titled “What it costs”A snapshot adds nothing to the pool when it is taken: it records the state the disks already have. What it comes to hold is the data the machine writes over while it is kept. There is no per-snapshot size to show, because qcow2 does not report one: the machine’s snapshots page shows a single Disk files figure, the root disk plus attached volumes, with the snapshot data inside it. The Storage page counts the same bytes inside Instance disks and Data volumes, and has no category of its own for snapshots.
Nothing is reserved up front. Admission checks one thing: that the pool is still above its 10% free-space floor. Under it, the capture is refused rather than allowed to fill the pool. A full pool stops writes for every machine on the host, not just this one, which is why the floor is a host-wide rule rather than a per-machine one.
Nothing is capped by count. There is no limit on how many snapshots a machine may keep, and no snapshot of yours is evicted to make room. A fixed count is the wrong control here: one snapshot of a large busy volume costs more than a dozen of small quiet ones. What limits you is the pool, and the Storage page warns you when free space in it goes under the floor. Cached base images are a separate matter: the host does delete those to claw space back, and the next machine that needs one re-downloads it.
When you run out of room you delete a snapshot yourself. Virtainer Free will not silently remove your oldest rollback point to make space. Deleting one removes its data from the disk files, frees what it held, and leaves the machine untouched; its child snapshots move up to become children of its parent. You can do it while the machine runs.
Removing a volume, a NIC, or a network that a snapshot depends on asks you to
confirm first: the refusal names the snapshots that depend on it, and the request
has to acknowledge them (acknowledge_snapshots=true) to go through. A later
restore of one of those snapshots fails before it touches a disk, and says what
is missing.
Reinstalling a machine discards its snapshots along with its disk, because there is nothing left for them to be rolled back onto.