Volumes¶
A volume is a VOLUME in your image that contemper backs with a real,
persistent disk: sized, named, formatted on first boot, and mounted
through /etc/fstab like any other filesystem. Every target attaches a
volume as a blank block device — even Incus, which also offers
filesystem volumes — because a block device is what every provider can
offer, and the guest prepares it the same way regardless of which
provider handed it over.
Declaring a volume¶
VOLUME /data
LABEL io.contemper.volume./data.size="10GiB"
VOLUME is what makes a path a volume at all; the label gives it a
size. Without a size, convert still records the volume — unsized — but
deploy refuses to create it until one is supplied, either by adding the
label or by passing --volume /data=10GiB at deploy time (which also
overrides a label's size, if you want a different size for one
deployment). contemper never guesses a size. The path must be absolute
and may not contain control characters such as a newline or tab.
A size declared by a label, and so recorded in a bundle, is capped at
1 TiB, because both come from the image. A larger volume is sized with
--volume at deploy time, which has no cap.
Sizes accept the same units everywhere in contemper: 10GiB, 512MiB,
or a bare byte count.
Names¶
A volume's identity is its path. contemper derives a name from it
automatically — the leading / dropped, the rest lowercased, / turned
into -, and anything still outside [a-z0-9-] dropped, shortened with
a hash suffix if that's still longer than 16 characters (the ext4 label
limit). /data becomes data; /var/lib/my-app becomes
var-lib-my-app.
The name is what ends up on disk — the ext4 label, and (where the target lets contemper choose) the disk's serial — but the path is what contemper and the volume helper reason about. Give a volume an explicit name only when you need one:
LABEL io.contemper.volume./data.name="appdata"
The usual reason is keeping the same volume attached across an image
change that moves the path (/data renamed to /var/lib/app, say): the
derived name would change and contemper would treat it as a different
volume, but an explicit name carries over unchanged. A shorter, friendlier
label is a fine reason too. Two volumes that resolve to the same name,
explicit or derived, fail the conversion, naming both paths.
fstab¶
contemper appends one line per volume to /etc/fstab:
LABEL=data /data ext4 defaults,nofail 0 2
A space, tab or backslash in the mount point is written as an octal
escape (\040, \011, \134), as fstab(5) requires.
nofail means a missing or not-yet-formatted volume doesn't block boot;
that only actually happens if you skip the volume helper (below) on an
image with no other way to prepare the disk. Opt out of the fstab lines
entirely — if your init system mounts volumes its own way — with the
image label io.contemper.fstab="false" or convert --no-fstab.
Bootloader images get one more appended line, for the ESP, which the
same opt-outs cover; see Mounting the ESP in the
guest.
On systemd, nofail by itself would do more than that: it also drops
the generated mount unit's implicit ordering before local-fs.target
(systemd.mount(5)), so boot could reach local-fs.target — and every
ordinary service, which is ordered after it by default — before the
volume is actually mounted. The systemd variant of the volume helper
(below) restores that ordering with a small boot-time generator, without
giving up what nofail is for: a missing disk still only delays boot by
its device timeout, never fails it.
How the volume helper works¶
Formatting a blank disk has to happen somewhere, and only a local hypervisor lets contemper create disks itself — every real provider hands the VM a blank block device. So it always happens in the guest, on every boot, by one rule:
- an ext4 filesystem already carrying the expected label is reused, unmounted and untouched — this is what lets a new image version keep running on the data an earlier version left behind;
- a disk whose first and last MiB are all zero is blank, and gets formatted, labelled, and — see Seeding a volume from the image below — seeded from the image's own content at that path, if it has any;
- anything else is left alone and logged as a mismatch, so the failure mode is always an unformatted, unmounted volume — never lost data.
contemper ships this rule as a first-boot helper, merged into your image
automatically. Concretely, it's a support image like any other (see
Support images), published at
ghcr.io/contemper-project/volumes-support, with the actual logic in one
POSIX shell script and two variants — one per init system — that each
carry only the small integration needed to run it at boot:
io.contemper.branch.init-system.openrc.requires.files=/sbin/openrc
io.contemper.branch.init-system.openrc.image=ghcr.io/contemper-project/volumes-support-init-system-openrc@sha256:...
io.contemper.branch.init-system.systemd.requires.files=/usr/lib/systemd/systemd
io.contemper.branch.init-system.systemd.image=ghcr.io/contemper-project/volumes-support-init-system-systemd@sha256:...
The variant images are named by digest, not a floating tag, so a given
volumes-support:v1 always names an exact, reproducible pair of variant
images.
Each published image (the base and both variants) carries a signed build provenance attestation, checked with:
$ gh attestation verify oci://ghcr.io/contemper-project/volumes-support:v1 --repo contemper-project/contemper
When it merges. Only when your image declares at least one volume
(no volumes, no helper — a plain qemu build never gains this layer). If
you also pass --support, that image and its own resolved variants merge
first; the volume helper merges after, so a user's own support-image
customizations are never shadowed by it. --volume-helper <ref>
substitutes a different image (hack/e2e-volumes.sh points this at a
locally built one); --no-volume-helper skips it entirely, in which case
you are responsible for getting volumes formatted and mounted — the
fstab lines and /etc/contemper/volumes are still written regardless,
since they describe the image's declared volumes, not who acts on them.
Which variant wins. Exactly like any support image's branches: the
openrc variant wins if /sbin/openrc exists in your merged image
before any support layers are merged, systemd wins if
/usr/lib/systemd/systemd does, and neither existing fails the
conversion — naming the branch and pointing at --no-volume-helper — so
a build never silently produces unprepared volumes. This also means an
image needs mkfs.ext4 (from e2fsprogs — the same package that provides
the boot-time fsck.ext4 most images already need) and dd/od
(from busybox or coreutils) available for the helper to do its job;
convert checks and fails naming whatever's missing. Unlike fsck.ext4,
e2label/tune2fs/dumpe2fs live in a separate e2fsprogs-extra
package on Alpine, so the helper deliberately doesn't use them: it reads
the label it needs straight out of the ext2/3/4 superblock with dd and
od instead (see the script itself,
support/volumes-support/base/usr/lib/contemper/format-volumes, for how).
/sbin/openrc matches equally whether your image keeps /sbin as a
real directory or, as on a merged-/usr distro (Fedora, Arch, current
Debian/Ubuntu), a symlink to /usr/sbin: predicate resolution follows a
symlink anywhere in the path, not only at the very end.
What lands in the guest. The base image contributes the script
itself, at /usr/lib/contemper/format-volumes. The winning variant adds
just its integration: OpenRC gets /etc/init.d/contemper-volumes
(depend() { before fsck localmount }) and its enablement symlink
/etc/runlevels/boot/contemper-volumes; systemd gets /etc/systemd/system/contemper-volumes.service
(Before=local-fs-pre.target) and its enablement symlink under
local-fs-pre.target.wants/ — pre-created in the image, since a support
image only ever drops in files and never runs systemctl enable or
anything else. Either way it runs before local filesystems are checked
and mounted, so the
label is right by the time /etc/fstab's LABEL= lines are resolved.
The systemd variant also adds a small generator,
/etc/systemd/system-generators/contemper-volumes
(systemd.generator(7)): at boot, for each line in
/etc/contemper/volumes it computes the fstab-generated mount unit's
name (systemd-escape --path --suffix=mount <mountpoint>) and drops in
a Before=local-fs.target override for it — the ordering nofail (see
fstab above) otherwise removes. It changes nothing else about
the mount, so a genuinely missing disk still only delays boot; it never
fails or logs anything if /etc/contemper/volumes or systemd-escape
isn't there, or if a directory can't be written, so it can never itself
break boot.
The systemd unit adds DefaultDependencies=no and
After=systemd-udev-trigger.service: by the time that trigger's own oneshot
has run, udev is already active and every disk already attached — root and
volumes alike, since every target hands them all over upfront rather than
hot-plugging — has had its boot-time ("coldplug") event replayed for udev to
process, without the (deprecated) full udevadm settle. Device discovery
itself needs no udev at all — /sys/block/*/serial is a sysfs attribute the
kernel populates at probe time — but formatting a blank disk does rely on
udev to make the new label visible as /dev/disk/by-label/<name>. udev
normally relabels a freshly formatted disk through the "watch" rule in
60-block.rules, an inotify watch on the block device that fires a
synthetic "change" event when a writer (mkfs.ext4 included) closes it —
the kernel itself raises no such event on close — and that watch is only
armed once udev has already processed the device's own add/change event.
systemd-udev-trigger.service having run only guarantees those events were
queued, not that udev finished processing them before format-volumes
runs, so the watch may not exist yet when mkfs.ext4 closes the device and
no relabel ever happens. format-volumes closes that gap itself: right
after a successful mkfs.ext4 it runs udevadm trigger --action=change on
the device (and udevadm settle to wait for the label), rather than
relying on the watch. udevadm is optional — an image without it (Alpine's
mdev-based one) simply skips this, since OpenRC's localmount mounts by
LABEL through blkid at mount time and never needs a udev-created
symlink. The LABEL=<name> mount unit systemd generates from fstab is
bound to the resulting udev-created device unit rather than attempted once
and given up on, so it simply waits for the label to appear once
format-volumes has triggered it.
convert also writes /etc/contemper/volumes, one line per declared
volume (name serial-pattern fs mountpoint), which the script reads with
while read -r — never sourced.
Finding a disk. The script matches each volume's serial pattern
against /sys/block/*/serial (for local-qemu and Incus alike, this is
just the volume's name — see --target's serial handling); the matching
block device is what gets checked and, if needed, formatted.
Logging, and never failing boot. Every decision is logged — to the console, and to syslog if one is already running — naming the volume, the disk, and what happened and why. The script always exits 0: a mismatch or a formatting failure is visible in the log, never a reason to stop the machine from booting.
What's recorded. The merged helper and its winning variant show up
next to any --support image, in their own volumeHelper field in
contemper.json and their own volume-helper.* lines in
/etc/contemper/build — ref, digest, and the resolved variant, so a
disk's exact provenance (support image and volume helper alike) is
always readable from the bundle or the guest itself.
Seeding a volume from the image¶
A volume's mount point doesn't have to start out empty in the image: if
you COPY or otherwise write files to /data before (or after)
declaring VOLUME /data, that content stays part of the image like any
other file. Without seeding, mounting a blank volume there on first boot
would simply hide it — the same surprise Docker's bind mounts would give
you, which is exactly why Docker seeds a named volume from the image
the first time it's used instead. contemper does the same: the first
time a volume's disk is blank (see How the volume helper
works above), right after formatting it
and before it is mounted, the volume helper copies whatever the image
has at that mount point onto the new filesystem with cp -a, preserving
ownership, permissions, timestamps, symlinks and special files (and, on
images with GNU coreutils, extended attributes). The volume's root
directory also takes over the owner and mode of the image's directory,
even when that directory is empty, so an image that only prepares an
empty mount point owned by its service user gets a volume that user can
write to. A reused volume — one already carrying the right label —
is never touched by this, or by anything else: whatever changes a
running appliance made to it stay exactly as they are across every
later redeploy or image update, seeded content included.
If seeding fails (a mount error, or a copy that runs out of space on the volume), the helper logs it and boot carries on: the volume stays formatted and mounts with whatever was copied before the failure. It is not retried on a later boot, since by then the disk carries its label and counts as reused.
A volume nested under another one (/data and /data/cache) works the
same way: the outer volume is seeded with the inner one's mount point
(and, as in Docker, whatever the image has below it, which the inner
volume then hides), so the inner volume has a directory to mount on.
Opt a volume out of seeding with its own label, the same
io.contemper.volume.<path>.* scheme every other per-volume setting
uses:
LABEL io.contemper.volume./data.seed="false"
An opted-out volume still gets formatted and mounted exactly as before;
it just starts empty, with mkfs.ext4's root-owned root directory, even
when the image has content at that path. Opting out the outer one of two
nested volumes also leaves out the inner one's mount point, which then
has to be created some other way before the inner volume can mount.
convert records every opted-out volume's name (not its path) in
/etc/contemper/volumes-noseed, one per line — written only when at
least one volume actually opts out, so a missing file means "seed
everything" to the volume helper, the same as an image converted before
this existed. convert also reports, on stderr, what it worked out for
each declared volume — seeded, opted out, or nothing there to seed in
the first place:
✔ /data → data 10.0 GiB
└ seed: image has content at /data, copied onto the volume on first format
Seeding only ever happens once, the first time a volume's disk is formatted — it is not a sync: an image update that adds or changes files at the volume's path has no effect on a volume contemper has already formatted and (optionally) seeded, the same way any other change to a reused volume's expected content doesn't. Convert still reports what a fresh volume would get seeded with, since that's the only time it matters.
Deploying locally¶
deploy --to local-qemu keeps one qcow2 file per volume in a
per-instance state directory,
${XDG_STATE_HOME:-~/.local/state}/contemper/local-qemu/<instance>,
created on first deploy and reused on every later one — the disk survives
even though the root disk itself resets on every boot (snapshot=on).
The instance defaults to the source image's repository name, without its
tag, so redeploying a new tag of the same image reuses the same volumes;
--name overrides it, for running more than one instance of the same
image side by side.
The first deploy records which image the instance's volumes belong to,
and a later deploy of a bundle from a different image under the same
instance name is refused instead of attaching those volumes. A bundle
you received from someone else gets its own --name, so it never sees
the volumes of another instance. The record comes from the bundle's own
manifest, so it catches two different images sharing an instance name by
accident, but not a manifest edited to claim another image. When a deploy reuses existing volumes,
it prints a line naming them.
$ contemper deploy --to local-qemu _out/my-appliance-v2.aarch64/
📁 instance my-appliance ~/.local/state/contemper/local-qemu/my-appliance
✔ /data → data created (10.0 GiB)
contemper never resizes an existing volume disk. If a later deploy asks for a different size than the disk already on file, it fails rather than silently growing or shrinking it — remove the file yourself (or deploy without changing the size) to start over.