Thank you for visiting!
My little window on internet allowing me to share several of my passions
Categories:
- VoidLinux
- vdcron
- OpenBSD
- FreeBSD
- ZFS
- Tunnel
- fapws
- Nvim
- Firewall
- got
- PEKwm
- Zsh
- VM
- High Availability
- My Sysupgrade
- Nas
- VPN
- DragonflyBSD
- Alpine Linux
- Openbox
- Desktop
- Security
- yabitrot
- nmctl
- Tint2
- Project Management
- Hifi
- Alarm
Most Popular Articles:
Last Articles:
Recreating OpenBSD's daily and weekly reports on Void Linux with ZFS: the checks I run and why
Posted on 2026-10-04 17:26:00 from Vincent in VoidLinux
On OpenBSD, /etc/daily and /etc/weekly quietly watch over the machine and mail you a report. On FreeBSD, periodic(8) does the same. I love that model, but when I wanted ZFS on my laptop, Void Linux turned out to have better drivers than FreeBSD for my hardware. The catch is that Void ships no periodic maintenance at all. So I wrote two small POSIX shell scripts, one daily and one weekly, that mimic the BSD reports and add checks specific to ZFS and to a laptop. They log everything, compare today with yesterday, and interrupt me only for real problems such as a degraded pool, a failing disk, or a kernel installed without its ZFS module. This post walks through each check, the command behind it, and why it is there.
I can put both at the top of void-zfs-maintenance-blog.md if you want. I can also write a shorter, punchier title, for example "Daily and weekly maintenance on Void Linux with ZFS, the BSD way", which matches the current heading of the post.
Daily and weekly maintenance on Void Linux with ZFS, the BSD way
I run OpenBSD on most of my machines, and on my laptop it works fine. But I wanted ZFS on that laptop, and for that Void Linux has better drivers than FreeBSD does on my hardware. So the laptop runs Void, with a root pool called vpool, a few boot environments under vpool/ROOT, and a separate vpool/home.
What I missed immediately was the quiet reassurance of the BSD periodic scripts. OpenBSD runs /etc/daily, /etc/weekly and /etc/monthly from cron, and mails root a report (see daily(8), which documents all three). FreeBSD does the same through periodic(8). You rarely read the report, but when something is wrong it is the first place it shows up. Void has nothing like that: cron is only a scheduler, and the directories it might run are empty. Even the choice of cron implementation is up to you.
So I wrote two POSIX shell scripts, one daily and one weekly, that try to give me the same safety net, adapted to a ZFS laptop. This post goes through what they check and why, and shows the command behind each check, with its options and pipes explained. The tests matter more than the code, and most of them translate to any system. Man page links go to the Void Linux manual pages (man.voidlinux.org) for Linux and Void tools, and to the OpenBSD and FreeBSD manuals for the BSD ones.
Design: a report, plus a few things worth interrupting you for
Both scripts print a plain-text report to standard output, section by section, and I send that to a log file. That is the BSD model, and it keeps the output easy to grep and easy to diff between days. Every section starts with a tiny helper that prints a banner:
section() { printf '\n==== %s ====\n' "$1"; }
have() { command -v "$1" >/dev/null 2>&1; }
section gives the log its structure. have asks the shell whether a command exists (command -v prints the path and returns success if it does), with all output thrown away, so that optional tools such as smartctl or xcheckrestart can be skipped cleanly when they are not installed.
On top of the report, a small number of conditions are treated as major, meaning I should know about them today and not whenever I next read the log. Those raise a desktop notification through a small helper of my own, which is not interesting here. The important design decision is the line between the two categories. A failed login or an orphaned package is information and belongs in the log. A failing disk, a degraded pool or a kernel update that is waiting to be installed deserves an interruption. If everything interrupts you, you stop looking, so the list of major conditions is short.
Two kinds of notification behave differently. Conditions that stay true until you fix them, like an unhealthy pool, notify on every run. One-off events, like a particular kernel update becoming available, notify only once. The script remembers a key for each in a small state file:
alert_once() {
key=$1; shift
if grep -qxF "$key" "$ALERTED"; then
echo "ALERT (already notified): $*"
elif alert "$@"; then
echo "$key" >> "$ALERTED"
fi
}
grep -qxF searches the state file quietly (-q, only the exit status matters) for a line that matches the key exactly (-x whole line, -F fixed string, no regular expression). If the key is already there, nothing pops up. If not, the notification is sent, and only when it was sent successfully is the key appended. A new kernel version has a different key, so it produces a new notification, but the same one does not repeat every morning.
Most of the other state lives in /var/backups/maintenance, where the "what changed since yesterday" checks keep their baselines. This idea comes straight from OpenBSD's daily script, which diffs important files against a saved copy and mails you the difference.
The daily script
Pool health and capacity
The first section is ZFS itself:
zpool list vpool
zpool status -x vpool
zpool list -H -o health vpool
zpool list gives the one-line overview. zpool status -x prints only pools with problems, and for a healthy pool just the sentence "pool 'vpool' is healthy", which makes it easy to test for in a case statement. The third command reads the health column directly: -H (scripted mode) removes the header and separates columns with tabs, and -o health selects only that column. The script insists on the word ONLINE; anything else is a notification.
Capacity gets its own test, with a threshold of 80 percent. ZFS is a copy-on-write filesystem and gets noticeably unhappy as a pool fills up, so the warning belongs well before the pool is actually full.
zpool list -H -o name,capacity vpool |
awk -v max=80 '{ c = $2; sub("%", "", c); if (c + 0 >= max) print "pool " $1 " is " $2 " full" }'
The pipe feeds a line like vpool 13% into awk. -v max=80 passes the threshold in as a variable. Inside, $2 is the second field (13%), sub("%", "", c) strips the percent sign from a copy, and c + 0 forces a numeric comparison. Only if the number reaches the limit does awk print a message, so a healthy pool produces no output at all.
For the same reason I stopped using df for ZFS filesystems. df reports per-dataset figures that say little about the pool, and an inode check is meaningless here because ZFS allocates inodes dynamically. What remains of df covers the filesystems that are not ZFS, which on this laptop means only the EFI system partition:
df -hP -x zfs -x tmpfs -x devtmpfs -x squashfs -x efivarfs |
awk 'NR > 1 && ($5 + 0) >= 85 { printf "%s is %s full; ", $6, $5 }'
-h gives human-readable sizes, -P forces the POSIX output format so that every filesystem stays on one line, and each -x excludes a filesystem type. In the awk part, NR > 1 skips the header line, $5 is the Use% column (+ 0 again turns 16% into the number 16), and $6 is the mount point. A full ESP gets a warning at 85 percent, because a kernel update that cannot write to it is a nasty way to find out.
Snapshots
The script then creates a daily snapshot of the active boot environment and of the home dataset. The active dataset is not hard-coded; it is whatever is mounted on /:
root_ds=$(findmnt -n -o SOURCE /)
findmnt shows what is mounted where, -n leaves out the header, and -o SOURCE prints only the device or dataset, here something like vpool/ROOT/void. This keeps the script working after I switch to another boot environment. The snapshot itself is created with zfs snapshot like this:
zfs snapshot "$ds@daily-$(date +%Y-%m-%d)"
Before that, the script checks with zfs list -H -t snapshot -o name <snapshot> whether it already exists, so running it twice in one day does nothing the second time, which makes it safe to test by hand.
Old snapshots are pruned, keeping the last seven, with zfs list and zfs destroy:
zfs list -H -t snapshot -o name -s creation -d 1 "$ds" |
grep "^$ds@daily-" | head -n -7 |
while read -r old; do zfs destroy "$old"; done
-t snapshot lists snapshots instead of filesystems, -o name prints only their names, -s creation sorts them by creation time (oldest first), and -d 1 limits the listing to this dataset's own snapshots, without those of child datasets. The grep keeps only names that start with the dataset name followed by @daily-. head -n -7 is a GNU extension that prints everything except the last seven lines, which are exactly the snapshots to delete, and the while read loop destroys them one by one.
Only snapshots matching the daily- prefix are ever candidates. This matters because I already have a separate script that takes snapshots and replicates them to my NAS with incremental zfs send and zfs receive. Those snapshots use different names, so the two jobs never touch each other's work.
That leads to the most useful snapshot test: it looks at the newest snapshot that is not one of the daily ones and checks its age.
zfs list -Hp -t snapshot -o name,creation -s creation -d 1 "$ds" |
grep -v "^$ds@daily-" | tail -1 |
awk -v now="$(date +%s)" '{
age = int((now - $2) / 86400)
printf "latest other snapshot: %s (%d days old)\n", $1, age
if (age > 10) print "WARNING: no recent weekly snapshot, check the NAS job"
}'
The new option here is -p, which prints exact, parsable values, so the creation time comes out as a Unix timestamp instead of a formatted date. grep -v inverts the match and drops my own daily snapshots, tail -1 keeps the newest remaining one, and awk subtracts its timestamp from the current time (date +%s, seconds since the epoch) and divides by 86400 seconds per day. If the weekly NAS snapshot is more than ten days old, something is wrong with the backup job. A backup that silently stopped working is exactly the type of failure that stays invisible until the day you need it, so the daily script watches the other script for me.
Services
Void uses runit, so the services check loops over /var/service and asks sv status about each one:
for s in /var/service/*; do
out=$(sv status "$s" 2>&1)
case $out in
*down:*"normally up"*) echo "$out" ;; # and remember it for the notification
run:*) ... ;;
*) echo "$out" ;;
esac
done
sv status prints lines such as run: /var/service/sshd: (pid 812) 3600s or down: /var/service/foo: 5s, normally up. The case patterns test for those words. A service that is down but normally up is a notification. A service that was stopped on purpose is still listed in the report, but it does not raise one. For a running service, the script extracts the uptime in seconds with sed: sed -n 's/.*) \([0-9][0-9]*\)s.*/\1/p', which prints only the digits between ) and the trailing s (-n plus the p flag means only lines where the substitution matched are printed). A service that started less than a minute ago is listed as "RECENTLY STARTED", because a service that keeps restarting looks perfectly healthy in a single snapshot and only shows its problem in the uptime counter.
Processes, logins and the kernel log
A few cheap checks follow. Zombie processes are found with ps:
ps -eo stat,pid,ppid,comm | awk '$1 ~ /^Z/'
-e selects all processes and -o chooses the columns, with the state first. awk prints only the lines whose first field starts with Z, the state of a zombie.
The authentication logs are searched for trouble:
grep -iE 'failed password|invalid user|authentication failure' /var/log/socklog/auth/current | tail -20
-i ignores case, -E enables extended regular expressions so that | means "or", and tail -20 keeps the last twenty matches. This section is for reading, not alerting, since on a laptop behind NAT it is mostly my own typos.
The kernel log is checked in two ways. The first prints the error-level messages from dmesg since boot:
dmesg --level=err,crit,alert,emerg | grep -v "Unknown key identifier" | tail -20
--level filters by severity, and grep -v removes one known harmless udev complaint about an unknown key identifier from the hardware database. Filtering known noise is worth the effort: a report that always contains one meaningless warning trains you to ignore the warnings section. The second check is stricter and searches the whole log for patterns that really indicate trouble:
dmesg | grep -E 'I/O error|Buffer I/O|nvme.*timeout|ata[0-9.]+: .*(failed|error)|Machine check|Oops:|BUG:' | tail -5
That regular expression covers I/O errors, NVMe timeouts, SATA link problems, machine check exceptions and kernel oopses. A hit is a notification, but only once per distinct message: the key used for alert_once is a checksum of the matching lines, computed with cksum and cut, as in cksum | cut -d' ' -f1 (cut splits on a space with -d' ' and keeps the first field, the checksum itself). The same old error does not interrupt me every day, but a new one does.
Disk health
If smartmontools is installed, each disk gets a SMART health check:
smartctl -H /dev/nvme0n1 | grep -iE 'overall-health|SMART Health Status'
smartctl -H asks the drive for its overall health verdict, and the grep extracts the single line that carries it, which has a different wording on SATA and NVMe. Anything other than a pass is a notification. For the NVMe drive the script goes one step further and reads two attributes:
smartctl -A /dev/nvme0n1 | awk -F: '/Critical Warning/ { gsub(/[ \t]/, "", $2); print $2 }'
smartctl -A /dev/nvme0n1 | awk -F: '/Percentage Used/ { gsub(/[ \t%]/, "", $2); print $2 }'
-A prints the attribute table. awk -F: splits each line on the colon, a pattern in slashes selects the matching line, and gsub deletes spaces, tabs and the percent sign from the value so that only 0x00 or a bare number remains. A non-zero critical warning notifies immediately. Wear notifies once when the drive reaches 90 percent of its rated life. A drive that still reports "PASSED" can be quite far along its wear curve, so the overall verdict alone is a weak signal.
Configuration drift
Four files are tracked: /etc/passwd, /etc/group, /etc/sudoers and the SSH daemon configuration. The script keeps a copy of each in the state directory, and compares it with the live file:
if ! diff -u "$saved" "$f"; then
cp -p "$f" "$saved"
alert "$f changed, see the daily log for the diff"
fi
diff -u prints a unified diff and returns a non-zero status when the files differ, which is what the if ! tests. In that case the diff is already in the report, and cp -p refreshes the baseline, preserving mode and timestamps, so tomorrow's run starts clean. The first run simply records the baselines. This is the OpenBSD approach, and it is one of the best features of that daily report. The diff in the log tells me whether the change was me.
Updates, and the one check I care most about
The updates section deliberately stays small:
updates=$(xbps-install -Sun 2>&1)
n=$(printf '%s\n' "$updates" | awk '$2 == "update" { n++ } END { print n + 0 }')
printf '%s\n' "$updates" | awk '$2 == "update" { print $1 }' | grep -E '^(linux[0-9.]*|zfs|dkms)-[0-9]'
xbps-install -Sun combines -S (synchronise the repository index), -u (update) and -n (dry run), so nothing is installed. Each pending update is a line whose second field is the word update, so awk can count them (n + 0 makes the result print 0 when there are none). I only print the number, because forty-five pending updates are not news. What is news is a new kernel or ZFS release, so the third command prints the names of pending packages and grep -E keeps those that start with linux, zfs or dkms followed by a dash and a digit. The digit is deliberate: it keeps out packages such as linux-firmware-amd, which do not need a module rebuild. Any match triggers a one-time notification.
The next section is the one that justifies running ZFS on Void with a scrutiny the BSDs never required. On Void, ZFS is built as an out-of-tree module with DKMS, separately for each installed kernel. If a kernel update arrives and the module build fails, the next reboot lands in a kernel that cannot import the pool, and on a system with ZFS as its root that is a very bad morning. So the script prints the DKMS status and then verifies the result on disk:
dkms status
for d in /usr/lib/modules/*; do
k=${d##*/}
[ -e "/boot/vmlinuz-$k" ] || continue
find "$d" -name 'zfs.ko*' | grep -q . || echo "WARNING: kernel $k has NO zfs module"
done
${d##*/} is shell parameter expansion that strips everything up to the last slash, leaving the kernel version. The [ -e ... ] || continue line skips module directories that have no matching kernel image in /boot, so leftovers do not cause false alarms. find ... -name 'zfs.ko*' looks for the module file (the * covers compressed variants such as .ko.xz), and grep -q . succeeds only if find printed at least one line. Any kernel without a module is a notification that says, in effect, do not reboot into this. It is a simple file-existence test and it covers the most dangerous failure in this setup.
Last, the script checks whether a reboot is needed:
[ ! -d "/usr/lib/modules/$(uname -r)" ]
If the module directory of the running kernel (its version comes from uname -r) no longer exists, a newer kernel replaced it and the running one is orphaned. If xcheckrestart (from the xtools package) is available, the script also runs it to list processes that still use libraries that were replaced on disk, which is the closest Void has to the "you should restart these services" hint from the BSDs.
The weekly script
The weekly script covers things that do not need to happen every day, and a few that are expensive.
Battery
Because this is a laptop, the first section reads the battery straight from sysfs:
full=$(cat /sys/class/power_supply/BAT0/energy_full)
design=$(cat /sys/class/power_supply/BAT0/energy_full_design)
pct=$((full * 100 / design))
energy_full is what the battery can hold today and energy_full_design is what it could hold when new; some drivers expose charge_full and charge_full_design instead, and the script tries both. The shell's $(( )) does integer arithmetic, so the result is the health as a percentage. The cycle count is printed too when cycle_count exists, and a one-time notification fires if health falls below 60 percent. Batteries rarely fail suddenly, so a weekly look at the trend is plenty.
Scrubs
ZFS needs periodic scrubs to detect and repair silent corruption, and this is where laptop-awareness matters. First the script shows the result of the last scrub, using zpool status:
zpool status vpool | sed -n '/scan:/,/config:/p' | sed '$d'
The first sed -n '/scan:/,/config:/p' prints only the range of lines from the one containing scan: to the one containing config:; the second sed '$d' deletes the last line of that range, so the unwanted config: header is dropped. If that text says a scrub repaired errors, matched with grep -E 'scrub repaired .* with [1-9][0-9]* errors', I get a notification once per distinct result.
Whether a new scrub is due is decided by a stamp file:
find "$STATE/scrub-vpool" -mtime -28 | grep -q .
find -mtime -28 matches a file modified less than 28 days ago, so a successful match means the stamp is recent and no scrub is due. If it is due, the script checks that the machine is on mains power before starting:
[ "$(cat "$d/type")" = "Mains" ] && [ "$(cat "$d/online")" = "1" ]
zpool scrub vpool && touch "$STATE/scrub-vpool"
Every entry in /sys/class/power_supply has a type file, and the adapter reports Mains with online set to 1 when plugged in. A scrub reads the entire pool and I do not want it draining the battery or waking the disk on the train. If I am on battery, the scrub is postponed and the script says so. The stamp is only touched when zpool scrub really started, thanks to &&. It also never starts a second scrub on top of a running one, which it detects with grep -q 'scrub in progress' on the status output.
TRIM
For SSDs, fstrim does not work on ZFS, so the script uses the pool-level equivalent:
zpool get -H -o value autotrim vpool
zpool trim vpool
zpool get reads a pool property, with -H and -o value again giving just the bare value. If autotrim is already on, the pool trims continuously and nothing is needed; otherwise the script starts a one-off zpool trim.
Boot environments and space
Boot environments are one of the best things about ZFS on a root filesystem, and also one of the easiest to forget about. The script prints the active environment and lists all of them, followed by a space breakdown:
zfs list -r -o name,used,refer,creation vpool/ROOT
zfs list -r -o space vpool
-r recurses into child datasets and -o picks columns. The special column group space expands to the available space and a split of the used space into snapshots, the dataset itself, refreservations and children, which shows how much of the usage is really snapshots. The script never deletes anything. Old boot environments are a decision for a human, but the weekly list makes it hard to forget that they exist.
Package housekeeping
The package part is mostly reporting, with one exception:
xbps-remove -O -y # clear obsolete files from the package cache
xbps-query -O # list orphaned packages
vkpurge list # list kernels that could be removed
xbps-pkgdb -a # check the integrity of all installed packages
xbps-query -m > /var/backups/maintenance/xbps-manual.list
xbps-remove -O removes obsolete files from the cache and -y answers the confirmation question, which is safe and automatic. xbps-query -O lists orphans, packages that were installed as a dependency and are no longer needed by anything. vkpurge list shows old installed kernels, and I leave the removal to myself. xbps-pkgdb -a checks every installed package against its metadata, and if its output contains the word "error" I get a notification, once per distinct output. The last command writes the list of manually installed packages (-m) to a file, which is a cheap way to be able to rebuild the system after a disaster.
Security and indexes
The weekly script scans for setuid and setgid files and compares the list with last week's, the same diff-against-a-baseline pattern as the configuration check. Because ZFS puts / and /home in different datasets, a naive find / -xdev would skip home entirely, so the script first asks findmnt for every mounted ZFS filesystem:
mounts=$(findmnt -rn -t zfs -o TARGET)
find $mounts -xdev -type f \( -perm -4000 -o -perm -2000 \) | sort > "$STATE/suid.new"
diff -u "$STATE/suid.list" "$STATE/suid.new"
findmnt -r uses raw output without tree drawing, -n drops the header, -t zfs selects only ZFS mounts, and -o TARGET prints the mount points. The variable is deliberately left unquoted in the find call so that the shell splits it into / and /home. -xdev stops find from crossing into other filesystems, -type f limits it to regular files, and the escaped parentheses group two permission tests joined by -o (or). -perm -4000 matches files with the setuid bit set, and -perm -2000 the setgid bit; the leading dash means "at least these bits". The result is sorted so that diff compares like with like, and a new setuid binary appearing between two weeks is rare enough that it is worth a notification.
Finally the script refreshes the locate and man databases with updatedb and mandb -q (-q suppresses the progress output), a leftover from the BSD weekly script that I still like.
Scheduling and limits
I trigger both scripts from my own small cron wrapper and keep them in a directory under /root, away from files that packages manage. Whatever you use, schedule them for a time when the laptop is likely to be awake. Plain cron does not catch up on jobs that were missed while the lid was closed.
These scripts reflect my setup, and they have limits worth stating. The pool name and the home dataset are variables at the top and should be changed for another system. The log paths assume the socklog setup common on Void. The regular expressions that decide which dmesg lines are serious are a judgment call, and will need tuning. A few options, such as head -n -7, are GNU extensions, which is fine on Void but would need changing on a BSD. I also run them on one machine only. I would rather you read them as a list of ideas than as a package.
What I like most about the result is how closely it follows the philosophy of the BSD originals: do nothing clever, report on everything, compare today with yesterday, and stay quiet unless something really needs you. Void gives you a fast, minimal base and leaves the housekeeping to you. These two scripts are my answer, with a couple of ZFS and laptop-specific checks added on top.
Conclusion
Since my VoiLinux is running on my laptop, I trigger those 2 scripts via vdcron, my specific cron.
Logs are available into the vdcron log file.
In some critical situations, my very specific version of daily and weekly scripts send messages to the screen via xcowsay.