Announcement

New LTS for FreeBSD is now available in an initial release  Learn More

Klara

Key Article Takeaways 

  • Buying a SLOG without checking for sync writes wastes money. Measure with zpool iostat -yr first.  
  • RAIDZ expansion adds capacity, not a new RAIDZ level. Old data keeps its original parity ratio until rewritten.  
  • Fast Dedup and Direct IO both punish blanket adoption. Scope them to the workload that actually needs them, quota or no quota.  
  • A special vdev needs redundancy matching the pool, and snapshots need automated pruning. Skip either and the pool degrades quietly until it doesn't. 

In Part 1 of Biggest ZFS Misconfigurations and How to Fix Them, we covered the classics: wrong ashift, recordsize misaligned with the workload, overly wide RAIDZ where it should have been mirrors, being overly optimistic about dedup, and pools full to the brim. Those mistakes are old enough to vote. This time we cover issues of a newer vintage. OpenZFS 2.3 shipped many new and useful capabilities (RAIDZ expansion, Fast Dedup, Direct IO), and every new capability arrives with a brand-new way to hold it wrong. We'll also pay off two promises from Part 1: the SLOG people keep buying as a write cache, and the special vdev that quietly becomes a single point of failure. 

As before: we’ll cover why it hurts, how to detect it, and what the fix will cost. 

1. The SLOG You Bought as a Write Cache 

A separate log device accelerates exactly one thing: synchronous writes. Async writes, which are most writes on most systems, things like copying files and other day-to-day tasks, aggregate in RAM and flush to the disks in large transaction groups. They never touch the SLOG. Before spending money on expensive low-latency, high-endurance NVMe hardware, find out whether you have sync writes at all: 

$ zpool iostat -yr tank 5     # print histogram of request sizes broken down by 
                              # read vs write, and sync vs async 

Under your usual load, how many of the writes are sync_write compared to async_write? Remember that writes under the “agg” column are multiple smaller writes aggregated into a larger, more efficient write, compared to individual “ind” writes. 

If you run NFS, databases, or VMs that fsync constantly, asking for guarantees that the data is on disk before continuing, then a proper SLOG (low-latency NVMe with power-loss protection) can completely transform the performance of your pool by reducing the write latency. If you don't, it does nothing at all, doesn’t even get sent any commands, and that money could have been spent on RAM instead.

The shortcut version of this mistake is worse. zfs set sync=disabled makes benchmarks fly by lying to applications: ZFS acknowledges the write, then keeps it in RAM for up to several seconds before it reaches stable storage. Power loss in that window and the pool survives intact, but your database's last acknowledged transactions are gone. If sync=disabled makes your workload fast, that measurement is telling you to buy the SLOG.

Bonus tip: With OpenZFS 2.4, if you have a special (metadata) vdev, it can act as a SLOG, with the new “embedded special log” feature. This can be very helpful since most systems have a limited number of NVMe drive slots. Instead of dedicating a pair of those slots to a SLOG, which only needs to be a dozen GiB in size, a special vdev can set aside some space to act as a SLOG and use the rest for metadata acceleration. However, remember that redundancy of a special vdev is even MORE important; see misconfiguration number 5. 

2. RAIDZ Expansion Is Not a Do-Over

OpenZFS 2.3 was the first release to ship with RAIDZ expansion, a feature many small deployments waited over a decade for: grow a RAIDZ vdev one disk at a time. 

$ zpool attach tank raidz1-0 da4    # 4-wide raidz1 becomes 5-wide 

The misconfiguration is in expecting more than the feature promises. Expansion reflows existing data across the wider stripe, but pre-expansion blocks keep their original data-to-parity ratio; only newly written blocks use the new width. Expand a 5-wide RAIDZ2 to 6-wide, and old blocks are still 3 data + 2 parity, so the full capacity gain arrives gradually, as data is rewritten, not the moment the reflow finishes. The new contiguous free space is available immediately, but the existing data is using the less efficient ratio. Plan for that or rewrite the data deliberately after expanding. This process can be helped by the new zfs rewrite command, introduced in OpenZFS 2.4.0 and included in 2.3.4 and later. 

Know what RAIDZ expansion doesn't do: it won't change the RAIDZ level (a Z1 stays a Z1), it won't remove disks, and it won't rescue the wide-RAIDZ-instead-of-mirrors layout we discussed in Part 1, because the IOPS profile doesn't change. And since the reflow reads and rewrites all allocated space, expanding a pool that is already at 90% capacity during business hours is choosing pain. Expansion is a planning tool for growing pools, not a repair tool for misdesigned ones. 

3. Fast Dedup: Better Math, But Still Math

Fast Dedup, which Klara built together with TrueNAS and donated to OpenZFS, fixes most of what made legacy dedup a trap: faster lookups, a logged table that batches updates, and for the first time, operator controls. The new misconfiguration is reading "fixed" and enabling it everywhere. 

The rules from Part 1 still apply: dedup only pays when the data actually contains duplicates, and the table still consumes memory. What's new is that you can finally govern it: 

$ zfs set dedup=on tank/vms              # still per-dataset, still scoped 
$ zpool set dedup_table_quota=10G tank   # hard ceiling on DDT size (RAM usage)
$ zpool ddtprune -p 20 tank              # evict oldest 20% of single-ref entries 

That quota is the difference between the old failure mode (the table quietly outgrows RAM and pool performance falls off a cliff) and a bounded system that degrades predictably. Be aware that there is no transition from old dedup. Fast Dedup applies to newly created dedup tables on pools with the feature enabled. A pool that has been running legacy dedup keeps the legacy table layout; there is no in-place upgrade mechanism. Turning the feature on does not retroactively repair any of your existing DDT problems.

4. Direct IO Where the ARC Belonged 

Also new in 2.3 is Direct IO, a dataset property that lets reads and writes bypass the ARC. 

$ zfs set direct=standard tank/pgsql    # honor O_DIRECT if the app asks 

On very fast NVMe pools, serving a workload that does its own caching, the ARC's extra memory copy is measurable overhead, and Direct IO removes it. The previous sentence has three qualifiers, all of which matter. The misconfiguration is direct=always on a general-purpose dataset because a benchmark on someone else's NVMe array said it helped. For most workloads, the ARC is the reason ZFS feels fast, and bypassing it means every read pays full media latency. Direct IO also expects properly aligned I/O sized to the dataset's recordsize; applications that don't cooperate lose the benefits. 

The safe deployment: leave direct=standard so applications that explicitly request O_DIRECT (databases configured for it) get it, and everything else keeps using the ARC. Reach for always only after measuring your workload on your hardware and being sure it is worth the tradeoffs. 

5. The Special Vdev That Isn't Special Enough 

Allocation classes let you put metadata (and optionally small blocks) on faster devices: 

$ zpool add tank special mirror nda0 nda1 
$ zfs set special_small_blocks=32K tank/projects 

Used well, a special vdev makes a big HDD pool feel dramatically faster, because metadata operations stop burning up IOPS with seeking. The misconfiguration is treating it like a cache. It is not a cache. It holds the only copy of your metadata: lose the special vdev and the pool is trashed. We have seen RAIDZ2 pools, built to survive concurrent dual disk failures, wearing a single consumer SSD as a special vdev. That pool's real fault tolerance is zero. The special vdev must match the pool's redundancy standard: mirror it, and on a RAIDZ2 pool, consider a 3-way mirror. 

L2ARC deserves its cameo here, because it is the opposite device with the same misunderstanding. L2ARC is safe to lose but rarely the right first move: its headers consume RAM, so bolting a cache device onto a RAM-starved system makes performance worse. Max out memory first. Persistent L2ARC survives reboots now, which fixed the old cold-cache complaint, but it didn't change the arithmetic. 

6. Snapshot Sprawl 

Automated snapshots without automated retention and pruning is how pools fill up in slow motion. Every deleted file whose blocks a snapshot references stays allocated, and a dataset showing modest usage can be pinning terabytes. 

$ zfs list -t snapshot -o name,used -s used | tail   # the expensive ones 
$ zfs holds -r tank@migration-2024                   # find forgotten holds 

The fix is policy as code: a snapshot tool with retention tiers (hourly for a day, daily for a month, monthly for a year, whatever your recovery objectives actually require) and a periodic report of snapshot-pinned space. When a pool from Part 1's "running it full" section shows up, snapshots are the cause about half the time. 

There is another trap here as well. Even if the snapshots are not consuming huge amounts of space from deleted blocks, there are scaling tradeoffs. Each dataset keeps a list of its snapshots; when that list is fewer than 1000 snapshots, it is a very simple list, and everything is fine. As you exceed that number, ZFS has to add indirection, breaking the list into leaf blocks, each containing some of the snapshots. As the lists grow, the time it takes to add, remove, and list snapshots gets a little longer. It never impacts the speed of reading or writing data, but if you don’t prune old snapshots just because you are not low on space, you can make listing those snapshots a painfully slow operation. Even if the wait doesn’t bother you, it eats up IOPS that could be put to better use. 

The Pattern 

Across many situations, the story is the same. The ZFS defaults are sane; problems arise when changes are made without understanding what they will do. Always measure before changing anything to see if the change will be beneficial. Enabling a feature without the conditions existing that could benefit from that feature will not provide any benefit;  dedup without duplicate data, a SLOG without sync writes, Direct IO without an app that caches for itself, a special vdev without redundancy, etc. The 2.3 features are excellent tools; a disciplined sysadmin measures first, then scopes the feature to the workload that justifies it, and knows which decisions can be revisited and which, like ashift, you will have to live with forever. 

If your enterprise runs on OpenZFS, ensuring that your storage is performing as well as possible matters. If you're not 100% confident about the state of your storage, consider a ZFS health check with Klara. If you would like help designing, configuring, or measuring your new or existing ZFS pools, consider Klara’s Storage Design service or a Performance Analysis and Tuning engagement. Get the experts who designed the features to help you make the most of them. 

  

 

 

Back to Articles