Btrfs/ZFS/bcachefs under workloads classic benchmarks skip

(bartosz.fenski.pl)

46 points | by farlight 1 hour ago

12 comments

  • fenio 9 minutes ago
    The author of the benchmark here. I went over some comments and I'll try to tackle them here. I'm pretty clear that GH runner based benchmark is far from perfect due to noisy neighbours etc. Thus every test first is running so called calibration... to reject completely unreliable VMs. I'm fully aware that this can't completely fix the issue. Can limit it but not fix. But as of now there are 593 runs recorded so average should still be quite meaningful.

    Having that said I'm desperately trying to get REAL hardware to run that benchmark. With some successes ;)

    Few months ago I got Hetzner machine from Kent Overstreet and I was able to finish 3 runs before machine died... Results: https://bartosz.fenski.pl/modern-fs-benchmark/real-hw/

    Currently I've got even more interesting machine with tons of disks and I'm running new set of benchmarks but it's really in its initial stage.

    https://bartosz.fenski.pl/modern-fs-benchmark/sas-hdd/ 2nd run in progress... one run on REAL hardware takes much more time than on GH runner so it's slow.

    But this new hardware has also so many disks that the plan is to try also more complex, tiered cache topologies. I'm working on it.

    I'm happy to answer any other questions, sources of every piece of this benchmark are freely available and I'm not saying they are 100% correct. I'm open to improvements.

    • koverstreet 4 minutes ago
      I went back and forth with Hetzner a couple times, I think we just got a bad machine :)

      I've been saying it for months, but eventually I'm going to move the automated builds off the 48 core monster and we'll be able to use that for automated perf testing too. The machine we just got has spindles for EC perf testing, but the Hetzner monster has very high end enterprise ssdd.

      Also, just got done with the Rust for Linux conference, still not home but here's slides that still need reformatting: https://evilpiepirate.org/~kent/Kangrejos-2026-bcachefs.pdf

  • Farmadupe 59 minutes ago
    > CI runs use loop devices on shared ephemeral VMs (one VM per filesystem): compare shapes and ratios, not absolute MB/s. Each job records a host-calibration anchor — see the table.

    I think if you're not using baremetal for such tests, it's likely that the results are simply not comparable at all? What if another tenant is also using the disk?

    • walrus01 53 minutes ago
      It's a fair point but it's also possible the person running the tests has a dedicated test hypervisor for this , so that different configurations of filesystems and VMs can be created and destroyed quickly in an automated manner.

      If it's something as simple as a KVM hypervisor that only runs 1 test VM at a time (with no other load from anything else other than the basic systemd daemons, ssh daemon etc running on the hypervisor), the results could be very close to bare metal.

      I can see it being very time consuming and annoying to do repeated manual bare metal OS installs and new partitioning/filesystem creation for such a large variety of tests.

      The author does also say that performance isn't really the main thing but rather, data integrity:

      https://github.com/fenio/modern-fs-benchmark

      • toast0 42 minutes ago
        > I can see it being very time consuming and annoying to do repeated manual bare metal OS installs.

        Well don't do that then. There's lots of other options. Probably the simplest is a single bare metal install on a simple filesystem on one device. run the filesystems under test on other storage dedicated to testing.

        You could also boot into a network install and use local storage exclusively for testing.

        • Farmadupe 39 minutes ago
          Yes exactly. The issueThe epherrality of the VMs isn't an issue, it's the _shared_ part that's the concern here. Going by the fact that the kernel is listed as "kernel 7.0.0-1012-azure" I feel like it's a fair risk that there may have been noisy neighbours.
      • Farmadupe 43 minutes ago
        > compare shapes and ratios, not absolute MB/s

        In this case, given that the author's own disclaimer (above) already disclaims the numeric readings, I'm not sure how it's possible to make any inference on "shapes and ratios" derived from the numeric readings.

  • magicalhippo 11 minutes ago
    > Every push/2-hourly cron builds each filesystem across 4 loop devices backed by sparse files, runs the suite, and publishes a results table in the job summary plus JSON artifacts.

    I get that real hardware costs (author mentions EUR 70 a month for a suitable server), but without at least a baseline snapshot comparison run between real hardware, both SSD and HDD, and the sparse file-backed loop devices, it's hard to take much away from this.

    Sadly the AI apocalypse isn't making stuff like this easy to do as a hobby.

  • loeg 1 hour ago
    What is md-raid10 doing that is so much worse than lvm-raid10? In terms of "I/O" and "responsiveness." It's not really obvious to me from either the linked page or https://github.com/fenio/modern-fs-benchmark . In principle they should be similar?
  • blop 45 minutes ago
    I think the reviews should also include the social aspect of these filesystems...

    There is and have been many promising and exciting FS to replace the old boring ones, but for storage you not only want to avoid technical issues but also maintainer(s) drama...

    • koverstreet 14 minutes ago
      Why do people keep bringing up drama?

      The community infighting has sucked, but that's a thing that matters primarily for maintainers.

      I think most users just want something that works.

      • Skunkleton 7 minutes ago
        Related username?

        To answer the original question, most people who care about their filesystem at all care about its stability. Not just "does it work now" but also "will it work and improve over time". Infighting puts the future at risk.

  • skerit 1 hour ago
    Oh, so bcachefs is doing pretty well.
    • tarruda 1 hour ago
      Except for the fact that the developer has sabotaged the project into being removed from mainline?
      • tombert 32 minutes ago
        It's relatively easy to get it working as a kernel module at least. I got it set up on a NixOS box without too much trouble.
      • AceJohnny2 49 minutes ago
        that's not necessarily a sabotage.
        • eikenberry 27 minutes ago
          Sabotage might not be the best word, but it hurt trust and adoption.
          • koverstreet 11 minutes ago
            It's just been a lot less drama within the project since the split.

            I do have a lot more pull requests to merge than I did before. I don't know if you want to count "Kent isn't reviewing Pars fast enough" as drama :)

      • irusensei 27 minutes ago
        That might be the best thing happened to the project since now development can happen at its own pace without the clicky bait influencers.

        In fact they delivered the erasure coding for parity raid back in march this year.

        The thing is that as soon as you seriously give a chance to Bcachefs you see how good it is. I can only tell you that mixing different device tiers and having a per-file/directory replication setting is a god send specially in these times where storage costs more than gold.

  • sippingabonedry 28 minutes ago
    So two filesystems that are essentially shunned from the Linux kernel and permanent second-class citizens, and one that was removed from Red Hat and has a questionable history of reliability. Oh boy which do I choose?

    I'm saying ZFS on another OS.

  • farlight 1 hour ago
  • blop 1 hour ago
    For peace of mind I'm still using zfs (since the last 15+ years) but I'm definitely not impressed by the performance...
    • slyfox125 1 hour ago
      Different tools for different jobs; use ZFS for your data store and ext4 for your primary drive.
      • blop 50 minutes ago
        yes indeed, zfs for my nas basically
  • markhahn 34 minutes ago
    what does "integrity" fail mean in the first table? that the case didn't recover from the 2G corruption?
  • Farmadupe 1 hour ago
    @farlight assuming that you're the creator do you think you'd be able to rework the HTML/CSS? I'm sure you've got good data but speaking on behalf of my eyeballs, the results page is... hard to read!
    • Farmadupe 1 hour ago
      > 2G of random garbage is written directly onto one member device (behind the filesystem's back, offset 1G — python injector; uutils dd mis-seeks on dm devices), caches dropped, then a full scrub: btrfs scrub -B, zpool scrub + wait, bcachefs scrub, md/lvm sync-action 'check' (which can only COUNT mismatches — no checksums to know which copy is right).

      I'm not sure that nuking 2G of the underlying block device is a recoverable error on any filesystem that I'm aware of? Can you confirm if any ofthe filesystems really came out of the other side in a usable state after scrubbing?

      -----

      > Trivial-op p99, idle (ms) # A trivial operation — one 4k write + fsync every 200ms (like a shell appending history or an editor updating its swap file) — run alone for 10s. p99 of the fsync completion

      In fact, if it's OK for me to ask, are any of the metrics tht you used standard industry metrics? It looks like several of the tests are bypassing the kernel's page cache? -- which I worry may fall into the trap of "I modified the system to be unrepresentative of reality and then tested it".

      ----

      > kernel 7.0.0-1012-azure

      Can you confirm if you tested on a bare metal machine? were you the only tenant?

      • vlovich123 1 hour ago
        1 device out of the replica set I’m assuming so all of them should recover.
      • hlieberman 1 hour ago
        The integrity check is only on the tests which are either RAID or the filesystem equivalent.
    • fenio 2 minutes ago
      what exactly would you like to improve?