Saving 100 terabytes of memory by optimizing 1.1.1.1's DNS cache

(blog.cloudflare.com)

49 points | by TangerineDream 42 minutes ago

3 comments

  • irdc 19 minutes ago
    This is why system programming still matters.

    Looks like they're missing the obvious optimisation of putting the record data right after the CacheEntry members instead of allocating memory separately though. But that might just be me as a C-programmer talking and not be all that easy in Rust.

  • strenholme 5 minutes ago
    With my own MaraDNS, I aggressively optimized the memory usage of blacklist entries by having a single really big malloc() to allocate the memory for the entries, then traversing that memory block for potentially blacklisted entries.

    When I was using one malloc() per entry, a large blacklist took up 237 megabytes of memory. The same blacklist, once optimized to be loaded with a single malloc() call, only took up 9.5 megabytes of memory.

    https://samboy.github.io/blog/entries/MaraDNS.html#BlogEntry...

  • eviks 22 minutes ago
    > Once we store a DNS response in the cache, however, we never modify it again. The capacity field serves no purpose, but still costs 8 bytes per Vec

    Were there no design discussions/reviews when the system was setup to catch trivial things like this?

    • lbriner 18 minutes ago
      It is often not worth optimising in the early days. You don't know how popular it will become, you might not know how many DNS records you will hold, it was possibly written in an earlier language and ported as-is.

      At the point someone queries the 100TB of RAM, then maybe it is worth revisiting but even that has risks. You have to design the migration path, have fallback mechanisms etc.

      • eviks 8 minutes ago
        It's also often that you can avoid all those future migration/fallback risks and pains if you invest a little bit of design thinking upfront.

        So how would you decide which path to take in situations like this?

    • mhitza 19 minutes ago
      Premature optimization argument fits right in. Now that memory is up to 10x more expensive it is worth considering optimizing programs with large memory footprint.
      • eviks 15 minutes ago
        How does that fit? What would be the evil of not wasting memory for many years at 1x?
        • jgrahamc 13 minutes ago
          One of the "evils" of premature optimization is how much time you spend on the optimization vs. the benefit you get from it. If your goal is correctness and shipping fast and you're not memory constrained then spending time using the least amount of memory is a waste of time specifically because you want to ship fast.

          Another interesting thing that happens is you don't necessarily know what form your actual optimizations will need to take. Later when your systems grow you discover the suboptimal parts you hadn't optimized for.

          Very early on at Cloudflare I worked on part of the DNS infrastructure that took DNS records from the UI and got them in a state for actual authoritative serving. The system had been constructed anticipating Cloudflare having millions of customers with unique domains, but it had not been constructed for a single customer with a single domain with millions of records. This caused a periodic slow down in DNS record updating while the system churned on that one customer.

          In a different job I worked on a piece of optimization software that needed to keep track of "node" A is reachable from node "B". This had been implemented as a matrix (literally a malloced NxN matrix of ints storing 0 or 1) which worked really well for small systems. But you'd be out of memory really fast on a large project. I replaced the matrix with a hash table and all was good because the matrix was actually really sparse.

          • stickfigure 1 minute ago
            Absolutely true, but I will say that LLMs have changed the equation somewhat.

            With a rather short prompt, claude/codex will take your code, write a harness, profile it, build experiments, profile those, and give some pretty solid advice which one to pick. It's the kind of goal-directed, bite-sized job that LLMs excel at. Extremely low-commitment.

            Except for the whole "making changes in production at scale" problem, of course.

        • gbear605 12 minutes ago
          Engineers are expensive, especially good system engineers who are trained in your code base. Very possible that this just hadn't gotten to the top of the priority list.
          • eviks 3 minutes ago
            I don't understand why you need training on your code base to design a cache format for read only vs rw workloads, but anyway yours is a comment about neglect, not the "evil" that would happen if you did that design
    • micromacrofoot 22 minutes ago
      it was working so no one thought to check