RIP, vector database

(turbopuffer.com)

61 points | by razin 57 minutes ago

6 comments

  • gopalv 13 minutes ago
    > This write amplification is large enough that our efforts to tune indexing throughput have started to hit diminishing returns.

    > don't key on the ANN address. That is precisely the change turbopuffer v3 makes. As you can imagine, it is not a trivial change.

    This is a direct parallel to how Postgres and Mysql built indexes.

    Your design choice went from a Postgres design pattern to a Mysql one. The difference is the reindexing cost vs the lookup cost - Postgres optimized for lookup and Mysql does for indexing on writes. Or more accurately, Postgres was better with good schema design using joins & mysql was optimized for a bad design with less normalization where many indexes exist for the same table.

    Postgres always points an index to a row-id within postgres which is an arbitrary value which changes on each update.

    Mysql, always assuming the storage engine is pluggable, points to the primary index entry and adds an extra indirection to the lookup.

    This means that you point the mysql index to a stable id, so unless you go update the primary key for a row, you won't have to update the indexes for all the attribute lookups you might have made to data.

    I don't do databases any more that much, but the design for NIMBLE file format has a lot of quirks which are relevant to this specific idea.

    But the old Uber post about switching from Postgres to Mysql to prevent index amplification[1] is a direct mirror to this post.

    [1] - https://www.uber.com/us/en/blog/postgres-to-mysql-migration/

    • phoghed 7 minutes ago
      > mysql was optimized for a bad design

      TIL I should have been using mysql the whole time

      • woadwarrior01 4 minutes ago
        Richard Gabriel's "Worse is better" vibes.
  • gk1 7 minutes ago
    Vector databases were always more about retrieval than either vectors or data storage. But the term stuck all too well and companies held on to it a tad too long. Sorry :)
  • drewlanenga 10 minutes ago
    the multi-vector duplication thing makes sense, copying every attribute once per vector explodes quickly. what's the new primary index?
  • OutOfHere 34 minutes ago
    It would be nice to have a page that actually loads. This one doesn't. RIP.
    • throwawy0352 22 minutes ago
      Loads really fast for me. (MacBook Air, average internet)

      If you still have issues, try https://web.archive.org/web/20261001100105/https://turbopuff...

      • wilj 12 minutes ago
        It has a pagespeed insights score of 55 and noticeably sluggish on my m3 max.

        And what's with the throwaway account for this one comment? Is this becoming reddit with throwaway shills now?

        • throwawy0352 0 minutes ago
          Yes, I get big money from the web archive to promote their services. It's the new scheme that shills like me go for.

          The reason is that I have no account at HN and rarely comment. I create a new account a few times a year because I don't remember, or care about my previous account.

          I could have made an account named john2042 and you would not think twice. Instead I let people know upfront what type of account this is. Quite the opposite of what a true shill would do.

        • phoghed 6 minutes ago
          fucking shills, making helpful comments and promoting seemingly nothing, what's this place coming to?
    • syndacks 30 minutes ago
      loads just fine on my $10k laptop with 10g internet here in NYC
      • alexjplant 17 minutes ago
        Takes 11 seconds to load on Firefox on Linux with 3G-level throttling enabled in Dev Tools.
      • uproarchat 22 minutes ago
        Also loads fine on my beater in the sticks :)
    • swedishPerson1 22 minutes ago
      [dead]
    • jasonmp85 18 minutes ago
      [dead]
  • sreekanth850 33 minutes ago
    I find very little reason to use a pure vector database for enterprise retrieval. We built an enterprise retrieval engine on top of a SQL database with native vector support, and the flexibility is something we cannot ignore. Vector similarity is just one query primitive alongside full text search, filters, joins, ordering and normal relational predicates. Tenant/app/collection isolation becomes part of the query itself. ACLs, document versions, categories, metadata constraints and temporal filters are ordinary predicates rather than something you have to bolt onto a vector store. SQL is already going to be part of almost any enterprise system. Adding a separate vector database introduces another moving part and syncing two system whenever you update your data is the most difficult thing to get right.
  • blakeashleyjr 32 minutes ago
    This sounds like the Postgres vs. InnoDB argument 10 years later. Postings pointed at physical location (the ANN slot), so every SPFresh rebalance rewrote every index touching that doc. InnoDB solved this by pointing secondary indexes at the PK and eating an extra lookup on read. Curious what that extra lookup costs you when it's an S3 GET instead of a B-tree hop.

    "Updating one vector can move hundreds of attributes and their indexes" is basically Uber's 2016 Postgres write amplification post, but for search. Same fix too: stop pointing indexes at where the row lives.

    So ANN becomes a secondary index that points at a doc ID, and vector search now needs a hop to complete. Do clusters keep their own copy of the vectors so the search itself stays local, and only result fetch pays the indirection? Otherwise cold p99 seems like it gets worse.