What Happens When the Cost of Intelligence Drops 100x

(catalystneuro.com)

59 points | by bkd9 2 hours ago

17 comments

  • Balooga 24 minutes ago
    Jevons Paradox [1]

    > when technological improvements that increase the efficiency of a resource's use lead to a rise, rather than a fall, in total consumption of that resource.

    [1] - https://en.wikipedia.org/wiki/Jevons_paradox

    Las Vegas replaced the expensive incandescent lighting on the strip with cheaper to run LED equivalents. But the costs didn't come down because they were able to add more lights and larger displays.

    I think the same will happen with tokens. As the cost of tokens comes down, these models will just consume more tokens.

    • bee_rider 1 minute ago
      Definitely not going to argue against the Jevons paradox in general, it is observed in various cases.

      The lightbulb thing seems different though? Or at least it is a specific subset. Lights in Las Vegas are sort of an advertisement, right? In the sense that having the brightest or most interesting (or whatever) lights draw attention to your show, casino, hotel, whatever. It’s kind of a zero sum game in that the different shops are competing for the finite attention of a more-or-less set number of tourist. I think part of the Jevons paradox is that society generally finds more useful applications of the newly cheap thing. If the thing’s only purpose is to compete better in a competition with a set prize (all of the tourists’ money), that’s constrained in some way.

    • AndrewKemendo 14 minutes ago
      Turns out Grey Goo was just thermal paste
    • perching_aix 19 minutes ago
      Sounds like a variation on the induced demand principle: https://en.wikipedia.org/wiki/Induced_demand

      Related:

      - Parkinson's law: "Work expands to fill the available time." https://en.wikipedia.org/w/index.php?title=Parkinson%27s_Law

      - Lewis–Mogridge position: "Traffic expands to meet the available road space." https://en.wikipedia.org/wiki/Lewis%E2%80%93Mogridge_positio...

      And I pretty much just plain agree, this is exactly what will happen.

  • andai 22 minutes ago
    > Reading everything becomes the default. At a cent per document, a model can read every paper

    I love how in our day "reading everything" means "the computer reads it for me".

    I expect soon the computer will be able to go on bicycle rides, and spend time with my wife.

    • sssilver 17 minutes ago
      Hasn't the computer already been spending time with your wife?
      • smugtrain 11 minutes ago
        Asymmetric multiprocessing with your mom
      • lubujackson 13 minutes ago
        You're absolutely right!
    • bellowsgulch 20 minutes ago
      Futurama did it first!
  • jbotdev 11 minutes ago
    I think speed is actually going to be a bigger factor than cost. Even projects where “money is no object” often hit a wall with LLM response times.

    Sure you can speed things up with parallel work under subagents, but as with parallelizing traditional computational tasks, there are diminishing gains.

    I keep hearing people saying just change the way you work to trust long-running agents and multi-task more, because they’re too slow to work with interactively for many use cases. I think that’s painful in a world where we expect humans to still heavily guide and interact with agents for their day-to-day work.

  • AnotherGoodName 1 hour ago
    I think a big one is robotics. A robot can today fold your laundry. It takes ~10mins per item. Seriously. It takes a long time to process the image find the corner move the claw to the corner of the shirt and attempt to straighten before folding.

    Robots right now generally move at glacial speeds. You might have seen robots doing flips in semi controlled environments but watch how slowly they open doors etc. processing time is a major bottleneck.

    • Alien1Being 28 minutes ago
      Xiaomi robots do this in double digit seconds.

      Still slow compared to humans, but Chinese robots will be as successful as Chinese EVs, phones and solar panels.

    • segmondy 34 minutes ago
      You must not have been paying attention to development with robots, there are many videos of robots moving really fast in "non controlled environments"
      • airstrike 25 minutes ago
        Yes, I think I saw one in a video titled "Robocop"
      • TheAceOfHearts 19 minutes ago
        Honestly, most of the videos I've seen of robots moving around quickly aren't actually doing anything useful. We've had really impressive tech demos for the past 15 years of robots dancing and jumping around. But I don't want a dancing robot, I want a robot to make me a BLT, wash my dishes, take out the trash, and fold my laundry.

        The most recent video which actually impressed me was a demonstration from Gemini Robotics 2, where a robot was shown autonomously removing the bag from a trash can and folding the loops closed in real time.

        I don't follow robotics advances closely so it's possible I'm just ignorant, do you know any autonomous robotics demonstrations of useful activities that you would suggest checking out?

    • andai 40 minutes ago
      Wait til the robots get on Cerebras, it'll set your pants on fire.
      • LoganDark 36 minutes ago
        Only if the fire manages to escape your wallet!
    • CodingJeebus 31 minutes ago
      Have LLMs improved at being able to process physics-based problems and environments? I remember that issue being discussed around generative gaming a while ago but I hadn't heard much about it recently.
      • AnotherGoodName 4 minutes ago
        Transformers in general are huge but if you say LLM you’re specifically saying the language model transformer. Robots use vision transformers
  • MichaelNolan 35 minutes ago
    100x seems like an underestimate. Even with no model improvements, we should see that sort of reduction. Looking at TSMC’s margins, Nvidia’s margins, and OAI/Anth (alleged) margins on inference, there is a room for a 100x reduction.

    Right now all three of those are at abnormally high levels. Competition will come for all three.

  • brotchie 57 minutes ago
    There’s still 50-500x cost reduction in “this is only an engineering problem” low hanging fruit from specialized chips to run the models + improved distillation.

    Entirely feasible that by 2031, Fable 5 (or greater) intelligence level models will run cool on smart phones, if not sooner.

    • andai 30 minutes ago
      I saw a 1B model yesterday that was fine tuned on Fable output. I found that hilarious, but it did actually make all the scores go up.

      (Actually talking to it, it was about as coherent as you'd expect, i.e. 3/10)

      The floor for "actually usable model" keeps dropping though. (Seems to be about 27B right now?)

  • newAccount2025 18 minutes ago
    I’m loving small models. The gemma4 26/31b models have been deeply impressive on weird prose analysis tasks that I am working on. Nova-micro is really stupid but is extremely fast when it’s smart enough to do something. I’m trying to be disciplined about able to evaluate quality vs cost everywhere for real systems built on this stuff. I probably need to get off Bedrock because it’s missing a lot of other little models that might be good competitors.
  • nchmy 21 minutes ago
    I've been working with the chinese open models for 4 months. They are more than capable for a tiny fraction of the cost of the frontier ones. And yet they also continue to get significantly better and (Deepseek's recent price increase aside) cheaper. Its hard to fathom how the truly frontier stuff will be able to compete long-term.
    • jostmey 19 minutes ago
      And why won’t the frontier models continue to become better? The open models are getting better but so are the frontier models. The frontier models might remain in a constant race to remain ahead
      • HappyPanacea 4 minutes ago
        Diminishing returns on both intelligence and training, mostly
      • ForHackernews 6 minutes ago
        They are running out of novel, clean training data and compute. There is probably a limit to how much improvement can be squeezed out of LLMs. Recent improvements have been more about orchestration and "reasoning" loops (i.e. iteratively feeding context back through the model).
  • LordHumungous 21 minutes ago
    > Reading everything becomes the default. At a cent per document, a model can read every paper in a field, every record in an archive, every email, or every message in a support queue as a matter of routine

    Pretty much already happening

  • andai 32 minutes ago
    A year ago I had an aha moment, when I realized that for my purposes, Gemini Flash was not only 9x cheaper, but 3x faster than Gemini Pro, while producing identical output. Who's the best model now!

    For a lot of tasks, even small models have saturated them a while ago, and then going cheaper and faster is just pure gains.

    For coding I also prefer to do it interactive/realtime, micro-prompting, surgical edits, which the small models can handle just fine.

    And then at the top, the real question is consistency. Not "can they do it" but "reliably enough that you don't need to constantly double check everything." (In my experience, not quite there yet, although it's getting way better.)

  • Multiplayer 28 minutes ago
    This means a great de-risking is happening for the costs of deploying somewhat autonomous agents. This has profound implications on the timeline of deployable personal agents. Cost was a significant factor for many people during the OpenClaw frenzy, specifically when they let their agents run somewhat wild. It will become much more palatable, or already has, to install whatever the next generation of token consuming autonomous systems will be.
  • bdhdhduuyd 31 minutes ago
    Personally I still see LLMs as very advanced search engines which lack intelligence. To me it seems that the cost of getting data is reduced by LLMs, not the cost of intelligence. I mean: we tell the model what we want to achieve, and the model responds with the right data in de form of code in seconds.

    That's why 'stackoverflow programmers' will have a hard time competing with LLMs but engineers are still needed for their intelligence.

    Well that's just my 2 cents.

    • dboreham 10 minutes ago
      I think this viewpoint fails to understand what "intelligence" is. The idea must be that intelligence is some special thing that only humans have. So when machines couldn't do jack s... we said "it's the Turing test". When machines blew through the Turing test we said "that was just prediction..not really intelligence, that's different".

      It's not different. The delusion humans have is that intelligence is special and magical. It's not. It's just nature's prediction machine. A very fancy version to be sure. But not qualitatively different .

      All statements that "oh but it'll never be able to do that" will prove false.

  • swatcoder 15 minutes ago
    The short and real answer is that we drown in slop. As and if compute stops being a bottleneck and the "floor" (as the author calls it) plummets on the cost of inference capability, storage, bandwidth, and human attention become the bottlenecks. Spam, slop, scams, phishing, and inundation are let loose and scale faster than any constructive capabilities can even be imagined, let alone curated.

    It's lovely to think about all the cool things one could do with cheap and capable inference in isolation, and scary to think about what a malicious party might acheive were the capability ceiling to grow unbounded, but the systemic reality we're destined to encounter between those dreams is just the complete saturation of every digital network and networked service with slop and waste.

  • bkd9 1 hour ago
    Author here. I made these plots because I had been searching for them for months and never found quite what I wanted: how the cheapest way to reach a fixed capability level has moved over time. Artificial Analysis publishes enough data to reconstruct it. If someone knows of a source that already tracks this, with historical prices, please share.
    • andai 18 minutes ago
      Thanks for the effort you put into this.

      I wonder if it might drive the point even further if the graph scales were linear? Or maybe the progress has been so great that this would make the graphs unreadable?

      https://xkcd.com/1162

  • smallnix 26 minutes ago
    [dead]
  • promptsphere 49 minutes ago
    [flagged]
  • qsera 49 minutes ago
    Idiocy becomes rampant!
    • perching_aix 2 minutes ago
      [delayed]
    • andai 17 minutes ago
      AGI ≈ Ralph × Infinite persistence

      Just like real life!