Ember-1

(fireworks.ai)

95 points | by gmays 1 hour ago

16 comments

  • intothemild 30 minutes ago
    So they trained a model on open weights, and then aren't releasing the weights... am I reading this right?
    • DonsDiscountGas 9 minutes ago
      It happens. Most open licenses aren't GPL style copyleft.
    • netvarun 20 minutes ago
      Technically kimi k-3 weights license is not open weight (it has a lot of restrictions). I would classify it as ‘weight open’ similar to the bsl and fsl ’source open’ licenses.
      • Evidlo 9 minutes ago
        weight available
  • jamienk 33 minutes ago
    Ignoring for the moment issues of what "counts" as open, won't open models rapidly advance due to stuff like this in ways that it's less possible for the proprietary ones to do? This is exactly how Linux & Wikipedia, for example, overtook their "frontiers", right?
    • andsoitis 22 minutes ago
      > Ignoring for the moment issues of what "counts" as open, won't open models rapidly advance due to stuff like this in ways that it's less possible for the proprietary ones to do? This is exactly how Linux & Wikipedia, for example, overtook their "frontiers", right?

      I suspect the advantage that catapulted Linux ahead of the establishment was less technical potential and talent and more organizational advantage. That's not to diminish the technical talent of the Linux crew, but them being unencumbered gave them more degrees of freedom. The rest is history.

      So as long as the AI companies don't succumb to "big company" dynamics, they can outlead. To wit: Open AI and Anthropic are kicking Google's ass.

      • jamienk 15 minutes ago
        Diff people have diff motives to experiment, then new work is done on top of stuff that "hits" in a way no one anticipated. Then work gets piled on top in a way that might make it hard to port
        • andsoitis 14 minutes ago
          > then new work is done on top of stuff that "hits" in a way no one anticipated.

          Indeed. And when you have freedom to play, you are able to find new stepping stones that you didn't anticipate. And you can combine stepping stones in new ways to make new discoveries.

          Greatness cannot be planned.

    • segmondy 22 minutes ago
      No, because close labs/models borrow but don't contribute back.
    • intothemild 29 minutes ago
      Yes, absolutely, but only if people keep contributing in the open.
      • jack_pp 22 minutes ago
        not necessarily, just knowing something is possible will motivate others to achieve it somehow. Which is why there are so many LLMs and OAI doesn't have a monopoly
  • nico 11 minutes ago
    > The problem: thinking models think too much

    This is partly the appeal of Jev et al; having a quick model for simple tasks, that doesn’t require that much thinking

    It’s amazing all the workflows that models like that can unlock. And yes, classifiers and other ML models have been around for a while for these types of tasks, but Jev has made it easy and cheap to play and experiment. This in turn, is incentivizing people to try them for a bunch of stuff, unlocking creativity and producing a lot of new cool (and eventually potentially very useful) applications

    • demibabs 1 minute ago
      What are the useful applications of Jev so far? Not to sound dismissive, I just haven’t seen what people are using it for yet.
  • andsoitis 35 minutes ago
    > The problem: thinking models think too much

    Analysis paralysis stifles not just human intelligence, but other intelligences too.

  • netvarun 23 minutes ago
    Off topic:With sol pricing drop tbh kimi k3’s value prop has not been that great. For our internal use case/testing/benchmarks sol come out with way better quality and much cheaper costs. Kimi really needs to drop their pricing (I heard it’s set by them across all the neoclouds) Sol is at 2/10 vs kimi’s 3/15
    • nostrebored 1 minute ago
      Agreed, I think the only place where it’s still interesting is ui design. Visually kimi and muse feel much nicer than frontier models to me, but maybe it’s an artifact of everything terrible being Claude Design
    • drob518 9 minutes ago
      Agreed. Even on the open weight side, GLM 5.3 has roughly equivalent performance to Kimi K3 for less than half the cost.
  • tomrod 34 minutes ago
    Well done, and great iteration.

    The pareto frontier needs clearer distinction. Benchmarks miss half the story. What, if any, capability is lost by the token reduction (for example, was it like super awesome at Golang before and now kind of sucks? that kind of distinction).

    • drob518 7 minutes ago
      Unfortunately, it’s hard to make a chart of that.
  • dbuxton 12 minutes ago
    Do they mean Opus 5.5 or Opus 5?
  • themgt 11 minutes ago
    The result? Ember-1 set a new Pareto frontier for Bedside Bench across both open and closed models including GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5 on cost/task.

    "Pareto": 8 hits

    "Opus 5.5": zero hits

    • wmf 3 minutes ago
      Obviously this research was done before 6.0 Sol and Opus 5.5 came out. Your point stands that the frontier moves quickly and small gains can be eclipsed quickly.
  • tdhz77 31 minutes ago
    Does anybody know if this would be a good model for creative writing?
  • logicallee 14 minutes ago
    This is really interesting. I think the Fireworks Serverless Training infrastructure they used to develop it is also unique and needed. Except if someone works at one of a handful of the largest labs, it is very difficult to set up or try any sort of training pipeline. The managed training infrastructure makes it available to more people.
  • erichocean 34 minutes ago
    Need this done for DeepSeek, ideally one of the Flash models.
    • drob518 6 minutes ago
      And GLM. Both Deepseek 4.1 Flash and GLM 5.3 Flash are quote verbose when thinking.
    • atemerev 30 minutes ago
      If you have the compute, I have the expertise.
  • monkey_monkey 35 minutes ago
    I don't think the article mentions Pareto frontier enough.

    Also, did I miss a memo? Suddenly every article on AI seems to be talking about the Pareto frontier - or have I just not been paying attention?

    • DonsDiscountGas 7 minutes ago
      They want it to be the best at something. And it's obviously not the absolute smartest. So here we are.
    • AnodicElegy 23 minutes ago
      I guess they figure "best bang for your buck" comes off a little too colloquial.
  • ls612 32 minutes ago
    On the smaller end, Quen 3.8, while being extraordinarily capable for a small local model, also suffers from extreme thinking. I wonder if the techniques described here generalize to other models too.
    • spijdar 25 minutes ago
      I suspect it might generalize to other large models, but I don't think Qwen3.8 27B is one of them. Kimi K3 is a 2.8 trillion parameter model, and I suspect that is playing a big role in being able to reduce the length of CoT without taking a hit in quality.

      That's just vibes, though.

  • esafak 33 minutes ago
    It looks like it would be similar to GLM 5.3 Flash, had they tested it...
  • justmeeew 14 minutes ago
    [dead]
  • huflungdung 35 minutes ago
    [dead]