10 comments

  • giamma 3 minutes ago
  • WithinReason 18 minutes ago
    Mixed signals, here it's performing below even GPT-5.4 Nano:

    https://livebench.ai/

    while here it outperforms Fable by a significant margin:

    https://oxalpha.com/

    but if the latter is true, will people still say it was "distilled" from Fable?

    • sunbum 14 minutes ago
      the 2nd website is not official, just something someone slopped together for some reason.
    • epolanski 8 minutes ago
      GLM 5.3 was a great model, so this would be strange to release a regressed model
      • ImprobableTruth 4 minutes ago
        It's probably GLM 5.3 flash, so weaker but cheaper.
  • esskay 21 minutes ago
    I'd be interested to know what was going on with it during the public test as there were numerous reports of it improving considerably at tasks it was asked to do early on in the test compared to later in it.
    • daveyoung 16 minutes ago
      Two potentials from my pov:

      1. Just variance in pass@K. If you prompt any model multiple times you'll see a large variance. N=1, but I find chinese open source models have a higher variance than higher-RL'd models like fable/opus.

      2. They legitimately shipped a new RL checkpoint over the 7 days, which I find hard to believe.

      I am leaning towards 1.

    • rfoo 17 minutes ago
      lol don't shout out the obvious
  • tosh 13 minutes ago
    my guess is this is a small model punching way above its weight

    on toy benches it made quite a few mistakes but was able to fix all of them on its own

    (meaning more tokens, more turns, more tool calls — but same outcome as gpt 5.6 sol)

  • garo-pro 1 hour ago
    Unfortunately I can't find sources other than this for now but this seems to be legit.
    • mohsen1 2 minutes ago
      > The company on Wednesday confirmed speculation that the Ox Alpha model is a new iteration of its GLM series and said it will release the weights for it tonight, in response to queries by Bloomberg News.

      Seems legit.

      It's really hard to know how good it is. So much hype around it.

  • j_maffe 29 minutes ago
    Anyone has a link to a report of its capabilities? I can't find a reliable source.
    • vblanco 21 minutes ago
      completely vibes based, but ive been using it to port Mindustry game from Java to C# with agents, and its been working for 50 hours (its 15-20 tks so super slow inference). Its done a fantastic work and its almost finished now. Better results than deepseek flash and gpt luna by a mile on this kind of long term work. Less good than gpt sol or opus. We dont know the param count but my guess is 200-300 range.
    • daveyoung 14 minutes ago
      likely a distilled glm 5.3 that will punch within 20% of that at 2-3x less size. you'll find that capability is typically very jagged on models that are distilled
  • dgellow 34 minutes ago
    Do we know the size of the model?
  • daveyoung 5 minutes ago
    [dead]
  • hncsiocp9x 19 minutes ago
    [dead]