Decisions API is in public beta

(developers.openai.com)

45 points | by chiefstorm 1 hour ago

12 comments

  • simonw 9 minutes ago

      curl https://api.openai.com/v1/decisions \
        -H "Authorization: Bearer $(llm keys get openai)" \
        -H "Content-Type: application/json" \
        --data '
      {
        "model": "gpt-6-luna",
        "input": [{
          "role": "user",
          "content": [
            {"type": "input_text", "text": "I am angry about the new product feature"}
          ]
        }],
        "questions": [{
          "type": "predicate",
          "name": "complaint",
          "instructions": "Is this a complaint?"
        }, {
          "type": "predicate",
          "name": "compliment",
          "instructions": "Is this a compliment?"
        }]
      }'
    
    Returned:

      {
        "model": "gpt-6-luna",
        "answers": [
          {
            "type": "predicate",
            "name": "complaint",
            "probability": 0.91
          },
          {
            "type": "predicate",
            "name": "compliment",
            "probability": 0.06
          }
        ],
        "usage": {
          "input_tokens": 310,
          "input_tokens_details": {
            "cached_tokens": 0,
            "cache_write_tokens": 0
          },
          "output_tokens": 0,
          "output_tokens_details": {
            "reasoning_tokens": 0
          },
          "total_tokens": 310
        }
      }
    
    That https://api.openai.com/v1/decisions endpoint is notable because usually when OpenAI define an endpoint like that it ends up as a defecto standard for other providers.
  • Topfi 4 minutes ago
    Ran my decisions evals (still rudimentary, less than 600 calls (UI component selection, chat charting, tag selection, PKM stuff) against Jev and Mercury Decide. Jev because it has replaced my mt0 efforts by sheer force of affordability and Mercury Decide because I do like dLLM efforts (and I'd like to use fewer model providers if possible).

    Preliminary of course, but seems to be slower than Jev and slower similar to Mercury Decides measurements in numbers, though not in growing linearly with the amount of input (346 p50 and 860 p95), less confident concerning my ambiguous UI component and charting assessment tasks (have very specific use cases for these models which Luna often fails to meet at 0.6 and lower confidence), lead to a few failed calls which neither competitor had and more expensive than either to boot by a factor of 3,1 times.

    Basically slower, more expensive and less capable than Jev, roughly on par with Mercury Decide.

    Seems a bit undercooked overall and I'd rather frontier-labs don't jump on bandwagons until they can offer something competitive. In fairness, I have yet to test image input, maybe that makes all the difference.

  • TSiege 11 minutes ago
    The response to Jev should be the nail in the coffin over whether or not the AI business is a commodity market.

    Out of no where Jev appeared as the next round of the price wars. Jev showed the value of System One models. A fast yes/no/confidence score not only is cheaper but also often all people want. Open source versions flood hugging face and now the big players are giving up a potentially big driver of output tokens to keep customers and race to the bottom price wise.

    If I were OpenAI or Anthropic I’d be racing to make their products as sticky as possible bc ppl will flock to what’s cheapest otherwise.

    • gobdovan 1 minute ago
      > If I were OpenAI or Anthropic I’d be racing to make their products as sticky as possible bc ppl will flock to what’s cheapest otherwise.

      Hopefully people will flock to whatever product is making its mission to be commodity and the easiest to replace. Really don't want another free ingress, 100$/TB egress Cloud situation.

  • sidcool 4 minutes ago
    Jev really shook up the industry. This seems obvious in hindsight
  • mohsen1 16 minutes ago
    Since it is fast and understand images, I wonder if it can play video games. I have a harness setup for the LLM play EA FC but even the fastest LLMs are too slow for it. I need to try this with Decisions API
    • binlog 14 minutes ago
      One of the examples on the docs page is it playing a video game. Doubt it’ll be able to run anything complex though.
  • mritchie712 16 minutes ago
    it already supports image inputs, which was the first big gap I found in Jev.
  • esafak 45 minutes ago
    You knew it was going to happen! Benchmarks or it didn't happen.
  • lab14 57 minutes ago
    How is the pricing vs Jev?
    • jerrygenser 19 minutes ago
      $0.10/mm input vs. $0.042/mm input. Both free output.
  • peterson_lock 17 minutes ago
    Can we use this through subscription?
  • Imustaskforhelp 1 hour ago
    This rather didn't take long for OAI to create*, I remember people giving opinions and discussions that it won't take too long and that openAI should do it[0], so looks like they were right.

    Interesting to see where all this leads us and if other major labs follow suit

    Edit: decisions voice looks really interesting as well[1]

    [0]: https://news.ycombinator.com/item?id=49802161: OpenAI is well positioned to fast-follow Jev

    [1]: https://developers.openai.com/api/docs/guides/decisions-voic...

  • zane_shu 33 minutes ago
    [flagged]
  • dvt 12 minutes ago
    I genuinely do not understand why anyone would pay OpenAI for this. Running something comparable to Jev is pretty trivial. The whole point of paying for ChatGPT is because OpenAI has a bunch of warehouses that can run a zillion-parameter model.

    Running a decision model is way easier and much cheaper. Are they really just trying to capitalize on the hype here? It feels like they really have absolutely zero moat.

    • mediaman 9 minutes ago
      Why would I run it myself? It's $0.10 per million tokens. Dirt cheap. (Jev is even cheaper.)

      You could ask the same question about why anyone would rent a VPS. I can just run my own hardware, it's just a computer!

      Buy vs rent is not just about what's possible, it's about what's economic.

    • csharpminor 3 minutes ago
      If you're in an enterprise that already has a procurement agreement with OpenAI, this means you don't have to onboard another vendor. Bucket platform strategy.
    • simonw 8 minutes ago
      Depends on the quality of the results. These things are driven by text prompts. If it turns out the OpenAI one returns better quality results than open weight variants they'll be rewarded by the market.

      Anyone using a decision model like this is going to have to spin up their own evals - these are far harder to vibe-check than regular text output LLMs.

    • TSiege 8 minutes ago
      There isn’t a moat in the sense of self hosting but you need a reason for people who don’t want that to stay on your platform. Customers save time and effort managing payments easier this way. However it’s a race to the bottom price wise.

      Going to be all about branding and platform stickiness for OpenAI to make investors and creditors whole.

    • super256 7 minutes ago
      [delayed]
    • tmhall 8 minutes ago
      For my use case it will cost like $11 a month and we already have OpenaAI keys and accounts with billing in place. I don't want to run my own model infra and I don't want to get permission to set up an account with typesafe.ai
    • jcims 6 minutes ago
      If you work for a company that has a 3 to 6 month onboarding period for new vendors and a lifetime commitment to maintain a whole bunch of vendor management horseshit for as long as that relationship exists, it makes a ton of sense.

      Add in a bunch of model governance and oversight for anything you train yourself and it’s pretty much a slam dunk deal.