Are AI Labs Pelicanmaxxing?

(dylancastillo.co)

86 points | by dcastm 2 hours ago

17 comments

  • simonw 19 minutes ago
    This is fantastic

    I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations.

    Catching a lab cheating specifically on my one dumb benchmark would be really funny.

    Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than anything I was considering.

    His conclusion:

    > Nothing jumped out at me. I couldn’t find a case where the pelican-bicycle images looked noticeably better than the rest of that model’s grid.

    • gilleain 14 minutes ago
      Perhaps also vary the bird? Wikipedia tells me pelicans are in the order _Pelecaniformes_ so shoebills or herons might do.
  • mauvehaus 3 minutes ago
    > All 21 pelican-bicycle images, across all seven labs, face right. No other animal/vehicle combination does that.

    > However, facing right is common: 60% of all 1,008 images do it. How common depends on the animal and the vehicle, and bicycles are one of the two vehicles where it’s strongest

    Of course the pelican on the bicycle is facing right. The drivetrain on a bicycle is on the right side. If you want any representation of a bicycle that shows the drivetrain you're going to show the right side of it if you want to do so without the frame occluding it. It's an excellent bet that their training data reflects this.

    Citation: https://www.rei.com/c/bikes

  • dllu 8 minutes ago
    I feel like getting LLMs to spit out an SVG is akin to getting a human artist to draw something by just reciting a list of coordinates. It's insanely hard and unnatural.

    Image generation models nowadays can easily generate a photorealistic pelican riding a bicycle, where the bicycle has perfect structure. But it is, of course, only a raster image.

    It seems that we're missing a kind of step to decompose an image into a list of instructions (say, SVG paths, or even brush strokes with a real brush) to reproduce it properly. Doing so would probably need a true understanding of the structure of the scene, which is something that AI still struggles with to this day.

  • stusmall 33 minutes ago
    I'm glad someone ran the numbers on this. Every single Simon Willison post of an SVG is followed with someone dismissing it saying "I'm sure they train on it by now." This is despite a good blog post with sound logic on how easy that is to catch. [1] Glad to see someone took the time for a quantitative analysis of dumb little animals riding dumb little bikes.

    1. https://simonwillison.net/2025/Nov/13/training-for-pelicans-...

  • scosman 4 minutes ago
    join me in building the ideal training set for pelicans riding bicycles: https://github.com/scosman/pelicans_riding_bicycles
  • apwheele 8 minutes ago
    So this is not my experience at all for asking about simple SVG icons for web-pages. Here is one of the examples I have tried for in the past, make a simple cartoon SVG knife for a map icon for a crime map.

    https://x.com/CrimeDecoder/status/2080008114615537766

    Can see the images for ChatGPT/Claude (Sonnet 5), and Gemini are all quite bad.

    Jagged edge of LLMs. How do you explain being able to generate very complicated shapes in the Pelican example but cannot make a much simpler icon without just alluding to it is in the training data?

  • Wowfunhappy 40 minutes ago
    > The more plausible story is SVGmaxxing

    Exactly--and you have to ask yourself at this point what "maxxing" really means, since "get better at drawing SVGs" is a useful skill.

    • beering 20 minutes ago
      Really awful how the AI labs are skillmaxxing /s

      Pelicans aside, we need to remember that benchmarks are the only good quantitative way we have of comparing models. If someone has complaints about “benchmaxxing”, please ask them to contribute a better benchmark! It is valuable work and very appreciated.

  • simonw 12 minutes ago
  • Rooster61 18 minutes ago
    I find it humorous that the animal + plane combo appears to be such an outlier. I assume this is due to the models assuming the user mean plain and misspelled it in the prompt.
    • NitpickLawyer 9 minutes ago
      GLM has 2 combos of "on a plane" literally sitting inside a plane, with a window and a bit of wing showing. That's funny.
  • jonatron 32 minutes ago
    OK, so we've done animal_vehicle, how about new SVG ideas each time? I just tried "make an SVG of a man sitting in a chair at a computer behind a desk" which gives more interesting results than the animalVehicle test.
    • ninju 18 minutes ago
      There probably good set of images of that description already so it does exercise the inference capability of the model
  • johndough 51 minutes ago
    Another point for consideration: Specialized SVG models create way better looking pelicans riding a bicycle. (E.g. Refract V4: https://jumpshare.com/s/8liB7Aiuoo3yucbWGXjZ mirror: https://postimg.cc/McV70p84 )
    • solarkraft 36 minutes ago
      That’s an impressive image, but what a mistake it was to click the second link (on mobile without an ad blocker). I wouldn’t send it to anyone I respect ...
    • ACCount37 27 minutes ago
      The name is "Recraft V4", and from looking it up: yeah, it sure seems like whatever black magic they use for SVG generation kicks ass.
  • tomas789 46 minutes ago
    Having an objective score is quite difficult. Maybe it would be better to do a pairwise comparison and calculate ELO?
    • NitpickLawyer 6 minutes ago
      Just click through the models. At a glance (and highly subjective) I don't see anything jumping out as oom worse than anything else. I only noticed a model placing the animal inside a plane (with seat and small window) but other than that, they all seem similar inside each model to me.
    • javier123454321 26 minutes ago
      If you want to, go ahead, but it seems to me the author already exceeded the energy expenditure that this question warranted.
  • andy99 57 minutes ago
    If an AI researcher was going to pelicanmaxx, they would almost certainly apply the augmentations mentioned in the article during training, e.g. randomly selecting animals and conveyances. You’d want a model that generalizes well, just sfting in that specific prompt would be pretty bush league for a frontier lab.

    I don’t have any reason to believe they are gaming the benchmark, just saying. I do find the idea of a data labeller having to generate thousands of svgs of different animals on different modes of transportation quite funny though.

    • cute_boi 33 minutes ago
      At this point, I think there are so many pelican images in the pretraining data that drawing a pelican no longer makes sense as a model evaluation task.
  • j45 26 minutes ago
    The models definitely seem to pay attention to the tests.

    Since the tests can be generally gamed with directing descriptions at it non-deterministically, there's a greater chance the questions solution can be found.

    Of course, hopefully the models are instead adding patterns and types of questions as well and it makes the models more capable, but it may be limited in how it transfers to other types of questions in breadth or depth.

  • dcchambers 57 minutes ago
    It's incredible that each model has it's own style that remains relatively consistent throughout all of the different generated examples.
  • cute_boi 35 minutes ago
    https://playcode.io/blog/macbook-svg-benchmark

    I think we should stop using pelican benchmark.

    • dllu 1 minute ago
      I disagree with this in the blog post:

      > Every single one is a pelican, on a bicycle, first try. When every student gets an A, the exam has stopped grading.

      Numerous pelicans and their bikes are clearly horribly malformed. In fact none of the bike frames are correct. Fable and Opus come close, but the top of the diamond is disconnected in Fable's case and the head tube is misaligned with the front fork in Opus's case.

      And of course, as the parent post shows, labs don't actually seem to be training on the pelican bike case.

  • sbseitz 47 minutes ago
    I wish I could downvote this for Pelicanmaxxing lmao.