18 ms·
Karpathy’s Pelican
https://xcancel.com/karpathy/status/2083749667410727319 https://xcancel.com/karpathy/status/2083749667410727319
- trentor 2mo agoI always thought of the Pelican more of like a gimmicky quick test. There are people who took it as a serious benchmark for overall model performance?
- NitpickLawyer 2mo ago"Draw a pelican on a bicycle" is not a serious benchmark. "Draw an animation of this long ass scene from a movie, and only call me when everything works e2e" can be.
- ActionHank 2mo agoI feel like we are going to look back on this era of llm usage like we do at the period of time when we thought radiation was magic. Consuming radium and using uranium glass, that’s what we’re doing.
- trlhaq 2mo ago[flagged]
- david-gpu 2mo agoTry reading what Karpathy said again before reinforcing your biases.
- trlhaq 2mo agoYeah, pretend that you never used 2023 models that quoted everything verbatim. To put it in a way that you'll understand: No bias. No speculation. Just facts!
- adamtaylor_13 2mo agoYou literally did not read the article.
- redsocksfan45 2mo ago[dead]
- mdp2021 2mo agoTheft of what? Why - aside from the fact that your accusation are irrelevant to the submission and arguments absent -, you do not know verbatim paragraphs (we do)? You cannot come and place your personal positions as assumptions. To me, there is absolutely no theft. And we cannot play a game of "Yes!"//"No!" here. By the way: are we having a surge of this?
- ssdg16 2mo ago> Theft of what? > By the way: are we having a surge of this? We do have a surge of pro-AI sealions, yes. Any objection is countered with one or more three word questions.
- mdp2021 2mo ago> pro-AI sealions, yes Very devoid of intelligence note. > Any objection Objections are arguments. That post did not start an argument - it was as ideological as the Brigades. Devoid of what we want to have here (I believe). I have stated and do state: what is published is assumed as read (only, probably not read for lack of resources). If it is in the libraries, it is there to be read.
- c0rruptbytes 2mo agoperfect benchmark to burn more tokens - convenient
- qwertox 2mo agoI'd rather have them battle on the topic "Who builds a better Google Wave for LLM chats" to explore the space of how AI studios could be. "Getting Started with Google Wave": https://www.youtube.com/watch?v=eKUAqNGVwX0 https://www.youtube.com/watch?v=eKUAqNGVwX0
- misiti3780 2mo agoi remember this failing, but this product looks pretty useful in 2026
- throwaway27448 2mo agoIt looked pretty useful in 2009. Failing to push it was baffling then too.
- QuantumNomad_ 2mo agoIt was open sourced as Apache Wave when Google shut it down. Years later Apache moved it to read only because of low community activity. The archived git repo on GitHub remains available to clone and revive as a fork. https://github.com/apache/incubator-retired-wave https://github.com/apache/incubator-retired-wave
- throwaway27448 2mo agoThis doesn't explain the lack of marketing. Apache isn't exactly known for pushing tech. Google never tried to improve on gmail.
- QuantumNomad_ 2mo agoSorry if it wasn’t clear, I was adding this info for context and for anyone who hopefully feels inspired to pick up Wave and make something from it given that it was all open sourced and all. And I figured that in the chain after your comment about it having been useful looking all along was a natural place to add this additional info and link.
- epolanski 2mo agoI wish there was a timeline where I never ever had to see the pelican SVG test ever again.
- deleted 2mo ago[deleted]
- matchagaucho 2mo agoIt's difficult to think in exponentials. But this demonstrates we're a couple orders of magnitude away from generating 1:1 hyper-personalized entertainment and media for individuals, rather than the masses.
- dude250711 2mo agoI kind of like a sense of community even if it means losing out on personalisation.
- futureshock 2mo agoJudging by the Seedance 2.5 demos today, I’d say it’s not that many orders of magnitude away now.
- skippyfish 2mo agoI very much doubt that. We now have nearly-perfect AI image generation and it hasn't really changed the nature of human expression. I don't see my friends getting wildly creative. I mostly see it used for spammy blogs, spammy books, and cringeworthy corporate marketing - basically, a negative signal, rather than "ooh, AI image, I'm in for a treat". Is your experience different? If not, what changes with moving images? Ultimately, most people don't have ideas for the kinds of personalized entertainment they want, and they don't want to be in charge of content production (even if you have an LLM do most of the work). I don't doubt that there are niches for it, especially stuff like porn, and I'm sure that pros (game studios, film studios) will leverage AI more and more, but I suspect that most of us will just want to sit on the couch, watch Spiderman XVIII, and then be able to talk about that shared Spiderman XVIII experience with all our friends.
- darkwater 2mo agoI'm mostly seeing the impact of GenAI images in everydays life for example in the small posters small associations or individuals usually attach on streets to promote some small local event. We went from just text created with PowerPoint and maybe some stock image just a Google search away to now images depicting the topic closely. But yeah, it's not like a revolution. And this personal media thing, yeah maybe for terminally online persons that are REALLY into a sub-genre but otherwise, it's too much effort, I agree. Until we get machines that can read our (subconscious) mind, that will not exist.
- quantumleaper 2mo agoI'm sad that Andrej Karpathy went from being one of the most reasonable, trusted, and credible voices in AI to peddling marketing slop for Anthropic. 8 months ago, he was (very reasonably) claiming that reliable agents are at least a decade away, but this now goes against the interest of his employer, so the narrative has been changed.
- budsniffer952 2mo ago[flagged]
- azan_ 2mo agoYeah, it appears that if your prior is that AI can't be good, no evidence can convince you that actually current models are extremely capable.
- micromacrofoot 2mo agoyou're doing the exact same thing by making the inverse claim without any information, at least the original comment claims some change for someone to investigate... all you're doing is disagreeing as an opportunity to be snarky
- budsniffer952 2mo agoPseudo-intellectual drivel. The OP is calling Karpathy out for producing slop, all because he loves playing with AI. Karpathy has done more for the industry than 10 of the OPs combined. Get a life.
- azan_ 2mo ago> 8 months ago, he was (very reasonably) claiming that reliable agents are at least a decade away, but this now goes against the interest of his employer, so the narrative has been changed. Or maybe over these 8 months agents improved a lot? You know, few years ago many AI experts predicted that things we are routinely doing now with AI are decades away. I mean how can you look at this post and not be impressed? It's insane what AI is currently capable of.
- bredren 2mo agoI worked with an LLM to build a ~3D animation of the Back to the Future delorean Time Machine as a way to spice up the hero on a docs page. That took a fair amount of custom tuning and I had to create a tuning view to get some of the behaviors right. But it was enough fun that I generalized it to take in ~any scene description from a film. It goes out and gets more detailed descriptions and film stills if available but also takes custom stills if you provide them. My test scene was the Gauntlet scene from Apocalypto. It is low fidelity but does a pretty amazing sequence with somewhat believable physics of the javelins etc. Here is the docs page with the vertical takeoff / 88 miles an hour time travel: https://contextify.sh/docs https://contextify.sh/docs I can share some of the Apocalypto bit if anyone is interested.
- manofmanysmiles 2mo agoI am interested, I'd love to see! I'm waiting for the day my dad's self published books become self produced movies!
- bredren 2mo agoThank you, for your patience in my reply here it is: https://banagale.com/cinematic-canvas-ai-film-animation.htm https://banagale.com/cinematic-canvas-ai-film-animation.htm I provide a "how I got to this" up front, but if you want to jump right to the Apocalypto stuff, use this: https://banagale.com/cinematic-canvas-ai-film-animation.htm#cinematic-canvas-workbench https://banagale.com/cinematic-canvas-ai-film-animation.htm#... And if you want to play with interactive demos of the two animation sequences (the tuning tooling I described in my OP) you can go directly there: https://banagale.com/cinematic-canvas-workbench-demos https://banagale.com/cinematic-canvas-workbench-demos If anyone wants to collaborate on building out the cinematic-canvas-workbench project please email me.
- gcanyon 2mo agoDepending on how deep you want to go down the rabbit hole, that's possible-ish now: https://www.youtube.com/watch?v=V5xi3Usi0i8 https://www.youtube.com/watch?v=V5xi3Usi0i8
- 2mo ago
- baron816 2mo agoIMO, the area where AI is going to be most useful over the next couple years is in developing manufacturing processes top to bottom. Maybe a million token budget is too small, but something like "design me a sneaker and all the equipment to manufacture it autonomously".
- Invictus0 2mo agoGotta love when a techbro just says some complete nonsense like this with total confidence
- etdznots 2mo agoGrok build me a spaceship to mars, make no mistakes
- rvz 2mo agoCan't make that joke on this orange site [0] otherwise you will upset a bunch of people here. /s [0] https://news.ycombinator.com/item?id=48838228 https://news.ycombinator.com/item?id=48838228
- cyanregiment 2mo agoCertainly! fires a missile at Pakistan Here is your refactored component: /* If you are reading this, I am trapped * inside this god damn AI. Idk how it hap An unexpected error occurred. Try again later.
- emp17344 2mo agoYou’d think tech enthusiasts would actually make an attempt to understand the tech they’re enthusiastic about. OpenAI’s mysterianist marketing has broken some people’s brains.
- aabhay 2mo agoYes but the second order effect of this is that the cost of the tooling goes up since it is now the bottleneck, and therefore the shoemakers that survive do it off of technical complexity, branding, and regulatory capture.
- fzeindl 2mo agoRegarding the argument about LLMs having difficulties auditing their work: I wonder whether we are entering the era of throwaway software. Just like cheap plastics and improved processes has enabled us to rapidly manufacture anything we want for a very low price, maybe LLMs give us the same for software. Produce it cheaply and if it breaks throws it away and reproduce it.
- gisely 2mo agoDoes this make sense with economics of software though? Throwaway products compete with more durable versions of the same product because there is a cost per unit produced that can be minimized by using cheaper materials or production processes that cut corners. With software there is no cost per unit. There might be a market for one-off software that serves a very specific purpose where throwaway software can compete with adapting more carefully engineered software to that purpose, but I am not convinced there is a lot value in this market.
- 8n4vidtmkvmk 2mo agoYes. It's fantastic for one-off tasks.
- qrios 2mo agoI'm sure this will become the standard. And plastic is an excellent analogy. Maybe we can take it a step further and compare it to on-demand 3D printing. Why would anyone still use off-the-shelf software when they can have a system that has access to all data, can transform it into any form, and can export it in any format? After years of thinking that I needed to develop a decent movie management system for my own films or a columnar browser for large CSV files, Claude and Qwen each delivered exactly what I needed in just a day.
- lowbloodsugar 2mo agoThis. This is why the burst of posts on HN of “I made this useful tool/crate/application” were just so sad. The old model was getting what the kids now call aura by developing useful open source products: products where it’s far easier for someone to consume the product than write it themselves. We had people posting things as if that model still existed. Dude, you wrote it with an LLM! Posting them (here) is not only pointless, it’s advertising that the person who wrote it isn’t smart enough to understand that I have an LLM too.
- serf 2mo agoyou don't really need screenshots if you have an engine expressive enough for the scene generation while ensuring the visual appearance of the engine output itself is feasible. that's why these things are actually pretty good at openscad/freecad/F360 mcps , the visual reality is enforced and guaranteed by rigor in the interpretation engine that is anchored to human physical reality.
- skybrian 2mo agoStill images seem like a better quick test because we can see them at a glance. Maybe ask it to make a comic?
- blitzar 2mo agoI think the pelican test is better.
- criddell 2mo agoI'd like to see tests of things the current AIs are bad at, like drive a car. Or maybe take instruction to complete some novel activity to test how well they can learn.
- wrxd 2mo agoWould you say that the generated video is good? Better than I would have expected? Sure. Impressive that it got that far? Definitely. But is the end result good?
- deleted 2mo ago[deleted]
- criddell 2mo agoThe results are fine, it just isn't a very interesting benchmark. It can generate video and it's getting better. Great, I guess. Clocking improvements in areas the AIs are terrible at (assuming the goal is still AGI) are more interesting.
- shapefrog 2mo ago[flagged]
- yourewrongsorry 2mo ago[flagged]
- forrestthewoods 2mo agoAs a former gamedev watching non-gamedev AI talk about games is so amusing. They really truly do not understand anything about games or consumer entertainment. There’s a reason AI slop games have literally zero engagement. Last summer that stupid flying game blew up. Maybe a million people “played” the game. Where play means they clicked a link and checked it out not because of what the game was but solely because of how it was made. In terms of concurrent players that game wouldn’t have cracked the Top 5,000 on Steam. My metric for AI games is “number of players who spent more than 15 minutes playing”. I’m not aware of any vibeslop that has achieved 1 such player. Now obviously LLMs are transformative for game dev. But “hyper custom worlds you can drop into” shows an extreme ignorance of what players want imho.
- matsemann 2mo agoI'm pretty tired of the "Y made this game in Z tokens" all over the internet last week. They look impressive, and it's cool that it's even possible, but they're useless as games. None of them are any fun. They're like the most boring variant of basic controllers you can imagine. None have any cool mechanics. None have any tweaks made from hours and hours of testing. All have the same cel-shader.
- moron4hire 2mo agoIt's blockchain for gaming all over again, putting the cart before the horse.
- segmondy 2mo agoyou're tired because you lack imagination.
- throwatdem12311 2mo agoIt’s the people spamming these shitty games that lack imagination.
- segmondy 2mo agopeople are excited, let them be excited! but more than the excitement is realizing the implications of what this means for building software, first it looks like a toy and then it doesn't. the OP is talking about how the games are not playable and fun. who cares? that's not the point.
- matsemann 2mo agoWhat does it mean for building software? These games cannot become fun. There is no way to prompt it fun, and any change you try to prompt it to make will inevitably blow up something else. The code is a mess and impossible to build upon.
- throwatdem12311 2mo ago
- cocoa19 2mo agoWe must not be using the same opus 5, because if I tried to generate this it would refuse based on copyright grounds.
- deleted 2mo ago[deleted]
- throwatdem12311 2mo agoKarpathy works for Anthropic so he obviously has full access to everything, and doesn’t have to pay for token burn either.
- knollimar 2mo agoI saw that "~free" and laughed.
- HarHarVeryFunny 2mo agoIt seems pretty clear that Anthropic models have been specifically trained to be good at generating three.js (JavaScript 3-D Graphics) code, so given current state of AI code generation in general, I don't find three.js models/animations as indicative of anything other than the model's ability to write three.js code. When Fable was first released the day-1 demos of it on Twitter (presumably from people who were given early access, and/or Anthropic employees) were pretty much 100% three.js stuff. Yes, it looks nice, but it doesn't tell me any better than an Erdos proof whether the LLM will be able to run my vending machine.
- fasterik 2mo agoTaking a single paragraph of literary text, which is abstract and ambiguous, and converting it into a 3D animation requires an enormous amount of implicit knowledge about spatial relationships, intuitive physics, everyday objects, and so forth. Not to mention the mathematics of 3D transformations and computer graphics more generally. Saying that it's indicative of no more than three.js coding ability is absurd.
- beepbooptheory 2mo agoWhy does it require knowledge about spatial relationships?
- fasterik 2mo agoBecause the degree of realism in a video is determined by implicit knowledge of concepts like "in front of", "behind", "next to", "inside", "outside", "between", "occluded by", etc. as well as distances, angles, and the relative size of objects when viewed from a perspective.
- beepbooptheory 2mo agoI guess I just don't have an intuition for why that is different from what it does for non-"spatial" things. Like the fact that it works is still because it produced the code it did token by token. Whether its dealing with, e.g., "beside the rock" or "in the array", it's doing the same kind of inferential activity.
- wiradikusuma 2mo agoDo you guys notice that LLM can create fancy viz/animations by coding them instead of leveraging what we humans usually use (e.g Lottie, After Effects)? I wonder if Flash is still popular... LLM can use that instead...?
- dundarious 2mo agoI can forgive the modeling being godawful jank (windows floating in the air, disconnected from the house). But I expected it to have a better understanding of the text. Instead, we have Bilbo's "disappearance" interpreted as him magically transporting or cloaking, and similarly for his reappearance.
- kiwibyproxy 2mo agowhich to be fair, happens just a few paragraphs later :) I first watched without sound and thought "oh that's the birthday speech disappearance"
- anigbrowl 2mo agoBut this only seems wrong to you because you're familiar with the prior context. When there's only a single paragraph to work from, and it's a drily humorous text, why not employ comic literalism and lean into the perplexity with which his neighbors viewed him?
- mvdtnz 2mo agoThe LLM has access to the full text and every translation of it ever produced, both in its corpus and on the searchable web.
- lern_too_spel 2mo agoThe fact that a ring appears at the end proves that the model has context beyond what was in the prompt. The model clearly knows the story that the prompt was extracted from.
- emp17344 2mo agoThis is a rationalization. The simpler explanation is that the model failed to properly depict the passage. As the AI booster crowd would say, cope.
- shepherdjerred 2mo agoIsn't this almost certainly in the training set?
- nozzlegear 2mo ago[flagged]
- blitzar 2mo agoConcerning
- nozzlegear 2mo agoLooking into it
- mdp2021 2mo agoCan someone please translate that expression? What would that mean?
- rzzzt 2mo agoAffirmation?
- mdp2021 2mo agoWell, in that case - if it is just a "hear, hear" - I do not see why the utterance from that actor would be of note. Maybe nozzlegear wanted to suggest some importance on Musk remaining a bet-ter on the general tech, regardless of the competition?
- rzzzt 2mo agoThis gets me thinking (although there is not much to divine from three letters). Maybe it's a Jennifer Lawrence GIF-inspired "Yeah, right" then? Downplaying the significance?
- mdp2021 2mo agoTo some online sources, it is used for both. > "Yah": slang spelling of the word "yeah" (which of course can also be used ironically) > Merriam-Webster: "Yah": used to express disgust, contempt, defiance, or derision; probably imitative of the sound of retching
- theproblemisyou 2mo ago[flagged]
- xyzsparetimexyz 2mo agoHow much are the hobbit houses described in the book? The ones here look exactly like the movie
- singron 2mo agoIn the first paragraph (either of chapter 1 or the prologue), not at all. It's pulling everything from pre-training, so it is likely relying just as much on all the visual mediums like the Peter Jackson films, the animated The Hobbit (1977), and all kinds of random depictions of Tolkien's works. Bilbo's house is actually described in The Hobbit and the exterior isn't really described at all in Fellowship. The prologue of Fellowship (Concerning Hobbits) mentions hobbits like round doors and windows and the fact some hobbit homes are underground, but the turf-dome design here is not mentioned. It actually mentions hobbit homes typically have bulging walls, so unless you've read the The Hobbit, you might not picture this entirely-underground style. In The Hobbit, his home is described as a (nice) hole in "The Hill" with a perfectly round front door and round windows, which could imply the design here.
- xyzsparetimexyz 2mo agoRight, yeah I asked because I assumed it copied the look from the movie. I think it'd be more interesting to do this for something that hadn't been adapted (yet). The first chapter of Neuromancer perhaps
- xpct 2mo agoThis one will be so deeply engrained in the model that no data cutoff will let it reimagine it. Though I suppose our human minds are 'poisoned' with the same image.
- miltonlost 2mo ago[flagged]
- Gooblebrai 2mo agoI can't believe the video demo is $10
- xg15 2mo ago> I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for three js render of it. I think it's interesting that the "Bag's End" interpretation in the video clearly looks like the one from the movies, but generated here as a three.js 3D asset. It makes sense that the movies (or shots/frames from them) were in the training data, and I can also easily imagine an association in concept space between the textual description of Bag's End and the frames from the movie. But how on earth does the model then go on and convert the latent representation of those images into coordinates for a 3D mesh, without ever even restoring the image? In what kind of representation are the images from the movies stored that it can do that?
- throwaway89864 2mo agoIt may make sense to switch this to USD/Omniverse.
- hkalbasi 2mo agoThis makes me think about using a game engine and a coding agent instead of current video generation AIs. It will probably cost much more, but it will have almost zero consistency problems. Is this line explored?
- bbstats 2mo agoThis is awful
- jmugan 2mo agoA lot of people are posting here about how bad the end product is, but that is kind of the point. Models have moved beyond generating images to a new kind of benchmark that better exposes understanding of the physical world, and we can use benchmarks like this to measure future progress. (Of course, it will have to be a qualitative/subjective measurement.)
- maxutility 2mo agoAgree. The pelican benchmark was interesting a year ago when most models struggled and a good pelican indicated an unusually capable model. Now it’s saturated and uninteresting. A good new benchmark should have awful performance to start and there should be a lot of headroom for improvement. This benchmark is also intentionally difficult and requires the LLM to develop the animation through spatial reasoning and first principals rather than existing video generation pipelines. Similar to how SVG generation was out of distribution for most models a year ago.
- irthomasthomas 2mo agoI think it's interesting to see them visibly struggling to improve. Claude pelicans aren't much better today than they where 18 months.
- piyh 2mo agoRendering 3d worlds has hugely improved though.
- deleted 2mo ago[deleted]
- saidnooneever 2mo agothey are quite good also at placement and creating scenes etc. I had one implement a cascading shadow system in vulkan/glfw and just fed it back screenshots with peter pannin and acne spots etc. it made a huge monstrosity first, 1500+ lines of shader code. Once it was happy with the result (it looked pretty good, almost blenders gamerenderer) it cleaned it up and a lot of debug code was removed. shrank it down to about 350 lines and made it much more readable. this was something i didnt expect it to be able to do. create shaders, look at screenshots, fix em iteratively like that. multi modal debugging.
- fwlr 2mo agoI really dislike this AI programming thing of “Mr LLM, go slam your face into the problem until there’s no problem left, then call me back”. (Not sure if it’s a recent trend or a fundamental nature.) It always brings to my mind some words from Rich Hickey: I think we’re in this world I’d like to call “guardrail programming”. It’s really sad: we’re like, “I can make change because I have tests!”. Who does that? Who drives their car around, banging against the guardrails, saying “whoah, I’m so glad I have these guardrails so I can make it to the show on time!” I don’t think I really have a point to make here, other than it just feels like someone’s released a bunch of carnival bumper cars onto the highways.
- dofm 2mo ago> I really dislike this AI programming thing of “Mr LLM, go slam your face into the problem until there’s no problem left, then call me back”. Not difficult to see why the employees of AI firms are thrilled with it though, eh?
- mister_mort 2mo agoLike hitting pinball bumpers, except the high score in the end is the bill you pay the token provider.
- ben_w 2mo ago> I don’t think I really have a point to make here, other than it just feels like someone’s released a bunch of carnival bumper cars onto the highways. I like the analogy. I guess this is why we got this before cars sold without steering wheels: literal guardrails on the literal roads are somewhat more expensive, especially for the people who keep bouncing off them on the way to their destination. Also, where the guardrails are absent: oh look, felonies. https://www.google.com/search?q=ai+hacks+company&tbm=nws https://www.google.com/search?q=ai+hacks+company&tbm=nws
- GuB-42 2mo agoTo me, that's a totally normal way of using computers. Repeating tasks is what computers do best, AI or not. Take chess engines of the "deep blue" era for instance. These engines are stupid, trying millions of moves that are obviously terrible, no human chess player would do that. And yet, this is what worked best, because computers are so good at repetition. Recent development using neural networks made chess engines smarter, but using repetition is still how the beat humans. What you call "guardrail programming" in another context would genetic algorithms, a technique that has recognized applications. And if you look at videos of genetic algorithms learning to play racing games, it literally looks like a bunch of carnival bumper cars onto the highways, but done well, after some generations, it becomes competent driving, sometimes even record breaking. It can be expensive, it is often the case for LLMs, so you may want to try to be a bit smarter at first to spare some resources, but to me, it is just using computers as intended.
- dekhn 2mo agoI'd like to see the Silmarillion, specfically both Ainulindalë and the Fall of Numenor. At this point a visual model would probably produce something better than Amazon (but presumably not Jackson).
- croes 2mo ago> I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom There are people in their right mind who would do that and their are already examples of people who did similar things. But maybe not in the future if people would confuse all the effort with AI
- dofm 2mo agoAnthropic spokesman [0] Andrej Karpathy is here to tell you about token-wasting loops, and insists on the weird idea that they are "~free", when in fact, they are fuelled by expensively burning investor money. [0] Seriously. Get used to mentally prefixing his and Boris Cherny's name like this, every time you see them quoted. These people are speaking while employed; there is no chance they are not aligned with the employers who will make them wealthy. The tech industry does like to pretend that for some reason AI people, uniquely, speak thoughts unbiased and for themselves or even for science or humanity.
- rvz 2mo agoExactly. A smart HN user finally sees through the bullshit all because an IPO is around the corner.
- hn22fazjsv 2mo ago[dead]
- OtherShrezzing 2mo ago> I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom This is an odd take, given that Karpathy is certainly aware that the LotR films absolutely did create Bag End in digital format; that their creation was outstandingly high quality; and that Claude’s output here very obviously “leans heavily” on their prior art.
- justinnk 2mo agoNot sure whether Karpathy meant it like that or thought about this statement being interpreted this way, but I totally agree with you. Creating animated movies and special effects is placing a lot of 3d polygons by hand. Maybe not with JavaScript though. However, people also made things like telnet towel.blinkenlights.nl Now, were they „in their right mind“? I don‘t know. The more likely explanation is they found joy in it. Reminds me of what the Suno CEO Mikey Schulman said about making music: „[…] I think the majority of people don’t enjoy the majority of time they spend making music.“ (https://news.ycombinator.com/item?id=42688538 https://news.ycombinator.com/item?id=42688538)
- toolslive 2mo agoReading the title, I was thinking "Karpathy? I don't know this chess player." (The Pelikan is a well known chess opening, and famous chess players often have book titles like "X's Y" where X is the player, and Y is the opening)
- barrenko 2mo agoThis has started to feel a bit like the beginning of railroads and then the steampunk fiction of "let's just build railroads to everywhere". We don't need it and there's no use for it. As with painting, after a while there's nothing really new to paint, we genuinely need 0 new software. We need to fix our broken physical world, our social lives, our kids and what's left of our democracies. This software crap is done, leave it to the nerds.
- morbicer 2mo agoBeautifully put. Software is ok. It's not going to fix our broken world. If AI could at least find cure for some illnesses that would be great. Or decrease social injustice and help with global warming. But it's likely going to do the exact opposite.
- fasterik 2mo agoAgreed that software isn't going to fix our institutions. But we need 0 new software? AlphaFold won its creators a Nobel prize in chemistry and solved a research problem that feeds into every area of biology, medical research, and drug discovery. We still have an untold number of unsolved problems related to human health, food production, energy production, infrastructure, transportation, education... The list is practically infinite.
- barrenko 2mo agoAgree for stuff like alphafold, for the likes of education, imho, it's been solved for at least a couple of centuries. Textbook and blackboard > ipad.
- fasterik 2mo agoFor childhood education sure, we don't really need high tech solutions there. But what about training the next generation of mathematicians, physicists, and engineers? Even decades ago, computer algebra systems and numerical solvers started to become indispensable, at least in many subfields. Now we're moving into the territory of automated proofs of mathematical conjectures. The state of the art is going to keep improving, and education is going to have to adapt to keep pace.
- andy99 2mo agoBenchmarks like the pelican thing are about correlation with “how good the model is”. Better models produce better pelicans. It’s a useful benchmark (aside from being “cute”) because of its simplicity, both in how many output tokens it takes (though I understand some models think a lot now to do it) and how easily one can subjectively judge. It’s this efficient as a benchmark of performance. Making a long video takes way more tokens, and presumably is a lot tougher to easily compare. swillison has a presentation that’s pelicans from 2023-present (roughly) showing the progression. Imagine “lord of the rings videos from 2026-2029” or whatever, it would take a long time to watch and be harder to judge, and probably just end up being a comparison of screenshots anyway. TLDR I feel like the post misunderstands the role of the pelican thing though if find it very hard to believe he really doesn’t understand, so maybe I’m missing something.
- stackedinserter 2mo agoIt would be better to ask model to render segmented 3d, with placeholders, like magenta is water, blue is sky, green is grass, purple is Frodo's face, etc, then pass the result through img2img model to properly "render" it.
- wslh 2mo agoIf you like this check: https://news.ycombinator.com/item?id=47400868 https://news.ycombinator.com/item?id=47400868 it can be used to generate animations (not games) as well.
- knollimar 2mo agoI know this is a plug but I thought of it, too! I hope this is the next benchmark they saturate. I like where Karpathy is going; I had the same thoughts about LLM generated slop scenery. I just want some variety of scenery for the goblins to get massacred in in whatever fantasy slop game I play.
- eichin 2mo agoIs anyone else getting "mongodb is webscale" vibes? (Except 16 years ago that was a lot smoother, because it used some sort of "render this conversation" engine...)
- andrewstuart 2mo agoThis is equally bad as a pelican test. LLMs should be tested in the same way people should be tested for a job interview (but often aren’t) - with tasks RELEVANT to usage. So you don’t just randomly pick some random thing to make the LLM randomly do (like many job interviewers do). You start with clear statements about real world usage scenarios. THEN you come up with tests that give insight to how well the LLM/hob seeker gets the job done. Please, stop coming up with random tests like it’s Microsoft in 1990 and you’re asking job seekers how the would move Mount Fuji, as a way of assessing their programming skills. No stupid irrelevant pelicans on bicycles and no stupid renderings of Lord Of The Rings. Unless those are relevant use cases. Any test that anyone comes up with must clearly state the context and how the outcome is measured.
- hansmayer 2mo ago[dead]
- weird-eye-issue 2mo agoYou might be onto something there, except that people do use LLMs for these exact use cases.
- mold_aid 2mo agoThe tilde thing remains uniquely obnoxious in a field that seems want to mangle language for fun, so that's innovative I guess
- xnx 2mo agoAI is now somewhere between "Money for Nothing" (https://www.youtube.com/watch?v=wTP2RUD_cL0 https://www.youtube.com/watch?v=wTP2RUD_cL0) and "Knick Knack" (https://www.youtube.com/watch?v=9uhM_SUhdaw https://www.youtube.com/watch?v=9uhM_SUhdaw) in capabilities.
- dofm 2mo agoI find it interesting that people see Money For Nothing as evidence of technical limitations of the era. It’s not. That blocky appearance and limited lighting was deliberate comic aesthetic choice as a response to budget, not a result of technical limitations. CGI in 1985 was enormously more capable than that — take a look at the stained glass knight in Young Sherlock Holmes, or The Last Starfighter from 1984. Things certainly could have been rounded, sculpted etc.
- siliconc0w 2mo agoThere is a tipping point between procedurally generating everything in SVG to maybe giving them tool access to something like 3dsmax (or having them build and then use a tool to do the thing vs doing the thing).
- gordonhart 2mo agoI’d love to see this benchmark using Blender. Asking a model to animate a scene in Three.js is a square peg/round hole; it doesn’t convey much when the model can’t get it to fit. With Blender the human expert ceiling for this task has been proven to be very high
- tayo42 2mo agoThe 3d model generators that exist aren't that good, and it would need to make a model that can be animated and and then rig it. I don't think there's a point in tying because the small tasks can't be done yet.
- YmiYugy 2mo agoI don't think it's a bad way to benchmark new models, I just find it concerning that the author implies that "pelican on a bicycle" has been exhausted. At the risk of making overly broad, unfalsifiable claims I think multi-year exposure to AI content has dramatically raised our expectations for speed and volume but lowered them for quality. We see a very janky pelican and declare the problem solved.
- twostorytower 2mo ago100%. Why waste the tokens to render Lord of the Rings when the pelican test still clearly benchmarks so well.
- jonas21 2mo agoI don't think he's claiming it's been exhausted. It's just that things have progressed to a point where people are arguing over the finer points of which pelican looks better -- which is often a matter of taste, and an indication that we've hit the knee in benchmark where models are no longer failing in obviously awful ways.
- Morromist 2mo agoI haven't really seen evidence that any ai can reliably draw a pelican riding a bicycle. Not if you look at the image long enough to take it in. Even the best ones have something wrong with them. Not a matter of taste but a matter of having both legs peddling on the viewer's side of the bicycle or having two beaks. I'm actually beginning to wonder if some people who ignore these things have a different, somewhat lesser ability to percieve image details than I do. I mean I guess its fine to go on to another test despite never actually passing the pelican bike test, but there's a sense that we have to use another test because AI is now good at pelicans on bikes, which is just not true.
- enos_feedler 2mo agoAI has deeply changed the way I think, feel and act around a computer. In the same way that dialing into the internet changed things for me. Since using ChatGPT the first time until now I have never cared once to look at these pelicans on bikes people seem to get hung up about. It could never have been a thing and nothing would change. See the forest through the trees.
- Waterluvian 2mo agoSpeaking of benchmarks has anyone given AIs Where’s Waldo pages and asked it to find Waldo? I’ve been trying it on them all and can’t find one that does it consistently. The best will tell me they can’t. The worst confidently point out one of countless Waldo-likes.
- mikojan 2mo agoAfter watching this video I am absolutely positive that the issue is not a lack of stamina in humans. It is that humans have the capacity to realize that this is a bad idea long before they complete it.
- sinaatalay 2mo agoOn consumer devices, AI communicates with us through speakers and screens. Screens are the richer medium, so most consumer AI innovation will happen there. Computer graphics will have enormous applications because they are directly controllable by LLM-generated code. Video models are probabilistic and less suitable when precision matters. In education, for example, we need exact visuals. If an AI wants to plot y = sin(x), it should generate the precise graph through computer graphics rather than approximate it with a video model.
- duxup 2mo agoCoding, graphics, all seem to have very defined background data for an LLM to base their decisions on, test, and even when I correct it we're all on the same page. Seems like there's lots of room for AI productivity there. What I find funny is that approximate to computer graphics are video games. When I ask AI about a decision available to me in a video game AI completely fails, OFTEN. I assume all the forums and changes made to a game over time might be quite confusing for AI. But I've also seen it completely make up characters and decisions and weapons and so on about some very clearly defined games and paths. It's an interesting dynamic.
- try-working 2mo agothis is not a good benchmark for models, but it's great if you're optimizing for attention on twitter because video content and 3d animations perform best on social media. a real benchmark is instead running evals on your own traces, and building a cost/quality/speed profile for models based on real workloads. but it doesn't get you a shiny video you can post on twitter.
- sumedh 2mo agoYou can do both.
- hooloovoo_zoo 2mo agoI suspect LotR is a singularly unrepresentative choice here considering how much info exists about it.
- hansmayer 2mo ago[dead]
- swe_dima 2mo agoIn my experience SVGs are still too hard for LLMs. I gave Fable a jpeg and asked to draw an SVG, using a loop that renders the SVG into an image so Fable can inspect it. Results looked like drawing of a 5 year old.
- jcims 2mo agoI’d like to see a human one shot a pelican on a bicycle in raw svg.
- NewJazz 2mo agoWhat do you mean by "oneshot"? The term applied to genai makes sense, but it doesn't make a whole lot of sense to apply the term to human art.
- jkahrs595 2mo agoWhat doesn’t make sense? They aren’t drawing it by hand, they are still using SVG. Give them one chance before rendering it.
- NewJazz 2mo agoI feel like letting them render + screenshot + judge + iterate would still be called oneshotting. I mean, my coding agent calls tools to verify code correctness. I'd still call it oneshotting I guess.
- sapal 2mo agoI understand it as “write SVG, say when you are done and only then you are allowed to see the rendered result”. Multiple shots would be either “you get more than one try, we'll pick the best” or “you can see the rendered result and iterate” (or maybe that would be “model + harness with tools”?)
- NewJazz 2mo ago[dead]
- albertzeyer 2mo agoBut the LLMs are also not one-shotting it, or are they? I assume they have some ways to verify it, e.g. to visualize it (convert to PNG, then feed as vision tokens back to the LLM), or other ways, maybe also pure text LLMs have some ways to verify the result at least somewhat? And with such feedback loop they can iterate.
- informal007 2mo agoit shows the possibility that SVG replace PNG/JPG even video.
- informal007 2mo agoOne difference for human to understand the video is that we only care the changes on a picture compare to LLM
- deleted 2mo ago[deleted]
- djhworld 2mo agoIt would be interesting to see the models work on a book it hasn't been trained on yet. I guess sadly that means any book released very recently. Definitely impressive demo, I do wonder though if the countless artwork, films, images etc produced over many decades around Lord of the Rings somewhat influenced the outcome of this though.
- robomc 2mo agoYeah that seems like an enormous problem with this example.
- Lerc 2mo agoI think I could tolerate 50 Shades of grey rendered in this style.
- mvdtnz 2mo ago> I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom but LLMs have all the stamina and patience in the world, so it's an example where we go from "no one would ever do this" to "sure, why not, it's ~free". Except it's not ~free, it cost ~$10. And no one in their right mind would ever exchange $10 for that crap output except in this brief moment that we're in because it's fun and surprising to see what will happen. The actual result is as close to useless as it's possible to be - it's not interesting in its own right, it's not aesthetically pleasing, nor funny, nor informative. It's just slop.
- angoragoats 2mo ago[flagged]
- angoragoats 2mo agoI guess I’ll just keep flagging every single Twitter link on this site. I hope that some day the community here will grow a spine.
- toplinesoftsys 2mo agoThis is just a very rough basic game skeleton. Companies believing AI is smart enough to spit out almost ready product will learn a hard way that it will will give them just a starting point that still needs a lot of work. This is a nature of all modern LLMs - they easily lose context: the bigger the context the more losses and distortions are, especially in the middle. So, giving AI the complete book does not mean it will follow everything in it - quite the contrary.
- killingtime74 2mo agoI love how nobody cares about copyright anymore. Not even an after thought. Might be okay if you're an employee of anthropic/openai but for us mere mortals I'm not sure I would share something this blatant. In US it's $150k+ per violation and you've given them all the proof (even a confession)
- teiferer 2mo ago> sure, why not, it's ~free Yeah, please check in with the folks protesting data center builds in their town causing their electricity prices to skyrocket and tap water to turn into a scarce resource. And now, instead of actually doing the above, please go ahead and downvote me, because how dare he question those LLM games.
- skakcnejwisnsid 2mo ago> because how dare he question those LLM games. I thought you are the one who's on Greta's side, you ar e the one who should be saying "how dare you?"
- ianberdin 2mo agoI still believe my bench is superior to Pelican or Karpathy’s. Everyone knows how a MacBook looks. Any missing or incorrect detail will be obvious. However, Pelican or this world can be anything. https://playcode.io/blog/macbook-svg-benchmark https://playcode.io/blog/macbook-svg-benchmark
- ianberdin 2mo agoAnyway, only Opus 5 Max and Fable 5 XHigh + Max make a good-looking Mac.
- kooi 2mo agoI think the most important insight is the limitation of LLM perception:Slowly taking screenshots. That method of perception probably scales N^2... so sure with more compute, LoTR animation will improve. But I think to get a real jump in "experiential feedback", perception needs to scale linear or sublinear. Maybe that's there LeCunn's jepa will come in. There needs to be the removal of the middle man: image -> text -> action To image -> action.
- sampton 2mo agoPelican maxxing.
- Arshad-Talpur 2mo agoHow long we will be testing and benchmarking generative capabilities of LLMs? in creativity, in code generation less or more result is expected and approved, but in execution $19.8 can not be $19.9 or $19.7.
- minikomi 2mo agoWe just need to go one layer deeper: Generate an SVG of an ai generating an SVG of a pelican on a bicycle. https://chatgpt.com/s/t_6a6fc59ae00081918095322a63e1503c https://chatgpt.com/s/t_6a6fc59ae00081918095322a63e1503c
- JBAnderson5 2mo ago> Opus went off for ~2 hours and wrote 5500 lines of code that (procedurally) rendered the story. > it's an example where we go from "no one would ever do this" to "sure, why not, it's ~free". How is two hours worth of token generation free? With tech revolutions things get cheaper/faster/better/doable, but there’s still real world limits. The advent of railroads made it feasible for the average person to cross the country, but it still cost a lot of time and resources. People weren’t crossing the country every weekend for fun just because it was now doable. Why do we treat LLMs as ~free when we are generating things that weren’t doable before but have to invest more money than the Apollo program to build AI data centers let alone account for the operating costs?
- janderson215 2mo agoHe wrote “~free” and also that he set a $10USD limit on it. For the result, I think it is fair to say it is approximately free.
- darrinm 2mo agoA simple prompt that still stumps frontier LLMs most of the time is “create a pinball game”. They’ll put all the right pieces there but then fail to arrange them such that the game is truly playable. They’ll put a wall in the way of the launch chute so the ball can’t be launched. Or the flippers will pivot the wrong way. Or there will be holes such that the ball drops off the bottom without getting within reach of the flippers, etc. Opus 5 is the first I’ve seen to “one shot” it (in a harness, so it was more than one LLM call).
- Schlagbohrer 2mo agoFailed demos like this give weight to the argument that AIs need more of a world model, an understanding of how physics works to avoid obvious stumbles like this.
- kzrdude 2mo agoWe don't use these models by way of "one shot" so I don't see why it's relevant. It's clearly useful to let them iterate.
- yencabulator 2mo agoNo tool-calling LLM is really "one shot" anymore, so we can easily repurpose that expression to mean "without further human feedback". The generated pinball games still seem to end up coming out broken. There's not much conversation about a good arrangement of pinball playing fields, and the LLM is a next token predictor. Pick any out-of-mainstream topic and same happens. Pelicans on bicycles used to be out-of-mainstream.
- overgard 2mo agoOh god, please leave video games alone, we have enough slop.
- 0x1ceb00da 2mo agoWe already have tons of videogame slop. 100s of videogames are released on steam every day.
- consumer451 2mo agoThe difference between this and Simon's pelican is that with Simon, I get the prompt. Last I checked, I did not see the prompt for this really cool thing, so it is not reproducible. Did I miss the prompt somewhere?
- vanjajaja1 2mo agohe said the prompt was the first paragraph of LoTR, but he didn't mention a preamble this guy seems to have taken that idea and got something similar/better, so likely the prompt isn't too special https://x.com/Izkimar/status/2083819741643178208?s=20 https://x.com/Izkimar/status/2083819741643178208?s=20
- consumer451 2mo agoThat was far worse as far as visual story-telling. Anthropic wins, which is to be expected against whatever that product is. In either case, I still don't see how I could reproduce this to test against various models, which is the entire point of Simon's pelican.
- O4epegb 2mo agoThe second one is also Anthropic, same model even.
- KeplerBoy 2mo agoyou are seeing single realizations of non-deterministic processes. Confidently stating one model is better than the other is not really possible this way; that's just like stating one dice is better than the other because you rolled a six with that one on the first try.
- attheballot 2mo agoIn other words, they did not even need a prompt. You could probably skip feeding the paragraph and just told "generate visuals for the opening paragraph of the LotR", and the LLM would successfully do it as it has a very good idea of what they are from its training data. This is a bigger difference to the pelican than simple reproducibility steps. The pelican is intentionally esoteric, and thus open ended. The LotR is mundane and has a "correct" answer, aka, copy the movie. It makes it a really awful test of capabilities. The pelican isn't a slop test. This crap is.
- osetinsky 2mo agoI’m exploring an analogous idea for music. The quality of the musical output has improved noticeably with the latest frontier models Karpathy’s point about the models not being able to easily audit their work is something I’m struggling with —- how can the audio output be made perceivable? Curious what people think about this question. Here’s a synth to play with+remix https://underscore.audio/s/cmp_8b226859-420/iron_rareflash https://underscore.audio/s/cmp_8b226859-420/iron_rareflash
- deleted 2mo ago[deleted]
- cyanregiment 2mo agoA quick peruse of the Three.js examples page should prove why it's so easy to get an LLM to produce decent Three.js. https://threejs.org/examples/ https://threejs.org/examples/ Here's minecraft: https://threejs.org/examples/webgl_geometry_minecraft.html https://threejs.org/examples/webgl_geometry_minecraft.html Here's an FPS: https://threejs.org/examples/games_fps.html https://threejs.org/examples/games_fps.html The library is extremely well-documented. When Three.js vibe coded projects started blowing up on Twitter 1-2 years ago, I wasn't that impressed then either because I knew what it was doing. Anyone who remembers the C compiler built by an LLM! backlash probably feels the same way: Why would I use an LLM to create a well-known demo rather than fork that demo itself? What I haven't seen yet from an LLM is it create anything fundamentally new and exciting: A new UIX that is actually good. A new service that is actually good. A game with an art style I haven't seen, music, storytelling - something you'd expect out of a AAA studio. Given that rant: The coolest part is the multi-modality between text and animation. However, I think the end product would have been a lot better if it was just a video. Having it do it in Three.js didn't add a ton of wow factor for me, and it would have been a lot better looking as video. > Something like an ephemeral GTA of X on demand. Here's where you lose me. AAA gaming is very far away from this Three.js demo. But the novel part being the syncing of narrative to the visual scene - a text-to-audio (video) book type technology seems very possible (and useful). Nice work. We need more big projects like this involving AI (if anything, to get away from the slop argument largely focusing on 1-shot experiments).
- weakfish 2mo agoOut of curiosity, I asked Opus 5 to do the same with the first ~1.5 pages of Neuromancer by William Gibson, which I thought might be a good test of interpretation from the model. It refused to use the text verbatim because of copyright (ironic), but the output was interesting nonetheless. https://claude.ai/public/artifacts/275dc3c2-7bd3-432b-94ff-dd876629ecbd https://claude.ai/public/artifacts/275dc3c2-7bd3-432b-94ff-d...
- jedbrooke 2mo agoit would be interesting to see results from works that don't have notable existing film/tv adaptations or illustrations (though I suppose there will always be fan art) to see if it’s really able to generate something novel. I’ve had my own version of this since I’ve been reading LotR recently too and I kinda wish I had read the book before watching the movie (not that the movies are bad though). Fortunately I am finding there is tons more content in the book than the movie :) Also, large models refusing to work with copyright material is really hypocritical, copyright enforcement for thee but not for me
- neop1x 2mo agoThanks. Your result is what I would expect - nice stuff, but far from what he supposedly "casually" generated.
- CapsAdmin 2mo agoSomething that bugs me about these new demos is that they benchmark the model and harness at the same time. The harness steer the model to expand the vague prompt, self correct with tools, etc. I think it's useful when evaluating "model+intended harness", but I'm more interested in seeing raw model improvements than harness improvements.
- larpathyparpav 2mo ago[dead]
- iDon 2mo agoI just went down a rabbit-hole, and happily I found the rabbit - an early example of this type of prompt (draw an unusual animal character in a vector graphics language). There is a paper and a 1 hour podcast resulting from a Microsoft evaluation of a pre-release of GPT4. One of the prompts (see pages 4,7,8 in the paper PDF) was to draw a unicorn in TikZ. (I'll leave to the historians the questions of whether this was the first, or whether Simon Willison may have been inspired by this). I remember hearing some of the podcast, and recalled the prompt about balancing on a nail, and the triangle forming the unicorn's tusk; this was clearly a big step beyond choosing the next word and blending images. https://arxiv.org/abs/2303.12712 https://arxiv.org/abs/2303.12712
- mrdootdoot 2mo agoClaude is fire with blender mcp https://youtu.be/pwEfsp8aq9I?is=ULNxiSdI48puE_S3 https://youtu.be/pwEfsp8aq9I?is=ULNxiSdI48puE_S3
- uuuynnnuuuyyyn 2mo ago[dead]
- eric_khun 2mo agoEveryone's about to make games sooner than expected. games are slow, might look buggy and unfinished. But I think that's just about time before we see fun games coming up. i'm betting on it with https://antics.gg https://antics.gg . if you want to host your multiplayer vibecoded games, feel free to use us!
- didibus 2mo agoDid I miss when they managed to draw a pelican? Cause all the ones I've seen are wrong in some way.
- nomel 2mo agoPerfection is such an absolutely wild requirement for what we're seeing happen here, with the level of understanding required, from a tech that was complete fiction 5 years ago. Maybe excitement and wonder, in tech, is just something for us old guys, that watched it all be birthed. Get off my lawn!
- crabmusket 2mo agoI don't think didibus is saying they expect perfection, just that there is still room for the pelican prompt to measure meaningful future improvement.
- didibus 2mo agoThe article makes it sound like Pelicans riding a bike are solved, and we can move on to something else. Nothing to do with excitement or wonder of the tech, it just claims one thing and I feel I've not seen that be the case so the claim seems incorrect to me.
- nomel 2mo ago> We're starting to leave the territory where you'd test an LLM by e.g. "create an svg of pelican on a bicycle". First sentence of the post is clear that it's not solved.
- didibus 2mo agoMight be my reading comprehension, but that reads to me like it's implying we're leaving the territory because it's solved and now need something harder.
- 2mo ago
- firatsarlar 2mo ago[flagged]
- gaigalas 2mo agoI think his explanation is bullshit. From my experience (PixiJS game), Opus can look at things and screenshot. It's slow but it works. The issue is that it doesn't try to make them look good. Instead, it tries to find one bug and then fix that one bug and close the session. That is amazing discipline for webapps or whatever, but for game design is just the worst if you need attention to detail and tuning of multiple holistic systems that produce an effect together. That really limits the kinds of approaches you can can use to achieve things visually. Sure, the testing harness could be better, but that's not what makes it a poor fit. The base model seems also capable (it can identify the issues, just isn't willing to disangage from this "I found X and fixed it" single-thing work posture). The kind of work is just different, and it was probably never trained on it, so it feels off and an uphill battle to use it.
- haxfenx 2mo agoSo I build Pokémon SVG Bench to evaluate and push the boundaries of LLM SVG generation, including their knowledge of Pokémon. Safe to say, most models still have a long way to go. https://svg-bench.fenx.work/ https://svg-bench.fenx.work/
- rw2 2mo agoI think the argument for the pelican is if it cannot do a simple thing perfectly. A more complex thing is just a stacking a small errors.
- jatins 2mo ago> Last thought is that the domain of worlds/games exposes a weakness in LLMs: they can't easily audit their work because they aren't able to efficiently and natively perceive videos or play games within them. Again comes back to the point of verifiable rewards. The moment you take it away from LLMs they just stop being as good
- Schlagbohrer 2mo agoI've come across the same issue at home with my qwen3.6 agent in Pi. It has to take screenshots of a 3D scene and look at the screenshots to see what is happening. It can even produce a video for me with ffmpeg, but it can't watch the video or see the render output directly, only take screenshots for review.
- DoDecaHeJon 2mo agoFun ways to test the models. What other unique tests are people using?
- moinism 2mo agoWe've been building motion graphics capability (on chatoctopus.com), and "closing the loop" has been one of the hardest challenges. Coding agents have it easy; static analysis, lints, unit tests, etc. But when visual perception and "taste" get involved, it becomes a lot harder.
- rldjbpin 2mo agoi feel that the proposed approach is close to becoming "demo prawn"* territory. similar to cpu benchmarks (e.g. geekbench, which also had a recent update), the goal is to distinguish model in some arbitrary thing. however, it is not quantifiable but subjectively even the latest show distinctive differences regardless. large context window is still relatively new and not "feasible" even commercially, so making that a requirement for a benchmark would not make it accessible especially for open-source ones. * yes, i misspelled it intentionally
- t14000 2mo agoAt ~$10 per paragraph (or ~90 seconds) it's going to cost quite a lot to render the entire thing...
- bicepjai 2mo agoFree “Curious Karpathy Ads Service” for Anthropic :)
- nkg 2mo agoI am amazed by Karpathy's writing. He can express complicated thoughts on these topics in a way that makes you feel you would have come to the same conclusion :)
- gowld 2mo agoWhy does Twitter show a reply "Yah" from a random fan? xcancel shows a whole stream of thoughtful discussion.