4 ms·
he said the prompt was the first paragraph of LoTR, but he didn't mention a preamble this guy seems to have taken that idea and got something similar/better, s
by vanjajaja1 2mo ago
he said the prompt was the first paragraph of LoTR, but he didn't mention a preamble
this guy seems to have taken that idea and got something similar/better, so likely the prompt isn't too special
https://x.com/Izkimar/status/2083819741643178208?s=20 https://x.com/Izkimar/status/2083819741643178208?s=20
- consumer451 2mo agoThat was far worse as far as visual story-telling. Anthropic wins, which is to be expected against whatever that product is. In either case, I still don't see how I could reproduce this to test against various models, which is the entire point of Simon's pelican.
- O4epegb 2mo agoThe second one is also Anthropic, same model even.
- KeplerBoy 2mo agoyou are seeing single realizations of non-deterministic processes. Confidently stating one model is better than the other is not really possible this way; that's just like stating one dice is better than the other because you rolled a six with that one on the first try.
- attheballot 2mo agoIn other words, they did not even need a prompt. You could probably skip feeding the paragraph and just told "generate visuals for the opening paragraph of the LotR", and the LLM would successfully do it as it has a very good idea of what they are from its training data. This is a bigger difference to the pelican than simple reproducibility steps. The pelican is intentionally esoteric, and thus open ended. The LotR is mundane and has a "correct" answer, aka, copy the movie. It makes it a really awful test of capabilities. The pelican isn't a slop test. This crap is.