9 ms·
Just tested the new Opus 4.6 (1M context) on a fun needle-in-a-haystack challenge: finding every spell in all Harry Potter books. All 7 books come to ~1.75M to
by ck_one 8mo ago
Just tested the new Opus 4.6 (1M context) on a fun needle-in-a-haystack challenge: finding every spell in all Harry Potter books.
All 7 books come to ~1.75M tokens, so they don't quite fit yet. (At this rate of progress, mid-April should do it ) For now you can fit the first 4 books (~733K tokens).
Results: Opus 4.6 found 49 out of 50 officially documented spells across those 4 books. The only miss was "Slugulus Eructo" (a vomiting spell).
Freaking impressive!
- zamadatix 8mo agoTo be fair, I don't think "Slugulus Eructo" (the name) is actually in the books. This is what's in my copy: > The smug look on Malfoy’s face flickered. > “No one asked your opinion, you filthy little Mudblood,” he spat. > Harry knew at once that Malfoy had said something really bad because there was an instant uproar at his words. Flint had to dive in front of Malfoy to stop Fred and George jumping on him, Alicia shrieked, “How dare you!”, and Ron plunged his hand into his robes, pulled out his wand, yelling, “You’ll pay for that one, Malfoy!” and pointed it furiously under Flint’s arm at Malfoy’s face. > A loud bang echoed around the stadium and a jet of green light shot out of the wrong end of Ron’s wand, hitting him in the stomach and sending him reeling backward onto the grass. > “Ron! Ron! Are you all right?” squealed Hermione. > Ron opened his mouth to speak, but no words came out. Instead he gave an almighty belch and several slugs dribbled out of his mouth onto his lap.
- ck_one 8mo agoThen it's fair that id didn't find it
- sobjornstad 8mo agoI have a vague recollection that it might come up named as such in Half-Blood Prince, written in Snape's old potions textbook? In support of that hypothesis, the Fandom site lists it as “mentioned” in Half-Blood Prince, but it says nothing else and I'm traveling and don't have a copy to check, so not sure.
- zamadatix 8mo agoHmm, I don't get a hit for "slugulus" or "eructo" (case insensitive) in any of the 7. Interestingly two mentions of "vomit" are in book 6, but neither in reference to to slugs (plenty of Slughorn of course!). Book 5 was the only other one a related hit came up: > Ron nodded but did not speak. Harry was reminded forcibly of the time that Ron had accidentally put a slug-vomiting charm on himself. He looked just as pale and sweaty as he had done then, not to mention as reluctant to open his mouth. There could be something with regional variants but I'm doubtful as the Fandom site uses LEGO Harry Potter: Years 1-4 as the citation of the spell instead of a book. Maybe the real LLM is the universe and we're figuring this out for someone on Slacker News a level up!
- xiomrze 8mo agoHonest question, how do you know if it's pulling from context vs from memory? If I use Opus 4.6 with Extended Thinking (Web Search disabled, no books attached), it answers with 130 spells.
- clanker_fluffer 8mo agoWhat was your prompt?
- petercooper 8mo agoOne possible trick could be to search and replace them all with nonsense alternatives then see if it extracts those.
- andai 8mo agoThat might actually boost performance since attention pays attention to stuff that stands out. If I make a typo, the models often hyperfixate on it.
- jazzyjackson 8mo agoA fine instruction following task but if harry potter is in the weights of the neural net, it's going to mix some of the real ones with the alternates.
- ozim 8mo agoExactly there was this study where they were trying to make LLM reproduce HP book word for word like giving first sentences and letting it cook. Basically they managed with some tricks make 99% word for word - tricks were needed to bypass security measures that are there in place for exactly reason to stop people to retrieve training material.
- ck_one 8mo agoDo you remember how to get around those tricks?
- meroes 8mo agoWhat is this supposed to show exactly? Those books have been feed into LLMs for years and there's even likely specific RLHF's on extracting spells from HP.
- rvz 8mo ago> What is this supposed to show exactly? Nothing. You can be sure that this was already known in the training data of PDFs, books and websites that Anthropic scraped to train Claude on; hence 'documented'. This is why tests like what the OP just did is meaningless. Such "benchmarks" are performative to VCs and they do not ask why isn't the research and testing itself done independently but is almost always done by their own in-house researchers.
- jaco6 8mo ago[dead]
- muzani 8mo agoThere was a time when I put the EA-Nasir text into base64 and asked AI to convert it. Remarkably it identified the correct text but pulled the most popular translation of the text than the one I gave it.
- majewsky 8mo agoSucks that you got a really shitty response to your prompt. If I were you, the model provider would be receiving my complaint via clay tablet right away.
- noumenon1111 8mo agoImagine you ordered the new Claude Opus and instead you got Gemini telling you to glue the cheese on your pizza...
- deleted 8mo ago[deleted]
- guluarte 8mo agoyou can get the same result just asking opus/gpt, it is probably internalized knowledge from reddit or similar sites.
- ck_one 8mo agoIf you just ask it you don't get the same result. Around 13 spells were missing when I just prompted Opus 4.6 without the books as context.
- guluarte 8mo agobecause it is non-deterministic, i just did and got 50 spells. prompt: finding every spell in all Harry Potter books in the first 4 books, only the officially documented spells across those 4 books, there are 50, without using websearch Here's my best compilation from memory, organized by book. I'll be upfront: getting to exactly 50 is tricky without cross-referencing, so some might be debatable or I might miss a couple. Here goes: --- *Book 1 - Philosopher's/Sorcerer's Stone* 1. *Lumos* - produces light 2. *Nox* - extinguishes Lumos 3. *Alohomora* - unlocks doors 4. *Wingardium Leviosa* - levitation 5. *Petrificus Totalus* - full body-bind 6. *Locomotor Mortis* - leg-locker curse 7. *Incendio* - produces fire 8. *Reparo* - mends broken objects *Book 2 - Chamber of Secrets* 9. *Expelliarmus* - disarms opponent 10. *Rictusempra* - tickling charm 11. *Tarantallegra* - forces target to dance 12. *Serpensortia* - conjures a snake 13. *Finite Incantatem* - ends spell effects 14. *Obliviate* - erases memories 15. *Aparecium* - reveals invisible ink 16. *Immobulus* - freezes target in place 17. *Homorphus* - mentioned by Lockhart (werewolf cure claim) 18. *Peskipiksi Pesternomi* - Lockhart's failed pixie spell *Book 3 - Prisoner of Azkaban* 19. *Expecto Patronum* - produces a Patronus 20. *Riddikulus* - repels a Boggart 21. *Waddiwasi* - shoots objects at target 22. *Impervius* - waterproofing charm 23. *Dissendium* - opens secret passage 24. *Ferula* - conjures bandages/splint 25. *Mobilicorpus* - moves a body 26. *Lumos Maxima* - intense light *Book 4 - Goblet of Fire* 27. *Accio* - summoning charm 28. *Avada Kedavra* - killing curse 29. *Crucio* - Cruciatus curse (torture) 30. *Imperio* - Imperius curse (control) 31. *Stupefy* - stunning spell 32. *Engorgio* - enlarges target 33. *Reducio* - shrinks target 34. *Sonorus* - amplifies voice 35. *Quietus* - reverses Sonorus 36. *Morsmordre* - conjures the Dark Mark 37. *Priori Incantatem* - reveals last spell cast 38. *Deletrius* - erases magical residue 39. *Densaugeo* - enlarges teeth 40. *Furnunculus* - causes boils 41. *Impedimenta* - slows/stops target 42. *Reducto* - blasts solid objects 43. *Diffindo* - severing charm 44. *Relashio* - releases sparks/grip 45. *Orchideous* - conjures flowers 46. *Avis* - conjures birds 47. *Point Me* - Four-Point Spell (compass) 48. *Ennervate* - revives stunned person 49. *Protego* - shield charm 50. *Conjunctivitis Curse* - affects eyesight (Krum on the dragon) --- A few caveats: some of these (like Lumos Maxima, Homorphus, Peskipiksi Pesternomi) are borderline since they're either mentioned rather than properly cast, or might be film additions that bleed into memory. The Conjunctivitis Curse is described but its incantation isn't explicitly given in the text. And Protego might technically first appear with its incantation in Book 5 during DA practice rather than Book 4. If you want, I can turn this into a spreadsheet or document with columns for spell name, effect, who casts it, and which chapter.
- bartman 8mo agoHave you by any chance tried this with GPT 4.1 too (also 1M context)?
- LanceJones 8mo agoAssuming this experiment involved isolating the LLM from its training set?
- grey-area 8mo agoOf course it didn't. Not sure you really can do that - LLMs are a collection of weights from the training set, take away the training set and they don't really exist. You'd have to train one from scratch excluding these books and all excerpts and articles about them somehow, which would be very expensive and I'm pretty sure the OP didn't do that. So the test seems like a nonsensical test to me.
- adarsh2321 8mo ago[dead]
- golfer 8mo agoThere's lots of websites that list the spells. It's well documented. Could Claude simply be regurgitating knowledge from the web? Example: https://harrypotter.fandom.com/wiki/List_of_spells https://harrypotter.fandom.com/wiki/List_of_spells
- ck_one 8mo agoIt didn't use web search. But for sure it has some internal knowledge already. It's not a perfect needle in the hay stack problem but gemini flash was much worse when I tested it last time.
- joshmlewis 8mo agoI think the OP was implying that it's probably already baked into its training data. No need to search the web for that.
- deleted 8mo ago[deleted]
- viraptor 8mo agoIf you want to really test this, search/replace the names with your own random ones and see if it lists those. Otherwise, LLMs have most of the books memorised anyway: https://arstechnica.com/features/2025/06/study-metas-llama-3-1-can-recall-42-percent-of-the-first-harry-potter-book/ https://arstechnica.com/features/2025/06/study-metas-llama-3...
- ribosometronome 8mo agoCouldn't you just ask the LLM which 50 (or 49) spells appear in the first four Harry Potter books without the data for comparison?
- viraptor 8mo agoIt's not going to be as consistent. It may get bored of listing them (you know how you can ask for many examples and get 10 in response?), or omit some minor ones for other reasons. By replacing the names with something unique, you'll get much more certainty.
- muzani 8mo agoThere's a benchmark which works similarly but they ask harder questions, also based on books https://fiction.live/stories/Fiction-liveBench-Feb-21-2025/oQdzQvKHw8JyXbN87 https://fiction.live/stories/Fiction-liveBench-Feb-21-2025/o... I guess they have to add more questions as these context windows get bigger.
- dwa3592 8mo agohave another LLM (gemini, chatgpt) make up 50 new spells. insert those and test and maybe report here :)
- irishcoffee 8mo agoThe top comment is about finding basterized latin words from childrens books. The future is here.
- kybernetikos 8mo agoI recently got junie to code me up an MCP for accessing my calibre library. https://www.npmjs.com/package/access-calibre https://www.npmjs.com/package/access-calibre My standard test for that was "Who ends up with Bilbo's buttons?"
- TheRealPomax 8mo agoThat doesn't seem a super useful test for a model that's optimized for programming?
- dom96 8mo agoI often wonder how much of the Harry Potter books were used in the training. How long before some LLM is able to regurgitate full HP books without access to the internet?
- IhateAI 8mo ago[flagged]
- siwatanejo 8mo ago> All 7 books come to ~1.75M tokens How do you know? Each word is one token?
- koakuma-chan 8mo agoYou can download the books and run them through a tokenizer. I did that half a year ago and got ~2M.
- huangmeng 8mo agoyou are rich
- matt_lo 8mo agouse AI to rewrite all the spells from all the books, then try to see if AI can detect the rewritten ones. This will ensure it's not pulling from it's trained data set.
- gbalduzzi 8mo agoNeat idea, but why should I use AI for a find and replace? It feels like shooting a fly with a bazooka
- miohtama 8mo agoBazooka guarantees the hit
- xenodium 8mo agoI like LLMs, but guarantees in LLMs are... you know... not guaranteed ;)
- throwaway290 8mo agoI think that was the point
- jack_pp 8mo agoit's like hiring someone to come pick up your trash from your house and put it on the curb. it's fine if you're disabled
- luckydata 8mo agodo you know all the spells you're looking for from memory?
- wickedsight 8mo agoYou could just, you know, Google the list.
- dr_dshiv 8mo agoComparison to another model?
- dudewhocodes 8mo agoThere are websites with the spells listed... which makes this a search problem. Why is an LLM used here?
- bilekas 8mo agoIt's just a benchmark test excersize.
- hereonout2 8mo agoI was playing about with Chat GPT the other day, uploading screen shots of sheet music and asking it to convert it to ABC notation so I could make a midi file of it. The results seemed impressive until I noticed some of the "Thinking" statements in the UI. One made it apparent the model / agent / whatever had read the title from the screenshot and was off searching for existing ABC transcripts of the piece Ode to Joy. So the whole thing was far less impressive after that, it wasn't reading the score anymore, just reading the title and using the internet to answer my query.
- nobodywillobsrv 8mo agoYes I have found that grok for example actually suddenly becomes quite sane when you tell it to stop querying the internet And just rethink the conversation data and answer the question. It's weird, it's like many agents are now in a phase of constantly getting more information and never just thinking with what they've got.
- bestham 8mo agoTouché, that is what we humans are doing to some degree as well.
- Szpadel 8mo agobut isn't it what we wanted? we complained so much that LLM uses deprecated or outdated apis instead of current version because they relied so much on what they remembered
- nobodywillobsrv 8mo agoTo be clear, what I mean is that grok will query 30 pages and then answer your question vaguely or wrongly and then ask for clarification of what it meant and then it goes and requeries everything again ... I can imagine why it might need to revisit pages etc and it might be a UI thing but it still feels like until you yell at it to stop searching for answers to summarise it doesn't activate it's "think with what you got" mode. I guess we could call this gathering and then do your best conditional on what you found right now.
- deleted 8mo ago[deleted]
- grey-area 8mo agoSurely the corpus Opus 4.6 ingested would include whatever reference you used to check the spells were there. I mean, there are probably dozens of pages on the internet like this: https://www.wizardemporium.com/blog/complete-list-of-harry-potter-spells https://www.wizardemporium.com/blog/complete-list-of-harry-p... Why is this impressive? Do you think it's actually ingesting the books and only using those as a reference? Is that how LLMs work at all? It seems more likely it's predicting these spell names from all the other references it has found on the internet, including lists of spells.
- sigmoid10 8mo agoMost people still don't realize that general public world knowledge is not really a test for a model that was trained on general public world knowledge. I wouldn't be surprised if even proprietary content like the books themselves found their way into the training data, despite what publishers and authors may think of that. As a matter of fact, with all the special deals these companies make with publishers, it is getting harder and harder for normal users to come up with validation data that only they have seen. At least for human written text, this kind of data is more or less reserved for specialist industries and higher academia by now. If you're a janitor with a high school diploma, there may be barely any textual information or fact you have ever consumed that such a model hasn't seen during training already.
- joenot443 8mo ago> even proprietary content like the books themselves This definitely raises an interesting question. It seems like a good chunk of popular literature (especially from the 2000s) exists online in big HTML files. Immediately to mind was House of Leaves, Infinite Jest, Harry Potter, basically any Stephen King book - they've all been posted at some point. Do LLMS have a good way of inferring where knowledge from the context begins and knowledge from the training data ends?
- rendx 8mo ago> It seems like a good chunk of popular literature (especially from the 2000s) exists online in big HTML files Anna's Archive alone claims to currently publicly host 61,654,285 books, more than 1PB in total.
- hansmayer 8mo ago> Just tested the new Opus 4.6 (1M context) on a fun needle-in-a-haystack challenge: finding every spell in all Harry Potter books. Clearly a very useful, grounded and helpful everyday use case of LLMs. I guess in the absence of real-world use cases, we'll have to do AI boosting with such "impressive" feats. Btw - a well crafted regex could have achieved the same (pointless) result with ~0.0000005% of resources the LLM machine used.
- ActionHank 8mo agoThe books were likely in the training data, I don't know that it's that impressive.
- SebastianSosa 8mo agonow thx to this post (and the infra provider inclination to appeal to hacker news) we will never know if the model actually discovered the 50 spells or memorized it. Since it will be trained on this. :( But what can you do, this is interesting
- psychoslave 8mo agoAh and no one thrown TOAC in it yet?
- polynomial 8mo agoYou need to publish this tbh
- deleted 8mo ago[deleted]
- kmacdough 8mo agoWhat are we testing here? It feels like a very odd test because it's such an unreasonable way to answer this with an LLM. Nothing about the task requires more than a very localized understanding. It's not like a codebase or corporate documentation, where there's a lot of interconectedness and context that's important. It also doesn't seem to poke at the gap between human and AI intelligence. Why are people excited? What am I missing?
- kylehotchkiss 8mo agoI love the fun metric. My hope is that locally run models can pass this test in the next year or two!
- matt-p 8mo agoNow try it without giving it the books as context. I'm sure it probably knows there are 49.
- deleted 8mo ago[deleted]
- deleted 8mo ago[deleted]
- big-chungus4 8mo agoNow edit the books and replace all spell names with different ones, and try again
- big-chungus4 8mo agoBy work I meant with
- NeroVanbierv 8mo agoJust did a similar experiment but outside the harry potter universe to remove the training bias. It worked well! > ChatGPT: "Generate a two page short story like harry potter, but don´t mention anyting harry potter related. make up 4 unique spells in the story that are used" Response see https://chatgpt.com/share/698af9cd-f628-800d-9250-b260f1478cf1 https://chatgpt.com/share/698af9cd-f628-800d-9250-b260f1478c... > Claude: "What unique wizarding spells can you find in this story? [story]" Response = https://i.imgur.com/Jzzs3PC.png https://i.imgur.com/Jzzs3PC.png