4 ms·
Honest question, how do you know if it's pulling from context vs from memory? If I use Opus 4.6 with Extended Thinking (Web Search disabled, no books attached)
by xiomrze 8mo ago
Honest question, how do you know if it's pulling from context vs from memory?
If I use Opus 4.6 with Extended Thinking (Web Search disabled, no books attached), it answers with 130 spells.
- clanker_fluffer 8mo agoWhat was your prompt?
- petercooper 8mo agoOne possible trick could be to search and replace them all with nonsense alternatives then see if it extracts those.
- andai 8mo agoThat might actually boost performance since attention pays attention to stuff that stands out. If I make a typo, the models often hyperfixate on it.
- jazzyjackson 8mo agoA fine instruction following task but if harry potter is in the weights of the neural net, it's going to mix some of the real ones with the alternates.
- ozim 8mo agoExactly there was this study where they were trying to make LLM reproduce HP book word for word like giving first sentences and letting it cook. Basically they managed with some tricks make 99% word for word - tricks were needed to bypass security measures that are there in place for exactly reason to stop people to retrieve training material.
- ck_one 8mo agoDo you remember how to get around those tricks?
- djhn 8mo agoThis is the paper: https://arxiv.org/abs/2601.02671 https://arxiv.org/abs/2601.02671 Grok and Deepmind IIRC didn’t require tricks.
- eek2121 8mo agoThis really makes me want to try something similar with content from my own website. I shut it down a while ago because the number of bots overtake traffic. The site had quite a bit of human traffic (enough to bring in a few hundred bucks a month in ad revenue, and a few hundred more in subscription revenue), however, the AI scrapers really started ramping up and the only way I could realistically continue would be to pay a lot more for hosting/infrastructure. I had put a ton of time into building out content...thousands of hours, only to have scrapers ignore robots, bypass cloudflare (they didn't have any AI products at the time), and overwhelm my measly infrastructure. Even now, with the domain pointed at NOTHING, it gets almost 100,000 hits a month. There is NO SERVER on the other end. It is a dead link. The stats come from Cloudflare, where the domain name is hosted. I'm curious if there are any lawyers who'd be willing to take someone like me on contingency for a large copyright lawsuit.
- camdenreslink 8mo agoThe new cloudflare products for blocking bots and AI scrapers might be worth a shot if you put so much work into the content.
- prawn 8mo agoFurther, some low effort bots can be quickly handled with CF by blocking specific countries (e.g., Brazil and Russia, for one of my sites).
- lobsterthief 8mo agoI work for a publisher that serves the Chinese market as a secondary market. Sucks that we can’t blanketly do this since we get hammered by Chinese bots daily. We also have an extremely old codebase (Drupal) which makes blanket caching difficult. Working to migrate from Cloudfront to Cloudflare at least
- pron 8mo agoThis reminds me of https://en.wikipedia.org/wiki/Pierre_Menard,_Author_of_the_Quixote https://en.wikipedia.org/wiki/Pierre_Menard,_Author_of_the_Q... : > Borges's "review" describes Menard's efforts to go beyond a mere "translation" of Don Quixote by immersing himself so thoroughly in the work as to be able to actually "re-create" it, line for line, in the original 17th-century Spanish. Thus, Pierre Menard is often used to raise questions and discussion about the nature of authorship, appropriation, and interpretation.
- ck_one 8mo agoWhen I tried it without web search so only internal knowledge it missed ~15 spells.