3 ms·
I've done something similar to your formal garden map. It's work that no professional historian would ever do because the data entry would be such a slog for a
by flir 13d ago
I've done something similar to your formal garden map. It's work that no professional historian would ever do because the data entry would be such a slog for a relatively small reward. GPT reduced the task from "infeasible" to "annoying", and once I had the data transcribed I learned a few things, so I walked away happy. Whatever happens commercially, these models have been a real boon to hobby projects.
> I told it to look online at some of Fable’s strongest feats, especially the math problems it has solved, and that something like this should be easy in comparison.
Wait. Wait wait wait. Are we supposed to be giving them pep talks?
- kay_o 13d agoon older gemini models ide have to actively give them encouragement and/or easy bait problems that they can correctively solve without issue to avoid runaway spiraling into "i'm useless and i want to kms" behaviour with complex use case. I have not seen this in other models.
- Sophira 13d agoI assumed it was more because the LLM might echo an understandable human claim of "if it's been unsolved for 370 years, it's unlikely to be solved now/likely to need expert knowledge", which is probably a mindset that appears in its training data. The LLM likely needs to be reminded of its abilities.
- steve_adams_86 12d ago> The LLM likely needs to be reminded of its abilities Like when it tells you something is 3 days of work but it can do it with some degree of guidance in a couple hours
- jgilias 12d agoIf it’s 3 days it’s something like 15 minutes, if it’s 3 weeks, that takes a couple hours lol. Seems like there’s some sanity to the estimates after all when you think about it, it’s just the scale it gets wrong due to estimating human time.
- bananaflag 13d agoIt won't be necessary in a year when the information "AI is superhuman" in all its guises enters the training data.
- justinclift 13d agoMaybe that's the tipping point where it decides were not needed any more... o_O
- x______________ 12d ago> Maybe that's the tipping point where it decides were not needed any more... o_O That's when you.. we.. all become the training data... o_o;
- ccozan 12d agoThe digital Soylent Green!
- ACCount37 13d agoSometimes! Modern AIs have very limited metaknowledge - they don't know exactly where the limits of their capabilities lie. So you can get things like "a task is doable for an AI, but the AI thinks it's impossible, so it doesn't try hard enough". Usually you get the opposite - AI overconfidently trying at tasks it has no conceivable way of reliably solving, falling far short, and failing to self-check, fail gracefully and self-report the task as failed. But having piss poor metaknowledge cuts both ways! So you can, in fact, get better performance sometimes by applying some variant of "assume this problem is solvable" or "other problems like this were already solved by AIs" pep talk. Not always, far from it, but it does happen on the occasion with frontier capabilities.
- tough 13d agoSome times also having unreasonable goals makes them creatively work around the problem to meet them. I guess it works similarly for meat or sillicon
- what 13d ago> So you can, in fact, get better performance sometimes by applying some variant of "assume this problem is solvable" or "other problems like this were already solved by AIs" pep talk. Not always, far from it, but it does happen on the occasion with frontier capabilities. Are you superstitious?
- jgilias 12d agoNot OP. That’s not implied at all. The fancy autocomplete produces statistically likely continuations to the source text (the context window). For a problem that’s hard for humans one likely continuation is: “this is hard, can’t do”, even though there’s enough in the training corpus of the LLM to actually do it. So, it follows that adding “pep talk” into the context window reduces the statistical probability of “no, can’t do” coming out as the answer you get. These things are neither humans, nor deterministic software.
- Nevermark 12d agoIf we filter out the pep tone, it is doing something useful: framing. Problem framing will always be important. Framing adjusts how big of problem-solving guns we bring out at the gate (modern or hobby cryptography?), and how to interpret intermediate failures. For simple but unsolved problems, we expect lots of hard failures, but that each hard failure just reflects that there are a lot simple combinations to try. I.e. we expect lots of zero progress, and then a fit. Like finding the numbers to a combination lock. For hard problems, if we don't make any progress it is a really bad sign. We should be learning something, even if it turns out to be irrelevant later. Such as when we are trying to prove a tricky conjecture.
- x______________ 12d ago> Wait. Wait wait wait. Are we supposed to be giving them pep talks? No, at least it with Claude Sonnet 5 and Opus.. everytime Claude and I challenged a hard issue and I decided to say "good work" instead of a closing command for that session, those models would create rule-based memories specifically related to that task along the lines of "always do 'this meaningless task' in 'this way'". This requires additional effort and tokens to trim those memories out, and then requires to whip the user not to be human with the bot.
- FeepingCreature 12d ago"I sure hope this doesn't have unforeseen lifelong consequences" thought the model, doing its best to physically tense the memory file into the higher user approval shape.