5 ms·
You're not seeing this the right way. You are saying the equivalent argument of: "Look at how much hand-holding this processor needs. We had to give it step by
by rytill 4y ago
You're not seeing this the right way. You are saying the equivalent argument of: "Look at how much hand-holding this processor needs. We had to give it step by step instructions on what program to execute. We are still a long way from computers automating any significant aspect of society."
LLMs are a primitive that can be controlled by a variety of higher level algorithms.
- Imnimo 4y agoThe "higher level algorithm" of "how to do abstract thought" is unknown. Even if LLMs solve "how to do language", that was hardly the only missing piece of the puzzle. The fact that solving the language component (to the extent that ChatGPT 'solves' it) results in an agent that needs so much hand-holding to interact with a very simple simulated world shows how much is left to solve.
- famouswaffles 4y agoYou've been told it doesn't need that much handholding. https://arxiv.org/abs/2303.11366 https://arxiv.org/abs/2303.11366 https://arxiv.org/abs/2303.17651 https://arxiv.org/abs/2303.17651 Why insist otherwise ?
- Imnimo 4y agoI don't understand what you intend these papers to demonstrate. Surely the fact that the level of hand-holding they propose (both Self-Refine and Reflexion offload higher-order reasoning to a hand-crafted process) is so helpful even on extremely simple tasks demonstrates that a great deal of hand-holding is required for complex tasks. That these techniques improve upon the baseline tells us that ChatGPT is incapable of doing this sort of simple higher-order thinking internally, and the fact that the augmented models still offer only middling performance on the target tasks suggests that "not that much handholding" (as you describe them) is insufficient.
- famouswaffles 4y agoMiddling performance ? Do you actually understand the benchmarks you saw ? assuming you even read it. 88% of human eval is not middling lmao. Fuck, i really have seen everything.
- Imnimo 4y agoI don't see a benchmark in either paper that shows "88% of human eval". Which table or figure are you looking at?
- famouswaffles 4y agoIt's with reflexion https://twitter.com/johnjnay/status/1639362071807549446 https://twitter.com/johnjnay/status/1639362071807549446
- Imnimo 4y agoBut this is not raw Reflexion (it's not a result from the paper, but rather from follow-on work). The project uses significantly more scaffolding to guide the agent in how to approach the code generation problem. They design special prompts including worked examples to guide the model to generate test cases, prompt it to generate a function body, run the generated code through the tests, off-load the decision of whether to submit the code or to try to refine to hand-crafted logic, collate the results from the tests to make self-reflection easier, and so on. This is hardly an example of minimal hand-holding. I'd go so far as to say this is MORE handholding than the paper this thread is about.
- famouswaffles 4y agoI guess we just have different meanings of hand holding then.
- famouswaffles 4y agofor me, an unsupervised pipeline is not handholding. the thoughts drive actions. If you can't control how those thoughts form or process memories then i don't see what is hand holding about it. a pipeline is one and done.
- deleted 4y ago[deleted]
- losteric 4y agoFraming LLMs as primitives is marketing-speak. These are high-level construction for specific runtimes, which are difficult to test and subject to change at anytime.
- catlifeonmars 4y agoHah. Sounds like qubits.
- rytill 4y agoDoes a primitive definitely need to be easy to test or deterministic?