5 ms·
I would say the strawberry/o1 hype was even worse than the GPT-5 hype There were months worth of articles on how strawberry is considered almost dangerous inte
by shmatt 2y ago
I would say the strawberry/o1 hype was even worse than the GPT-5 hype
There were months worth of articles on how strawberry is considered almost dangerous internally its so smart. I know we only got -mini and -preview but...this doesn't feel like AGI
- lordswork 2y agoTo be fair, o1 is a major breakthrough in the field. If other AI labs can't crack scaling useful inference compute, OpenAI will maintain a big lead.
- lolinder 2y agoIsn't o1 just applying last year's Tree of Thoughts paper in production? Is there any reason to believe that the other companies will struggle to implement their own? https://github.com/princeton-nlp/tree-of-thought-llm https://github.com/princeton-nlp/tree-of-thought-llm
- impossiblefork 2y agoI don't think it's tree of thoughts at all. I think it's as they say: reinforcement learning applied to cause it to generate a relatively long 'reasoning trace' of some kind from which the answer is obtained through summarisation. I think it's likely a cleverly simplified version of QuietSTaR, with no thought tokens, just one big generation to which the RL is applied. The way I believe it's trained in practice is as follows: they have a bunch of examples, some at the edge of GPT-4s ability to answer, some beyond it, some that GPT-4 can answer if you're lucky with the randomness. Then they give it one of these prompts, generate a fairly long text, maybe 3x the length of the answer, and summarize that to produce the final answer. Then they use REINFORCE to reward the generated texts that increase the probability of the summary being correct.
- WJW 2y agoNot be be nitpicky, but being the first to deploy recent academic research papers to production should count as a breakthrough IMHO.
- lordswork 2y agoThere seems to be several components involved: tree of thoughts, MCTS, RL, high quality chain of thought data, possibly multiple models.
- njtransit 2y agoo1 seems like it’s basically 4o with some chain of thought bolted on. Personally, I don’t consider chain of thought a breakthrough, let alone a major one.
- lordswork 2y agoIt's much more than CoT. I suspect it will take other labs some time to replicate.
- petesergeant 2y ago> o1 is a major breakthrough Is it? I feel like if you don't care about the cost it's pretty easily replicable on any other LLM, just with a lang-chain sort of approach
- lordswork 2y agoThen why hasn't it been done?
- tim333 2y agoSomeone gave the various models an IQ test and the previous ones scored 80-90 so a bit dim compared to humans, and o1 got 120, quite bright for a human. https://www.reddit.com/r/ClaudeAI/comments/1fhwfyl/openai_1o_gets_120_iq_on_norway_mensa_iq_test https://www.reddit.com/r/ClaudeAI/comments/1fhwfyl/openai_1o... this may have consequences for how useful it is.
- aunty_helen 2y agoCoT can be _easily_ achieved using langgraph in a similar manner. There’s no “scaling of inference” it’s just prompting, all the way down.
- lordswork 2y agoIt's not just CoT
- jsheard 2y ago> There were months worth of articles on how strawberry is considered almost dangerous internally its so smart. Like clockwork, every time they need to drum up excitement: https://www.theverge.com/2019/11/7/20953040/openai-text-generation-ai-gpt-2-full-model-release-1-5b-parameters https://www.theverge.com/2019/11/7/20953040/openai-text-gene...
- noobermin 2y agoI was bashing my head into walls since 2017 or so when people were saying AI will eat the world and we have to worry about non-alignment and I felt insane realizing no one else even asked if it was manufactured hype. People in my life to this day are still falling for these tactics despite, to me, seeming bright regarding everything else. To be clear, it is true that transformers did change things but the merchants are still over selling it and everyone else laps it up while not meta-thinking about it for even one second.
- ben_w 2y agoIt may be hype, but there's plenty of solid logic behind the general case. There's also a huge range of practical demonstrations of non-aligned, monomaniacal and not particularly smart intelligences, that literally eat humans: bacteria. (Also lions and tigers and bears, if you want to insist that evolution doesn't itself count as intelligence).
- tim333 2y agoIt's like in the 1990s - (a) the internet will revolutionize commerce so (b) you should buy Webvan. One can be true without the other and (a) took a while. But a lot of people bought Webvan and similar back in the day.
- CephalopodMD 2y agoTotally agree. It took me a full week before I realized that the Strawberry/o1 model was the mysterious Q* Sam Altman has been hyping up for almost a full year since the openai coup, which... is pretty underwhelming tbh. It's an impressive incremental advancement for sure! But it's really not the paradigm shifting gpt-5 worthy launch we were promised. Personal opinion: I think this means we've probably exhausted all the low hanging fruit in LLM land. This was the last thing I was reserving judgement for. When the most hyped up big idea openai has rn is basically "we're just gonna have the model dump out a wall of semi-optimized chain of thought every time and not send it over the wire" we're officially out of big ideas. Like I mean it obviously works... but that's more or less what we've _been_ doing for years now! Barring a total rethinking of LLM architecture, I think all improvements going forward will be baby steps for a while, basically moving at the same pace we've been going since gpt-4 launched. I don't think this is the path to AGI in the near term, but there's still plenty of headroom for minor incremental change. By analogy, i feel like gpt-4 was basically the same quantum leap we got with the iphone 4: all the basic functionality and peripherals were there by the time we got iphone 4 (multitasking, facetime, the app store, various sensors, etc.), and everything since then has just been minor improvements. The current iPhone 16 is obviously faster, bigger, thinner, and "better" than the 4, but for the most part it doesn't really do anything extra that the 4 wasn't already capable of at some level with the right app. Similarly, I think gpt-4 was pretty much "good enough". LLMs are about as they're gonna get for the next little while, though they might get a little cheaper, faster, and more "aligned" (however we wanna define that). They might get slightly less stupid, but i don't think they're gonna get a whole lot smarter any time soon. Whatever we see in the next few years is probably not going to be much better than using gpt-4 with the right prompt, tool use, RAG, etc. on top of it. We'll only see improvements at the margins.
- basch 2y agochat has become a limiting factor. its both too linear, and hard to revise. it's hard to undo parts of the conversation that poison it. it's too hard to save important bits that shouldn't be forgotten or drowned out. its not word processor like enough. i envision the next generation of these products being multipane by default. on the left I have a chat, in the center I have a whiteboard, and on the right I have a rendered document. throwing a clip of something onto the whiteboard makes it modifiable by chat. "store this, categorize it, summarize it, place it in the document." whatever comes next needs to function more like onenote or obsidian. just to use an example from current events today, lets say I want to make a Parody Dossier on Walz, similar to todays Vance leak. I should be able to describe the project to chat. It builds a document structure. I tell it we are going to scrape all of the internets jokes on Walz at a bbq or other non-scandals. I should be able to quickly click through a table of contents, and "chat" with each paragraph. "This one needs fleshing out, this one needs summarization." As we scrape a reddit post, we want to incorporate not only the original post, but all the best comments. I should be a be able to "chat with the document editor" and put together a 200 page document in the amount of time it took me to write this post, just by describing what is and isnt working, and dragging and dropping. chat, the simplicity of it, understanding of complex sentences, and multi sentence conversations was a UI paradigm leap. it went well past keyword search and the new command line of the internet. its a great first step, and a nice reset after a decade of interface stagnation, but now its ubiquity and simplicity, like the search box, is clouding peoples imagination and ability to dream up the next new interactive interface, which I expect to involve more mouse and visual relationships. tldr: the llm is a component of the next generation interface, not the entire interface itself.
- axpvms 2y agospeaking of which, try asking ChatGPT how many r's are in strawberry