5 ms·
This article is misleadingly conflating short-term analyses with long-term outcomes. The idea that "reasoning doesn't matter" in the long-term is absolutely as
by mrshadowgoose 3y ago
This article is misleadingly conflating short-term analyses with long-term outcomes.
The idea that "reasoning doesn't matter" in the long-term is absolutely asinine. Human-level general reasoning is obviously one of the coveted goals of AI research.
It remains unclear how the open-source community is ever going to amass the tens of millions required to train foundational models. And if it somehow does, no sane government will permit uncontrolled research towards AGI.
- logicchains 3y ago>And if it somehow does, no sane government will permit uncontrolled research towards AGI If there was an effective, distributed means for training LLMs, enough people are passionate about LLMs that the only way governments could stop it is if every country in the world turned into communist China with respect to internet restrictions.
- api 3y agoDistributed training without insane bandwidth requirements is the holy grail here. It has to work in terms of work units like Folding@Home.
- mrshadowgoose 3y ago> If there was an effective, distributed means for training LLMs Is it reasonable to believe that this is possible? Distributed training requires extremely high-bandwidth and low-latency interconnect. The internet ain't that. Believe me, I dearly want the open-source community to "win". A future where only governments have a monopoly on AI research is absolutely terrifying. Given the known parameters though, that future seems inevitable.
- max51 3y ago>Is it reasonable to believe that this is possible? Distributed training requires extremely high-bandwidth and low-latency interconnect. The internet ain't that. The same can be said about piracy and bittorrent. The community torrents is so large that I would bet a lot money they have more computing+bandwidth resources combined than openAI by a wide margin, probably orders of magnitude.
- TeMPOraL 3y agoWhy low latency? Is the latency here a meaningfully annoying limit in training, that can't be reasonably offset by just adding more compute nodes to the network? Latency is the limit in the end, but I feel that there's plenty of easy-ish wins to have in redesigning the architecture and training approach to make it an irrelevant in practice.
- danShumway 3y ago> The idea that "reasoning doesn't matter" in the long-term is absolutely asinine. Human-level general reasoning is obviously one of the coveted goals of AI research. I don't know if I agree with this. Human-level reasoning is what researchers care about, but an AI that is controllable and that is consistent is way more valuable than an AI that is smart. And I think the general point here is that capabilities have outpaced control. There is a huge gulf between the capabilities of current AI and the ability to actually manipulate and utilize that AI. If you could somehow theoretically build an AI that was half as smart as GPT-4 but that was completely immune to prompt injection in every single situation, it would be more useful than GPT-4. See also Stability AI vs Midjourney, etc... Midjourney is far more capable but it simply doesn't matter -- input methods and control methods and the ability to fine-tune are more important than base model capabilities. Current models are quite capable for the tasks they're being used for, the reason they fall over and the reason why it's difficult to use them in those tasks is specifically because of the lack of control and reliability; and making them smarter seems to be only making them harder to control. If you go long-long-term then we get into science fiction territory and it's easy to say that human-level general reasoning is the highest priority. But that's because when you think about that long-term you are not thinking about tradeoffs. You're assuming a theoretical world where human-level reasoning is perfectly controllable and doesn't give wildly inconsistent results that make it useless for critical tasks and that the safeguards you need to put around it don't make it harder to work with than a human being. And yeah, if you can have literally everything, sure, you want human-level reasoning. But there are a lot of things in AI that matter a lot more than human-level reasoning and it's skipping a lot to say "human level reasoning is the most important" and to just assume that the other issues will get sorted out. Most businesses using AI are not using it because they're invested in replicating humanity, they're using it because they want to accomplish a specific task. If an AI replicates humanity but is bad at that task, businesses will go with the tool that's good at that task instead. And researchers will be disappointed because they want AGI, but successful businesses don't choose their tech stack based on what makes researchers happy. ---- > It remains unclear how the open-source community is ever going to amass the tens of millions required to train foundational models. I also think this is a little over-confident. This assumes that tens of millions are always going to be necessary to train foundational models, which I don't think is a safe bet to make. My impression looking at some of the more targeted work people are doing is that better curated data sets that are more focused on specific tasks may end up being straight-up better to train with than "the entire Internet". To add to that, I don't think it's a safe bet that model knowledge won't at some point be fully transferable between models or that these foundational models won't become a commodity. I mean, heck, we don't even know if the model weights are under copyright. It is very feasible that some kind of collective Open model might end up being good enough and that everyone just kind of standardizes on that as a base and builds on top of it. If that becomes useful enough that the companies investing into OpenAI decide "meh, we'll just invest resources/GPUs/etc to the Open version" then there will be a point where no VC-backed competitor will then be able to outpace the speed of those contributions, because they'll be racing alone against the entire market contributing to a single Open base. Even if none of that happens, it's also worth noting that the way we currently train LLMs is biased towards inefficiency, in part because research is primarily conducted by companies who can afford to be inefficient. But it's not a safe bet to assume that we won't find a better way to train models that uses less data and that doesn't try to recreate reasoning capabilities out of pure syntactic language -- a learning process that basically no living intelligent agent follows. Humans don't develop emergent reasoning from language, they learn language after developing reasoning and by mapping that language to real-world experiences; this is basically the complete opposite to how LLMs approach training. If different training methods get discovered, will they have the same ridiculous data requirements? I don't think I can say with complete confidence that they will. ---- > And if it somehow does, no sane government will permit uncontrolled research towards AGI. My issue here is that no sane government would trust AGI to a private corporation either. The realistic outcomes here are that either the government won't interfere, in which case Open Source communities will be able to do their own research the same as private companies, or the government will heavily interfere, in which case probably only state-developed AGIs will exist. A lot of businesses would love to have selective regulation, but the idea that AI is too dangerous to trust to hobbyists but is not too dangerous to put in the hands of people like Elon Musk or Sam Altman is ludicrous. OpenAI/Google aren't even responsible currently with their existing LLMs, I can only imagine how horribly they'd handle an actual AGI. If the government is currently shrugging its shoulders over the dumpster fire that is current LLM safety mechanisms, I'm not sure why they'd suddenly start caring about Open Source communities.