4 ms·
A few important things to remember here: The best engineering minds have been focused on scaling transformer pre and post training for the last three years bec
by iandanforth 2y ago
A few important things to remember here:
The best engineering minds have been focused on scaling transformer pre and post training for the last three years because they had good reason to believe it would work, and it has up until now.
Progress has been measured against benchmarks which are / were largely solvable with scale.
There is another emerging paradigm which is still small(er) scale but showing remarkable results. That's full multi-modal training with embodied agents (aka robots). 1x, Figure, Physical Intelligence, Tesla are all making rapid progress on functionality which is definitely beyond frontier LLMs because it is distinctly different.
OpenAI/Google/Anthropic are not ignorant of this trend and are also reviving or investing in robots or robot-like research.
So while Orion and Claude 3.5 opus may not be another shocking giant leap forward, that does not mean that there arn't giant shocking leaps forward coming from slightly different directions.
- joe_the_user 2y agoTesla are all making rapid progress on functionality which is definitely beyond frontier LLMs because it is distinctly different Sure, that's tautologically true but that doesn't imply that beyondness will lead to significant leaps that offer notable utility like LLMs. Deep Learning overall has been a way around the problem that intelligent behavior is very hard to code and no wants to hire many, many coders needed to do this (and no one actually how to get a mass of programmers to actually be useful beyond a certain of project complexity, to boot). People take the "bitter lesson" to mean data can do anything but I'd say a second bitter lesson is that data-things are the low hanging fruit. Moreover, robot behavior is especially to fake. Impressive robot demos have been happening for decades without said robots getting the ability to act effectively in the complex, ad-hoc environment that human live in, IE, work with people or even cheaply emulate human behavior (but they can do choreographed/puppeteered kung fu on stage).
- hobs 2y agoAnd worth noting that Tesla faked a ton of its robot footage already, they might be making progress but their physical human robotics does not seem advanced at the moment.
- ben_w 2y agoIndeed. Even assuming the recent robot demo was entirely AI, the only single thing they demonstrated that would have been noteworthy was isolating one voice in a noisy crowd well enough to respond; everything else I saw Optimus do, has already been demonstrated by others. What makes the uncertainty extra sad, is that a remote controllable humanoid robot is already directly useful for work in hazardous environments, and we know they've got at least that… but Musk would rather it be about the AI.
- hereme888 2y agoAre we humans so different? Why do you wear what you wear? People emulate their older siblings, and so learn behavior. LLMs can create new programs, after having initially learned similar examples from others. Likewise for AI media.
- knicholes 2y agoOnce we've scraped the internet of its data, we need more data. Robots can take in video/audio data 24/7 and can be placed in your house to record this data by offering services like cooking/cleaning/folding laundry. Yeah, I'll pay $20k to have you record everything that happens in my house if I can stop doing dishes for five years!
- triyambakam 2y agoOr get a dishwashing machine?
- hartator 2y agoWhy 5 years?
- fifilura 2y agoFive years, that's all we've got. https://en.m.wikipedia.org/wiki/Five_Years_(David_Bowie_song) https://en.m.wikipedia.org/wiki/Five_Years_(David_Bowie_song...
- bredren 2y agoBecause whatever org fills this space will be working on ARR.
- exe34 2y agothat's when the robot takes his job and he can't afford the robot anymore.
- twelve40 2y ago> OpenAI has announced a plan to achieve artificial general intelligence (AGI) within five years, an ambitious goal as the company works to design systems that outperform humans.
- knicholes 2y agoNo real reason. I just made it up. But that's kind of my reasonable expectation of longevity of a machine like a robotic lawnmower and battery life.
- eli_gottlieb 2y ago>The best engineering minds have been focused on scaling transformer pre and post training for the last three years The best minds don't follow the herd.
- slashdave 2y agoI hear what you are saying, but "innovation" is also often used to excuse some rather badly engineered concepts
- demosthanos 2y ago> that does not mean that there arn't giant shocking leaps forward coming from slightly different directions. Nor does it mean that there are! We've gotten into this habit of assuming that we're owed giant shocking leaps forward every year or so, and this wave of AI startups raised money accordingly, but that's never how any innovation has worked. We've always followed the same pattern: there's a breakthrough which causes a major shift in what's possible, followed by a few years of rapid growth as engineers pick up where the scientists left off, followed by a plateau while we all get used to the new normal. We ought to be expecting a plateau, but Sam Altman and company have done their work well and have convinced many of us that this time it's different. This time it's the singularity, and we're going to see exponential growth from here on out. People want to believe it, so they do, and Altman is milking that belief for all it's worth. But make no mistake: Altman has been telegraphing that he's eyeing the exit, and you don't eye the exit when you own a company that's set to continue exponentially increasing in value.
- deleted 2y ago[deleted]
- lcnPylGDnU4H9OF 2y ago> Altman has been telegraphing that he's eyeing the exit Can you think of any specific examples? Not trying to express disbelief, just curious given that this is obviously not what he's intending to communicate so it would be interesting to examine what seemed to communicate it.
- tim333 2y agoYeah, listening to him last week he seemed very unlike that https://www.youtube.com/watch?v=xXCBz_8hM9w&t=2324s https://www.youtube.com/watch?v=xXCBz_8hM9w&t=2324s
- sincerecook 2y ago> That's full multi-modal training with embodied agents (aka robots). 1x, Figure, Physical Intelligence, Tesla are all making rapid progress on functionality which is definitely beyond frontier LLMs because it is distinctly different. Cool, but we already have robots doing this in 2d space (aka self driving cars) that struggle not to kill people. How is adding a third dimension going to help? People are just refusing to accept the fact that machine learning is not intelligence.
- warkdarrior 2y ago> Cool, but we already have robots doing this in 2d space (aka self driving cars) that struggle not to kill people. How is adding a third dimension going to help? If we have robots that operate in 3D, they'll be able to kill you not only from behind or from the side, but also from above. So that's progress!
- akomtu 2y agoMy understanding is that machine learning today is a lot like interpolation of examples in the dataset. The breakthrough of LLMs is due to the idea that interpolation in a 1024-dimensional space works much better than in a 2d space, if we naively interpolated English letters. All the modern transformers stuff is basically an advanced interpolation method that uses a large local neighborhood than just few nearest examples. It's like the Lanczos interpolation kernel, using a 1d analogy. Increasing the size of the kernel won't bring any gains, because the current kernel already nearly perfectly approximates an ideal interpolation (a full dataset DFT). However interpolation isn't reasoning. If we want to understand the motion of planets, we would start with a dataset of (x, y, z, t) coordinates and try to derive the law of motion. Imagine if someone simply interpolated the dataset and presented the law of gravity as an array of million coefficients (aka weights)? Our minds have to work with a very small operating memory that can hardly fit 10 coefficients. This constraint forces us to develop intelligence that compacts the entire dataset into one small differential equation. Btw, English grammar is the differential equation of English in a lot of ways: it tells what the local rules are of valid trajectories of words that we call sentences.
- tick_tock_tick 2y ago
- rafaelmn 2y ago>There is another emerging paradigm which is still small(er) scale but showing remarkable results. That's full multi-modal training with embodied agents (aka robots). 1x, Figure, Physical Intelligence, Tesla are all making rapid progress on functionality which is definitely beyond frontier LLMs because it is distinctly different. Tesla is selling this view for almost a decade now in self-driving - how their car fleet feeding training data is going to make them leaders in the area. I don't find it convincing anymore
- torguyvg46787 2y agoThe approaches are very limited, and it's essentially artificial artificial AI (and need a lot of human teleop demos). At CoRL last week, the progress has noticeably plateaued. Roboticists notably were pessimistic that scaling laws will apply to robotics because of the embodiment issues.
- Dunedan 2y agoWhile one could argue whether Tesla or another company is the leader in this space, don't all promising self-driving approaches rely on this paradigm?
- mvdtnz 2y ago> The best engineering minds have been focused on scaling transformer pre and post training for the last three years because they had good reason to believe it would work, and it has up until now. Or because the people running companies who have fooled investors into believing it will work can afford to pay said engineers life-changing amounts of money.
- slashdave 2y agoThe improvements in transformer implementation (e.g. "Flash Attention") have saved gobs of money on training and inference, I am guessing most likely more than the salary of those researchers.
- slashdave 2y ago> Tesla are all making rapid progress on functionality The lack of progress with self driving seems to indicate that Tesla has a serious problem with scaling. The investment in enormous compute resources is another red flag (if you run out of ideas, just use brute force). This points to a fundamental flaw in model architecture.
- airstrike 2y agoThe gap from the virtual world of software and the brutally uncompromising nature of physical reality is wider than most people seem to accept. It's almost like saying "we've already visited every place on Earth, surely Mars is just around the corner now"