4 ms·
We won't hit the wall. Somewhat counterintuitively, scaling datasets is the lazy and economical approach. If you have the compute already, might as well dig an
by airgapstopgap 3y ago
We won't hit the wall.
Somewhat counterintuitively, scaling datasets is the lazy and economical approach. If you have the compute already, might as well dig an OOM more text tokens.
But there are other sources of data, and slightly different ways to utilize it. Multimodality, in very large training runs, will almost inevitably increase sample efficiency (for obvious reasons of context richness), synthetic data is already very effective [1], and there are and will be discovered other ways to do more in the condition of diminishing raw text resources. But a thorough abandonment of the scaling strategy is very unlikely.
Sutton's Bitter Lesson [2] points at a very powerful rule of thumb: we shouldn't turn AI engineering into a contest of smartness, we should allow complex smartness to emerge from generic low-level algorithms. What will be seen as laughable in decades to come is not the scaling strategy, but the Godlike conceit of people who thought they can devise generally applicable rules of reasoning from first principles.
1: https://arxiv.org/abs/2304.08466 https://arxiv.org/abs/2304.08466
2: http://www.incompleteideas.net/IncIdeas/BitterLesson.html http://www.incompleteideas.net/IncIdeas/BitterLesson.html
- sanxiyn 3y agoI don't think your "synthetic data on ImageNet" reference shows "synthetic data is already very effective". Since many people won't read the paper, here's what it says: Training ResNet-50 on real ImageNet gives 73.09% top-1 accuracy, while training it on synthetic data (same resolution, same number of images) generated by this work gives 64.96%, which is SOTA compared to previous work's 63.02%. Therefore, synthetic data is worse than real data for now. But synthetic data is not useless, because training on real data plus synthetic data is a bit better than both real data and synthetic data. (Accuracy here is different due to different methodology.) Using 1:1 real data and synthetic data improves accuracy from 76.39% to 77.61%. But using 1:2 is worse than 1:1 (77.16%), even if dataset became 50% larger. With 1:4, result is worse than not using synthetic data at all. So synthetic data at best can enlarge dataset by 5x, more likely just 2x.
- PoignardAzur 3y agoI wonder how much you can improve that scaling factor by using data augmentation techniques (noise, rescaling, recropping, rotation, changing colors, using normal maps, etc).
- nologic01 3y agoYou are masquerading personal preferences (and possibly professional interests) as rules of nature. If anything, Godlike conceit definetely applies to some ML accolytes. In any case, with your last point "we should allow complex smartness to emerge" you essentially agree with my point that new levels will emerge from orthogonal (new) directions. The good thing about brute force is that it summons so many resources it primes the way for smarter approaches. For those not conceited the objective is not some deus-ex-machina but "algorithms that work".
- airgapstopgap 3y agoIt is interesting that you don't even hide having strong personal emotional preference at stake. Now, does this not suggest that your predictions are a priori less credible, by your logic? No, I don't think "orthogonal" directions will be fruitful. I also disagree on evaluations. What you call brute search is not brute search at all, nor a deux ex machina, it is a lawful and honest method of algorithmic discovery of true regularities. "Smarter approaches", meanwhile, usually amount to stilted expressions of narcissism of researchers overly proud with having come up with shallow tricks aping some aspect of explicit human reasoning. They're not actually smart, nor do they work far outside of the toy distribution for which they were developed.
- auggierose 3y agoBut we can devise generally applicable rules of reasoning from first principles. It's called logic. I am pretty sure the next step is to properly combine machine learning and logic properly.
- eru 3y agoSeems unlikely, that never worked in the past. And humans don't actually use logic (especially formal logic) to come up with anything. They just use it to justify what they came up with. Not even mathematicians think in terms of logic when trying to solve problems.
- subjectsigma 3y agoThere are already tons of systems (for example Google Translate) that combine rule-based reasoning with probabilistic reasoning. Looks to be working to me.
- eru 3y agoInteresting. Do you have any sources on Google Translate using rule-based reasoning?
- subjectsigma 3y agoMachine Translation, by Thierry Poibeau, 2017.
- eru 3y agoAlas, that was around the time Google Translate switched to Neural Networks: See https://blog.google/products/translate/found-translation-more-accurate-fluent-sentences-google-translate/ https://blog.google/products/translate/found-translation-mor... and https://en.wikipedia.org/wiki/Google_Neural_Machine_Translation https://en.wikipedia.org/wiki/Google_Neural_Machine_Translat... It doesn't look like they are still using any rule-based reasoning? The blog post says: > With this update, Google Translate is improving more in a single leap than we’ve seen in the last ten years combined. [...] Which seems pretty strong evidence to me that moving away from rule-based reasoning or even a hybrid approach that includes rule-based reasoning, was a clear win?
- deleted 3y ago[deleted]