10 ms·
Nah, it's already being done for GPT-3's competitors and will likely be done soon for GPT-4's competitors https://arstechnica.com/information-technology/2023/0
by sounds 4y ago
Nah, it's already being done for GPT-3's competitors and will likely be done soon for GPT-4's competitors
https://arstechnica.com/information-technology/2023/03/you-can-now-run-a-gpt-3-level-ai-model-on-your-laptop-phone-and-raspberry-pi/ https://arstechnica.com/information-technology/2023/03/you-c...
- systemvoltage 4y agoCurious why even companies at the very edge of innovation are unable to build moats? I know nothing about AI, but when DALLE was released, I was under the impression that the leap of tech here is so crazy that no one is going to beat OpenAI at it. We have a bunch now: Stable Diffusion, MidJourney, lots of parallel projects that are similar. Is it because OpenAI was sharing their secret sauce? Or is it that the sauce isn’t that special?
- elevaet 4y agoI think it's because everyone's swimming in the same bath. People move around between companies, things are whispered, papers are published, techniques are mentioned and details filled in, products are backwards-engineered. Progress is incremental.
- PaulHoule 4y agoGoogle got a patent on transfomers but didn't enforce it. If it wasn't for patents you'd never get a moat from technology. Google, Facebook, Apple and all have a moat because of two sided markets: advertisers go where the audience is, app makers go where the users are. (There's another kind of "tech" company that is wrongly lumped in with the others, this is an overcapitalized company that looks like it has a moat because it is overcapitalized and able to lose money to win market share. This includes Amazon, Uber and Netflix.)
- mgfist 4y agoI don't think this is strictly true, though it's rare. The easiest example is the semiconductor industry. ASML's high end lithography machines are basically alien and cannot be reproduced by anyone else. China has spent billions trying. I don't even think there's a way to make the IP public because of how much of it is in people's heads and in the processes in place. I wonder how much money, time and ASML resources it would take to stand up a completely separate company that can do what ASML does assuming that ASML could dedicate 100% of their time in assisting in training the personnel at said company.
- PaulHoule 4y agoYeah, this is probably also true for TSMC, Intel and ARM. Look how slow progress is on RISC-V on the high end despite RISC-V having the best academic talent.
- throwaway2037 4y agoI would also add Samsung semi to that list. As I understand, for the small nodes, everyone is using ASML. That's a bit scary to me. About RISC-V: What does you think is different about RISC-V vs ARM? I can only think that ARM has been used in the wild for longer, so there is a meaningful feedback loop. Designers can incorporate this feedback into future designs. Don't give up hope on RISC-V too soon! It might have a place in IoT which needs more diverse compute.
- kybernetyk 4y ago>despite RISC-V having the best academic talent. academic performance is a bad predictor for real world performance
- varjag 4y agoIt's a decent predictor of real world performance just not a perfect one.
- pclmulqdq 4y agoUnfortunately, RISC-V, despite the "open source" marketing, is still basically dominated by one company (SiFive) that designs all the commercial cores. They also employ everyone who writes the spec, so the current "compiled" spec document is about 5 years behind the actual production ISA. Intel and others are trying to break this monopoly right now. Compare this to the AI ecosystem and you get a huge difference. The architecture of these AI systems is pretty well-known despite not being "open," and there is a tremendous amount of competition.
- shiftingleft 4y ago
- light_hue_1 4y ago> Google got a patent on transfomers but didn't enforce it. Google's Transformer patent isn't relevant to GPT at all. https://patents.google.com/patent/US10452978B2/en https://patents.google.com/patent/US10452978B2/en They patented the original Transformer encoder-decoder architecture. But most modern models are built either only out of encoders (the BERT family) or only out of decoders (the GPT family). Even if they wanted to enforce their patent, they couldn't. It's a classic problem with patenting things that every lawyer warns you about "what if someone could make a change to circumvent your patent".
- varjag 4y agoYou can't tell unless you read the claims thoroughly. Degenerate use cases can be covered by general claims.
- light_hue_1 4y agoIndeed. I read the claims. You can too. They're short.
- varjag 4y agoAre you kidding? There are 30 claims, it's an hours' work to make complete sense of how these work together and what they possibly do/do not cover. I've filed my own patents so have read thru enough of prior art and am not doing it for a pointless internet argument.
- versteegen 4y agoIANAL. I looked through the patent, not just the Claims. I certainly didn't read all of it. But while it leaves open many possible variations, it's a patent for sequence transduction and it's quite explicit everywhere that the system comprises a decoder and an encoder (see Claim 1, the most vague) and nowhere did I see any hint that you could leave out one or the other or that you could leave out the encoder-decoder attention submodule (the "degenerate use-case" you suggested). The patent is only about sequence transduction (e.g. in translation). Now an encoder+decoder is very similar to a decoder-only transformer, but it's certainly an inventive step to make that modification and I'm pretty sure the patent doesn't contain it. It does describe all the other pieces of a decoder/encoder-only transformer though, despite not being covered by any of the claims, and I have no idea what a court would think about that since IANAL.
- sokoloff 4y agoOr, Amazon, Uber, and Netflix have access to so much capital based on investors' judgment that they will be able to win and protect market share by effective execution, thereby creating a defensible moat.
- Tanjreeve 4y agoI think his point was that If that moat doesn't exist without the ongoing context of more money being thrown at it then it isn't a moat.
- panzi 4y agoIsn't MidJourney a fork of Stable Diffusion?
- light_hue_1 4y agoIt's because moving forward is hard, but moving backward when you know what the space of answers is, is much easier. Once you know that OpenAI gets a certain set of results with roughly technology X, it's much easier to recreate that work than to do it in the first place. This is true of most technology. Inventing the telephone is something, but if you told a competent engineer the basic idea, they'd be able to do it 50 years earlier no problem. Same with flight. There are some really tricky problems with counter-intuitive answers (like how stalls work and how turning should work; which still mess up new pilots today). The space of possible answers is huge, and even the questions themselves are very unclear. It took the Wright brothers years of experiments to understand that they were stalling their wing. But once you have the basic questions and their rough answers, any amateur can build a plane today in their shed.
- zamnos 4y agoI agree with your overall point, but I don't think that we'd be able to get the telephone 50 years earlier because of how many other industries had to align to allow for its invention. Insulated wire didn't readily or cheaply come in spools until after the telegraph in the 1840's. The telephone was in 1876 so 50 years earlier was 1826.
- hnick 4y agoYou didn't mention it explicitly but I think the morale factor is also huge. Once you know it's possible, it does away with all those fears of wasted nights/weekends/resources/etc for something that might not actually be possible.
- taneq 4y agoI'm not sure how "keep the secret sauce secret and only offer it as a service" isn't a moat? Here the 'secret sauce' is the training data and the trained network, not the methodology, but the way they're going, it's only a matter of time before they start withholding key details of the methodology too.
- kybernetyk 4y agoLuckily ML isn't that complicated. People will find out stuff without the cool kids at OpenAI telling them.
- kybernetyk 4y ago>Or is it that the sauce isn’t that special? Most likely this.
- sounds 4y agoOpenAI can't build a moat because OpenAI isn't a new vertical, or even a complete product. Right now the magical demo is being paraded around, exploiting the same "worse is better" that toppled previous ivory towers of computing. It's helpful while the real product development happens elsewhere, since it keeps investors hyped about something. The new verticals seem smaller than all of AI/ML. One company dominating ML is about as likely as a single source owning the living room or the smartphones or the web. That's a platitude for companies to woo their shareholders and for regulators to point at while doing their job. ML dominating the living room or smartphones or the web or education or professional work is equally unrealistic.
- raducu 4y agoI also expect a high moat, especially regarding training data. But the counter for the high moat would be the atomic bomb -- the soviets were able to build it for a fraction of what it cost the US because the hard parts were leaked to them. GPT-3 afik is an easier picking because they used a bigger model than necessary, but afterwards there appeared guidelines about model size vs. training data, so GPT-4 probably won't be as easily trimmed down.
- hoseja 4y agoThe sauce really doesn't seem all that special.
- dr_dshiv 4y agoBecause we are headed to a world of semi-automated luxury socialism. Having a genius at your service for less than $1000 per year is just an insane break to the system we live in. We all need to think hard about how to design the world we want to live in.
- siva7 4y agoYou can have the most special sauce in the world but if you're hiding it in the closet because you fear that it will hurt sales of your classic sauce then don't be surprised with what will happen (also known as Innovators Dilemma)
- usrbinbash 4y ago> Or is it that the sauce isn’t that special? The sauce is special, but the recipe is already known. Most of the stuff things like LLMs are based on comes from published research, so in principle coming up with the architecture that can do something very close, is doable to everyone with the skills to understand the research material. The problems start with a) taking the architecture to a finished and fine tuned model and b) running that model. Because now we are talking about non-trivial amounts of compute, storage and bandwidth, so quite simple resources suddenly become a very real problem.