5 ms·
Saying “hey don’t go down the path we are on, where we are making money and considered the best in the world.. it’s a dead end” rings pretty hollow.. like “don’
by xt00 3y ago
Saying “hey don’t go down the path we are on, where we are making money and considered the best in the world.. it’s a dead end” rings pretty hollow.. like “don’t take our lunch please?” Might be a similar statement it feels..
- whywhywhywhy 3y agoEveryone hoping to compete with OpenAI should have an "Always do the opposite of what Sam says" sign on the wall.
- Freire_Herval 3y ago[dead]
- famouswaffles 3y agoIt's a pretty sus argument for sure when they're scared to release even parameter size. although the title is a bit misleading on what he was actually saying. still, there's a lot left to go in terms of scale. Even if it isn't parameter size(and there's still lots of room here too, it just won't be economical), contrary to popular belief, there's lots of data left to mine
- thewataccount 3y agoNah - GPT-4 is crazy expensive, paying 20$/mo only get's you 25messages/3hours and it's crazy slow. The api is rather expensive too. I'm pretty sure that GPT-4 is ~1T-2T parameters, and they're struggling to run it(at reasonable performance and profit). So far their strategy has been to 10x the parameter count every GPT generation, and the problem is that there's diminishing returns everytime they do that. AFAIK they've now resorted to chunking GPT through the GPUs because of the 2 to 4 terabytes of VRAM required (at 16bit). So now they've reached the edge of what they can reasonably run, and even if they do 10x it the expected gains are less. On top of this, models like LLaMa have shown that it's possible to cut the parameter count substantially and still get decent results (albiet the opensource stuff still hasn't caught up). On top of all of this, keep in mind that at 8bit resolution 175B parameters (GBPT3.5) requires over 175GB of VRAM. This is crazy expensive and would never fit on consumer devices. Even if you use quantization and use 4bit, you still need over 80GB of VRAM. This definitely is not a "throw them off the trail" tactic - in order for this to actually scale the way everyone envisions both in performance and running on consumer devices - research HAS to be on improving the parameter count. And again there's lots of research showing its very possible to do. tl;dr: smaller = cheaper+faster+more accessible+same performance
- haxton 3y agoI don't think this argument really holds up. GPT3 on release was more expensive ($0.06/1000 tokens vs $0.03 input and $0.06 output for GPT4). Reasonable to assume that in 1-2 years it will also come down in cost.
- thewataccount 3y ago> Reasonable to assume that in 1-2 years it will also come down in cost. Definitely. I'm guessing they used something like quantization to optimize the vram usage to 4bit. The thing is that if you can't fit the weights in memory then you have to chunk it and that's slow = more gpu time = more cost. And even if you can fit it in GPU memory, less memory = less gpus needed. But we know you _can_ use less parameters, and that the training data + RLHF makes a massive difference in quality. And the model size linearly relates to the VRAM requirements/cost. So if you can get a 60B model to run at 175B's quality, then you've almost 1/3rd your memory requirements, and can now run (with 4bit quantization) on a single A100 80GB which is 1/8th the previously known 8x A100's that GPT-3.5 ran on (and still half GPT-3.5+4bit). Also while openai likely doesn't want this - we really want these models to run on our devices, and LLaMa+finetuning has shown promising improvements (not their just yet) at 7B size which can run on consumer devices.
- ericmcer 3y agoYeah I am noticing this as well. GPT enables you to do difficult things really easily, but then it is so expensive you would need to replace it with custom code for any long term solution. For example: you could use GPT to parse a resume file, pull out work experience and return it as JSON. That would take minutes to setup using the GPT API and it would take weeks to build your own system, but GPT is so expensive that building your own system is totally worth it. Unless they can seriously reduce how expensive it is I don't see it replacing many existing solutions. Using GPT to parse text for a repetitive task is like using a backhoe to plant flowers.
- beachy 3y ago> For example: you could use GPT to parse a resume file, pull out work experience and return it as JSON. That would take minutes to setup using the GPT API and it would take weeks to build your own system, but GPT is so expensive that building your own system is totally worth it. True, but an HR SaaS vendor could use that to put on a compelling demo to a potential customer, stopping them from going to a competitor or otherwise benefiting. And anyway, without churning the numbers, for volumes of say 1M resumes (at which point you've achieved a lot of success) I can't quite believe it would be cheaper to build something when there is such a powerful solution available. Maybe once you are at 1G resumes... My bet is still no though.