14 ms·
Could you train a ChatGPT-beating model for $85k and run it in a browser?
- TMWNN 4y agoHey, that means it can be turned into an Electron app!
- brrrrrm 4y agoThe WebGPU demo mentioned in this post is insane. Blows any WASM approach out of the water. Unfortunately that performance is not supported anywhere but chrome canary (behind a flag)
- raphlinus 4y agoThis will be changing soon. I believe Chrome M113 is scheduled to ship to stable on May 2, and will support WebGPU 1.0. I agree it's a game-changing technology.
- agnokapathetic 4y ago> My friends at Replicate told me that a simple rule of thumb for A100 cloud costs is $1/hour. AWS charges $32/hr for an 8xA100s (p4d.24xlarge) which comes out to $4/hour/gpu. Yes you can get lower pricing with a 3 year reservation but thats not what this question is asking. You also need 256 nodes to be colocated on the same fabric -- which AWS will do for you but only if you reserve for years.
- sebzim4500 4y agoMaybe they are using spot instances? $1/hr is about right for those.
- thewataccount 4y agoAWS certainly isn't the cheapest for this, did they mention using AWS? Lamdba Labs is 12$/hr for 8xA100's, and there's others relatively close to this price on demand, I assume you can get a better deal if you contact them for a large project. Replicate themselves rent out GPU time so I assume they would definitely know as that's almost certainly the core of their business.
- IanCal 4y agoLambda labs charges about 11-12/hr for 8xA100.
- celestialcheese 4y agolambdalabs will let you do on-demand 8xa100 @ 80GB VRAM/GPU for $12/hr, or reserved @ $10.86/hr 8xA100 @ 40gb for $8/hr Replicate friend isn't far off.
- pavelstoev 4y agomodel-depending, you can train on lesser (cheaper) GPUs but system-level optimizations are needed. Which is what we provide at centml.ai
- lxe 4y agoKeep in mind that image transformer models like stable diffusion are generally smaller than language models, so they are easier to fit in wasm space. Also. you can finetune llama-7b on a 3090 for about $3 using LoRA.
- bitL 4y agoOnly for images. People want to generate videos next and those models will be likely GPT-sized.
- Metus 4y agoThere is a video model making the rounds on /r/stablediffusion and it is just a tiny bit larger than Stable Diffusion.
- isoprophlex 4y agoYou're not kidding! it's far from perfect, but pretty funny still... https://www.reddit.com/r/StableDiffusion/comments/126xsxu/nightmare_continues_octopus_dinner https://www.reddit.com/r/StableDiffusion/comments/126xsxu/ni... Too bad SD learned the Shutterstock watermark so well, lol
- bitL 4y agoIt's cool though not very stable in details over temporal axis.
- Metus 4y agoOf course the quality is horrible relative to a proper video, it just illustrates that txt2vid might not need 100B+ parameters.
- danielbln 4y agoGenerative image models don't use transformers, they're diffusion models. LLMs are transformers.
- ultrablack 4y agoIf you could, you should have done it 6 months ago.
- munk-a 4y agoI mean - is there a developer alive that'd be unable to write the nascent version of Twitter? I think that Twitter as a business exists entirely because of the concept - the code to cover the core functionality is absolutely trivial to replicate. I don't think this is a very helpful statement because actually finding the idea on what to build is the hard part - or even just believing it's possible. The company I work at has been using NLP for years now and we have a model that's great at what we do... but if you asked if we could develop that into a chatbot as functional as chatgpt two years ago you'd probably be met with some pretty heavy skepticism. Cloning something that has been proven possible is always easier than taking the risk building the first version with no real grasp of feasibility.
- rspoerri 4y agoSo cool it runs on a browser /sarcasm/ i might not even need a computer. Or internet when we are at it. It either runs locally or it runs on the cloud. Data could come from both locations as well. So it's mostly technically irrelevant if it's displaying in a browser or not. Except when it comes to usability. I don't get it why people love software running in a browser. I often close important tools i have not saved when it's in a browser. I cant have offline tools which work if i am in a tunnel (living in Switzerland this is an issue) . Or it's incompatible because i am running LibreWolf. /sorry to be nitpicking on this topic ;-)
- ftxbro 4y ago> I don't get it why people love software running in a browser. If you read the article, part of the argument was for the sandboxing that the browser provides. "Obviously if you’re going to give a language model the ability to execute API calls and evaluate code you need to do it in a safe environment! Like for example... a web browser, which runs code from untrusted sources as a matter of habit and has the most thoroughly tested sandbox mechanism of any piece of software we’ve ever created."
- rspoerri 4y agoOSX does app sandboxing as well (not everywhere). But yeah, you're right i only skimmed the content and missed that part.
- rspoerri 4y agoThinking about it... I don't know exactly about the browser sandboxing. But isn't it's purpose to prevent access to the local system, while it mostly leaves access to the internet open? Is that really a good way to limit and AI system's API access?
- simonw 4y agoThe same-origin policy in browsers defaults to preventing JavaScript from making API calls out to any domain other than the one that hosts the page - unless those other domains have the right CORS headers. https://developer.mozilla.org/en-US/docs/Web/Security/Same-origin_policy https://developer.mozilla.org/en-US/docs/Web/Security/Same-o...
- JasonZ2 4y agoDoes anyone know how the results from a 7B parameter model with bloomz.cpp (https://github.com/NouamaneTazi/bloomz.cpp https://github.com/NouamaneTazi/bloomz.cpp) compares to the 7B parameter Alpaca model with llama.cpp (https://github.com/ggerganov/llama.cpp https://github.com/ggerganov/llama.cpp)? I have the latter working on a M1 Macbook Air with very good results for what it is. Curious if bloomz.cpp is significantly better or just about the same.
- munk-a 4y agoA wonderful thing about software development is that there is so much reserved space for creativity that we have huge gaps between costs and value. Whether the average person could do this for 85k I'm uncertain of - but there is a very significant slice of people that could do it for well under 85k now that the ground work has been done. This leads to the hilarious paradox where a software based business worth millions could be built on top of code valued around 60k to write.
- prerok 4y agoNit: not to write but to run. The cost of development is not considered in these calculations.
- nico 4y ago> This leads to the hilarious paradox where a software based business worth millions could be built on top of code valued around 60k to write. Or the fact that software based businesses just took a massive hit in value overnight and cannot possibly defend such high valuations anymore. The value of companies is quickly going to shift from tech moats to brands. Think CocaCola - anyone can create a drink that tastes as good or better than coke, but it's incredibly hard to compete with the CocaCola brand. Now think what would have happened if CocaCola had been super expensive to make, and all of a sudden, in a matter of weeks, it became incredibly cheap. This is what happened to the saltpeter industry in 1909 when synthetic saltpeter was invented. The whole industry was extinct in a few years.
- ftxbro 4y agoHis estimate is that you could train a LLaMA-7B scale model for around $82,432 and then fine-tune it for a total of less than $85K. But when I saw the fine tuned LLaMA-like models they were worse in my opinion even than GPT-3. They were like GPT-2.5 or like that. Not nearly as good as ChatGPT 3.5 and certainly not ChatGPT-beating. Of course, far enough in the future you could certainly run one in the browser for $85K or much less, like even $1 if you go far enough into the future.
- icelancer 4y agoYeah, the constant barrage of "THIS IS AS GOOD AS CHATGPT AND IS PRIVATE" screeds from LLaMA-based marketing projects are getting ridiculous. They're not even remotely close to the same quality. And why would they be? I want the best LLMs to be open source too, but I'm not delusional enough to make insane claims like the hundreds of GitHub forks out there.
- robertlagrant 4y ago> I want the best LLMs to be open source too How do you do this without being incredibly wealthy?
- mejutoco 4y agoPooling resources a la SETI@home would be an interesting option I would love to see.
- simonw 4y agoMy understanding is that can work for model inference but not for model training. https://github.com/bigscience-workshop/petals https://github.com/bigscience-workshop/petals is a project that does this kind of thing for running inference - I tried it out in Google Collab and it seemed to work pretty well. Model training is much harder though, because it requires a HUGE amount of high bandwidth data exchange between the machines doing the training - way more than is feasible to send over anything other than a local network connection.
- version_five 4y agoIf you have ~100k to spend, aren't there options to buy a gpu rather than just blow it all on cloud? How much is an 8xA100 machine? 4xA100 is 75k, 8 is 140k https://shop.lambdalabs.com/deep-learning/servers/hyperplane/customize https://shop.lambdalabs.com/deep-learning/servers/hyperplane...
- dekhn 4y agoyou're comparing the capital cost of acquiring a GPU machine with the operational cost of renting one in the cloud. Ignoring the operational costs of on-prem hardware is pretty common, but those costs are significant and can greatly change the calculation.
- sounds 4y agoRemember to discount the tax depreciation for the hardware and deduct any potential future gains from either reselling it or using it.
- version_five 4y agoFor a server farm, sure, for one machine, I don't know. Assuming it plugs into a normal 15A circuit, and you have a we-work or something where you don't pay for power, is the operational cost of one machine really material?
- dekhn 4y agoit's hard to tell from what you're saying: you're planning on putting an ML infrastructure training server on a regular 15A circuit, not in a data center or machine room? And power is paid for by somebody else? My thinking about pricing doesn't include that option because I wouldn't just hook a server like that up to a regular outlet in an office and use it for production work. If that works for you- you can happily ignore my comments. But if you go ahead and build such a thing and operate it for a year, please let us know if there were any costs- either dollar or in suffering- associated with your decision [edit: adding in that the value of this machine also suggests it cannot live unattended in an insecure location, like an office] signed, person who used to build closet clusters at universities
- ushakov 4y agoNow imagine loading 3.9 GB each time you want to interact with a webpage
- KMnO4 4y agoYeah, I’ve used Jira.
- neilellis 4y ago:-)
- sroussey 4y ago10yrs from now models will be in the OS. Maybe even in silicon. No downloads required.
- swader999 4y agoThe OS will be in the cloud interfacing into our brain by then. I don't want this btw.
- pessimizer 4y agoNot in mine. I don't even want redhat's bullshit in there. I'm not installing some black box into my OS that was programmed with motives that can't be extracted from the model at rest.
- sroussey 4y agoiOS already has this to a degree, for a couple of years.
- make3 4y agoAlpaca uses knowledge distillation (it's trained on outputs from OpenAI models). It's something to keep in mind. You're teaching your model to copy an other model's outputs.
- thewataccount 4y ago> You're teaching your model to copy an other model's outputs. Which itself was trained on human outputs to do the same thing. Very soon it will be full Ouroboros as humans use the model's output to finetune themselves.
- visarga 4y ago> You're teaching your model to copy an other model's outputs. That's a time honoured tradition in ML, invented by the father of the field himself, Geoffrey Hinton, in 2015. > Distilling the Knowledge in a Neural Network https://arxiv.org/abs/1503.02531 https://arxiv.org/abs/1503.02531
- fzliu 4y agoI was a bit skeptical about loading a _4GB_ model at first. Then I double-checked: Firefox is using about 5GB of memory for me. My current open tabs are mail, calendar, a couple Google Docs, two Arxiv papers, two blog posts, two Youtube videos, milvus.io documentation, and chat.openai.com. A lot of applications and developers these days take memory management for granted, so embedding a 4GB model to significantly enhance coding and writing capabilities doesn't seem too far-fetched.
- ChumpGPT 4y agoI'm not so smart and I don't understand a lot about ChatGPT, etc, but could there be a client side app like Folding@home that would allow millions of people to give processing power to train a LLM?
- holloworld 4y ago[dead]
- whalesalad 4y agoAre there any training/ownership models like Folding@Home? People could donate idle GPU resources in exchange for access to the data, and perhaps ownership. Then instead of someone needing to pony up $85k to train a model, a thousand people can train a fraction of the model on their consumer GPU and pool the results, reap the collective rewards.
- ftxbro 4y agoYes there is petals/bloom https://github.com/bigscience-workshop/petals https://github.com/bigscience-workshop/petals but it's not so great. Maybe it will improve or a better one will come.
- whalesalad 4y agoReally interesting live monitor of the network: http://health.petals.ml http://health.petals.ml
- polishdude20 4y agoI wonder how they handle illegal content. Like, if you're running training data on your computer, what's to stop someone else's data that is illegal, from being uploaded to your computer as part of training?
- riedel 4y agoI read that it is only scoring the model collaboratively but it allows some fine-tuning I guess. Getting the actual gradient descent to parallelize is more difficult because one needs to average the gradient when using data/batch parallelism. It becomes more a network speed than GPU speed problem. Or are LLMs somehow different?
- ellisv 4y agoThat’d be cool but I don’t think most idle consumer GPUs (6-8GB) would have large enough memory for a single iteration (batch size 1) of modern LLMs. But I’d love to see more federated/distributed learning platforms.
- fswd 4y agoThere is somebody finetunin 160m rwkv4 on alpaca on the rwkv discord, I am out of the office and can't link but the person posted in prompt showcase channel
- buzzier 4y agoRWKV-v4 Web Demo (169m/430m params) https://josephrocca.github.io/rwkv-v4-web/demo/ https://josephrocca.github.io/rwkv-v4-web/demo/
- lmeyerov 4y agoIt seems the quality goes up & cost goes down significantly with Colossal AI's recent push: https://medium.com/@yangyou_berkeley/colossalchat-an-open-source-solution-for-cloning-chatgpt-with-a-complete-rlhf-pipeline-5edf08fb538b https://medium.com/@yangyou_berkeley/colossalchat-an-open-so... Their writeup makes it sounds like, net, 2X+ over Alpaca, and that's an early run The browser side is interesting too. Browser JS VMs have a memory cap of 1GB, so that may ultimately be the bottleneck here...
- SebJansen 4y agodoes the 1gb limit extend to wasm?
- jesse__ 4y agoWASM is specified to have 32-bit pointers, which is 4GB. AFAIK browser implementations respect that (when I did some nominal testing a couple years ago)
- jesse__ 4y agoI thought the memory limit (in V8 at least) was 2GB due to the GC not wanting to pass 64 bit pointers around, and using the high bit of a 32-bit offset for .. something I now forget ..? Do you have a source showing a JS runtime with a 1GB limit?
- jesse__ 4y agoUPDATE: After a nominal amount of googling around it appears valid sizes have increased on 64-bit systems to a maximum of 8GB, and stayed at 2GB on 32-bit systems, for FF at least. I guess it's probably 'implementation defined' https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Errors/Invalid_array_length https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe... https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/ArrayBuffer#resizing_arraybuffers https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
- lmeyerov 4y ago
- GartzenDeHaes 4y agoIt's interesting to me that LLaMA-nB's still produce reasonable results after 4-bit quantization of the 32-bit weights. Does this indicate some possibility of reducing the compute required for training?
- alecco 4y agoInteresting blog but the extrapolations are way overblown. I tried one of the 30bn models and it's not even remotely close to GPT-3. Don't get me wrong, this is very interesting and I hope more is done in the open models. But let's not over-hype by 10x.
- captaincrowbar 4y agoThe big problem with AI R&D is that nobody can keep up with the big bux companies. It makes this kind of project a bit pointless. Even if you can run a GPT3-equivalent on a web browser, how many people are going to bother (except as a stunt) when GPT4 is available?
- simonw 4y agoAn increasingly common complaint I'm hearing about GPT3/4/etc is people who don't want to pass any of their private data to another company. Running models locally is by far the most promising solution for that concern.
- adeon 4y agoThe ones that can't use the GPT4 for whatever reason. Maybe you are a company and you don't want to send OpenAI your prompts. Or a person who has very private prompts and feel sketchy about sending them over. Or maybe you are an individual who has a use case that's too edgy for OpenAI or a silicon valley corporate image. When Replika shut down people trying to have virtual boyfriend/girlfriends on their platform, their reddit filled up with people who mourned like they just lost a partner. I think it's important that alternative non-big bux company options exist, even if most people don't want to or need to use them.
- psychphysic 4y agoThose are seriously niche use cases. They exist but can they fund gpt5 level development?
- r00fus 4y agoMost corporations/governments would prefer to keep their AI conversations private. Definitely mainstream desire, not niche.
- psychphysic 4y agoWho does your government and corporate email? In the UK it's all either Gmail (for government) and Outlook (NHS). For compliance reasons they simply want data center certification and location restrictions. If you think a small corp is going to get a big gov contract outside of a nepo-state you're in for a shock.
- breck 4y agoJust want to say SimonW has become one of my favorite writers covering the AI revolution. Always fun thought experiments with linked code and very constructive for people thinking about how to make this stuff more accessible to the masses.
- jedberg 4y agoWith the explosion of LLMs and people figuring out ways to train/use them relatively cheaply, unique data sets will become that much more valuable, and will be the key differentiator between LLMs. Interestingly, it seems like companies that run chat programs where they can read the chats are best suited to building "human conversation" LLMs, but someone who manages large text datasets for others are in the perfect place to "win" the LLM battle.
- thih9 4y ago> as opposed to OpenAI’s continuing practice of not revealing the sources of their training data. Looks like that choice makes it more difficult to adopt, trust, or collaborate on the new tech. What are the benefits? Is there more to that than competitive advantage? If not, ClosedAI sounds more accurate.
- astlouis44 4y agoWebGPU is going to be a major component in this. Modern GPU's prevalent in mobile devices, desktops and laptops, is more than enough to do all of this client side.
- pavelstoev 4y agoTraining a ChatGPT-beating model for much less than $85,000is entirely feasible. At CentML, we're actively working on model training and inference optimization without affecting accuracy, which can help reduce costs and make such ambitious projects realistic. By maximizing (>90%) GPU and platform hardware utilization, we aim to bring down the expenses associated with large-scale models, making them more accessible for various applications. Additionally, our solutions also have a positive environmental impact, addressing the excess CO2 concerns. If you're interested in learning more about how we are doing it, please reach out via our website: https://centml.ai https://centml.ai
- nope96 4y agoI remember watching one of the final episodes of Connections 3: With James Burke, and he casually said we'd have personal assistants that we could talk to (in our PDAs). That was 1997 and I knew enough about computers to think he was being overly optimistic about the speed of progress. Not in our lifetimes. Guess I was wrong!
- skybrian 4y agoI wonder why anyone would want to run it in a browser, other than to show it could be done? It's not like the extra latency would matter, since these things are slow. Running it on a server you control makes more sense. You can pick appropriate hardware for running the AI. Then access it from any browser you like, including from your phone, and switch devices whenever you like. It won't use up all the CPU/GPU on a portable device and run down your battery. If you want to run the server at home, maybe use something like Tailscale?
- simonw 4y agoThe browser thing is definitely more for show than anything else - I used it to help demonstrate quite how surprisingly lightweight these models can be.
- v4dok 4y agoCan someone at the EU, the only player in this thing with no strategy yet just pool together enough resources so the open-source people can train models. We don't ask much, just give compute power
- 0xfaded 4y agoNo, that could risk public money benefitting a private party. Feel free to form a multinational consortium and submit a grant application to one of our distribution partners under the Horizon program though. Now, how do you plan to create jobs and reduce CO2?
- PeterisP 4y agoYes, there are a bunch of government-funded supercomputers or clusters which can be obtained for public research needs (based on an evaluation of which projects are likely to bring the most benefit), and are used, among other things, to train large language models. E.g. some interesting Swedish models got trained on https://www.nsc.liu.se/systems/berzelius/ https://www.nsc.liu.se/systems/berzelius/ .
- nwoli 4y agoWhat we need is a RETRO style model where basically after the input you go through a small net that just fetches a desired set of weights from a server (serving data without compute is dirt cheap) and is then executed locally. We’ll get there eventually
- tinco 4y agoCan anyone explain or link some resource on why these big GPT models all don't incorporate any RETRO style? I'm only very superficially following ML developments and I was so hyped by RETRO and then none of the modern world changing models apply it.
- nwoli 4y agoOpenai might very well be using that internally who knows how they implement things. Also emad retweeted a RETRO related thing a bit back so they might very well be using that for their awaited LM, here’s hoping
- deleted 4y ago[deleted]
- Tryk 4y agoWhy doesn't someone just start a gofundme/kickstarter with the goal of funding the training of an open-source ChatGPT-capable model?
- cj 4y agoCreate a clone of OpenAI that pledges to remains open and remains not for profit. That could do really well via crowd funding with the right spin/marketing behind it.
- gessha 4y agoAnd when everyone buys in, you go private everything and reap the benefits. Brilliant!
- cj 4y agoExcept, don’t go private this time. Appoint an external board not controlled by the CEO and write it into your bylaws.
- UncleEntity 4y agoWhere’s the money in that?
- cj 4y agoNowhere obvious. I think there are plenty of rich people who would benefit indirectly from running an OpenAI-like non-profit.
- gessha 4y agoI’m a pessimist about corporate doing the right thing unless it’s economically aligned to do so. “This time” happens with open source where there’s typically no economical incentive and people are doing it for the heck of it.
- gessha 4y agoWe need a DAWNBench* benchmark for training ChatGPT the fastest and cheapest. * https://dawn.cs.stanford.edu/benchmark/ https://dawn.cs.stanford.edu/benchmark/
- cavisne 4y agoThere is a minimum cluster size to get good utilization of the GPU’s. $1 an hour per chip might get you one A100 but it won’t get you hundreds clustered together.
- captainmuon 4y agoI guess companies like OpenAI and Google have no incentives to make models use less resources. The compute required, and of course also their training data, is their moat. If you accept that your model knows less about the world - it doesn't have to know about every restaurant in mexico city or the biography of every soccer player around the world - then you can get away with much fewer parameters and much less training data. Then you can't query it like an oracle about random things anymore, but you shouldn't do that anyway. But it should still be able to do tasks like reformulating texts, judging simularity (by embedding distance), and so on. And TFA mentions it also, you could hook up your simple language model with something like ReAct to get really good results. I don't see it running in the browser, but if you had a license-wise clean model that you can run on premises on one or two GPUs, that would be huge for a lot of people!
- dr_dshiv 4y agoLong Speculative Post on Small Models Hypothesis 1: With better logical thinking (an API call away!), I bet you could train a GPT based on a “small” initial dataset. Why shouldn’t multilingual wikipedia/wiktionary and libgen be enough? That’s what, like less than 10% of the OpenAI training? /s Hypothesis 2: Data sets of philosophical dialogues could help efficiently develop AI reasoning skills. Socratic thinking in Plato and Xenophon represented a powerful new mode of critical thinking. Maybe some Student-Teacher-Student template of dialogue could be powerful in developing useful datasets for AI training. What is the utility of different AI reflective loops for generating training data? (References appreciated if you know any) One possibility to test is a chain of Analyze, Evaluate and Apply loops, applied over and over? “analyze the above piece of text, then evaluate it, then apply to everyday life.” Now, on HN, many have expressed concern that GPT trained on GPT-GPT conversations is going to result in very misaligned models. Like a copy machine degradation, do we want training data from the AI being trained on the AI? But, on the other hand, it is possible that supporting reflective thought is a good idea in AI (we generally value reflective thought) or a bad idea (maybe the reflection will somehow turn it evil, or at least misaligned). Design Question: how might we create useful training data through a process of structuring AI-AI dialogue? “Student-Teacher-Student” conversations seem like they could be good as a useful mode of dialogue. Previously, I’ve finetuned GPT with the complete works of Plato and I was able to generate interesting new dialogues. But the question is whether new dialogues could produce useful data. Perhaps I could use GPT4 to read a part of Plato and then try to autocomplete another part of Plato. Or, as above, use a piece of Platonic dialogue as a target, then use an Analyze, Evaluate, Apply chain on it. We could use methods like these over and over again to make a large dataset about philosophical reasoning. We could have human ratings of the reasonableness of the dialogue output. If a Socratic structure of thinking could read the complete works of Plato over and over again, commenting, countering and synthesizing— with human oversight (RLHF), perhaps we could develop a small module for philosophical reasoning. It might still need millions of conversations, though. But, perhaps by reflecting philosophically by itself, it could produce a sufficiently large dataset that enabled a sophisticated small model with very open resources. And, you’d still need the human preference training RLHF to get it to interact well—and I think it also needs some world model. In any case, I think making smaller and smaller models is a good idea, it sounds fun. TL;DR 1. AI training has philosophically interesting implications 2. Philosophical reasoning is valuable to develop in AI 3. Good philosophical reasoning might be a key benchmark for small models. These models don’t need to know everything but perhaps they could learn what they don’t know. 4. Reading a lot of Plato over and over could be a great way to train GPT that it doesn’t know a lot. 5. What kind of AI-AI dialogues might produce training data that is useful for training small models?
- d4rkp4ttern 4y agoEveryone seems to assume that all the “tricks” behind training ChatGPT are known. The only clues are in papers from ClosedAI like the InstructGPT paper. So we assume there is Supervised Fine Tuning, then Reward Modeling and finally RLHF. But there are most likely other tricks that ClosedAI has not published. These probably took years of R&D to come up with, others trying to replicate ChatGPT would need to come up with these tricks on their own. Also curiously the app was released in late 2022 while the knowledge cutoff is 2021 — I was curious why that might be, and one hypothesis I had was that it may have been because they wanted to keep the training data fixed while they iterated on numerous methods, hyperparameter tuning etc. All of these are unfortunately a defensive moat that ClosedAI has.