22 ms·
Karpathy's MinGPT
- deleted 6y ago[deleted]
- minimaxir 6y ago> huggingface/transformers has a language-modeling example. It is full-featured but as a result also somewhat challenging to trace. E.g. some large functions have as much as 90% unused code behind various branching statments that is unsued in the default setting of simple language modeling. I don't understand this criticism of Transformers. Doesn't tracing (in both TorchScript and ONNX forms, which Transformers supports for exporting) just take the relevant model graph and freeze it? I don't think either contains the somewhat-weighty performance code.
- activatedgeek 6y agoI think the argument here is about pedagogy not performance.
- karpathy 6y agominGPT is actually quite performant too, the min refers to breadth of supported functionality (eg the absence of support for various additional conditioning, exotic masking, masked LMs, finetuning, pruning, etc).
- activatedgeek 6y agoSure thing! I only meant to imply the relative ordering of considerations.
- minimaxir 6y agoGPT training performance on the CPU is funny. The vocab size and context window size have a massive effect on both speed and accuracy.
- deleted 6y ago[deleted]
- cs702 6y agoFWIW, I took that to mean "code path is challenging [for a human being] to trace."
- master_yoda_1 6y agoI agree with this. Changes for Deep learning models is so fast there is no point maintaining a super reusable code.
- fabmilo 6y agoTry to read the code and understand how it works and you will find it very challenging to interpret. But not just that even the documentation is very sparse and hard to read. Compare that to the open-ai code, is so coincise and easy to read. There is mastery in doing that, deep mastery. Few repositories on tensorflow or pytorch organization get to that level.
- MiroF 6y agoHonestly the library doesn't seem that hard to understand, although it can be under documented at times - I found looking through the source very helpful.
- sillysaurusx 6y agoAgreed re: OpenAI's GPT implementation. It took roughly a year to appreciate how simple it is. https://github.com/openai/gpt-2/blob/0574c5708b094bfa0b0f6dfe3fd284d9a045acd9/src/model.py#L147-L173 https://github.com/openai/gpt-2/blob/0574c5708b094bfa0b0f6df... Especially compared to StyleGAN, BERT, or pretty much anything else. I used to hate the OpenAI GPT codebase: zero comments? no classes? What does "mlp" even mean? But over time, I find myself reverting to their style.
- xiphias2 6y agoIt's fun to find the sources to see how GPT output correlates with the input data: Input texts: Go, rate thy minions, proud insulting boy! Hither to London, to be crown'd our king. / Welcome, sweet prince, to London, to your chamber. / Post you to London, and you will find it so / Now to London, To see these honours in possession. How will the country for these woful chances Misthink the king and not be satisfied! Output: Go, rating to London, with all these woful chances Misthink the king and not be satisfied!
- master_yoda_1 6y agoDid you understand the correlation? And what part of the source code is doing it? Or you had just fun without any understanding?
- xiphias2 6y agoGPT is using multi-headed attention, so of course it's not as simple as putting a few texts together, but I was still interested in finding some similar texts (that can be done because the training data is only 1MB).
- master_yoda_1 6y agowhy don't use TF-IDF and use just elastic search?
- sillysaurusx 6y agoThat's a really interesting idea. Could you go into detail about how you're searching for similar texts using GPT? It's true that the probability distribution is a sort of "edit distance". And GPT has already been used for text compression: https://bellard.org/nncp/gpt2tc.html https://bellard.org/nncp/gpt2tc.html so it seems not too far of a stretch to use it for similarity matching. (Sure, perhaps there are more efficient or more effective techniques than using GPT for this, but I like the idea and am curious how it works.)
- 6y ago
- master_yoda_1 6y agoNice but i am scared people write in their resume that they train a GPT model from scratch. And when asked in detail they will accept just ran minGPT without understanding it. This is the AI story now a days. Best solution is to ignore minGPT and write your own version.
- codezero 6y agoThat seems fine. How many reddit clones have we seen? For what it's worth, as a hiring manager who hires technical people, but not software engineers, these kinds of side projects can really help folks with a less formal education. The upside is that if they DO understand it or at the very least learned something interesting or useful while working on the project it will help them a lot in an interview. Folks who do what you're worried about exist, and they just don't get hired by people who aren't impressed with the shallow use of some new technology. There are also plenty of companies where that person will be fine, not really need to use GPT to do their job, and everyone will be happy anyways :)
- master_yoda_1 6y agoHow you feel if somebody claim they know java just because they can call elastic search api (or put any tool written in java). Are you going to hire them and put them on a project which involve coding in java?? the issue is false claim on resume which shows no integrity. On other extreme do you go to a doctor for surgery who make false claim about being done surgery.
- codezero 6y agoSome companies are looking for someone who knows enough Java to use an ElasticSearch API - it's about knowing your audience, and I think a lot of new folks just don't know how to target their applications. For what it's worth I've gone back and forth over this in my career, and I do think it has a bit to do with the level of your assertion, but there's an amount of naïveté that creeps into resumes especially since advice is so widely varied, do you brag, sell yourself, show real projects, or your past titles. Anyways, there's no right way, and I've decided that the job seeker is in a position of weakness to employers and the industry in general, and if people are seeking to better themselves - great. When I interview and hire people, I'm the one who screens out people who lie or are incapable of doing their job, what they put on their resume is just part of the process. Your medical example is a bad one, sorry, I'm not going to engage with it, there are gatekeepers in almost all industries, and I'm not proposing changing them by saying that someone putting Java on their resume is not the same as a doctor lying about their qualifications. One, because systems exist to vet those qualifications, and two because they're simply not on the same spectrum. How many people have you screened, hired, and for how many roles? Is this a problem you've run into or experienced professionally, or just something that annoys you? I've experienced both. When I ran a satellite, I employed interns and most were CS majors who claimed to know C, C++, and Java - I asked them to write strcpy in C as my interview, none of them could do it, and these people were juniors in college, so no, I'm not too worried about what people put on their resume because it's just not representative of their actual abilities ever. If someone claims to know Java and gets a job with no actual test of their skill, that's the manager's fault.
- mindfulplay 6y agoIf this is the same person that also is in charge of Tesla's "auto" pilot, then it says a lot. His work has personally led to the deaths of many Tesla owners completely out of their control. Due to pure careless, Silicon Valley high-horse bullshit. Can we please stop hyping this dude. This AI/ML crap that comes out of silicon valley for critical faculty has to stop.
- dang 6y agoPersonal attacks are not ok on Hacker News. Also, please don't fulminate in comments generally. We want curious conversation, not flamewar. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- refulgentis 6y agothe Tesla Autopilot rewrite due in 6-10 weeks that'll enable self-driving cars, promised by Elon on Friday, must be going well if the head of Autopilot is playing around with language ML models and open sourcing them https://twitter.com/elonmusk/status/1294374864657162240?s=20 https://twitter.com/elonmusk/status/1294374864657162240?s=20
- adamnemecek 6y agoHe's allowed to do other things. Also, maybe his part of the project is done. Also the thing might be delayed.
- refulgentis 6y agoI agree, I do side projects on the weekend to decompress, but...not a great look in 4 years and running of consumer fraud edit: rolling my eyes at downvotes, we have a CEO and CTO selling "full self driving" for several thousand dollars, promising cross-continental trips by end of year, promising you're buying a money printing robotaxi, telling you to buy today before they raise prices because fair market value is $150K. You have direct evidence of this as recently as 72 hours ago. If you're still splitting hairs to infer good faith, who is benefiting from that?
- lacker 6y agoI downvoted because complaining about Tesla is irrelevant to the topic, not to express pro-Tesla sentiment.
- refulgentis 6y agoKarpathy is the head of Tesla autopilot, and the headline is Karpathy's MinGPT.
- danudey 6y agoThat has nothing to do whatsoever with Tesla's "4 years and running of consumer fraud". Whatever your thoughts on Tesla's public image or business behaviors, it's not especially relevant to the discussion on machine learning.
- modeless 6y agoI think many people don't appreciate just how simple state-of-the-art neural net techniques are. It's great to have an illustration of just how little code you need to get the results that have amazed people. You could say that it relies on PyTorch which is a lot of code, but most of the complexity comes from the need to do GPU/TPU/SIMD acceleration. A truly from-scratch CPU-only implementation in C would still not be a large amount of code.
- typon 6y agoIt's kind of like saying Redis is just an dict. The optimizations in the entire software/hardware stack is what makes these innovations possible. A C implementation would not work in reality for large models.
- modeless 6y agoHere is a C implementation of GPT-2. https://bellard.org/nncp/gpt2tc.html https://bellard.org/nncp/gpt2tc.html I don't disagree that hardware acceleration is key in enabling these models, but I still find it interesting how simple the core techniques are.
- read_if_gay_ 6y agoIs there anything Bellard hasn’t done?
- ypcx 6y agoA few more transformer implementations that I’ve found: https://github.com/pbloem/former/blob/master/former/transformers.py https://github.com/pbloem/former/blob/master/former/transfor... https://github.com/openai/blocksparse/blob/master/examples/transformer/enwik8.py https://github.com/openai/blocksparse/blob/master/examples/t... https://github.com/google/trax/blob/master/trax/models/transformer.py https://github.com/google/trax/blob/master/trax/models/trans...
- fpgaminer 6y agoThis is really cool to see; I'm glad karpathy shared the work. The internet at large has been really good at taking opaque machine learning research and elucidating the details and recreating the results. I've seen a few posts/repositories/etc doing that for GPT-3 as well but man, 175B parameters is just so far out of reach for hobbyists. It's really a shame. In time further research will likely make language models more efficient to where something GPT-3-like can be trained at the hobbyist level. Probably a blended model like SHA-RNN or something with dedicated memory in its architecture, so the model isn't burning precious weights on remembering e.g. Lincoln's birthday. In the meantime though it makes me sad that something as impressive as GPT-3 is solely the toy of corporations.
- GaryNumanVevo 6y agoWe recently trained GPT-3 (SMALL) at work on our GPU cluster for fun, took 4 days across a couple dozen machines... Millions of dollars in CAPEX and OPEX just for one model
- stainforth 6y agoYou're saying your project costed millions of dollars, or the big boys' projects did?
- shmageggy 6y agoIf "4 days across a couple dozen machines" cost millions, something is very wrong.
- deleted 6y ago[deleted]
- T-A 6y agoNot if it was a couple dozen of these machines: https://www.hardwarezone.com.sg/tech-news-nvidia-dgx-a100-supercomputer-super-performance-fight-covid-19 https://www.hardwarezone.com.sg/tech-news-nvidia-dgx-a100-su...
- deleted 6y ago[deleted]
- iagovar 6y agoI've been delaying introducing myself to this field seriously, but this feels attractive to me. I have a 32GB Dual Xeon through RDP, would you recommend running this in local?
- abakus 6y agoIf you just want to understand the Transformer, here is a clean implementation: https://github.com/blue-season/pywarm/blob/master/examples/transformer.py https://github.com/blue-season/pywarm/blob/master/examples/t...
- chronolitus 6y agoand here's a breakdown of the architecture: http://dugas.ch/artificial_curiosity/GPT_architecture.html http://dugas.ch/artificial_curiosity/GPT_architecture.html
- odnes 6y agoThese 4 videos (~45 mins) do an excellent job at explaining attention, multi-headed attention, and transformers: https://www.youtube.com/watch?v=yGTUuEx3GkA https://www.youtube.com/watch?v=yGTUuEx3GkA
- mark_l_watson 6y agoThis is great. Years ago his RNN example helped me a lot.
- bigtex 6y agoDoes it stop Tesla's from ramming into red fire trucks?
- haolez 6y agoIs it possible to build an useful transformer without investing millions of dollars?
- grandmczeb 6y agoPeople have built passable translation systems with transformers using a single high end GPU.
- martythemaniak 6y agoI wonder, is a community trained model feasible? As in, get a few tens of thousands of dev to run a seti@home type app on their GPUs during the night, and at the end you get access to the 175B trained model. If it cost 5m to train, but IIRC that was estimated at cloud gpu prices, if you're using spare capacity you're just paying for electricity.
- thewarrior 6y agoSeems doable to me
- karpathy 6y agoFun idea. GPT @ Home :D. Scatter of the inputs would be very cheap as they are tiny LongTensors (sequences of indices), but the Gather of the gradients seems like a bottleneck. These models can be quite large. Maybe each worker only communicates back some sparse or potentially precision-reduced gradients? In the limit, I recall papers that were able to only communicate one bit per dimension. May also be possible to further reduce the number of weights by weight sharing, or e.g. with HyperNetworks.
- londons_explore 6y agoI built this a few years ago: https://github.com/Hello1024/shared-tensor https://github.com/Hello1024/shared-tensor It does updates to weights based on 1 bit precision updates each iteration. It would be fairly trivial to go to less than 1 bit precision too - simply set some threshold (eg 3), and wherever the difference between the weight on the server and the client is greater than 3, transmit a binary "1", else send a binary "0". Then entropy code all the resulting binary. By adjusting the threshold up and down, you trade off the size of the data to send Vs precision.
- 0-_-0 6y agoI read a paper that did exactly what you describe but of course I can't find it now...
- 6y ago
- fpgaminer 6y agoJust started a run of play_char on my 8GB 2070. Had to drop batch size to 128. Getting ~2.2 iterations per second, so it looks like it's going to take two hours for training to finish. I don't expect my training to differ from karpathy's, but I'm curious to play around with the trained model a bit. Already ran the math one and got the ~same results.
- fpgaminer 6y agoIn the committed notebook the final training loss for play_char was 0.30588. Yet my training got down to 0.02638. Odd. Either way the resulting model seems to be just as good/bad. Like char-rnn it's amazing to see it spell words from scratch consistently. It has a good grasp on structure and even a passable grasp on grammar. But also like char-rnn it lacks any ability to form coherent sentences. EDIT: I'm running it on the IMDB dataset now ... just to see.
- fpgaminer 6y agoRunning for two epochs on the IMDB dataset (133MB corpus) it only got to a loss of 1.1. Likely the regularization is too high (I didn't tweak the hyperparameters at all, and assume regularization was quite high for the limited tinyshakespeare corpus). Either way, it at least started to learn more grammar: Prompt: This is my review of Lord of the Rings. > I can't tell why the movie is a story with a lot of potential the main reason I want to see a movie that is Compared to the Baseball movie 10 Both and I can say it was not just a bad movie.
- studentdev 6y agoFrom the GitHub Readme "The rest of the complexity is just being clever with batching (both across examples and over sequence length) so that training is efficient." What kind of complexities is he talking about? Is it simply the complexity of having a batch dimension? (compared to more simple single input code)
- kamalkraj 6y agoA TensorFlow re-implementation of mingpt https://github.com/kamalkraj/minGPT-TF https://github.com/kamalkraj/minGPT-TF