39 ms·
NanoChat – The best ChatGPT that $100 can buy
https://x.com/karpathy/status/1977755427569111362 https://x.com/karpathy/status/1977755427569111362
- oblio 1y agoI wonder, if something like this were trained on Wikipedia, could it become a reliable local Wikipedia search engine, basically?
- simonw 1y agoI don't think so. Training on documents is not a great way of building a search engine for those for the information in those documents, because the training process mixes all of that information together in ways that detach the individual words from the source documents they came from. As usual, if you want an LLM to be able to help search a corpus of text the best way to achieve that is to teach it how to use a search tool against that text.
- victor106 1y ago> the best way to achieve that is to teach it how to use a search tool against that text. Any examples of this?
- simonw 1y agoI've seen this called "agentic RAG" by some people. The easiest way to get a local demo is with Claude Code or Codex CLI. They know how to use grep, and you can set them loose on a folder full of text files and tell them to use grep to answer questions - it can work really well. I just tried this in "claude --dangerously-skip-permissions": > Use Python and AppleScript to find Apple Notes that mention UPS ... and fell down a rabbit hole of optimizations because my Notes collection is HUGE, but it got there in the end!
- daft_pink 1y agoWow, how do we sign up for the Eurekalabs course and how much does it cost?
- huseyinkeles 1y agoKarpathy says nanochat will become the capstone project of the course LLM101n being developed by Eureka Labs. I guess it’s still a work in progress? Couldn’t find any other information elsewhere.
- Schiphol 1y agoA bit more info [here](https://github.com/karpathy/LLM101n https://github.com/karpathy/LLM101n)
- karpathy 1y agoStill under development, remaining work includes tuning nanochat (current state being solid v0.1) and finalizing the in-between projects so that students can "unlock" all complexity that hides underneath: `torch.Tensor`, `torch.dist`, `.backward()`, '.compile()`, etc. And then the more ops heavy aspects.
- BrokenCogs 1y agoWhat's the pricing for the course/EurekaLabs? P.s. thanks for all you're doing
- karimf 1y agoI've always thought about the best way to contribute to humanity: number of people you help x how much you help them. I think what Karpathy is doing is one of the highest leverage ways to achieve that. Our current world is build on top of open source projects. This is possible because there are a lot of free resources to learn to code so anyone from anywhere in the world can learn and make a great piece of software. I just hope the same will happen with the AI/LLM wave.
- viccis 1y agoI recommend his ANN/LLM from scratch videos to people a lot because not only is he a clear instructor, but his code tends to be very Pythonic and just the right balance of terse but readable (not counting the Pytorch vectorization stuff, but that's not his fault, it's just complex). So I think people benefit just from watching and imitating his code style.
- croes 1y agoI‘m afraid the technology will do more damage because many people will abuse it for fake news and misinformation.
- IntrepidPig 1y agoYeah it feels similar to inventing the nuke. Or it’s even more insidious because the harmful effects of the tech are not nearly as obvious or immediate as the good effects, so less restraint is applied. But also, similar to the nuke, once the knowledge on how to do it is out there, someone’s going to use it, which obligates everyone else to use it to keep up.
- shafyy 1y agoIf it only were so easy
- deleted 1y ago[deleted]
- bkettle 1y agoThis free tradition in software is I think one of the things that I love so much, but I don't see how it can continue with LLMs due to the extremely high training costs and the powerful hardware required for inference. It just seems like writing software will necessarily require paying rent to the LLM hosts to keep up. I guess it's possible that we'll figure out a way to do local inference in a way that is accessible to everyone in the way that most other modern software tools are, but the high training costs make that seem unlikely to me. I also worry that as we rely on LLMs more and more, we will stop producing the kind of tutorials and other content aimed at beginners that makes it so easy to pick up programming the manual way.
- deleted 1y ago[deleted]
- flakiness 1y agoEureka Labs: https://github.com/EurekaLabsAI https://github.com/EurekaLabsAI What a prolific person Andrej is. It's been more than amazing to follow along!
- jackphilson 1y ago[flagged]
- nsriv 1y agoControlling culture, yes but wild pivot to mention that criminal alongside Karpathy.
- jackphilson 1y agoI mean just an example. He obviously wasn't the most ethical person. Depends how you do it
- IOT_Apprentice 1y agoNeither are Stalin, Netanyahu, Pol Pot, Hitler, Charles Manson et al. Way to derail the conversation. Focus on the positive people and their legacy of time, sharing, positive energy and contributions to society
- jackphilson 1y agonot derailing, just pointing out effective ways of producing good which is what i was responding to. i think its good for people to be aware of this. those people are all examples of people who have influenced culture for bad. you can do it for good: bryan johnson, civil rights leaders, leftist streamers. andrew tate was just the most effective, recent, and obvious one which is why I pointed him out.
- cultofmetatron 1y agonot a particularly ethical guy and I wouldn't hold him up as a example of morality but the guy hasn't actually been found guilty YET. Multiple courts have tried. You'd think that for a guy under as much scrutiny as him that they would have SOMETHING to pin him on by now. Innocent until PROVEN guilty is a foundational legal precedent for a reason.
- deleted 1y ago[deleted]
- andrewmcwatters 1y ago[flagged]
- TheAceOfHearts 1y agoHere's the announcement post [0] from Karpathy, which provides a bit of additional context. [0] https://x.com/karpathy/status/1977755427569111362 https://x.com/karpathy/status/1977755427569111362
- dang 1y agoThanks - we'll put that in the toptext as well
- swyx 1y ago> Thank you to chief LLM whisperer Alec Radford for advice/guidance. oh man an Alec x Andrej podcast would BREAK THE INTERNET... just saying... going from glory days of GPT1 to now building GPT3? in 4 hours
- codybontecou 1y agoPlease oh please. This would be perfect.
- mhitza 1y agoShould be "that you can train for $100" Curios to try it someday on a set of specialized documents. Though as I understand the cost of running this is whatever GPU you can rent with 80GB of VRAM. Which kind of leaves hobbyists and students out. Unless some cloud is donating gpu compute capacity.
- portaouflop 1y agoIf I have let’s say 40gb RAM does it not work at all or just take twice as long to train?
- typpilol 1y agoWon't work at all. Or if it does it'll be so slow since it'll have to go to the disk for every single calculation so it won't ever finish.
- karpathy 1y agoIt will work great with 40GB GPU, probably a bit less than twice slower. These are micro models of a few B param at most and fit easily during both training and inference.
- utopcell 1y agoHow low can this go? Can this run on a 5090 card (32GiB)?
- JonathanFly 1y agoSet nproc_per_node-1 instead of 8 (or run the training script directly instead of using torchrun) and set device_batch_size=4 instead of 32. You may be able to use 8 with a 5090, but it didn't work on my 4090. However it's way slower than expected, one H100 isn't 250x the 4090, so I'm not sure it's training correctly. I'll let it run overnight and see if the outputs make any sense, maybe the metrics are not accurate in this config.
- Havoc 1y ago>If your GPU(s) have less than 80GB, you'll have to tune some of the hyperparameters or you will OOM / run out of VRAM. Look for --device_batch_size in the scripts and reduce it until things fit. E.g. from 32 (default) to 16, 8, 4, 2, or even 1. That sounds like it could run on a 24gb GPU. Batch size of 8 would imply 20gb mem, no? ...presumably just takes forever
- zipy124 1y agoYes, you can always stream data when training or doing inference on models when vram is lacking but the slow down is extremely noticeable. This is the case for CPU code too and is why optimising for bandwidth is so critical in high-performance computing. Your ability to compute is almost always substantially larger than your bandwidth. An Avx512 capable CPU with a suitable amount of cores is easily capable of doing multiple terabytes of fp64 operations per second, but is typically limited by memory bandwidth, GPUs with LLMs have just broadened this knowledge to more people. A fun consequence of the fact that CPUs got faster at a rate quicker than memory is look up tables of pre-computed values used to be common optimisations in code, but now it is almost always quicker to re-compute them than to retrieve a pre-computed value from memory for common use-cases.
- JonathanFly 1y ago> Batch size of 8 would imply 20gb mem, no? I'm running it now and I had to go down to 4 instead of 8, and that 4 is using around 22-23GB of GPU memory. Not sure if something is wrong or if batch is only scaling part of the memory requirements. (Edit: I restarted running the training script directly instead of torch run, and 8 still doesn't fit, but 4 is now using 16-17 instead.) On my 4090 the tok/sec is 523, which is 1/2000 of the 1,000,000 tok/sec of the 8 80GB H100s. That feels too slow so maybe something is wrong. The 4090 is about 1/3 of the raw compute. I'm sure there's other losses from less batching but even if it were 1/10ths as fast, I'd expected something more like 1,000,000 / 10 / 8 so at least 10,000 tok/sec.
- Havoc 1y agoThanks for investigating. Sounds like throwing some dollars at a cloud gpu makes more sense then
- faxmeyourcode 1y agoThis weekend I just cracked into nanoGPT (https://github.com/karpathy/nanoGPT https://github.com/karpathy/nanoGPT), an older but fabulous learning exercise where you build and train a crappy shakespeare GPT with ~0.8M parameters on a cpu. Results are about what you'd expect from that, they suck, but you can start to feel the magic, especially if you're not a deep learning professional and you just want to poke around and hack on it. I started writing up a blog post on my weekend with nanoGPT but it's not done yet... Would have been great to link to here lol oh well
- andrewljohnson 1y agothe shakespeare code tuned a little with different training data does a good job of generating Magic The Gathering commander decks
- dmarcos 1y agoI like the idea of specific-purpose toy models. How did you tune the code and what dataset you used?
- SeanAnderson 1y agowould love more details on this. this is exactly the type of project I'd like to dabble in to get more up to speed.
- vunderba 1y agoFWIW, there was a pretty popular post on HN around generating MTG cards using AI a couple years back but I believe that their approach was a fine-tune on an existing LLM. https://news.ycombinator.com/item?id=37427854 https://news.ycombinator.com/item?id=37427854
- astrange 1y agoPeople have been doing this for a while. https://x.com/roborosewater https://x.com/roborosewater https://bsky.app/profile/roborosewaterm.bsky.social https://bsky.app/profile/roborosewaterm.bsky.social You can see the invention of RLHF/ChatGPT here because text generation suddenly became much more coherent and also much less interesting. You have to go back to older tech for surrealism because nobody will let you see the good stuff (the base models).
- CountGeek 1y agoSo could I in practice train it on all my psychology books, materials, reports, case study and research papers and then run it on demand on a 1xH100 node - https://getdeploying.com/reference/cloud-gpu/nvidia-h100 https://getdeploying.com/reference/cloud-gpu/nvidia-h100 whenever I have a specialised question?
- zipy124 1y agoYou could but it would be significantly worse than fine-tuning or RAG with a pre-trained model, or using a smaller model since your dataset would be so small.
- leokeba 1y agoYou could do that indeed, but the performance would be abysmal. For this kind of use-case, it would be a LOT better to use a small pre-trained model and either fine-tune it on your materials, or use some kind of RAG workflow (possibly both).
- dmix 1y ago> it would be a LOT better to use a small pre-trained model and either fine-tune it on your materials, or use some kind of RAG workflow (possibly both). I noticed NewRelic has a chat feature that does this sort of thing, it's scoped very narrowly down to their website and analytics DSL language, and generates charts/data from their db. I've always wondered how they did that (specifically in terms of set up the training/RAG + guardrails). It's super useful.
- simonw 1y agoYou might be able to figure that out just by asking it - see if you can get it to spit out a copy of the system prompt or tell you what tools it has access to. The most likely way of building that would be to equip it with a "search_docs" tool that lets it look up relevant information for your query. No need to train an extra model at all if you do that.
- gojomo 1y ago
- cyanydeez 1y agoif the AI bubble is anything to be compared to, how is 100$ worth anything in GPT terms.
- computer23 1y agoHas the word ChatGPT become generic? This has nothing to do with OpenAI's ChatGPT.
- huflungdung 1y ago[dead]
- simonw 1y agoIt's a reasonable shortcut for what this project provides: training code, inference code and a ChatGPT-style web interface for chatting with the model.
- sbassi 1y agoWhich data uses for training?
- eranation 1y agoI think he mentioned somewhere he used fineweb (I assume this one https://huggingface.co/datasets/HuggingFaceFW/fineweb https://huggingface.co/datasets/HuggingFaceFW/fineweb)
- simonw 1y agokarpathy/fineweb-edu-100b-shuffle: https://huggingface.co/datasets/karpathy/fineweb-edu-100b-shuffle https://huggingface.co/datasets/karpathy/fineweb-edu-100b-sh... Which is derived from HuggingFaceFW/fineweb-edu: https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu HuggingFaceTB/smol-smoltalk: https://huggingface.co/datasets/HuggingFaceTB/smol-smoltalk https://huggingface.co/datasets/HuggingFaceTB/smol-smoltalk And extra fine-tuning on portions of: cais/mmlu: https://huggingface.co/datasets/cais/mmlu https://huggingface.co/datasets/cais/mmlu openai/gsm8k: https://huggingface.co/datasets/openai/gsm8k https://huggingface.co/datasets/openai/gsm8k allenai/ai2_arc: https://huggingface.co/datasets/allenai/ai2_arc https://huggingface.co/datasets/allenai/ai2_arc
- dinkblam 1y agofrom their promotional material: >> Why is the sky blue? > The sky is blue due to an optical illusion called the Rayleigh Scattering Rayleigh Scattering is not an illusion but an effect. > […] particles are made up of tiny blue and violet particles that cause the light to bend in a particular way. ugh. no, there are no "tiny blue" particles in the sky.
- simonw 1y agoThat was the point. That example is meant to demonstrate that the model that trained for 4 hours can imitate a conversation but isn't actually anywhere close to being useful.
- kragen 1y agoWhere did you find that?
- Ono-Sendai 1y agonot sure why you are being downvoted. That 'explanation' of Rayleigh scattering is just wrong.
- simonw 1y agoBecause that explanation is expected to be wrong. This is a partially trained tiny model, Andrej shared that obviously incorrect explanation to emphasize that the model is not trustworthy or useful at that stage.
- sammyd56 1y agoI'm doing a training run right now (started 20min ago). You can follow it at https://api.wandb.ai/links/sjd333-none/dsv4zkij https://api.wandb.ai/links/sjd333-none/dsv4zkij Will share the resulting model once ready (4 hours from now) for anyone to test inference.
- royosherove 1y agoCool. Is there a simple "howto" on running this repo with training on W&B for a programmer like me who has never done model training flows? Maybe you could share the steps you took?
- sammyd56 1y agoThere's not much to it... it took longer to spin up the cloud machine than it did to kick off the training run. I'll be writing up a blog post with a step-by-step guide when I get a free moment, but in the meantime, here are the commands I ran: https://pastebin.com/sdKVy0NR https://pastebin.com/sdKVy0NR
- royosherove 1y agoAh I was missing the WANDB_RUN env var. so did not get any logs. thanks!
- Lerc 1y agoThe comment beside the first chart >Our main measure of progress. Bits per byte is, per Karpathy, "a much better measure than just the typical cross-entropy loss, because it further normalizes the loss on each token by the number of bytes of that token, making the metric tokenizer-invariant". Is so blindingly obvious, that I'm ashamed to think that I didn't think do it when trialing my own tokenizer approach on tinystories. I might go back and have a look at how well my tokenizer compared to how well I imagined it compared.
- typpilol 1y agoWhy hasn't anyone made a tokenizer that's 1 character per token. Is it because it requires an insane amount of compute? Or would the loss of efficiency make it dumber then modern tokenizers?
- efficax 1y agoTry ~300k for an 8xH100 lol
- samus 1y agoAndrej Karpathy slays again by spreading knowledge about this important subject to the people!
- lebimas 1y agoI see Karpathy, I click
- sieve 1y agoNice! His Shakespeare generator was one of the first projects I tried after ollama. The goal was to understand what LLMs were about. I have been on an LLM binge this last week or so trying to build a from-scratch training and inference system with two back ends: - CPU (backed by JAX) - GPU (backed by wgpu-py). This is critical for me as I am unwilling to deal with the nonsense that is rocm/pytorch. Vulkan works for me. That is what I use with llama-cpp. I got both back ends working last week, but the GPU back end was buggy. So the week has been about fixing bugs, refactoring the WGSL code, making things more efficient. I am using LLMs extensively in this process and they have been a revelation. Use a nice refactoring prompt and they are able to fix things one by one resulting in something fully functional and type-checked by astral ty.
- danielmarkbruce 1y agoUnwilling to deal with pytorch? You couldn't possibly hobble yourself anymore if you tried.
- sieve 1y agoIf you want to train/sample large models, then use what the rest of the industry uses. My use case is different. I want something that I can run quickly on one GPU without worrying about whether it is supported or not. I am interested in convenience, not in squeezing out the last bit of performance from a card.
- danielmarkbruce 1y agoYou wildly misunderstand pytorch.
- sieve 1y agoWhat is there to misunderstand? It doesn't even install properly most of the time on my machine. You have to use a specific python version. I gave up on all tools that depend on it for inference. llama-cpp compiles cleanly on my system for Vulkan. I want the same simplicity to test model training.
- earthnail 1y agoThis is absolutely fantastic. I really can't wait for the final course to be live. It's in the "shut up and take my money" category. I had so much fun with the nanoGPT videos.
- deleted 1y ago[deleted]
- RobGR 1y agoThis is an LLM trained using a $100 budget to RENT access to graphics cards. It's not about what you could do BUYING hardware for $100.
- danielmarkbruce 1y agoNowhere does he suggest he is buying hardware.
- HelloMcFly 1y agoOnce the LLM is trained you don't need the rented hardware anymore.
- montebicyclelo 1y ago> nanochat is also inspired by modded-nanoGPT Nice synergy here, the lineage is: Karpathy's nano-GPT -> Keller Jordan's modded-nanoGPT (a speedrun of training nanoGPT) -> NanoChat modded-nanoGPT [1] is a great project, well worth checking out, it's all about massively speeding up the training of a small GPT model. Notably it uses the author's Muon optimizer [2], rather than AdamW, (for the linear layers). [1] https://github.com/KellerJordan/modded-nanogpt https://github.com/KellerJordan/modded-nanogpt [2] https://kellerjordan.github.io/posts/muon/ https://kellerjordan.github.io/posts/muon/
- deleted 1y ago[deleted]
- varunneal 1y agoMuon was invented by Keller Jordan (and then optimized by others) for the sake of this speedrunning competition. Even though it was invented less than a year ago, it has already been widely adopted as SOTA for model training
- tbalsam 1y agoThis is the common belief but not quite correct! The Muon update was proposed by Bernstein as the result of a theoretical paper suggesting concrete realizations of the theory, and Keller implemented it and added practical things to get it to work well (input/output AdamW, aggressive coefficients, post-Nesterov, etc). Both share equal credit I feel (also, the paper's co-authors!), both put in a lot of hard work for it, though I tend to bring up Bernstein since he tends to be pretty quiet about it himself. (Source: am experienced speedrunner who's been in these circles for a decent amount of time)
- varunneal 1y agoI think it's good to bring up Bernstein & Newhouse as well as Yuchen Jin, Jiacheng You and the other speedrunners who helped iterate on Muon. But I think it's very fair to call Keller Jordan the main author of Muon of its current form. I'm also in the speedrunning community though maybe not as long as you have
- wyldfire 1y agoI would love to take an existing open-weight model and fine-tune it with specific training data along these lines. Can I do that with Qwen or GLM? Is there a ~simple recipe for doing that?
- tdhz77 1y agoThese are the time of community posts that are legendary.
- kragen 1y agoThis is really inspiring! Does anyone have some example of how well or poorly it performs on some example prompts?
- kragen 1y agoSimon.incutio.com points out that there are screenshots on https://xcancel.com/karpathy/status/1977755430093980034 https://xcancel.com/karpathy/status/1977755430093980034.
- tehnub 1y agoInteresting exchange on the use of AI coding tools: curious how much did you write the code by hand of it? Karpathy: Good question, it's basically entirely hand-written (with tab autocomplete). I tried to use claude/codex agents a few times but they just didn't work well enough at all and net unhelpful, possibly the repo is too far off the data distribution. https://x.com/karpathy/status/1977758204139331904 https://x.com/karpathy/status/1977758204139331904
- oblio 1y agoWe're still not ready for ouroboros.
- gyomu 1y ago> the repo is too far off the data distribution ah, this explains why these models have been useless to me this whole time. everything i do is just too far off the data distribution!
- SchemaLoad 1y agoEverything is unless your app is a React todolist or leatcode questions.
- SeanAnderson 1y agoor a typical CRUD app architecture, or a common design pattern, or unit/integration test scaffolding, or standard CI/CD pipeline definitions, or one-off utility scripts, etc... Like 80% of writing coding is just being a glorified autocomplete and AI is exceptional at automating those aspects. Yes, there is a lot more to being a developer than writing code, but, in those instances, AI really does make a difference in the amount of time one is able to spend focusing on domain-specific deliverables.
- positron26 1y agoIt has gotten to the point that I don't modify or write SQL. Instead I throw some schema and related queries in and use natural language to rubber duck the change, by which point the LLM can already get it right.
- dabockster 1y agoThe title is extremely misleading - you have to rent time on an H100 cluster to get it to work. It is not on-device, and thus not truly $100. I was really excited, too, until I looked through the readme files and the code.
- simonw 1y agoIt's about training a model from scratch for $100.
- arkmm 1y agoWhat's misleading about that? You rent $100 of time on an H100 to train the model.
- mynameisjoseph 1y agoI feel same. The title looks like I could have on-deivce ChatGPT with $100 forever. I couldn't imagine it's about training the model by myself.
- simonw 1y agoSince the resulting model is only ~561M parameters you could run it on a Raspberry Pi that costs less than $100.
- rpdillon 1y agoThe title is saying you can train your own model for $100. That part is true: the $100 goes to the cloud provider to rent you $250k of hardware for four hours. Then you can run that model on whatever hardware you have lying around, because it's really small.
- zoba 1y agoI’m very excited for this. An early question I have: what would need to be done to make this a “thinking” model?
- JKCalhoun 1y ago"The fastest way to feel the magic is to run the speedrun script speedrun.sh, which trains and inferences the $100 tier of nanochat. On an 8XH100 node at $24/hr, this gives a total run time of about 4 hours." I am clueless and don't understand this. Where is the $100 being spent? Some sort of API you have to pay to access? Some sort of virtual hardware you have to rent access to?
- llleeeooo 1y agoRenting 8 H100s would cost you about 24/h
- simonw 1y agoH100s are expensive NVIDIA GPUs, each costing about $30,000. 8XH100 means you have 8 of those wired together in a big server in a data center somewhere, so around a quarter of a million dollars worth of hardware in a single box. You need that much hardware because each H100 provides 80GB of GPU-accessible RAM, but to train this model you need to hold a LOT of model weights and training data in memory at once. 80*8 = 640GB. ~$24/hour is how much it costs to rent that machine from various providers.
- deleted 1y ago[deleted]
- KnowledgeWeaver 1y agoAh, but this is nice project. I'll start hacking once it's easier to fine-tune it with own documents for specific questions. What plaques me, though, is how you prevent the model from answering questions it was not trained for?
- cat_plus_plus 1y agoEnd to end training is a different beast, but finetuning and inference of impressive LLMs like QWEN3 can be done on pretty run of the mill hardware like Apple Silicon macs and gaming PCs if anyone wants a personalized assistant with character. Just ask AI how to finetune AI using unsloth (if using NVIDIA) or MLX (for apple) and it will give you ready to run python scripts.
- chipsrafferty 1y agoWould love to hear some metrics on training it on your personal computer rather than a "cloud GPU box". I don't care if it takes 3 months to train if I have something good, offline, and free(ish, but just pay electric bills)
- zoba 1y agoI’d also be interested in this. Especially for Macs
- ComputerGuru 1y agoEach H100 can do 60 TFLOPS of f32 operations, while a single RTX 3080 can do roughly half that (just under 30). So complete back-of-the-envelope answer would be 16x as long (since nanochat is targeting four hours with 8xH100) 64 hours isn’t too bad at all! (An RTX 2080 can only do 10 TFLOPS for fp32, so that would be again 3x as long.)
- lostmsu 1y agoThis is going to be the single most powerful boost to my indie research efforts in years. Thank you, Andrej!
- yieldcrv 1y ago> nanochat is designed to run on a single 8XH100 node
- jmspring 1y ago8XH100 nodes start at ~$450ish/day. Not sure about the $100 part. I need to dig into the post.
- markr1 1y ago$100 to teach us all how to build an LLM, this is what open education should look like.
- saivishwak 1y agoVery cool project! Hopefully it will propel SLM development
- mips_avatar 1y agoThanks Andrej for putting this up. Your videos gave me the confidence to work full time on LLMs last year after I left Microsoft
- jumski 1y ago100$ to train a sort of talkable model in 4 hours? wow
- deleted 1y ago[deleted]
- spacecadet 1y agoBuilt so many nano AIs over the last several years. I have played with nanoGPT, its ok. Just hype for Kpathy... So many tiny LLMs out there now that run on cheap SOCs. Try SmolVLM512, runs fine on a sub $100 pi.
- simonw 1y agoYou're misunderstanding the project. This isn't about an LLM that runs on $100 hardware. It's about a usable LLM that costs $100 to train from scratch.
- spacecadet 1y agoNo I get that. Having trained my own small LLMs and for much less than $100.
- desaiguddu 1y agoI am building a product similar to DataGPT https://datagpt.com/ https://datagpt.com/ and Julius.ai - will this help in that?
- simonw 1y agoNot at all. This project is for learning how LLMs work and how to build them from first principles. If you want to solve problems that aren't "how do I build an LLM from scratch" this isn't the right path for you.
- alex000kim 1y agoI created this PR to make it easier for folks to train and serve it on any cloud (or their own K8s): https://github.com/karpathy/nanochat/pull/18 https://github.com/karpathy/nanochat/pull/18
- homicipher 1y ago[dead]
- megadragon9 1y agoLove the educational value of this "nano-sized" project. This reminded me of the from-scratch project I created to learn about deep learning libraries, neural networks all the way to LLMs like GPT-2 using just Numpy and Python [1]. Learning is done by "re-inventing the wheel" yourself, one step at a time :) [1] https://github.com/workofart/ml-by-hand https://github.com/workofart/ml-by-hand