8 ms·
Karpathy's contribution to teaching around deep learning is just immense. He's got a mountain of fantastic material from short articles like this, longer writin
by gchadwick 11mo ago
Karpathy's contribution to teaching around deep learning is just immense. He's got a mountain of fantastic material from short articles like this, longer writing like https://karpathy.github.io/2015/05/21/rnn-effectiveness/ https://karpathy.github.io/2015/05/21/rnn-effectiveness/ (on recurrent neural networks) and all of the stuff on YouTube.
Plus his GitHub. The recently released nanochat https://github.com/karpathy/nanochat https://github.com/karpathy/nanochat is fantastic. Having minimal, understandable and complete examples like that is invaluable for anyone who really wants to understand this stuff.
- throwaway290 11mo agoAnd to all the LLM heads here, this is his work process: > Yesterday I was browsing for a Deep Q Learning implementation in TensorFlow (to see how others deal with computing the numpy equivalent of Q[:, a], where a is an integer vector — turns out this trivial operation is not supported in TF). Anyway, I searched “dqn tensorflow”, clicked the first link, and found the core code. Here is an excerpt: Notice how it's "browse" and "search" not just "I asked chatgpt". Notice how it made him notice a bug
- stingraycharles 11mo agoFirst of all, this is not a competition between “are LLMs better than search”. Secondly, the article is from 2016, ChatGPT didn’t exist back then
- code51 11mo agoI doubt he's letting LLM creep in to his decision-making in 2025, aside from fun side projects (vibes). We don't ever come across Karpathy going to an LLM or expressing that an LLM helped in any of his Youtube videos about building LLMs. He's just test driving LLMs, nothing more. Nobody's asking this core question in podcasts. "How much and how exactly are you using LLMs in your daily flow?" I'm guessing it's like actors not wanting to watch their own movies.
- danielbln 11mo agohttps://news.ycombinator.com/item?id=45788753 https://news.ycombinator.com/item?id=45788753
- mquander 11mo agoKarpathy talking for 2 hours about how he uses LLMs: https://www.youtube.com/watch?v=EWvNQjAaOHw https://www.youtube.com/watch?v=EWvNQjAaOHw
- code51 11mo agoVibing, not firing at his ML problems. He's doing a capability check in this video (for the general audience, which is good of course), not attacking a hard problem in ML domain. Despite this tweet: https://x.com/karpathy/status/1964020416139448359 https://x.com/karpathy/status/1964020416139448359 , I've never seen him citing an LLM helped him out in ML work.
- soulofmischief 11mo agoYou're free to believe whatever fantasy you wish, but as someone who frequently consults an LLM alongside other resources when thinking about complex and abstract problems, there is no way in hell that Karpathy intentionally limits his options by excluding LLMs when seeking knowledge or understanding. If he did not believe in the capability of these models, he would be doing something else with his time.
- strogonoff 11mo agoOne can believe in the capability of a technology but on principle refuse to use implementations of it built on ethically flawed approaches (e.g., violating GPL licensing laws and/or copyright, thus harming open source ecosystem).
- soulofmischief 11mo agoWhat you see as copyright violation, I see as liberation. I have open models running locally on my machine that would have felled kingdoms in the past.
- confirmmesenpai 11mo agowhat you did here is called confirmation bias. > I think congrats again to OpenAI for cooking with GPT-5 Pro. This is the third time I've struggled on something complex/gnarly for an hour on and off with CC, then 5 Pro goes off for 10 minutes and comes back with code that works out of the box. I had CC read the 5 Pro version and it wrote up 2 paragraphs admiring it (very wholesome). If you're not giving it your hardest problems you're probably missing out. https://x.com/karpathy/status/1964020416139448359 https://x.com/karpathy/status/1964020416139448359
- away74etcie 11mo agoYes, embedding .py code inside of a speedrun.sh to "simplify the [sic] bash scripts." Eureka runs LLM101n, which is teaching software for pedagogic symbiosis. [1]:https://eurekalabs.ai/ https://eurekalabs.ai/
- kubb 11mo agoI was slightly surprised that my colleagues, who are extremely invested in capabilities of LLMs, didn’t show any interest in Karpathy’s communication on the subject when I recommended it to them. Later I understood that they don’t need to understand LLMs, and they don’t care how they work. Rather they need to believe and buy into them. They’re more interested in science fiction discussions — how would we organize a society where all work is done by intelligent machines — than what kinds of tasks are LLMs good at today and why.
- teiferer 11mo agoWhich is terrible. That's the root of all the BS around LLMs. People lacking understanding of what they are and ascribing capabilities which LLMs just don't have, by design. Even HN discussions are full of that. Even though this page literally has "hacker" in its name.
- kubb 11mo agoI’m trying not to be disappointed by people, I’d rather understand what’s going on in their minds, and how to navigate that.
- deleted 11mo ago[deleted]
- tim333 11mo agoI see your point but on the other hand a lot of conversations go: A: what will we do when AI do all the jobs, B: that's silly LLMs can't do the jobs. The thing is A didn't say LLM, they said AI as in whatever that will be a short while into the future. Which is changing rapidly because thousands of bright people are being paid to change it.
- HarHarVeryFunny 11mo agoThe trouble is that "AI" is also very much a leaky abstraction, which makes it tempting to see all the "AI" advances of recent years, then correctly predict that these "AI" advances will continue, but then jump to all sorts of wrong conclusions about what those advances will be. For example, things like "AI" image and video generation are amazing, as are things like AlphaGo and AlphaFold, but none of these have anything to do with LLMs, and the only technology they share with LLMs is machine learning and neural nets. If you lump these together with LLMs, calling them all "AI", then you'll come to the wrong conclusion that all of these non-LLM advances indicate that "AI" is rapidly advancing and therefore LLMs (also being "AI") will do too ... Even if you leave aside things like AlphaGo, and just focus on LLMs, and other future technology that may take all our jobs, then using terms like "AI" and "AGI" are still confusing and misleading. It's easy to fall into the mindset that "AGI" is just better "AI", and that since LLMs are "AI", AGI is just better LLMs, and is around the corner because "AI" is advancing rapidly ... In reality LLMs are, like AlphaFold, something highly specific - they are auto-regressive next-word predictor language models (just as a statement of fact, and how they are trained, not a put-down), based on the Transformer architecture. The technology that could replace humans for most jobs in the future isn't going to be a better language model - a better auto-regressive next-word predictor - but will need to be something much more brain like. The architecture itself doesn't have to be brain-like, but in order to deliver brain-like functionality it will probably need to include another half-dozen "Transformer-level" architectural/algorithmic breakthroughs including things like continual learning, which will likely turn the whole current LLM training and deployment paradigm on it's head. Again, just focusing on LLMs, and LLM-based agents, regarding them as a black-box technology, it's easy to be misled into thinking that advances in capability are broadly advancing, and will rise all ships, when in reality progress is much more narrow. Headlines about LLMs achievement in math and competitive programming, touted as evidence of reasoning, do NOT imply that LLM reasoning is broadly advancing, but you need to get under the hood and understand RL training goals to realize why that is not necessarily the case. The correctness of most business and real-world reasoning is not as easy to check as is marking a math problem as correct or not, yet that capability is what RL training depends on. I could go on .. LLM-based agents are also blurring the lines of what "AI" can do, and again if treated as a black box will also misinform as to what is actually progressing and what is not. Thousands of bright people are indeed working on improving LLM-adjacent low-hanging fruit like this, but it'd be illogical to conclude that this is somehow helping to create next-generation brain-like architectures that will take away our jobs.