Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
karpathy
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
31.
▲
by
karpathy
2y ago
I think at this point docs should start to be written not for humans but for LLMs, i.e. package all of it up with the --help into one giant txt file for easy attachment to an LLM when asking the question you'd like. Imo it's a rel
32.
▲
by
karpathy
2y ago
Can this be put in analogy to arithmetic and calculators? People had to be a lot better at mental math and calculator removed the pressure. You could imagine making similar arguments that losing the ability would be disastrous. The reasons
33.
▲
by
karpathy
2y ago
(H100 SXM is 1000 TFLOPS, *2 is from "with sparsity", which is not used here.)
34.
▲
by
karpathy
2y ago
fwiw I totally understand the sentiment! it's actually a bit sad to me that so much of our content is moving from the shared, open web to platforms like twitter, unfortunately there seems to be too much value add around built-in discov
35.
▲
by
karpathy
2y ago
That is 100% my intention and hope and I think we are very close to deleting all of that. Right now on master, I am already only using Python for the tokenization preprocessing. In principle the requirements for llm.c should be extremely mi
36.
▲
by
karpathy
2y ago
The 350M model I trained last night was 30B tokens, 14 hours, ~$200. Conveniently, 300B is exactly 10X the tokens so ~$2K would be the estimate. You'd have to wait 140 hours on one box though. Getting an H100 box instead of A100 will a
37.
▲
by
karpathy
2y ago
sounds good. both work, (though) I think HN has a bit of an anti-twitter bias.
38.
▲
by
karpathy
2y ago
Zero To Hero doesn't make it all the way to a chatbot, it stops at pretraining, and even that at a fairly small scale or character-level transformer on TinyShakespeare. I think it's a good conceptual intro but you don't get t
39.
▲
by
karpathy
2y ago
Yes definitely. Related tweet of mine: https://x.com/karpathy/status/1760388761349927356?lang=en 1. Build the thing 2. Build the ramp Currently on step 1 :). It helps to build it first so you know where you are go
40.
▲
by
karpathy
2y ago
My understanding and suspicion is mostly less than you think. Llama 3 architecture has the following changes on GPT-2: 1. delete the absolute positional encoding and replace with RoPE 2. delete all biases in all layers (in LayerNorms, they
41.
▲
by
karpathy
2y ago
The baseline is definitely PyTorch (or JAX), and indeed something like nanoGPT. I just never got nanoGPT "past the finish line" of really crossing the t's and dotting the i's and reproducing the models with as much care
42.
▲
by
karpathy
2y ago
Hi HN the main (more detailed) article is here https://github.com/karpathy/llm.c/discussions/481 Happy to answer questions!
43.
▲
by
karpathy
2y ago
The paper mentions some reasons why these quick fix ideas are not as simple as it sounds. For example many rare tokens are “intermediate” merges inside the BPE algorithm, shorter prefixes of longer words. The long word is common, but its ea
44.
▲
by
karpathy
3y ago
Hah why is this on HN today? Update from 2024: https://github.com/tysam-code/hlb-CIFAR10 Train to 94% on CIFAR-10 in <6.3 seconds on a single A100. Or ~95.79% in ~110 seconds (or less!) https://paperswith
45.
▲
by
karpathy
3y ago
I wrote this, which might be a bit helpful: https://github.com/karpathy/llm.c/blob/master/doc/layernorm/... But if you don't have the background, I'd recommend my YouTube videos, see
46.
▲
by
karpathy
3y ago
Very impressive looking! Just wanted to caution it's worth being a bit skeptical without benchmarks as there are a number of ways to cut corners. One prominent example is heavy model quantization, which speeds up the model at a cost of
47.
▲
by
karpathy
3y ago
Hi all, it's still too early to look at this code :) I wanted to put up my alpha version to start the feedback going a bit. I'm recording a video alongside where we build minBPE and expect that to be a lot more useful, coming out
48.
▲
by
karpathy
3y ago
In principle easy and possible, just not exactly useful. Would just involve adding the backward pass. But I’m not sure that this is something many people would want.
49.
▲
by
karpathy
3y ago
Yes :(
50.
▲
by
karpathy
3y ago
Still training. I will put it in readme
51.
▲
by
karpathy
3y ago
It’s not supposed to infer beyond max seq len right now, it’s undefined behavior. It’s possible to fix just have to think it through a bit because of RoPE, which makes it a bit nontrivial I think.
52.
▲
by
karpathy
3y ago
Yay fun to see it make its way to HN :) It turns out that my original checkpoint runs _way_ faster than I expected (100 tok/s) on MacBook Air M1 with -O3 when compiling, so I am now training a bigger 44M model, which should still runni
53.
▲
by
karpathy
4y ago
:(
54.
▲
by
karpathy
4y ago
Yes but in the same way as saying that computers are just a Markov chain.
55.
▲
by
karpathy
4y ago
there were no non-competes, garden leaves or etc., i just took some time for myself and then ~late last year started to feel an itch again.
56.
▲
by
karpathy
4y ago
rough steps: 1. collect a very large dataset, see: https://www.lesswrong.com/posts/6Fpvch8RR29qLEWNH/chinchilla... . scrape, de-duplicate, clean, wrangle. this is a lot of work regardless of $. 2. get on a call wi
57.
▲
by
karpathy
4y ago
I was also a bit surprised that the Chinchilla numbers and tables don't reproduce and that there are calculation bugs in the paper (e.g. the FLOPs calculation in the paper is wrong), especially because the paper has been so impactful i
58.
▲
by
karpathy
4y ago
Ty agree, most people practically speaking will be interested in finetuning rather than from-scratch pretraining. I currently have some language about it in readme but I agree this should get more focus, docs, examples, etc.
59.
▲
by
karpathy
4y ago
Wow, fun to find this trending on HN this morning! I am currently also working on the associated video lecture (as the next episode of my video lecture series here https://karpathy.ai/zero-to-hero.html ), where I will build
60.
▲
by
karpathy
4y ago
Elon also understands deep neural nets a lot more than I think people imagine. He starts with good intuitions and mental models, but also actively asks for technical deep dives, and has very good retention. E.g. I recall teaching him about
More ›