12 ms·
RustGPT: A pure-Rust transformer LLM built from scratch
- bigmuzzy 1y agonice
- techsystems 1y ago> ndarray = "0.16.1" rand = "0.9.0" rand_distr = "0.5.0" Looking good!
- kachapopopow 1y agoI was slightly curious: cargo tree llm v0.1.0 (RustGPT) ├── ndarray v0.16.1 │ ├── matrixmultiply v0.3.9 │ │ └── rawpointer v0.2.1 │ │ [build-dependencies] │ │ └── autocfg v1.4.0 │ ├── num-complex v0.4.6 │ │ └── num-traits v0.2.19 │ │ └── libm v0.2.15 │ │ [build-dependencies] │ │ └── autocfg v1.4.0 │ ├── num-integer v0.1.46 │ │ └── num-traits v0.2.19 () │ ├── num-traits v0.2.19 () │ └── rawpointer v0.2.1 ├── rand v0.9.0 │ ├── rand_chacha v0.9.0 │ │ ├── ppv-lite86 v0.2.20 │ │ │ └── zerocopy v0.7.35 │ │ │ ├── byteorder v1.5.0 │ │ │ └── zerocopy-derive v0.7.35 (proc-macro) │ │ │ ├── proc-macro2 v1.0.94 │ │ │ │ └── unicode-ident v1.0.18 │ │ │ ├── quote v1.0.39 │ │ │ │ └── proc-macro2 v1.0.94 () │ │ │ └── syn v2.0.99 │ │ │ ├── proc-macro2 v1.0.94 () │ │ │ ├── quote v1.0.39 () │ │ │ └── unicode-ident v1.0.18 │ │ └── rand_core v0.9.3 │ │ └── getrandom v0.3.1 │ │ ├── cfg-if v1.0.0 │ │ └── libc v0.2.170 │ ├── rand_core v0.9.3 () │ └── zerocopy v0.8.23 └── rand_distr v0.5.1 ├── num-traits v0.2.19 () └── rand v0.9.0 () yep, still looks relatively good.
- cmrdporcupine 1y agolinking both rand-core 0.9.0 and rand-core 0.9.3 which the project could maybe avoid by just specifying 0.9 for its own dep on it
- Diggsey 1y agoIt doesn't link two versions of `rand-core`. That's not even possible with rust (you can only link two semver-incompatible versions of the same crate). And dependency specifications in Rust don't work like that - unless you explicitly override it, all dependencies are semver constraints, so "0.9.0" will happily match "0.9.3".
- eximius 1y agoThis doesn't sound right. If A depends on B and C - B and C can each bring their own versions of D, I thought?
- steveklabnik 1y agocan does not mean must. Cargo attempts to unify (aka deduplicate) dependencies where possible, and in this case, it can find a singular version that satisfies the entire thing.
- Diggsey 1y agoWithin a crate graph, for any given major version of a crate (eg. D v1) only a single minor version can exist. So if B depends on D v1.x, and C depends on D v2.x, then two versions of D will exist. If B depends on Dv1.2 and C depends on Dv1.3, then only Dv1.3 will exist. I'm over-simplifying a few things here: 1. Semver has special treatment of 0.x versions. For these crates the minor version depends like the major version and the patch version behaves like the minor version. So technically you could have v0.1 and v0.2 of a crate in the same crate graph. 2. I'm assuming all dependencies are specified "the default way", ie. as just a number. When a dependency looks like "1.3", cargo actually treats this as "^1.3", ie. the version must be at least 1.3, but can be any semver compatible version (eg. 1.4). When you specify an exact dependency like "=1.3" instead, the rules above still apply (you still can't have 1.3 and 1.4 in the same crate graph) but cargo will error if no version can be found that satisfies all constraints, instead of just picking a version that's compatible with all dependents.
- 0xffff2 1y ago
- imtringued 1y agocargo tree llm v0.1.0 (RustGPT) ├── ndarray v0.16.1 │ ├── matrixmultiply v0.3.9 │ │ └── rawpointer v0.2.1 │ │ [build-dependencies] │ │ └── autocfg v1.4. │ ├── num-complex v0.4.6 │ │ └── num-traits v0.2.19 │ │ └── libm v0.2.15 │ │ [build-dependencies] │ │ └── autocfg v1.4.0 │ ├── num-integer v0.1.46 │ │ └── num-traits v0.2.19 () │ ├── num-traits v0.2.19 () │ └── rawpointer v0.2.1 ├── rand v0.9.0 │ ├── rand_chacha v0.9.0 │ │ ├── ppv-lite86 v0.2.20 │ │ │ └── zerocopy v0.7.35 │ │ │ ├── byteorder v1.5.0 │ │ │ └── zerocopy-derive v0.7.35 (proc-macro) │ │ │ ├── proc-macro2 v1.0.94 │ │ │ │ └── unicode-ident v1.0.18 │ │ │ ├── quote v1.0.39 │ │ │ │ └── proc-macro2 v1.0.94 () │ │ │ └── syn v2.0.99 │ │ │ ├── proc-macro2 v1.0.94 () │ │ │ ├── quote v1.0.39 () │ │ │ └── unicode-ident v1.0.18 │ │ └── rand_core v0.9.3 │ │ └── getrandom v0.3.1 │ │ ├── cfg-if v1.0.0 │ │ └── libc v0.2.170 │ ├── rand_core v0.9.3 () │ └── zerocopy v0.8.23 └── rand_distr v0.5.1 ├── num-traits v0.2.19 () └── rand v0.9.0 ()
- tonyhart7 1y agois this satire or does I must know context behind this comment???
- stevedonovan 1y agoThese are a few well-chosen dependencies for a serious project. Rust projects can really go bananas on dependencies, partly because it's so easy to include them
- obsoleszenz 1y agoThe project only has 3 dependencies which i interpret as a sign of quality
- leoh 1y agoI don't know if OP intended satire, but either way it is an absurd comment. Think about how "from scratch" this really is.
- worldsavior 1y agoThis doesn't mean anything. A project can implement things from scratch inefficiently but there might be other libraries the project can use instead of reimplementing.
- Charon77 1y agoAbsolutely love how readable the entire project is
- yieldcrv 1y agoNever knew Rust could be that readable. Makes me think other Rust engineers are stuck in a masochistic ego driven contest, which would explain everything else I've encountered about the Rust community and recruiting on that side.
- jmaker 1y agoNot sure what you’re alluding to but that’s just ordinary Rust without performance or async IO concerns.
- GardenLetter27 1y agoMost Rust code looks like this - only generic library code goes crazy with all the generics and lifetimes, due to the need to avoid unnecessary mallocs and also provide a flexible API to users. But most people aren't writing libraries.
- cmrdporcupine 1y agoDon't underestimate what some programmers trying to prove their cleverness (or just trying to have fun) can do if left unchecked. I think most Rust code does indeed look like this but I've seen plenty of projects that go crazy with lifetimes and generics juggling where they don't have to.
- emporas 1y agoIt is very procedural/object oriented. This is not considered good Rust practice. Iterators make it more functional, which is better, more succinct that is, and enums more algebraic. But it's totally fine for a thought experiment.
- deleted 1y ago[deleted]
- ndai 1y agoI’m curious where you got your training data? I will look myself, but saw this and thought I’d ask. I have a CPU-first, no-backprop architecture that works very well on classification datasets. It can do single‑example incremental updates which might be useful for continuous learning. I made a toy demo to train on tiny.txt and it can predict next characters, but I’ve never tried to make an LLM before. I think my architecture might work well as an on-device assistant or for on-premises needs, but I want to work with it more before I embarrass myself. Any open-source LLM training datasets you would recommend?
- kachapopopow 1y agohuggingface has plenty of openai and antrophic user to assistant chains, beware there are dragons (hallucinations), but good enough for instruction training. I actually recommend distilling kimi k2 instead for instruction following capabilities.
- electroglyph 1y agohttps://huggingface.co/datasets/NousResearch/Hermes-3-Dataset https://huggingface.co/datasets/NousResearch/Hermes-3-Datase...
- Snuggly73 1y agoTo my untrained eye, this looks more like an instruct dataset. For just plain text, I really like this one - https://huggingface.co/datasets/roneneldan/TinyStories https://huggingface.co/datasets/roneneldan/TinyStories
- hadlock 1y agoThe training data is contained inside main.rs; it looks like about 50 statements about general facts. Probably to keep training time down. If you go off-script things fall apart pretty quickly: Enter prompt: hello Model output: What are eclipses ? Assistant : Eclipses occur when one celestial body moves into the shadow of another </s> Enter prompt: what are facts Model output: eclipses Assistant : Eclipses Eclipses What What was when What through you ? through you meet through using of What was What tall ? of What thank help you explain using , and do you a patterns you you a patterns through air in tall you help metal metal </s> Enter prompt: how are mountains formed? Model output: I ? ' I ' : Of : in happy Hi wind in yeast altering it it </s>
- kachapopopow 1y agoThis looks rather similar to when I asked an AI to implement a basic xor problem solver I guess fundementally there's really only a very limited amount of ways to implement this.
- Goto80 1y agoNice. Mind to put a license on that?
- thomask1995 1y agoLicense added! Good catch
- untrimmed 1y agoAs someone who has spent days wrestling with Python dependency hell just to get a model running, a simple cargo run feels like a dream. But I'm wondering, what was the most painful part of NOT having a framework? I'm betting my coffee money it was debugging the backpropagation logic.
- taminka 1y agolowkey ppl who praise cargo seem to have no idea of the tradeoffs involved in dependency management the difficulty of including a dependency should be proportional to the risk you're taking on, meaning it shouldn't be as difficult as it in, say, C where every other library is continually reinventing the same 5 utilities, but also not as easy as it is with npm or cargo, because you get insane dependency clutter, and all the related issues like security, build times, etc how good a build system isn't equivalent of how easy it is include a dependency, while modern languages should have a consistent build system, but having a centralised package repository that anyone freely pull to/from, and having those dependencies freely take on any number of other dependencies is a bad way to handle dependencies
- itsibitzi 1y agoWhat tool or ecosystem does this well, in your opinion?
- taminka 1y agoany language that has a standardised build system (virtually every language nowadays?), but doesn't have a centralised package repository, such that including a dependency is seamless, but takes a bit of time and intent i like how zig does this, and the creator of odin has a whole talk where he basically uses the same arguments as my original comment to reason why odin doesn't have a package manager
- zoobab 1y ago"a standardised build system (virtually every language nowadays?)" Python packages still manage poorly dependencies that are in another lang like C or C++.
- enricozb 1y agoI did this [0] (gpt in rust) with picogpt, following the great blog by jaykmody [1]. [0]: https://github.com/enricozb/picogpt-rust https://github.com/enricozb/picogpt-rust [1]: https://jaykmody.com/blog/gpt-from-scratch/ https://jaykmody.com/blog/gpt-from-scratch/
- trackflak 1y ago[dead]
- abricq 1y agoThis is great ! Congratulations. I really like your project, especially I like how easily it is to peak at. Do you plan on moving forward with this project ? I seem to understand that all the training is done on the CPU, and that you have next steps regarding optimizing that. Do you consider GPU accelerations ? Also, do you have any benchmarks on known hardware ? Eg, how long would it take to train on a macbook latest gen or your own computer ?
- thomask1995 1y agoHI! OG Author here. Honestly, I don't know. This was purely a toy project/thought experiment to challenge myself to learn exactly how these LLMs worked. It was super cool to see the loss go down and it actually "train". This is SUPER far from a the real deal. Maybe it could be cool to see how far a fully in memory LLM running on CPU can go.
- ramon156 1y agoCool stuff! I can see some GPT comments that can be removed // Increased for better learning this doesn't tell me anything // Use the constants from lib.rs const MAX_SEQ_LEN: usize = 80; const EMBEDDING_DIM: usize = 128; const HIDDEN_DIM: usize = 256; these are already defined in lib.rs, why not use them (as the comment suggests)
- sloppytoppy 1y ago[flagged]
- deleted 1y ago[deleted]
- tialaramex 1y agoFor the constants is it possible the author didn't know how? I remember in my first week of Rust I didn't understand how to name things properly, basically I was overthinking it.
- vlovich123 1y agoLots of signs this is an LLM-generated project. All the emojis in the README are a hint as well.
- tayo42 1y agoFrom his reddit post https://old.reddit.com/r/rust/comments/1nguv1a/i_built_an_llm_from_scratch_in_rust_just_ndarray/ne8hu1m/ https://old.reddit.com/r/rust/comments/1nguv1a/i_built_an_ll...
- ericdotlee 1y agoDo you think vibe coded rust will rot the quality of language code generally?
- adastra22 1y agoThese things will be corrected over time.
- jlmcgraw 1y agoSome commentary from the author here: https://www.reddit.com/r/rust/comments/1nguv1a/i_built_an_llm_from_scratch_in_rust_just_ndarray/ https://www.reddit.com/r/rust/comments/1nguv1a/i_built_an_ll...
- Snuggly73 1y agoCongrats - there is a very small problem with the LLM - its reusing transformer blocks and you want to use different instances of them. Its a very cool excercise, I did the same with Zig and MLX a while back, so I can get a nice foundation, but since then as I got hooked and kept adding stuff to it, switched to Pytorch/Transformers.
- capestart 1y ago[dead]
- Emma_Schmidt 1y ago[dead]
- zenlot 1y agoRust == stars in GitHub.
- bionhoward 1y agoThat time to first token is impressive, it seems like it responds immediately
- ericdotlee 1y agoThis is incredibly cool, but I wonder when more of the AI ecosystem will move past python tooling into something more... performant? Very interesting to already see rust based inference frameworks as well.
- lutusp 1y agoIt would have been nice to see a Rust/Python time comparison for both development and execution. You know, the "bottom line"?
- amoskvin 1y agogreat job! which model does it implement? gpt-2?
- yobbo 1y agoVery nice! Next thing to add would be numerical gradient testing.
- tripplyons 1y agoIs that where you approximate a partial derivative as a difference in loss over a small difference in a single parameter's value? Seems like a great way to verify results, but it has the same downsides as forward mode automatic differentiation since it works in a pretty similar fashion.
- yobbo 1y agoYes, the purpose is to verify the gradient computations which are typically incorrect on the first try for things like self-attention and softmax. It is very slow. It is not necessary for auto-differentiation, but this project does not use that.
- chcardoz 1y agosuper fun!! I am running it right now and going to use it to train on a corpus of my own writing to make a gpt of myself.
- om8 1y agoHave a similar project. Also written in rust, runs in a browser using web assembly In-browser demo: https://galqiwi.github.io/aqlm-rs https://galqiwi.github.io/aqlm-rs Source code: https://github.com/galqiwi/demo-aqlm-rs https://github.com/galqiwi/demo-aqlm-rs
- selinkocalar 1y agoThe memory safety guarantees in Rust are probably useful here given how easy it is to have buffer overflows in transformer implementations. CUDA kernels are still going to dominate performance though. Curious about the tokenization approach - are you implementing BPE from scratch too or using an existing library?
- farhanhubble 1y agoI've not written a single line of Rust ever, but I have occasionally looked under the hood of Tensorflow, Pytorch etc. and have been a machine learning practitioner for several years. The succinctness of the interfaces surprised me!