3 ms·
Here's a Rust version in case anyone's curious what it would look like. It also clocks 106 tokens/second in release mode. https://github.com/garrisonhess/llam
by gman_ 3y ago
Here's a Rust version in case anyone's curious what it would look like. It also clocks 106 tokens/second in release mode.
https://github.com/garrisonhess/llama2.c/blob/517a1a3e487f315200d6ccffe92b2fd00e2575aa/src/main.rs https://github.com/garrisonhess/llama2.c/blob/517a1a3e487f31...
- wtarreau 3y agoAs often with Rust, someone transliterates something that already exists just because they can, without providing any benefit at all. Sometimes it even results in fragmenting the community efforts to improve the project.
- voz_ 3y agoCan you chill? Stuff like this is super useful. The original c file is educational, so is this. And now by having it two ways, we have a tiny little Rosetta Stone for folks that wanna learn.
- jheuel 3y agoMight also help as input to the next iteration of LLMs
- gman_ 3y agoLooks like you spoke too soon, I'm clocking 340+ tokens per second with my improved Rust implementation, compared to 106 with the original C. That being said, I didn't share this for any reason other than to share ideas and promote learning. Cheers
- elteto 3y agoYou should list your email in your HN profile. That way the Internet could check with you to see if you approve whenever someone starts a new personal project.
- l-m-z 3y agoAnother random (self) plug for a rust version, this uses the candle ML library we've been working on for the last month and can be run in the browser. https://laurentmazare.github.io/candle-llama2/index.html https://laurentmazare.github.io/candle-llama2/index.html The non-web version has full GPU support but is not at all minimalist :)
- ianpurton 3y agoReally nice.