5 ms·
Have you ever wanted to boot and inference a herd of 1000's of Virtual baby Llama 2 models on big ass enterprise servers? No? Well, now you can! (Almost baremet
by AMICABoard 3y ago
Have you ever wanted to boot and inference a herd of 1000's of Virtual baby Llama 2 models on big ass enterprise servers? No? Well, now you can! (Almost baremetal)
Also drop the binary portable run.com cosmocc build on any OS and run! Truly portable. (Soon baremetal)
Special Thanks & Credits:
llama2.c - @karpathy
cosmopolitan - @jart
unikraft - @unikraft
Would love to hear your feed back here!
- dingdingdang 3y agoWould love some examples, practical a-z stuff; how does one make a couple of these run on a local (yes, large) server for instance?
- Datagenerator 3y agoThis and also how to let them talk to each other and follow the conversation after setting up only the initial topic - that would be awesome!
- AMICABoard 3y agoIf you add it as an issue, I'll address it in future. I like the idea, but with the current toy size small models it won't be too fun (I did a manual try). Once larger models run at good pace, it would be absolutely cool.
- Datagenerator 3y agoThanks and love the reference to the wonderful blinking guru meditation we used to see on our Amiga
- AMICABoard 3y agoI grew up with the Amiga 500, it was my first love, how can I forget those times!
- AMICABoard 3y agoIt's just the beginning, still optimization and some figuring out to do. This fork is based on karpathy's llama.c and we try to mirror its progress and add our patches to it to add performance, be binary portable or run as unikernel. However there is a catch, this doesn't currently infer the 7b or bigger meta llama 2 models yet. It's too slow and memory consuming. My plan is to get to a stage where we can actually infer larger models at a comfortable speed like llama.cpp / ggml does, add GPU acceleration along the way. Also this doesn't have a web api, I'll be adding that in the next update, then it would actually make sense to deploy it on a server to test it out. Right now you would have to manually spawn vm instances with qemu like this: qemu-system-x86_64 -m 256m -accel kvm -kernel L2E_qemu-x86_64 or qemu-system-x86_64 -m 256m -accel kvm -kernel L2E_qemu-x86_64 -nographic and that's not very practical especially as there is no web api yet. So see this more as a tech preview - release early release often thing. Yeah as I'll get time, I'll be adding a build for firecracker and also write instructions to spawn 100's of baby llama 2 kvm qemu / firecracker builds on a powerful server. Thank you for your interest. As per your suggestion, a comprehensive howto is planned. Feel free to add any issue / wants / suggestions to https://github.com/trholding/llama2.c https://github.com/trholding/llama2.c , I'll address those as I get time. I'm stuck with bigger IRL projects, but if there is deep interest from the community I'll be sure to spend more time on this.
- deleted 3y ago[deleted]