3 ms·
I have not measured pre-fill, but it's said to be around 1000. It feels very snappy and unlike my experience with running 27B models the performance stays pret
by UncleOxidant 17d ago
I have not measured pre-fill, but it's said to be around 1000.
It feels very snappy and unlike my experience with running 27B models the performance stays pretty flat even as the context increases. Unfortunately, we don't know how Halogen is doing this because it's closed source, but I think AMD should offer that guy some $$$ because he's done a lot of good work getting more performance out of Strix Halo.
- syntaxing 17d agoHalogen as in this right? https://github.com/peonist-ai/halogen-flash-server https://github.com/peonist-ai/halogen-flash-server
- UncleOxidant 17d agoyes
- syntaxing 17d agoThanks! I really like how the author packaged everything into a container. Definitely going to give it a go over the weekend!