Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
gitpusher42
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
31.
▲
by
gitpusher42
2mo ago
There was an ai winter for very long time. The math for NNs was already here, but not enough compute/data I saw a pretty cool project to run an llm on an esp32 device https://github.com/slvDev/esp32-ai
32.
▲
by
gitpusher42
2mo ago
Thank you! afaik ollama relies on llama.cpp and mmap. mmap loads pages on demand and doesn't use the same explicit cache or parallel reads like my engine. Most likely ollama/llama.cpp will be way slower in this case
33.
▲
by
gitpusher42
2mo ago
not with this engine. Kimi is a very different model. You can try to check HN later, I believe someone will build engine for this model for edge devices
34.
▲
by
gitpusher42
2mo ago
Thank you! Let me know how it goes and share your tok/s results
35.
▲
by
gitpusher42
2mo ago
I tried it. madvise didn't make mmap better than pread. I also tested F_RDADVISE, it helped on short decodes, but somehow got worse on longer decodes. Not very clear why, most likely problem somewhere at APFS and it is closed source an
36.
▲
by
gitpusher42
2mo ago
m5 device also uses the same approach. the same 16 cache slots. and experts are evicted from memory as needed. And activity monitor shows 2gb usage for m5 pro (24gb btw)
37.
▲
by
gitpusher42
2mo ago
Uh, not really, unfortunately. Asahi uses Vulkan for gpu, but kernels for this project are metal.
38.
▲
by
gitpusher42
2mo ago
Yeah, correct! You can set up this engine to get more expert cache slots (e.g 32 instead of 16) to get a better hit rate and better tok/s. it will be 3.5gb instead of 2gb.
39.
▲
by
gitpusher42
2mo ago
I double checked m2 logs. Cache hit rate is about 59-69%. 250-320MB went through `pread` per generated token. It is 3gb/s during this i/o phase.
40.
▲
by
gitpusher42
2mo ago
Thank you very much! I think it will throttle quite soon, but I haven't tried runs longer than 30minutes with this engine. However, there is no constant load on ssd or gpu. i/o and gpu work are alternating and there is a brief idl
41.
▲
by
gitpusher42
2mo ago
For Mac I would start from MLX engine. For exact model choice it is better to check bench results, and select model based on your need. A lot of good feedback about Qwen3.6, but I haven't used it in my tasks
42.
▲
by
gitpusher42
2mo ago
Thank you! Text is not the main part of this repo. The main part is the technology and the list of experiments (and some useful knowledge I got from this project, haha). I have always been bad at writing or editing text (in both my native l
43.
▲
by
gitpusher42
2mo ago
I am not native, my English is far away from perfect. I am using LLMs for checking my texts or grammar. I always trying to edit it properly, but sometimes I missing parts like that because I don't really have this "language feelin
44.
▲
by
gitpusher42
2mo ago
I think it depends on usage pattern. You trade speed for lower memory usage. Maybe engine specialisation and faster SSDs is the future for local inference, who knows
45.
▲
by
gitpusher42
2mo ago
Thanks! I wanted to add hugging face token field to speed up model download, but then I realised that people might not trust to give their tokens And yeah, local models are better for security, at least your conversation stay on the machine
46.
▲
by
gitpusher42
2mo ago
This approach will only work for MoE models. There is a Qwen 35b-a3b. You just need to do GPU stop after router and read the requested experts to ram. And it is possible to build similar engine for this model (or feel free to adopt my engin
47.
▲
by
gitpusher42
2mo ago
Thank you! Thankfully this is purely software engineering problem. And as usual there is no free lunch. You trade speed for lower memory usage
48.
▲
by
gitpusher42
2mo ago
Size reduction is mostly based on Experts size. And it is limited by SSD speed. Check for Colibri and Flash-Moe, they are doing similar things with bigger models, but tok/s is not high
49.
▲
by
gitpusher42
2mo ago
Switching from mmap to parallel pread. From 0.5tok/sec to almost 4tok/sec. Running GPU work while reading missed experts also helped a lot, 4.4 -> 4.7
50.
▲
by
gitpusher42
2mo ago
Thanks for sharing! SSD read speed is the biggest limiting factor here, unfortunately
51.
▲
by
gitpusher42
2mo ago
I am relying quite heavily on system caching and pread. And yeah, M5 is a way faster and I can guess Mac can cache something, even if process stays under 2gb. It was 83ms read per token for M2 and 12ms on M5 pro. Total is 163ms/tok vs
52.
▲
by
gitpusher42
2mo ago
My friend just compared this project to running Cyberpunk on a very old machine at 16 FPS, huh
53.
▲
by
gitpusher42
2mo ago
Thank you! Share your tok/s results later
54.
▲
by
gitpusher42
2mo ago
Thank you! That’s useful. I might try lowering the minimum version later. The 2.4x prefill improvement will only work on the apple10 GPU family. The M1 uses apple7 as I remember
55.
▲
by
gitpusher42
2mo ago
You’re right, Gemma isn’t the best model for coding (afaik more "everyday tasks" related). My first idea was to use Qwen, but its architecture was much more complex to implement in this stack. I chose Gemma so I wouldn’t spend all
56.
▲
by
gitpusher42
2mo ago
Yeah! Check the Colibri and Flash-MoE projects. They’re already doing that. https://github.com/danveloper/flash-moe https://github.com/JustVugg/colibri
57.
▲
by
gitpusher42
2mo ago
It is super cool! Diffusion Gemma was released around the middle of my project, and I seriously considered switching to it. But I decided to finish the project as it was. I believe it would be a perfect match! Feel free to use any parts of
58.
▲
by
gitpusher42
2mo ago
I think I first saw Flash-MoE ( https://github.com/danveloper/flash-moe ) in April. Huge respect to them, it was a big inspiration for this project!
59.
▲
by
gitpusher42
2mo ago
My first version used plain `mmap`. On the 8 GB M2, a cold 3.36 MB expert took 10 ms with mmap and 2.8 ms with `pread`. The full simulation was 0.50tok/s for `mmap` vs 4 tok/s for `pread` With `mmap`, OS loads pages reactively as
60.
▲
by
gitpusher42
2mo ago
What exact specs do you have? It might be because it's the 256 GB version. afaik, those versions have much slower memory bandwidth than the 512 GB models My friend tried it on an M4 MacBook Pro and got 25–27 tok/s
More ›