2 ms·
Very nice indeed. Model weights in RAM blocks distributed all over a big FPGA: should be super helpful at minimizing RAM bandwidth bottlenecks. To say nothing o
by RetroTechie 2mo ago
Very nice indeed. Model weights in RAM blocks distributed all over a big FPGA: should be super helpful at minimizing RAM bandwidth bottlenecks. To say nothing of latency.
But model(s) implemented are clearly too small to be useful as a 'chat partner'. Tried a couple of sentences - replies is just some gibberish coming out.
This really needs a bigger FPGA, or some other application(s) where a tiny LLM does actually useful work. Barring that, generated tokens/sec is kind of a meaningless measure imho.
- imtringued 2mo agoObviously at 20k tokens per second your primary goal would be some sort of time series model running at 20KHz for processing sensor data and I'd say the model might even be too big for that. You could probably run a bunch of sensors at once.