4 ms·
Gemma 2B 60 tk/s WebGPU on M1
- lccccc 3y agoThis looks far faster than TVM's one?!
- mjks 3y ago[flagged]
- karmasimida 3y agoBut what good is 2B for? I didn’t see a use case tbh
- lccccc 3y agoWill Google add performant 7B support? 7B's quality is good enough for a serious Web APP.
- reqo 3y agoFine-tuning on a niche task I guess! E.g. a bunch of LoRAs that can be easily swapped depending on the task!
- karmasimida 3y ago2B's hallucination will be glorious I feel
- rvz 3y agoThis is the problem. There isn't any use-cases for this other than being a toy. Even on the web.
- deleted 3y ago[deleted]
- qrian 3y agoIt only says 'affor affor affor...(repeated 100+ times)' for me. I am not sure how to debug this.
- mjks 3y agoBut is it saying "'affor affor affor...(repeated 100+ times)" really fast? But actually, did you download the linked model from Kaggle ("gemma-2b-it-gpu-int4.bin") and upload it into the demo? It is working fine for me out of the box.
- qrian 3y agoyes, I loaded the 1.35GB .bin file.
- hackerrrrx 3y agoAre u using 'gpu' one? what's your platform. my mbp2019 seems good
- magemgem 3y agoTry a different prompt
- qrian 3y agonope, all my prompts lead the the same answer.
- miohtama 3y agoIt thinks it’s a pokemon.
- impjdi 3y agothis can happen if you try to feed a cpu model to gpu inference. make sure you download the gpu model if you are using gpu accelerated ml inference.
- hackerrrrx 3y agoWow!