3 ms·
Prompt processing on Macs is VERY slow...
by pulse7 15d ago
Prompt processing on Macs is VERY slow...
- CoastalCoder 15d ago> Prompt processing on Macs is VERY slow... That sounds like a contradiction. /J
- larodi 15d agowhich engine? cause there are dozens already, and some are quite fast.
- embedding-shape 15d agoWhich one is the fastest and how fast is it? Even the "fast" ones doesn't seem to even reach close to consumer NVIDIA GPUs released years ago when it comes to prompt processing.
- anonreplier 15d agolarge model and slow, or speed and a teeny model, make your choice
- embedding-shape 15d agoOr you get N of RTX Pro 6000 and get large and fast models :) Make your choice, and do it before the prices go up even more.
- larodi 15d agoindeed. make up your mind, and also "slow/fast" mean very little given plethora of models to choose from. fast for what, slow for what. fast with which harness, etc...
- anonreplier 15d agowell yes if money is no object, why not
- ekianjo 15d agonone of them are Nvidia GPU FAST
- thecolorblue 15d agoBut more memory is available so a larger model can be loaded. It depends on the use case which is better. Any chat or voice model will have better UX with nvidia but document or code generation will be better with apple.
- ekianjo 13d agoit's not a binary thing. At some point Apple becomes too slow with very large models. If you can just run a model at 1 token per second and it takes 30 mins to process a long context, it's useless