4 ms·
DSV4 Flash 0731 already runs on RTX 4090 24GB + 128GB system RAM at a usable tok/s and quantization.
by 127 2mo ago
DSV4 Flash 0731 already runs on RTX 4090 24GB + 128GB system RAM at a usable tok/s and quantization.
- Gecko4072 2mo agoYou personally? Just curious. Context window is also a factor and ram isn’t really cheap. Sparks are assembled units which I like.
- dannyw 2mo agoFor the same price as a DGX Spark here (A$8499) I can buy roughly 544GB of DDR5-5200MHz from retail; which on a quad channel platform would deliver ~160gb/s real world; and ~320gb/s with octa channels (Xeon, Threadripper Pro). If you can afford it or somehow find a used unit, you can go Epyc for 12 channels. 8/12 channel DDR5 will beat DGX Spark in inference/decode even without a GPU of any kind, as it’s memory bandwidth bound, and the Spark tops out at ~240gb/s real world. With some optimisation and maths, it’s entirely plausible to ach You are paying an extraordinary amount of money for the convenience of a super small unit, with still mediocre software support, but at least a community. Expect to be crawling through forum posts regularly, as SM121/Spark has many quirks and ecosystem issues still. Please don’t pay another 70-80% gross margins on top of already inflated DRAM prices unless you need. The Spark IS really nice if you want to test out ConnectX or if you really need something small and compact and quiet. Also consider: used Adas or even Ampere NVIDIA workstation GPUs can come with a lot of VRAM and be “reasonable”, with CUDA.
- danielEM 2mo agoBeen investigating these multichannel AMD based platforms last year and seem like none of them can in real scenarios utilize anywhere close to their theoretical bandwidth.
- kybernetikos 2mo agoI've run it with a large context window on 256GB ram + 4090. It wasn't super fast, but it was manageable and it completed the tasks I gave it well.
- xiconfjs 2mo agoHow many tps and size of ctx?