5 ms·
AITemplate's original designer is here. We quit Meta in January and start HippoML (https://hippoml.com/ https://hippoml.com/). We just disclosed our new engine'
by antinucleon 3y ago
AITemplate's original designer is here. We quit Meta in January and start HippoML (https://hippoml.com/ https://hippoml.com/). We just disclosed our new engine's performance on LLM: https://blog.hippoml.com/large-language-model-inference-from-datacenter-to-edge-ed2f94da4a81 https://blog.hippoml.com/large-language-model-inference-from... On Apple M2 Max our new engine encode/decode is 13.8X/2.4X faster than llama.cpp
- ralfd 3y agoWhat is your planned business model here?
- antinucleon 3y agoWe will disclose more details very soon.
- brucethemoose2 3y agoVery interesting. Is 8bit/4bit support in the works? Will it work with bitsandbytes out of the box? Speedy inference is great, but in practice many users are running the biggest ~4-bit LLM that will fit into their RAM/VRAM pool these days. This is why llama.cpp is so good, its (AFAIK) the only implementation that will split a 4 bit quantized model so easily.
- antinucleon 3y agoYes. We support >= 1bit <= 16bit models out of box for various of models.
- mhh__ 3y agoReally doesn't surprise me that much. Llama.cpp seems like an OK first passs but I assume there is loads of time left on the table in terms of graph optimizations optimizing for the memory hierarchy properly.
- brucethemoose2 3y agoIt also doesn't use Apple GPUs at all. Its 100% CPU inference, with some CUDA/OpenCL (but no metal and no zero-copy) offload at the moment.
- antinucleon 3y agoIt is actually non-trivial to get GPU run fast, especially on SoC with strong CPU like M2.
- hutzlibu 3y agoGPU programming in general is definitely not trivial, as I can confirm with struggeling to learn WebGPU right now. But it really depends on the problem, simple math operations on lots of data is usually indeed trivially faster. Like AI mostly is with math on matrices. Or for example I just implemented a simple 2D raycastsolver in wgsl and as a first project it is totally not optimized - but even on my old laptop with crappy integrated GPU, but (relativly) fast CPU - I can now do 10000 raycasts per frame easily, while the cpu (wasm!) struggles with 500. The raw power of the gpu is really awesome. But every step is hard and debugging a mess. Which is why only a handful of people seems do be doing it. But now would probably be a good time, to get into it.. as I think gpu compute just has started and will get big.
- paulmd 3y agoI've been out of the space for a long time, and it's possible you know these already, but these are a couple weird tricks that can help: * Radix sort is your friend. Fun fact, O(n log n) is not the fastest a sort can run, it's the fastest a comparison-based sort can run. Radix sort runs in O(N) time, and in fact parallelizes extremely well on GPUs. Extremely. They are great at it. And there are in-place radix sorts too, just a bit slower (same asymptotic performance tho). * "Insert an element into this collection" style steps can be replaced by a sort and a prefix-sum operation. If you know the offset of the first element with key K, and you know the offset of the first element with key J, you know the offset and size of that "collection" for K within a flat array ("size(K) = offset(J) - offset(K)"). Both of these run super fast in parallel and if you can tweak your problem around to be some kind of sorting operation that usually produces good speedups like this. Easiest way to get a speedup from everything I've heard, if your program works fast from a lookup table it's a good approach. * Recomputing (or duplicate computation) is often much faster than storing intermediate results. "Procedural generation" is interesting because you can re-compute the generation step on demand. Random123 is also very nice compared to a (CuRand) mersenne twister/etc - why are you, a believer in the cryptographic maxim that hash(key, msg) is uncorrelated to hash(key, msg+1), still storing RNG state? Being able to play back arbitrary parts of a datastream at will is incredibly powerful, you can fastforward and rewind through the data previously used to interact with an item, as long as you know the epoch of the interaction you want for a particular key. And because computation is cheaper than memory, and memory bandwidth - it's really actually practically free in program time terms to just do some math. This is a form of data compression and performance enhancement. * Generally you must understand the idea of divergence and memory coalescing/alignment to keep those lanes executing. And it is highly preferable to use sorts and prefix scans and butterfly operations (reduction, etc) even within a warp, because traditional "mutex/atomic" paradigms don't work well with 100k threads. But this is just the programming idioms of this particular platform, I am sure LISP is similar too in terms of "oh that's how you do that" once you're accustomed. ("This thread would like to write 3 items to an output array, can I get an alignment into a datablock the warp group is going to allocate from a global counter via atomic_add and write to" is another common task element in many parallel workloads. That is the same prefix-sum idiom, or a butterfly reduce, just at a warp level. And then when all the warps have finished, you can sort this unsorted output by a key array to align it, and start processing it again...) * Cuda Unbound (CUB) provides reference implementations of all of these operations, and it handles portability across CUDA generations. Having good primitives makes a big difference, most of the workload isn't business logic, it's library calls, so using a library that just does it right gives you a lot of bang for buck. Don't write it yourself. * Texture maps aren't just for graphics, they are a black box that lets the GPU perform 2D and 3D coalescing and some interpolation. * Intelligent use of constants memory is another one, probably as is the use of CPU memory. If a value will be seldom accessed, you can probably stuff it into host memory and just accept the slowdown. Or you can store only epochs on the GPU and recompute intermediate values as needed. Try to ensure that all threads in a warp will do it too (host access vs recomputing). * Raytracing is of course impervious to all of this (so don't worry too much that you can't magically hammer a speedup out of it, nobody really can). You can accelerate the raycasting and intersection testing (and AMD and NVIDIA and Intel all do this differently) but as a general matter rays are completely random and uncoalesced and divergent, they bypass almost all of the ideal GPGPU "hot paths". Ray sorting/shader execution reordering is something that needs hardware assistance, and Intel and NVIDIA both have hardware along these lines. The idea of Intel of making a facility for async future/promise dispatch for sparse tasks (and then sorting the shaders to get good coalescing/etc) is really neat and they've said it's going to come to GPGPU. https://youtu.be/SA1yvWs3lHU?t=289 https://youtu.be/SA1yvWs3lHU?t=289 * You can, however, use your rays more efficiently. And that's an area of active focus for everyone. And I think more efficient use of TAAU samples is probably where raster is going too.
- huevosabio 3y agoAny idea how hippo, AI Template and TVM compare in performance?
- antinucleon 3y agoHippo is faster than AITemplate, and supports more generative models. We haven't compared vs TVM, but for absolute token/s on M2 Max, Hippo is able to run decoding on LLAMA with datacenter level GPUs performance (with other SW).
- huevosabio 3y agoThanks, I've added myself to the waitlist. Please let us know when this can be tried!
- choppaface 3y agoHow does Hippo compare with TensorRT?
- philipturner 3y agoThe difference in bandwidth between M2 Max and data centers GPUs isn't that much (less than a factor of 5). The difference in compute is much, much larger. If you only have fast GEMV kernels, and not fast GEMM kernels, you're basically locked into an inference engine that can only run GPT-style transformers efficiently. However, it can *technically* support all the models, but at what ALU utilization?
- thewataccount 3y agoDo you know how it's speed compares to exllama, specifically with an nvidia gpu by chance?
- antinucleon 3y agoWe haven't compared yet.
- yeison 3y agoDid Facebook invest in this. Is that why it's under Facebookincubator?
- antinucleon 3y agoWe developed AITemplate majorly for Meta's focus at that time, eg Ads/Ranking need. For HippoML is startup we are building for Generative AI. HippoML is not using AITemplate.
- yeison 3y agoWill this have some similarities to what Mojo is trying to solve?
- antinucleon 3y agoMojo is trying to create a new language to solve the problem, and specialized for CPU. We are using a more pragmatic way to solve GPU AI computation problem.
- jph00 3y agoMojo is not at all specialised for CPU. It sits on top of MLIR and excellent support for all major accelerators is planned.
- philipturner 3y agoFor bandwidth-bound problems like large language models, you could also solve it with properly written CPU kernels (Mojo) and usage of the AMX accelerators for compute-intensive parts. I'd be more interested if they had a GPU port of Stable Diffusion, which is really compute-intensive. That's where the GPU has a major advantage over CPU, on lower-end chips like M1/M2 and M1/M2 Pro. > specialized for CPU Mojo's SIMD execution model should map directly to GPUs. Instead of writing shaders thinking about a single GPU thread, you access the assembly ISA directly, thinking about an entire warp/simdgroup. That's how I think when writing SIMD-group matrix kernels anyway.
- sroussey 3y agoWould it work with instructor-xl or similar which is designed for embeddings and retrieval? On device for privacy is key.
- antinucleon 3y agoYes