3 ms·
This is an interesting idea, with the main non-trivial win really being vectorized GBDT inference. Instead of serially going down a DT, you can convert it to ve
by vladf 6y ago
This is an interesting idea, with the main non-trivial win really being vectorized GBDT inference. Instead of serially going down a DT, you can convert it to vectorizable GEMM code.
By a cursory look at hummingbird/ml/operator_converters/_tree_commons.py this seems to just be doing a GraphBLAS-style graph traversal with dense GEMM, which strikes me as resulting in incredibly redundant computation. I think this would only give you acceleration for very small trees (edit: granted, I think that's what the defaults are for a lot of these packages).
I'd be interested in seeing how this stacks up against:
* native GPU execution for XGB/LGBM
* a _sparse_ GEMM implementation of their algorithm
* the classical reference for DT vectorization, Quickscorer, see https://github.com/hpclab/quickscorer https://github.com/hpclab/quickscorer
- interesaaat 6y agoIn the technical report (https://scnakandala.github.io/papers/TR_2020_Hummingbird.pdf https://scnakandala.github.io/papers/TR_2020_Hummingbird.pdf) table 8 we compared against NVIDIA FIL library. From their blogpost (https://medium.com/rapids-ai/rapids-forest-inference-library-prediction-at-100-million-rows-per-second-19558890bc35 https://medium.com/rapids-ai/rapids-forest-inference-library...) they claim to be faster than the original XGB\LGBM GPU implementation, so we decided to pick FIL as GPU baseline. Regarding the GEMM strategy: you are right, it only works for shallow trees. In Figure 7 we compared GEMM against the other strategies implementing typical tree traversal. At some point we had an implementation using sparse GEMM, but only for sparse input feature vectors.