6 ms·
Decision trees – the unreasonable power of nested decision rules
- xmprt 7mo agoInteresting website and great presentation. My only note is that the color contrast of some of the text makes it hard to read.
- thesnide 7mo agoexactly my thought. and here thr reader view of FF is a godsend. having 'accessible' content is not only for people with disabilities, it also help with bad color taste. well, at least bad taste for readable content ;)
- defanor 7mo agoThe FF reader view here starts from "We just saw how a Decision Tree", gobbling up half the article. Simply disabling CSS works better. Though in both cases, it seems that ordering might be a bit mixed up.
- Jaxon_Varr 7mo ago[dead]
- moi2388 7mo agoThat was beautifully presented!
- fooker 7mo agoFun fact - single bit neural networks are decision trees. In theory, this means you can 'compile' most neural networks into chains of if-else statements but it's not well understood when this sort of approach works well.
- Almondsetat 7mo agoDo you know of any software that does this? Or any papers on the matter? It could be a fun weekend project
- tomashubelbauer 7mo agoMade me think of https://github.com/xoreaxeaxeax/movfuscator https://github.com/xoreaxeaxeax/movfuscator. Would be definitely cool to see it realized even if it would be incredibly impractical (probably).
- fooker 7mo agoI think any quantization approach should work, not an expert on this.
- deleted 7mo ago[deleted]
- smokel 7mo ago> single bit neural networks are decision trees. I didn't exactly understood what was meant here, so I went out and read a little. There is an interesting paper called "Neural Networks are Decision Trees" [1]. Thing is, this does not imply a nice mapping of neural networks onto decision trees. The trees that correspond to the neural networks are huge. And I get the idea that the paper is stretching the concept of decision trees a bit. Also, I still don't know exactly what you mean, so would you care to elaborate a bit? :) [1] https://arxiv.org/pdf/2210.05189 https://arxiv.org/pdf/2210.05189
- lioeters 7mo agoClosest thing I found was: Single Bit Neural Nets Did Not Work - https://fpga.mit.edu/videos/2023/team04/report.pdf https://fpga.mit.edu/videos/2023/team04/report.pdf > We originally planned to make and train a neural network with single bit activations, weights, and gradients, but unfortunately the neural network did not train very well. We were left with a peculiar looking CPU that we tried adapting to mine bitcoin and run Brainfuck.
- fooker 7mo ago> I still don't know exactly what you mean Straight forward quantization, just to one bit instead of 8 or 16 or 32. Training a one bit neural network from scratch is apparently an unsolved problem though. > The trees that correspond to the neural networks are huge. Yes, if the task is inherently 'fuzzy'. Many neural networks are effectively large decision trees in disguise and those are the ones which have potential with this kind of approach.
- fc417fc802 7mo ago> Training a one bit neural network from scratch is apparently an unsolved problem though. I don't think it's correct to call it unsolved. The established methods are much less efficient than those for "regular" neural nets but they do exist. Also note that the usual approach when going binary is to make the units stochastic. https://en.wikipedia.org/wiki/Boltzmann_machine#Deep_Boltzmann_machine https://en.wikipedia.org/wiki/Boltzmann_machine#Deep_Boltzma...
- zelphirkalt 7mo agoDecision trees are great. My favorite classical machine learning algorithm or group of algorithms, as there are many slight variations of decision trees. I wrote a purely functional (kind of naive) parallelized implementation in GNU Guile: https://codeberg.org/ZelphirKaltstahl/guile-ml/src/commit/25cb709db2d5863c92df05967a71e02f2f9aa00f/decision-tree.scm https://codeberg.org/ZelphirKaltstahl/guile-ml/src/commit/25... Why "naive"? Because there is no such thing as NumPy or data frames in the Guile ecosystem to my knowledge, and the data representation is therefore probably quite inefficient.
- srean 7mo agoWhat benefit does numpy or dataframes bring to decision tree logic over what is available in Guile already ? Honest question. Guile like languages are very well suited for decision trees, because manipulating and operating on trees is it's mother tongue. Only thing that would be a bit more work would be to compile the decision tree into machine code. Then one doesn't have traverse a runtime structure, the former being more efficient. BTW take a look at Lush, you might like it. https://lush.sourceforge.net/ https://lush.sourceforge.net/ https://news.ycombinator.com/item?id=2406325 https://news.ycombinator.com/item?id=2406325 If you are looking for vectors and tensors in Guile, there is this https://wedesoft.github.io/aiscm/ https://wedesoft.github.io/aiscm/
- boccaff 7mo agotree algorithms on sklearn use parallel arrays to represent the tree structure.
- zelphirkalt 7mo agoI think data frames are quite memory efficient and can store non-uniform data types (as can vectors in Guile). Generally, a ton of work has gone into making operations on data frames fast. I don't think a normal vector or multi-dimensional array can easily compete. Data frames are probably also compiled to some quite efficient machine code. Not sure whether Guile's native data structures can match that. Maybe they can. Also I think I did not optimize for memory usage, and my implementation might keep copies of subsets of data points for each branch. I was mostly focused on the algorithm, not that much on data representation. Another point, that is not really efficiency related, is that data frames come with lots of functionality to handle non-numeric data. If I recall correctly, they have functionality like doing one-hot encoding and such things. My implementation simply assumes all you have is numbers. There might also be efficiency left on the table in my implementation, because I use the native number types of Guile, which allow for arbitrarily large integers (which one might not need in many cases) and I might even have used fractions, instead of inexact floats. I guess though, with good, suitable data structures and a bit of reworking the implementation, one could get a production ready thing out of my naive implementation, that is even trivially parallelized and still would have the linear speedup (within some bounds only, probably, because decision trees usually shouldn't be too deep, to avoid overfitting) that my purely functional implementation enables. Thanks for the links!
- kqr 7mo agoExperts' nebulous decision making can often be modelled with simple decision trees and even decision chains (linked lists). Even when the expert thinks their decision making is more complex, a simple decision tree better models the expert's decision than the rules proposed by the experts themselves. I've long dismissed decision trees because they seem so ham-fisted compared to regression and distance-based clustering techniques but decision trees are undoubtedly very effective. See more in chapter seven of the Oxford Handbook of Expertise. It's fascinating!
- ablob 7mo agoI once saw a visualization that basically partitioned decisions on a 2D plane. From that perspective, decision trees might just be a fancy word for kD-Trees partitioning the possibility space and attaching an action to the volumes. Given that assumption, the nebulous decision making could stem from expert's decisions being more nuanced in the granularity of the surface separating 2 distinct actions. It might be a rough technique, but nonetheless it should be able to lead to some pretty good approximations.
- srean 7mo agoYou have this thing a little backwards that it is unintentionally hilarious. Decision trees predate KD trees by a decade. Both use recursive partitioning of function domain a fundamental and an old idea.
- lokimedes 7mo agoWhen I worked at CERN around 2010, Boosted Decision Trees were the most popular classifier, exactly due to the (potential for) explainability along with its power of expression. We had a cultural aversion for neural networks back then, especially if the model was used in physics analysis directly. Times have changed…
- wodenokoto 7mo agoAre boosted decision trees the same as a boosted random forest?
- boccaff 7mo agoshort answer: No. longer answer: Random forests use the average of multiple trees that are trained in a way to reduce the correlation between trees (bagging with modified trees). Boosting trains sequentially, with each classifier working on the resulting residuals so far. I am assuming that you meant boosted decision trees, sometimes gradient boosted decisions trees, as usually one have boosted decision trees. I think xgboost added boosted RF, and you can boost any supervised model, but it is not usual.
- hansvm 7mo agoThe training process differs, but the resulting model only differs in data rather than code -- you evaluate a bunch of trees and add them up. For better or for worse (usually for better), boosted decision trees work harder to optimize the tree structure for a given problem. Random forests rely on enough trees being good enough. Ignoring tree split selection, one technique people sometimes do makes the two techniques more related -- in gradient boosting, once the splits are chosen it's a sparse linear algebra problem to optimize the weights/leaves (iterative if your error is not MSE). That step would unify some part of the training between the two model types.
- srean 7mo ago> Times have changed… This makes me a little concerned -- the use of parameters rich opaque models in Physics. Ptolemaic system achieved a far better fit of planetary motion (over the Copernican system) because his was a universal approximator. Epicyclic system is a form of Fourier analysis and hence can fit any smooth periodic motion. But the epicycles were not the right thing to use to work out the causal mechanics, in spite of being a better fit empirically. In Physics we would want to do more than accurate curve fitting.
- srean 7mo agoA 'secret weapon' that has served me very well for learning classifiers is to first learn a good linear classifier. I am almost hesitant to give this away (kidding). Use the non-thresholded version of that linear classifier output as one additional feature-dimension over which you learn a decision tree. Then wrap this whole thing up as a system of boosted trees (that is, with more short trees added if needed). One of the reasons why it works so well, is that it plays to their strengths: (i) Decision trees have a hard time fitting linear functions (they have to stair-step a lot, therefore need many internal nodes) and (ii) linear functions are terrible where equi-label regions have a recursively partitioned structure. In the decision tree building process the first cut would usually be on the synthetic linear feature added, which would earn it the linear classifier accuracy right away, leaving the DT algorithm to work on the part where the linear classifier is struggling. This idea is not that different from boosting. One could also consider different (random) rotations of the data to form a forest of trees build using steps above, but was usually not necessary. Or rotate the axes so that all are orthogonal to the linear classifier learned. One place were DT struggle is when the features themselves are very (column) sparse, not many places to place the cut.
- ekjhgkejhgk 7mo ago> (ii) linear functions are terrible where equi-label regions have a partitioned structure. Could you explain what "equi-label regions having a partitioned structure" mean?
- srean 7mo agoI missed a word "recursively", that I have edited in my original comment now. Consider connected regions in the domain that have the same label. Much like countries on a political map. The situation where this has a short description in terms of recursive subdivision of space, is what I am calling a partitioned structure. It's really rather tautological.
- ekjhgkejhgk 7mo ago"recursively partitioned" sounds like a fractal to me. Not sure what you really mean.
- huqedato 7mo agoRandom forests on the same site: https://mlu-explain.github.io/random-forest/ https://mlu-explain.github.io/random-forest/
- s2l 7mo agoAnd others: https://mlu-explain.github.io/ https://mlu-explain.github.io/
- deleted 7mo ago[deleted]
- EGreg 7mo agoIsn’t that exactly how humans (and even animals) operate? Human societies look for actual major correlations and establish classifications. Except with scientific-minded humans, we often also want, to know the why behind the correlations. David Hume got involved w that… https://brainly.com/question/50372476 https://brainly.com/question/50372476 Let me ask a provocative question. What, ultimately, is the difference between knowledge and bias?
- srean 7mo agoTo a certain degree yes. https://en.wikipedia.org/wiki/Esagil-kin-apli https://en.wikipedia.org/wiki/Esagil-kin-apli In this Mesopotamian text, diagnostic rules are structured as a nest of if then else rules. So I have been told, not that I have read it myself.
- getpokedagain 7mo agoI worked (professionally) on a product a few years ago based upon decision tree and random forest classifiers. I had no background in the math and had to learn this stuff which has payed dividends as llms and AI have become hyped. This is one of the best explanations I've seen and has me super nostalgic for that project. Gonna try to cook up something personal. It's amazing how people are now using regression models basically all the time and yet no-one uses these things on their own.
- jjcc 7mo agoI worked on a product which was the best ID reader in the world at the time 25 years ago. The OCR engine was based on Decision tree and "Random Forest" (I suspect the name did exist) with only 3 trees. It was very effective as a secret weapon of the competitiveness. I tried to train a NN with a framework called SNNS(Stuttgart Neural Network Simulator) as the 4th tree complement to the existing 3. Today, hand writing OCR is a "hello world" sample in Tensorflow.
- mistrial9 7mo agoin the interest of understanding, is there any code or similar for the approach? does that OCR run anywhere today?
- jjcc 7mo agoThe technology was developed by my predecessor during late 90s when microprocessors was much less powerful, and the resolution of image sensor was low. The relatively high accuracy based on those conditions was a critical factor to use Decision Tree as OCR engine. It's used till 2007 when I left my company. I don't think it would survive afterwards due to quick change in technology. Even the desktop OCR applications at the time didn't use Decision Tree because the CPU was much more powerful. The DT OCR engine was competitive only under special use case.
- getpokedagain 7mo agoThat's awesome and based on my experience I'm not shocked this went well. I'm not sure what the features would be in this but I am assuming they could be specific pixel combinations or other things which would be easily labeled in a few ways. I hope you had fun with it. My previous project was far from that. https://healthverity.com/audience-manager/ https://healthverity.com/audience-manager/ I had a lot of fun, really the last fun project I've had. I hope you had fun as well.
- ayhanfuat 7mo agoI am surprised r2d3's visual intro is not referenced here (https://r2d3.us/visual-intro-to-machine-learning-part-1/ https://r2d3.us/visual-intro-to-machine-learning-part-1/). I think it was the first (if not first, maybe most impactful) example for scroll triggered explainers.
- jebarker 7mo agoThe killer feature of DTs is how fast they can be. I worked very hard on a project to try and replace DT based classifiers with small NNs in a low latency application. NNs could achieve non-trivial gains in classification accuracy but remained two orders of magnitude higher latency at inference time.
- levocardia 7mo agoAlso, decision trees (but not their boosted or bagged variants) are easy (well, easy-ish) to port manually to an edge device that needs to run inference. Small vanilla NNs are as well, but many other popular "classical" ML algorithms are not.
- bawis 7mo ago>> but many other popular "classical" ML algorithms are not Examples ?
- deleted 7mo ago[deleted]
- ssttoo 7mo agoI just wish we’d stop with the “unreasonable” click-bite. Cheapens an otherwise excellent article, like “7 x (number 6 will surprise you)” of yesteryear
- bobek 7mo agoWow. This page is actually a product of LLM [0]. So they can produce useful stuff after all :) [0]: https://news.ycombinator.com/item?id=47195123 https://news.ycombinator.com/item?id=47195123
- jadengeller 7mo agoNo, you misread
- bobek 7mo agoAnd you are absolutely correct. I've seen the DT page thanks to the linked HN submission (actually comment [1]. And incorrectly associated the DT article incorrectly today. Thank you. [1]: https://news.ycombinator.com/item?id=47200131 https://news.ycombinator.com/item?id=47200131
- hkbuilds 7mo agoDecision trees are underrated in the age of deep learning. They're interpretable, fast, and often good enough. I've been using a scoring system for website analysis that's essentially a decision tree under the hood. Does the site have a meta description? Does it load in under 3 seconds? Is it mobile responsive? Each check produces a score, the tree aggregates them. Users understand why they got their score because the logic is transparent. Try explaining why a neural network rated their website 73/100. Decision trees make that trivial.
- srean 7mo agoYou sir are keeping alive an old tradition. 1K BC old. https://en.wikipedia.org/wiki/Esagil-kin-apli#The_Sakikk%C5%AB_(SA.GIG) https://en.wikipedia.org/wiki/Esagil-kin-apli#The_Sakikk%C5%...
- jwilber 7mo agoLooks great, if I remember correctly the author mentioned that the spring color palette + fruit examples are a bit tongue-in-cheek.
- uoaei 7mo agoThey're rough approximations of conditioned distributions qua Bayesian statistics. You may treat them as such. Hence why random forests have such enduring power.
- levocardia 7mo agoOne of the ML textbooks (ESL maybe?) I read described decision trees as (paraphrasing) "really great - they are interpretable, fast to fit, work on lots of different types of data and outcomes, insensitive to scaling and distributional issues, don't have too many tuning parameters...except they just don't work very well." That latter problem can be solved with bagging or boosting, though you are bargaining away many of the other advantages.
- deleted 7mo ago[deleted]