5 ms·
I might have missed something, but if they only have 2 features, simply plotting the data out will make any trend very clear.
by nacc 10y ago
I might have missed something, but if they only have 2 features, simply plotting the data out will make any trend very clear.
- kehrlann 10y agoAgreed. Also, they don't discuss how they chose their features... It seems the problem was already solved before even applying ML to it. Maybe the example is too naive ?
- sergey-obukhov 10y agoHtml size and html tags count was a natural choice. If it didn't work out the next step would be to try something else. You're right that it's a very naive example and that in a way it was solved before any ml was applied. The surprising part for me was that both features turned out to be interchangeable i.e. any of them could be used. I would expect html tags count to be much more accurate / reliable, etc. Another interesting part for me was the threshold. It's somewhat clear that it should be somewhere between 20 and 1 sec probably but where exactly?
- stevetrewick 10y agoWould the person or persons down voting all the author of TFA's comments mind explaining why? Seems a bit off.
- blt 10y agoThe slowest part of any parser is the lexer - guessing whatever processing they do on the parse tree structure is insignificant by comparison.
- sergey-obukhov 10y agoRe-pasting my comment here where it belongs as a reply: It could be a great idea - to plot the data on 2 axes (x - html length or html size, y - processing time, if I understand this correctly). It's simple and elegant. I'll try that. It could be though that the chart will get messy with all this data points. One of the reasons I like the percentiles approach is that it makes it clear that there is a trade-off between message processing time and the number / percentage of messages we can process.
- minimaxir 10y ago> It's simple and elegant. The OP likely made that comment because plotting the data is often done before the fancy machine learning as a part of the exploratory data analysis. (especially with a low number of variables!)
- IanCal 10y agoSet a low alpha to the points on the chart of there are loads.
- gmfawcett 10y agoIanCal's comment about setting the alpha is a good one. Another easy option is to make a heatmap of 2D bin counts. Since you're using R already, you can use ggplot for this: http://docs.ggplot2.org/current/geom_bin2d.html http://docs.ggplot2.org/current/geom_bin2d.html
- prashnts 10y agoFor higher dimensions, a RadViz plot with lower alpha is really useful too. [0] http://docs.orange.biolab.si/2/widgets/rst/visualize/radviz.html#description http://docs.orange.biolab.si/2/widgets/rst/visualize/radviz....
- deleted 10y ago[deleted]
- Terr_ 10y agoI'm reminded of Anscombe's Quartet as an example of how a little visualization can reveal things that might be hidden behind summary statistics. https://en.wikipedia.org/wiki/Anscombe's_quartet https://en.wikipedia.org/wiki/Anscombe's_quartet