3 ms·
Large scale online ML is able to learn from terafeature datasets. The accuracy is often equal or slightly worse than in-memory techniques. Vowpal Wabbit is mad
by blauwbilgorgel 12y ago
Large scale online ML is able to learn from terafeature datasets. The accuracy is often equal or slightly worse than in-memory techniques.
Vowpal Wabbit is made to scale. You set a fixed bitsize and words and n-grams are hashed. So if you expect 2^32 unique words you set the bitsize to around 32. More data is usually better. Linear speed-ups by adding parallel machines.
I too think that HN's user patterns could be different than other web estates. With large scale spam filters like at Yahoo mail, I believe they employ two (or more) models: One fitted on your inbox, and one fitted on everyone's inbox. That ensemble model should be able to specialize on your behavior, yet still be able to detect general spam that it has already seen in other boxes.
>Is this the basic assumption that close documents also have close information entropy?
Yes, that is the gist of it. The better the compressor, the closer NCD will approximate NID.
A simple principle: Compressors do a better job on repeating data patterns. If two files or documents share data patterns, then adding these together and compressing, will result in a smaller filesize, than if you concatenate and compress two files that don't share any data patterns.
NCD works on text, but not as good as other algo's for NLP. Sometimes PAQ (very slow, but efficient compressor) is used on genome data, or bzip on binary files like virusses. I don't think it will be practical here, since for a comparison every other file would need to be concatenated and compressed. If not using a fast compressor like Snappy or Gzip this would take a while, over a simple cosine distance between tokens.
- 2510c39011c5 12y ago>The accuracy is often equal or slightly worse than in-memory techniques. Slightly worse? because it has to go to different nodes to do the match, and hence incurs a slight delay which then incurs a bit inconsistency of the sample base? > I too think that HN's user patterns could be different than other web estates. That's a better generalization... perhaps some more mining could be done based on this, such as the social clustering ; and also perhaps construct different language/wording models for the different interest/area groups (by posts) -- after all, there is no language model that fits them all... And as in terms of spam detection, I assume if an ID, or originating IP, replies to almost every post, then it's unlikely the quality of his post would be high...Well, this returns to that fundamental question, how do you define "spam" in the space of HN? > With large scale spam filters like at Yahoo mail, I believe they employ two (or more) models: One fitted on your inbox, and one fitted on everyone's inbox. That ensemble model should be able to specialize on your behavior, yet still be able to detect general spam that it has already seen in other boxes. Interesting...It would be interesting to see how often the two models produce inconsistent results...Then if you biased towards one, then what's the use of the other one? Or perhaps they devised some strategy to combine the two models... >NCD works on text, but not as good as other algo's for NLP. Sometimes PAQ (very slow, but efficient compressor) is used on genome data, or bzip on binary files like virusses. I don't think it will be practical here, since for a comparison every other file would need to be concatenated and compressed... perhaps NCD and NID are theoretically charming...but perhaps they are only good for pure coding without much context to depend on...there are just tons of information outside the analysis target, especially a natural language, that would be hard to take into the calculation of the entropy model...By the way, in terms of the similarity analysis based on compressed binary, there was a paper (peHash) published a few years ago that was interesting.... https://www.usenix.org/legacy/event/leet09/tech/full_papers/wicherski/wicherski.pdf https://www.usenix.org/legacy/event/leet09/tech/full_papers/... Another issue with the entropy is, it's easy to inject some tokens to manipulate the frequency...In this case, some preprocessing should be in place...
- adw 12y agoRe Yahoo! - http://arxiv.org/pdf/0902.2206.pdf http://arxiv.org/pdf/0902.2206.pdf (which uses VW). VW is the (open-source) state of the art. If you want a linear learner, use it.