4 ms·
This sounds interesting! Can you give any reference to your publication or URL?
by HerrMonnezza 7y ago
This sounds interesting! Can you give any reference to your publication or URL?
- graycat 7y agoOkay, I'll give in: Nearly any good academic research library should have and be able to find the paper: "A Real-Time System-Adapted Anomaly Detector", Information Sciences, volume 115, pages 221-259. The core math is some measure preserving transformations of the data -- get to move data around without changing probability distribution -- for any distribution. The usual place see measure preserving is in ergodic theory. Recall, my work is distribution free which means need make no assumptions at all about probability distributions. In the context of computer and network performance data, distribution-free is nearly essential. After moving the data around, get to do some counting that leads to calculating false alarm rate. So the paper is some applied probability, that is, does some probability calculations in the service of some, call it, mathematical statistics. The simplest view is that it's all just nearest neighbors detection -- if get a point too far from the old points, then raise an alarm. Since I wrote the paper, that idea has become common. But apparently what has not become common is how to adjust and calculate false alarm rate, and in practice that is really important. Indeed, for any detection, my work can report the lowest false alarm rate for which the (real time) observation would still be a detection -- so get to have not just a detection but, intuitively, a <i>seriousness</i> measure. "Joe, this is a detection, and it will remain a detection with a false alarm rate of one in a billion. Put down your pizza, and let's investigate that." So, right, could display a strip chart with such data. For calculating false alarm rate there is a simple approach that gives the right answer, but the simple approach makes an impossible assumption of independence. Well, the right answer is still true; it's just that an actual proof of false alarm rate, where don't assume independence, is a bit tricky mathematically. This math is a close cousin of not independence but a weaker assumption exchangeability where often can get the same results as if had independence. My math may be something of a rationalization of part of resampling as pursued by P. Diaconis and B. Efron at Stanford. In an intuitive sense, we almost have independence. Hmm -- approximate independence is not a big field. There is a paper by M. Talagrand, A New Look at Independence that may be able to shed some light on my paper, find another way to do the calculations and show that the same calculations hold approximately, maybe with some bounds, in other situations. Talagrand is one heck of a good mathematician, a student of Choquet, a member of Bourbaki. It would be nice to have the most powerful statistical hypothesis tests as from the classic result of Neyman-Pearson; alas, in the context we don't have enough data to use Neyman-Pearson. But I did come up with a weaker but still useful sense in which my techniques do yield the most powerful test. In practice, the test will likely be seen as nicely powerful, that is, even when false alarm rate has been selected to something small, the detection rate is still, likely in practice, about the highest can hope for for that selected false alarm rate. There is an issue of how to make the computations fast -- so, need something out of computational geometry. I worked up a technique that should be reasonably fast. The solid state drives of today should do wonders for this technique. The computational geometry and the most powerful test derivation are not in the paper. There's more that can be done. Might also use the work for other anomaly detection problems. One nice point is, although might be using lots of variables, don't encounter "over fitting". What we were doing with expert systems was mostly just thresholds on one variable at a time. Well, in some cases, we know so much about the variable that that is okay. But there can be thousands of variables, collected by HP, Microsoft, etc., and we can't know just from personal knowledge or experience just what thresholds make good sense. And we have nothing solid on false alarm rate or how to adjust it -- yes, we could get an approximation by going back through the data. And, then we have still less on detection rate. If want a server farm that is reliable and secure, about the best you can hope for, then maybe my work would be one of the important tools and means. This stuff about "multi-dimensional" is serious and important. E.g., suppose we have three threshold detectors. Then whether we wanted to or not, we just assumed that the 3D region of normal, healthy, not an anomaly or sick, performance is a 3D box. The box does not promise to fit reality well. Make the box too small, and get too many false alarms. Make the box bigger and can reduce the false alarm rate, but since we still have a box that does not fit reality well we are stuck with detection rate too low and a "poor" detector. Well, can think a little and intuitively convince yourself that the region of normal (healthy) performance could be fractal, e.g., the Mandelbrot set, and still the math should work. To me, for monitoring server farms and networks, my little paper totally blew out of the water with the doors blown off nearly everything we were doing with expert systems. Let us all know what utility you find in the paper!
- maga_2020 7y agoOnly link to your paper that I was able to find is https://dblp.org/rec/journals/isci/Waite99 https://dblp.org/rec/journals/isci/Waite99 but there are no PDFs that I could download without payin 37.95 $USD at ( https://www.sciencedirect.com/science/article/pii/S0020025598100646 https://www.sciencedirect.com/science/article/pii/S002002559... ) I used K-means clustering in various domains (access control, fraud control, network data plane quality controls), so wanted to check your work. One problematic area for K-means, at least from my point of view, is that it is not really possible to easily trace what features/variables are mostly responsible for the anomaly.
- graycat 7y agoOne way to get a copy of the paper would be to make a photocopy at a library or pay a service to do that. Explanation is a problem, also for my work. Broadly, for a first cut, pick a target system want to monitor. Then use variables that are for the target, the whole target, nothing but the target or some such. Then if get a detection, start diagnosis, looking for cause, with the target. For two targets that interact, have detectors for each and then one more for variables from both of them. Broadly, but crudely, have a hierarchy of detectors and then when start getting detections chase down the tree of the hierarchy, like the patient is sick so, look at the major measures and the major organs, etc. When find a suspicious organ, zoom in, drill down, etc. Finally discover the problem is a USB cable that is loose from vibrations from a cooling fan or some such?? In this tree, might want to have somewhat coordinated false alarm rates. One motivation for my work was a cluster for transactions. One day one computer in the cluster got a little sick in the head and was throwing away all its incoming transactions. So, to the load leveling it looked not very busy and was getting nearly all the transactions and, thus, essentially ruined the work of the whole cluster. And that's not the only case of a cluster getting totally sick due to just one computer in the cluster getting sick. So, I was hoping that getting data from all the computers in the cluster would raise an alarm from the cluster and data from each of the computers in the cluster would essentially do the diagnosis of which computer in the cluster was sick. Likely more could be done. E.g., might want to do some data scaling. I don't yet have a big server farm, and I don't know anyone who does who cares enough about anomaly detection to use my work. Expert systems? At one time, yup. My work? Nope. If my startup does well, then, sure, I'll deploy my detector, lots of instances. My startup isn't anomaly detection. At one time, I wrote lots of VCs about my anomaly detection work, that it was from IBM's Watson lab, that it was published in Information Sciences, that I had work on fast algorithms, that I had good results on real data, etc. I got back just nothing. I guessed that the target customers would be high end shops so that I would have to come in with a highly polished product, with lots of good data handling utility tools, some first class hand holding, etc., all of which would be expensive before the first sale. So, I gave up on the idea that I could do a startup from my anomaly work. There might be a way; if I ran a large, really serious server farm, I have at least a little project pursuing my work and/or related work. But, apparently mostly that's not the way server farm management works. So, I picked another problem for a startup, one where I could get a good solution and bring it to good revenue with just my own efforts as sole, solo founder. I did that and am about to go for an alpha test. Somehow anomaly detection just doesn't get people very interested. Once I gave a talk, and some in the audience mentioned that could use my work for fraud detection in, say, credit cards. Well, maybe, but that audience had nothing to do with credit cards. At one time there were some Soviets in England watching some government offices and keeping track of lights, comings, goings, etc. They guessed that if a war was on the way, this data would show anomalies and early warning. Yup, the Soviets saw the broad issue. I don't know if they used anything like my work or not. So, maybe the secret is not just some good work in applied probability with some useful results but publicity, hype, fads, group think, a movement, etc., even if what are selling is just total nonsense. My work is on the shelves of the research libraries. I did my part. If people want anomaly detection, there's some good work there. I've told the Sand Hill Road people, the Hacker News audience, several organizations, e.g., the main NASDAQ server farm at Trumbull, CT, etc. My experience is that people will start to act, react, if they have a problem that, like a tight shoe, really hurts, and then start to work on the problem. If some early work has the shoe not pinch so much, then they will f'get about that work and that problem and concentrate on something else, even if there's lots of money to be made in better solutions to the original problem. There is a time lag problem: A lot of the math is on the shelves of the libraries, but only much more recently has suitable computing been available. Only a tiny fraction of people in computing now, where the computing has the potential, know that old math that's been sitting there waiting on the computing. E.g., when I first did my work on anomaly detection, some people said that a lot of computing would be needed. Well, yup. The response was so obvious I didn't have to make it: The needed computing was coming. Well, now, say, with solid state disks, big server farms, lots of system monitoring data gathering, etc., now the time is right for actual usage. But, nope, interest is really low! The startup I'm pursuing now is much more promising. I don't have to get a few bureaucratic, cautious, conservative, CIOs all fired up. Instead I just have to please a lot of Internet users a little bit each, and that should be much easier. The CIOs are running big server farms, and they are working. So, they don't have a shoe that pinches very much. So, they are not motivated to do anything new or different. Okay, lesson learned.