4 ms·
Yes, Facebook for cat owners is absurd; it was supposed to be a hoot and a parody of a silly idea. Your description of trying over and over is common but shoul
by HilbertSpace 16y ago
Yes, Facebook for cat owners is absurd; it was supposed to be a hoot and a parody of a silly idea.
Your description of trying over and over is common but should not be the only way.
Broadly another way is really the old entrepreneurship paradigm: Find a nasty problem that many people have and that has no good solution and where a large fraction of the people would be eager to have a good solution.
Look for a good solution. If actually find one, then go forward, implement the solution, and offer it to the people who wanted it.
Then there's the old saw that "Whatever you are working on, at least 20 other people are working on the same thing with the same ideas.". Well, this claim is silly: Anyone who has done much peer-reviewed original research can see clearly that, a large fraction of such research, especially once the work has been reviewed and accepted for publication, is unique with no one else doing anything very close.
E.g., once I was working in a research project applying artificial intelligence (AI) to system and network monitoring and management. Well, one of the first needs is doing well detecting problems in real time. Such detection is for either (A) old problems seen before or (B) new problems, zero day, never seen before. Assume that whenever we see a B problem we implement corrections and, thus, convert it to an A problem and solve it so that we never see it again. So, we are left with detecting B problems.
I thought that the AI techniques we were using were junk.
Indeed, clearly, as we monitor, there are two ways to be wrong, (1) a false alarm where we say that the system is sick when it is healthy and (2) a missed detection of a real problem where we say that the system is healthy when it is sick. So, clearly we are now necessarily close to statistical hypothesis testing with Type I error (false alarms) and Type II error (missed detections).
Then we are necessarily close to the classic Neyman-Pearson result on the way, for each rate of false alarms, to get the lowest possible rate of missed detections.
Well the whole field of zero day monitoring had not yet gotten even this far, which is really just a junior level course in mathematical statistics.
So, we want to do a hypothesis test. Okay but for the large literature of such tests, we have two issues: First, from server farms and networks, we can collect data on many variables, not just one. So, we want to be multi-dimensional. Second, especially being multi-dimensional, we have no hope of knowing the probability distribution of, say, a healthy system. So, we want our work to be distribution-free.
Well, can look through the literature, especially, say, E. Lehmann, and find nothing on multidimensional, distribution-free tests.
So, one Saturday I put my feet up and created a large, new family of such tests. I wrote out theorems and proofs to justify what I was doing. I wrote some corresponding software. Then I had some data from a complicated server farm, washed it through my software, and saw that I was getting what my theorems said. Then I did a long series of Monte Carlo tests with some very complicated data; my detection techniques worked just as intended.
So, get to select false alarm rate in advance and get that rate exactly. There is not enough data to get all of Neyman-Pearson, but in a powerful sense, asymptotically, for whatever false alarm rate is selected the techniques give the lowest possible rate of missed detections.
So, my work is progress in zero day monitoring of complex systems and networks. I published the work.
Got to tell you, history since I did that work and published it shows clearly that I was the only person in the world doing anything like what I did.
Is there a business in this, say, to be sold to HP, Microsoft, IBM, EMC, Cisco, or some such? Maybe, but my current project is easier to do and more valuable.
The lesson is, broadly, if really have something new and advanced, the chances that someone else is doing the same thing are small.
Of course, what I'm really talking about is applied math as the crucial, core 'secret sauce' to get a much better solution to a nasty problem and not just routine software for just some intuitive idea.
Then this is a broad area of opportunities: In column A list nasty problems people would like to have solved. In column B list some applied math techniques, old or new, that take in data and spit out results. Then find a good pair, a problem from column A and a solution from column B where the solution is much better than anything else for the problem and likely difficult to duplicate or equal. Now write the corresponding software and proceed with little risk of anyone else doing the same thing. The key is making the project one in applied math, not just computing or computer science.
Where'd I get this paradigm? Sure: The US DoD has been doing such things with great success all the way back to WWII, and I started my career around DC in DoD work.
Can this paradigm work in Web 2.0 and consumer-facing Internet? I do believe so!