4 ms·
One of the issues that I see with this approach is that for most applications the scaling issue is data scaling rather than CPU. Here's why I think you will ru
by moonpolysoft 18y ago
One of the issues that I see with this approach is that for most applications the scaling issue is data scaling rather than CPU. Here's why I think you will run into scaling issues:
Lets assume you have a gigabit line out of your colo. Lets also assume that your average game client is on a cable modem with a 1megabit connection. That gives you capabilities to stream work units to 1024 clients simultaneously as an upper limit. In keeping with an average 1 megabit client, it will take 3 minutes to stream down a 20 megabyte work unit, maxing their connection. 1024 concurrent clients * 20 megabytes = 20 gigabytes. So you're looking at 3 minutes of overhead transfer out, and likely another 3 minutes or so of overhead transfer of results back from the client. So that's approximately 6 minutes per gigabyte just in transmission overhead. And it gets worse as clients are added to the system, since you would need scale out your datacenter just to handle coordinating all of the clients. Which begs the question: why aren't all those servers just doing the dang work already?
That kind of overhead limits this technique's usefulness only to applications which have relatively high computational complexity and relatively small amounts of data. And those applications do exist, however they're pretty far from the day to day needs of most companies. Sun found this out the hard way with their Sun Grid project, which last time I checked was a failure. Sorry, I really wish you the best of luck.
- westside1506 18y agoYep, we quite aware of all of these issues. My background is in HPC and my previous company was a successful exit to a major oilfield services firm. My software and its descendants are used on nearly 100,000 CPUs. We are definitely focused on the applications that have extremely high compute/io ratios. In general, these boil down to either high compute problems with no real data or problems where the data can be shared between multiple work units. An example of the latter is stock market analysis - the nodes download stock data for a few stocks and stay busy running different combinations for a very long time.
- moonpolysoft 18y agoSo that brings up a good issue, namely the proprietary nature of the code you're distributing. For instance, stock market analysis firms zealously guard their algorithms. Let's say I'm a potential customer: How will you protect my bytecode from being stolen by competitors when it has to be run on unmanaged machines in the wild?
- westside1506 18y agoGood question. There really isn't anything that stops someone from reading and trying to interpret the byte code. There is some protection in the fact that you never know what type of work unit is happening at a given time. The algorithms are typically each snippets of code instead of full applications, so someone would need to piece together quite a lot of information. Of course, if someone is particularly concerned, they can run a jar obfuscator.