7 ms·
Microsoft Supercharges Bing Search With Programmable Chips
- l31g 12y agohttp://research.microsoft.com/apps/pubs/default.aspx?id=212001 http://research.microsoft.com/apps/pubs/default.aspx?id=2120...
- jacquesm 12y agoI'd rather have better results than faster results. Faster is only important once you have the quality problem worked out, first make it good and then make it fast has been a long time mantra. The reason is that it is usually very expensive to make something really fast because optimizing code is hard and expensive (case in point they use custom hardware here). The upside is that they're doing something innovative but if Bing really wants to steal marketshare from Google they have to improve on their quality, not on their speed. I'd rather see them take 10 seconds and deliver an absolutely perfect answer than 0.001 second and deliver something not on par with Google but 10 times faster. Impressive to see them backing an exotic solution like this though, and if and when they do get it to be better than Google it may pay off. Are there any developments like this underway at Google?
- sharemywin 12y agoI just wish they'd get rid of that annoying image on their homepage or at least let me turn it off.
- jacquesm 12y agoI googled that for you ;): http://www.bing.com/?rb=0 http://www.bing.com/?rb=0
- samirahmed 12y agoThere was no mention of where exactly these would go. I doubt it would be on machines serving response online ... since the bottle neck is often in IO. Being able to index, process and learn data faster can lead to faster iteration and improve the speed of batch or offline jobs which in turn could improve the relevance.
- jacquesm 12y ago> There was no mention of where exactly these would go. From the article: "The system takes search queries coming from Bing and offloads a lot of the work to the FPGAs, which are custom-programmed for the heavy computational work needed to figure out which webpages results should be displayed in which order. " That looks like they're in the interactive path somewhere.
- mtdewcmu 12y ago>>The upside is that they're doing something innovative but if Bing really wants to steal marketshare from Google they have to improve on their quality, not on their speed. Well, there are two ways to increase bottom-line profits. One is to increase revenue by stealing market share. The other is to decrease costs. This would apparently help cut costs by decreasing the number of servers and saving on electricity.
- ewzimm 12y agoTheir quality is sometimes really bad. Check out the first result for their own portal to all their online services: http://www.bing.com/search?q=portal.microsoftonline.com http://www.bing.com/search?q=portal.microsoftonline.com edit: Interestingly, for me it went from an "offer not available" link to urlwebz to the actual sign-in in the hour since I posted that. If nothing else, the search results are changing fast.
- l31g 12y agohttp://www.theregister.co.uk/2014/06/16/microsoft_catapult_fpgas/ http://www.theregister.co.uk/2014/06/16/microsoft_catapult_f...
- valarauca1 12y agoSounds like there is a market niche starting to develop for FPGA in server applications. I'm not saying rush out and make PCIe powered and communicating FPGA's I'm just saying there maybe a market developing for it. Especially with good open source dev tools.
- sliverstorm 12y agoI think you would just about have to start from square one if you want an open source FPGA toolchain. I don't believe such a toolchain exists at all right now. So what I'm saying is, forget about developing a PCI-e FPGA board. If you want an open source toolchain for it, you better start there, because that's going to be 99.9% of the effort.
- th0ma5 12y agoSeems like I read about Google doing this almost 10 years ago. I know that IBM has the Netezza product which also uses FPGAs for accelerating queries.
- chollida1 12y agoThis has been going on in the HFT space for a number of years. FPGA's are used to parse data feeds as the sheer volume of quotes overwhelms most systems. In fact after moving the networking stack into user land and using inifiniband networking gear, its probably the third most common optimization I've seen/heard of for HFT systems. Here's a quick, but surprisingly accurate description of a common HFT setup: http://www.forbes.com/sites/quora/2014/01/07/what-is-the-technology-stack-like-behind-a-high-frequency-trading-platform/ http://www.forbes.com/sites/quora/2014/01/07/what-is-the-tec... Some one had asked about hte number of quotes that need to be parsed. From forbes... > Mr. Hunsader: The new world is now a war between machines. For some perspective, in 1999 at the height of the tech craze, there were about 1,000 quotes per second crossing the tape. Fast forward to 2013 and that number has risen exponentially to 2,000,000 per second. Keep in mind that the "tape" is the slow SIP line that exchanges use to keep prices in sync and show customers that don't use the exchanges direct feeds. ie it aggregates all the quotes from all venues and throws a way alot as they can't be parsed in time or didn't change the top level quote. With 40+ venues at which a HFT fund can get feeds from 2,000,000 second is a fraction of what a cutting edge HFT would have to parse to keep up with all venues. The typical setup is that you'll run strategies across multiple machines so you have the gateway machine that directs the quote to the appropriate machine. The biggest problem is the speed at which the quotes arrive. Unlike a web request, that you can take 300 milliseconds to parse and return, if you don't parse and respond to the quote in under 10-20 micro seconds you've already lost. So the FPGA transition is to make sure there is never a back log of quotes or any pauses in the handling of bursty quotes. This can't be overstated enough. Margins are squeezed so tightly now that your algo will appear to be working fine until a big burst of quotes happen and your machines can't keep up and when the dust settles in 20 seconds, you'll find you lost $5000, which might be your entire day's profit from that one symbol/algo pair.
- MrBuddyCasino 12y agoI'm curious if these kinds of setups are also used in other high-freq scenarios. For instance, I could imagine using techniques like userland network stack and reserving cores exclusively in services like WhatsApp. I think they're currently on a highly customized Erlang stack and are able to handle huge numbers of queries per machine. Any insiders here with a good background story?
- dmmalam 12y agoIs there any breakthrough in programming these things? From a quick glance at the paper it seem like the kernels are still hand written in Verilog. Though there seem to be some significant software infrastructure in integrating the FPGAs into cluster management systems. I think easily and uniformly programming disparate compute devices (CPUs, SIMD, GPUs, FPGAs, ISPs, DSPs, and eventually quantum) is the next BIG problem in programming languages. Several Haskell projects seem promising, but these still tend to be nice DSLs that generate verilog or shaders. On most mobile SoCs the CPU usually takes an increasingly smaller part of the die; there are 10gbe network cards with FPGAs on them; and we've got parrella. The hardware exists, we sorely need the next breakthrough programming environment.
- xamlhacker 12y agoAltera and Xilinx are starting to support compiling OpenCL to FPGAs. Not sure how efficient that is, but that is at least a step towards unified programming environment of various types of devices.
- sliverstorm 12y agoCPUs, GPUs, DSPs, FPGAs... they are so different, it's hard to say they ever could be programmed uniformly.
- 14113 12y agoI think they could - in fact I'm starting a PhD in a similar area soon! It's a matter of providing a high level enough programming language (e.g. Haskell) and a smart enough compiler that can automatically parallelise sections, and with the right middleware/compiler back end it should be possible!
- sliverstorm 12y agoI buy CPU+GPU unification (in fact I highly anticipate it) and I also buy that a DSP could function as a dynamic coprocessor, as they are often programmed in C. But my day job revolves around HDLs, and it is my opinion that a higher level language isn't the answer. Fifty years from now it might be, but state-of-the-art HDL compilers just aren't good enough yet. It's like C compilers a few decades ago, where you had to insert some inline ASM in your code here and there because the compilers were still developing. So I guess what I'm saying is you can't target CPU, GPU, DSP, & FPGA in one compiler until we can master targeting FPGA even just by itself.
- azakai 12y agoActually, I already find bing quite fast. Comparing to google search, bing results tend to load a little faster but to be a little lower in quality. FPGAs may make bing twice as fast as it already is, but I don't feel like it needs to be faster. Although, I guess if its faster they can trade that off for more work done and so better results, perhaps.
- l31g 12y agoIf they make Bing 2X faster, then they can roughly cut the amount of servers they need by half. They measure "speed" by number of requests that can be fulfilled in Z amount of time.
- zackmorris 12y agoI've been ranting about the inadequacies of mainstream processors for almost twenty years. I remember even back in the late 90s, seeing processors that were 3/4 cache memory, with barely any transistors used for logic. It's surely worse than that now, with the vast majority of logic gates on chips just sitting around idle. To put it in perspective, a typical chip today has close to a billion transistors (the Intel Core i7 has 731 million): https://en.wikipedia.org/wiki/Transistor_count https://en.wikipedia.org/wiki/Transistor_count A bare minimum CPU that can do at least one operation per clock cycle probably has between 100,000 (SPARC) and 1 million (the PowerPC 602) transistors and runs at 1 watt. So chips today have 1,000 or 10,000 that number of transistors, but do they run that much faster? No of course not. And we can even take that a step further, because those chips suffered from the same inefficiencies that hinder processors today. A full adder takes 28 (yes, twenty eight) transistors. Could we build an ALU that did one simple operation per clock cycle with 1000 transistors? 10,000? How many of those could we fit on a billion transistor chip? Modern CPUs are so many orders of magnitude slower than they could be with a parallel architecture that I’m amazed data centers even use them. GPUs are sort of going the FPGA route with 512 cores or more, but they are still a couple of orders of magnitude less powerful than they could be. And their proprietary/closed nature will someday relegate them to history, even with OpenCL/CUDA because it frankly sucks to do any real programming when all you have at your disposal is DSP concepts. I really want an open source billion transistor FPGA running at 1 GHz that doesn’t hold my hand with a bunch of proprietary middleware, so that I can program it in a parallel language like Go or MATLAB (Octave). There would be some difficulties with things like interconnect but that’s what things like map reduce are for, to do computation in place rather than transferring data needlessly. Also with diffs or other hash-based algorithms, only portions of data would need to be sent. And it’s time to let go of VHDL/Verilog because it’s one level too low. We really need a language above them that lets us wire up basic logic without fear of the chip burning up. And don’t forget the most important part of all: since the chip is reprogrammable, cores can be multi-purpose, so they store their configuration as code instead of hardwired gates. A few hundred gates can reconfigure themselves on the fly to be ALUs, FPUs, anything really. So instead of wasting vast swaths of the chips for something stupid like cache, it can go to storage for logic layouts. What would I use a chip like this for? Oh I don’t know, AI, physics simulations, formula discovery, protein folding, basically all of the problems that current single threaded architectures can’t touch in a cost-effective manor. The right architecture would bring computing power we don’t expect to see for 50 years to right now. I have a dream of someday being able to run genetic algorithms that take hours to complete in a millisecond, and being able to guide the computer rather than program it directly. That was sort of the promise with quantum computing but I think FPGAs are more feasible.
- samfisher83 12y agoThis is cool and all, but instead of spending money on this project why not try to improve their search engine or just not spend this money since bing loses so much money. I don't mind waiting half a second extra for my search results. It seems more like their thinking is we got a lot of engineers we are paying a bunch of money to lets do some project.
- Scaevolus 12y agoThis move saves money, since the servers process queries more efficiently-- “Right off the bat we can chop the number of servers that we use in half,” Burger says. "Just improving their search engine" isn't some simple task. Google has a head-start measured in thousands of man-years. Closing that gap takes a great number of smart people a great amount of time. I assume that Google is improving slower than Bing, since their algorithms and systems are more mature and closer to the "asymptotically ideal search engine".
- l31g 12y agoI think this project was funded not only because it helps Bing, but it also accelerates pretty much any large-scale data center application. So in theory, this system could help applications that do not exist yet.