38 ms·
Developer Preview – EC2 Instances with Programmable Hardware
- ranman 10y agoIf you don't click through to read about this: you can write an FPGA image in verilog/VHDL and upload it... and then run it. To me that seems like magic. HDK here: https://github.com/aws/aws-fpga https://github.com/aws/aws-fpga (I work for AWS)
- adamdecaf 10y agoIs that repo going to be made public? It looks to be private right now.
- pjmlp 10y ago+1 for VHDL. :)
- grandalf 10y agoThis is very awesome. Could you add some more thoughts on the tooling and the development workflow? Is it possible to target the Xilinx hardware using only open source (or AWS proprietary) tools? Or is Vivado still required for advanced stuff?
- RandomOpinion 10y agoThe press release says: "This AMI includes a set of developer tools that you can use in the AWS Cloud at no charge. You write your FPGA code using VHDL or Verilog and then compile, simulate, and verify it using tools from the Xilinx Vivado Design Suite (you can also use third-party simulators, higher-level language compilers, graphical programming tools, and FPGA IP libraries)." So basically, buying a copy of Vivado is the minimum. There aren't any open source tools that directly output Xilinx FPGA bitstreams that I know of.
- aseipp 10y agoIt looks like the FPGA Developer AMI includes Vivado and a license explicitly for use on these platforms (look at the PuTTY screenshot in the blog post; it has a customized MOTD). You just need to set up the license server that Vivado will use and point it to the right license. So I guess the real question is: what exactly is granted by the Vivado license on these AMIs? Do we get things like SDSoC, SDAccel, etc, and all the libraries? [1] The blog seems to imply you can program these things with OpenCL too (AKA SDAccel), so I'm guessing that these features are all enabled, but details about the included Vivado license in the AMI would be nice. [1]: https://www.xilinx.com/products/design-tools/vivado.html#buy https://www.xilinx.com/products/design-tools/vivado.html#buy
- aseipp 10y agoVivado is required for all advanced features and programming Xilinx chips in general; like the sibling post said, there is no open FPGA toolchain implementation for Xilinx devices, especially for extremely high end ones like the ones being offered on the F1 (I expect they'd run at like, several thousand USD per device, on top of a several thousand dollar Vivado license for all the features). It doesn't look like there's much AWS proprietary stuff here, though we'd have to wait for the SDK to be opened properly to be sure. I imagine it's mostly just making all of the stuff prepackaged and easily consumable for usage, and maybe some extra IP Cores or something for common stuff, and lots of examples. If you're already using Vivado I imagine using the F1/Cloud won't introduce any kind of major changes to what you expect.
- cardigan 10y agoThis is really cool. Do you think it will be possible to run MongoDB on an FPGA anytime soon?
- Something1234 10y agoI really hope that is sarcasm.
- gnofehufo 10y agoI'm currently working on this. Speedup around 2x for most operations. Not kidding, quite a few startups are currently trying to optimize typical data operations with special algorithms.
- wmf 10y agoAren't there other software databases that are already more than 2x faster than Mongo and don't lose data?
- gnofehufo 10y agoMaybe, I'm not talking about Mongo specifically. You can find 'equivalents' to CPU data structures for FPGAs and speed up operations on/with them while still saving power. There's lots of trouble with how buffers are used and memory is accessed. So it's not a trivial task, but IF you can optimize generic data structures and replace the existing ones you basically have 2x the speed or half the energy consumption for any DB.
- brian-armstrong 10y agoBut what's the developer time/cost for that?
- gnofehufo 10y agoTotally depends on your use case. From the blog post: > From here, I would be able to test my design, package it up as an Amazon FPGA Image (AFI), and then use it for my own applications or list it in AWS Marketplace. As a user of those Marketplace images, you just look at the hourly fees. Your team needs to set this up, of course, and replace the old stuff, e.g. MongoDB with a new, sped-up FPGA-MongoDB. (And you'd need to fix some new bugs.) If time is super-critical to you, e.g. if you're working with analytics: do you really need to speed up your processing pipeline? E.g. processing stuff not once but twice per day? If yes, then you'd better off having people on your team who understand all this and are able to fix and implement stuff themselves. Second scenario would be quite a bit more expensive, but still, FPGAs aren't rocket science and there's no way around them in the future.
- zyngaro 10y agoExactly what I thought. This is amazing. FPGA is commonly used in embedded systems to perform application specific tasks and now application developers have access to this power too. I guess many machine learning application might take profit of that power instead of using comparatively very expensive graphics hardware.
- ranman 10y agoIf you guys are curious about these announcements I'll be recapping them and going into more detail on twitch.tv/aws at 12:30 pacific
- brian-armstrong 10y agoHuh? Isn't Twitch just for gaming content?
- frikk 10y agoNope. Twitch is excellent for all kinds of live content.
- brian-armstrong 10y agohttps://www.twitch.tv/p/rules-of-conduct https://www.twitch.tv/p/rules-of-conduct "All content that is neither gaming-related nor permitted under the rules for Twitch Creative Conduct is prohibited from broadcast."
- deleted 10y ago[deleted]
- Crosseye_Jack 10y agoI've seen many people programming on Twitch https://www.twitch.tv/directory/game/Creative/programming https://www.twitch.tv/directory/game/Creative/programming While its mainly Game dev or game dev related its not limited to game dev stuff. From their FAQ https://help.twitch.tv/customer/portal/articles/2176641 https://help.twitch.tv/customer/portal/articles/2176641 Examples of what you can broadcast on Twitch Creative: ... Programming and coding Software and game development Web development EDIT: It seems that re:invent is being streamed on twitch anyway.
- brian-armstrong 10y agoThis is a product announcement, though
- eliben 10y agoI'm not sure what you mean by the "magic" part here, can you please clarify? [background: many years of writing VHDL specifically for FPGAs, using various dev boards and custom boards]
- ebrewste 10y agoThe magic part is the thing we have gotten used to with the cloud -- virtual hardware you never see and rent by the minute. Imagine having an FPGA idea and not needing to make board, pay for a dev board, or even find a dev board in your lab... Like your idea and need more? Spin up 100 more right now...
- orbifold 10y agoI'm very curious if/how you have managed to make the developer experience sane and enjoyable. I've experience with a FPGA cluster of ~800 FPGAs and it definitely does not get used to its full potential because of the tooling around it.
- cottonseed 10y agoThis is so awesome, I can't even. I wrote arachne-pnr [0] to learn about FPGAs to get ready for this day. Just signed up, can't way to play with these! I hope the growing popularity of FPGAs for general-purpose computing will help push the vendors to open up bitstreams and invest in open-source design tools. [0] https://github.com/cseed/arachne-pnr https://github.com/cseed/arachne-pnr
- makapuf 10y agoWow Clifford is that you ? I hope this, exciting as it may be, won't make you leave open fpga efforts for the dark side (saw your talk last Fosdem, was very exciting)
- aseipp 10y agoCotton is the author of arachne-pnr. Clifford is the author of Yosys and IceStorm, which are all separate projects. Not the same person. FWIW, Clifford has recently started reversing the bits of the modern Xilinx FPGA series. So, stay tuned for a Xilinx IceStorm-equivalent sometime down the road (a few years, probably...)
- cottonseed 10y agoNo, Clifford is cliffordvienna on HN. He wrote Yosys (and amazing piece of software) and did the iCE40 reverse engineering (amazing work). I wrote the place and router, arachne-pnr.
- makapuf 10y agoAnd kudos for that.
- noselasd 10y agoSo it's tied to the PCIe bus - how do you interact with your FPGA once you programmed it - are there general drivers you can use, or do you also have to create a linux driver to talk to your FPGA ?
- scott_wilson46 10y agoXilinx provide software drivers and IP for PCIe DMA and memory mapped interfaces. These are fairly easy to integrate (probably not the best for latency though - I've developed my own but I require a specific use case - low latency but don't care about bandwidth).
- cma 10y agoHow do FPGAs compare with GPUs for the inference stage of Deep Learning algorithms? Can they accelerate it a lot?
- deleted 10y ago[deleted]
- nl 10y agoNo, but they do use less power: To the best of our knowledge, state-of-the-art performance for forward propagation of CNNs on FPGAs was achieved by a team at Microsoft. Ovtcharov et al. have reported a throughput of 134 images/second on the ImageNet 1K dataset [28], which amounts to roughly 3x the throughput of the next closest competitor, while operating at 25 W on a Stratix V D5 [30]. This performance is projected to increase by using top-of-the-line FPGAs, with an estimated through- put of roughly 233 images/second while consuming roughly the same power on an Arria 10 GX1150. This is com- pared to high-performing GPU implementations (Caffe + cuDNN), which achieve 500-824 images/second, while con- suming 235 W. Interestingly, this was achieved using Micros oft- designed FPGA boards and servers, an experimental project which integrates FPGAs into datacenter applications. https://arxiv.org/pdf/1602.04283v1.pdf https://arxiv.org/pdf/1602.04283v1.pdf
- shaklee3 10y agoThat's hard to compare. Typically FPGAs are doing fixed-point math, so they can do more operations with less power. GPUs have traditionally done floating point. However, with the new Pascal architecture, certain cards (P4/P40) support 8-bit integer dot products, which give a massive boost in performance/W. It's still fairly high at 250W, but that's for an entire card with 24GB of memory. You'd have to compare that to an FPGA with that much memory on a PCIe card if you're doing apples to apples. Something like this is appropriate for comparison: http://www.nallatech.com/store/fpga-accelerated-computing/pcie-accelerator-cards/nallatech-385a-arria10-1150-fpga/ http://www.nallatech.com/store/fpga-accelerated-computing/pc...
- ap22213 10y agothat repository is 404?
- grandalf 10y agoHmm this still isn't public. Any ETA?
- jakozaur 10y agoSo know anyone can run their High Frequency Trading business on their side :-P. So much easier than buying hardware. Also deep learning works sometimes similarly. It's easier to play with on AWS with their hourly billing than buying hardware for many use cases.
- zitterbewegung 10y agoThe latencies from AWS servers to the exchanges probably would make HFT applications unfeasible.
- spullara 10y agoNot when you use Amazon's new regions, us-fin-1, that is within the exchange's datacenter. /s?
- grandalf 10y agoThat would actually be tremendously disruptive! Superb idea.
- brilliantcode 10y agofor a moment I got super excited and thought us-fin-1 was real then saw that trailing slash indicating sarcasm. maybe we'll see High Frequency Trading For The Masses sort of situation in the future that wipes out profits for existing guys although it seems unlikely seeing how arbitrage opportunities are all automated by large capital holders.
- theocean154 10y agoYeah you need to be in the colo. Also these aren't on the network card, the cpu introduces too much latency
- wyldfire 10y ago> Today we are launching a developer preview of the new F1 instance. In addition to building applications and services for your own use, you will be able to package them up for sale and reuse in AWS Marketplace. Wow. An app store for FPGA IPs and the infrastructure to enable anyone to use it. That's really cool.
- perlgeek 10y agoI guess this will be a game changer for FPGA-mineable digital currencies. Maybe not for Bitcoin, because people have invested heavily into dedicated mining hardware, but I'm interested to see what it'll do for the smaller altcoins.
- mi100hael 10y ago> Maybe not for Bitcoin, because people have invested heavily into dedicated mining hardware The thing is, it seems like people always invest heavily into dedicated hardware when using FPGAs. I'll be interested to see what people actually end up using this service for.
- wmf 10y agoFor any cryptocurrency that's profitably mineable on AWS the difficulty immediately increases to the point that it's no longer profitable.
- deleted 10y ago[deleted]
- klagermkii 10y agoWould love to know what that gets priced at per hour, as well as if they plan to have smaller FPGAs available while developing.
- prashnts 10y agoFor my institute this is going to be _really_ useful for Genomics data processing because we can't justify buying expensive hardware for undergrad research. Using a FPGA hardware over cloud sounds almost magical!
- brian-armstrong 10y agoYou can't justify buying it but you can justify renting it? Has your department heard of amortization?
- prashnts 10y agoRenting for a short period of time vs. buying the hardware are very different IMO.
- op00to 10y agoMost research finance departments are absolutely horrified at OpEx because any strange non-capital expenditure makes them look less efficient than the next research institute. This comes in handy when two labs are up for a grant, and they are equally qualified. The more efficient institute gets the grant. You can imagine asking for the lab credit card for EC2 time is not met with enthusiasm.
- CamperBob2 10y agoThat's interesting. So, buying a lot of expensive, soon-to-be-obsolete hardware makes your lab more attractive?
- op00to 10y agoExactly. I haven't worked at my old lab for more than 5 years, but they are still advertising on their web page the systems I built when employed there. Woo! 2010-era blade servers!
- r00fus 10y agoWouldn't bandwidth/transfer costs basically nullify the computing gains? I know someone who used to be in genomics and cloud-anything was priced-out due to transfer costs.
- irq-1 10y agoOVH is testing Altera chips - ALTERA Arria 10 GX 1150 FPGA Chip https://www.runabove.com/FPGAaaS.xml https://www.runabove.com/FPGAaaS.xml
- _nrvs 10y ago_NOW_ things are getting really interesting!
- majke 10y agoBitcoin mining. WPA2 brute forcing. Maybe someone will finally find the triple-des password used at adobe for password hashing. The possibilities are endless :)
- deleted 10y ago[deleted]
- problems 10y agoMining is unlikely, with bitcoin at least. Bitcoin passed the FPGA stage and moved onto ASICs many years ago. There are some alt coins that are currently best mined on GPUs though and this may change that or put their claims to a real test.
- rphlx 10y agoThe boards used for this preview do not have enough memory bandwidth to pose even a modest threat to the latest batch of memory-hard GPU PoW algos.
- brendangregg 10y agoVery interesting. I'd still like to see the JVM pick up the FPGA as a possible compile target, that way people could run apps that seamlessly used the FPGA where appropriate. I have mentioned this to Intel, who are promoting this technology (and also have a team that contributes to the JVM), but so far no one is stating publicly that they are working on such a thing.
- baybal2 10y agojava hello world will not fit even into a 10 gigagate chip
- theatrus2 10y agoBecause the model is so different there would be no benefit.
- brendangregg 10y agoIntel already have a compression library as a proof of concept that shows a large benefit. The JVM compiler knows A) how many instructions each method is and B) how CPU hot it is. Just with a compression library, the compiler could identify very hot and very small methods and test them on the FPGA, in parallel to normal execution, and measure the performance difference, and switch to the FPGA if it was beneficial (which may be for <1% of methods). I believe the JVM already has much of the infrastructure to do such parallel method tests.
- technological 10y agoQuick Question: If anyone wants to learn programming an FPGA is learning C only way to go ? how hard is to learn and program in verilog/VHDL without electrical background ? If anyone suggests links or books, please do Thank You
- ranman 10y agoI have a physics background but not an EE background. I found verilog pretty easy to grasp. VHDL took me a lot longer. To get some basic ideas I always recommend the book code by charles petzold: https://www.amazon.com/Code-Language-Computer-Hardware-Software/dp/0735611319 https://www.amazon.com/Code-Language-Computer-Hardware-Softw... It walks you through everything from the transistor to the operating system. (Apparently I need to add that I work for AWS on every message so yes I work for AWS)
- technological 10y agoThank you
- pjmlp 10y agoNo, you can also go the Ada way with VHDL. One key difference to keep in mind for digital programming is that everything happens in parallel, unless explicitly serialized, which is the opposite of the usual software development most people know about.
- grandalf 10y agoI found VHDL much easier than C or verilog, I think it has to do with how your brain is wired.
- ktta 10y agoI would suggest Digital Design by Morris Mano[1]. It'll start off with basic intro from digital gates to FPGAs itself! And you really don't need any EE background for this book. This book starts from absolute basics and it'll also teach you Verilog along the way. And verilog is used more in the industry than VHDL(which more popular in Europe and in the US army for some reason). I'm surprised where you got the idea of using C to program FPGAs, are you thinking of SystemC or OpenCL (they're both vastly different from each other) I'm really surprised a sibling comment recommended the code book. It really meant to be a layman's reading about tech. It's a great book but it won't teach you programming FPGAs. [1]: https://www.amazon.com/Digital-Design-Introduction-Verilog-HDL/dp/0132774208 https://www.amazon.com/Digital-Design-Introduction-Verilog-H...
- jordz 10y agoAzure will be next I guess. They're already using FPGA based systems to power Bing and their Cognitive Services.
- XnoiVeX 10y agoThat's just anecdotal. No one has seen it. The Wired article sounded like content marketing.
- dgacmu 10y agoThey've published papers about it -- https://www.microsoft.com/en-us/research/publication/configurable-cloud-acceleration/ https://www.microsoft.com/en-us/research/publication/configu... -- they're giving talks about it -- Mark Russinovich was here a few weeks ago with a very long talk. Doug Burger and Derek Chiou are leading a lot of these efforts, and they're absolutely for real. I'm not sure I agree with them that this is the right path forward (but they're smart and know their stuff, so I'm probably wrong), but it's absolutely for real.
- RossBencina 10y ago> Xilinx UltraScale+ VU9P fabricated using a 16 nm process. > 64 GiB of ECC-protected memory on a 288-bit wide bus (four DDR4 channels). > Dedicated PCIe x16 interface to the CPU. Does anyone know whether this is likely to be a plug-in card? and can I buy one to plug in to a local machine for testing?
- smilekzs 10y agoEven if it does, this can easily sell for $10k+.
- errordeveloper 10y agoYeah, the point is that you should need to buy any hardware even for development, which is the biggest win to me!
- aseipp 10y agoBut having the hardware is vital. You have to test your design a lot. You're still going to need Vivado (which isn't cheap) and you'll need instance time to test the design on the real hardware with real workloads, along with any syntheiszable test benches you want to run on the hardware. The pricing structure of the development AMI is going to be meaningful here, because it clearly includes some kind of Vivado license. It might not be as cheap as you expect, and you need to spend a lot of time with the synthesis tool to learn. The F1 machines themselves are certainly not going to be cheap at all. If you want to learn FPGA development, you can get a board for less than $50 USD one-time cost and a fully open source toolchain for it -- check my sibling comments in this thread. Hell, if you really want, you can get a mid-range Xilinx Artix FPGA with a Vivado Design Edition voucher, and a board supporting all the features, for like $160, which is closer to what AWS is offering, and will still probably be quite cheap as a flat cost, if you're serious about learning what the tools can offer: http://store.digilentinc.com/basys-3-artix-7-fpga-trainer-board-recommended-for-introductory-users/ http://store.digilentinc.com/basys-3-artix-7-fpga-trainer-bo... -- it supports almost all of the same basic device/Vivado features as the Virtex UltraScale, so "upgrading" to the Real Deal should be fine, once you're comfortable with the tools.
- koolba 10y agoJust wait till this gets combined with Lambda.
- dx034 10y agoHow would they do that? Since the FPGAs are not shared, I don't see how you could use it for very short-lived instances.
- koolba 10y agoIf the spin up time is fast enough then they could do it. Alternatively if it's active enough then there would a stream of requests processed by the same, already loaded, FPGA.
- the_duke 10y agoI'd be interested in practical use cases that come to your mind (like someone who commented about genomics data processing for a university). What could YOU use this for professionally? (I certainly always wanted to play around with an FPGA for fun...)
- ktta 10y agoMachine Learning, most likely. See this: https://news.ycombinator.com/item?id=13074021 https://news.ycombinator.com/item?id=13074021
- scott_wilson46 10y agoMonte Carlo sims for options pricing? I've done this before on FPGA, might have a go at doing it for this instance as a fun exercise to test the concept!
- dx034 10y agoNot sure if that makes sense with the offer that Amazon has. The machines are huge, so either you're pricing a huge amount of options at a very high speed (which you'd probably do in-house with FPGAs that you own), or you'll be much cheaper using a good machine locally. Never found MC sims to be a bottleneck regarding time, but YMMV I guess?
- scott_wilson46 10y agoI've heard (although admittedly never seen in practice) that some places take a long time for this sort of things (running over a cluster of computers overnight). If you could do the same job on a single F1 instance in say an hour then I think that would be compelling! Bearing in mind that simple experiments I did showed an improvement of around 100x for this sort of task over a GPU.
- adamnemecek 10y agoDoes this mean that ML on FPGA's will be more common? Can someone comment on viability of this? Would there be speedup and if so would it be large enough to warrant rewriting it all in VHDL/Verilog?
- ktta 10y agoYes, definitely to your first and last two questions! It's not as viable as it resulting in a large scale FPGA movement anytime soon since the the industry and academia is heavy experienced with using GPUs. The software and libraries on GPUs, like CUDA, TensorFlow and other open source libraries are very mature and are optimized for GPUs. There will have to be libraries in Verilog (I for one I'm hoping to be a part of this movement for some time now, so I'd love it if anyone can guide me to anything going on) There are some major to minor hurdles. Although some of them might not seem like much[0], here they are: 1. Till now deep learning/machine learning researchers have been okay with learning the software stack related to GPUs and there are widespread tutorials on how to get started, etc. Verilog/VHDL is a whole different ball game and a very different thought process. (I will address using OpenCL later) 2. The toolchain being used is not open source and it's not really hackable. Although that is not that important in this case, since you're starting off writing gates from scratch, there will be problems with licensing, bugs that will be fixed at snail's pace (if ever) till there will be a performant open source toolchain (if ever, but I have hope in the community). You'll have to learn to give up at a customer service rep if you try to get help, unlike open source libraries where to head to github's issue page and get help quickly with the main devs. 3. Although this move will make getting into the game a lot easier, it will still not change the fact that people want to have control over their devices and it will take time for people to realize they have to start buying FPGAs for their data centers and use them in production, which has to happen sometime soon. Using AWS's services won't be cost effective for long term usage, just like GPUs instances(I don't know how the spot instance siutation is going to look with the FPGA instances). This comes with it's own slew of SW problems and good luck trying to understand what's breaking what with the much slower compilation times and terribly unhelpful debugging messages. 4. OpenCL to FPGA is a mess. Only a handful of FPGAs supported using OpenCL. So this has lead to there being little to no open source development surrounding OpenCL with FPGAs in mind. And no the OpenCL libraries for GPUs cannot be used for FPGAs. More likely as from scrach rewrite. There should be a LOT more tweaking done to get them to work. OpenCL to FPGA is not as seamless as one might think and is ridden with problems. This will again, take time and energy by people familiar with FPGAs who have been largely out of the OSS movement. Although I might come of as pessimistic, I'm largely hopeful for the future in the FPGA space. This move isn't great news just because it lowers the barrier, but introduces a chip that will be much more popular and now we have a chip for which libraries can focus their support on, compared to before, when each dev had a different board. So you'll have to get familiar with this -- Virtex Ultrascale+ XCVU9P [1] And also, what might be interesting to you is that, Microsoft is doing a LOT on research on this. I think all of the articles on MS's use of FPGAs can explain better than I can in this comment. Some links to get you started: MS's blog post: http://blogs.microsoft.com/next/2016/10/17/the_moonshot_that_succeeded/ http://blogs.microsoft.com/next/2016/10/17/the_moonshot_that... Papers: https://www.microsoft.com/en-us/research/publication/accelerating-deep-convolutional-neural-networks-using-specialized-hardware/ https://www.microsoft.com/en-us/research/publication/acceler... Media outlet links: https://www.top500.org/news/microsoft-goes-all-in-for-fpgas-to-build-out-cloud-based-ai/ https://www.top500.org/news/microsoft-goes-all-in-for-fpgas-... https://www.wired.com/2016/09/microsoft-bets-future-chip-reprogram-fly/ https://www.wired.com/2016/09/microsoft-bets-future-chip-rep... I'd suggest started with the wired article or MS's blog post. Exciting stuff. [0]: Remember that academia moves at a much slower pace in getting adjusted to the latest and greatest software than your average developer. The reason CUDA is still so popular although it is closed source and you can only use nvidia's GPUs is that it got in the game first and wooed them with performance. Although OpenCL is comparably performant(although there are some rare cases where this isn't true), I still see CUDA regarded as the defacto language to learn in the GPGPU space. [1]: https://www.xilinx.com/support/documentation/selection-guides/ultrascale-plus-fpga-product-selection-guide.pdf#VUSP https://www.xilinx.com/support/documentation/selection-guide...
- anujdeshpande 10y agoHere's a post by Bunnie Huang, from a few months ago saying that Moore's law is dead and we will now have more of such stuff - http://spectrum.ieee.org/semiconductors/design/the-death-of-moores-law-will-spur-innovation http://spectrum.ieee.org/semiconductors/design/the-death-of-... Pretty interesting read. Also, kudos to AWS !
- krupan 10y agoFor complex designs the simulator that comes with the Vivado tools (Mentor's modelsim) is not going to cut it. I wonder if they are working on deals with Mentor (or competitors Cadence and Synopsys) to license their full-featured simulators. Even better, maybe Amazon (and others getting into this space like Intel and Microsoft) will put their weight behind an open source VHDL/Verilog simulator. A few exist but they are pretty slow and way behind the curve in language support. Heck, maybe they can drive adoption of one of the up-and-coming HDL's like chisel, or create one even better. A guy can dream...
- Cyph0n 10y ago> For complex designs the simulator that comes with the Vivado tools (Mentor's modelsim) is not going to cut it. It's now called QuestaSim I believe. But are you sure it can't handle simulating large designs? If yes, what is the full-featured software from Mentor that can? > Heck, maybe they can drive adoption of one of the up-and-coming HDL's like chisel Chisel isn't a full-blown HDL from what I understand; it's only a DSL that compiles to Verilog. In other words, you'd still need a Verilog simulator to actually run your design.
- krupan 10y agoQuesta is the full blown tool. Modelsim is a step down and that's what comes with FPGA tools. Usually the version of modelsim that Xilinx and Altera ship is crippled performance wise.
- scott_wilson46 10y agoNowadays, I don't believe you need a paid-for simulator like Questa, VCS, etc. I am developing verilog in my day job for FPGA's using icarus verilog (an open source simulator)which works fine for fairly large real world designs (I am also using cocotb for testing my code) and supports quite a lot of system verilog too.
- LeifCarrotson 10y agoAs someone who has little experience with FPGAs beyond some experiments with a Spartan-6 dev board that mostly involved learning to write VHDL and building a minimal CPU, I found the simulator to be of limited use. My tiny projects were small enough that the education simulator was plenty fast. It was nice when I didn't have the board available, and occasionally, the logic analyzer was useful when I didn't understand what my code was doing to a data structure. But usually, it was just a lot easier to simply flash the board and run the thing. What's the use of a simulator when you can spin up an AWS instance and run your program on a real FPGA?
- krupan 10y agoThe traditional EDA tool companies (Mentory, Cadence, Synopsys) all tried offering their tools under a could/SaaS model a few years back and nobody went for it. Chip designers are too paranoid about their source code leaking. I wonder if that attitude will hamper adoption of this model as well?
- CamperBob2 10y agoChip designers are too paranoid about their source code leaking. It's more an issue of being able to reproduce an existing build later on. You can't delegate ownership of the toolchain to the "cloud" (read: somebody else's computer) if you think you'll ever need to maintain the design in the future.
- gricardo99 10y agoI'm not so sure that is the issue. Currently you delegate ownership of the toolchain to the EDA vendor. Sure you have tools installed locally on your machines, but the tools typically have licenses that expire, so there's never a guarantee you can build it later with the exact same toolchain. Also EDA vendors end-of-life tools at some point, so even if you pay, that tool won't exist for ever, and the license will not be renewable. I do think the issue with cloud is the concern over IP. There are not a lot of EDA vendors, so the chances that your competitor is also using that same EDA vendor is pretty high. I think companies are pretty wary of using a cloud hosted service where you could literally be running simulations on the same machines as your competitors. Can you imagine some cloud/hosting snafu resulting in your codebase being accessible by your competitors? EDA companies also sell ASIC/FPGA IP, and VIP (verification IP), so there's also a pretty clear conflict of interest if they have access to your IP. So, if you're really paranoid, imagine the EDA vendors themselves picking through your IP and repackaging/reselling it as IP to other customers (encrypted of course so you can't readily identify the source code)?
- gricardo99 10y agothe EDA tools need your source code (HDL) to simulate or synthesize the design. But with these F1 instances, potentially the model doesn't have that problem. You develop/design an FPGA solution (some type of accelleration), then you provide it as a service. You don't expose your source code to your end customer, or the EDA tool companies. You do however, potentially expose your source code to Amazon. But possibly not, if you do your design/testing on EDA tools under your control, then deploy FPGA build packages to the F1 instances for hardware testing.
- ktta 10y agoIf anyone is wondering how the FPGA board looks like https://imgur.com/a/wUTIp https://imgur.com/a/wUTIp
- PeCaN 10y agoGood lord that is beautiful. What a massive FPGA.
- alexforencich 10y agoAre they actually using that one, or is that just a board that happens to have that particular FPGA on it?
- ktta 10y agoThe FPGA in the image is the retail version. But it's more than likely that amazon is using the same one since they don't modify the GPUs although they purchase them on a much larger scale.
- alexforencich 10y agoHow do we know they are using that particular board from bittware as opposed to a board from a different manufacturer or even an in-house design? The linked article does not mention bittware or the board part number.
- jeffnappi 10y agoHere's an even better view: http://www.bittware.com/xilinx/wp-content/uploads/sites/5/2016/10/XUPP3R.png http://www.bittware.com/xilinx/wp-content/uploads/sites/5/20... via http://www.bittware.com/xilinx/product/xupp3r/ http://www.bittware.com/xilinx/product/xupp3r/ Thanks OP
- n00b101 10y agoThis is huge
- brilliantcode 10y agowow. that's what was going through my mind reading this article but it quickly dawned upon me (and sad) that I probably won't be able to build anything with it as we are not solving problems that require programmable hardware but euphoric nonetheless to see this kind of innovation coming from AWS.
- SEJeff 10y agoAre these custom fpgas or an Altera or Xylinx?
- AlphaWeaver 10y agoIt appears to be a Xylinx.
- alexforencich 10y agoLooks like they are using Xilinx Ultrascale+ FPGAs.
- mozumder 10y agoAnyone have a hardware ZLIB implementation that I can drop into my Python toolchains as a direct replacement for ZLIB to compress web-server responses with no latency? Could also use a fast JPG encoder/decoder as well.
- wmf 10y agoGiven that your EC2 Web server is limited to 20 Gbps, you're probably better off using Intel zlib and choosing the right compression level tradeoff. If you're willing to pay a fortune for 100 Gbps of zlib then the FPGA might be more appropriate. For JPEG the GPU instances might be better.
- mozumder 10y agoThe problem is the latency associated with software Zlib, on the order of several milliseconds for a typical web response, and the CPU usage the entails, thereby limiting web request-response throughput.
- fpgaminer 10y agoWhy stop there? Hack your kernel to deliver network packets directly to the FPGA and then implement the whole server stack in the FPGA. Why settle for response times on the order of milliseconds when you can get nanoseconds? But seriously, I'm open to ideas for technologies that you or anyone else needs implemented for these instances. Would make an interesting side business for me. EDIT: I should point out that I'm an experienced "full-stack" engineer when it comes to FPGAs. I've implemented the FPGA code and the software to drive them. None of this software developed by "hardware guys" garbage.
- mozumder 10y agoSpeaking as a hardware guy, I think that's the ultimate goal as well :) Been planning a NIC card that directly serves web apps via HDL for a while now...
- kylek 10y agoI'm not totally up to date on it, but the RISC-V project has a tool (Chisel) that "compiles" to verilog... Interesting times for sure!
- emmelaich 10y agoAlso checkout clash-lang.org which takes an almost Haskell language to VHDL or Verilog or others.
- mmosta 10y agoFPGA Instances are a game changer in every way. Let this day be known as the beginning of the end general-compute infrastructure for internet scale services.
- Sanddancer 10y agoI'm surprised that no one has linked to http://opencores.org/ http://opencores.org/ opencores yet. They've got a ton of vhdl code under various open licenses. The project's been around since forever and is probably a good place to start if you're curious about fpga programming.
- 1024core 10y agoI'm a total FPGA n00b, so here's a dumb question: what can you do with this FPGA that you can't with a GPU? OK, here's a concrete question: I have a vector of 64 floats. I want to multiply it with a matrix of size 64xN, where N is on the order of 1 billion. How fast can I do this multiplication, and find the top K elements of the resulting N-dimensional array?
- deelowe 10y agoI can't answer the GPU comparison question, but I can answer the question of what you "can" do on a FPGA. Here are some example cores for FPGAs: http://opencores.org/projects http://opencores.org/projects Hopefully, by browsing that list, you can see how FPGAs aren't really directly comparable to something like a GPU.
- cheez 10y agoFGPA = Field Programmable Gate Array. Basically, you can create a custom "CPU" for your particular workflow. Imagine the GPU didn't exist and you couldn't multiply vectors of floats in parallel on your CPU. You could use a FPGA to write something to multiply a vector of floats in parallel without developing a GPU. It would probably not be as fast as a GPU or the equivalent CPU, but it would be faster than doing it serially. Another way to put it: you can create a GPU with a FPGA, but not vice versa.
- 1024core 10y agoThanks. But what's the capacity of this particular FPGA? How much can it "do" ? Surely it can't emulate a dozen Xeons; so what's the upper bound on what can be done on this FPGA?
- Ceriand 10y agoIs there direct DMA access to/from the network interface bypassing the CPU?
- alexforencich 10y agoDoesn't look like it from the article. That could be very interesting, but there could be network architecture constraints that prevent Amazon from providing that from the get-go. And it wouldn't be used in all cases, so that could burn a lot of switch ports. Seems like they're targeting more compute offload and less network appliance.
- fpgaminer 10y agoThese FPGAs are absolutely _massive_ (in terms of available resources). AWS isn't messing around. To put things into practical perspective my company sells an FPGA based solution that applies our video enhancement technology in real-time to any video streams up to 1080p60 (our consumer product handles HDMI in and out). It's a world class algorithm with complex calculations, generating 3D information and saliency maps on the fly. I crammed that beast into a Cyclone 4 with 40K LEs. It's hard to translate the "System Logic Cells" metric that Xilinx uses to measure these FPGAs, but a pessimistic calculation puts it at about 1.1 million LEs. That's over 27 times the logic my real-time video enhancement algorithm uses. With just one of these FPGAs we could run our algorithm on 6 4K60 4:4:4 streams at once. That's insane. For another estimation, my rough calculations show that each FPGA would be able to do about 7 GH/s mining Bitcoin. Not an impressive figure by today's standards, but back when FPGA mining was a thing the best I ever got out of an FPGA was 500 MH/s per chip (on commercially viable devices). I'm very curious what Amazon is going to charge for these instances. FPGAs of that size are incredibly expensive (5 figures each). Xilinx no doubt gave them a special deal, in exchange for the opportunity to participate in what could be a very large market. AWS has the potential to push a lot of volume for FPGAs that traditionally had very poor volume. IntelFPGA will no doubt fight exceptionally hard to win business from Azure or Google Cloud. * Take all these estimates with a grain of salt. Most recent "advancements" in FPGA density are the result of using tricky architectures. FPGAs today are still homogeneous logic, but don't tend to be as fine grained as they were. In other words, they're basically moving from RISC to CISC. So it's always up in the air how well all the logic cells can be utilized for a given algorithm.
- neurotech1 10y agoAny thoughts on why AWS/Xilinx didn't go for a mid-range FPGA to help validate customer requirements? My guess is that Amazon will have to be very careful not to price themselves out of the market, for mid-range Deep Learning based cloud apps. Wild guestimate but I think it'll cost more than $20/hr for each instance.
- fpgaminer 10y agoBased on my speculation, and to make a long analysis short: fewer, bigger FPGAs are better in the cloud from a user experience perspective than more, smaller FPGAs. The big applications are all going to consume as much FPGA fabric as they can (machine learning, data analysis, etc). Even "mid-range" Deep Learning will consume these FPGAs like candy. Non-deep learning will too; they can always just go more parallel and get the job done faster. Amazon is betting on the fact that they can get better pricing than anyone else. They probably can. No one else will be buying these FPGAs in quantities Amazon will if these instances become popular (within their niche). So for the medium sized players it'll be cheaper to rent the FPGAs from Amazon, even with the AWS markup, than to buy the boards themselves. Especially for dynamic workloads where you're saving money by renting instead of owning (which is generally the advantage of cloud resources). That's my guess anyway.
- jsh3d 10y agoThis is amazing! We have been developing a tool called Rigel at Stanford (http://rigel-fpga.org http://rigel-fpga.org) to make it much easier to develop image processing pipelines for FPGA. We have seem some really significant speedups vs CPUs/GPUs [1]. [1] http://www.graphics.stanford.edu/papers/rigel/ http://www.graphics.stanford.edu/papers/rigel/
- petra 10y agoAre the speedups enough to negate the much higher cost of an FPGA vs a GPU ?
- kidlj 10y agoTry to comment
- huntero 10y agoGiven that the Amazon cloud is such a huge consumer of Intel's X86 processors, even using Amazon-tailored Xeon's, it's surprising that Amazon chose Xilinx over the Intel-owned Altera. These Xilinx 16nm Virtex FPGA's are beasts, but Altera has some compelling choices as well. Perhaps some of the hardened IP in the Xilinx tipped the scales, such as the H.265 encode/decode, 100G EMAC, PCI-E Gen 4?
- rphlx 10y agoStratix10 (the large, Intel 14nm family) was delayed, delayed, delayed, and delayed some more. Last I heard it was supposed to be in high-prio customer hands by end of 2016, but unclear if that meant "more eng samples" or the actual, final production parts. Either way Xilinx beat them to market by approx 3-6 months AFAICT.
- lisper 10y agoNewbie question: What do verilog and VHDL compile down to, i.e. what is the assembly/machine language for FPGAs?
- anilgulecha 10y agologic-gate layout?
- JensSeiersen 10y agoA logic-gate/register netlist, i.e. a digital schematic of your design. This is done by a synthesizer program. It is then mapped to the available resources of your chosen FPGA, by a mapping program. Now you have the logic equivalent schematic using the FPGAs resources. Then the netlist is place-and-routed to fit it into the FPGA. If the design is to large/complex or the timing requirements to strict (to high a clock frequency), this phase can fail. This phase can also take many hours to complete, even on fast computers.
- alexforencich 10y agoBinary FPGA configuration instructions - block RAM contents, routing switch configruation, register configuration an initial state, PLL/DCM configuration, and of course LUT contents. That's the final result of the toolchain, ready to get sent to the FPGA via JTAG or written into a configuration flash chip. It's the FPGA equivalent of machine code. For the higher level object file or assembly language, that would be a netlist - essentially a digital representation of a schematic. The HDL is transformed into a netlist, then the netlist is optimized and the components converted from generics to device-specifc components, then the placement and routing is determined, and finally a 'bit' file is generated for actually configuring the FPGA. This process can take several hours for a large design.
- jasoncchild 10y agoOh man...this is freaking awesome!