5 ms·
Were these guys a very small shop? Any info on which exchanges they bidded on? The RTB company I worked at was considered "small" and we were handling around
by getsat 11y ago
Were these guys a very small shop? Any info on which exchanges they bidded on? The RTB company I worked at was considered "small" and we were handling around 300,000 requests/second with 15ms max time per request spent in our systems between just Google and Facebook's exchanges.
- chupy 11y agoI suspect a lot of people are interested in the tech used when handling 300k QPS and I wish you could give more details about it.
- getsat 11y agoI'll write a blog post on it later called "So you want to build a realtime bidding system?" and submit it to HN. :)
- ddorian43 11y agoI did this little presentation on doing a toy adserver with python: http://slides.com/dorianhoxha/ad-serving-with-python-uwsgi-lmdb-capnproto/ http://slides.com/dorianhoxha/ad-serving-with-python-uwsgi-l... It wasn't rtb (just normal adserver). Wonder what your comments would be?
- getsat 11y agoI gave your presentation a quick look. Maxmind - good. Not like there's many choices here, though. :) Instead of "waking up to sync logs", consider using something like NSQ to emit events as they happen. You can scale the number of servers/processes generating messages and the number of workers consuming those messages (and committing them to your database) very easily. You could also replace the writing of the transaction log with a NSQ event. It lets you avoid having to write and scale the log shipping stuff. We precalculated which ads a given user was eligible for and a separate process was contacted when a bid request came in to get the info for the ad to show. We never had to do anything funky to do geotargeting exclusion at scale. Instead of having your adserver connect to a database, have a separate process generate a working set (as JSON or whatever you fancy), compress it, and ship it to the adserver periodically. The adserver can just do a straight load from the file every minute or whatever interval you'd like. If the file's mtime is too old, raise an alert and stop serving ads if necessary. Keeping things separate and simple lets you scale more simply. Our working sets were on average about 2gb uncompressed and they could be loaded in a few seconds (C++/JSON and later Go + JSON). Seems like it was a fun project and I hope you learned a lot!
- cmrdporcupine 11y agoAt a previous job we did something similar with separate file sets describing ad inventory, but kept them in a native binary format and mmap'd them.
- ddorian43 11y agoHow do you do load-balancing there? Meaning, can 1 box of nginx/haproxy handle that many requests ?
- chupy 11y agoEven if a single box would be capable to handle that many requests, you should never have just one box to be the main entry point to your server (failsafe). I also suspect he refers to 300k distributed across multiple datacenters.
- getsat 11y agoYes, you are correct. The 300k was total across all datacenters. 85% of that was in the US, 10% in Europe, and 5% in Asia.
- getsat 11y agoAnycast will spread incoming requests out to N physical machines, then you can do your layer 3 load balancing, then you can do your SSL termination and HTTP load balancing. We didn't use anycast, though. I suggested it to our CTO many times and it would have saved us over $5,000/mo in DNS costs, but it never got done.
- lpgauth 11y agoHi, this article is from 2012, we're listening to much more inventory these days.
- getsat 11y agoThat's great!