4 ms·
Peanuts? So how much should they have paid for a capable solution then?
by cnd47 10y ago
Peanuts? So how much should they have paid for a capable solution then?
- sjwright 10y agoI don't know how much, but that barely seems like enough to pay for the hardware to deal with 10+ million visitors in one evening. (Apparently they aren't allowed to use cloud infrastructure for data safety / privacy reasons.)
- ComputerGuru 10y agoOnly if you use some "hip" and bloated web framework/platform. Tell software could handle that without breaking a sweat on minimal hardware.
- josephg 10y ago> that barely seems like enough to pay for the hardware to deal with 10+ million visitors in one evening. Really? The specs: The online census form was a pretty simple client side app. I'm going to say, I could whack something together in much less than 200k of JS. For each household the site needs to: 1. Serve the client side app. The app had no images beside the ABS header. I'll be generous and say 200k of static assets including the client side JS. 2. Accept & validate the code households received on their postal forms 3. Save the final result (a single record) back to the databases. A single append operation. Being generous, 20k of JSON data per household. When I filled in the census on Thursday night I noticed it made no effort to normalize addresses or anything else. Umless I'm missing something (or IBM turned off half the features after the DOS) the site was that simple. 1. 200k of static assets * 10 million households = 2TB of traffic. AWS would serve 2TB of traffic out of cloudfront for a whopping $200. If you don't trust 'the public cloud' your self-run data center could serve a tiny payload which downloaded the JS from cloudfront and verified its hash before running it. Or just serve all 2TB of traffic directly. 2. Accepting the unique code on the posted form is easy - just have a few servers which store the database of all posted codes in RAM and respond with 'yes' or 'no'. The codes for 10M households will fit in less than a gig of ram. 2. Saving the final result is pretty trivial. You need to accept 200 gigs of JSON data in an append-only journal. The peak write load would be around 4 million requests / hour, or cumulative ~1000 writes per second, or ~20MB/second of data. For an append-only store backed by an SSD this isn't much load at all. If the numbers weren't so low I'd suggest using a hand-written journal, but honestly most off the shelf databases could handle this write load without breaking a sweat. And you could easily shard it if you wanted to, and just aggregate back the data after its been collected. (You might want to actively geo-replicate the data immediately too, while its being collected just in case something happens). In total, I'd feel comfortable running this entire site out of a few dedicated machines in just about any data center in the country. If you're feeling paranoid geo-replicate it, and actively stream received census submissions to a couple offsite backup locations. Total whiteboarding time: about 20 minutes. Might make a good interview question. Of course IBM was also paid for consultancy time, and engineering time to design & implement the site itself, but even factoring in healthy margins I can't help but feel our government was ripped off by a factor of ~10-20x.
- dmoy 10y agoWas it a single form with no support for save & come back later to finish?
- DougWebb 10y agoEven if it supported saving and coming back later, that barely complicates things. The form itself would need the ability to redisplay with fields pre-populated, which you'd want anyway if you've got any server-side validation or to handle submit errors with an option to retry. For the database, you'd want to use caching. When the form is submitted you write the data locally and queue it for replication to a central server. To redisplay you'd try to read it locally and if you don't find it you'd read from the central server. You'd probably want to use the household's code for sharding so that you're going to the same local database each time. I can't imagine it'd be necessary to spread the load out geographically. I'd have a single datacenter / cloud equivalent set up with load-balanced machines for the front end and the database tier, a full fail-over backup in case of critical failure of the main datacenter, and a separate database server for the master which the others replicate to and clone from. This is a six-figure project, for sure, except maybe for time spent dealing with governmental red-tape.
- rbobby 10y agoTo me it feels like you've left out an awful lot of stuff. First off is project mgmt and coordination. On the government side there's going to be a small host of people involved, not just a single "customer". Meetings and more meetings. Somewhere between 50% and 100% of someone's time is going to be spent just dealing with them. And that's on top of the normal PM effort (so this could easily be 2 full time people). You've also left off the generation of the mail outs. This means a data load, an id assignment, and a "print file" for the printer/mailer. High coordination costs (PM now dealing with the gov't and the print shop). Security. Ouch. Need to make some effort to ensure the mail out codes can't be guessed easily... but are still small enough/readable enough that folks can use them (i.e. a GUID would be perfect except for the low usability). Further there ought to be some sort of post processing that attempts to ensure the data is good (i.e. no one played silly buggers and guessed codes). On the application side "just in case something happens" isn't good enough. Once the incoming census data is accepted it must not be lost. The statistician would go ballistic (what do you mean you don't know how much data you lost?) and there's really no way to recover lost data without sending out new mailings. On the back end is the final "give the stats guy the data" step. Conceptually just a giant CSV file... but how likely is that really? Also somewhere/somehow someone needs to build a list of households that have not completed the survey (and reporting on how many/what percentage and eventually producing lists of households for census staff to visit... which is probably handled internally by the existing systems... but it's another entire interface). And all the design back and forth of how the website should look (and it has to be properly accessible). Not really hard to do on IBM's side... but the gov't side is going to be a nightmare of micromanaging UX specialists. Oh and general security... it's a web app so there's all that. Hopefully the testing would be handled by a separate organization... which means another organization to coordinate with (PM is gonna be busy). And of course all the QA you can stand. Plus... there's legal costs (that contract ain't gonna review itself). Plus... there's the upfront costs of the actual bidding process. Can't bill for it... but this does mean that every successful bid by IBM needs to be higher to cover these sunk costs (i.e. if IBM wins 20% of all its bids then the costs for 80% of the failed bids need to be covered by the winning 20%... if you see what I mean). And now IBM needs to pay for some PR/damage control. In the end I'm not so sure that $9 million is so very out of whack. And damn it all if I didn't forget to include something for planing/supporting defending against denial of service attacks...