28 ms·
The implementation of the UK Covid-19 dashboard
- CraigJPerry 5y agoI like the UI design in figure 1. There’s no crap in the way of the data but i don’t feel overwhelmed either. My eyes can scan across sections and it feels natural, theres no firehose effect. I like the thought that’s gone into showing the % vaccinated in the top right. I like the dashed underlines telling me that some explainatory text is available. I think the page looks inoffensive but is clearly focussed on being informative. I wish more data repositories took care and attention towards how data is represented.
- cronin101 5y agoI think GOV.UK in general does a very good job of making important information readily available. It’s however extra jarring switching between the clean/efficient style of the online messaging and the underlying public services/offices that are still held together by chewing gum and a fax machine, whenever you have some issue that can only be resolved by persuading a human to stamp some paper (Visas, public records, etc.)
- spaniard89277 5y agoAt least you got that. Try Spain, when the tax agency uses behavioral data and puts incentives for tax agents, but you can't get a f*ing appointment for pretty much any service as they use the most stupid appointment system ever, and 90% of the state services are subcontracted to the lowest bidder making everything a PITA to use. All the fancy tech to get in your pockets. For everything else, go f*k yourself.
- cameronh90 5y agoIt was a lot worse a decade ago. The GDS have done a pretty good job of unifying the design language and methodology of disparate government departments, but of course, it is a huge job. It clearly involves just as much cultural and organisational overhaul as it does technology work. Most recently I found the DVLA license renewal was one of those ugly backwaters (albeit still fully online), but their license check code generator is great. For real terrible stuff, check out local council websites.
- one-more-minute 5y agoIn general the gov.uk website is stellar. Lots of good info, and plenty of thought has gone into making everything clear, accessible and pleasing to the eye. Things generally work in a way that they really don't on most public org websites. The team behind it blog a fair bit about how they've made these things happen. https://insidegovuk.blog.gov.uk https://insidegovuk.blog.gov.uk The only downside is that they often send you to sites run by other, significantly less competent bodies (looking at you, student loans company).
- jsmith99 5y agoThey had an early win getting a effective and consistent UI accross government sites, but digitising the underlying processes is a work in progress. Once you click that beautiful and accessible green submit button, it's not impossible a printer in Swansea whirs to life to print your answers out into the original paper form.
- hunter321 5y agohttp://ehmipeach.defra.gov.uk/ http://ehmipeach.defra.gov.uk/ >Using the right browser - only use Microsoft Edge PEACH is compatible with Microsoft Edge, but only when Internet Explorer mode is enabled. For guidance on setting up Internet Explorer mode in Edge, follow this link to instructions on the Microsoft Support Page. Yeah, there's still some way to go.
- samhw 5y ago> it's not impossible a printer in Swansea whirs to life to print your answers out into the original paper form Speaking of which: my mum went through the form to renew her Oyster card the other day, and, having finally completed it, it generated a PDF form and told her to print it out and take it to the Post Office. So yeah, there are definitely some Jira tickets still on the left.
- samwillis 5y agoWhile I almost entirely agree with you, having had to report positive LFT Covid tests and do the subsequent “test and trace” form this week I feel that while the UI and UX is nice the flow of those forms are really quite pore. There are repetitive/redundant questions plus inconsistent and out of date advise. It’s probably more an effect of the difficult moving target of changing rules and advice, but for what is an important government process everyone is coming into contact with I though it would be better. Maybe wishful thinking though.
- eggy 5y agoYes, I thought the same thing in finding it easier on my eyes/quick perception of the site. I do think the UK and some other countries do a better job of presenting data compared to the CDC. It's pretty much agreed that the rate of unvaccinated people vs. vaccinated people winding up in hospital beds is several times higher, however, all the CDC data presented is only rates. I want tallies or counts, and I cannot find them. For instance, on Ontario, Candada's site[1], the vaccinated are 74% vs. the unvaccinated's 26% of COVID hospitalizations. Most non-technical people think the hospitalizations of COVID patients is like over 90%. It's because more and more people are vaccinated, even with a lower rate of hospitalizations, the numbers are higher. Also, it's interesting to see on the Ontario site that COVID hospitalizations consist of 56% directly for COVID, and 44% were admitted for other reasons and then tested positive for COVID once hospitalized. The case is more telling for ICU with 81% admitted for COVID, and 19% for other reasons. I am trying to play with raw data more for refreshing my munging skills than making a point or fodder to add to the COVID noise. I have been coding since 1978, played with neural nets, GAs, and GP in the late 1980s, but I don't code or do data analysis for a living right now (other than buisness strategy reports that require some basic analysis). There's a lot of data out there, and it can get very confusing. I am back to using R/RStudio from a brief stint using Julia/Pluto notebooks and previously using Python/Jupyter notebooks. I even did a toy DSEIR model in J back in April 2020 based on previous work by a couple of people, which I plan on updating to April[2]. I am going to try and do some Lisp work, and I think I will settle on RStudio and Lisp for more genomic/bioinformatic stuff (yes, I know biolisp has been supplanted by python, however, Lisp is having a renaissance in symbolic-related areas of ML again like NLP). BTW, in what language was GPT implemented, not API languages, but what PL(s) was used to create the code - C++, Java, Go? I may be bad at navigating the CDC website, but I can't seem to get the dataset of numbers of hospitalizations by vaccination status, only rates or pre-filtered data. I do remember downloading raw data that seemed to have it (over 1.8gb, I think), but I can't seem to find it. I'd appreciate a link if anyone has it. [1] https://covid-19.ontario.ca/data/hospitalizations#hospitalizedAndICU https://covid-19.ontario.ca/data/hospitalizations#hospitaliz... [2] https://github.com/phantomics/april https://github.com/phantomics/april
- amelius 5y agoI would have preferred to see it implemented as a spreadsheet in the cloud.
- StringyBob 5y agoWell, you can just grab the data very easily: https://coronavirus.data.gov.uk/details/download https://coronavirus.data.gov.uk/details/download
- rob_c 5y agoThe data is exportable but you then have to deal with 8MB+ csv files. Still it's better than it being in docx I suppose... But then this would let you perform statistical analyses on _their_ data, I'm not sure they're such a big fan of that...
- aembleton 5y agoWhat format would you prefer it to be in than CSV? There are CSV libraries in pretty much every language and any spreadsheet can import them. 8MB really isn't that hard to handle these days.
- trebligdivad 5y agoI do wonder if the power of a distributed database is really needed here; it gets ~1 update a day, so there's no need to have clever consistency stuff. Most of the queries relate to either today's data (you open the map and zoom in to see how doomed your area is today), or the graphs showing a standard set of history (e.g. cases over the last year). You'd think you could extract that data to be static and not require database queries, and only fire up the database for the tiny proportion that go digging in history.
- lloydatkinson 5y agoI had the same thoughts and then it was confirmed how insane this setup is part way through: “At the time of writing, the Citus distributed database cluster adopted by the team on Azure is HA-enabled for high availability and has 12 worker nodes with a combined total of 192 vCores, ~1.5 TB of memory, and 24 TB of storage. (The Citus coordinator node has 64 vCores, 256 GB of memory, and 1 TB of storage.)” That’s beyond overkill for something that as you say could be generated statically a couple of times a day.
- sharken 5y agoMy suspicion is that since this has to do with COVID, there is no real limit on what the cost should really be. As for using the setup for other things, that seems less likely given this expensive setup.
- deleted 5y ago[deleted]
- tgv 5y ago> generated statically a couple of times a day. That would require actual work instead of selling an overpriced generic solution.
- smarx007 5y agoDid you look at the 3 different (non-trivial) APIs they are offering on top of the dashboard? Though I have a hard time understanding why use PostgreSQL instead of ClickHouse, for example.
- raesene9 5y agoThe covid dashboard I've found the nicest to use, is actually not one of the official ones but instead this https://www.travellingtabby.com/scotland-coronavirus-tracker/ https://www.travellingtabby.com/scotland-coronavirus-tracker... . The information is presented clearly and it's easy to see what's going on, although in my case the main reason is the breakdown for Argyll & Bute, which isn't a focus area for the national ones!
- zlib 5y agohttps://github.com/publichealthengland https://github.com/publichealthengland powered by a lot of Python and Ruby, nice.
- nojito 5y agoAnd a bunch of F# https://github.com/publichealthengland/coronavirus-dashboard-summary https://github.com/publichealthengland/coronavirus-dashboard...
- mdm12 5y agoNice catch! It surprises me how much more popular F# is in Europe compared to the US. I finally got a professional F# gig in the states (\o/), but there were very few options. It makes me wonder, are universities in Europe providing a more functional-first approach to CS education, or is something else going on?
- jll29 5y agoNever seen a job ad on anything sharp other than C# in Europe. There are occasionally LISP and Clojure jobs from what I can tell. (It's also hard to find people on the talent side. I needed a Haskell developer with NLP skills in 2005, and could not find one so we had to port our codebase to Java.)
- mdm12 5y agoIt's anecdotal, but on the F# Software Foundation's slack workspace[1], 4 out of the last 5 postings in the #jobs channel were in Europe. No doubt, any company that picks a niche programming language as their business's lingua franca is taking a risk. For me, though, that is an indicator that they care about quality and do not have a culture of treating engineers as replaceable assembly line parts. [1] https://fsharp.org/guides/slack/ https://fsharp.org/guides/slack/
- intricatedetail 5y ago
- rob_c 5y agoYes. This was never pure ONS data. It was always processed/massaged/tampered with. Although horrifyingly it did allow you to gauge the Westminster mindset and understand what Draconian measures they were planning to introduce sightly ahead of time because the data would have to reflect this when lord Boris got up on the podium...
- rob_c 5y agoDespite the downvotes the only incorrect statement here is that Boris has a lordship. Go export the ONS data yourself and analyse it unless your too lazy to see what I mean
- teh_klev 5y ago> Go export the ONS data yourself and analyse it unless your[sic] too lazy to see what I mean If you're going to make such claims then the onus is you to provide the evidence.
- cameronh90 5y agoSurely the onus is on the person claiming it's tampered?
- rob_c 5y agoNot tampered processed and misrepresented. Two my where you can remove the logarithmic plots on the MS data. Tell me where you can get the data behind this great firewall of information other than FOI requests and do gain some basic literacy in stats to spot some of the obvious oddities in the datasets (they're there)
- llimos 5y ago> From the beginning of the COVID-19 pandemic, the United Kingdom (UK) government has made it a top priority to track key health metrics and to share those metrics with the public. According to Dominic Cummings (ex-adviser to the PM), this isn't true at all - one of their biggest failings early on was to not have the data and not see the priority in getting it.[1] [1] https://news.sky.com/story/dominic-cummings-hearing-the-inside-story-of-the-timeline-of-the-weeks-before-covid-lockdown-12317517 https://news.sky.com/story/dominic-cummings-hearing-the-insi... : He added later that there was no data system at that point, and he needed to use his iphone as a calculator to make predictions about the extent to which infections would spread, which he then wrote down on a white board.
- zarzavat 5y agoIt's worth noting for those who don't follow UK politics (I don't recommend doing so), that Dominic Cummings was fired as the top advisor, and like a jealous ex he is determined to bring down the government by any means possible. So he is an unreliable source to say the least. Although the government seems to have left enough rope to hang themselves without Cummings needing to invent anything.
- Closi 5y agoThe government was tracking key health metrics and sharing them at the point Dom was talking about using his iPhone. For instance on the same date, the government did it's first daily briefing and shared infection and death metrics with the public (see https://www.bbc.co.uk/news/uk-51901818 https://www.bbc.co.uk/news/uk-51901818). Tracking key health metrics and sharing those metrics with the public doesn't mean that there is modelling about the extent to which infections would spread - although we also know that the imperial modelling was released a day later, so while he may have been using his iPhone to make predictions there were also academic teams modelling this that were collaborating with the government at the time (see https://www.imperial.ac.uk/news/196234/covid-19-imperial-researchers-model-likely-impact/ https://www.imperial.ac.uk/news/196234/covid-19-imperial-res...). It's also not clear what a 'data system' is in this context - there was clearly an effort to very quickly put something in place to capture data (because it couldn't wait a few weeks/months), but a more robust analytics system will inevitably take more than a few weeks to put in place if not already in place pre-pandemic (a lot of this is about how NHS trusts are structured in the UK, which operate fairly independently). It's not clear to me how quickly is realistic to implement what Dom thought was suitable in terms of a 'data system', particularly as I'm not particularly clear on his requirements (he seems to want an element of forecasting built into this system for instance?), so without knowing what the requirements are can we be confident that what he wanted to build was possible to build, test and implement in his expected timeline? So I don't think there is a clear contradiction here (and in fact, I think the evidence points to the fact that the statement in the article is probably correct).
- axiosgunnar 5y ago> Government project > awarded to Microsoft Hey Europe, want to stop being several decades behind in IT compared to US/China? One simple trick: Ban FAANG from public procurement in Europe! It‘s a no-brainer really. Buy locally, ideally giving small companies and startups a chance. You will have to do it anyway very soon if you want your privacy laws to be taken seriously. There might be a couple of months of friction while buerocrats have to find new procurement partners, but that's it. And then the European tech scene will rise.
- tjungblut 5y agoThe UK isn't part of the EU anymore.
- rob_c 5y agoBut frankly we're very similar in this regard, and given we might get dropped from H24 we're not in a good position tech or science wise heading into a recovery... We'll no doubt award yet more govt projects to the tech oligarchs of the west and praise students for using their toys...
- Doctor_Fegg 5y agoMaybe GP has edited their comment, but it says "Hey Europe" and we are very definitely still part of Europe.
- azalemeth 5y ago
- rob_c 5y agoGiven the original meetings were all about producing a dynamic simulation dashboard which would allow people and politicians to understand the impact of various measures on lives saved... Typical this turned into a pro cloud puff piece that frankly shows a serious amount of over design for what should be a data filtering/processing step to any reasonable "data scientist". And if I'm having to say a data scientist could do it better you know you got it wrong...
- axiosgunnar 5y ago> people and monsters monsters?
- rob_c 5y agoFixed to politicians. Freudian slip and a half there...
- glogla 5y agoI skimmed the article and it seems interesting. On the data side, they have ~7.5 billion total records and they add in 55 million new a day. On the web side, they have ~1 million daily unique users and 100k concurrent users at peak ("concurrent" means "in one minute" is seems). I'm no expert on the web part, but I'm kind of curious why they went with the design they did for the data part. The design, and the chosen technologies make me think they treated it more like a normal web app, not like a dashboard. I would expect OLAP database, not a sharded Postgres, and the data model feels very OLTP to me as well. Or maybe is that because it's mostly time series and not traditional data model? I'll have to go through the article in more detail.
- GraemeMeyer 5y agoA few other commenters have pointed out the same thing - I’m wondering if it’s simply the skill set they had on hand when the need arose
- mslot 5y agoOLAP stores are relatively fast at answering a single query on a large data set, but basically none of them can handle high throughput with subsecond response times (e.g. when the whole country checks statistics for their own postcode at 4pm). OLTP stores are relatively bad at aggregating across a lot of data. Analytics dashboards with many users, a lot of ever-changing data, and many different views exist in a gray area between OLAP and OLTP often referred to as real-time analytics or operational analytics. The queries are usually somewhat lighter / less ad-hoc / more indexed than in OLAP, but there can be hundreds or thousands of them per second with different filters and aggregations. There are some specialized real-time analytics databases like Druid. Citus (used in the article) allows you to run such workloads at scale on PostgreSQL.
- greatgib 5y agoFirst, when you see public officials doing a blog post on "Microsoft.com" website instead of on a public website, you know that something fishy is going on... On the other side, I have the feeling that this thing that clearly over-engineered. Just look at data their diagram... If I'm not wrong there is one writer and multiple reader for the data, or at least multiple writers on one side and multiple readers on another side, without a need for "real time" consistency. So, this thing could probably have been better splitted to not have the use for "scaled" databases
- sdoering 5y ago> First, when you see public officials doing a blog post on "Microsoft.com" website instead of on a public website, you know that something fishy is going on... The article states it was written by Claire Giordano from San Francisco. Not sure where you got the UK Government official from. To me it read like a b2b marketing piece and showcase. Kind of: We can power this, so we can power your BI dashboard as well. Taking this into account it was a nice write up and from a data analyst's and consultant's pov interesting to read.
- Rastonbury 5y agoComing from consulting, this is exactly what it is - from a pure engineering standpoint it may be lacking but if you've read any other 'case study' targeted as business people this goes really deep. Normally it's SEO fodder crammed with jargon and buzzwords
- arsome 5y agoIt definitely had a bit of a marketing tone, but it was focused on what's ultimately an open source product you can run pretty much anywhere, from a different cloud provider to your own bare metal, they just happened to use it on Azure.
- greatgib 5y agoIndeed I did not check the bio of the article writer, but when you read: <<As a result, the GOV.UK Coronavirus dashboard became one of the most visited public service websites in the United Kingdom.>> You don't expect the gov UK dashboard to be done us consultants...
- snthd 5y agoElsewhere[0] Microsoft have redefined "Open Source" to not include the right to redistribute, or to host on a cloud service. So while there's nothing wrong here with calling an MIT project open souce, it's not inconsistent with their own definition, and useable as propaganda. [0] https://azure.microsoft.com/en-gb/services/developer-tools/data-studio/#documentation https://azure.microsoft.com/en-gb/services/developer-tools/d... >Is Azure Data Studio open source? >Yes, the source code for Azure Data Studio and its data providers is open source and available on GitHub. The source code for the front-end Azure Data Studio, which is based on Microsoft Visual Studio Code, is available under an end-user license agreement that provides rights to modify and use the software, but not to redistribute it or host it in a cloud service. The source code for the data providers is available under the MIT license.
- chrisseaton 5y agoIt’s still useful to have the source open for reference even if you can’t redistribute though.
- nojito 5y agoHow is the most common open source license not open source?
- samhw 5y agoI mean, whatever the truth may be, you're begging the question somewhat by calling it an 'open source license'.
- skylanh 5y agoBSDL, GPL, MIT License. This has been an argument for at least 25 years that I've been around this stuff.
- snthd 5y agoBoth sides of that argument agree that the right to redistribute is fundamental to FLOSS. Microsoft is defining their product, which you can't redistribute[0], as "Open Source". [0] https://github.com/microsoft/azuredatastudio/blob/4012f26976957f6f099ba4b4864a0ea1f558d8c5/LICENSE.txt https://github.com/microsoft/azuredatastudio/blob/4012f26976...
- jll29 5y agoIt's funny to read about a dashboard with TBs of memory and distributed DBs when on HN, people pride themselves on getting Web servers to run on floppy disk based systems. Joking aside, I liked the description of the dashboard, and generally speaking the UK's government Web sites are better quality, support open data more, are easier to read and navigate than other European countries from what I have seen. This includes this dashboard, which looks clean, simple and functional. I was waiting for the big SQL Server advertising language and positively suprised that the article is very tech agnostic. I did all seem to be rather over-engineered, but Microsoft needs to make some money and government agencies don't generally have wizards from HN working for them, so I can live with an occasionally over-engineered system as long as important systems are working and remain up. The most mysterious part for me was why one would put JSON inside relational tables?
- wnolens 5y ago> The most mysterious part for me was why one would put JSON inside relational tables? Cheap and easy way to permit a flexible schema for some part of the data. Performance tests probably showed that for their specific query workload, any slow down from parsing/lack of index was fine.
- 2143 5y agoThe UK covid dashboard[2] mentioned in the article also has a "simple" version.[1] I love it when websites have a simple text version. [1] https://coronavirus.data.gov.uk/easy_read https://coronavirus.data.gov.uk/easy_read [2] https://coronavirus.data.gov.uk/ https://coronavirus.data.gov.uk/
- boomskats 5y agoThis feels like a marketing-led attempt to shoehorn Citus into something topical and shareable, having realised that they've barely talked about it since acquiring the company a couple of years ago. I'm all for Citus, but cmon. Overkill.
- adenner 5y agoIt is miles better than my state's local Covid Dashboard (https://coronavirus.iowa.gov/ https://coronavirus.iowa.gov/) that is updated fully once a week on Wednesdays. On Monday and Friday they simply post a screenshot of a pixelated version of a summary page only.
- 2dvisio 5y agoWonder why they haven’t used powerBI for that, but deep inside I know why...
- idleherb 5y agoLooking at their database schema (table "auth_user"), it looks like the user passwords are stored unencrypted and without salt?