4 ms·
Article is Paywalled but its basically about this twitter thread from Bloom Labs: https://twitter.com/jbloom_lab/status/1407445604029009923 https://twitter.com
by codeulike 5y ago
Article is Paywalled but its basically about this twitter thread from Bloom Labs:
https://twitter.com/jbloom_lab/status/1407445604029009923 https://twitter.com/jbloom_lab/status/1407445604029009923
This technical bit is interesting - although the data had been 'deleted' from the Sequence Read Archive* web app by the original submitter, this tweet explains that they were able to recover the data via storage.googleapis.com:
https://twitter.com/jbloom_lab/status/1407445615248691201 https://twitter.com/jbloom_lab/status/1407445615248691201
I discovered that even though the files were deleted from archive itself, they could be recovered from the Google Cloud at links like https://storage.googleapis.com/nih-sequence-read-archive/run https://storage.googleapis.com/nih-sequence-read-archive/run... (5/n) ...
So technical question for HNers - what lives at storage.googleapis.com usually? Was that like a cloud mirror or was it more like the 'delete' function in the web app was only removing things from the index but leaving the deleted stuff accessible?
* Sequence Read Archive seems to live at https://www.ncbi.nlm.nih.gov/sra https://www.ncbi.nlm.nih.gov/sra
- colechristensen 5y agoThey were supposedly storing artifacts publicly available on GCP's object store, an efficient way to do things for distribution of non-secure large pieces of data. "delete" deleted the reference to these objects but the objects were kept around. (This is a not-so-bad practice, if you're hacked and somebody tries to wipe everything, or some bug or fat finger deletes everything, you've deleted references to data not actual data)
- geoduck14 5y agoYou really should clean up your soft deleted objects. They are a perfect place for hackers to get data that you thought was gone.
- colechristensen 5y agoIf you are in the business of publishing scientific data, maybe you want people to be able to find "deleted" data, or at least you're not too concerned with really intensely enforcing people's abilities to delete something they have published.
- ve55 5y agoMost likely the latter: the server-side code in charge of deleting data did not make a call to their storage api to also remove the object itself. There's a good chance that is intentional and serves as a soft-deletion function, such that it could be reverted (or the data otherwise used) if needed.
- pitched 5y agoThat was my read on this too and Dr. Bloom accidentally hacked the NIH. The next question though is whether they’ll change this or not? Is the guarantee that anyone can retract at any time important enough to make the db useful? Will the Chinese government mandate no one there ever use it again now?
- nitrogen 5y agoDr. Bloom accidentally hacked the NIH. In case this is why this comment was downvoted, it's worth remembering that others have been charged with CFAA violation for basically the same thing.
- BlueTemplar 5y agoReminds me of this guy that got fined $4k by the court for something similar ? https://arstechnica.com/tech-policy/2014/02/french-journalist-fined-4000-plus-for-publishing-public-documents/ https://arstechnica.com/tech-policy/2014/02/french-journalis...
- jimmyswimmy 5y agoActing in the sense of as a devil's advocate here, but this guy went and rooted around a data repository for stuff that wasn't his... there was a case a few years ago of a guy who scripted the AT&T site to download similarly accessible data [0] and got 4 years jail for it. We can trot out the well-known Swartz story [1] here as well. Though I'm happy to have the information available in this particular case, I can't really say this is different in terms of either the letter or spirit of the law. So: when should we expect to hear this guy being charged with CFAA violation? I'm not holding my breath, and I hope he doesn't, but what he did seems as clear a violation of the law as the two examples listed here. [0] https://www.wired.com/2013/03/att-hacker-gets-3-years/ https://www.wired.com/2013/03/att-hacker-gets-3-years/ [1] https://en.wikipedia.org/wiki/United_States_v._Swartz https://en.wikipedia.org/wiki/United_States_v._Swartz
- 22c 5y agoIt was also discussed here yesterday: https://news.ycombinator.com/item?id=27598222 https://news.ycombinator.com/item?id=27598222 https://www.biorxiv.org/content/10.1101/2021.06.18.449051v1 https://www.biorxiv.org/content/10.1101/2021.06.18.449051v1 Unless this is a new instance of mysterious sequence deletion.