6 ms·
Datasets for Machine Learning
- Smerity 8y agoMy original comment was meant for a separate HN article on machine learning and I posted in the wrong tab. My apologies.
- pilooch 8y agoI had the same reaction. I don't like it too much when sites copy up information and only link to original content at the bottom of the page. The collection is good though, it's sad that it looks like it is stealing from the sources.
- rerx 8y agoHow is this related to the article on gengo.ai?
- pilooch 8y agoOops, missread for https://modelzoo.co/ https://modelzoo.co/
- rerx 8y agoTo train machine translation models parallel corpora in many languages are provided on the WMT conference site: http://www.statmt.org/wmt17/translation-task.html http://www.statmt.org/wmt17/translation-task.html and previous years
- rahimnathwani 8y agoFrom the title 'The 50 Best Free Datasets...' I was expecting a curated list of datasets. But the list has mix of individual datasets, and sites that provide/host datasets :(
- benhamner 8y agoBen from Kaggle. Open up the ~50 different individual datasets linked in separate tabs, and then quickly flip through all of them trying to get a sense of what each one is. That experience will demonstrate one of the main challenges we're aiming to solve by making Kaggle Datasets your default place to publish data online (https://www.kaggle.com/datasets https://www.kaggle.com/datasets)
- logancg 8y agoThis is a great idea Ben, and I appreciate the work you do. Do you see Kaggle datasets as a tool to encourage better data formatting, or are you also thinking about building tools for automatically visualizing, cleaning, and organising data?
- benhamner 8y agoAll of the above, and more! One thing I'm really excited about that we're about to release is a much better explorer for tabular data (automated histograms, sorting/filtering/showing the data, and the like). We also encourage sharing analytics code and visualizations that users create on the data back to the community. For example, see all these visualizations and insights in StackOverflow's developer survey data linked from https://www.kaggle.com/stackoverflow/stack-overflow-2018-developer-survey/kernels https://www.kaggle.com/stackoverflow/stack-overflow-2018-dev...
- bhnmmhmd 8y agoI've heard that Kaggle data sets encourage people to do "supervised" ML only. Is that true?
- hideo 8y ago(Not Ben, but - ) outside of academia, the main thing that seems to encourage people to do supervised ML is that it's the only thing that seems to work. I haven't really heard of any success stories with using unsupervised techniques for most common ML applications.
- bhnmmhmd 8y agoCan these datasets be used for academic and research purposes?
- logancg 8y agoThe link at the bottom should be emphasized: https://github.com/awesomedata/awesome-public-datasets https://github.com/awesomedata/awesome-public-datasets It is a very expansive collection of datasets, some well-prepped for ML and most not (which is part of the fun of it, anyways).
- kokimame 8y agoFor audio, LibriSpeech, M-AILABS, LJ-Speech, VCTK, TIMIT, Mocha-Timit, VoxForge, Blizzard Challenge, and so on.
- andy-wu 8y agoSurprised that CIFAR wasn’t mentioned under Images. I feel like that’s one of the standards, even more so than some of the ones that are listed.
- codemetro53 8y agoHere is a dataset for abstractive summarization created from Reddit . Dataset https://zenodo.org/record/1168855#.WyJG3I7pdhE https://zenodo.org/record/1168855#.WyJG3I7pdhE Paper http://aclweb.org/anthology/W17-4508 http://aclweb.org/anthology/W17-4508
- greentuna 8y agoDoes anyone know of good datasets for Concept Drift analysis?
- fwdpropaganda 8y agoCan't open this website.
- welly 8y agoClick on the link.
- fwdpropaganda 8y agoDone, what now?
- mrphilroth 8y agoSecurity industry related datasets always seem to be omitted from this type of thing. Please check out the excellent http://www.secrepo.com/ http://www.secrepo.com/.
- danso 8y agoTwo sources that are missing: opendatanetwork.com: this is effectively a Google for public Socrata data portals, and for me, the best way to discover datasets across different municipalities. For example, when I was interested in trying to replicate the NYT's "Do ‘Fast and Furious’ Movies Cause a Rise in Speeding?" [0] article, it was pretty easy to find a bunch of other traffic/motor vehicle violation datasets with opendatanetwork's search. Enigma public (https://public.enigma.com https://public.enigma.com): a huge collection of scraped public datasets, including flattened versions of data that originally comes in annoying-to-parse, such as U.S. lobbying disclosures [1] [0] https://www.nytimes.com/2018/01/30/upshot/do-fast-and-furious-movies-cause-a-rise-in-speeding.html https://www.nytimes.com/2018/01/30/upshot/do-fast-and-furiou... [1] https://public.enigma.com/datasets/lobbying-disclosures-lobbyists-2013/f3ce179f-9171-4754-9f71-71d7596d900a?&filter=%2B%5B%3E%5Blobbyist%5D%5D https://public.enigma.com/datasets/lobbying-disclosures-lobb...
- nahom1 8y agoHere are 100s of more open datasets for anyone to use: https://dataturks.com/projects/trending https://dataturks.com/projects/trending
- mohi13 8y agoHere are 1000s of more open datasets for anyone to explore, use or build upon: https://dataturks.com/projects/trending https://dataturks.com/projects/trending
- loisaidasam 8y agoInspired by this post, I was looking for a fun way to browse datasets randomly, which led me to build this Kaggle Random Dataset Generator: https://news.ycombinator.com/item?id=17313374 https://news.ycombinator.com/item?id=17313374 Thanks Gengo!