4 ms·
Thanks for the feedback. This is likely a "not quite yet" vs. "never". Definitely understand the motivations from a user standpoint for not needing to login to
by benhamner 9y ago
Thanks for the feedback. This is likely a "not quite yet" vs. "never".
Definitely understand the motivations from a user standpoint for not needing to login to download.
There's some non-obvious benefits we get as a small team by requiring login, in addition to new user growth. Bandwidth for hosting data can be large, and it's easier to reason about and prevent abuse in the context of authenticated users.
We do enable previewing the dataset while logged out, and the preview functionality will become more full-featured.
- mynewtb 9y ago> new user growth That sounds like you are inflating your user counts with garbage accounts.
- has2k1 9y agoAs a company that deals with data analysis, Kaggle can surely tell between different levels of user participation. I interpret it charitably as "some fraction of users forced to register in order to download data go on to become active users of the platform".
- random4369 9y ago> As a company that deals with data analysis, Kaggle can surely tell between different levels of user participation. Kaggle outsources data analysis to bright students willing to work for a tiny fraction of their actual worth. Their business model has nothing to do with their own analytics talent.
- has2k1 9y ago> Their business model has nothing to do with their own analytics talent I recall — some years ago — when they were advertising for Jobs, some of the criteria for potential technical fitness was a high placement in some of the competitions. Plus, I have engaged in a few of the competitions and you have to have a team of capable data scientists to design them. They are not clueless people.
- prepend 9y agoI’m sure they can. But they can also choose not to distinguish a high quality user from a low quality one when promoting their site. For example, this post - http://blog.kaggle.com/2017/06/06/weve-passed-1-million-members/ http://blog.kaggle.com/2017/06/06/weve-passed-1-million-memb... - talks about a million users. But does not say how many are low quality, inactive ones. Or even describe how they filtered out low quality users from their analysis.
- omg_ketchup 9y agoWelcome to marketing?
- hyperocular 9y agoA torrent option would go a long way to offsetting hosting costs on more popular items. Also I appreciate the preview option, it's good to see what is in a dataset before committing to downloading and extracting hundreds of gigabytes.
- gertye 9y agoAFAIK, Kaggle is a part of Google, therefore is not operated independently by a "small team". You just actively try not to be associated with the behemoth.
- solarkraft 9y agoKaggle has been a part of Google for about a year. https://techcrunch.com/2017/03/07/google-is-acquiring-data-science-community-kaggle/ https://techcrunch.com/2017/03/07/google-is-acquiring-data-s... I do wonder why they're only giving it a "small team" when they could give them lots of resources and advanced abuse protection - though this could just be an allocation issue, it seems like a wasted PR opportunity.
- prepend 9y agoPart of the best practices of public data involve unrestricted downlod - https://project-open-data.cio.gov/principles/ https://project-open-data.cio.gov/principles/ There are certainly benefits to you as a host if you restrict access to as a host. But if you have bandwidth concerns, perhaps just link to the source. As it is now, I may be confused that you are using open data as a way to drive user growth.
- QasimK 9y agoThanks for your response. I spoke slightly evocatively because this has caused issues for me and hence I didn’t like to see it advertised as “public”. The biggest issue is that I cannot share my work with others (easily) if I use a dataset from Kaggle. For example, ideally I just want to have a notebook online somewhere which can be instantly downloaded and run by anyone. Having to (automatically) download a dataset is a hinderance, and having to create a kaggle account first is an outright blocker. On the other hand using, for example, IPFS or a torrent would be better, because you can reference the dataset using a global identifier and anyone can easily get access to it.
- shepardrtc 9y agoI disagree with the parent. You've taken the time to organize and host these datasets. The least we can do is create an account to download them.
- diggan 9y agoSure, but then maybe don't call it "public" as people will think it can be downloaded without creating an account. Calling something public but require account creation is misleading.
- omg_ketchup 9y agoI'm not sure I'd agree- the data is accessible to anyone who wants it for free. There's no restrictions on creating the account. Radio is free, but you hear advertisements. Probably better to create an account than to have product placement in the actual data.
- diggan 9y agoThe restriction I'm talking about is creating the account. "Public" (at least for me) does not mean that I need to agree to some lengthy "terms of service", "privacy policy" and create an account. Public means it's public and can be accessed from curl or my browser of choice without signing a contract. Not sure were you are from, but where I'm from (Sweden), public radio (not free but public) and public TV is free of advertisement and does not require me to sign up for an account to be able to listen to/watch it. That's what I call public. - https://www.kaggle.com/terms https://www.kaggle.com/terms - https://www.kaggle.com/about/privacy https://www.kaggle.com/about/privacy
- shepardrtc 9y agoWhen Kaggle is saying "public" dataset, they're implying the origin of the dataset is public. Meaning, the datasets were created by various groups/companies/institutions and made available for the general public. Kaggle is simply hosting them again. They're doing us a service by organizing them all into one location and eating the bandwidth costs. My argument is that in return for that service to us, the least we can do is create an account with them.