4 ms·
Institutional Books: A 242B token dataset from Harvard Library's collections
- strangecasts 1y ago[flagged]
- rickydroll 1y ago[flagged]
- gshubert17 1y agoEdit: Two responses, https://news.ycombinator.com/item?id=44252450 https://news.ycombinator.com/item?id=44252450 and https://news.ycombinator.com/item?id=44252408 https://news.ycombinator.com/item?id=44252408, seem to be dupes. As rickydroll states, the time stamps and id numbers show it to be the first.
- rickydroll 1y agoIt's a copy of mine. Look at the timestamps.
- rudedogg 1y agoI meant it as a joke - to blatantly steal your comment since you said copyright is evil.
- rickydroll 1y agoSo you were cosplaying an LLM trained on my comments :-) Good one.
- 01HNNWZ0MV43FF 1y agoThe training sets should be public then
- rickydroll 1y agoYes they should
- deleted 1y ago[deleted]
- ks2048 1y agoSeems a strange comparison - I don't think anyone claims "search engines" should be a repository of cultural memory.
- rickydroll 1y agoNobody intended "search engines" to be a repository of cultural memory. They became that because they were built on content and information that encompassed cultural memory, and people used them for that purpose. How many times have you told someone to Google that instead of giving them a URL? Training sets are currently built on the same information, and now chatbots are a different way to query for that information. So, in the same way as with search engines, chatbots have become another repository of cultural memory. At some time in the future, people will come to believe that if it's not in a search engine or a chatbot, it doesn't exist, which to me is why it's vital to put everything we know into a training set in addition to archiving it someplace that will survive a Carrington-level blast from the sun. IMO, making multiple copies of archives of everything we know supersedes copyright.
- rudedogg 1y ago[flagged]
- SloopJon 1y agoAlthough this is characterized as 1.0, it is governed by the Terms of Use for Early-Access, which are quite limiting, including: "You may use the Service solely for noncommercial purposes."
- ninjin 1y agoIt really is rather peculiar to me. They frame it like this (emphasis mine): "With the preliminary publication of this dataset, we further seek to establish a community-led process to grow, improve, and use institutional data in ways that strengthen the knowledge ecosystem and assert the importance of ongoing stewardship of training data from the originating knowledge institutions themselves. To this end, we are experimenting to find the best way to release this data in a manner that facilitates collaboration. We encourage input on this process to guide the full publication of this and future dataset dataset releases, beginning with the following decisions: * At preliminary launch, we have published the metadata, including experimental metadata, in full for anyone to access and use. * At preliminary launch, we have published the dataset including OCR-extracted text under a noncommercial license, and with a 'click-through' that requires users to accept this license, additional terms of use, and to share basic contact information with us so that we can engage the community in its early use. * At preliminary launch, we have chosen to postpone the release of the raw scan images, though we will share them liberally with researchers and libraries who wish to review them. While we know AI developers and researchers are eager for more raw materials, we believe this minor friction can help build the relationships and norms necessary to grow a collaborative community." It is the fruit of their labour (well, the digitalisation is), so it is up to them to license it as they see fit. But it feels odd to me that they seem to want to be in control to this degree. In open source and my own research field, the pattern we tend to follow is to release freely, observe, and then build relationships rather than holding a "license gun" towards the head of potential collaborators. Lastly, I have only skimmed the pre-print, but I noted no commitment to a final license either. Not even a direction for it. Thus, as a natural language processing researcher I will stay clear from this dataset for the time being and hope the licensing situation improves.
- ks2048 1y ago“noncommericial” seems pernicious to me lately. I can see why people reach for it, but it really seems hard to define (there are many ways to profit off of something without simply selling it directly).
- timhigins 1y agohttps://huggingface.co/datasets/institutional/institutional-books-1.0-metadata https://huggingface.co/datasets/institutional/institutional-... https://huggingface.co/datasets/institutional/institutional-books-1.0 https://huggingface.co/datasets/institutional/institutional-... https://github.com/instdin/institutional-books-1-pipeline https://github.com/instdin/institutional-books-1-pipeline https://www.institutionaldatainitiative.org/institutional-books https://www.institutionaldatainitiative.org/institutional-bo...
- Frummy 1y agoAIs lizard brain will be 60% 1800s apparently, it might act like a villainous steampunk anglosaxon twirling a mustache in moments of survival, or at least some blend of those values while playing 5d chess. Read it H G Wells "World brain" to calm it down like a fond childhood memory
- DaSHacka 1y agoThis would be the funniest possible future, and a very distinct possibility depending on how the NYT lawsuit turns out in regards to IP holder rights versus AI "copyright laundering".
- deleted 1y ago[deleted]
- adt 1y agohttps://lifearchitect.ai/datasets-table/ https://lifearchitect.ai/datasets-table/