6 ms·
A multimodal dataset with one trillion tokens
- Olesya000 2y ago[dead]
- Olesya000 2y ago[dead]
- punnerud 2y agoMore info on the Salesforce blog: https://blog.salesforceairesearch.com/mint-1t/ https://blog.salesforceairesearch.com/mint-1t/
- j7ake 2y agoWow did not expect sales force to be behind this. It’s basically free advertising for technical people to join sales force.
- jszymborski 2y agoSalesforce has long been involved in publishing quality NLP papers, especially during Stephen Merity's tenure. Smerity's papers are some of my favourite. Check out https://ar5iv.labs.arxiv.org/html/1708.02182 https://ar5iv.labs.arxiv.org/html/1708.02182 And my all-time favourite https://ar5iv.labs.arxiv.org/html/1911.11423 https://ar5iv.labs.arxiv.org/html/1911.11423
- nighthawk454 2y agoHey thanks for those Smerity links, hadn't run across his work yet, second one in particular looks great
- jszymborski 2y agoGlad you liked it. How could you go wrong with a paper that starts with > Language has been a thorn in humanity’s side since we evolved a complex enough audio and graphics processing unit to grunt, let alone write cryptocurrency whitepapers or opinion columns.
- 0xDEADFED5 2y agoSalesforce produced one of the best Llama-3(8B) finetunes, IMO: SFR-Iterative-DPO-LLaMA-3-8B-R Hopefully they do something with Llama-3.1
- gotaran 2y agoI’m skeptical of the caliber of talent at Salesforce given the unusable state of their core product.
- paxys 2y agoThe people building CRM software aren't also the ones doing AI research. The two have nothing to do with each other.
- gotaran 2y agoThe ones doing AI research would be working at more prestigious institutions.
- ericjmorey 2y agoWhy make such a ridiculous assumption?
- hluska 2y agoDo you have a point or are you just insulting people out of some twisted definition of fun?
- inkyoto 2y agoSalesforce have been on a shopping spree for quite a few years now. They have purchased MuleSoft (an integration platform) and Slack amongst the others. Salesforce is anything but a CRM software company nowadays.
- optimalsolver 2y agoHow effective would modeling raw byte sequences be, with the individual bytes as the "tokens", and a vocabulary of 256 elements? You could then train on any kind of digital data.
- akrymski 2y agoForget bytes, go for bits. Vocab of size 2. At a theoretical level all of AI comes down to a classifier that is able to predict the next bit given a string of bits. Check out Tsetlin Machines. At some point we will be doing it in hardware. https://byte-gpt.github.io/ https://byte-gpt.github.io/
- kulikalov 2y agoSounds inefficient. It’s like predicting the boiling point of a kettle by measuring the speed of individual molecules of water.
- BizarroLand 2y agoThat would be surprisingly easy with 1st year calculus as long as you were willing to accept a small degree of inaccuracy.
- nodja 2y agoSomewhat inefficient for text, very inefficient for images, specially if you work in pixel space. The max context a model today has been trained is 1M tokens, which takes up a lot of memory. Even if context was not an issue, to generate a 1000x1000 image would take ~3 hours on 100token/s inference. Google has trained an encoder/decoder LLM on bytes called ByT5[1] [1] https://huggingface.co/google/byt5-xxl https://huggingface.co/google/byt5-xxl
- Tostino 2y agoI think the work on multi-token prediction[0] within a single turn could be a significant development that makes byte-level tokenization models more practical. This approach allows the model to predict multiple tokens in parallel, potentially addressing the efficiency concerns raised about byte-level models. By predicting multiple tokens simultaneously, it could significantly speed up inference time, especially for tasks that require generating large amounts of data (like images). This could help mitigate the performance bottleneck mentioned in the parent comment about generating a 1000x1000 image. [0] https://ar5iv.labs.arxiv.org/html/2404.19737 https://ar5iv.labs.arxiv.org/html/2404.19737
- brianjking 2y agoWhat's the license though?
- dpifke 2y agoFrom https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML#license https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML#l...: We release MINT-1T under a CC-BY-4.0 license, designating it primarily as a research artifact. While the dataset is freely available, users are responsible for ensuring its legal use in commercial settings. Users must independently verify compliance with applicable laws before employing MINT-1T for commercial purposes. Same page includes this caveat: Potential Legal and Ethical Concerns: While efforts were made to respect robots.txt files and remove sensitive information, there may still be content that individuals did not explicitly consent to include.
- paxys 2y agoAh yes, the "if you get busted for copyright violations it's not our problem" license.
- sva_ 2y agoDoes it make sense to measure a dataset in tokens? Shouldn't it be tokenizer-agnostic? I.e. the OpenAI tokenizer encodes about ~4 characters per token, but I could also have a tokenizer that does 1 character per token leading to a ~4x increase in token count (relative to the OpenAI tokenizer.)
- anas-awadalla 2y agoHello! Totally agree that tokens will be model dependent. We chose to calculate tokens using the GPT-2 tokenizer as that is a common metric used by other datasets like fineweb. So this should roughly give you a sense of how large the data is in comparison to others. We report other metrics too like number of documents and number of images.
- reverius42 2y agoHow does the GPT-2 tokenizer deal with non-text input? This dataset is multimodal but I thought GPT-2 was text only.
- ks2048 2y agoIt looks like it contains data from CommonCrawl and ArXiv. It's not clear what kind of processing they did, but sometimes these releases seem like just repackaging existing datasets with your name own name on them. It's not hard to get bulk downloads from these sources directly. I thought CommonCrawl truncated files at 1MB. I wonder if the PDFs for CommonCrawl were re-fetched from the URLs. That could be useful if they provide simple way to get those full files.
- anas-awadalla 2y agoHello! Creator of MINT here. We do a lot of pre-processing of commoncrawl (which in its raw form isn’t all that useful for training models). This includes heuristics to remove low quality text and images and deduplicating documents, paragraphs, and images. All of these are crucial to achieve good training performance. On your point regarding PDFs, we actually don’t constraint ourselves to the 1MB files and do our own downloading of PDFs!
- ks2048 2y agoI see. Thanks for the reply. I opened one of the tar files and see now how it has extracted the text into json files.
- supermatt 2y agoI havent trained any LLMs, so please accept my comment with all the naivety with which it is given - but in the "examples of MINT multimodal documents" graphic at the top of the README, it feels to me as though the labeling (for the images on the left) couldn't be much worse? Is this normal for these datasets? How are we able to build such powerful models with such poor quality data?
- zsyllepsis 2y agoI think the labels could be much, much worse. They could contain straight noise, just completely random text - not even words. They could also contain plausible, factual text which otherwise has no relationship with the text. I think most commonly image datasets like this consist of images and their captions, with the presumption that the content author had _some_ reason of associating the two. The goal of the model is to learn that association. And with a _lot_ of examples, to learn nuanced representations. In the third image, for example, we see some kind of text on a material. The caption mentions "Every year he rides for someone we know, touched by cancer". Perhaps the model is fed another example of bicycle races, with similar imagery of racing bibs. Perhaps its fed another of a race that specifically mentions it's a charity ride to raise money for cancer. Perhaps.... You get the idea. Alone, each example provides only vague connections between the image and the caption. But when you have a ton of data it becomes easier to separate noise from a weak signal.
- whiplash451 2y agoDeep learning is robust to massive label noise [1] Not to say that data quality does not matter, but these noisy sets are still very useful. [1] https://arxiv.org/abs/1705.10694 https://arxiv.org/abs/1705.10694
- sigmoid10 2y agoMinor correction: Deep learning using gradient descent is incredibly robust to noise. If you know the mathematics, this also makes sense intuitively: gradients of incorrect labels will generally point in random directions, whereas the "truth" points in a specific direction (and I explicitly mean truth in the sense of what is portrayed consistently as fact in the dataset, not the real world truth). So when you accumulate gradients, you will end up with a net effect that moves weights only towards the consistent answers. Since gradient descent is by far the most popular algorithm, it's easy to conflate these two things. But there are other approaches that don't treat noise so well.
- EGreg 2y agoLicense: None Means we can’t legally use it?
- thomashop 2y ago```We release MINT-1T under a CC-BY-4.0 license, designating it primarily as a research artifact. While the dataset is freely available, users are responsible for ensuring its legal use in commercial settings. Users must independently verify compliance with applicable laws before employing MINT-1T for commercial purposes.```
- benreesman 2y agoSalesforce quietly does some truly tier-one stuff. They don’t showboat it which makes them seem more, not less, serious at least from my seat. They use Bazel and shit, which is an acid test for being professionals, it’s a real shop. The Magnificent 7 are about to get the taste slapped out of their mouth by skittish momentum guys and their chattels on Sand Hill Road. I look forward to the space this week will create for shops like Salesforce.
- wsc981 2y agoSo, I read the blog post and checked the Github page, but not a clear picture here for me. I am still kinda new to the LLM space. What would the use-case be for this model? What are the advantages over something like Llama?
- Tepix 2y agoIt's a dataset to train models, not a model.
- naveen99 2y agoCopyright and intellectual property are directly at odds with these types of efforts, and has been losing to linux, gnu, github, wikipedia, mit open courseware, youtube, LLMs and their datasets. But copyright did slay Napster, PirateBay, anna’s archive etc…
- larodi 2y agoThis all be quite dated in 10-20 years now. Common information will be free as it was in the 90s, but valuable information will then probably cost even more. And 99.9% times illegal to obtain or possess.
- littlestymaar 2y agoUntil individual countries start realizing that protecting copyright is costing them lots of potential economic growth coming from IA, and the IA business start lobbying more than the copyright business, at which point the law would just change. Intellectual property is a fairly recent invention in economic history, and it only happened because it benefited the elite. If the balance of power changes so will the law.
- layer8 2y agoIA?
- littlestymaar 2y agoAI obviously, a typo whose likelihood is increased by the fact that it's spell “IA” in my language.
- aswegs8 2y agoObviously
- thesz 2y agoIntellect Amplifier?
- stealthcat 2y agoMarketed as “multimodal” but actually texts and images. Multimodal dataset should be multimedia: text, audio, images, video, and optionally more like sensor readings and robot actions.