7 ms·
DeepSeekMath 7B achieved 51.7% on MATH benchmark
- mdp 3y agoRelated paper - https://arxiv.org/pdf/2402.03300.pdf https://arxiv.org/pdf/2402.03300.pdf
- rgbrgb 3y agoSupports commercial use! Interesting what's unsupported: - In any way that violates any applicable national or international law or regulation or infringes upon the lawful rights and interests of any third party; - For military use in any way; - For the purpose of exploiting, harming or attempting to exploit or harm minors in any way; - To generate or disseminate verifiably false information and/or content with the purpose of harming others; - To generate or disseminate inappropriate content subject to applicable regulatory requirements; - To generate or disseminate personal identifiable information without due authorization or for unreasonable use; - To defame, disparage or otherwise harass others; - For fully automated decision making that adversely impacts an individual’s legal rights or otherwise creates or modifies a binding, enforceable obligation; - For any use intended to or which has the effect of discriminating against or harming individuals or groups based on online or offline social behavior or known or predicted personal or personality characteristics; - To exploit any of the vulnerabilities of a specific group of persons based on their age, social, physical or mental characteristics, in order to materially distort the behavior of a person pertaining to that group in a manner that causes or is likely to cause that person or another person physical or psychological harm; - For any use intended to or which has the effect of discriminating against individuals or groups based on legally protected characteristics or categories.
- ronsor 3y agoThe irony is that anyone who was going to do those things isn't going to care about a license anyway.
- BossingAround 3y agoTrue, but at least the author wouldn't be liable.
- zamadatix 3y agoThe MIT license covers liability more broadly and tightly in a single paragraph.
- declaredapple 3y agoYou either wouldn't be liable anyway (not responsible for what people use it for) or will still be held liable (took no measures to prevent malicious use).
- austin-cheney 3y agoWho would qualify as a user that seeks to violate applicable laws and yet is somehow identified as an official part of some legally recognized military? Furthermore, how would anybody know? As a dumb Army guy if I were doing military research I would just keep it on my private military internet that does not exist for non-military users.
- wand3r 3y agoIts virtue signaling. I know its over used, but seriously, who is intentionally harming minors BUT unwilling to break a ToS contract?
- nurettin 3y agoThe Devil?
- zeusk 3y agoMeta/Instagram probably
- austin-cheney 3y agoProcessed food vendors come to mind, but I get your point.
- paulddraper 3y agoFacebook lawyers
- godelski 3y agoDoes anyone know how much spoilage are in these datasets? Common crawl has a lot of websites in it, including Reddit and Stack*. I'm certain there are lots of questions in those datasets and we want to differentiate recall from problem solving (often confused). I have a deep distrust when using large datasets like this given a common one with 60 authors assumed writing leet code style programs by hand would mean they wouldn't appear in the training data (github) and didn't even bother to check. It's really hard to sanitize datasets of this size and deduplication is a much harder task than many realize. https://arxiv.org/abs/2107.03374 https://arxiv.org/abs/2107.03374 https://arxiv.org/abs/2303.09540 https://arxiv.org/abs/2303.09540
- ianbutler 3y agoSome have a lot and those models are often ignored (except by lay people or hobbyists which is a different problem), but many serious submissions from serious groups for benchmarks like this check for contamination to specifically avoid the problem you’re suggesting. Process for decontamination has been outlined by many groups so you can often check it out.
- godelski 3y agoSo I am a ML researcher. Note that part of my comment is specifying how difficult it actually is to ensure lack of spoilage. The second paper I link is actually a pretty good proof of this. Though I'll say that I wish they had been a bit more explicit about how a random pruning significantly improves results. Because that is quite the result in of itself, given that the datasets they look at are already filtered. Dedupe is fucking hard. So I'm not looking for a handwavy "trust me" I'm looking for the explicit vetting processes applied to these specific datasets. It's incredibly important to know the limits of your tools and that includes datasets (as well as metrics).
- deleted 3y ago[deleted]
- ianbutler 3y ago
- deepseekfake 3y agoI have spoken to team members, and they all say the results of this and coder are very, very much leakage (no suprisse given the result!!)
- godelski 3y agoThat's good to know, and better to admit. Earns a lot of respect, at least for me. Recall is still a pretty useful task. I just wish more would be less afraid to admit spoilage.
- pclmulqdq 3y agoThere's a good chance that's also true for GPT-4 given how they train. Without known completely new evals, it's hard to say that any LLM benchmark results aren't leakage.
- CuriouslyC 3y agoIf you're trying to prove the model has reasoning abilities, ask it the question in a language other than English, even better give it multiple sentences in different languages and tell it to answer the question without first translating the sentences.
- godelski 3y agoThat's not a great metric and is going to be incredibly language dependent. For example, the European languages all have a lot of similarities and so it should be unsurprising that a model trained on English can do pretty well on French and German. But then if you are to ask it a language that is fairly disjoint (say Chinese) then you are held back by the lack of language data from that dataset (or you run into the exact same issue as previously). It's definitely a metric worth trying, but we also must recognize the extreme limits of it too. Evaluation is quite difficult and the better our models perform the more difficult evaluation actually becomes. Anyone saying otherwise is likely trying to sell you something.
- godelski 3y ago