100 ms·
Facebook LLAMA is being openly distributed via torrents
- mdaniel 4y agoI was expecting it to be a newly created GitHub account, but no, seems they're willing to roll the dice on whatever the outcome is from this
- EamonnMR 4y agoGonna be interesting to see if Facebook tries to tell people they can't use this because it's stolen (when it was presumably built using data taken without permission.)
- ed 4y agoUnlike many llm’s this was trained using public training sets (and cited in their paper), to let anyone with the $$$ independently generate the weights
- zoranzv 4y agoGood app
- hsuduebc2 4y agoCan you point me a way to run this locally please?
- ema_dev 4y ago[dead]
- progbloging 4y agocool
- deleted 4y ago[deleted]
- rnosov 4y agoJust to make it clear, does this torrent include model weights?
- generalizations 4y agoIt contains weights for all four model sizes, apparently. This definitely saves on bandwidth costs. :)
- WithinReason 4y agoFolder structure for the 2 smaller models look like this: LLAMA │ tokenizer.model │ tokenizer_checklist.chk │ ├───13B │ checklist.chk │ consolidated.00.pth │ consolidated.01.pth │ params.json │ └───7B checklist.chk consolidated.00.pth params.json
- HarHarVeryFunny 4y agoSo what is content of those various files? Does this include the full models themselves, or just the weights ?
- WithinReason 4y agoThe pth file seems to be a model and weights, saved as described here: https://pytorch.org/tutorials/beginner/saving_loading_models.html#save-load-entire-model https://pytorch.org/tutorials/beginner/saving_loading_models... .chk file is am md5 hash of the file, the .json file contains this for the 7B model: {"dim": 4096, "multiple_of": 256, "n_heads": 32, "n_layers": 32, "norm_eps": 1e-06, "vocab_size": -1}
- HarHarVeryFunny 4y agoThanks, so from that PyTorch doc it seems that pickle format has the filenames of the model classes, but not the classes themselves. I'm sure someone will figure it out though!
- Laaas 4y agoWhat's the point of the form if it's freely accessible? This might be revolutionary in the LLM field, as Stable Diffusion was to DALL-E.
- rnosov 4y agoThe linked page is just a pull request, the actual repository readme doesn't mention torrent option at all.
- adamgordonbell 4y ago[flagged]
- ot 4y agoIf they were a Meta employee they wouldn't have had to sign the CLA.
- aaronharnly 4y agoI wouldn’t jump to that conclusion — the first and last names are not uncommon, and the GitHub user has some attributes (eg Haskell; geographic ties) they do not share with the LinkedIn profile you link. Unless you have strong evidence of their identity I’d suggest rethinking this.
- fwlr 4y agoThe Christopher King user who submitted the pull request has a git repo in C++ called “Final - My Homework” made in 2015; the Christopher King you link to completed his BA in Computer Science in 2005. I strongly suspect they’re different people who happen to share the same name.
- fwlr 4y agoThe user who submitted the pull request is not part of Meta or Facebook Research, and the users who signed off on reviewing the changes don’t appear to be either. I highly doubt Meta will approve the pull request. The models are being distributed by torrent by someone with access to the models, not by Meta themselves as far as we know. They likely still intend to distribute via the form. This is just someone publicizing the torrent link by being cheeky on GitHub. (As they didn’t reply to my request for the model - I specified it was for personal use and my use case was “I think it would be fun to run it on my own hardware” - I appreciate this little stunt a great deal!)
- KaiserPro 4y agoold school opensource, which is a bit surprising from meta. I wonder how they managed to square that with legal. Someone must have been very good friends with Zuck.
- papruapap 4y agois it?[0] The worst offender is AMZ, all the rest big tech are pretty open-source friendly. 0:https://opensource.fb.com/projects/ https://opensource.fb.com/projects/
- LilyFrenchPants 4y ago> old school opensource, which is a bit surprising from meta Aren't you a cheeky lad? Metea turned out lots of open-source database systems: * RocksDB * Hive * Presto * Cassandra * Velox LFP
- KaiserPro 4y agoI should be more precise: Getting anything that could produce, look like, or smell anything like misinformation out of meta is very hard (for good reason!) My friends have had repeated push back for various papers because they are ML based and could be in the same room as something that could possible be used by miscreants. And here we have a LLM that can spit out all sorts of things that are misinformation like. If their department tried to launch something like Galactica they would have been slapped down and told to think again about what they were doing in life.
- havkom 4y agoIs this warez?
- marginalia_nu 4y agoSure I'll download TeamMysticAvengers-meta-llm-x-cars-movie-model-x-angelina-jolie-naked-xxx-2023.zip.exe.torrent
- ocimbote 4y agoThanks for this late 1990s moment. My back stopped hurting while I was reading this :p
- grog_tremor 4y agoJust a sec, need to find the crack on astalavista
- napsterbr 4y agoOr a keygen with this soundtrack: https://youtube.com/watch?v=foYc1cVkyKk https://youtube.com/watch?v=foYc1cVkyKk
- optymizer 4y agokeygens with music! How could I forget. Thanks for reviving some good old memories.
- jasongill 4y agoThis comment took me back, haven't been to astalavista.box.sk in a LONG time - looks like the site is still online in some form (but without the old black and green color scheme)
- repple 4y agothis, and the soundtrack in the sibling reply is giving me fantastic series of flashbacks. thank you
- deleted 4y ago[deleted]
- KierPrev 4y agoWhy is Meta open sourcing its AI through torrent? Or am I understanding it all wrong
- rnosov 4y agoit's a pull request from ChristopherKing42. He is unlikely to be associated with Meta.
- drbscl 4y agoSending via HTTP will incur bandwidth costs. Torrents massively reduce this cost in the long run by making it P2P. Edit: maybe in this case it's a leak though
- Technotroll 4y agoTorrent can also be really, really fast with enough seeders, even giving CDNs a run for their money.
- deleted 4y ago[deleted]
- IceWreck 4y agoTheyre giving it to universities for free. Someone got access and then made PR with a link to the torrent
- gorbypark 4y agoIt seems like the model has been leaked (not by Meta) and is being distributed via a torrent. Someone has created a PR to the repo as a joke, suggesting that instead of filling out a form and waiting to be granted access (which is the official way to get access to the model), that you could just download it via the torrent.
- RobotToaster 4y agoFor those who didn't check the github discussion, I don't think this pull request came from a Facebook employee, lol.
- ot 4y agoIn case it's not clear what's happening here (and from the comments it doesn't seem like it is), someone (not Meta) leaked the models and had the brilliant idea of advertising the magnet link through a GitHub pull request. The part about saving bandwidth is a joke. Meta employees may have not noticed or are still figuring out how to react, so the PR is still up. (Disclaimer: I work at Meta, but have no relationship with the team that owns the models and have no internal information on this)
- deleted 4y ago[deleted]
- IanCal 4y agoIt's not even clear someone has leaked the models. A random person has put a download link on a PR, it could be anything.
- sebzim4500 4y agoIs there anything stopping anyone from using this for commercial purposes? I know that when you fill in the google form you need to agree to noncommercial use, but someone downloading this will never have agreed to that licence agreement.
- injidup 4y agoI don't know. Is there anything stopping you using the latest Miley Cyrus album for commercial purposes if you downloaded it via torrent and never agreed to any licencing terms?
- RobotToaster 4y agoIANAL, but I imagine it's a legal grey area if the weights can be copyrighted? Works produced by purely mechanical means don't normally meet the threshold of originality.
- londons_explore 4y agoand, copyright rarely bites if you use something without publishing/redistributing it. It would be like playing copyrighted music in your office without permission. Perhaps technically illegal, but your customers will never know what music your Devs were listening to...
- colinmorelli 4y agoI am quite confident that using this model for commercial purposes will, if detected, land you in quite a legal quagmire that almost certainly sides in favor of Meta. And even if it did not, Meta certainly has a more capable legal team with more cash to spend than the average HN user.
- basch 4y agoHas any photo or art ever been found to have been illegal due to pirated photoshop?
- 4y ago
- kaszanka 4y agoHere is the magnet link for posterity: magnet:?xt=urn:btih:ZXXDAUWYLRUXXBHUYEMS6Q5CE5WA3LVA&dn=LLaMA
- q1w2 4y agoGreat, now how do I run it? Do I need a GPU with over 65GB RAM?
- rnosov 4y agoGenerally, you'll need multiply model size by two to get required amount of video RAM. There are 4 sizes, so you might get away with even smaller GPU for say 13B model.
- version_five 4y agoTry this, it's for running llms that won't fit in the gpu: https://github.com/FMInference/FlexGen https://github.com/FMInference/FlexGen
- gpm 4y agoCurrently that looks like it only supports facebook's opt and galactica models. Though they do appear to plan to add support for more models.
- bioemerl 4y agoNope, more like 111gb
- psychphysic 4y agoThanks not working for me... Not that I could run it if I downloaded it.
- 4bpp 4y agoSince the point seems to be lost on some of the early commenters, this appears to be a cheeky PR by someone unaffiliated with Facebook, suggesting that they put a magnet link to (what seems to be) a leak of the model weights along with the previously existing invitation to apply to receive them on their own page.
- Manjuuu 4y agoThat guy does not seem to have anything to do with Facebook... interesting.
- Beaver117 4y agoFunny. iirc some of the big tech (I think it was Google?) use torrents internally to deploy very large images to servers. Piracy is not the only use case!
- jpgvm 4y agoIronically that is Facebook that used torrent for binary distribution. (no idea if it's still the case, that was a very long time ago).
- ithkuil 4y agoIt's not just food very large images. It's also useful for moderately large images/packages being deployed to many many many servers.
- regularfry 4y agoIt was used for years to distribute World of Warcraft updates. No idea if it still is.
- EamonnMR 4y agoUsed it to download linux distro images back when the size of an install CD was huge. Good times.
- gpm 4y agoNow that I think about it I wonder why we don't see it being used to distribute packages for linux distros. Seems more flexible than the current mirror system.
- Retr0id 4y agoThere seem to be a lot of confused commenters here. This is the content of an as-yet-unmerged pull request, and presumably not something that Facebook approves of.
- fancyfredbot 4y agoopt-175B weights are already openly available as I understand. Hugging-face also has openly available weights for a 176B parameter LLM called Bloom. Is LLAMA offering something over and above these?
- rihegher 4y agoaccording to Facebook Llama beats GPT3 on multiple benchmarks with smaller models that can be fine tune on a single A100 GPU EDIT: correcting the type of GPU
- px43 4y agoYeah, their recent papers show the smaller LLAMA models outperforming the major LLMs today, and they also have bigger models. This isn't just an alternative, it's a multi order of magnitude optimization. https://aibusiness.com/meta/meta-s-llama-language-model-outperforms-openai-s-gpt-3 https://aibusiness.com/meta/meta-s-llama-language-model-outp...
- koheripbal 4y agoCan I spend $5K and run it at home? What GPU(s) do I need?
- gpm 4y agoIn principal you can run it on just about any hardware with enough storage space. It's just a question of how fast it will run. This readme has some benchmarks with a similar set of models (and the code has support for even swapping data out to disk if needed): https://github.com/FMInference/FlexGen https://github.com/FMInference/FlexGen And here are some benchmarks running OPT-175B purely on (a very beefy) CPU machine. Note that the biggest llama model is only 65.2B: https://github.com/FMInference/FlexGen/issues/24 https://github.com/FMInference/FlexGen/issues/24
- px43 4y agoAs the models proliferate, I guess we'll be finding out soon. The torrent has been going pretty slow for me for the past couple hours, but it looks like there are a couple seeders, so eventually it'll hit that inflection point where there are enough seeders to give all the leechers full speed downloads. Looking forward to the YouTube videos of random tinkerers seeing what sort of performance they can squeeze out of cheaper hardware.
- deleted 4y ago[deleted]
- aent 4y agoFor anyone wondering, it includes 4 models: 7/13/30/65 billion parameters, the smallest one is 14Gb, the largest one is 131GB, all four are 235Gb.
- mlboss 4y agoIs it possible to run the smallest one on a consumer gpu with 24gb ram ?
- Tepix 4y agoRunning it is easy but you'll probably want to finetune it, too
- rihegher 4y agoI would be surprised if you can't. The smallest weight file is 14gb apparently
- eightysixfour 4y agohttps://github.com/facebookresearch/llama/blob/main/FAQ.md#3 https://github.com/facebookresearch/llama/blob/main/FAQ.md#3 Looks like it needs 14gb for weights and it isn't clear what the minimum size for the decoding cache is, but it defaults to settings for 30gb GPUs.
- MacsHeadroom 4y agoIn int8 7B needs only 9GB of VRAM and 13B needs only 20GB on a single GPU. https://github.com/oobabooga/text-generation-webui/issues/147#issuecomment-1454841458 https://github.com/oobabooga/text-generation-webui/issues/14...
- MacsHeadroom 4y agoYou can do even better!. You can run the second smallest one (better than GPT-3 175B) on 24GB of vram, ie LLaMA-13B. https://github.com/oobabooga/text-generation-webui/issues/147#issuecomment-1454841458 https://github.com/oobabooga/text-generation-webui/issues/14...
- version_five 4y agoAre there any official checksums available? I'm happy to see this, even if it's an unsanctioned stunt, because I think it's really pathetic of meta to want to gatekeep their "open" model. But ML models generally can execute arbitrary code, I'd want to make sure it's the real version at least.
- zb3 4y agoBut ML models generally can execute arbitrary code Is it the case if we're only talking about weights? I thought the rest is actually "open".
- px43 4y agoMy understanding is that weights are normally stored as pickled python blobs, which means arbitrary code execution as they are unpickled.
- Yoric 4y agoIf it's PyTorch, it can definitely contain and execute arbitrary code. One of the reasons I'm not a huge fan of PyTorch.
- londons_explore 4y agoThey could contain arbitrary code... But typically do not. That means that with the right viewer application it will be trivial to know for sure. It isn't like a multi gigabyte game for example, where knowing if there is any malicious code could easily be a multi-month reverse engineering project to get to the answer of 'probably not, but we don't have time to check every byte with a fine tooth comb'
- zb3 4y agoI only found this picklescan[0] serving this purpose, but it doesn't seem to be a finished project. [0] - https://github.com/mmaitre314/picklescan https://github.com/mmaitre314/picklescan
- kif 4y agoI wonder what the memory requirements would be to run such a large model. I'd love to be able to run this model, alas my MacBook can barely run toy models.
- px43 4y agoHell, I'd love to be able to buy a $30k server to run these models. I think to run BLOOM required something more along the lines of a $200k server.
- DeathArrow 4y agoNo need to spend $30k, use Azure or AWS.
- wincy 4y agoYep. It’s expensive to spin up an A100 80GB instance but not THAT expensive. Oracles cloud offering (first thing to show up in google search I know you probably won’t use them and it seems extra expensive) is $4.00 per hour. If you are motivated to screw around with this stuff there’s definitely options.
- gnramires 4y ago> If you are motivated to screw around with this stuff there’s definitely options. Erm, for inference that is. Training is definitely out of question for individuals I believe (unless you use much smaller models?).
- dilyevsky 4y agoGCP spot price for A100 80g gpu is only $1.25 and they give you $300 of credit when you open a new acc
- dragonwriter 4y agoUnless its for something you want to happen whenever and don't mind be dumped in process, shouldn't we look at on-demand, not spot, prices?
- controversial97 4y agoThe torrent is 224GB total, a load of 13 to 16GB .pth files
- ddtaylor 4y agoFWIW this information was already freely available via DHT scrapers like btdig [1] I think everyone at Facebook knows that torrents aren't secret and the Google form is basically a legal tool to shield them from liability while making litigation against anyone misusing the model easier. [1]: https://btdig.com/b8287ebfa04f879b048d4d4404108cf3e8014352/llama https://btdig.com/b8287ebfa04f879b048d4d4404108cf3e8014352/l...
- londons_explore 4y agobtdig blocked in the UK and many other countries. Use a USA VPN for access.
- sebzim4500 4y agoI'm in the UK and can view that link without a VPN.
- Retr0id 4y agoConsider yourself very lucky that your ISP doesn't suck. I'm in the UK and the link won't load, due to TLS SNI filtering (Virgin Media).
- throwaway3245 4y agoI'm surprised BT is fine to view but not Virgin.
- mavhc 4y agoTOR browser fixes all
- Retr0id 4y agoNot really.
- 4y ago
- pavlov 4y agoIt’s interesting that these models are both massively expensive to produce and self-contained to a degree that you can distribute the end product in a torrent. This has not been the case for most commercial software for the past 20 years, during the cloud era. If you could steal a dump of random Facebook source code, it would be 99% useless because it’s so closely tied to the infrastructure. There’s almost nothing you could usefully run on your own PC or server VM. But these ML models are like neutron stars of computation density. You can’t really peek inside to see what’s going on either. An unknown stolen model’s properties would need to be discovered by experimentation.
- heap_perms 4y ago> are like neutron stars of computation density I really like that expression.
- seydor 4y ago> massively expensive to produce and self-contained to a degree that you can distribute the end product in a torrent. So, like movies or software
- phyalow 4y agoOr Microsoft Office...
- justinclift 4y agoOr Linux distro's.
- bayindirh 4y agoOr BSD distros.
- Brian_K_White 4y agoThese are actually trivial and silly examples. The bulk of really valuable commercial code is not self contained or portable like those. Where is the torrent with a runnable copy of paypal, or amazon?
- Felminor 4y agoWas to expected. Anyhow I do remember a post of a person stating this will never happen but it's just a web form and request for describing of what type of research you do Of course it will be leaked
- popcorncowboy 4y agoYeah, Meta must have had a plan for "when this gets leaked" because they put up only the flimsiest of foils. As per other comments the most likely is simply that they could shield themselves (and plausibly litigate with grounds) while ensuring that the model escapes into the wild to wreak its chaos against MS (OAI) and big G. This way they can see what's what from the safety of their shielded bubble and make a more informed call about changing the license to something more permissive if it looks like the strategic wins against their enemies would be worthwhile. Win win win. (Except for the leaker, that was an unfortunate own goal, they're going down).
- wunderland 4y agoIn case it was unclear, the person who submit the pull request does not work for Facebook and is teasing them here.
- LoveMortuus 4y ago~220 GB :O That's quite big!
- londons_explore 4y agoNeeds ~200GB of graphics ram to run... Not many people will get this running!
- 988747 4y agoWhat do you mean "big"? fits on the average laptop :)
- eigenvalue 4y agoI'm not surprised-- I recently suggested that someone might try to pull an Aaron Swartz with the LLAMA weights (i.e., release them in an uncontrolled way similar to how Aaron attempted to release the JSTOR database). It's quite misleading for FB to claim that they are being so open, but then hoard the weights and only release it to a few academics. If the paper is to be believed, this is a major development, allowing you to get close to GPT3 performance on a single GPU (at least for inference on the smallest model). Clearly some renegade academic feels the same way.
- sebzim4500 4y ago>Clearly some renegade academic feels the same way. Or someone pretending to be a renegade academic. It's not like there is a KYC process.
- catchnear4321 4y ago> It's quite misleading for FB to claim that they are being so open, but then hoard the weights and only release it to a few academics. I mean at least they didn’t pick a name that heavily implied they were, are, and always will be open. Then do the opposite. You know, like OpenAI? So now we got some weights I guess.
- vintermann 4y agoIt was already the most open language model in its class, given that the code for training and inference was available and it only used public data for training. For Google and OpenAIs offerings, have fun reimplementing it from descriptions in the paper (including small crucial details that they may have left out), training it for a month, and then wondering if the implementation or the training data is the reason your model isn't as good as theirs.
- ImprobableTruth 4y agoThe training code is not available.
- 4y ago
- pimterry 4y agoI give it a week before we see tools for subtly watermarking your secret LLM's weights, so you can trace leaks like this later.
- londons_explore 4y agoWatermarking the weights is trivial. Watermarking the output is also possible, but more complex and with a statistical success rate Vs performance tradeoff.
- xuhu 4y agoIf you find an AI generated response online, and ask GhatGPT if it was the author, it says "it was probably written by a human". But we all know there is a split infinitive here, and an archaic form there, and it knows. But it won't tell us.
- avisser 4y agoI love the idea that LLMs will get watermarked in a way where you can ask them who they were built for and they just tell you.
- eigenvalue 4y agoCould already have happened in these weights. Reminds me of when the movie studios started projecting random dot patterns during movies to try to catch which theaters were leading to bootlegs. Their approach was essentially defeated by pirates sourcing multiple versions and combining them. In this case, I suspect you could add a small normally distributed random number to some random subset of the weights and it would have very little impact on performance but would corrupt any watermark beyond recognition.
- Tiberium 4y agoThe original 4chan thread seems to indicate that the leaker verified that his hashes matched with another person who had access to the weights, to make sure that the weights aren't watermarked [0] 0: https://boards.4channel.org/g/thread/91848262#p91849855 https://boards.4channel.org/g/thread/91848262#p91849855
- Tiberium 4y agoIt seems that the leak originated from 4chan [1]. Two people in the same thread had access to the weights and verified that their hashes match [2][3] to make sure that the model isn't watermarked. However, the leaker made a mistake of adding the original download script which had his unique download URL to the torrent [4], so Meta can easily find them if they want to. [1]: https://boards.4channel.org/g/thread/91848262#p91850335 https://boards.4channel.org/g/thread/91848262#p91850335 [2]: https://boards.4channel.org/g/thread/91848262#p91849717 https://boards.4channel.org/g/thread/91848262#p91849717 [3]: https://boards.4channel.org/g/thread/91848262#p91849855 https://boards.4channel.org/g/thread/91848262#p91849855 [3]: https://boards.4channel.org/g/thread/91848262#p91850503 https://boards.4channel.org/g/thread/91848262#p91850503
- causi 4y agoJust a warning to readers, I would not recommend clicking 4chan links while at work.
- VadimPR 4y agoDoes this mean that with big enough compute capacity - say, Petals https://github.com/bigscience-workshop/petals https://github.com/bigscience-workshop/petals which distributes the model over the internet over GPUs - we can run LLAMA?
- linearalgebra45 4y agoHypothetically, what would the consequences be if I ran this on my university's computing cluster?
- elcomet 4y agoNone
- htrp 4y agoEither you get a nice invitation to collaborate on research with one of your uni's professors..... or you get sent to academic/disciplinary review and probably suspended for the semester.
- cma 4y agoWhy? Model weights aren't copyrighted and they didn't protect it as a trade secret.
- hn_20591249 4y agoSeems a valid use of resources if you have a way to vaguely associate it to some academic side-project, just don't start monetizing the output and beware the wrath of stressed out PhDs if you use too much capacity.
- transitivebs 4y agoSeeding...
- Aissen 4y agoIt's nice that it's downloadable without filling a form (even though it should have been the default), a leak was bound to happen. The license is quite restrictive anyway: see RESTRICTIONS on https://forms.gle/jk851eBVbX1m5TAv5 https://forms.gle/jk851eBVbX1m5TAv5
- sebzim4500 4y agoIf someone just decides to use the torrent and ignore those restrictions it might finally establish precident for if you can copyright model weights.
- dougmwne 4y agoBut even if you could copyright them, once you do some fine-tuning, they are not the same model weights!
- charcircuit 4y agothat would be a derivative work
- dougmwne 4y agoThat’s a legal unknown. And it’s also a technical unknown how you would even determine it was descended from the same model in a way that would hold up in court.
- sebzim4500 4y agoI would imagine that the weights of a finetuned model are highly correlated with the original weights. Having said that, simply permuting the neurons would make it way harder to match them up, I can't think of a straightforward way to reverse it.
- counttheforks 4y agoYou mean like how the model itself is a derivative work of tons of copyrighted content? If the original model can sidestep the issue of being trained on copyrighted content, then it should be fair game to train a new model off of a copyrighted model.
- rvnx 4y agoThe most logical thing would be for Archive.org to distribute these weights
- binarymax 4y agoHas anyone managed to download this yet using the magnet link? Is it well seeded?
- unethical_ban 4y agoLet's say I wanted to use this for... whatever. How do I do it? I bookmarked some "AI for beginners" youtube videos. No, I'm not trolling. The jargon and the ideas around LLMs is completely foreign to me. I have no idea how they work.
- electrosphere 4y agoI would like to know too.
- turmeric_root 4y agoclone this and point the script(s) to your downloaded model files: https://github.com/facebookresearch/llama/ https://github.com/facebookresearch/llama/
- madmod 4y agoWould there be some way to “launder” the model to make it plausibly viable for commercial use? Train a new model with the weights of this model with some kind of noise added to make it hard to tell what it is based on?
- ImprobableTruth 4y agoDistillation would be the ideal way (especially because it also has efficiency gains), but as far as I know distillation for LLMs is kinda unproven. Honestly though, even if you just finetune it, which you will want anyway for any serious commercial application, it's essentially impossible to determine the origin.
- bertday 4y agoRandomly perturbing the weights and then finetuning would probably make it impossible. If someone had access to the finetune dataset and you didn’t add noise, they could see if the finetuning curves intersect. I guess in practice, it’ll look suspicious if you have an identical model architecture and have similar performance.
- m3kw9 4y agoTill someone puts up a site to test it
- ok123456 4y agoMaybe this is an intentional leak to damage OpenAI. A supposedly better model by some accounts that strikes right at the heart of their business plan of selling access for $250k/year. One month of access to their service could buy a machine capable of running this leaked model. Facebook nerfs a potential upstart competitor to keep current big-tech cartel stable. Maybe this is a bit conspiratorial, but we live in the age big-tech and big-conspiracy.
- sebzim4500 4y agoIMO it's way more likely that some random guy on 4chan leaked it than it being some vast conspiracy.
- tinyspacewizard 4y agoNot a conspiracy at all. See also IE, Android, Kubernetes...
- slig 4y agoI missed the one about K8S, do you have any resources?
- GuB-42 4y agoI am not aware of Android and Kubernetes being leaks, they were open source from the start. For Android, openness was a big marketing point. I am not aware of IE leaks, and if there were leaks, hackers searching for exploits would be probably be the most interested, and that would be a bad thing for Microsoft. The problem with leaks is that they don't come with a license, you don't have the right to use them for any legitimate purpose. No one who could afford a 250k/year license would touch that leak as it could get them in big trouble.
- ok123456 4y agoAny links about IE, Android and Kubernetes? I'm not up on these being ops.
- 4y ago
- underlines 4y ago- how much vRAM needed to run each model parameter size? - any inference optimization we can use similar to StableDiffusion, to bring down the vRAM requirements? I only know about these: - use 8bit precision - https://github.com/bigscience-workshop/petals https://github.com/bigscience-workshop/petals - https://github.com/FMInference/FlexGen https://github.com/FMInference/FlexGen - https://github.com/microsoft/DeepSpeed https://github.com/microsoft/DeepSpeed Anything that could bring this to a 10GB 3080 or 24GB 3090 without 60s/it per token?
- joker99 4y agoIf I may tack on a question as someone with zero clue of ML: when, if ever, will someone like me be able to run this on a Mac Studio with a M1 Ultra and 128GB of ram?
- terafo 4y agoAs far as I can tell you can do it right now, at least for small 13B model, not sure about bigger models.
- eightysixfour 4y agoI don't believe they could, need CUDA and more VRAM...
- terafo 4y ago128 gigs is more than enough to load 13B model into. Pytorch has M1 support for some time now so CUDA isn't required.
- eightysixfour 4y agoDoes M1 use system memory as VRAM as well?
- 4y ago
- throw14082020 4y agoIn case anyone was wondering, the torrent contains 219.01 GiB. More specifically, the 65B parameters, is 121GB, the 30B parameters is 60.59GB, and so on.
- speedylight 4y agoLegally speaking is it a good idea to download these models this way?
- EamonnMR 4y agoMy compliance brain says no, but the fact that models get trained with data they obtain without explicit permission makes says that finders keepers would be the relevant case law.
- onetokeoverthe 4y agoGood thing ive had decades of real relationships and sex. Knew the net would probably squash print and privacy the first minute i logged into aol. Who knew it would breed a generation of robot loving losers?
- xg15 4y agoHow horrible! Is there a torrent link so I can be sure to never accidentally download it?
- AustinDev 4y agoSee github link in OP :p
- DesiLurker 4y agothere is a magnet link here somewhere, i think
- lopkeny12ko 4y agoIf the model is open source, who cares? This is good for the community; no need to go through Meta's opaque approval process.
- gpm 4y agoRecent comment in this discussion thread of the PR > looks like some people have been complaining about the link. it will need more seeders before we can merge into main from someone claiming to be > Research Scientist at Facebook AI Research. Working on [...] and who has previously merged pull requests for a repo under https://github.com/facebookresearch https://github.com/facebookresearch (I'm going to leave their name out of this... because it feels like that comment might come back to bite them)
- deleted 4y ago[deleted]
- BeFlatXIII 4y agoGood. Information deserves to be free.
- ComplexSystems 4y agoThe smallest model (7B) is supposed to outperform GPT-3. Does anyone have any idea what hardware is needed to run this?
- coolspot 4y ago7B would require at least 14GB VRAM in 8 bit precision. 28GB in 16 bit precision.
- JimmyRuska 4y agoSupposedly double the model size so 14gb. RTX 4090 might be able to handle it. You can use lambdalabs to rent a server gpu for one of the larger models.
- TaylorAlexander 4y agoI don't know if it matters but the 7B parameter checkpoint is 13.5GB in size. Someone with 24GB VRAM struggled to run it: https://github.com/facebookresearch/llama/issues/55 https://github.com/facebookresearch/llama/issues/55
- throwaway1851 4y agoNo, the 13B model outperforms GPT-3. Judging from the metrics published in the paper, it does look like the 7B model is not far off from GPT-3 however.
- CapShoyo12 4y agoI'm excited, but having trouble running Llama on my local machine, has anyone managed this?
- AbusiveHNAdmin 4y ago[dead]
- alfalfasprout 4y agoWarning: do not use this for commercial purposes. While the weights may be available now, it's a lawsuit waiting to happen if you try to use this at work. See the original license: "a. Subject to your compliance with the Documentation and Sections 2, 3, and 5, Meta grants you a non-exclusive, worldwide, non-transferable, non-sublicensable, revocable, royalty free and limited license under Meta’s copyright interests to reproduce, distribute, and create derivative works of the Software solely for your non-commercial research purposes. The foregoing license is personal to you, and you may not assign or sublicense this License or any other rights or obligations under this License without Meta’s prior written consent; any such assignment or sublicense will be void and will automatically and immediately terminate this License."
- flangola7 4y agoAnd where did I sign my name to that agreement?
- alfalfasprout 4y agoLicense agreements/terms of use don't require signature usually. Consent is implied by downloading. That's also the case when you eg; clone a repo, download a file, etc. I'm anti DRM + restrictions as much as the next guy but just trying to save folks from a bad time if meta comes knocking after seeing corporate IPs downloading the weights.
- grrowl 4y agoThis might be a bit of an assumption, but it seems likely Meta is willing to lose more on lawyers than you'd be willing to ever spend.
- alfalfasprout 4y agoThey have a ton of in-house counsel too lol. Little of it would be incremental spend.
- MrStonedOne 4y ago
- politician 4y agoLLMs invalidate the concept of copyright to such a degree that I find it impossible to see this torrent as theft.
- Madmallard 4y agoThere's pretty much no point in downloading this right? It cannot be run with any fidelity on any consumer end gpu
- WithinReason 4y agoLooks like the weights are legit, I got the 7B model to generate some text (on a single GPU). Using the 1st prompt from the script it generated this: [I believe the meaning of life is] to be happy, and it is also to live in the moment. I think that is the most important thing. I'm not really a party girl. I'm not a girl's girl. I have a really small group of close girlfriends and that's all I need. I believe in equal rights for everyone. I'm not a rebel. I don't really rebel against anything. I'm a very traditional girl, very loyal. I'm a mum's girl and I'm a dad's girl. People have a right to know what's going on. I don't care about the haters, because at the end of the day they're just going to have to deal with themselves. I've been getting more and more into fashion since I was about 16. I know I'm a little different, but so what? I think that's good. I don't think you should be like everyone else. It's my birthday, and I'll cry if I want to. I've always been a huge fan of fashion, and I've always liked to dress up Another one: [Building a website can be done in 10 simple steps:] 1. Defining Goals 2. Your Branding and Web Presence 3. Defining Your Marketing Strategy 4. Creating Your Website 5. Your Website Design 6. Your Website Development 7. Your Website Launch 8. Your Website’s Content 9. Your Website’s Conversion Rate 10. Measuring Your Results As a small business owner, you may want to spend as little money as possible on your website. But if you want to see a positive ROI, you will need to spend some money. Defining goals is critical when building a website. You should know what you want to accomplish with your website. You need to know what you want your website to achieve. You need to know who you want to convert to a customer. You need to know how you want to reach your goals. You need to know what the timeframe is for your website goals. You need to know what you want to get out of your website. When building a website, you need to clearly define your goals. Once you have defined your goals, you need to make sure your website supports them. If you want to reach your goals, you
- aghack 4y agoCan this be finetuned?
- fiat_fandango 4y agoI wonder if anyone is legitimately concerned that mirrored downloads might contain malicious payloads?
- happycube 4y agoWho here didn't see this leak coming?