22 ms·
Grok
- sqreept 3y agoWhat are the languages supported by it?
- cyanydeez 3y agoTweets.
- atleastoptimal 3y agoI think everyone should realize the following realities of the LLM market 1. For sub-SOTA LLM's, distribution/marketing is more important than having a proprietary lock on capabilities. Open sourcing is a benefit for the firm, distincct from goodwill 2. For SOTA LLM's, keeping it closed and proprietary is the strategic play If grok were SOTA Elon never would have open sourced it. It's not even SOTA within XAI. This is a marketing play to win public sentiment against OpenAI.
- mlindner 3y agoIf it's better than any other open source LLM does that even matter? (I say "if" because I don't know.)
- keepamovin 3y agoI recall Elon saying something like this in an interview so I think it’s less of a deceptive take then perhaps your comment suggest. I think he said something like proprietary AI tech is going to be one year to 18 months ahead of where open source tech is which will follow on like one year to 18 months later. Suggesting that he’s aware of this dynamic and he’s not trying to conceal or misrepresent that. In other words, perhaps this was SOTA one year to two years ago?
- atleastoptimal 3y agoWhich is correct. The point I'm going for is not against Elon but against his obedient fans and knee-jerk OpenAI haters who claim that they should, by natural obligation, do the "right thing" and open source all their models, and Elon open sourcing grok is him "leading by example" and being the hero that OpenAI can't.
- keepamovin 3y agoInteresting. That point didn't come across in your original comment. I recommend you state it next time at the end. Often times stuff that seems obvious to us / yourself / people who know about something -- can go unstated in stuff you say that otherwise references specific points at hand -- and omits these general, but enlightening/useful perspectives/priors, which it would be good to share. This is not only for you specifically just a general reminder for all of us including me.
- atleastoptimal 3y agoI think that's true though my original comment I feel was sufficient in its claim and implicit assumptions. Basically I feel people's feelings about Elon vary a lot but are anchored by 3 general categories. > 1. Elon Musk is a messianic savior who is perfectly selfless and always does the right thing. Every business decision he makes is for the maximal good of humanity > 2. Elon Musk is a typical CEO who does typical CEO things, serving his own interests, except he's better at marketing his own image and is much more outspoken > 3. Elon Musk is an irredeemable evil who always does objectively wrong things My first comment was implicitly addressed to people in the 1 camp trying to bring them into the 2 camp (which is where I am).
- keepamovin 3y agoAlright, it just didn't come across for me, haha! :) I guess sometimes those implicit assumptions really are too implicit! I think it's good to err on the side of expressing them, because you can't assume someone else thinks the same way you do. That's what I've learned anyway. Hahahaha! :) Reading your comment again with your explanation it is clear that's what you're doing. Although, regarding your desires to present a balanced view and to persuade, I have an idea. It probably sounds like I have no idea what I'm talking about, but I think your OG comment would perhaps benefit from sounding a little bit more friendly toward Elon (not to the messianic savior level haha), but the way it sounds to me is Elon is being deceptive here and presenting it as goodwill when it's not. However, I think the truth is there's a little bit of both, right? There's good will but it's also strategic. I get if you don't think so, tho, no worries! Haha! :) Your OG comment sounds to me like Elon's just Machiavellian, and I get where you're coming from to remind the people who think he's a savior, but if you're point is not to go "against Elon" as you said, it might be good to acknowledge the good that he does. At least, that way -- whether or not you believe that acknowledgment -- if you hope to bring over people who think that way, you'll probably need to appeal to how they think, rather than just dose them with the truth you see, because then they'll shut it out, if there's nothing they can relate to. Although, if I haven't convinced you even a bit here, then maybe you shouldn't listen to me about persuasion because I guess I don't know how to do this myself. At least not effectively, or here with you. Haha!:) But if you do feel a little bit convinced then maybe consider it for next time to help your persuading people back to a more balanced view? :) But then, there's the question of if such a thing is even possible. If people have an particular view, it could be challenging to change it, as confirmation bias means you'll ignore evidence even when it expands your worldview. Hahaha! :) This was a funny conversation. I think we somehow skirted around the important point tho that OpenAI could in fact open source some of its older models, could it not? Musk is a typical CEO who does typical CEO things, serving his own interests, except he's better at marketing his own image and is much more outspoken, but there might also be a bit of truth to what the fanboys say about OpenAI in that it seems they do have some room to "open source" their non-SOTA stuff, or what am I missing?
- deleted 3y ago[deleted]
- sashank_1509 3y agoIn all the debate about open source I don’t think people realize, this model is most likely not reproducible ever again even given the code. Here’s what you need to reproduce the model: 1. An exact snapshot of the data used, many companies don’t have this, you have rough dataset versions but remember if even 1 token is different, the model produced won’t be the same. 2. Data must be sent to the training algorithm in the exact same order as it was originally. So every data loader needs to be with a fixed random seed. 3. All the probabilistic parts of your model needs to have a fixed random seed. Here I’m thinking of stuff like dropout and for autoregressive models you might be sampling your previous output, you have to ensure they are properly seeded. Generally you do see fixed seeds in academic papers but it’s easy to miss stuff especially in distributed training jobs. 4. Here’s another interesting thing, you start your training job on 1000 GPUs and then suddenly 4 GPUs fail. What do you do? There might be deterministic ways to solve this but the standard approach is to discard all updates that that GPU was going to do and restart that GPU from scratch. You can see why this is a problem? Now if you want to reproduce this training you need to disable those GPU at the same time in the new training job to make this work. I suspect there are even more things I didn’t think of that will make this model unique and irreproducible by training for eternity, almost like a human brain? In fact the notion of exact reproducibility in the world of LLMs is silly, there is only approximate reproducibility, (models with similar scores in benchmarks) but nothing exact. That said I can see the value of releasing source code but I’m completely fine with grok not releasing it. Source code can reveal tricks that have not been published in papers yet that a company discovered to improve their model. Seeing the performance of Grok, I’m pretty confident there isn’t any great tricks to be found in their code so I don’t really care, I would be pretty curious about OpenAI’s or Anthropic’s source code though.
- Grimblewald 3y agoWhich is why I don't buy into the LLMs don't have personal opinions schtick. Each LLM by virtue of the factors you've mentioned will have its own unique 'perspective', if you will, on a variety of topics. I think it's more correct to say everything a LLM says is it's personal opinion rather than it being some objective truth or something.
- mvkel 3y agoThis feels like a "now we can say we're open" PR play rather than contributing much value to the open source community. What is the practical use of this repo?
- deleted 3y ago[deleted]
- joydeep314 3y agoModel weights on huggingface: https://huggingface.co/xai-org/grok-1 https://huggingface.co/xai-org/grok-1
- aussieguy1234 3y agoHow hard would it be for an open source group to fine tune this into a chatbot?
- ilaksh 3y agoHas anyone outside of x.ai actually done inference with this model yet? And if so, have they provided details of the hardware? What type of AWS instance or whatever? I think you can rent like an 8 x A100 or 8 x H100 and it's "affordable" to play around with for at least a few minutes. But you would need to know exactly how to set up the GPU cluster. Because I doubt it's as simple as just 'python run.py' to get it going.
- zone411 3y agoIf you're just looking to test it out, it's probably easiest to wait for llama.cpp to add support (https://github.com/ggerganov/llama.cpp/issues/6120 https://github.com/ggerganov/llama.cpp/issues/6120), and then you can run it slowly if you have enough RAM, or wait for one of the inference API providers like together.ai to add it. I'd like to add it to my NYT Connections benchmarks, and that's my plan (though it will require changing the prompt since it's a base model, not a chat/instruct model).
- logicchains 3y ago>it's probably easiest Cheapest maybe, but easiest is just to rent a p4de.24xlarge from AWS for a couple hours to test (at around $40/hour..).
- zone411 3y agoI'd expect more configuration issues in getting it to run on them than from a tested llama.cpp version, since this doesn't seem like a polished release. But maybe.
- v9v 3y agoThe NYT Connections benchmark sounds interesting, are the results available online?
- zone411 3y agoGPT-4 Turbo: 31.0 Claude 3 Opus: 27.3 Mistral Large: 17.7 Mistral Medium: 15.3 Gemini Pro 1.0: 14.2 Qwen 1.5 72B Chat: 10.7 Claude 3 Sonnet: 7.6 GPT-3.5 Turbo: 4.2 Mixtral 8x7B Instruct: 4.2 Llama 2 70B Chat: 3.5 Nous Hermes 2 Yi 34B: 1.5 The interesting part is the large improvement from medium to large models. Existing over-optimized benchmarks don't show this. - Max is 100. 267 puzzles, 3 prompts for each, uppercase and lowercase - Partial credit is given if the puzzle is not fully solved - There is only one attempt allowed per puzzle, 0-shot. - Humans get 4 attempts and a hint when they are one step away from solving a group I hoped to get the results of Gemini Advanced, Gemini Pro 1.5, and Grok and do a few-shot version before posting it on GitHub.
- nasir 3y agoI'd be very curious to see how it performs especially on inputs that's blocked by other models. Seems like Grok will differentiate itself from other OS models from a cencorship and alignment perspective.
- porkbeer 3y agoSo far that is quite a low bar. But balancing is a thing nontheless, lest we end up with Tay again.
- cl3misch 3y agoLove the minimal repo, magnet link, and stating "open weights" instead of "open source". Refreshing!
- TheDudeMan 3y agoElon says open source: https://twitter.com/elonmusk/status/1767108624038449405?s=46&t=62CEo1y35cyP8jxNy7m2CQ https://twitter.com/elonmusk/status/1767108624038449405?s=46...
- greenpizza13 3y agoIf we just stop looking at Elon, he will lose his power. Why oh why do we keep giving him attention? There are plenty of great models out there that _aren't_ backed by maniacs.
- rafaelero 3y agoWhen those great role models are able to build a profitable spaceship company from the ground up I am sure we will pay attention to them.
- shantnutiwari 3y agoThose of us who dont spend all our time in LLMs-- whats this about? Whats the big deal and why is it on the front page at #1?
- kayge 3y agoI think this paragraph from an earlier Wired article [1] sums it up pretty well: "After suing OpenAI this month, alleging the company has become too closed, Elon Musk says he will release his “truth-seeking” answer to ChatGPT, the chatbot Grok, for anyone to download and use." [1] https://www.wired.com/story/elon-musk-no-choice-open-chatbot-grok/ https://www.wired.com/story/elon-musk-no-choice-open-chatbot...
- solumunus 3y agoThe deeper reason is that he’s throwing his toys out of the pram after failing to convince OpenAI to become part of Tesla.
- ArunRaja 3y agoIs this grok open sourcing really a big deal? How is this move beneficial for grok per se? Does it build trust as in other opensource products..?
- deleted 3y ago[deleted]
- tosh 3y agoblog post: https://x.ai/blog/grok-os https://x.ai/blog/grok-os * 314B parameters (86B active at a time) * mixture of experts 8 (2 active at a time) * weights and architecture licensed under Apache 2.0 (edit:) announcement blog post from last year with benchmarks compared to Claude 2, GPT-3.5 and GPT-4: https://x.ai/blog/grok https://x.ai/blog/grok (edit2:)TL;DR: somewhat comparable to GPT-3.5, Mixtral and Qwen-1.5-72B in capability but way larger than the open weight models
- TOMDM 3y agoMixtral is also comparable to gpt 3.5 and open. At 8x7B it's also a fraction of the size. Are there any benchmarks comparing Mixtral to Grok?
- tosh 3y agoMixtral announcement is here: https://mistral.ai/news/mixtral-of-experts/ https://mistral.ai/news/mixtral-of-experts/ Mixtral looks more economical @ capability to size (similar also for Qwen 1.5 72b)
- OkGoDoIt 3y agoIs a model so huge that’s only at the level of GPT 3.5 actually good? That seems incredibly inefficient to me.
- fwlr 3y agoOpenAI is valued at 90 billion and all they do is make GPT; Twitter is valued at 40 billion and this was essentially a vanity side-project by a cowboy CEO. Presuming that benchmarks and general “it’s about the level of 3.5” is accurate, it’s inefficient, but not incredibly inefficient imho
- pelorat 3y ago> Twitter is valued at 40 billion WAS vaulued at 44B. Now? Maybe 5 billion.
- extheat 3y agoAt 8x86B, looks like the largest open model yet by far. Would be interesting to hear how many tokens it's been trained on. Especially important for higher param models in order to efficiently utilize all those parameters.
- zone411 3y agoIt's actually not the largest. https://huggingface.co/google/switch-c-2048 https://huggingface.co/google/switch-c-2048 is 1.6T parameters.
- WeMoveOn 3y agobut is switch c even usable? iirc the training set was nowhere near enough for a model of that size to be coherent in a conversation
- p1esk 3y agoIt’s not 8x86B. Total number of parameters is 314B. Perhaps it’s 8x39B to fit on a single 8xA100 (40GB) server?
- dheera 3y agoThey all do this marketing bull. Mixtral has an 8x7B model but it's actually 46.7B, not 56B params. Kinda similar to how 4K displays are 3840 pixels wide, not true 4K which would be 4096. Marketing people called it 4K, not engineers.
- guitarlimeo 3y agoI've always thought of 4K as "4x FullHD". In that way it makes sense.
- mavhc 3y agoTV and Digital Cinema have different standards, because of course they do
- deleted 3y ago[deleted]
- hubraumhugo 3y agoWhen will we reach an upper limit/dimishing returns in terms of number of parameters and mixture of experts?
- andy99 3y agoWe may have already - data is more important than anything else which is why nobody has beat GPT4 yet. Throwing more parameters or more compute at the problem only gets you so far. But Grok was never a contender so there is room to improve on it. It is one of the biggest models open sourced as mentioned, so will be interesting to take a look at for sure.
- squigz 3y agoI think Groq is something else?
- lambdaba 3y agoClaude 3 has *decisively* beat GPT-4, I wonder how all their attributes compare.
- stainablesteel 3y agoi like some of claudes answers better, but it doesnt seem to be a better coder imo
- simonw 3y ago
- rvnx 3y agoOne subtle thing: Musk said "open-source", we got "open-weights" instead (still better than nothing though, so it's greatly appreciated).
- paulgb 3y agoDumb question: what should open-source mean in the context of something like this? Open access to the training data and training pipeline as well?
- CharlesW 3y agoIt's not a dumb question, and the answer is "yes".
- dudus 3y agoIf you release that instead of the binary weights you can be both more open and less useful for users. Fun
- zeroCalories 3y agoCome on, that's not reasonable to expect from a company, or useful for indie hackers. Having weights that can be used however you like is enough for most people, even large companies.
- schoen 3y agoMaybe it should be called something else? "Openly-licensed"? Just because the model weights are not really "source" (either as a matter of intuition or for example following the OSI "preferred form in which a programmer would modify the program" definition).
- zeroCalories 3y agoSure, but I don't want to train anyone's model from scratch. Realistically, I can't download all the training data, or run the pipeline, or train the model. Making all of that available to me would be a massive burden on the company too, so they simply won't do it. If I'm able to fine-tune it, that's enough for me, and imo, that fits with the spirit of open/free software. We have to understand that this is fundamentally a different thing than something like the Linux kernel, and closer to something like an industrial project. The output is just a bunch of numbers instead of something physical.
- nylonstrung 3y agoFor what reason would you want to use this instead of open source alternatives like Mistral
- zozbot234 3y agoIsn't this Apache licensed? Regardless, you can run multiple models concurrently on the same input using well-known ensemble techniques. (Not to be confused with mixture-of-experts, which is more like training a single model where only a few blocks are chosen to be active at any given time - a kind of sparsity.)
- tlb 3y agoNot super easy if they have different tokenizers.
- rvnx 3y agoMistral opened their weights only for very small LLaMA-like model.
- MallocVoidstar 3y agoI'm pretty sure Mixtral outperforms Grok-1 and uses much less memory to do it
- elfbargpt 3y agoI'm a little out of touch, is there a way to see how Grok measures up to other models?
- amrrs 3y agoBenchmarks here https://x.ai/blog/grok https://x.ai/blog/grok
- refulgentis 3y ago
- gardenhedge 3y ago> Due to the large size of the model (314B parameters), a machine with enough GPU memory is required to test the model with the example code What type of machine do you need to play around with this?
- a_wild_dandan 3y agoA single 192GB M2 Mac using a 4-bit quant would work.
- 317070 3y agoProbably a machine with about 628 GB of GPU memory. (2 bytes per parameter) So 8xH100 (80Gb each) should do it.
- Marlinski 3y agoI suppose it can be quantizised
- anigbrowl 3y ago'Chunky beast, needs 320 Gb VRAM likely 4 bit, likely is being run 8 bit on 8 x 80 Gb GPUs.' -Emad
- pogue 3y agoCan someone explain why the weights are posted via a Bittorrent magnet link? I have no way to check the size at the moment, but isn't that a bit unusual? There's also only 21 seeders right now according to https://checker.openwebtorrent.com/ https://checker.openwebtorrent.com/
- raydev 3y agoSpreads the burden/cost of distributing a 300+GB file.
- seydor 3y agomy optimistic explanation is we are going back to the 2000s internet , but probably we are not
- fzzzy 3y agoLet's hope so.
- ur-whale 3y ago> Can someone explain why the weights are posted via a Bittorrent magnet link? I think the best way to get an answer to that question is to try to host it yourself and see what happens.
- whywhywhywhy 3y agoBecause Bittorrent is an outstanding tech for delivering large files, more I think about it the more I'm surprised it wasn't taken advantage of more.
- Marlinski 3y agoit's been criminalized to hell by IP holders and hollywood. Such a shame they killed the best tech of the previous decade. Could have revolutionized how we distribute content, approach CDN and even streaming.
- 3y ago
- bbor 3y agoHonestly the most interesting part is taking a peek at the kind of AI researcher working for Twitter after the objectively messy layoffs and subsequent crunch. I notice neither of them has Twitter mentioned on their GitHub, which is prolly for the best to avoid harassment lol. Code wise, excited to see if this could grow into anything! I think it’s pretty clear that Grok didn’t have nearly enough investment to be a top model so Elon “sacrificed” it on a whim in his schoolyard spat with OpenAI, but I’m not complaining. I’ve always took Elon on his word that he truly is worried about centralization of AI, and I don’t think any of the emails released by his schoolmate Altman dissuade me of that. So I have some reasonable hope that he uses some of his immense resources to start “fighting the good fight” here with Le Cun
- cma 3y ago>taking a peek at the kind of AI researcher working for Twitter He made a separate company for this.
- paxys 3y agoNeither of them works at Twitter. xAI is a separate company, and only uses Twitter’s data to train.
- bbor 3y agoThanks for the correction! I know, I just don’t believe in corporations so the distinction is slight
- mattxxx 3y agoI respect the openness here! This is the future that I want to see
- trog 3y agoIs it open if it doesn't include the training data? Genuine question - I am not familiar enough with the terms and technology to know. But my understanding is the weights is just a more or less static collection of data that has been (to paraphrase Ted Chiang) lossily compressed from the actual raw training data. Without the training data to thoroughly evaluate what is in there, the only way you can figure it out is through experimentation - e.g. running it up in a chatbot and asking it questions. Is this roughly correct or am I misunderstanding what you can do with the weights?
- deleted 3y ago[deleted]
- foreverobama 3y ago[dead]
- giancarlostoro 3y agoFully agree. People will trash talk it due to Musk but lets not forget the engineers who poured hours of their lives into building this and are continuing to do so.
- sprobertson 3y ago> engineers who poured hours of their lives into building this Not to mar these specific engineers, but that's an empty phrase that can be said about anything ever built. It doesn't somehow make the idea or implementation good.
- giancarlostoro 3y agoThe phrase merely means dont just overlook something because someone else who did not even labour over the end result.
- moralestapia 3y agoWell, he delivered.
- paxys 3y agoPartially. Open weights is not open source.
- xcv123 3y agoThe architecture of the model is open source. Not just the weights. You can run the entire thing locally.
- gfodor 3y agoIn machine learning models the term open source has been largely accepted to mean sharing weights and, if necessary, inference code. You can argue if this is an abuse of the term but everyone does it, and saying someone didn’t deliver if they used it and published weights would probably mean saying the same about mistral, meta, etc.
- asadotzler 3y agoYes. So say the same thing about them Open source has a definition and abusing that hurts all of us except the billionaires.
- moralestapia 3y agoI get the "open source" argument, but what is the issue here? If you are able to reproduce the thing in its entirety and you're given no restrictions on its use, it seems compatible with the spirit of open sourcing things.
- 2devnull 3y agoFrom issues: “Well the magnet file contains a 300GB checkpoint “ That’s why they are using a torrent I suppose.
- deleted 3y ago[deleted]
- stale2002 3y agoHey, asking any experts here, what are their first thoughts in the significance of this? IE, is this comparable to any other model released, or are there significant metric differences that make it better for certain usecases? The only thing I see, of the top of my head, is that it is a very large model, and I don't think any models of similar size have been released.
- whimsicalism 3y agoseems like a large undertrained model, not that exciting imo compared to mixtral it is also not the biggest model oss, switch transformer was released years ago and is larger and similarly undertrained
- brucethemoose2 3y agoTests are not out yet, but: - It's very large, yes. - It's a base model, so its not really practical to use without further finetuning. - Based on Grok-1 API performance (which itself is probably a finetune) its... not great at all.
- Me1000 3y agoNot an expert by any means, but I like learning about this stuff and I play with a lot of open weight models. I’d say the significance is that it happened. It’s by far the largest open weight model I’ve seen. But I’m not sure why you’d use it over a model like Mixtral, which seems to perform about the same at like 1/6th the size. But I welcome any contribution to the open weight LLM community. Hopefully people will learn something interesting with this model. And I hope they keep releasing new versions!
- MichaelRazum 3y agoIf I may ask, how do you load such big models? 300gb seems like a lot to play around with.
- Me1000 3y agoYou're right, this model is going to be too big for most people to play around with. But to answer your question I have a 128GB of RAM in my M3 MacBook Pro, so I can use most of that for GPU inferencing. But still, this model is going to need to be heavily quantized for me to be able to use it. (fwiw, I probably wont try this one) In the next week or two I expect we'll see a GGUF version of the weights (might need to wait for a patch to llama.cpp first), and someone will release super small quantizations of it. I suspect my computer might be able to run a 3 bit quant, but it might need to go down to 2 bits to have any kind of reasonable context length. But with quants that small I'd expect the model's performance to degrade well below that of Mixtral, so it probably isn't really even worth using. But we'll see; quantization is weird, some models perform better than others when quantized.
- simonw 3y agoIs there a model card anywhere? I'd like to know what it was trained on.
- LZ_Khan 3y agoHow are people's experience with this model? Having the most weights is one thing but being a better model than the 70B models is another.
- labrador 3y agotbh, I've never seen anyone share anything interesting produced by Grok. I see plenty of posts on X and reddit of people sharing amazing things that GPT-4 and now Claude 3 Opus can do. Grok can roast people. That's pretty much all I've seen. I'd love to proven wrong if someone cares to share something interesting produced by Grok.
- swalsh 3y agoI use grok all the time to find tweets or ask about trends on Twitter. For that it's better than what used to exist. But its not a great model outside that narrow use case.
- arduanika 3y agoCODE_OF_CONDUCT.md has only five words. :)
- schappim 3y ago"Be excellent to each other."
- troupo 3y ago[flagged]
- TMWNN 3y ago>Which is ironic given Musk's own behaviour You mean, like immediately responding to Ukraine's plea for Starlink and funding it on his own for months? So much so, that in February 2023 its government called Musk "one of the biggest private donors of our future victory"? <https://www.pravda.com.ua/eng/news/2023/02/9/7388696/ https://www.pravda.com.ua/eng/news/2023/02/9/7388696/>
- troupo 3y ago[flagged]
- arandomusername 3y ago> pushing increasingly unhinged conspiracy theories like what?
- machiaweliczny 3y agoIf they are so behind they could make it open source instead of open weights and get some help.
- deleted 3y ago[deleted]
- nicce 3y agoFully open-source means also providing open access to their data sets? Which is the only valuable thing Twitter (X) has left.
- heyoni 3y agoAnd the one thing they are vehemently protecting from scrapers and other entities. Even nitter threw in the towel.
- EastSmith 3y ago> Which is the only valuable thing Twitter (X) has left. reply They have a very valuable user base (all kinds of world leaders for example), so the data is not the only valuable thing they have.
- nicce 3y agoI don’t see difference here. Userbase and their social networks and interactions is the data. They don’t have much value from advertising point of view anymore.
- sroussey 3y agoThat’s actually more valuable. Twitters data of small format text is awful for training. Best to just exclude it. There are hundreds of millions of people on Twitter, and a few of them are very smart. I don’t see how that helps here though.
- Takennickname 3y ago
- simonw 3y ago"Base model trained on a large amount of text data, not fine-tuned for any particular task." Presumably the version they've been previewing on Twitter is an instruction-tuned model which behaves quite differently from these raw weights.
- seccode 3y agoIt would be cool if these models had conversations with us where they ask questions. I think the future of AI is models that ask questions. There is so much data to be gained by doing this.
- swalsh 3y agoThat's just a matter of fine tuning
- seccode 3y agoDo you have an example model I could try that does this?
- ijustlovemath 3y agoThat "just" is doing some heavy lifting! GPT-4 is just a few matrix multiplications, how bad can their moat really be?
- swalsh 3y agoI'd bet a synthetic data set could do the job effectively.
- BoorishBears 3y agoNot sure what the snark here is for: It would be trivial to produce a dataset where the model asked you questions then fine-tune on that. People already do it with chain-of-thought and you could get away with a few dozen examples if you wanted to try this.
- littlestymaar 3y agoHow long before the Groq team sues for trademark violation? It's literally the purpose of trademark laws to make sure resembling names do not cause confusion in the mind of customers so it would be very surprising to see this situation persist.
- mlindner 3y agoGrok is a word in common parlance. So there's no way they could succeed in any suit. That's why the Groq team picked a modification of the word.
- littlestymaar 3y agoYou mean like Canvas®, Apple®, Windows® or Amazon®? Wanna try re-use these for your own business and see how it goes? There's nothing preventing you to trademark common words, it just must not be descriptive of your business.
- nostrebored 3y agoWould be a rough trademark enforcement case as “Grok” has been in common language for decades
- Findecanor 3y agoI myself have never heard it outside of "nerdy" circles... that is: people who would read science fiction. I personally am not entirely happy about the word (no matter how it is spelled) being used for a particular AI product. "Grok" to me means knowing a subject at a much deeper level than I think any AI is capable of at the present level of technology. But it would be passable to use it for a company name, to indicate that it is a goal to strive for.
- ben_w 3y agoGenerally agree, though I would say "knowing a subject at a much deeper level than any LLM is capable of", as AI more broadly also includes specialist models that are wildly super-human in narrow domains like chess and Go.
- philtar 3y ago[dead]
- orsenthil 3y agoI am not sure what open source models are accomplishing another than killing the lead from the competition (openai), only to give it to someone else who has expertise in the area of distribution. This will be yet another good addition to systems like Amazon BedRock.
- nateglims 3y agoI haven't seen anything about the larger architecture, but I think the value of grok is going to come from it's cheap access to twitter data for RAG etc.
- minimaxir 3y agoMany of the recent innovations in both LLM architecture and inference were only made possible through open models such as Llama 2 and Mistral 7B as a starting point for iteration and refinement, which in turn backpropagates (heh) back to the LLMs developers. It's a win-win for everyone. That's the power of open source.
- deleted 3y ago[deleted]
- geor9e 3y agoWell, look at the history. Google had an insurmountable lead, so Elon started OpenAI. Now OpenAI has an insurmountable lead too. So everyone else is starting in third place, or lower. David versus two Goliaths. If you try to become a third Goliath, you'll probably just get smashed. You're later to the game. In this situation, going scorched earth becomes a viable strategy. Slay the Goliaths. Become a hero to the masses. Attract the world's best talent who don't want to be associated with proprietary models. At that point you have a world class AI business with momentum towards AGI. And even if you're giving away last year's technology for free, the team you built is churning out new ideas that could be a financial bonanza one day. Shareholders are willing to pay for a long-term bet if the story is good.
- andre-z 3y agoThe only other Repository is a fork of Qdrant.
- captcanuk 3y ago"The implementation of the MoE layer in this repository is not efficient. The implementation was chosen to avoid the need for custom kernels to validate the correctness of the model." Or perhaps release your actual code AND the simplified implementation instead of hiding it and saying "you don't know her, she goes to a different high school"
- gfodor 3y agoAlways love it when someone gives away a gift and it’s not enough for people.
- captcanuk 3y agoNot just someone but the CEO of the company. He used HIS platform to say "This week, @xAI will open source Grok" (https://twitter.com/elonmusk/status/1767108624038449405 https://twitter.com/elonmusk/status/1767108624038449405) and they aren't doing that. What they delivered specifically says "We are releasing the base model weights and network architecture of Grok-1, our large language model."
- gordian-mind 3y agoSounds like they did what they said they would.
- redskyluan 3y agoThis seems not be a repo ready to open source. You only get weights, very less information about how the weights is trained and finetuned. But anyway, it always great to see more LLM weigts available.
- andrewstuart2 3y agoI would argue that there's no bar for open sourcing aside from "do you have the rights to do so." Some source or some public good is certainly better than none, and when the bar is low then you remove barriers to getting started, vs waiting until you have the time someday to "do it right."
- rezonant 3y agoWell what constitutes an "open source" model is still controversial and debatable-- lots of people on both sides of that argument.
- asadotzler 3y agoOpen source has had a useful agreed upon meaning for over 25 years. Maybe you're too young to understand why that matters but we're not.
- rezonant 3y agoI've been in the open source community for about 25 years so I doubt it. For what it's worth I would say a model should be fully reproducible to be open source, but that's not a decided consensus -- and AI models are sufficiently different than the source code / binary code distinction as to invoke discussion around defining it.
- deleted 3y ago[deleted]
- modeless 3y agoIs this the first major model to be natively FP8? I was wondering why people hadn't done it yet. Seems like a big win when hardware supports it.
- a_wild_dandan 3y agoNo, e.g. Yi-34B.
- modeless 3y agoAs far as I can tell Yi-34B is natively 16 bit float, the 8 bit version is quantized. https://huggingface.co/01-ai/Yi-34B#quantization https://huggingface.co/01-ai/Yi-34B#quantization