8 ms·
From the blog: https://arxiv.org/abs/2501.00663 https://arxiv.org/abs/2501.00663 https://arxiv.org/pdf/2504.13173 https://arxiv.org/pdf/2504.13173 Is there a
by okdood64 10mo ago
From the blog:
https://arxiv.org/abs/2501.00663 https://arxiv.org/abs/2501.00663
https://arxiv.org/pdf/2504.13173 https://arxiv.org/pdf/2504.13173
Is there any other company that's openly publishing their research on AI at this level? Google should get a lot of credit for this.
- Hendrikto 10mo agoMeta is also being pretty open with their stuff. And recently most of the Chinese competition.
- okdood64 10mo agoOh yes, I believe that's right. What's some frontier research Meta has shared in the last couple years?
- markisus 10mo agoTheir VGGT, Dinov3, and segment anything models are pretty impressive.
- robrenaud 10mo agoAnything with Jason Weston as a coauthor tends to be pretty well written/readable and often has nice results.
- tonyhart7 10mo ago"What's some frontier research Meta has shared in the last couple years?" the current Meta outlook is embarassing tbh, the fact they have largest data of social media in planet and they cant even produce a decent model is quiet "scary" position
- mirekrusin 10mo agoJust because they are not leading current sprint of maximizing transformers doesn't mean they're not doing anything. It's not impossible that they asses it as local maximum / dead end and are evaluating/training something completely different - and if it'll work, it'll work big time.
- johnebgd 10mo agoYann was a researcher not a productization expert. His departure signals the end of Meta being open about their work and the start of more commercial focus.
- woooooo 10mo agoThe start?
- DrewADesign 10mo agoI’ve long predicted that this game is going to be won with product design rather than having the winning model; we now seem to be hitting the phase of “[new tech] mania” where we remember that companies have to make things that people want to pay more money for than it costs to make them. I remember (maybe in the mid aughts) when people were thinking Google might not ever be able to convert their enthusiasm into profitability…then they figured out what people actually wanted to buy, and focused on that obsessively as a product. Failing to do that will lead to failure go for the companies like open AI. Sinking a bazillion dollars into models alone doesn’t get you shit except a gold star for being the valley’s biggest smartypants, because in the product world, model improvements only significantly improve all-purpose chatbots. The whole veg-o-matic “step right up folks— it slices, it dices, it makes julienne fries!” approach to product design almost never yields something focused enough to be an automatic goto for specific tasks, or simple/reliable enough to be a general purpose tool for a whole category of tasks. Once the novelty wears off, people largely abandon it for more focused tools that more effectively solve specific problems (e.g. blender, vegetable peeler) or simpler everyday tools that you don’t have to think about as much even if they might not be the most efficient tool for half your tasks (e.g. paring knife.) Professionals might have enough need and reason to go for a really great in-between tool (e.g mandolin) but that’s a different market, and you only tend to get a limited set of prosumers outside of that. Companies more focused on specific products, like coding, will have way more longevity than companies that try to be everything to everyone. Meta, Google, Microsoft, and even Apple have more pressure to make products that sanely fit into their existing product lines. While that seems like a handicap if you’re looking at it from the “AI company” perspective, I predict the restriction will enforce the discipline to create tools that solve specific problems for people rather than spending exorbitant sums making benchmark go up in pursuit of some nebulous information revolution. Meta seems to have a much tougher job trying to make tools that people trust them to be good at. Most of the highest-visibility things like the AI Instagram accounts were disasters. Nobody thinks of Meta as a serious, general-purpose business ecosystem, and privacy-wise, I trust them even less than Google and Microsoft: there’s no way I’m trusting them with my work code bases. I think the smart move by Meta would be to ditch the sunk costs worries, stop burning money on this, focus on their core products (and new ones that fit their expertise) and design these LLM features in when they’ll actually be useful to users. Microsoft and Google both have existing tools that they’ve already bolstered with these features, and have a lot of room within their areas of expertise to develop more. Who knows— I’m no expert— but I think meta would be smart to try and opt out as much as possible without making too many waves.
- astrange 10mo agoJust because they have that doesn't mean they're going to use it for training.
- bdangubic 10mo agooh man… just because they have data doesn’t mean they will serve you ads :) Geeeez
- tonyhart7 10mo ago"Just because they have that doesn't mean they're going to use it for training." how noble is Meta upholding a right moral ethic /s
- astrange 10mo agoA very common thing people do is assume a) all corporations are evil b) all corporations never follow any laws c) any evil action you can imagine would work or be profitable if they did it. b is mostly not true but c is especially not true. I doubt they do it because it wouldn't work; it's not high quality data. But it would also obviously leak a lot of personal info, and that really gets you in danger. Meta and Google are able to serve you ads with your personal info /because they don't leak it/. (Also data privacy laws forbid it anyway, because you can't use personal info for new uses not previously agreed to.)
- nl 10mo agoLlama 4 wasn't great, but Llama 3 was. Do we all forget how bad GPT 4.5 was? OpenAI got out of that mess with some miraculous post-training efforts on their older GPT-4o model. But in a different timeline we are all talking about how great Llama 4.5 is and how OpenAI needs to recover from the GPT 4.5 debacle.
- Aeolos 10mo agoAs a counterpoint, I found GPT 4.5 by far the most interesting model from OpenAI in terms of depth and width of knowledge, ability to make connections and inferences and apply those in novel ways. It didn't bench well against the other benchmaxxed models, and it was too expensive to run, but it was a glimpse of the future where more capable hardware will lead to appreciably smarter models.
- colesantiago 10mo agoTake a look at JEPAs (Video Joint Embedding Predictive Architecture), SAM (Segment Anything), etc for Meta's latest research. https://ai.meta.com/vjepa/ https://ai.meta.com/vjepa/ https://ai.meta.com/sam2/ https://ai.meta.com/sam2/ https://ai.meta.com/research/ https://ai.meta.com/research/
- UltraSane 10mo agoMeta just published Segment Anything 3 and along with a truly amazing version that can create 3D models posing like the people in a photo. It is very impressive.
- asim 10mo agoIt was not always like this. Google was very secretive in the early days. We did not start to see things until the GFS, BigTable and Borg (or Chubby) papers in 2006 timeframe.
- okdood64 10mo agoBy 2006, Google was 8 years old. OpenAI is now 10.
- vlovich123 10mo agoGoogle publishes detailed papers of its architecture once it’s built the next version. AI is a bit different.
- rcpt 10mo agoPage Rank
- mapmeld 10mo agoWell it's cool that they released a paper, but at this point it's been 11 months and you can't download a Titans-architecture model code or weights anywhere. That would put a lot of companies up ahead of them (Meta's Llama, Qwen, DeepSeek). Closest you can get is an unofficial implementation of the paper https://github.com/lucidrains/titans-pytorch https://github.com/lucidrains/titans-pytorch
- informal007 10mo agoI don't think model code is a big deal compared to the idea. If public can recognize the value of idea 11 months ago, they could implement the code quickly because there are so much smart engineers in AI field.
- jstummbillig 10mo agoIf that is true, does it follow this idea does not actually have a lot of value?
- deleted 10mo ago[deleted]
- fancy_pantser 10mo agoStudent: Look, there’s hundred dollar bill on the ground! Economist: No there isn’t. If there were, someone would have picked it up already. To wit, it's dangerous to assume the value of this idea based on the lack of public implementations.
- lukas099 10mo agoIf the hundred dollar bill was in an accessible place and the fact of its existence had been transmitted to interested parties worldwide, then yeah, the economist would probably be right.
- NavinF 10mo agoThat day the student was the 100th person to pick it up, realize it's fake, and drop it
- cubefox 10mo agoThe author is listed as a "student researcher", which might include a clause that students can publish their results. Here is a bit more information about this program: https://www.google.com/about/careers/applications/jobs/results/93865849051325126-student-researcher-2026 https://www.google.com/about/careers/applications/jobs/resul...
- embedding-shape 10mo ago> Is there any other company that's openly publishing their research on AI at this level? Google should get a lot of credit for this. 80% of the ecosystem is built on top of companies, groups and individuals publishing their research openly, not sure why Google would get more credit for this than others...
- bluecoconut 10mo agoBytedance is publishing pretty aggressively. Recently, my favorite from them was lumine: https://arxiv.org/abs/2511.08892 https://arxiv.org/abs/2511.08892 Here's their official page: https://seed.bytedance.com/en/research https://seed.bytedance.com/en/research
- hiddencost 10mo agoEvery Google publication goes through multiple review. If anyone thinks the publication is a competitor risk it gets squashed. It's very likely no one is using this architecture at Google for any production work loads. There are a lot of student researchers doing fun proof of concept papers, they're allowed to publish because it's good PR and it's good for their careers.
- jeffbee 10mo agoUnderrated comment, IMHO. There is such a gulf between what Google does on its own part, and the papers and source code they publish, that I always think about their motivations before I read or adopt it. Think Borg vs. Kubernetes, Stubby vs. gRPC.
- hustwindmaple 10mo agoThe amazing thing about this is the first author has published multiple high-impact papers with Google Research VPs! And he is just a 2nd-year PhD student. Very few L7/L8 RS/SWEs can even do this.
- Balinares 10mo agoI mean, they did publish the word2vec and transformers papers, which are both of major significance to the development of LLMs.
- HarHarVeryFunny 10mo agoMaybe it's just misdirection - a failed approach ? Given the competitive nature of the AI race, it's hard to believe any of these companies are really trying to help the competition.
- timzaman 10mo agolol you don't get it. If it's published it means it's not very useful
- okdood64 10mo agoWhat about the Attention paper?
- Palmik 10mo agoDeepSeek and other Chinese companies. Not only do they publish research, they also put their resources where their mouth (research) is. They actually use it and prove it through their open models. Most research coming out of big US labs is counter indicative of practical performance. If it worked (too) well in practice, it wouldn't have been published. Some examples from DeepSeek: https://arxiv.org/abs/2405.04434 https://arxiv.org/abs/2405.04434 https://arxiv.org/abs/2502.11089 https://arxiv.org/abs/2502.11089
- abbycurtis33 10mo ago[flagged]
- CGMthrowaway 10mo ago[flagged]
- elmomle 10mo agoYour comment seems to imply "these views aren't valid" without any evidence for that claim. Of course the theft claim was a strong one to make without evidence too. So, to that point--it's pretty widely accepted as fact that DeepSeek was at its core a distillation of ChatGPT. The question is whether that counts as theft. As to evidence, to my knowledge it's a combination of circumstantial factors which add up to paint a pretty damning picture: (1) Large-scale exfiltration of data from ChatGPT when DeepSeek was being developed, and which Microsoft linked to DeepSeek (2) DeepSeek's claim of training a cutting-edge LLM using a fraction of the compute that is typically needed, without providing a plausible, reproducible method (3) Early DeepSeek coming up with near-identical answers to ChatGPT--e.g. https://www.reddit.com/r/ChatGPT/comments/1idqi7p/deepseek_and_chatgpt_gave_me_same_answer_what/ https://www.reddit.com/r/ChatGPT/comments/1idqi7p/deepseek_a...
- grafmax 10mo agoThat’s an argument made about training the initial model. But the comment stated that DeepSeek stole its research from the US which is a much stronger allegation without any evidence to it.
- nickpsecurity 10mo agoArxiv is flooded with ML papers. Github has a lot of prototypes for them. I'd say it's pretty normal with some companies not sharing for perceived, competitive advantage. Perceived because it may or may not be real vs published prototypes. We post a lot of research on mlscaling sub if you want to look back through them. https://www.reddit.com/r/t5_3bzqh1/s/yml1o2ER33 https://www.reddit.com/r/t5_3bzqh1/s/yml1o2ER33
- govping 10mo agoWorking with 1M context windows daily - the real limitation isn't storage but retrieval. You can feed massive context but knowing WHICH part to reference at the right moment is hard. Effective long-term memory needs both capacity and intelligent indexing.