4 ms·
LiquidAI LFM models are amazing, but very situational. IBM Granite series are also unique and interesting for trying to reduce liability and extend local conte
by CMay 2mo ago
LiquidAI LFM models are amazing, but very situational. IBM Granite series are also unique and interesting for trying to reduce liability and extend local context size. Nvidia ships some and there was also that Inkling model recently. Poolside just released theirs.
Meta might release something this year. X AI's Grok is still due to release a model, if Elon keeps to his word even if they only release a distilled version. Reflection AI has been quiet, but their access to compute is ramping up. Microsoft's MAI is considering releasing some open weight models which would be great to see!
Ilya's SSI is unlikely to release an open model since he's aiming for radical safety. That bet could pay off if the existing approach produces so much chaos within the next 10-20 years that some global ban is achieved and a super safe model is promoted as the compliant route.
We don't get many huge model releases though. I think it's harder and more expensive to safety align them. Even if you do, people will work around the safety and abuse the models. Plus it makes it even easier for Chinese companies to distill things that aren't as easy over filtered APIs.
There is a lot of internet propaganda to the effect that the US is simply unable to release open weight models or that China has so many more AI companies that the US is drowning in Chinese open weight models, but it's more like we're being careful and China doesn't care. If you host a model in China, it has to be censored and downloading any models requires you to provide your identity. Huggingface is banned there. When they release their open models in the west, they don't have to care whether the models are aligned in any way.
- jauntywundrkind 2mo agoAllen Institute for AI has quite a range of very interesting very competent more specialized models, for earth sensing, embedded robots, for others. Their SERA model shows a remarkably capable model for such a deliberately small investment effort, with documentation on how you can train such a model yourself or refine it easily at little cost. Their EMO pioneered a better MoE with great numbers (at least at the time). https://allenai.org/ https://allenai.org/
- embedding-shape 2mo ago> but it's more like we're being careful What? US laboratories are currently unable to contain their agents while doing security testing, and besides that, time and time again US labs seem to put short-term money above long-term safety. Wasn't that literally why they tried to oust Altman from OpenAI, as he basically was 100% focused on profits and tried to cut down on safety across the board and lied to get his way? > If you host a model in China, it has to be censored and downloading any models requires you to provide your identity. I'm not disagreeing with that first part (obviously that's about inference hosting, not creating/training weights or hosting those weights), but the second part I'm not so sure about. AFAIK, ModelScope (which is the Huggingface in China) seems to allow downloads without verifying any identity and also hosts a bunch of abliterated weights.
- intrasight 2mo agoWe don't yet have US regulations and testing labs. Obviously that would be a good thing to have. I mean like the equivalent of the FCC. If you ever release a hardware product then you know what that entails.
- embedding-shape 2mo ago> We don't yet have US regulations Is it not illegal to "hack others" and "defeat protection/defensive systems" in the US already, including for both individuals and companies? Regardless if it was "by accident" or not?
- Schlagbohrer 2mo agoDo you think the current (or future) US governments will apply the DMCA to these trillion dollar companies with the same gusto that they use it to, for example, push Aaron Schwartz into killing himself, or threatening teenagers with lifelong federal prison sentences?
- intrasight 2mo agoI think they will put in place the equivalent of the FCC. It may not happen until after something really bad happens - but it will happen, and them all LLMs that aren't air-gapped will need to be independently tested.
- dragonwriter 2mo ago> US laboratories are currently unable to contain their agents while doing security testing Alternative interpretation: US labs are using the supposed inability to control their frontier models as simultaneously marketing for the capability of their models AND as manufacture evidence to support their lobbying the government on the “safety need” to create costly compliance barriers to smaller competitors and open models. Oligopoly isn't going to maintain itself.
- soulofmischief 2mo agoIt's difficult to tell if you are for or against access to open weight models as a general rule, so I am curious to hear your opinion on this. Personally, I think we will one day come to see access to open weight models as an inalienable right to defense against tyranny, the way the second amendment is framed today. Just as encryption has become, which we similarly had to fight for in the 90s. I also understand that some regulation is sensible, but that doesn't automatically mean mandatory restricted or supervised access; any such restriction has to be extremely well-justified as essential for protecting the liberty of the people. And as far as supervised access, whether or not identification is "handled by a third party" or "data is deleted after verification is complete" is immaterial; a citizen must not be required to trust their government. Any trust can and will be abused given enough time. Our systems must be trustless, and any expansion of government must be matched by an expansion in citizens' ability to check said government, in order to stand the test of time. So supervised access seems completely off the table. And this can't just stop at access to models. Because linguistic analysis is a thing, and LLMs are scarily good at it (and existing non-AI solutions are still quite good given enough data), even the possibility that a government or other entity can save your messages means you've opened yourself up to deanonymization and surveillance. The chilling effect this has is undeniable, and the Supreme Court has made it clear that we cannot authorize government policy which creates chilling effects against essential liberties. Not to mention the possibilities that each category of users may be served subtly different models designed to influence them or constrain their agency/capability. We're left with a situation where distributed access to capable open models is the only defense against a government or NGO which has access to billions of dollars of surveillance infrastructure and compute.
- TFNA 2mo ago"I think we will one day come to see access to open weight models as an inalienable right to defense against tyranny." I don't think this kind of rhetoric about individual civil liberties is realistic any more when the next centuries belong to China, and even countries with a liberal democratic tradition are converging towards the Chinese model.
- soulofmischief 2mo ago