3 ms·
I'm interested to see where they want to land performance-wise (i.e. which point they choose on the scaling curve) and the niche they want to carve. They have a
by lithobraking 2mo ago
I'm interested to see where they want to land performance-wise (i.e. which point they choose on the scaling curve) and the niche they want to carve. They have a decent ways to scale beyond trinity large, in paticular on posttrain/RL before they are competitive with open-weights, especially internationally.
Deepseek is explicitly banned [1] at LLNL and I wouldn't be suprised if there's a blanket ban on all Chinese models. But nowadays models like tera/luna could fill this area of the pareto front, and LANL already runs openai models on their clusters [2]. Maybe it's in custom SFT/RL, for instrument control or sensitive topics? But you'll still have to compete with frontier models + a harness.
I would have also liked to see a carrot tied to their offer. It'll be hard to get teams to contribute RL gyms or curated text. But throw in a "we'll fund a postdoc/student to do that" and I think you'd have teams scrambling to apply.
[1] https://hpc.llnl.gov/about-livermore-computing/ai-ml-lc/lc-llm-model-download-decision-guide?utm_source=chatgpt.com https://hpc.llnl.gov/about-livermore-computing/ai-ml-lc/lc-l...
[2] https://www.energy.gov/nnsa/articles/nnsas-los-alamos-national-laboratory-launches-frontier-ai-models-venado-supercomputer?utm_source=chatgpt.com https://www.energy.gov/nnsa/articles/nnsas-los-alamos-nation...
- cududa 2mo agoI'd actually suggest a great starting point would be a local command reviewer LLM. Could ostensibly be a modern AV type thing. Particularly seeing this lately has driven the need home deeper to me: https://x.com/chrisbanes/status/2085341561609425230?s=20 https://x.com/chrisbanes/status/2085341561609425230?s=20 An open weight tool call auto-reviewer, has all sorts of achievable scaling curve milestones.
- unethical_ban 2mo agoThat's interesting a locally hosted LLM would be banned. I'm assuming locally hosted is included. Do they think it's been trained to sabotage equipment?
- SyneRyder 2mo agoI don't think we know either way, but we do know at one point Anthropic would silently sabotage requests, Stuxnet style: https://simonwillison.net/2026/Jun/10/if-claude-fable-stops-helping-you/ https://simonwillison.net/2026/Jun/10/if-claude-fable-stops-... I can imagine if the US were already doing that as a safeguard, they would assume their "adversaries" (to use Anthropic language) were doing the same as well, whether that were true or not, and therefore would not trust those models even if locally hosted.
- petcat 2mo agoLocally running LLM means nothing because "open weight" models are still inscrutable. It's like bringing a dog home from the rescue and just hoping that it doesn't have the tendency to bite kids in the face. You just can't know. All you can do is try to add some new training telling it not to bite kids.