Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
modeless
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
91.
▲
by
modeless
2mo ago
This guy actually built an algae farm to test how much you need to breathe. https://youtu.be/AAbyUaLN2QA
92.
▲
by
modeless
2mo ago
Cost, weight, durability. For cameras, bandwidth.
93.
▲
by
modeless
2mo ago
I actually like Tau's videos a lot better than Google's. An agile robot getting out of a car carrying a tote is actually novel and difficult and useful. In contrast, what Google is showing here is mostly glorified pick and place w
94.
▲
by
modeless
2mo ago
I do not. A ban on capable open-weight models for an indefinite period of time falls into the category of bans on open-weight models. If you wanted Anthropic's statement to be true you would need to qualify "ban" or "ope
95.
▲
by
modeless
2mo ago
My point is that advocating a de facto ban on capable open source models is inconsistent with Dario's statement here that "Anthropic has never advocated for a ban on open-weights models." Call a spade a spade.
96.
▲
by
modeless
2mo ago
OK that is a position they could take but my point is that's inconsistent with "Anthropic has never advocated for a ban on open-weights models". What you're describing is a ban on capable open-weights models until some f
97.
▲
by
modeless
2mo ago
Any sufficiently capable open weights model would fail "safety" testing though, as any "safeguards" of the sort Anthropic likes could be removed. That's just another way of saying they want a ban on capable open sou
98.
▲
by
modeless
2mo ago
> All sufficiently capable models, open and closed, should go through mandatory safety testing What happens if a model fails the test? Surely one can use Kimi K3 for evil, somehow or other. What now? "Mandatory safety testing"
99.
▲
by
modeless
2mo ago
Massively simplifying here but when doing one rasterization pass you process a flat list of triangles, in order, once. When raytracing a camera view you query a database of triangles, at least once for each pixel. To avoid scanning the whol
100.
▲
Ilya Sutskever: "Time to scale that SSI"
(twitter.com)
3 points
by
modeless
2mo ago
|
0 comments
101.
▲
by
modeless
2mo ago
I imagine they built these a long time ago (anticipating fewer test flight problems) and have probably improved the design since then making these obsolete. They are going to build thousands so they can't be too precious.
102.
▲
by
modeless
3mo ago
They state the puzzle is "Witness-like" which I assume means that it follows the rules from the well-known puzzle game "The Witness" which Opus definitely knows.
103.
▲
by
modeless
3mo ago
Also they plan to catch the ship next time! Excitement guaranteed!
104.
▲
by
modeless
3mo ago
Lots of missing tiles still. Best case is it's all infant mortality and replacing just those tiles will result in a durable heat shield for many more flights. Worst case is every flight loses that many tiles and the ablative underneath
105.
▲
by
modeless
3mo ago
Personally I hope to use a lot less software in the future. Anything I want to know or do online I'll just ask an AI and it can wade through the sea of cookie popups or UI redesigns or whatever to get it done.
106.
▲
by
modeless
3mo ago
Successful Raptor on-orbit relight demo! That should clear the next flight to go to orbit and deploy operational satellites! Big milestone for the program. And the ship landing was so soft it didn't even explode! That's never happ
107.
▲
by
modeless
3mo ago
Fully agreed. At some point you have to draw a line and say that the rest is the responsibility of the kernel and hardware and user, and I think Fil-C drew that line in the right place.
108.
▲
by
modeless
3mo ago
I'm not sure about the speed of ASAN, it may be comparable but doesn't guarantee memory safety. It's only for catching mistakes and not secure against an adversary. Valgrind is dramatically slower.
109.
▲
by
modeless
3mo ago
The whole point of Fil-C is that it is fast enough to consider using in production for some applications while still guaranteeing memory safety. We already have ASAN and Valgrind and other tools for development purposes, that's not wha
110.
▲
by
modeless
3mo ago
From what I read the actual escape was through a proxy that allows downloading Python packages from the internet. It's not supposed to allow general internet access but the AI found a previously unknown vulnerability in it. That is har
111.
▲
by
modeless
3mo ago
Adding Fil-C-like runtime checks to Rust is definitely an interesting direction. As I mentioned upthread. It's not just the availability of the safe API that's interesting, though, but also the prohibition on using the unsafe API
112.
▲
by
modeless
3mo ago
Better include Rowhammer too. Maybe Fil-C should run a test and refuse to start on any system with bad RAM or unpatched CPU errata. It could also monitor the voltage to protect against undervolting attacks. And you'll need some cosmic
113.
▲
by
modeless
3mo ago
Fil-C has access to a memory safe API called mmap with a lot of the capabilities of the mmap system call, while Rust's safe subset does not. Rust allows you to use the mmap system call unsafely, while Fil-C does not. I feel these are b
114.
▲
by
modeless
3mo ago
Rust doesn't runtime validate that your usage of syscalls is memory safe, while Fil-C does. For example you can call mmap in Fil-C and it is still guaranteed to be memory safe, while in Rust you can easily violate memory safety by call
115.
▲
by
modeless
3mo ago
It will continue to be valuable as a cost and speed benchmark long after it is saturated at the high end. And they are already working on ARC-AGI 4 and thinking about going even farther.
116.
▲
by
modeless
3mo ago
Opus is cheaper than Fable. They could probably replace Fable with Opus but why? They would be churning customers to different models for no reason. Even if a model scores better on benchmarks it can always regress in your specific use case
117.
▲
by
modeless
3mo ago
Yes, I think it indicates real progress in fluid intelligence. Clearly these models are making huge strides in usefulness which are well correlated with their ARC-AGI scores. I don't think this is benchmaxxing. These companies are lock
118.
▲
by
modeless
3mo ago
Wow, 30% on ARC-AGI-3 for $20k total. Huge jump from GPT-5.6's 7.8% at $20k per task. I continue to believe ARC-AGI measures something different and important compared to other benchmarks.
119.
▲
by
modeless
3mo ago
Removing the "evil" profit motive has always been an argument for socialism. If it would work for healthcare, why not everything?
120.
▲
by
modeless
3mo ago
Some people thought it was a good idea to cap [administration cost + profit] of health insurance companies as a percentage of premiums. As a result, health insurance administrators have two ways to increase their compensation: they can comp
More ›