Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
vessenes
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
vessenes
5d ago
I was going to say the reverse - claude has been the less satisfying normalized by benchmark for me in the last year. Both astra and fable have their quirks, but I am 90% codex this year up from 10% last year.
2.
▲
by
vessenes
5d ago
Nice to see this release cadence increasing and some continued improvement in quality. I am guessing these models are basically still outcomes of the cursor team integrating with the massive amount of compute they now own: I’d imagine we wi
3.
▲
by
vessenes
7d ago
Agreed. Another difficulty here is there are not good benchmarks for this new architecture yet, so it’s easy to potshot and snipe, where jev seems to be pretty broadly intelligent/at least have had a lot of rl in different domains. We
4.
▲
by
vessenes
8d ago
A counterpoint - I was told a story by one of my professors in the late 1990s, about one of his professors -- he'd written a thesis, gotten hired somewhere like Princeton, and taught there for a few years as Dr. <Somebody>. One
5.
▲
by
vessenes
13d ago
Well, duh. If you could do this with Opus 4.8, we would know. When Astra’s successor is 2-3x better at math research, and the internal teams say “we believe we will get there,” I’m inclined to believe the insiders.
6.
▲
by
vessenes
13d ago
Sorry, I don't understand what this means. What does it mean?
7.
▲
by
vessenes
13d ago
I think that implies they're seeing unusual new subscription demand, yes? That's how I'd read it, not least because I worked through four resets this week on Astra, which is un unbelievable amount more inference than I'v
8.
▲
by
vessenes
13d ago
Prices are moving inference to highly profitable, oAI seems to have solved their training problems, and they're currently competing nicely with Anthropic on the coding side. I'd wait, too -- why fight this stuff out in public when
9.
▲
by
vessenes
14d ago
“We” does a lot of work here. You keep using that word. I don’t think that word means what you think it means.
10.
▲
by
vessenes
15d ago
Shenanigans is an inaccurate word; it implies underhanded behavior that's hidden / concealed. I think "to act so aggressively" is more balanced. Here's my answer: If you think we're getting to AGI in the next 9
11.
▲
by
vessenes
16d ago
Yep, it's possible. But, we have literally no idea how much a set of prompts would impact training as far as general usefulness. I don't think we even know if Tristan's said he allowed training on his prompting or not. This i
12.
▲
by
vessenes
16d ago
This is speculation right now. The idea would be that if someone used the product and granted training rights, which is the default for many subscription levels, then some knowledge would have been imparted into the general weights of the n
13.
▲
by
vessenes
17d ago
Thanks! It's just infra, so you could do what you wanted. The clients include a safety reminder, and the datastructures include "unsafe" in the name of the inputs passed around, but that's just a little hygiene. There re
14.
▲
by
vessenes
17d ago
I built what I think is a pretty good library and set of tools for this earlier this year -- https://github.com/corpollc/qntm (or `uvx qntm --help`); it includes a cli, python and typescript libraries, and works out of
15.
▲
by
vessenes
17d ago
Not arguing against open science - it's super valuable. I'm saying that pearl clutching by people reading Tao isn't useful, because it misses some long history which tells us that this kind of science has been seen as fundame
16.
▲
by
vessenes
17d ago
That’s not untrue. But it’s also a misstatement of mathematical history. Many leading mathematicians historically have been highly competitive — Gauss comes to mind. Woe betide the lesser intellect that sent Gauss some ideas. The Newton Lei
17.
▲
by
vessenes
18d ago
If those researchers did not opt out then training data might go in. I think it’s a courteous acknowledgement; as was reaching out and examining the direction of proofs themselves. At stake here is a particular mathematician dynamic - ego,
18.
▲
by
vessenes
19d ago
I can’t read your original comment, but when the chief scientist of a company that just released six month old software that can use Kicad on your computer to design a circuit board and have it created and shipped to you tells you he thinks
19.
▲
by
vessenes
19d ago
The situations aren’t equivalent - luckily in my opinion because the stakes with nuclear are much higher. von Neumann constructed a multinational game theory approach appropriate for weapons. AGI is a much harder problem to corral because t
20.
▲
by
vessenes
19d ago
I’m like a 2(.5?) there - I don’t think ASI will care about my kids better than I will for some definitions of better, for instance, and I feel very fuzzy and vague about what actual differences in qualia between me and ASI would yield in t
21.
▲
by
vessenes
19d ago
No. These parties are composed of people who most definitely think this way, and therefore will have distinct goals and interests when presented with opportunities. That’s reality quite aside from how a game theorist assesses the situation.
22.
▲
by
vessenes
20d ago
This is a good essay, and makes me hopeful. I’m on the record saying that it is extremely dangerous to slow down because the race for AGI is a zero-trust game — defections pay - and combined with a compounding returns model on defection,
23.
▲
by
vessenes
20d ago
Anthropic’s Mechinterp did some very fine work on this. TLDR - you can; you train a decoder on neuralese to english and then add a loss function for a roundtrip of english -> neuralese -> english (or possibly n -> e -> n? I don’
24.
▲
by
vessenes
21d ago
Looked to me like it missed the chain though.
25.
▲
by
vessenes
21d ago
Those tend to be under so-called ‘max’ modes, or ‘ultra’ - where you scale up inference time compute and then choose amongst your answers. Astra is new weights.
26.
▲
by
vessenes
21d ago
It’s already running in Hawthorne, CA
27.
▲
by
vessenes
21d ago
True, .. and. In this case, the original proof is considered rigorously checked, so finding a bug in the kernel would be nice to know about, but in my opinion would not take away from the accomplishment (FLT in lean using agents) nor the ma
28.
▲
by
vessenes
22d ago
I'd like to note that we should remember a formalized Lean proof does have value in that it enters the pantheon of true things other Lean proofs can rely on. Agreed that for the humans, descriptions and being able to 'grok' t
29.
▲
by
vessenes
22d ago
Pretty efficiently, apparently, since it saturated ARC-AGI-3 in half of the predicted time, and according to the Chollet blog post on the fly created dense DSLs to describe and analyze individual games.
30.
▲
by
vessenes
24d ago
Weave robotics (no affiliation) claims they are folding over 1,000 lbs of laundry a week in homes right now: https://www.weaverobotics.com . I have no idea how well these work.
More ›