Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
twotwotwo
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
twotwotwo
26d ago
The Jacobian one is great, and fits here -- 1) the counterexample came out as a tweet of a single expression which seems exceptionally hard to turn into something sensible, which crystallizes the lab's approach; 2) it links (like the
2.
▲
by
twotwotwo
26d ago
This is worsened by OAI/Ant's strategy of grabbing the glory and running rather than spending effort trying to to advance understanding of math. Tao, who is quite sophisticated in use of AI, has said a lot about this, including in
3.
▲
by
twotwotwo
2mo ago
Making things understandable is part of intelligence as much as producing the initial artifact is. Even if the proof checks out in Lean (or the code runs and passes QA) if it's a mess, it will be hard to use it to do anything further
4.
▲
by
twotwotwo
2mo ago
About the argument that a beam's strength:weight ratio worsens as it grows: do the artificial structures humans use, like house frames or bridges' trusses, recover strength:weight ratio? Would a theoretical giant with a skeletal s
5.
▲
by
twotwotwo
2mo ago
Hard to make big predictions, but it sure looks like at least this level of capability is going to be available in the open and relatively cheap to run. The 'floor' has gone up: today's model a bit behind SOTA is like model r
6.
▲
by
twotwotwo
2mo ago
One read is 1) they're getting a lot of traffic for Flash, 2) they've said they're updating Pro soon and expect that to lead to a traffic spike for Pro, but 3) that would leave them overloaded, so 4) they're going to rai
7.
▲
by
twotwotwo
2mo ago
The field of mathematics is smart about this and knows the difference between a pile of Lean code and understanding, and mathematicians try to get from the unintuitive explanations to something that makes more sense, e.g. https://
8.
▲
by
twotwotwo
2mo ago
Singer-songwriter Silvana Estrada named an album released last year Vendrán Suaves Lluvias (meaning There Will Come Soft Rains). She talked about how she and her brother read The Martian Chronicles as kids and she remembered the Sara Teasda
9.
▲
by
twotwotwo
2mo ago
Okay, a stretch for HN, but: singer-songwriter Silvana Estrada named an album released last year Vendrán Suaves Lluvias (meaning There Will Come Soft Rains). She talked about how she and her brother read The Martian Chronicles as kids and s
10.
▲
by
twotwotwo
2mo ago
I think this rates as too obvious to say among anyone remotely close to this, but worth noting there is a lot of distance between an exploit and any real confusion. Proofs aren't generally machine-read-only. An exploit of a proof syste
11.
▲
by
twotwotwo
2mo ago
This thread has some context. A proof-system researcher found some proof-system bugs and presented them a funny way: https://leanprover.zulipchat.com/#narrow/channel/270676-lean... A mathematically-inclined review
12.
▲
by
twotwotwo
2mo ago
There is a blog post waiting to be written (that I won't write) about the size/effort tradeoffs, and particularly how small models get some surprisingly good results with lots of turns and reasoning. DeepSWE will let you chart tur
13.
▲
by
twotwotwo
3mo ago
This is great--LLMs 'forgetting who they are' is one of the most uncanny things they do, and the note about why static benchmarks underperform human attackers is on point. One sort of wild idea: 'give words a color'. Tha
14.
▲
by
twotwotwo
4mo ago
Years ago work was bit by the analogous thing in MySQL. Like it usually does, it took a chain of events: - We wrote a cronjob to periodically DELETE for a retention policy on a table we'd just created. Most senior person on the team re
15.
▲
by
twotwotwo
4mo ago
To repeat, not a dig at FrontierCode, which is substantial progress in benchmarking. But I'd argue modeling the rest of process is tha(aaa)t valuable and becomes more so as coding capability progresses: Async agents interact on a l
16.
▲
by
twotwotwo
4mo ago
I'm liking the effort to make new, no-longer-saturated benchmarks. I'll also be a bit suspicious if some model aces it -- matching OSS maintainers' taste more often is a plausible improvement in quality but if they nail it
17.
▲
by
twotwotwo
4mo ago
If you worry about sending your data off for inference, Fireworks is one of the companies serving open models with solid performance and compliance/zero data retention sorted out. OpenCode supports them and many others. Cursor uses the
18.
▲
by
twotwotwo
4mo ago
We have a lot of synapses, but (agreeing with you) I don't find that sufficient to explain why humans (or animals!) do what we do. If you throw zillions of parameters at a problem with a weak architecture, you get really high-fidelity
19.
▲
by
twotwotwo
4mo ago
Whatever is the darker shade of blue in the bottom-right graph had a bump at the same time cost did. Perhaps that's output tokens (which include reasoning)?
20.
▲
by
twotwotwo
5mo ago
The fielded systems require something that wasn't there in the original model of zero-knowledge proofs. That could be as little as a trusted-enough public source of randomness: the prover makes their initial commitments, plays the veri
21.
▲
by
twotwotwo
5mo ago
It is kinda neat how the density can trickle down. When an individual SSD can hold tens of TBs, recent-gen drives can do millions of random reads/s each, and one socket can handle lots of RAM and many cores, it doesn't take the fa
22.
▲
by
twotwotwo
6mo ago
Kagi has it as an option in its Assistant thing, where there is naturally a lot of searching and summarizing results. I've liked its output there and in general when asked for prose that isn't in the list/Markdown-heavy "
23.
▲
by
twotwotwo
6mo ago
You could model more of the process: the dev's work as well as the model's, and the cost of catching a bug later or deploying it live. Those tasks push me further towards smaller tasks in general. (And they make the Gas Town type
24.
▲
by
twotwotwo
6mo ago
The topic of cooldowns just shifting the problem around got some discussion on an earlier post about them -- what I said there is at https://lobste.rs/s/rygog1/we_should_all_be_using_dependency... and here's
25.
▲
by
twotwotwo
6mo ago
There is nothing specific to the role-switching here (as opposed to other mistakes), but I also notice them sometimes 1) realizing mistakes with "-- wait, that won't work" even mid-tool-call and 2) torquing a sentence around
26.
▲
by
twotwotwo
6mo ago
I agree with the addition at the end -- I think this is a model limitation not a harness bug. I've seen recent Claudes act confused about who they are when deep in context, like accidentally switching to the voice of the authors of a p
27.
▲
by
twotwotwo
7mo ago
One potential application I briefly had hope for was really good power loss protection in front of a conventional Flash SSD. You only need a little compared to the overall SSD capacity to be able to correctly report the write was persisted,
28.
▲
by
twotwotwo
7mo ago
This is fascinating, and makes me wonder what other things that 'should' be impossible might just be waiting for the right configuration to be tried. For example, we take for granted the context model of LLMs is necessary, that al
29.
▲
by
twotwotwo
8mo ago
For folks that like this kind of question, SimpleBench ( https://simple-bench.com/ ) is sort of neat. From the sample questions ( https://github.com/simple-bench/SimpleBench/blob/main/simpl
30.
▲
by
twotwotwo
8mo ago
Yeah, this(-ish): there are shipping models that don't eliminate N^2 (if a model can repeat your code back with edits, it needs to reference everything somehow ), but still change the picture a lot when you're thinking about, say
More ›