Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dooglius
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
dooglius
5d ago
> Instead of finding a nerfed model, after six weeks of reconstructing wire logs, parsing transcripts, analyzing output tokenization, and staring at data, I found a much deeper issue. The model identity had remained the same, but the inf
2.
▲
by
dooglius
10d ago
It would appear no one else submitted the posts in question, so the most obvious answer is the account is finding interesting articles that no one else has found. It would be one thing if it won out of many dupe posters, but that doesn'
3.
▲
by
dooglius
10d ago
There's a good chance it'll be scrubbed now that it's frontpaged here
4.
▲
by
dooglius
13d ago
Tangent, but that's not in the movie. It was in Clarke's contributions to the script and novelization, but Clarke and Kubrick had a bitter falling out over different visions and Kubrick took out much of Clarke's stuff from th
5.
▲
by
dooglius
15d ago
Both of your links are basically ads, or at least the titles make them sound like it. > Lenovo showed a model that incorporates solid-state air cooling. That technology has been around for a while but it looks like it matured and is prod
6.
▲
by
dooglius
15d ago
I'm not sure what you mean by "verify" here. I could run the lean verifier as could anyone else. Maybe I could write my own proof checker and do a purely mechanical translation into my own thing, though I don't think tha
7.
▲
by
dooglius
15d ago
There have been several threads and developments on this over the past few days, including statements from the primary subjects involved. Third-hand instagram comments are not really the best source to be bringing in.
8.
▲
by
dooglius
15d ago
Weren't the agents massively parallel, whereas the lean verifier presumably is not? Also, I presume said agents were themselves running the verifier on their own parts many times.
9.
▲
by
dooglius
15d ago
I mean, I have a bachelor's in math and I don't imagine I could begin to understand either the human or LLM proofs without a massive investment of time and effort.
10.
▲
by
dooglius
16d ago
You would want to use your own benchmark, certainly, not something publicly known. But outside of that I don't think "this is a benchmark" is such an easy category to determine -- a benchmark should be similar to a typical sc
11.
▲
by
dooglius
16d ago
Run a benchmark with a large number of samples, rerun a few days later. Compare results, use statistics to see if there's a statistically significant difference.
12.
▲
by
dooglius
17d ago
Do you have hard evidence of this assertion?
13.
▲
by
dooglius
25d ago
Presumably the parent does not want to have to trust Aurora to do that
14.
▲
by
dooglius
25d ago
It isn't exactly hard for a bad actor to come up with that prompt
15.
▲
by
dooglius
25d ago
Out of curiosity, does this reset/modify the timestamps of the old posts? I was confused at reading comments from "hours ago" I thought for sure I'd read days ago.
16.
▲
by
dooglius
28d ago
I'm not advocating for this, but it's a example of a position reasonable enough I'd expect to see some level of support for. I'd view it as less extreme, for example, than requiring all new security-relevant code be in R
17.
▲
by
dooglius
28d ago
Ex: Security-critical code contributions should be scrutinized via state of the art tooling, including but not limited to fuzzers, linters, and adversarial LLM review. For non-security-critical code, use of LLMs is encouraged but not requir
18.
▲
by
dooglius
28d ago
The voting seems to have been pretty much linear to how pro-LLM they were. So it's interesting that all of the proposals were essentially anti-LLM and the chosen one was the mostly neutral, only slightly anti-LLM one. The absence of an
19.
▲
by
dooglius
29d ago
In fact, this is what happens in the case of laptop displays
20.
▲
by
dooglius
29d ago
I don't think you actually had to visualize the sheep for it to work, I never thought of that as being explicitly required
21.
▲
by
dooglius
1mo ago
What are you suggesting and how would it be different than how SIGFPE already works?
22.
▲
by
dooglius
1mo ago
I thought the point of counting sheep was to do something boring and repetitive that puts your brain into a relaxed state
23.
▲
by
dooglius
1mo ago
It's tricky because even if you've seen the pictures before, it's not obvious whether the formal problem is defined in terms of minimizing the size of the big square or maximizing the size of the little squares (which are equ
24.
▲
by
dooglius
1mo ago
Responding to a couple comments here: there is no picture or new arrangement of squares because those are _upper_ bounds for the problem. The best known arrangement, i.e. the best known upper bound, has not changed.
25.
▲
by
dooglius
1mo ago
Oh wow I remember having read/heard that it was helpful for memorization, never occurred to me to think that there was an impact for some listeners too.
26.
▲
by
dooglius
1mo ago
Unless your _ads_ are going to compete with your competitors' then that's not going to be effective. Ironically middlemen are probably solution here, you'd be better off going to trade shows and working with department stores
27.
▲
by
dooglius
1mo ago
See the last 3 sentences of GP's post
28.
▲
by
dooglius
1mo ago
The reference isn't intended to be interpreted as the correct answer, it's the answer given by a control model that was trained on a broader corpus.
29.
▲
by
dooglius
1mo ago
Isn't the underlying question proved impossible by Godel's incompletness theorem?
30.
▲
by
dooglius
1mo ago
I'm confused, the Opus 5 announcement said it was (outside a few special cases) better than Mythos/Fable, but the Opus prompt here seems to suggest the opposite?
More ›