Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
thatguysaguy
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
thatguysaguy
29d ago
part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing.
2.
▲
by
thatguysaguy
29d ago
presumably that's a safety evaluation not a training setting
3.
▲
by
thatguysaguy
3mo ago
Scott's parenting posts are some of his best
4.
▲
by
thatguysaguy
3mo ago
The contrast between the example screenshots and the standard internet behavior in the live demo is hilarious
5.
▲
by
thatguysaguy
4mo ago
The python overhead of launching big ML jobs is nontrivial, so I think speeding that up would be meaningful. (I mean the initial tracing and other setup, not things once the GPUs are actually doing the work).
6.
▲
by
thatguysaguy
5mo ago
It looks like people in this thread are confusing fleet utilization and MFU. If they're doing a lot of RL, it's really not surprising to see such low numbers.
7.
▲
by
thatguysaguy
5mo ago
That's b/c the people working on Gemini serving are in GDM.
8.
▲
by
thatguysaguy
5mo ago
What sort of workloads are you thinking of?
9.
▲
by
thatguysaguy
6mo ago
What is up with people saying you cannot prove a negative? Of course you can! (At least in formal settings) For example it's extremely easy to prove there is no square with diagonals of different lengths. I'm the hard end, Andrew
10.
▲
by
thatguysaguy
7mo ago
Ah dang. When I did this I also thought the length bug was intentional but I didn't figure it out before I started my new job, so I dropped the puzzle.
11.
▲
by
thatguysaguy
7mo ago
Maybe I missed something, but I see little evidence that there is a concerning ability to deanonymize. Many people post under a pseudonym but then link to their GitHub etc. In fact by construction the HN dataset _only_ consists of people w
12.
▲
by
thatguysaguy
8mo ago
You can just try other svgs, I got some pretty good ones. (*Disclaimer: I work for Google, but also I have zero idea about what they trained deepthink on)
13.
▲
by
thatguysaguy
10mo ago
TPUs predate LLMs by a long time. They were already being used for all the other internal ML work needed for search, youtube, etc.
14.
▲
by
thatguysaguy
10mo ago
I'm actually not talking about whether the PR works or was tested. Let's just assume it was bug-free and worked as advertised. I would say that even in that situation, they should not accept the PR. The reason is that no one is th
15.
▲
by
thatguysaguy
10mo ago
A big part of software engineering is maintenance not just adding features. When you drop a 22,000 line PR without any discussion or previous work on the project, people will (probably correctly) assume that you aren't there for the lo
16.
▲
by
thatguysaguy
11mo ago
It's a volunteer run project... Saying that they have a duty to do anything other than what they want is quite strange.
17.
▲
by
thatguysaguy
11mo ago
Verification via LLM tends to break under quite small optimization pressure. For example I did RL to improve <insert aspect> against one of the sota models from one generation ago, and the (quite weak) learner model found out that it
18.
▲
by
thatguysaguy
1y ago
FAIR is not older AI... They've been publishing a bunch on generative models.
19.
▲
by
thatguysaguy
1y ago
Back when BERT came out, everyone was trying to get it to generate text. These attempts generally didn't work, here's one for reference though: https://arxiv.org/abs/1902.04094 This doesn't have an expli
20.
▲
by
thatguysaguy
1y ago
I would recommend going and reading what the BlueSky leadership actually wrote, rather than this post's summary of it.
21.
▲
by
thatguysaguy
1y ago
Why would you think that deepseek is more efficient than gpt-5/Claude 4 though? There's been enough time to integrate the lessons from deepseek.
22.
▲
by
thatguysaguy
1y ago
37 billion bytes per token? Edit: Oh assuming this is an estimate based on the model weights moving fromm HBM to SRAM, that's not how transformers are applied to input tokens. You only have to do move the weights for every token during
23.
▲
by
thatguysaguy
1y ago
Joel's blog in general is an extremely great read. I highly recommend subscribing.
24.
▲
by
thatguysaguy
1y ago
At least part of is is that the capex for LLM training is so high. It used to be that compute was extremely cheap compared to staff, but that's no longer the case for large model training.
25.
▲
by
thatguysaguy
1y ago
I both got a job through such a thread, and have now seen the other side of the applicant pipeline. The average applicant (in general, idk about HN in particular) is not very strong! Especially true when you consider the alternative of pres
26.
▲
by
thatguysaguy
1y ago
> Do you think the students in that poll had really thought about the credibility of their university when voting? That's fair, and I'm not sure of course. I guess a more interesting question would be what if there was another
27.
▲
by
thatguysaguy
1y ago
I think the author doesn't understand the example correctly, although to be fair I don't think the professor put the most important option on there either. Imagine there are two schools, one gives all students a 95% in all their c
28.
▲
by
thatguysaguy
1y ago
Yeah I wouldn't make a deal like this with someone who is operating in bad faith... The cases I've seen of this are between public intellectuals with relatively modest amounts of money.
29.
▲
by
thatguysaguy
1y ago
Some people do actually have end of the world bets out but you have to structure it differently. What you do is the person who thinks the world will end is paid cash right now, and then in N years when the world hasn't ended they have
30.
▲
by
thatguysaguy
1y ago
Intuitive yes, but since P != PSPACE is still unproven it's clearly hard to demonstrate.
More ›