Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
naasking
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
55 ms
·
61.
▲
by
naasking
2mo ago
> We know exactly how attention layers work and how they produce the next word as well as draw them from larger feature spaces. This is not what's meant by the statements that we don't know how LLMs work. Explain why LLMs are s
62.
▲
by
naasking
2mo ago
The search space is far too large for a mere order of magnitude to make any difference at all.
63.
▲
by
naasking
2mo ago
> we know everything about how LLMs work No we don't. That we understand the low level mechanics of a system doesn't mean we understand how any high level phenomena emerge from those low level mechanics. This is as true for qua
64.
▲
by
naasking
2mo ago
I don't think that's correct. Increasing parameter count increases capabilities in all domains per the scaling laws. Models are larger than they were years ago, so capabilities in all domains must necessarily be better. This doesn
65.
▲
by
naasking
2mo ago
Bribes are not speech though, and advertising is.
66.
▲
by
naasking
2mo ago
It's interesting to claim that it would be difficult to explain why we punish successful crime more. A successful murder creates more suffering (victim's friends and family), and ostensibly the loss of a productive member of socie
67.
▲
by
naasking
2mo ago
It does matter though. If you want to murder someone by hitting them with a plushie, you're not going to get charged with attempted murder because it's not possible that that would ever work. There must be justification that the c
68.
▲
by
naasking
2mo ago
The margins for groceries are objectively thin. The only way to provide food at lower prices is to provide a worse good or service, eg. less variety, less quality, less availability, etc. You will see all of these outcomes in NYC, if anyone
69.
▲
by
naasking
2mo ago
Yes, but there is little evidence this had a meaningful effect on votes.
70.
▲
by
naasking
2mo ago
They're important everywhere of course, but especially on mobile. If AI researchers figure out how to offload knowledge and expertise from reasoning weights, then a core reasoning ASIC linked to the knowledge would totally rock.
71.
▲
by
naasking
2mo ago
> I can't see how the model could include the actual subjective human experience. People who say LLMs have subjective experience aren't saying they have human-type subjective experience. Nobody who sees an LLM express hunger wh
72.
▲
by
naasking
2mo ago
Even mechanistic models generate interesting discussion. How many years have we discussed Turing machines and the lambda calculus? Almost a century of great work came out of those. The reason I insist on mechanistic models is because the or
73.
▲
by
naasking
2mo ago
> Dunno about the parent commenter, but I personally interpret the concept as having a hidden representation of self that is continually tended to I don't see why an LLM could not have a sense of identity or personality while it
74.
▲
by
naasking
2mo ago
Making definitive claims about whether LLMs do or do not have specific properties absolutely does require precise definitions of those properties that can be used to evaluate those questions. Merely hand waving that LLMs didn't undergo
75.
▲
by
naasking
2mo ago
> The qualia themselves, even those that are quite abstract, are rooted in our physical presence and evolution. There is no objective evidence of qualia. All evidence of qualia are vocal or other expressions of belief in qualia. Percepti
76.
▲
by
naasking
2mo ago
> AI has no self-awareness What is your mechanistic model of self awareness that yields this conclusion? > It's a tool Does your model suggest that tools can't have self awareness?
77.
▲
by
naasking
2mo ago
What sentence from my post implies that?
78.
▲
by
naasking
2mo ago
All of them depend on temporarily suspending concurrency in order to synchronize. Even atomic exchange operations are like this at the hardware level.
79.
▲
by
naasking
2mo ago
> 3) and also by how Microsoft operates (e.g. certainly not using any modern C). Are you suggesting "modern C" is less prone to memory safety issues, and that if MS used "modern C" then that would meaningfully reduce
80.
▲
by
naasking
2mo ago
You wouldn't have to worry so much about running up to date software if memory safety were pervasive.
81.
▲
by
naasking
3mo ago
An opinion is a particular subject's belief [1]. That's literally the definition. "X believes Y at time T", is a fully specified fact and completely follows from the definitions of all terms involved. [1] a view, judgmen
82.
▲
by
naasking
3mo ago
This seems like a very long winded way to agree with my statement that opinions can only be considered facts when modelled as time series data.
83.
▲
by
naasking
3mo ago
Opinions can only be seen as facts when modelled as time series data. Verifying an opinion X at time T does not entail X at time T+1.
84.
▲
by
naasking
3mo ago
Comments sometimes help the LLM think through a problem. They are particularly useful if you have thinking disabled, eg. it's not uncommon to run Qwen3.6 locally, but its thinking is extremely verbose, so it's fast if you disable
85.
▲
by
naasking
3mo ago
> Can't help to think of a recent HN post about most AI-generated projects being abandoned within months. Why? That which can be created with little effort can be dismissed with little fanfare, since it can easily be recreated later
86.
▲
by
naasking
3mo ago
> If you can do a Rust rewrite with AI, I can create one as well. Translations are a lot more likely to be error-free and robust, as the original data structures and algorithms have been battle tested.
87.
▲
by
naasking
3mo ago
> Trump called for investigation and arrest on Comey, Cheney, Powell, ... despite the fact that they never crossed the line in expressing their criticism to Trump's decisions. I'm not a fan of Trump's governance, but none
88.
▲
by
naasking
3mo ago
> When, at the same time, "normal" people have absolutely no problem criticizing politicians in a "normal" way. So you agree that the US has more free speech because people can deviate from your completely arbitrarily
89.
▲
by
naasking
3mo ago
I'd say that one data point is not enough to object to, except if what's depicted is literally impossible, like the AI generated images of a female pope. Beyond those, the distribution of a set of generated images is what implies
90.
▲
by
naasking
3mo ago
> I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. Maybe you should base this assessment on more than just vibes. Grok came out pretty bal
More ›