Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dwohnitmok
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
dwohnitmok
2mo ago
[flagged]
32.
▲
by
dwohnitmok
2mo ago
As a sibling comment points out you don't strictly need positional embeddings for decoder-only causal transformers. You definitely need it for non-causal ones (e.g. the encoder of the original transformer paper!). And yes accumulation
33.
▲
by
dwohnitmok
2mo ago
"All translations in this collection are © Scott Weisman. All rights reserved, except as granted by the license below." Does copyright actually belong to Scott Weisman if all the words are outputs of LLMs? Is there any relevant ca
34.
▲
by
dwohnitmok
2mo ago
Vinge addresses this in the linked article. > Stan Ulam [28] paraphrased John von Neumann as saying: >> One conversation centered on the ever accelerating progress of technology and changes in the mode of human life, which gives th
35.
▲
by
dwohnitmok
2mo ago
Vinge also coined the term "the Singularity" ( https://accelerating.org/articles/comingtechsingularity ) what he describes as "an opaque wall across the future" once superhuman artificial intelligence
36.
▲
by
dwohnitmok
3mo ago
People do use textbooks like that all the time in the experimental setup tested (essentially an open book quiz). I agree there are important differences in how textbooks and LLMs are used in real life. This study didn't explore that at
37.
▲
by
dwohnitmok
3mo ago
This study is pretty bad. The comment ( https://news.ycombinator.com/item?id=48970182 ) on the other link with the direct PDF explains the problem well, which is that nothing here being tested is specific to AI systems. This
38.
▲
by
dwohnitmok
3mo ago
Roughly speaking it is the difference between having a contractor go out and do some work and having that same contractor first come up with a plan to do some work, run that by you, and then go out to do that work. Part of it is as a anothe
39.
▲
by
dwohnitmok
3mo ago
> Moreover, it seems the prompt included the technique used to solve the problem: I don't believe this is true. The author sent techniques he used, but I don't believe any of those were ultimately what GPT-5.6 used. GPT-5.6 als
40.
▲
by
dwohnitmok
3mo ago
The author also used GPT-5.6 to write the prompt. This did involve giving GPT-5.6 access to his previous work and a back and forth process (so definitely still used the author's expertise to some degree), but the prompt itself is also
41.
▲
by
dwohnitmok
3mo ago
> This is not a remark about AI, but there's something funny about mathematics in that every novel result is broadly perceived as a big deal. This isn't true using the level of originality you're implying with your softwar
42.
▲
by
dwohnitmok
3mo ago
Do you know what kind of compression ratios you get out of curiosity? Presumably it would be lower than this format (because this format is meant to be lossy not lossless) but very curious as a baseline.
43.
▲
by
dwohnitmok
3mo ago
Interesting comments by @gwern (and why this is interesting to me beyond just the stories themselves) > The most striking result of the contest for me is what I am calling “AI allegory steganography”: a large fraction of the stories turn
44.
▲
by
dwohnitmok
4mo ago
To be fair the article never says that the family who donated the land was the one who was suing.
45.
▲
by
dwohnitmok
4mo ago
Since this seems to be a misapprehension by a couple of commentators I'll put this as a top-level comment. The family bringing the lawsuit is not the family that donated the land.
46.
▲
by
dwohnitmok
4mo ago
My guess is standing. The family bringing the suit is not the family that donated the land.
47.
▲
by
dwohnitmok
4mo ago
Nice!
48.
▲
by
dwohnitmok
4mo ago
The current HN submission title ("AGI timelines shift with whichever lab is dominant") is very bad. It is neither the title of the article nor is it the thrust of the content. The title of the article is "How long until AI au
49.
▲
by
dwohnitmok
5mo ago
> I'm guessing (wildly) this was around 0.5M USD in compute time. That seems like an especially wild guess. If you take e.g. Opus 4.7 prices, and make the assumption that you are consuming roughly $30 for every million tokens of out
50.
▲
by
dwohnitmok
6mo ago
> apparently well evidenced view that Lu Xun's overwhelming coverage in popular media and secondary schooling neglects to point out his anti-character stance What do you mean by "apparently well evidenced view?" No I'
51.
▲
by
dwohnitmok
6mo ago
Good to know!
52.
▲
by
dwohnitmok
6mo ago
We talked about this years ago. This is very much taught in the PRC (and I believe Taiwan for that matter). I specifically gave you examples of standardized tests that go over this material. https://news.ycombinator.com/item
53.
▲
by
dwohnitmok
6mo ago
There are many ways for a project to no longer be worth the company's attention. E.g. it might be the case that total costs factoring in on-going engineering energy and money (which is quite different than just compute costs!) are too
54.
▲
by
dwohnitmok
6mo ago
This seems to have a healthy helping of AI editing help (if not fully generated by AI). The links don't quite go to the sources that they should and there's a lot of AI-isms. Anyways, the calculation for the costs seem crazy high
55.
▲
by
dwohnitmok
6mo ago
We are kind of talking past each other. I'm saying something simpler. This all goes back to the original point I made in reference to your reply to johnfn: >> The post is factoring in training costs, not just inference. It is not
56.
▲
by
dwohnitmok
7mo ago
Again, that is a statement about inference time costs, not training costs.
57.
▲
by
dwohnitmok
7mo ago
No it's not. Otherwise this part doesn't make sense > in fact, they actually compound the problem by encouraging significantly more usage because if eliminating training costs makes running the model above cost, the problem is
58.
▲
by
dwohnitmok
7mo ago
When's the last time you jailbroke a model? Modern frontier models (apart from Gemini which is unusually bad at this) are significantly harder to override their system prompt than this. Again, let's say the system prompt is "
59.
▲
by
dwohnitmok
7mo ago
@krackers gives you a response that points out this already happens (and doesn't fully work for LLMs). > The hypothetical approach I've heard of is to have two context windows, one trusted and one untrusted (usually phrased as
60.
▲
by
dwohnitmok
7mo ago
You are only looking at supply. Neither supply nor demand by themselves adequately describe prices (even in supply-demand 101 theory; in practice of course it gets significantly more complicated than just supply and demand). There are field
More ›