Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
gwern
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
91.
▲
by
gwern
6mo ago
The description by OP in comments like https://old.reddit.com/r/LegalAdviceUK/comments/1s92fql/my_s... seems to strongly imply that all the accounts were unconnected in a GSuite sense, and they are being
92.
▲
by
gwern
6mo ago
You'll love GreaterWrong, then: https://www.greaterwrong.com/posts/BJ4pnropWdnzzgeJc/i-am-de...
93.
▲
by
gwern
7mo ago
FWIW, we did consider a histogram heuristic, and I believe GreaterWrong still uses one rather than InvertOrNot.com. But I regularly saw images on GW where the heuristic got it wrong but ION got it right, so the accuracy gap was meaningful;
94.
▲
by
gwern
7mo ago
> It's absolutely true that there's a subset of raster images, like diagrams with white backgrounds and black lines, that would benefit from inversion. I could be wrong, but in my experience they're a minority, and the cos
95.
▲
by
gwern
7mo ago
> In practice they're rare enough that the per-page toggle handles them, but it's the honest limitation of the approach. I don't understand how you handle raster images. You simply cannot invert them blindly. So it sounds
96.
▲
by
gwern
7mo ago
Have you considered, since you can extract the images via the mask, selectively inverting them? One can fairly reliably use a small NN to classify images by whether they should be inverted or just dimmed, and I've used it with great su
97.
▲
by
gwern
7mo ago
A good wiki like MediaWiki supports various levels of visibility. For example, you could define a namespace for each group of readers like 'Family:'. Or use transclusions from subpages. (This might sound like a bit of a hassle but
98.
▲
by
gwern
7mo ago
> Detractors of AI are often accused of moving the goalpost, but I think your comment is guilty of the same. Before Claude Code, we had Cursor, Github Copilot, and more. Each of these war purportedly revolutionizing software engineering.
99.
▲
by
gwern
7mo ago
I was being sarcastic, because that point obviously also applies to the subjects in the experiment as well.
100.
▲
by
gwern
7mo ago
One might worry that it would increase the authors' confidence even following their LLM rewrite errors and reduce accuracy overall regardless of moderators.
101.
▲
by
gwern
7mo ago
End-to-end optimization in action! Although I'd've liked more than 1 example (pathfinding) here.
102.
▲
by
gwern
7mo ago
Or more precisely, isn't this reinventing notebooks (not the first JS-centric notebook either)?
103.
▲
by
gwern
7mo ago
Yes, it's greedy so may hit local optima. You can fit learning curves and try to extrapolate out to avoid that problem, to let you run long enough to be reasonably sure of a dead end, and periodically revive past candidates to run long
104.
▲
by
gwern
7mo ago
> The agent can theoretically come up with a protocol to run those same 12 experiments one-by-one and only then decide which branch to explore next - which I think would lead to the same outcome? At least in theory, adaptiveness should s
105.
▲
by
gwern
7mo ago
Except the power drill isn't being used to make a better chimpanzee.
106.
▲
by
gwern
7mo ago
A little disappointing. All about the history of bell curves, but I don't think it does a very good job explaining why the bell curve appears or the CLT is as it is.
107.
▲
by
gwern
7mo ago
After some chatting with GPT-5.4 Pro, I think the 10x claim might simply be because in the >90% load region, you get the famous unbounded delays. And with realistic loss functions like absolute or squared error (consider that in a queue
108.
▲
by
gwern
7mo ago
The 10x claim reminds me of various scaling laws in anthropology/sociology, especially the much-debated Haire cube-root scaling law (eg https://gwern.net/doc/sociology/1959-haire.pdf https://gwern.
109.
▲
by
gwern
7mo ago
I agree. That's why you should write as much as you can now, if you want to get it into the LLMs ( https://gwern.net/blog/2024/writing-online ). You never know when the window will slam shut and LLM training go
110.
▲
by
gwern
7mo ago
https://en.wikipedia.org/wiki/The_Spirit_Level_(Wilkinson_an...
111.
▲
by
gwern
7mo ago
Linux isn't a strong brand at all. Even nerds would struggle to tell you what a 'Linux' is (or as I prefer to call it, since that's every non-smartphone Linux I've ever used, 'GNU/Linux'). And I'
112.
▲
by
gwern
7mo ago
Some of the typos and poor typography might be accidental simply because that is so common, but the rest of it is surely deliberate. Remember that _American Psycho_ is repeatedly hinting to you throughout the movie that most, or even all, o
113.
▲
by
gwern
7mo ago
It's definitely very valuable, but for what AI model? How does any of that lead to AGI, or even just a good coding agent?
114.
▲
by
gwern
7mo ago
https://news.ycombinator.com/item?id=47134407 Thanks for writing it up!
115.
▲
by
gwern
7mo ago
List of examples: https://gwern.net/turing-complete It was probably unintentional, yeah, I don't recall any mentions of early printf being overloaded to do stuff, nor is it clear why you would do that since you're
116.
▲
by
gwern
7mo ago
The solution here seems to be to impose some constraint or requirement which means that literal copying is impossible (remember, copyright governs copies , it doesn't govern ideas or algorithms - that would be 'patents',
117.
▲
by
gwern
7mo ago
I would have to consider carefully if I thought I was a high-enough quality candidate that it would be interpreted as a countersignal rather than a signal. If I, gwern, specifically, were to apply, I might; because I know I am widely read
118.
▲
by
gwern
7mo ago
Also just differing levels of relevance. You don't talk with a businessman or investor or famous people in general because of their writing; if you made a list of relevant skills, 'proper spelling when quickly texting from a phone
119.
▲
by
gwern
7mo ago
My understanding was that Autoresearch was defined as training from scratch (since it's based on the nanogpt speedrun), not using any pretrained models. So it couldn't do anything like upcycling a pretrained model or the Franken
120.
▲
by
gwern
7mo ago
Adding, swapping, or duplicating layers has a long history (eg. StyleGAN, upcycling), and it was pointed out at least as far back as He et al 2015 (Resnets) that you could ablate or add more layers because they functioned more as just doing
More ›