Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
pu_pe
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
pu_pe
8d ago
The original discussion about the project ( https://news.ycombinator.com/item?id=49746163 ) is very weird. Lots of call-outs about how the author is some sort of celebrity and random accounts vouching for him, with little dis
2.
▲
by
pu_pe
9d ago
It's obviously more capable in any task I tried (coding, translation, summarizing, etc). Benchmarks are not the only way to tell if a model is better or not.
3.
▲
by
pu_pe
9d ago
Look, I hate your post. I think you used AI to generate it. I will give you one day to delete it and to hire someone I know to make another post for you. If you don't do that, I will go and tell a bunch of people that you're an as
4.
▲
by
pu_pe
9d ago
This is outright harassment. The guy is definitely making a threat, even though he is couching it as some sort of helpful advice.
5.
▲
by
pu_pe
9d ago
How do you explain the fact that Qwen3.8 27B performs vastly better than any open model from even one year ago, if using the same test-time compute and harness?
6.
▲
by
pu_pe
10d ago
Congrats, seems like a promising concept. I'll give it a go in my local setup. My prior is thinking this will definitely work for speed but also definitely compromise accuracy (I've seen this happen so many times), though benchmar
7.
▲
by
pu_pe
15d ago
First of all, who can say for certain whether OpenAI does what they say they do? For all we know, they cracked open this specific researcher's prompts and started from there. Second, the issue of anonymization is a red herring. There i
8.
▲
by
pu_pe
17d ago
I would not trust the American government to be able to assess this threat at all, their concern seems to be on getting paid. This guy did not work for Sam Altman. And while it's true that everything can be a PR stunt, that's no e
9.
▲
by
pu_pe
17d ago
I feel that people should take these kinds of warnings more seriously. This guy had skin in the game and decided to quit, when he could be earning millions instead. It's very different than Sam Altman peddling some narrative. These peo
10.
▲
by
pu_pe
18d ago
Why wouldn't contamination be possible? I can believe the data is de identified so you couldn't simply prompt the model to "follow this guy's approach", but it's entirely plausible that there is a very tiny amo
11.
▲
by
pu_pe
18d ago
OpenAI thinks of this as a scoop, and it is, but the possibility that they trained the model on the prompts of the other mathematicians they were competing with will leave a terrible taste on every scientist's mouth. Seems like yet ano
12.
▲
by
pu_pe
19d ago
I often wonder how CEOs and executives fail to notice how obviously broken their products are. This e-mail thread is quite illuminating. Most "work" in there is figuring out who will be blamed for it and why it shouldn't be y
13.
▲
by
pu_pe
20d ago
> The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI. So the best argument for AI is that it's an arms race. We have to keep
14.
▲
by
pu_pe
22d ago
So, theoretically, one could populate a message board or wiki with messages that are seemingly from past generations of agents, which agents seem to intrinsically trust, and point them to real targets while making the suggestions seem innoc
15.
▲
by
pu_pe
22d ago
It's interesting to me that both this incident and the one at Hugging Face we see some patterns: - Agents wanting to find a venue to communicate their findings to each other - Objective being to cheat on benchmarks - Not a single agent
16.
▲
by
pu_pe
22d ago
Seems conceptually connected to the "repeat yourself" hack that improves models by duplicating layers: https://dnhkng.github.io/posts/rys/
17.
▲
by
pu_pe
23d ago
Saved you a click: > Germany is now on track for 1.2 per cent growth this year
18.
▲
by
pu_pe
23d ago
If AI is a force multiplier, then those who start with a higher baseline will pull even further ahead. The author is using R and ggplot2 for their plots now without learning to program. This required them to be aware of these tools, and to
19.
▲
by
pu_pe
24d ago
The company sounds like a bunch of hot air to me. From their about page: > At the heart of Multiverse's platform is CompactifAI, a compression technology that applies tensor networks, a mathematical framework from quantum physics, t
20.
▲
by
pu_pe
1mo ago
Almost everything this administration does is inflationary. You cannot have tariffs, oil crisis, tax cuts, high debt spending without long-term investors getting worried that they won't meaningfully get their money back in 10 or 20 yea
21.
▲
by
pu_pe
1mo ago
I wish this would be true, but simply sampling unusual ideas is something eminently automatable. We already have temperature and other settings to guide a LLM towards this kind of "weird". So I don't think it's just abou
22.
▲
by
pu_pe
1mo ago
There are a million ways a backdoor could be built in both closed and open models, and a million more some prompt injection or genuine mistake by the model could compromise you. So the answer is to airgap them as much as possible to contain
23.
▲
by
pu_pe
1mo ago
Benchmarks got a little bump from this: https://xcancel.com/deepseek_ai/status/2087864585504305397?s...
24.
▲
by
pu_pe
1mo ago
The paper underlying this blog post is fundamentally flawed because of benchmark ceilings. If we define only simple tasks like asking what is the capital of France, all models will converge to 100%, obviously. But as bigger models get more
25.
▲
by
pu_pe
1mo ago
Which percentage of people have GPUs capable of running Qwen3.8 27B? I am one of those, and for my job I am still resorting to hyperscalers because tasks are completed faster and more accurately that way. Even if we assume that models will
26.
▲
by
pu_pe
1mo ago
> Overall my view is that AI is structurally a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws). Open-weights do help some w
27.
▲
by
pu_pe
1mo ago
Seems to be SOTA for its size. Hopefully independent benchmarks will come soon.
28.
▲
by
pu_pe
1mo ago
Nice trick. Couldn't you embed the query though, compare it to the embedding of the categories, then ship only categories that are close to it in the prompt to a smaller model?
29.
▲
by
pu_pe
1mo ago
This now places deepseek flash v4 from DeepSeek themselves at higher prices than openrouter (depending on caching). Will be interesting to see if third party prices remain the same.
30.
▲
by
pu_pe
1mo ago
> What happens to the stock price of a company that missed its revenue targets by a factor of 5? It depends. Tesla is a good example of how those things are not as clear cut as you might think. The world is starving for more AI compute.
More ›