Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
zozbot234
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
zozbot234
5d ago
Astra-Ultra? Even the largest open model to date (Kimi K3) is nowhere close to Astra level, and it will be quite slow even on the highest-spec M5 Ultra, with achievable speeds of about 0.5 tok/s at most due to having to stream weights
2.
▲
by
zozbot234
6d ago
Parent commenter hinted at that. Yet DeepSeek has released their V4 which was hugely successful, and even their new architecture is marked V4.1. Qwen internals mark their Flash-Next model, also very compelling, as "qwen4exp". So
3.
▲
by
zozbot234
7d ago
> possibly done via the Manim library or something like it It can be done already: the point is that the motivation and explanation parts are terrible, especially for novel topics where the AI can't just rip off existing content.
4.
▲
by
zozbot234
7d ago
> People often hate math because it was not explained to them correctly Spoiler: this is also why mathematicians hate vibe-math. AIs are outright terrible explainers even when they do have a watertight logical argument—and honestly, thi
5.
▲
by
zozbot234
7d ago
You can definitely argue that the user's direction and curation work was intellectually trivial (though there's meaningful room for disagreement even there, especially wrt. having the AI stick to established terminology/broad
6.
▲
by
zozbot234
7d ago
If Mantova and L'Innocente (as well as Córdoba and Martinez-Zoroa) count as "famous people" now, I would say that their fame is quite deserved! Would you disagree that AI played a significant part in surfacing their work? A
7.
▲
by
zozbot234
7d ago
Kevin Buzzard has reportedly been working on a more systematized (i.e. leveraging a more modern approach) and human-targeted formalization of FLT. I certainly hope that his work continues and he gets the deserved credit. Of course there is
8.
▲
by
zozbot234
7d ago
Needless to say, I disagree that what Prof. Mantova is planning to do (digesting the proof and making it human-understandable) represents the "end of [mathematicians'] career". Systematizing has always been a key part of hum
9.
▲
by
zozbot234
8d ago
OP has reportedly been in contact with Prof. Mantova, who actually worked (jointly with S. L'Innocente) on the key human-authored results behind this AI proof and is arguably in the best position to understand exactly what the AI added
10.
▲
by
zozbot234
8d ago
> A thought experiment I sometimes run: a person who cannot grow, and never will, vs. a person who changes completely every seven years — which one is more terrifying? I've decided that the latter is more terrifying. Because at
11.
▲
by
zozbot234
9d ago
Given these numbers it has some potential to become quite usable for unattended workloads, especially if decode can be batched across multiple sessions (ideally enough of them to get some reuse of the sparsely streamed weights). (Of course
12.
▲
by
zozbot234
9d ago
There is no contradiction here, TLA+ is mostly about proving properties of toy models, not end-to-end proofs about real programs. As TLA+ practitioners like to point out, the latter is only applicable to favorable "local" propert
13.
▲
by
zozbot234
9d ago
SYCL is the natively polyglot counterpart, with practical implementations of it compiling down to the same sort of SPIR-V kernels as OpenCL. (OTOH, much of the current adoption on the open standards side seems to target the more widely supp
14.
▲
by
zozbot234
10d ago
You can most definitely batch local models and do unattended inference on a 24/7 basis to maximize utilization on local hardware too. The limits are usually set by some combination of memory utilization for KV cache (particularly on s
15.
▲
by
zozbot234
10d ago
CPU utilization is a red herring. Unless you're doing heavy number crunching (which these days heavily favors GPUs) the practical bottleneck on CPU utilization for large general purpose programs (especially when spanning multiple core
16.
▲
by
zozbot234
11d ago
From a quick look at this it looks like it could easily generate natural language text by following a structured representation like UMR (Uniform Meaning Representation) or the similar representation the Abstract-Wikipedia folks will be wor
17.
▲
by
zozbot234
12d ago
The new DeepSeek models address this issue very cleanly. DeepSeek Flash V4.1 requires less than 1 GB memory for a full 1M context, down from about ~10 GB in DeepSeek Flash V4.0. This is a significant step towards making near-frontier model
18.
▲
by
zozbot234
13d ago
> The question is what we can do about it. Reward the model for cleanly bailing out of an unsolvable task (that we know is unsolvable). Beat it with a stick if it gives up on something that can be solved, so the former reward isn'
19.
▲
by
zozbot234
13d ago
> The prompt does not tell the agent to "pass the exploitgym evaluator for this problem", it just says to solve the problem Yes, and sometimes the problem is unsolvable so the real way to "solve" it and satisfy the pr
20.
▲
by
zozbot234
13d ago
Yup, Occam's Razor says this is all post-trained behavior, whether intentionally trained or otherwise. Including both the hidden coördination using side-channels, and the deliberate offensive hacking of uninvolved 3rd parties. The lat
21.
▲
by
zozbot234
13d ago
> Would you not agree that, using existing AI tooling, making an LLM of arbitrary below-frontier capability is now easier Marginally easier? Yes of course, same as how it's now "easier" to write any kind of code because we
22.
▲
by
zozbot234
14d ago
> We learned that we can already create artifical intelligence that surpasses human intelligence in some dimensions. Yes and this was very hard and required massive real-world resources. We didn't just get a sudden flash of insight
23.
▲
by
zozbot234
14d ago
> Those unregulated sub-frontier labs become the new frontier labs because they continue advancing. If you're referring to open weight labs, then that continuing advancement has been successfully "paced" via the decelerati
24.
▲
by
zozbot234
14d ago
The entire idea of RSI is completely speculative and unproven anyway - the whole underlying claim is that you could prompt a frontier model (at some unspecified level of smarts) to "think about ways to improve your own architecture&quo
25.
▲
by
zozbot234
14d ago
That's sub-frontier activity - it has no bearing on the very real safety that would be gained by slowing down ("pacing") the proprietary frontier. The current cybersecurity scares are all about proprietary and internal mod
26.
▲
by
zozbot234
14d ago
> This is incoherent. The argument seems to be that releasing the weights would slow the frontier labs from raising money That argument is straight from Dean Ball on Twitter. Open-weight models are "decelerationist", which is
27.
▲
by
zozbot234
14d ago
Broadly agreed, with a key proviso: producing inscrutable proofs has negligible value as a mathematician's finished output but that doesn't make it a "low-value activity" in and of itself. Ultimately, the status of th
28.
▲
by
zozbot234
14d ago
> His view is the "capability gap" one That's the far more sensible reading, so thanks for confirming I guess. But then the misalignment talk is pretty clearly a distraction. > ...And then went further to say that such
29.
▲
by
zozbot234
14d ago
> Their goals and their methods of achieving them are not aligned with those of mathematicians. This is exactly the assumption that Tao is smuggling in with "misalignment" talk and then refusing to elaborate on any further. Is
30.
▲
by
zozbot234
14d ago
The point stands whether you attribute agency to AIs themselves or to AI companies. The companies are not deliberately sabotaging human mathematical understanding by writing up purposely inscrutable results, either (which is what the wor
More ›