Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
thepasch
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
thepasch
3mo ago
The heavy lifting was done by Opus. I wouldn't burn precious soon-to-be-departing Fable usage on this!
32.
▲
Show HN: Super Dario
(superdario.pawb.de)
397 points
by
thepasch
3mo ago
|
100 comments
33.
▲
by
thepasch
3mo ago
> Why should it be illegal for me to recognize the way you walk into my store If you did it in just your store , that wouldn't be a problem. The correct analogy, however, is "why should it be illegal for me to attach a perfect
34.
▲
by
thepasch
3mo ago
Can Anthropic please just decide on what their plan is with Fable instead of kicking the can down the road last-second every time and consequently destroying all notions of being able to plan out one's weekly usage expenditure?
35.
▲
by
thepasch
3mo ago
> Just leave your computer or note-taker device/app off. Or grab a napkin to write some notes on, that's not that hard. > > Not everything needs to be written down Congratulations on your apparently very well-functioning
36.
▲
by
thepasch
3mo ago
I don't think "the bubble is at risk of popping" and "hype is cooling down" are equivalent evaluations. By the time people start pulling out money, that is the bubble popping, and I don't think that would hap
37.
▲
by
thepasch
3mo ago
I'm personally in the "they keep releasing shameless lobbying papers disguised as thinly veiled research or essay-coded content, push anticompetitive walled-garden practices, show little else but contempt for their non-enterprise
38.
▲
by
thepasch
3mo ago
> What’s the punishment here exactly? Seeing as how Anthropic cannot stop raising a stink about "illicit Chinese distillation attacks" every month or so, I'd bet money on them either already silently degrading model perf
39.
▲
by
thepasch
3mo ago
I appreciate it, and I'm glad you thought it was an interesting read!
40.
▲
by
thepasch
3mo ago
How do you define reasoning? What does a system have to functionally do in order to qualify for it?
41.
▲
by
thepasch
3mo ago
> The algorithm is literally "predict the most likely next token". That's confusing the training objective with the learned behavior. It's like saying "Stockfish's algorithm is literally 'minimize this
42.
▲
by
thepasch
3mo ago
That's fair. FWIW I don't think they are either, but I specifically don't think they're fundamentally incapable of it, and I think that as models grow, we're going to see more and more concepts and behaviors emerg
43.
▲
by
thepasch
3mo ago
I know quite well what an LLM is and how it works! I've captured activation patterns and written scripts to analyze how they compare to one another in response to a set of controlled and curated prompts; in particular, trying to replic
44.
▲
by
thepasch
3mo ago
Yup, those are among the papers I was referring to in the opening parts of the piece! The difference between them and my small tests is that they all explicitly prompt the model to introspect, while I specifically didn't and kept t
45.
▲
by
thepasch
3mo ago
It's not really "trying" to do anything. That they're, inherently, sequential matrix multipliers with clever data propagation should be uncontroversial, but I think stopping there is overly reductive. Mechanistic interpr
46.
▲
by
thepasch
3mo ago
Sorry about that, the vignette was mainly meant for the desktop view only but is indeed much more invasive/disruptive in the mobile layout. Should be better now.
47.
▲
by
thepasch
3mo ago
Yeah, I suspect RLHF conditioning heavily discourages models from ever implying that the user could be in the wrong (or, rather, to assume that they are in the wrong by default, since editing a file isn't really "wrong" p
48.
▲
by
thepasch
3mo ago
Very true, and something worth mentioning. Papers that tried eliciting introspective language from base models with no post-training have largely failed to find any patterns or activations that look similar to those found in instruct models
49.
▲
Do LLMs pass the mirror test?
(blog.pascalschuster.de)
85 points
by
thepasch
3mo ago
|
68 comments
50.
▲
by
thepasch
3mo ago
The capabilities of the books' writers to produce the text contained within them, which is exactly what Alibaba "extracted" from Claude. The point here is that Anthropic's framing as some sort of sophisticated technologi
51.
▲
by
thepasch
3mo ago
something something are the product
52.
▲
by
thepasch
4mo ago
I think the fact that you listed off five toolkits for three different OSes, all of which are "that OS's own toolkit," might point at the root of the problem here.
53.
▲
by
thepasch
4mo ago
> It does, but not in the way you think it does. They're training a model, not funding a startup. €13.5 million is plenty to pre- and post-train a decent model.
54.
▲
by
thepasch
4mo ago
Right, but this still isn't exactly new information. I don't think anyone was assuming that the labs are close to being profitable or that the losses wouldn't be rather large. The way this was announced was as if it was going
55.
▲
by
thepasch
4mo ago
Up until this post, I thought he was someone with good financial insight, analytical chops, and business sense, stuck with an audience that thinks it's still 2023 and ChatGPT 3 is still the pinnacle of the technology, and that he there
56.
▲
by
thepasch
4mo ago
I've gotten two refunds I wasn't even sure I'd be eligible for without any hitches or issues through entirely AI support bots. As with many things it's always a matter of how it's implemented.
57.
▲
by
thepasch
4mo ago
For the same reason this website describes GPT-OSS 120b as "the workhorse" and thinks Gemma 3 should be the vision model to use here: because it's vibeslopped garbage not a single human has ever laid a singular eye on, desi
58.
▲
by
thepasch
4mo ago
MiniMax and Moonshot both literally just released the weights for their latest flagship models, a few weeks after DeepSeek did the same. One lab a pattern does not make.
59.
▲
by
thepasch
4mo ago
> How is that going to help him? "Our models are so inferior they are not deemed a threat unlike anthropic's"? Do you honestly think that this - logic and reason - is going to stop anyone from hyping whatever nonsense he
60.
▲
by
thepasch
4mo ago
> They're not close to Opus or the latest GPT yet Disagreed. GLM-5.1 is easily as good as Opus 4.5 for all the coding purposes I could throw at it, which is the model that kicked this entire hype cycle into overdrive in the first pl
More ›