Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
stalfie
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
stalfie
7d ago
I think it is legitimate to argue that the process here does not fit the usual definition of "brute forcing". Traditionally, brute forcing would refer to something like a dictionary attack, where an algorithm tries to match all po
2.
▲
by
stalfie
7d ago
I think political lobbying for legal changes might be the only realistic pathway, in that current legislation (eg. GDPR in the EU) is extremely punitive even for minor violations. I personally am trying to float using local LLMs to create a
3.
▲
by
stalfie
7d ago
Well, from a Bayesian standpoint the odds of that seems to be pretty much zero, given that you would have to decipher an arbitrary message that coincidentally pointed to a real date and time that in retrospect turns out to be the correct ti
4.
▲
by
stalfie
7d ago
From context it doesn't look like this was the interpretation of "sidestepping" OP was using.
5.
▲
by
stalfie
7d ago
From the article: > Astra felt compelled to check its work and found that, in fact, the English cruiser HMS Canterbury arrived in Sevastopol on November 24, 1918, based on its original logs Then follows a picture of the original log pa
6.
▲
by
stalfie
7d ago
Hear hear! There are so many obvious improvements to how almost everything is done. For instance, in medicine review articles as a class of articles largely represent a giant waste of time. RCTs flatten all their gathered data during publis
7.
▲
by
stalfie
3mo ago
Oh I wasn't talking about humans. I should probably have pointed that out. There's some scenarios like catastrophic crop failure and so on that might lead to that, but frankly I doubt it. I was referring more to everything else. C
8.
▲
by
stalfie
3mo ago
Sure, here are some citations: https://arxiv.org/abs/2401.03910?utm_source=chatgpt.com https://www.frontiersin.org/journals/psychology/articles/10.... https://arxiv.org/p
9.
▲
by
stalfie
3mo ago
"Climate conspiracy"? Like you, mean, the conspiracy of climate scientists to publish facts to the best of their understanding? I don't know what exact strawman you're arguing against, although I'm sure you can alwa
10.
▲
by
stalfie
3mo ago
No one knows how LLMs work. We know how the architecture works, but almost nothing about why. Saying "statistical next token prediction" tells you about as much about LLMs as saying "action potential thresholds" tells yo
11.
▲
by
stalfie
3mo ago
I was referring more to the fact that no one predicted that next token predicting Transformers would go so far. Not about "AI" in general.
12.
▲
by
stalfie
3mo ago
No one imagined LLMs in their current format, it was simply a result of discovering that scaling compute and tokens produced better and better results with the Transformer architecture. The inventors of the Transformer architecture were wor
13.
▲
by
stalfie
3mo ago
10 years is a long time. 10 years ago the Transformer architecture didn't exist. I would call it moderately unlikely at best. At the very least, I would say it's likely that development will require an entirely different skillet 1
14.
▲
by
stalfie
3mo ago
There's at least one benchmark that attempts to measure this, but it has been running for a year plus so it's quite infrequently updated now. https://fiction.live/stories/Fiction-liveBench-Mar-25-2025/o..
15.
▲
by
stalfie
3mo ago
Ironically, the article points out that the original authors publisher actually put out two DMCA notices to google last year, apparently with no effect. I guess DMCA takedowns are only for the big fish fighting the good fight against car pi
16.
▲
by
stalfie
3mo ago
Mmmm, not sure I agree with this, although this is a topic where we would have to do a lot of groundwork to formulate our positions precisely in order to ensure we're actually discussing the same thing. My counterargument is that verif
17.
▲
by
stalfie
3mo ago
Once again, Russia turns out to be the reason we can't have nice things. War truly is a waste for everyone involved. Now that Russia is also helping North Korea to launch satellites (one so far), expect everything to get worse in the f
18.
▲
by
stalfie
3mo ago
Well, I'd argue that this depends on the field you're investigating. Sometimes you have a way to identify objective reality and sometimes you don't. In mathematics the majority of the field is verifiable in this way. Coding a
19.
▲
by
stalfie
3mo ago
I guess so. Just to be clear, I was talking about post-training methods for reasoning models here, not pre-training. I think "model as a judge" should actually do okay as a "sentiment analysis" style reward for expressin
20.
▲
by
stalfie
3mo ago
One thing I wonder about hallucinations, is that it seems on the surface that it is an easy problem for RLVR to target. Since you're already generating enormous amounts of reasoning traces which are verified by correct answers, just ha
21.
▲
by
stalfie
3mo ago
The criticism is also similar to those faced by Theranos. Survivorship bias is always a factor when looking backwards.
22.
▲
by
stalfie
3mo ago
Excluding the cost of X-ray/CT/MRI machines, operating them, getting people to them and through them, sometimes injecting contrast, and sometimes dealing with side effects of said contrast, radiologists, I think. You can scale all
23.
▲
by
stalfie
3mo ago
It's worse then that unfortunately. Even when invasive tests are positive, and we think we caught a cancer early, we know from population statistics that the reality is that often nothing would have happened. So we don't even trul
24.
▲
by
stalfie
4mo ago
Well, the Fable guardrails breaks this argument, as when you get booted down to Opus 4.8 it still happily responds (as does most other models after a "I'm not a doctor but..." hedge). So you get to press the big red button an
25.
▲
A Plea to the Labs: Let the Models Diagnose
(tangent.bearblog.dev)
2 points
by
stalfie
4mo ago
|
3 comments
26.
▲
by
stalfie
4mo ago
I got so frustrated with Fable refusing to make any medical diagnoses, which is the most recent iteration of a longer trend that has bothered me for years, that I made a blog to express how annoyed I am. Posting it here in case anyone is in
27.
▲
by
stalfie
4mo ago
Honestly, I have yet to see any evidence of data leak from private sources. I think one of the better example is "simple-bench", which at least used to be a low-key benchmark that I would assume would have been saturated quickly i
28.
▲
by
stalfie
4mo ago
Update in case anyone reads this comment ever again. I have found that I trigger the guardrails any time I ask for medical Q&A as a doctor, be it ECGs, case reports, and so on. But if I phrase it like I'm the patient ("help me
29.
▲
by
stalfie
4mo ago
Tried to benchmark ECG interpretation capabilities, and I hit the guardrails no matter what I do. Incredibly frustrating that medical performance seems to be a victim of "biological risk" guardrails.
30.
▲
by
stalfie
4mo ago
This article describes how Transformers work, but not really how LLMs work. Explaining the underlying architecture gives you about as much insight into how a modern LLM behaves as an breakdown of neuronal biochemistry and a few pathways doe
More ›