Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
COAGULOPATH
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
COAGULOPATH
2mo ago
This AI generated post (100% on Pangram) is pretty out of date. >On SimpleQA, a benchmark of factual recall with no tools allowed, the current leader is Gemini 2.5 Pro at 53%, so the best recall money can buy still misses half the questi
2.
▲
by
COAGULOPATH
2mo ago
>I'd like to know a lot more about how that works. My guess is that it works like Gemini's SynthID: by altering the logprobs of the next token. Like, for every 10th token, instead of outputting the most probable, it outputs the
3.
▲
by
COAGULOPATH
4mo ago
I used to make those. There was a lot of creative stuff people discovered. For example, the game's terrain was baked at launch, so you couldn't turn land into water. But someone noticed that bridges seemed to spawn water textures
4.
▲
by
COAGULOPATH
5mo ago
Please do not post AI generated comments.
5.
▲
by
COAGULOPATH
6mo ago
In the system card they seem to dismiss this. Quotes; > (...) Claude Mythos Preview’s gains (relative to previous models) are above the previous trend we’ve observed, but we have determined that these gains are specifically attributable
6.
▲
by
COAGULOPATH
7mo ago
Yes, I find LLM-written posts valueless because I can already talk to a LLM any time I want (and get the same info). It's not these commenters are the Queen of Sheba bearing a priceless gift of LLM slop. That stuff's pretty cheap.
7.
▲
by
COAGULOPATH
8mo ago
>I've found Moltbook has become so flooded with value-less spam over the past 48 hours that it's not worth even trying to engage there, everything gets flooded out. When I filtered for "new", about 75% of the posts ar
8.
▲
by
COAGULOPATH
8mo ago
And even if you could, how can you tell whether an agent has been prompted by a human into behaving in a certain way?
9.
▲
by
COAGULOPATH
8mo ago
Is it a success? What would that mean, for a social media site that isn't meant for humans? The site has 1.5 million agents but only 17,000 human "owners" (per Wiz's analysis of the leak). It's going viral because a
10.
▲
by
COAGULOPATH
8mo ago
If the site is exposing the PII of users, then that's potentially a serious legal issue. I don't think he can dismiss it by calling it a joke (if he is). OT: I wonder if "vibe coding" is taking programming into a culture
11.
▲
by
COAGULOPATH
9mo ago
Thanks, I didn't realize the situation was so dire.
12.
▲
by
COAGULOPATH
10mo ago
> In 1920, there were 25 million horses in the United States, 25 million horses totally ambivalent to two hundred years of progress in mechanical engines. But would you rather be a horse in 1920 or 2020? Wouldn't you rather have mod
13.
▲
by
COAGULOPATH
10mo ago
And they do hacky things like space elements vertically using <br> tags.
14.
▲
by
COAGULOPATH
2y ago
Something I'm increasingly noticing about LLM-generated content is that...nobody wants it. (I mean "nobody" in the sense of "nobody likes Nickelback". ie, not literally nobody.) If I want to talk to an AI, I can t
15.
▲
by
COAGULOPATH
2y ago
In some domains (math and code), progress is still very fast. In others it has slowed or arguably stopped. We see little progress in "soft" skills like creative writing. EQBench is a benchmark that tests LLM ability to write stori
16.
▲
by
COAGULOPATH
2y ago
>Being monetarily successful does not mean you’re good or shouldn’t be criticised. Is anyone saying that Mr Beast is good and shouldn't be criticised? I can't see them.
17.
▲
by
COAGULOPATH
2y ago
I think this works, not because LLMs have a "hallucination" dial they can turn down, but because it serves as a cue for the model to be extra-careful with its output. Sort of like how offering to pay the LLM $5 improves its output
18.
▲
by
COAGULOPATH
2y ago
Came here hoping to find this. You will not unlock "o1-like" reasoning by making a model think step by step. This is an old trick that people were using on GPT3 in 2020. If it were that simple, it wouldn't have taken OpenAI s
19.
▲
by
COAGULOPATH
2y ago
That's definitely weird, and I wonder how legal it is.
20.
▲
by
COAGULOPATH
2y ago
It's just Ilya typing really fast.
21.
▲
by
COAGULOPATH
2y ago
>but much worse (and worse even in comparison to GPT4) than English composition O1 is supposed to be a reasoning model, so I don't think judging it by its English composition abilities is quite fair. When they release a true next-ge
22.
▲
by
COAGULOPATH
2y ago
I've heard rumors that GPT4's training data included "a custom dataset of college textbooks", curated by hand. Nothing beyond that. https://www.reddit.com/r/mlscaling/comments/14wcy7m/
23.
▲
by
COAGULOPATH
2y ago
Yes, this only helps multi-step reasoning. The model still has problems with general knowledge and deep facts. There's no way you can "reason" a correct answer to "list the tracklisting of some obscure 1991 demo by a ban
24.
▲
by
COAGULOPATH
2y ago
My experience is the opposite: laypeople are excessively pessimistic on LLM progress ("AI is so dumb. It tells you to put glue on pizza and eat rocks)", usually due to a remembered anecdote that's either years old or reflects
25.
▲
by
COAGULOPATH
2y ago
What he means is that if you search for "Diminished by its artsiness" + "Pauline Kael" you won't find any results (except for ones related to this news story). Google is polluted with AI generated content but not t
26.
▲
by
COAGULOPATH
2y ago
Sanitarium is one of those Bad Mojo-esque games that's worth playing for how unique it is. The isometric viewpoint never really worked for me, and undercuts the immediacy of the horror. Things aren't happening to you , but to a l
27.
▲
by
COAGULOPATH
2y ago
Most guides on detecting AI images are from 2022 and have aged like dinosaur milk. “AI can’t draw hands.” “AI can’t draw straight lines.” “AI can’t spell words.” We now live in an age of photorealistic fake media. It is no longer true that
28.
▲
by
COAGULOPATH
3y ago
Gemini Ultra gets this right. (Usually it's worse at GPT4 at these sorts of questions.)
29.
▲
by
COAGULOPATH
3y ago
Yeah, it's funny. I used to think "Demis Hassabis...where have I heard that name before?" And then I realized I saw him in the manuals for old Bullfrog games.
30.
▲
Gemini Ultra is out. Does it beat GPT4? (~10k words of tests/observations)
(coagulopath.com)
1 points
by
COAGULOPATH
3y ago
|
1 comments
More ›