Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
porridgeraisin
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
31.
▲
Deep dive on evals for agents [video]
(youtube.com)
2 points
by
porridgeraisin
14d ago
|
0 comments
32.
▲
by
porridgeraisin
14d ago
A lot of streaming platform listen minutes has always been this type of filler music. Until now, you could not really disambiguate intentional listening and filler listening. Now you can and the filler listeners wont even notice[1]. A huge
33.
▲
by
porridgeraisin
15d ago
I think you're focusing only on the general coding agent aspect of LLMs. > If openrouter is only for people who "make their own benchmark" in a mission to roll something that's usually made by scientists with large bu
34.
▲
by
porridgeraisin
15d ago
You'd be surprised. Not enough teams still have their own evals. India[1] and the US[2] is my experience. To get many of them to understand the benefit of putting a couple of people to do data labelling and write a couple of verifiers
35.
▲
The Elephant in the Context Window
(twitter.com)
2 points
by
porridgeraisin
15d ago
|
0 comments
36.
▲
by
porridgeraisin
15d ago
There is also the aspect of model generations being robust under a particular ctx mgmt strategy/tailored harness. Only since this year are most models robust in this way. In many open models, the problem is that successive generations
37.
▲
by
porridgeraisin
16d ago
There are thousands of such techniques across different parts of the system. In ML, there are way too many ideas, and lots of people knowingly and unknowingly restate the same ideas. It's a new field, so even common language is not the
38.
▲
by
porridgeraisin
16d ago
Yep. a multiplicative LSTM to be exact.
39.
▲
by
porridgeraisin
16d ago
This is adapted from Microsoft research's YOCO. It was known for a while(2024!). Yes, credit to Deepseek for actually scaling it up and releasing a frontier flash LLM. Edit: the rest of this thread has become a US China infowar theory
40.
▲
by
porridgeraisin
16d ago
In india we use UPI. But people still use cards enough to be in decent enough terms with them for when you want to make a risky purchase and chargeback. In a sense, it's simply insurance. You pay extra 2% everywhere so you can dispute
41.
▲
by
porridgeraisin
16d ago
This. My extended familys shop really benefitted from moving away from cash. Now the only source of cash is people who are paying from their black money stashes.
42.
▲
by
porridgeraisin
17d ago
I phrased it badly just out of bed. I meant what you're saying. That it's OK if it's slop if it serves a business function. Edited
43.
▲
A Biography of Lee Holloway, the Architect of Cloudflare's Technology (Part 1)
(note.com)
103 points
by
porridgeraisin
17d ago
|
22 comments
44.
▲
by
porridgeraisin
17d ago
Because the problem has almost no value unto itself. The clay statement of navier stokes is not relevant to how CFD is done in practice. It's about what is non verifiable versus verifiable. The same way it produces "slop" cod
45.
▲
by
porridgeraisin
17d ago
It was behind the heavy 100$/mo+ plans. Clearly meant for enterprise and not consumer.
46.
▲
by
porridgeraisin
17d ago
Yep. But I suppose these are early forms of the product and who distributes the next, high PMF product the best is what matters. So I wouldn't write off others.
47.
▲
by
porridgeraisin
18d ago
Instinct exists with the same features. It is quite good. I was thinking they were gonna get acquired by meta 100%, but looks like meta has a competitor.
48.
▲
by
porridgeraisin
18d ago
I agree on that, my post was not meant to be opposing this.
49.
▲
by
porridgeraisin
18d ago
The for-case for this type of method is that this is economies of scale for mathematics. We are basically mass manufacturing math. Just like you have just 100 designers for a product selling millions of units, you will now need 100 mathemat
50.
▲
Agentic Video Understanding with Gemini
(twitter.com)
2 points
by
porridgeraisin
18d ago
|
0 comments
51.
▲
by
porridgeraisin
18d ago
It is not pseudoscience. Have you read about how it works?
52.
▲
Principia: Relational Physics Tests for Video Models
(principiabench.github.io)
1 points
by
porridgeraisin
18d ago
|
0 comments
53.
▲
Kenyans Made a Living Writing College Essays. Then A.I. Arrived
(nytimes.com)
3 points
by
porridgeraisin
19d ago
|
0 comments
54.
▲
by
porridgeraisin
19d ago
There is a reason imperative programming is so common. It is more amenable to poorly designed, under specified, iterative development. Most real world software is in that category naturally. If you're meticulously designing and enginee
55.
▲
by
porridgeraisin
19d ago
> it must perform it's normal autoregressive decoding to know what is the correct token in order to have something to compare with Correct except for the word "autoregressive". When you have to verify a sequence of tokens
56.
▲
by
porridgeraisin
19d ago
TBH, your username gives that part away atleast. My mind went to slavic when I saw the "ov" - it might be factually wrong, but thats waht happened.
57.
▲
Microsoft Tgrep: Trigram-indexed grep
(github.com)
4 points
by
porridgeraisin
19d ago
|
0 comments
58.
▲
Tensor Product Representations
(twitter.com)
1 points
by
porridgeraisin
20d ago
|
0 comments
59.
▲
LLM representations have implicit symbolic structure
(twitter.com)
10 points
by
porridgeraisin
20d ago
|
0 comments
60.
▲
by
porridgeraisin
21d ago
The noise is the main issue. I'm not american, but one of my relatives works in bernie sanders's office, and one of my friends (30ish) is involved somehow in the democratic party not sure how exactly. And what I hear from them is
More ›