Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
andy_xor_andrew
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
andy_xor_andrew
2mo ago
The article mentions you can always simply use a smaller, local, un-watermarked LLM to rephrase the original watermarked text. Which is true, sure. But if we're talking about deterministically taking some watermarked LLM output and hav
2.
▲
by
andy_xor_andrew
3mo ago
Did you read the post? It's not even that long. He explicitly mentions this...
3.
▲
Make invalid states unrepresentable (for your agents)
(debugti.me)
2 points
by
andy_xor_andrew
3mo ago
|
0 comments
4.
▲
by
andy_xor_andrew
3mo ago
This is a weird thing to call out, when there's so much else to talk about (price, specs, etc) buuuuuut- Check out the gameplay video partway down the page, where the two people are on the couch playing Cuphead. Right under "Your
5.
▲
by
andy_xor_andrew
4mo ago
I've used Hermes Agent in a container, and it worked ok. Little rought around the edges, but that's to be expected. But what I didn't understand... what benefit does it actually bring? On a default loadout, even after disabli
6.
▲
by
andy_xor_andrew
4mo ago
The "address lookup" strategy is really interesting, especially how it uses actual DNS: https://docs.iroh.computer/concepts/address-lookup https://github.com/Nuhvi/pkarr/
7.
▲
by
andy_xor_andrew
5mo ago
the build they use is from February, over two months old: https://github.com/ggml-org/llama.cpp/releases/tag/b8121 Which might not sound like much, but 2months in llm time is a long time, especially rega
8.
▲
by
andy_xor_andrew
6mo ago
I've been wondering about adaptive decoding! It seems obvious to me that at some points during decoding (reasoning, "creative thinking") you would want a higher temperature, while at other points (emitting syntactically corre
9.
▲
by
andy_xor_andrew
6mo ago
Not really? If you read it, there is no validation, no correctness signal, no verification, none of that. They're just passing in benchmark inputs, collecting the outputs (regardless of their quality), training on those outputs, and th
10.
▲
by
andy_xor_andrew
6mo ago
I find the branding to be a little odd. Like, it should be a github page with a README that says "here's how to use this." Like, the full explanation of this project is right there in the HN title: "The free AI already o
11.
▲
by
andy_xor_andrew
6mo ago
> x402 is an open, neutral standard for Internet-native payments. It lets anyone on the Internet easily charge, and any client pay on-demand, on a pay-per-use basis. A client, such as an agent, sends a HTTP request and receives a HTTP 40
12.
▲
by
andy_xor_andrew
8mo ago
Curious if someone could enlighten me- How much of these sorts of patches are specifically checking if a certain application is running, and then changing behavior to match what that application expects? And how much of it is simply better
13.
▲
by
andy_xor_andrew
8mo ago
if I'm not mistaken (and I very well may be!) my primary confusion with closures comes from the fact that: the trait they implement (FnOnce / Fn / FnMut) depends entirely upon what happens inside the closure. It will automati
14.
▲
by
andy_xor_andrew
9mo ago
no payoff whatsoever? I just asked Claude to do a task that would have previously taken me four days. Then I got up and got lunch, and when I was back, it was done. I would never make the argument that there are no risks. But there's a
15.
▲
by
andy_xor_andrew
9mo ago
I guess that makes this "standing on the shoulders of fabrications"
16.
▲
by
andy_xor_andrew
10mo ago
> This is a 30B parameter MoE with 3B active parameters Where are you finding that info? Not saying you're wrong; just saying that I didn't see that specified anywhere in the linked page, or on their HF.
17.
▲
by
andy_xor_andrew
10mo ago
> former Dean of Electronics Engineering and Computer Science at Peking University, has noted that Chinese data makes up only 1.3 percent of global large-model datasets (The Paper, March 24). Reflecting these concerns, the Ministry of St
18.
▲
by
andy_xor_andrew
1y ago
I truly, genuinely wanted to like Liquid Glass. I think the default reaction to ANY change in UX, even changes that are generally improvements, is: "I don't like this, it's different!" I thought that'd be the case f
19.
▲
by
andy_xor_andrew
1y ago
> In order to limit the impact of similar issues in the future, all sites on statichost.eu are now created with a statichost.page domain instead. This read like a dark twist in a horror novel - the .page tld is controlled by Google! htt
20.
▲
by
andy_xor_andrew
1y ago
Sure, of course I will trust the report as the source of truth. But I'm interested in the reporting. There are, you know, journalistic standards, which are considered kinda "journalism 101"! For instance, getting the basic fa
21.
▲
by
andy_xor_andrew
1y ago
I read the article (twice) and I still have the impression the pilot was in fact the one in the conference call Opening line: > A US Air Force F-35 pilot spent 50 minutes on an airborne conference call with Lockheed Martin engineers tr
22.
▲
by
andy_xor_andrew
1y ago
https://news.ycombinator.com/item?id=27529697
23.
▲
by
andy_xor_andrew
1y ago
I was wondering this as well. The green box could simply indicate it detected a face, using something like YOLO, or even a simpler technique like some point-and-shoot cameras use to decide where to focus (on faces, obviously).
24.
▲
LLMs Are Magic – Their Applications Should Be, Too
(debugti.me)
1 points
by
andy_xor_andrew
1y ago
|
0 comments
25.
▲
by
andy_xor_andrew
1y ago
Yeah, it's bizarre. Normally the pathway for this kind of thing would be: 1. theorized 2. proven in a research lab 3. not feasible in real-world use (fizzles and dies) if you're lucky the path is like 1. theorized 2. proven in a
26.
▲
Hacker News Clone (Microeval)
(artificialanalysis.ai)
2 points
by
andy_xor_andrew
1y ago
|
0 comments
27.
▲
by
andy_xor_andrew
1y ago
The article mentions AlphaGo/Mu/Zero was not based on Q-Learning - I'm no expert but I thought AlphaGo was based on DeepMind's "Deep Q-Learning"? Is that not right?
28.
▲
by
andy_xor_andrew
1y ago
the magic thing about off-policy techniques such as Q-Learning is that they will converge on an optimal result even if they only ever see sub-optimal training data. For example, you can use a dataset of chess games from agents that move tot
29.
▲
by
andy_xor_andrew
1y ago
in my experience, TTS has been a "pick two" situation: - fast / cheap to run - can clone voices - sounds super realistic from what I can tell, Chatterbox is the first that apparently lets you pick 3! (have not tried it myself
30.
▲
by
andy_xor_andrew
1y ago
It seems like the core innovation in the exploit comes from this observation: - the check for prompt injection happens at the document level (full document is the input) - but in reality, during RAG, they're not retrieving full documen
More ›