Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
robrenaud
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
Ask HN: Is there something like Google style guide for AI-only coded apps?
1 points
by
robrenaud
7mo ago
|
2 comments
32.
▲
by
robrenaud
8mo ago
Please serve well quantized models. If you can get 99 percent of the quality for 50 percent of the cost, that is most times a good tradeoff.
33.
▲
by
robrenaud
8mo ago
Cite a source. Your concrete claim is that, on average, for every $1 of subscription revenue on a monthly subscription, OpenAI and Anthropic were losing $11.50? It seems completely implausible. I could believe that if a $20 sub used every
34.
▲
by
robrenaud
8mo ago
I used to play very competitively, but I've been more chill recently. I just think it's a nice problem/dataset to work with, because of the depth of my understanding of the game.
35.
▲
by
robrenaud
8mo ago
I’ve been experimenting with a live win probability predictor for the 10-player arcade game Killer Queen. The goal is to predict the winner in a causal, event-by-event fashion. Right now I’m struggling to beat a baseline LightGBM model trai
36.
▲
by
robrenaud
8mo ago
A compiler that can turn cash into improved code without round tripping a human is very cool though. As those steps can get longer and succeed more often in more difficult circumstances, what it means to be a software engineer changes a lot
37.
▲
by
robrenaud
9mo ago
Previously, I made a live win probability model for the 5v5 arcade game Killer Queen Arcade from their game events API. Now I am trying to use that model to make: 1. A post game instant replay that shows the most important/pivotal mom
38.
▲
by
robrenaud
9mo ago
Here is research about doctors interpreting test results. It seems to favor GP's view that many doctors struggle to weigh test specificity and sensitivity vs disease base rate. https://bmjopen.bmj.com/content/bmjop
39.
▲
by
robrenaud
9mo ago
I suspect the models would be more useful but perhaps less popular if the semantic content of their answers depended less on the expectations of the prompter.
40.
▲
by
robrenaud
10mo ago
What are you embedding? Are you doing a geo restricted area (small universe?).
41.
▲
by
robrenaud
10mo ago
Paying the local currency with your own cards seems simple and works?
42.
▲
by
robrenaud
10mo ago
I learned Python circa 2000 as a 17 year old. It felt pretty easy to read and write, had minimal surprises, and it made writing simple programs easy. The batteries-includedness was great. It felt like it was designed by a smart guy for pra
43.
▲
by
robrenaud
10mo ago
Omg, it was so frustrating to say: Summarize recent working arxiv url And then it tells me the date is from the future and it simply refuses to fetch the URL.
44.
▲
by
robrenaud
10mo ago
What is harder, beating Lee Sedol at Go, or physically placing stones on a Go board? Which is closer to AGI? Because AlphaGo can only do one. AI could very well be better at formal theorem proving than fields medalists pretty soon. It wil
45.
▲
by
robrenaud
10mo ago
I think you underestimate how powerful lean is, and close it is to the tedious part of formal math. A theorem prover needs consult no outside resource. A formal math LLM-like generator need only consult the theorem prover to get rid of ha
46.
▲
by
robrenaud
10mo ago
Low scores on HLE and ARC AGI might be a good sign. They didn't goodhart their models. ARG AGI in particular doesn't mean much, IMO. It's just some weird hard geometry induction. I don't think it correlates well with
47.
▲
by
robrenaud
10mo ago
They are the LLM whisperers. In the same way Nagel knew what it was like to be a bat, Anthropic has the highest fraction of people who approximately know what it's like to be a frontier ai model.
48.
▲
by
robrenaud
10mo ago
Anything with Jason Weston as a coauthor tends to be pretty well written/readable and often has nice results.
49.
▲
by
robrenaud
10mo ago
I don't think that's a great analogy. LoRAs tend to be adapters bolted onto to systems by people other than the system designers, and they are low rank factorizations. There is nothing low rank or adapter here.
50.
▲
by
robrenaud
10mo ago
From what I've seen at neurips, in terms of most different but maybe viable, it would be this. https://sakana.ai/ctm/ In terms of a fresh perspective on designing learning systems, nested learning seems very inter
51.
▲
by
robrenaud
10mo ago
I agree that pass@k feels a bit weird for large k. But for LLMs, it's a decent proxy for "are the knowledge/skills/circuit necessary to solve the problem somewhere in the model". Note that choices for large k is o
52.
▲
by
robrenaud
10mo ago
RLVR is the more particular term of art in this domain. VR stands for verified rewards and is the single bit per rollout that is the heart of the post. Maybe we can convince dang to update the title.
53.
▲
by
robrenaud
10mo ago
He is essentially expanding upon an idea made by Andrej Karpathy on his podcast about a month prior. Karpathy says that basically "RL sucks" and that it's like "sucking bits of supervision through a straw". https:&
54.
▲
by
robrenaud
10mo ago
I had some json data that I wanted an annotation interface for. So I asked codex to put it into sqlite and make a little annotation webserver. It worked quickly/easily and without hassle. Sqlite supports queries over json-like objec
55.
▲
by
robrenaud
10mo ago
I just wouldn't. RL is nice in that it is handles messy cases where you don't have per example labels. How do you build a learned chess playing bot? Essentially the state of the art is to find a clever way of turning the problem o
56.
▲
by
robrenaud
11mo ago
I worked on a similiar problem about a year ago, on large dense models. https://www.lesswrong.com/posts/PkeB4TLxgaNnSmddg/scaling-sp... In both cases, the goal is to actually learn a concrete circuit inside a netw
57.
▲
by
robrenaud
11mo ago
Benchmarks need to change. There is a 4 choice choice question. Your best guess is the answer is B, at about 35% chance of being right. If you are graded on fraction of questions answered correctedly, the optimization pressure is simply t
58.
▲
by
robrenaud
11mo ago
Open source AI is just a lost term. It has been co-opted. If the weights are released, it's open source. Not because that makes sense, not because it's right, but because that's the unfortunate marketting term that has stu
59.
▲
by
robrenaud
11mo ago
LLMs do have some internal representations that predict pretty well when they are making stuff up. https://arxiv.org/abs/2509.03531v1 - We present a cheap, scalable method for real-time identification of hallucinated t
60.
▲
by
robrenaud
11mo ago
pytorch
More ›