Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ElFitz
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
61.
▲
by
ElFitz
4mo ago
I am glad to have checked that one out. I don’t know whether to be amazed or horrified.
62.
▲
by
ElFitz
4mo ago
> Claude Code is the cancer that will kill the patient, Boris is the the Kardashian version of Karpathy, with less business sense. I am not sure I understand.
63.
▲
by
ElFitz
4mo ago
That, especially the conclusion, is hilariously bad.
64.
▲
by
ElFitz
4mo ago
I’ve been experimenting with two things on this: - multi-model consensus, with multiple cross-review rounds. Obviously, the number of inference tasks explodes with the number of models. Led to some interesting results [^0]. - giving an agen
65.
▲
by
ElFitz
4mo ago
> With such eggregious trillions of dollars worth of money (basically the whole economy getting floated by tech), you are bound to see people within this do the grift playbook and talk about themselves and succeed and that has become the
66.
▲
by
ElFitz
4mo ago
I’m working on Descartes[^0]. First to help diagnose what’s wrong with a machine. I’ve started implementing actual background monitoring of the system, and next will be letting an agent build its own layers of tailor-made deterministic rule
67.
▲
by
ElFitz
4mo ago
From what I found, current state of the art on modelling "reaction space" with graphs that is to use "hypergraphs" where edges can lead to more than one node[^0]. But I am just someone who got curious; not even an amateu
68.
▲
by
ElFitz
4mo ago
> Feels complex like solving a Rubik's cube to write down synthesis steps but it is all a sequence of memorized tricks. Do Cannizaro if you want this, Bergmann to do that. I remember two years ago, when I actually got into using gra
69.
▲
by
ElFitz
4mo ago
> Ah, I see, sort of like figuring out the boundaries of your knowledge base and seeing if you have missed any connections between concepts? Yeah. And also resurface relevant old notes, ideas, and web clippings I have forgotten. Two rece
70.
▲
by
ElFitz
4mo ago
Isn't that what we call public debt?
71.
▲
by
ElFitz
4mo ago
The graph visualisation itself isn’t much use in my experience. But being able to tie related notes together, and see at the bottom of one which other notes reference it is interesting. Even more now that a LLM can take care of the actual t
72.
▲
by
ElFitz
4mo ago
I think it’s not the first time the US has used that sort of interpretation of the law. There’s this one[^0] but also, I believe, an older case, also involving Microsoft, about data in Ireland. But I can’t find it. [0]: https://h
73.
▲
by
ElFitz
4mo ago
These kinds of situations are why I gave my AI agents stray thoughts (automated insights / suggestions from a separate llm call with some curated context) that trigger on loop / rabbit hole detection. Quite a bit of false positive
74.
▲
by
ElFitz
4mo ago
Haha. Yes. Much smaller scale versions of this led me to joke with a coding agent that LLMs tended to converge towards "Large corporation infrastructure best practices" when designing cloud infrastructure, when it was only me work
75.
▲
by
ElFitz
4mo ago
I’ve been making Codex and Claude get their work reviewed by most recent best performing model of their own family, and each other’s, for months. On top of that, we have been running multi-model AI reviews on every PR through their respecti
76.
▲
by
ElFitz
4mo ago
> and I have the feeling that the harness is much more important than the consensus expectation. Is that really the consensus? There’s been a bit of literature lately on that. Can’t find the one about looking into whether or not the harn
77.
▲
by
ElFitz
4mo ago
I’ve had a similar experience. Pointing out past suboptimal / failing behaviours to new opus sessions would almost always actually create a sort of "anchoring bias" that would drive the agents towards exhibiting the failure m
78.
▲
by
ElFitz
4mo ago
I love hard caps, and am tired of cloud services not even offering those. May make sense for large companies. Makes no sense for hobbyists and small companies. But maybe that's the point?
79.
▲
by
ElFitz
4mo ago
That’s what evals are for. And there’s no reason evals can’t be done on multi-turn agents in a loop (or not): it’s pretty much what all these benchmarks do.
80.
▲
by
ElFitz
4mo ago
I remember Windows Phone 8 existed , but that’s pretty much it. And yes, that’s the big question: what’s in it for the app publishers? :/
81.
▲
by
ElFitz
4mo ago
A few months back someone reverse-engineered private ANE APIs and shown some significant performance improvements compared to CoreML and Metal, on both inference and training. - https://maderix.substack.com/p/inside-the
82.
▲
by
ElFitz
4mo ago
Oh, neat. Totally missed it, thanks!
83.
▲
by
ElFitz
4mo ago
Yes, but you can have this by providing surfaces as UI components to Siri. Not sure if they do that (yet), but no reason an app couldn’t expose "Here’s what you can use to present data of shape X", or "here’s a UI for process
84.
▲
by
ElFitz
4mo ago
I used to wonder what "apps" might become in an "App Intent-first" world. Bundles that provide data and capabilities to iOS and Siri? And perhaps libraries of UI components to display and interact with said data? But the
85.
▲
by
ElFitz
4mo ago
Wait. You worked on Guild Wars, Starcraft, Warcraft, and Diablo? This place is incredible.
86.
▲
by
ElFitz
4mo ago
I’m working on Descartes[^0]. First to help diagnose what’s wrong with a machine. Later to help manage and monitor it by letting an agent build layers of tailor-made deterministic rules and statistical models, a bit like the description of
87.
▲
by
ElFitz
4mo ago
Perhaps that’s it. I would tend to agree with his position, I think, but don’t appreciate being preached to. Even less so when I agree with what’s being said.
88.
▲
by
ElFitz
4mo ago
You can disagree. Sarcastically, or otherwise. But I think you may be reading more into my comment than I put there. I’m not attacking the piece. I’m not saying it’s right. I’m not saying it’s wrong. What I’m saying is, the tone made it har
89.
▲
by
ElFitz
4mo ago
I find it difficult to separate this piece’s tone from its content. The tone puts me off and makes it hard for me to judge it on its merits, despite some of the arguments seeming sound and well supported.
90.
▲
by
ElFitz
4mo ago
> Many website will request my personal physical address for trivial matters like billing or delivery. Some will even require it for no actual reason at all. Do I need to give my living address when I buy a sandwich? Then why would I nee
More ›