Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
weitendorf
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
31.
▲
by
weitendorf
2mo ago
IME this is a strong/reliable model smell, you typically see smarter and less benchmaxxed models' thinking traces spending more time exploring the solution space, and benchmaxxed models more time trying to refine/decide on th
32.
▲
by
weitendorf
2mo ago
I think it's much simpler than that, they are indeed trying to counter against sycophancy and models lying (the decision whether or not to lie to the human user, take shortcuts, etc. comes up in their thinking traces pretty commonly II
33.
▲
by
weitendorf
2mo ago
Frontier labs are not a monolithic entity. There is a clear self-verification/difficulty ramp in cybersecurity, and it is a very valuable as a skill both offensively and defensively. So it is absolutely certain that someone, somewhere,
34.
▲
by
weitendorf
2mo ago
I think what we really need is project/thread-scale continual learning. The problem is that the important parts of the conversation to you are the novel bits you just did, rather than all the context building the agent did to get to th
35.
▲
by
weitendorf
2mo ago
I do this a lot and you have to be really careful to clean these up or qualify/steer agents around them. They’ll often be very emphatically confident about some assumption or implication they made, and if another agent stumbles upon th
36.
▲
by
weitendorf
2mo ago
Claude Code makes agents reasonably aware of where their log files/history/etc are and get stored. Generally they’ll work with them without explicitly being told (especially to recover broken sub agents, corrupted sessions, etc) t
37.
▲
by
weitendorf
2mo ago
Right now we're going through multiple concurrent but slightly non-overlapping transitions where AI is almost good enough for something, and seeing a predictable surge of spam/bad quality crap people don't like, followed by
38.
▲
by
weitendorf
2mo ago
People are still figuring out the social norms surrounding "deepfakes" and I'm convinced there are some versions of the future where it becomes normalized for content under essentially fair use/free speech doctrine, an
39.
▲
by
weitendorf
2mo ago
I've been dabbling in this space on the application layer and have spent a few hundred myself experimenting with video models. It can be really entertaining/addicting to build with them because you're essentially pulling the
40.
▲
by
weitendorf
2mo ago
If you understand finance and aren’t specifically attempting to arb on that timescale, you actually want to participate in markets with those participants, because their presence gives you less variance/better price discovery on the sc
41.
▲
by
weitendorf
2mo ago
This already happened 10-20 years ago when personal finance got big on the Internet, it’s just taking a long time to play out. It was never about ROI anyway, just preservation of capital and peace of mind - makes a lot of sense in the analo
42.
▲
by
weitendorf
2mo ago
Are you sure this is Google's doing? Sites can tell Google what they want indexed and want not to show, and Google's crawler also is supposed to not hammer sites with unwanted indexing, and has quotas where if your site has a ma
43.
▲
by
weitendorf
2mo ago
It's true, but part of the problem is that there is an oversupply of people seeking high-skill jobs before they have had the time to actually build those skills. Whereas if you need to a guy to haul lumber around, and there's a gu
44.
▲
by
weitendorf
2mo ago
There are two that I know of, and neither seem to do a better job at solving my problem of providing a place to host and distribute open source software. I hosted my own forgejo server used by actual other people * very slightly* before it
45.
▲
by
weitendorf
2mo ago
99.9% of clients are behind NAT or working on code owned by someone other than just themselves. Unless your "decentralized" repos are on two devices you control on the same private network, or you think carrying hard drives around
46.
▲
by
weitendorf
2mo ago
It's usually not 24/7, it's that when I use just 1 agentic coding tool there's too much variance/hand-holding to go do something else, and also often a lot of downtime in between turns as they work. So it's a t
47.
▲
by
weitendorf
2mo ago
I'm not going to pretend I know how to do Github better than Github, but I almost can't believe how poor their support is for fine grained access tokens, org/individual credentials, etc. are in the same era as Copilot/VS
48.
▲
by
weitendorf
2mo ago
My experience with Codex is that it goes off to do its thing for 10-60 minutes and either nails it and comes back with everything done, or comes back with something that I almost can’t believe a near-SOTA model would think I wanted based on
49.
▲
by
weitendorf
2mo ago
It’s interesting how deep-fried LLMs are getting the more post-training they receive. They’re undoubtedly getting much smarter overall, but also much weirder. Before they were just trying to model our behavior, only really having us to lear
50.
▲
by
weitendorf
3mo ago
I open source as much software as I can because I want the models to train on it and get better at it!
51.
▲
by
weitendorf
3mo ago
“He’s only being good because he likes how it feels, or values goodness, or exists in a social context where doing the right thing is socially rewarded! That has no bearing on whether he, intrinsically, is good!” Not to get all philosophica
52.
▲
by
weitendorf
3mo ago
I think it’s one thing to give you something free forever, and another to deliberately foster dependency / suck all the oxygen out of the room just to abuse it. It’s also not a binary thing. You could truly start off with the noblest i
53.
▲
by
weitendorf
3mo ago
Free + Open is simply how you earn credibility and user trust when you don’t already have it. That doesn’t mean the trust is unearned once gained, or a bait and switch, or purely Machiavellian either btw. Consumers and businesses need cre
54.
▲
by
weitendorf
3mo ago
Because model providers are not optimizing for being indistinguishable from human text, and in fact, there is more value/demand in modeling a different distribution (ie an “agent” capable of producing vast amounts of concrete procedura
55.
▲
by
weitendorf
3mo ago
I think configurability depends on how important your tool is to the core job function or role being performed, where it becomes very valuable for helping them directly perform the tasks they and their employer value, vs how much it allows
56.
▲
by
weitendorf
3mo ago
It means the same thing to you, but not to the whole spectrum of people using AI. You literally see it on Reddit all the time where people are complaining about the same model either over-engineering or doing too much, vs it being requiring
57.
▲
by
weitendorf
3mo ago
You see this a lot with beginners, because until you’ve done the work long enough to truly know what works, you only really know what you have seen through other people’s performance of the work (to the degree it is even understandable and
58.
▲
by
weitendorf
3mo ago
Just got it working with codex in a container! FYI I think there is a bug most others will run into at the Codex:Muse interface. It's some kind of parsing or integration error due to what I think is codex not anticipating server-side t
59.
▲
by
weitendorf
3mo ago
I’ve spent my time very similarly working on my own voice stack project, but having also seen how non-developers use AI or experience technology in general, I truly think they are better served with a different UX and product than what we h
60.
▲
by
weitendorf
3mo ago
I think the bare truth is that the target audience for this product is not people who are highly particular about terminology in answers involving vector mathematics. It’s a different set of tradeoffs for users that don’t already have stron
More ›