Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
xpct
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
91.
▲
by
xpct
2mo ago
I don't think we're approaching the limit of deterministic prompt -> source code mapping any time soon. Small variability in prompts produces medium variability in outputs. Building on previous outputs only extends the variabil
92.
▲
by
xpct
2mo ago
Why is it draconian? If we take Dec-2025 as the point where the models became good for real work (agreed by many), not even a year has passed since then. The world hasn't changed, and there are still many open legal questions.
93.
▲
by
xpct
2mo ago
I wonder whether there are any obvious third party targets that would affect a large portion of unsuspecting LLMs. Perhaps the Wikipedia page of an unfolding geopolitical event, poisoning models which fetch it? Some other malleable websites
94.
▲
by
xpct
2mo ago
It's probably fuzzily fixable by including instruction authority levels in the training data. Can't expect much more than that, given that the model itself is fuzzy.
95.
▲
by
xpct
2mo ago
Lately, I've been having less and less success using the models for work. Simple things like: 'write this in a separate file' (writes it in the same file) 'format this aiming for 5 LoC' (emits newline after every co
96.
▲
by
xpct
2mo ago
> This is a reason why many people get upset when a movie adaptation is made I still haven't watched the Dune movies because of this reason. I liked the book a lot, and had a very personal image of what the world looked like, and wa
97.
▲
by
xpct
3mo ago
I imagine this varies by person, but thinking about a source file is already reminiscent of 2D space navigation, where the X axis is typically gated by conditions. In a similar vein, when I think about functions from different source files,
98.
▲
by
xpct
3mo ago
I think Google can play both sides here: they may end up a frontier lab if OpenAI/Anthropic dissolve, or a regular LLM provider, in which case they could be serving open models.
99.
▲
by
xpct
3mo ago
That's the type of culture that made me leave my last job. Feels very inhuman when you see your manager every week, and they still make you fill in robotic questionnaires for potential salary raise. It's deeply demotivating to lea
100.
▲
by
xpct
3mo ago
> Hard to see take-off stopping I think it's reasonable to assume that we're close to, or already at superhuman cybersecurity capabilities at certain domains. But reaching superhuman abilities at one domain doesn't guarant
101.
▲
by
xpct
3mo ago
More curiously, why did it feel the incentive to find the solutions? Would its CoT include "the only way to solve this is to download the test set", or would it include "I'd like to inspect a few entries from the test se
102.
▲
by
xpct
3mo ago
> I'm still undecided on if this that moment If it's a serious incident, then a post hoc with detailed description of the event is coming. So far, none of the companies have released anything close to it when describing their i
103.
▲
by
xpct
3mo ago
Well I'd say these are different risks. It's either tied to the agreement Sony has with the movie provider, or with the platform itself. Either one could pull out. Or, my point, the company could also go under. What is the agreeme
104.
▲
by
xpct
3mo ago
If you bought movies on a digital platform that would later go under (could be Sony one day), what would happen to your collection? Is it transferable in any way? If not, it's already a risk no matter which platform you use.
105.
▲
by
xpct
3mo ago
Please don't frame these as technological crimes, nobody has yet been prosecuted for that specifically.
106.
▲
by
xpct
3mo ago
I'd add that this is also how it used to work with Google search too: there's been facts that would show up as the first result, then you'd just forget it 10 minutes later and have to re-google it. I know this happened to me
107.
▲
by
xpct
3mo ago
Thank you for sharing. The way I reasoned about it myself: to make better predictions, we should know what type of outcomes are likely. We can express these outcomes by doing computations in some of the layers, and the training signal adjus
108.
▲
by
xpct
3mo ago
That's the first thing I noticed as well. Does that mean Africa and South America will have Starlink internet before Scandinavia?
109.
▲
by
xpct
3mo ago
I really don't understand why we feel the need to drop the review aspect. Programming with LLMs is a very non-linear process, there's no prompt-to-code mapping where editing a part of the prompt would produce an identical code con
110.
▲
by
xpct
3mo ago
What does it mean for coding to be solved, and software engineering to not have been solved? If it implies there's now a set protocol that can be followed to reach good results, how has that not been the case before? If it means that w
111.
▲
by
xpct
3mo ago
I think it's related that we as humans see when something becomes hard to reason about, and decide to refactor it. I'm not sure whether an LLM with full ownership of a codebase could do that, for its own benefit. I do see a worl
112.
▲
by
xpct
3mo ago
(If) something like the current LLM/agent paradigm remains in a few years, and companies settle down into their respective niches, I imagine more user-friendly tools will be built, with more control over subagent spawning, context, cac
113.
▲
by
xpct
3mo ago
I think it's useful for learning about unknown unknowns. If you don't have a clear direction, it's entirely fine to start with a university course then stop when you get a feeling for what you really need.
114.
▲
by
xpct
3mo ago
Why's that bad? It's valuable to discuss whether a presentation should have been a blog post, a video, or a tech talk. The same logic applies to LLM text. If people consistently describe LLM text as having low information density
115.
▲
by
xpct
3mo ago
I personally find that models are trending towards ignoring any instructions given to them, so depend more on vendor-instilled behavior. Anecdotal, but I had bad experiences with OAI's new 5.6.
116.
▲
by
xpct
3mo ago
I love the idea of this, but I don't think I could convince my friends to use it over mainstream platforms.
117.
▲
by
xpct
3mo ago
We wouldn't be having debates about it every other day if it was
118.
▲
by
xpct
3mo ago
On some levels I tried guessing, starting with different letters, and still got into some local minima in my brain where I couldn't guess the word, only for it to be something obvious like 'pound'.
119.
▲
by
xpct
3mo ago
Wow I suck! Played today's and a few others, and got ~4-6/18. I like the timer. For 4 letters it's enough time to guess, but I have no idea what they mean: 'Doby', 'Etas'
120.
▲
by
xpct
3mo ago
It sucks, but that ship has sailed. I'm not sure how they'll handle the issue of most new repos being AI generated, if they continue using new code for training. If the world accepts LLMs as a valid method for license washing, I d
More ›