3 ms·
> which is more accurately described as compressed efficiently onto latent space. The actual difference between solving compression+search vs novel creative sy
by photonthug 1y ago
> which is more accurately described as compressed efficiently onto latent space.
The actual difference between solving compression+search vs novel creative synthesis / emergent "understanding" from mere tokens is always going to be hard to spot with these huge cloud-based models that drank up the whole internet. (Yes.. this is also true for domain experts in whatever content is being generated.)
I feel like people who are very optimistic about LLM capabilities for the later just need to produce simple products to prove their case; for example, drink up all the man pages, a few thousand advanced shell scripts that are easily obtainable, and some subset of stack-overflow. And BAM, you should have a offline bash oracle that makes this tiny subset of general programming endeavor a completely solved problem.
Currently, smaller offline models still routinely confuse the semantics of "|" vs "||". (An embarrassing statistical aberration that is more like the kind of issue you'd expect with old school markov chains than a human-style category error or something.) Naturally if you take the same problem to a huge cloud model you won't have the same issue, but the argument that it "understands" anything is pointless, because the data-set is so big that of course search/compression starts to look like genuine understanding/synthesis and really the two can no longer be separated. Currently it looks more likely this fundamental problem will be "solved" with increased tool use and guess-and-check approaches. The problem then is that the basic issue just comes back anyway, because it cripples generation of an appropriate test-harness!
More devs do seem to be coming around to this measured, non-hype kind of stance gradually though. I've seen more people mentioning stuff like, "wait, why can't it write simple programs in a well specified esolang?" and similar
- skydhash 1y agoA naive thought: What you would get if you hardcode the language grammar and not let the training discern it, so instead of it, kinda like an expert system constraining its output?
- earnestinger 1y agoWhat kind of hard coding do you have in mind? How the technique would look like?
- skydhash 1y agoWe already know the keywords of the language and symbols from the standard library and othe major ones. As well as the rules of the grammar. So the weight can be biased against that. Not sure how that would work, though. I don’t think that would help with natural language to programming language, but that can probably help with patterns, kinda like a powerful suggestion engine.
- Xmd5a 1y agoword2vec meets category theory https://en.wikipedia.org/wiki/DisCoCat https://en.wikipedia.org/wiki/DisCoCat >In this post, we are going to build a generalization of Transformer models that can operate on (almost) arbitrary structures such as functions, graphs, probability distributions, not just matrices and vectors. https://cybercat.institute/2025/02/12/transformers-applicative-functors/ https://cybercat.institute/2025/02/12/transformers-applicati...
- ModernMech 1y agoI try to do something like this when I ask the AI to generate tests -- I'll cook up a grammar and feed it to the LLM in a prompt, and then ask it to generate strings from the grammar. It's pretty good at it, but it'll produce mistakes, which is why I'll write a parser for the grammar and have the LLM feed the strings it makes through the parser and correct them if they're wrong. Works well.