4 ms·
Yes, most people (including myself) do not understand how modern LLMs work (especially if we consider the most recent architectural and training improvements).
by pushedx 7mo ago
Yes, most people (including myself) do not understand how modern LLMs work (especially if we consider the most recent architectural and training improvements).
There's the 3b1b video series which does a pretty good job, but now we are interfacing with models that probably have parameter counts in each layer larger than the first models that we interacted with.
The novel insights that these models can produce is truly shocking, I would guess even for someone who does understand the latest techniques.
- measurablefunc 7mo agoWhat's the latest novel insight you have encountered?
- brookst 7mo agoNot the person you asked, and “novel” is a minefield. What’s the last novel anything, in the sense you can’t trace a precursor or reference? But.. I recently had a LLM suggest an approach to negative mold-making that was novel to me. Long story, but basically isolating the gross geometry and using NURBS booleans for that, plus mesh addition/subtraction for details. I’m sure there’s prior art out there, but that’s true for pretty much everything.
- measurablefunc 7mo agoI don't know, that's why I asked b/c I always see a lot of empty platitudes when it comes to LLM praise so I'm curious to see if people can actually back up their claims. I haven't done any 3D modeling so I'll take your word for it but I can tell you that I am working on a very simple interpreter & bytecode compiler for a subset of Erlang & I have yet to see anything novel or even useful from any of the coding assistants. One might naively think that there is enough literature on interpreters & compilers for coding agents to pretty much accomplish the task in one go but that's not what happens in practice.
- pushedx 7mo agoWhich agents are you using, and are you using them in an agent mode (Codex, Claude Code etc.)? The difference in quality of output between Claude Sonnet and Claude Opus is around an order of magnitude. The results that you can get from agent mode vs using a chat bot are around two orders of magnitude.
- measurablefunc 7mo agoThe workflow is not the issue. You are welcome to try the same challenge yourself if you want. Extra test cases (https://drive.proton.me/urls/6Z6557R2WG#n83c6DP6mDfc https://drive.proton.me/urls/6Z6557R2WG#n83c6DP6mDfc) & specification (https://claude.ai/public/artifacts/5581b499-a471-4d58-8e05-147ab7c3ef4e https://claude.ai/public/artifacts/5581b499-a471-4d58-8e05-1...). I know enough about compilers, bytecode VMs, parsers, & interpreters to know that this is well within the capabilities of any reasonably good software engineer but the implementation from Gemini 3.1 Pro (high & low) & Claude Opus 4.6 (thinking) have been less than impressive.
- Kim_Bruning 7mo agoPossibly a dumb question: but are you running this in claude code, or an ide, or basically what are you using to allow for iteration?
- measurablefunc 7mo agoI'm using Google's antigravity IDE. I initially had it configured to run allowed commands (cargo add|build|check|run, testing shell scripts, performance profiling shell scripts, etc.) so that it would iterate & fix bugs w/ as little intervention from me as possible but all it did was burn through the daily allotted tokens so I switched to more "manual" guidance & made a lot more progress w/o burning through the daily limits. What I've learned from this experiment is that the hype does not actually live up to the reality. Maybe the next iteration will manage the task better than the current one but it's obvious that basic compiler & bytecode virtual machine design in a language like Rust is still beyond the capabilities of the current coding agents & whoever thinks I'm wrong is welcome to implement the linked specification to see how far they can get by just "vibing".
- kennyloginz 7mo agoThere is prior art, so it’s not novel.
- auraham 7mo agoI highly recommend Build a large language model from scratch [1] by Sebastian Raschka. It provides a clear explanation of the building blocks used in the first versions of ChatGPT (GPT 2 if I recall correctly). The output of the model is a huge vector of n elements, where n is the number of tokens in the vocabulary. We use that huge vector as a probability distribution to sample the next token given an input sequence (i.e., a prompt). Under the hood, the model has several building blocks like tokenization, skip connections, self attention, masking, etc. The author makes a great job explaining all the concepts. It is very useful to understand how LLMs works. [1] https://www.manning.com/books/build-a-large-language-model-from-scratch https://www.manning.com/books/build-a-large-language-model-f...
- phreeza 7mo agoBut this is missing exactly the gap which OP seems to have, which is going from a next token predictor (a language model in the classical sense) to an instruction finetuned, RLHF-ed and "harnessed" tool?
- js8 7mo agoThe book has a sequel https://www.manning.com/books/build-a-reasoning-model-from-scratch https://www.manning.com/books/build-a-reasoning-model-from-s... It will give you an answer to the extent anybody can.