4 ms·
The most important thing to learn for most practical purposes is what the thing can actually do. There's a lot of fuzzy thinking around ML - "throw AI at it and
by GeneralMayhem 3y ago
The most important thing to learn for most practical purposes is what the thing can actually do. There's a lot of fuzzy thinking around ML - "throw AI at it and it'll magically get better!" Sources like Karpathy's recent video on what LLMs actually do are good anti-hype for the lay audience, but getting good practical working knowledge that's a level deeper is tough without working through it. You don't have to memorize all the math, but it's good to get a feel for the "interface" of the components. What is it that each model technique actually does - especially at inference time, where it needs to be well-integrated with the rest of the stack?
In terms of continued relevance - "deep learning", meaning, dense neural nets trained to optimize a particular function, haven't fundamentally changed in practice in ~15 years (and much longer than that in theory), and are still way more important and broadly used than the OpenAI stuff for most purposes. Anything that involves numerical estimation (e.g., ad optimization, financial modeling) is not going to use LLMs, it's going to use a purpose-built model as part of a larger system. The interface of "put numbers in, get number[s] out" is more explainable, easier to integrate with the rest of your software stack, and more measurable. It has error bars that are understandable and occasionally even consistent. It has a controllable interface that won't suddenly decide to blurt corporate secrets or forget how to serialize JSON. And it has much, much lower latency and cost - any time you're trying to render a web page in under 100ms or run an optimization over millions of options, generative AI just isn't a practical option (and is unlikely to become one, IMO).
I don't have a significant math or theoretical ML background, but I've spent most of the last 10 years working side by side with ML experts on infra, data pipelines, and monitoring. I'm not sure I could integrate the sigmoid off the top of my head, but that's not what's important - I've done it once, enough to have some idea how the function behaves, and I know how to reason about it as a black box component.
- mvkel 3y agoAwesome response, and a reasoned take.
- Breza 3y agoTerrific explanation, and it matches my experience running a data science team. I encourage my team to start with the simplest possible approach to every problem, which requires understanding how different algorithms work. Does this project require a t-test, XGBoost, a convolutional neural network, something else? What if we recode the dependent variable from numeric to binary?
- wilkystyle 3y ago> Sources like Karpathy's recent video on what LLMs actually do are good anti-hype for the lay audience Which video is this?
- bjacobt 3y agoI believe op means Intro to Large Language Model https://youtu.be/zjkBMFhNj_g?si=XQQ3p92ajuQYOyqN https://youtu.be/zjkBMFhNj_g?si=XQQ3p92ajuQYOyqN
- GeneralMayhem 3y agoYep, that's the one I meant - sorry, should have linked. His series on making a GPT from scratch is also great for building intuition specifically about text-based generative AI, with an audience of software developers.