4 ms·
This looks very interesting. The easiest way to navigate to the start of this series of articles seems to be https://www.gilesthomas.com/til-deep-dives/page/2
by andrehacker 1y ago
This looks very interesting.
The easiest way to navigate to the start of this series of articles seems to be
https://www.gilesthomas.com/til-deep-dives/page/2 https://www.gilesthomas.com/til-deep-dives/page/2
Now if I only could find some time...
- sitkack 1y agohttps://news.ycombinator.com/from?site=gilesthomas.com https://news.ycombinator.com/from?site=gilesthomas.com
- Tokumei-no-hito 1y agomaybe it renders differently on mobile but this was the first entry for me. you can use the nav at the end to continue to the next part https://www.gilesthomas.com/2024/12/llm-from-scratch-1 https://www.gilesthomas.com/2024/12/llm-from-scratch-1
- gpjt 1y agoAuthor here: I endorse this comment ;-) That's definitely the route I've optimised for for reading the series.
- badsectoracula 1y agoToo bad the book seems to be using Python and some external library like tiktokens just from chapter 2, meaning that it'll basically stop working next week or so, like everything Python, making the whole thing much harder to follow in the future. Meanwhile i learned the basics of machine learning and (mainly) neural networks from a book written in 1997[0] - which i read last year[1]. It barely had any code and that code was written in C, meaning it'd still more or less work (though i didn't had to try it since the book descriptions were fine on their own). Now, Python was supposedly designed to look kinda like pseudocode, so using it for a book could be fine, but at least it should avoid relying on external libraries that do not come with the language itself - and preferably stick to stuff that have equivalent to other languages too. [0] https://www.cs.cmu.edu/~tom/mlbook.html https://www.cs.cmu.edu/~tom/mlbook.html [1] which is why i make this comment (and to address the apparent downvotes): even if i get the book now i might end up reading it in 3-4 years. Stuff not working will be a major obstacle. If the book is good, it might end up been recommended by people 2-3 years from now and some people may end up getting it and/or reading it even later in time. So it is important for the book to be self-contained, at least when it comes to books that try to teach the theory/ideas behind things.
- andrehacker 1y agoMyeah, C and C++ have the advantage that the compilers support compile for old versions of the language. The languages are in much flux partly because of security problems, partly because features are added from other languages. That means that linking to external libraries using the older language version will fail unless you keep the old version around simply because the maintainer of the external library DID upgrade. Python is not popular in ML because it is a great language but because of the ecosystem: numpy, pandas, pytorch and everything built on those allows you to do the higher level ML coding without having to reinvent efficient matrix operations for a given hardware infrastructure.
- badsectoracula 1y ago(i assume with "The languages are in much flux" you meant python and not c/c++ because these aren't in flux) Yeah i get why Python is currently used[0] and for a theory-focused book Python would still work to outline the algorithms - worst case you boot up an old version of Python in Docker or a VM, but it'd still require using only what is available out of the box in Python. And depending on how the book is written, it may not even be necessary. That said there are other alternatives nowadays and when trying to learn the theory you may not need to use the most efficient stuff. Using C, C++, Go, Java, C# or whatever other language with a decent backwards compatibility track record (so that it can work in 5-10 years) should be possible and all of these should have some small (if not necessarily uberefficient) library for the calculations you may want to do that you can distribute alongside the book for those who want to try the code out. [0] even if i wish people would stick on using it only for the testing/experimentation phase and move to something more robust and future proof for stuff meant to be used by others
- andrehacker 1y ago"The languages are in much flux" you meant python and not c/c++ because these aren't in flux No I meant C++. 2011 14882:2011[44] C++11 2014 14882:2014[45] C++14 2017 14882:2017[46] C++17 2020 14882:2020[47] C++20 2024 14882:2024[17] C++23 That is 4 major language changes in 10 years. As a S/W manager in an enterprise context having to coordinate upgrades of multi-million LOC codebases for mandated security compliance, C++ is not the silver bullet in handling the version problem that exists in every eco system. As said, the compilers/linkers allow you to run in compatibility mode so as long as you don't care about the new features (and the company you work for doesn't) then, yes, C/C++ is easier for managing legacy code.