15 ms·
Tinygrad: A simple and powerful neural network framework
- emaro 4y agoI love this website. Their style tag literally is: <style> body { font-family:'Lucida Console', monospace } </style> Also look like a very cool project.
- arketyp 4y agoI like this too and I don't understand the downvotes. It says a lot about the philosophy of the project. Minimalist, bold, brutalist, no-frills first principles thinking. For better and worse.
- ivalm 4y ago“Almost 9k stars” is actually 7.3k stars… But otherwise very cool project :)
- learndeeply 4y agoThe code is very easy to read. Doesn't seem like there's data/model parallelism support for training, which will be important for real-world use.
- DeathArrow 4y agoI believe neural networks are over hyped sometimes. They are not always the best tool for the job. There are lots of other ML techniques such as SVM, naive Bayes, k-nearest neighbor, decision tree, logistic regression, random forest etc. nobody is using because they lack the hype factor. If something lacks some keywords like neural network, deep learning, reinforced learning, than it is deemed not cool.
- minimaxir 4y agoThe problems where traditional ML works best and the problems where Transformers or ConvNets work best are usually two different domains. AI is not a buzzword.
- DeathArrow 4y ago>The problems where traditional ML works best and the problems where Transformers or ConvNets work best are usually two different domains. Yes and we are using NN for everything.
- adamsmith143 4y agoBut we aren't. Outside of using AEs for embeddings and then feeding them through a boosted tree model I don't know anyone using NNs for tabular data. We all use XGBoost or Catboost, etc.
- learndeeply 4y agoI can't think of anything that neural nets can't beat, except small tabular data with boosted decision trees. Can you give some examples?
- zelphirkalt 4y agoExplicability is a big part of it It is often worth being a percent less accurat but having an explainable result.
- sakras 4y agoI must say they gained instant credibility with the minimalistic website given how fast it loaded. Code looks simple and easy to follow, and I love how the comments are constantly mentioning hardware characteristics, making maxing the hardware the goal. It seems that it’s trying to achieve this by jitting optimal code for the operations at hand rather than hand-optimizing kernels, and betting that the small number of operations will make tuning the codegen tractable. I haven’t kept up much with what’s happening in ML, but at least in the realm of columnar database engines, interpreting a series of hand-optimized kernels seems to be the dominant approach over compiling a vectorized query plan. Are compilers good enough at optimizing ML operations that specializing on input shape makes a difference over hand-tuned kernels?
- kklisura 4y agoIt's geohot. He comes with credibility. [1] [1] https://en.wikipedia.org/wiki/George_Hotz https://en.wikipedia.org/wiki/George_Hotz
- matesz 4y agoIf anybody is dealing with procrastination watch George Hotz live streaming 10h straight working on this library [1][2]. Does he take some supplements to do this? There is even 19.5h stream [3]. Actually I have local obs setup to record myself, just instead of streaming I do recordings for my own inspection. Important part is to do the inspection after. It works wonders. [1] https://youtu.be/GXy5eVwnL_Q https://youtu.be/GXy5eVwnL_Q [2] https://m.youtube.com/watch?v=Cb2KwcnDKrk https://m.youtube.com/watch?v=Cb2KwcnDKrk [3] no joke, 19.5h stream https://www.youtube.com/watch?v=xc0jGZYFQLQ https://www.youtube.com/watch?v=xc0jGZYFQLQ
- langsoul-com 4y agoWhat's the file size for your recordings? 5 hours of 720p would be huge.
- albert_e 4y agoYouTube lets you livestream from OBS but mark the stream as private. You get unlimited free storage of your streams for your personal use that way without the need for any local storage at all. I haven't come across any limits or downsides to this yet but happy to be corrected.
- matesz 4y agoThe bigger problem is recording taking too much cpu. That's why I don't record full work day, just chunks, whenever I feel I procrastinate. Youtube is an option here, I've tested it however cpu problem doesn't go away. My plan is to make this obs-ndi plugin work on ubuntu, so I will be able to record on ubuntu to take the load off of mac which is my primary laptop. PS. I forgot to read obs-ndi instructions properly, it works ok so now I can delegate regording to second laptop
- terafo 4y agoIt wouldn't be huge, I once recorded a week of me using my pc(so around 80 hours in total), and it was sub-100 gigs. It was 1080p with decent quality, don't remember FPS though.
- 4y ago
- fragmede 4y agoOf course, the stable diffusion tie-in is not to be missed! https://github.com/geohot/tinygrad/blob/master/examples/stable_diffusion.py https://github.com/geohot/tinygrad/blob/master/examples/stab...
- jamesrom 4y agotinygrad core is over 1000 loc now[1]. If anyone was looking for a fun weekend project :) https://github.com/geohot/tinygrad/blob/master/.github/workflows/test.yml#L9 https://github.com/geohot/tinygrad/blob/master/.github/workf...
- KptMarchewa 4y agoIt does achieve that by being most horizontally dense Python code I've ever seen.
- orf 4y agoWow https://github.com/geohot/tinygrad/blob/master/tinygrad/tensor.py https://github.com/geohot/tinygrad/blob/master/tinygrad/tens...
- deleted 4y ago[deleted]
- orlp 4y ago> It's extremely simple, and breaks down the most complex networks into 4 OpTypes: > > - UnaryOps operate on one tensor and run elementwise. RELU, LOG, RECIPROCAL, etc... > - BinaryOps operate on two tensors and run elementwise to return one. ADD, MUL, etc... > - ReduceOps operate on one tensor and return a smaller tensor. SUM, MAX > - MovementOps operate on one tensor and move the data around, copy-free with ShapeTracker. RESHAPE, PERMUTE, EXPAND, etc... > > But how...where are your CONVs and MATMULs? Read the code to solve this mystery. Ok, I was curious, so I read the code. The answer is that it represents a MATMUL as a 1x1 CONV. And it lied about CONV, which is a ProcessingOps.CONV and explicitly represented and implemented: https://github.com/geohot/tinygrad/blob/c0050fab8ff0bc667e40da11980f4ac4c21affda/tinygrad/llops/ops_cpu.py#L40 https://github.com/geohot/tinygrad/blob/c0050fab8ff0bc667e40... Quite the letdown of figuring out this 'mystery'.
- WithinReason 4y agoTo directly quote the source: # these are the llops your accelerator must implement, along with toCpu UnaryOps = Enum("UnaryOps", ["NOOP", "NEG", "RELU", "EXP", "LOG", "SIGN", "RECIPROCAL"]) BinaryOps = Enum("BinaryOps", ["ADD", "SUB", "MUL", "DIV", "POW", "CMPEQ"]) ReduceOps = Enum("ReduceOps", ["SUM", "MAX"]) MovementOps = Enum("MovementOps", ["RESHAPE", "PERMUTE", "EXPAND", "FLIP", "STRIDED", "PAD", "SHRINK"]) ProcessingOps = Enum("ProcessingOps", ["CONV"]) https://github.com/geohot/tinygrad/blob/caea34c52996cde2ed464d740ee9b00d50ddccf5/tinygrad/ops.py#L10 https://github.com/geohot/tinygrad/blob/caea34c52996cde2ed46... There is a MAX but not a MIN? Is that because max(x,y) = -min(-x,-y)? But then why is there a SUB? Why is there a RELU if it's only max(0,x)? Maybe MIN is just too rare to be worth implementing?
- georgehotz 4y agoMin is an HLOP. From: https://github.com/geohot/tinygrad/blob/master/tinygrad/tensor.py#L222 https://github.com/geohot/tinygrad/blob/master/tinygrad/tens... def min(self, axis=None, keepdim=False): return -((-self).max(axis=axis, keepdim=keepdim)) All folded together, no slower than MAX.
- eterevsky 4y agoHow is it compared to JAX? After TensorFlow and PyTorch, JAX seems very simple, basically an accelerated numpy with just a few additional useful features like automatic differentiation, vectorization and jit-compilation. In terms of API I don't see how you can go any simpler.
- cl3misch 4y agoHe mentioned in a recent stream that he dislikes the complexity of the XLA instruction set used by JAX. So it's less the user-facing API, and more the inner workings of the library.
- learndeeply 4y agoJAX is a DSL on top of XLA, instead of writing Python. Example: a JAX for loop looks like this: def summ(i, v): return i + v x = jax.lax.fori_loop(0, 100, summ, 5) A for loop in TinyGrad or PyTorch looks like regular Python: x = 5 for i in range(0, 100): x += 1 By the way, PyTorch also has JIT.
- eterevsky 4y agoI've just tried making a loop in a jit-compiled function and it just worked: >>> import jax >>> def a(y): ... x = 0 ... for i in range(5): ... x += y ... return x ... >>> a(5) 25 >>> a_jit = jax.jit(a) >>> a_jit(5) DeviceArray(25, dtype=int32, weak_type=True)
- koningrobot 4y agoIt definitely works, JAX only sees the unrolled loop: x = 0 x += y x += y x += y x += y x += y return x The reason you might need `jax.lax.fori_loop` or some such is if you have a long loop with a complex body. Replicating a complex body many times means you end up with a huge computation graph and slow compilation.
- eterevsky 4y ago
- bArray 4y agoHow does this compare on embedded systems for performance? For example PyTorch vs tinygrad, or Darknet vs tinygrad?
- alexmolas 4y ago> almost 9000 GitHub stars I wouldn't say that 7500 stars is almost 9000 stars ;)
- JacobiX 4y agoI love those tiny DNN frameworks, some examples that I studied in the past (I still use PyTorch for work related projects) : thinc.by the creators of spaCy https://github.com/explosion/thinc https://github.com/explosion/thinc nnabla by Sony https://github.com/sony/nnabla https://github.com/sony/nnabla LibNC by Fabrice Bellard https://bellard.org/libnc/ https://bellard.org/libnc/ Dlib dnn http://dlib.net/ml.html#add_layer http://dlib.net/ml.html#add_layer
- 37ef_ced3 4y agoAnd https://NN-512.com https://NN-512.com
- dedoussis 4y agoIt's funny that geohot/tinygrad chooses to not meet the PEP8 standards [0] just to stay on brand (<1000 lines). Black [1] or any other python autoformatter would probably 2x the lines of code. [0] https://peps.python.org/pep-0008/ https://peps.python.org/pep-0008/ [1] https://github.com/psf/black https://github.com/psf/black
- koningrobot 4y agoMore like 10x. Black is truly a terrible thing.
- kurisufag 4y agoto anybody experienced in writing functional-esque oneliners, PEP8 is an appalling waste of space
- gregjw 4y agoGeohot at it again, this guy nails everything.
- mhh__ 4y agoThis doesn't really nail anything at the moment. It used to nail simplicity but now its a mess IMO
- stephc_int13 4y agoI understand that the Python code is mostly driving faster low-level code, but I wonder how much time is effectively wasted by not using a lower-level language. From my experience with game engines, it often turns out to be a bad idea (for performance and maintainability) to mix C/C++ and Lua or C#.
- terafo 4y agoI would argue that there are performance *benefits* for a developer in running python code, due to how programs are run in python(Jupyter notebooks) you basically can change program on the fly, and not recompile and restart it, as you would do with compiled languages. And yeah, CPU does very very little in modern DL workloads and it is commonplace for CPU python code to be jitted and vectorized, so performance difference isn't as large as you would think.
- lynndotpy 4y agoThis is very true! Another benefit to interactivity is when exploring/using bad code. In academia, you'll often be importing the worst and least-well-documented code you've ever seen. Being able to interactively experiment with someones 500-line 0-documentation function is often a better path to understanding than directly reading the code.
- brrrrrm 4y agoDoesn’t really matter for large batch/large model training on GPUs that don’t need much coordination. But Python speed is one of the main motivations for a JS/TS based ML lib I’m working on: https://github.com/facebookresearch/shumai https://github.com/facebookresearch/shumai
- RektBoy 4y agoNo Bible quotes? I'm disappointed...
- passion__desire 4y agoI can't believe how can someone so accomplished believe in God.
- li4ick 4y agoLike Knuth? He even has a book about it: https://www.goodreads.com/book/show/484459.Things_a_Computer_Scientist_Rarely_Talks_About https://www.goodreads.com/book/show/484459.Things_a_Computer...
- RektBoy 4y agoI come from probably the most atheistic country in the world (CZ) Yet, I had made this bashing comment about bible. IMHO anyone can believe in whatever they want. Christ., Islam, anything. I (and I would say every friend of mine) don't care about what do you believe in, but if you publicly preach some religion, prepare to be made fun of, or take a stand and try defend it your religion with arguments. But no blind faith here. (Personally, if I like some religion it's Shinto.)
- thrtythreeforty 4y agoI think it's orthogonal. There are tons of smart people who believe in God. (Knuth has already been mentioned.) If God wanted, He could make himself apparent to everyone. Clearly that isn't the case; there is room to doubt or to believe no matter how smart or accomplished you are.
- Siira 4y agoThe Bible is a very specific version of a creator that the world already gives you enough evidence to disprove ten times. You could argue that the possibility of the Bible being divine is not zero, but it’s less than any epsilons people care about. The reason it endures are cultural, which is clear from its correlations with geography and demographics.
- brrrrrm 4y ago> It compiles a custom kernel for every operation, allowing extreme shape specialization. This doesn't matter. Just look at the performance achieved by CuDNN kernels (which back PyTorch), they're dynamically shaped and hit near peak. For dense linear algebra at the size of modern neural networks, optimizing for the loop bound condition won't help much. > All tensors are lazy, so it can aggressively fuse operations. This matters. PyTorch teams are trying to implement that now (they have LazyTensor, AITemplate, TorchDynamo), but I'm not sure of the status (it's been tried repeatedly). > The backend is 10x+ simpler, meaning optimizing one kernel makes everything fast. The first part of that sentence matters, the second part doesn't. Kernels are already fast and their reuse outside of being fused into each other (which you need a full linear algebra compiler to do) isn't very high. If you make sum fast, you have not made matrix multiplication fast even though MM has a sum in it. It just isn't that easy to compose operations and still hit 80+% of hardware efficiency. But it is easier to iterate fast and build a seamless lazy compiler if your backend is simple. You can pattern match more easily and ensure you handle edge cases without insanely complicated things like alias analysis (which PyTorch has to do).
- FL33TW00D 4y agoAny more writing on laziness in frameworks? I'm trying to implement it myself.
- brrrrrm 4y agoThe only thing I'd recommend is exposing "eval()" or something to let users tell you when they want you to evaluate things. It'll save a ton of time when it comes to hot-fixing performance and memory use issues. It's really hard to determine when to evaluate, and although it's a fun problem to figure out, it's nice to have an escape hatch for users to just tell you. (Flashlight has explored this and written about it here: https://fl.readthedocs.io/en/latest/debugging.html?highlight=eval#common-issues-and-pitfalls-with-the-arrayfire-backend https://fl.readthedocs.io/en/latest/debugging.html?highlight...) If you're interested, I've looked into symbolic laziness, which allows you to infer correct input sizes even when the constraints happen later. Can be useful for errors. https://dev-discuss.pytorch.org/t/loop-tools-lazy-frontend-experimenting-with-symbolic-laziness/372 https://dev-discuss.pytorch.org/t/loop-tools-lazy-frontend-e...
- neets 4y ago4 OpCodes, I think Geohot is taking a cue from his favorite intellectual's Curtis Yavin's Urbit project
- lr1970 4y agoAs it was recently discussed at length here on HN [0] (401 comments), George Hotz (the lead of tinygrad) is taking time off his self-driving startup comma.ai [1]. Curious if this would help or hurt tinygrad progress. [0] https://news.ycombinator.com/item?id=33406790 https://news.ycombinator.com/item?id=33406790 [1] https://comma.ai/ https://comma.ai/
- tucosan 4y agoCan someone from the ML crowd ELI5 to me what tinygrad does, how it plugs into an ML pipeline and what it's use cases are?
- gamegoblin 4y agoThere are libraries like tensorflow and PyTorch that allow the user to define their neural net in simple, readable Python code, and they internally "compile" and optimize your neural net to run on GPUs and such. Tinygrad is like a very, very lean PyTorch with a different philosophy -- it intends to keep the codebase and API surface very very small and focus most of its energy on optimizing the way the output neural net runs on physical hardware. The author, George Hotz, has observed in the last few years that neural net performance is hindered by lack of optimization here, particularly around memory accesses.
- lostmsu 4y agoIt was ok as an educational tool, but now they don't count GPU implementation in 1000 lines, so it is not small. Considering the code style it is closer to 20k+ lines when formatted and GPU code included. It also doesn't support bfloat16 so is doomed to be 2x slower.
- terafo 4y agoActual code of tinygrad is less than 5k lines. There is also 1600 lines of tests and around 2k lines of example models. And I didn't count unfinished support for geohot's own unfinished neural network accelerator(verilog for that accelerator sits in repo too), which is abandoned.
- lostmsu 4y ago> Actual code of tinygrad is less than 5k lines Yeah, not > Considering the code style I mean it is possible to read it, but I would not say it is optimized for it. Which I suppose betrays the goal.
- terafo 4y ago> Yeah, not Provide some evidence. I just ran tokei on freshly cloned tinygrad repo using arguments from[1]. Got 4854 lines of code, which is less than 5k lines. [1] tokei --exclude *.json --exclude accel/cherry --exclude test --exclude examples
- lostmsu 4y agoAfter running black it is more like 6.8k. And black does not format C and/or shader code. But even 5k is closer to 20k on the log scale than to the promised 1k.
- bullen 4y agoDoes anyone know of a neural network that is written in C and GLSL and that runs on normal OpenGL?
- kwant_kiddo 4y agoI think posts like this are only getting upvotes because George Hotz owns the project. I do see value in simple code, but the constraint of 1000 LOC makes little sense to me, especially when the code is formatted poorly. This will get downvoted, but reading the comments here I dont understand the (cult/respect) for him. Siding with the most successful CTF-team ever (PPP) he won defcon two times. He made a startup with funding that makes a cool 'niche' product. I just think a guy like Chris Lattner or Dave Cutler who made so much impact on real computing deserve so much more respect, but I guess that the norm here is to admire this guy.
- gamegoblin 4y agoYour list lacks the reason for his initial fame: iPhone and PS3 jailbreaking. And I think you're downplaying the achievements of Comma AI -- it may still be somewhat niche, but its product is better than Tesla Autopilot for highway driving (they aren't there on city driving / FSD yet), all with an absolutely tiny team.
- lostmsu 4y agoRe: Comma AI. This is what it tells me about my run-of-the-mill Toyota: > openpilot upgrades your Toyota Highlander Hybrid with automated lane centering at all speeds, and adaptive cruise control that automatically resumes from a stop. Both are annoying artificial limitations Toyota put presumably to avoid abuse by inattentive drivers. I mean it can't change lanes. What does it do exactly?
- gamegoblin 4y agoComma AI deliberately made lane change require a small bit of human intervention for safety reasons. The human hits the blinker and gives the wheel a tiny nudge in the direction, and then openpilot will complete the lane change and resume driving in the new lane. The theory is that at the current ability of software like Tesla and Comma has, it's probably a good idea for a human to be paying more attention during a lane change maneuver. Comma is of the opinion that the level of autonomy Teslas have is probably unnecessarily unsafe. Comma cares a lot about safety (e.g. they have much more sophisticated driver monitoring than Tesla). Lane change here: https://www.youtube.com/shorts/xm8DRwvLObQ https://www.youtube.com/shorts/xm8DRwvLObQ Since openpilot is open source software, there are of course forks that exist that remove these safety limitations and will lane change automatically.
- bfrankline 4y agoIf you care exclusively about numerical stability and performance, why _this_ set of operators (e.g., there’re plenty of good reasons to include expm1 or log1p and certainly trigonometric functions)? It’d be an interesting research problem to measure and identify the minimal subset of operators (and I suspect it’d look differently than what you’d expect from an FPU). If you care exclusively about minimalism, why not limit yourself to the Meijer-G function (or some other general-purpose alternative)?
- therealchiggs 4y agoThere's an interesting roadmap in the "cherry" folder of the git repo[0]. It begins by bringing up a design on FPGA and ends with selling the company for $1B+ by building accelerator cards to compete with NVIDIA: Cherry Three (5nm tapeout) ===== * Support DMA over PCI-E 4.0. 32 GB/s * 16 cores * 8M elements in on board RAM of each core (288 MB SRAM on chip) * Shared ~16GB GDDR6 between cores. Something like 512 GB/s * 16x 32x32x32 matmul = 32768 mults * 1 PFLOP @ 1 ghz (finally, a petaflop chip) * Target 300W, power savings from process shrink * This card should be on par with a DGX A100 and sell for $2000 * At this point, we have won. * The core Verilog is open source, all the ASIC speed tricks are not. * Cherry will dominate the market for years to come, and will be in every cloud. * Sell the company for $1B+ to anyone but NVIDIA [0] https://github.com/geohot/tinygrad/blob/master/accel/cherry/README#L22-L68 https://github.com/geohot/tinygrad/blob/master/accel/cherry/...
- deleted 4y ago[deleted]