21 ms·
The tiny corp raised $5.1M
- freediver 3y agoTrue entrepreneur. Having a vision and ignoring naysayers. Go George!
- nwienert 3y agoGotta have money to start to keep taking risks
- SXX 3y agoWhile it's not available to many of us (and especially those outside of US) it doesn't take that much money to just tinker with whatever you want. Just have to cut time wasted on news and politics, socialising and literally anything outsidie of being hacker or entrepreneur. Most people just would never take that risk and will stick to their well-paying job.
- nwienert 3y agoWas it pg who said it? Not sure but basically - middle class get about one shot. Upper class can keep shooting their whole lives. Of course lower class usually doesn’t get one. I took my one shot! Any more is too risky for quite a while.
- coffeebeqn 3y agoThe way they make money is for AMD to buy them out if they’re successful
- vruiz 3y agoWhich would make total sense for AMD if they pull it off.
- aeyes 3y agoWhy would they buy something that is open source? They could acqui-hire but Hotz doesn't strike me as a person that would stay at a big corp like AMD for a significant amount of time.
- vruiz 3y agoTo control it. If they are successful at making AMD cards competitive in AI, and I agree that what's missing it's only the software, that would create immense value for AMD. Too much to not have control over how it evolves. If they are successful it will not just be Hotz, he will hire other devs and an entire community will form around it. Sure they could just fork it and try to continue development themselves, but the community and momentum might very well not go with them.
- moralestapia 3y agoIt wouldn't surprise if those 5M came from AMD
- btown 3y agoIf/when tinygrad is successful then AMD acquiring control of the stewardship/direction-setting of the software that drives the incremental demand for their hardware is far more valuable than Hotz's talent itself.
- neom 3y agoSeems like a much better mission for George Hotz to go on than single handedly trying to fix Twitter.
- vippy 3y agoGeorge Hotz is a talented engineer but he absolutely does not have the social science background to "fix" Twitter.
- threeseed 3y agoBeing an engineer at Twitter is much more akin to an enterprise like a bank or telco. You need to be a team player and work across different groups e.g. Product, Testing, SRE etc in order to successfully get features into Production. So being a talented engineer is useful but having high emotional intelligence and being able to negotiate and collaborate is far more important.
- pankajdoharey 3y agoYou mean talk more and code less?
- FrustratedMonky 3y agoI'm surprised more people don't realize this. It happens to all big corp.
- confoundcofound 3y agoYou don’t need a social science background to fix Twitter.
- patentatt 3y agoIf they can achieve something competitive with CUDA for $5m, why hasn't AMD done it yet?
- nacs 3y agoAMD/ATI has always been terrible at drivers/software for their GPUs going back decades. It’s really as simple as that and it still hasn’t changed so nvidia is dominating them in AI as a result.
- 1letterunixname 3y agoTheir stuff was garbage back when there was VLB.
- crpowers 3y agoAMD has been tackling exactly the wrong problems. They poured their money into a porting solution for developers to take CUDA code and run it on their GPUs. I guess they didn't find it worth it to really compete. I doubt it's about being able to tackle this problem with $5m, but rather convincing the company they can win.
- 1letterunixname 3y agoI still don't understand what problem they're trying to solve in EPYC in the hypervisor space with encryption. They should've been adding tensor cores and neural acceleration to their CPUs. The need for headed graphics cards is moot and wasteful. NVIDIA solved this with the A100. NVIDIA may spin into a mainstream enterprise CPU and systems vendor as a sales channel for converged CPU-GPU solutions beyond what they're already doing.
- SXX 3y agoIf you talking of AMD SEV it's actually a useful technology. Confidential virtual machines not only protects you from possible spying on AWS or Azure, but also make it possible to have some decentralized / P2P compute more feasible. Of course nothing is perfect and you can never have 100% trust to someone else hardware, but it's defenetely step in right direction.
- reaperman 3y agoLove this overall. Wonderful. But I wouldn't say Cerebras failed just yet -- they're committing to OpenXLA which may provide a better dev-experience in the long run than Nvidia lock-in.
- tzhenghao 3y ago> I think the only way to start an AI chip company is to start with the software. The computing in ML is not general purpose computing. 95% of models in use today (including LLMs and image generation) have all their compute and memory accesses statically computable. Agree with this insight. One thing Nvidia got right was a focus on software. They introduced CUDA [1] back in 2007 when the full set of use cases for it didn't seem very obvious. Then their GPUs had Tensor cores, and more complementary software like TensorRT to take full advantage of them post deep learning boom. Right as Nvidia reported insane earnings beat too [2]. Would love more players in this space for sure. [1] - https://en.wikipedia.org/wiki/CUDA https://en.wikipedia.org/wiki/CUDA [2] - https://www.cnbc.com/2023/05/24/nvidia-nvda-earnings-report-q1-2024.html https://www.cnbc.com/2023/05/24/nvidia-nvda-earnings-report-...
- foobiekr 3y agoI know the founders of an AI chip company that taped out and got working chips on their first go. They got their chip done, it’s pretty solid. Chip has great perf and is super power efficient, a solid delivery. I knew they'd nail it and they did. The SW story is a train wreck, though. The problem basically was that they couldn’t hire any good SW people. As I said I know the founders. They are both genuinely decent guys, they put their own money in so they have some (well, minimal) skin in the game, and they know a ton of expert-level embedded and systems coders with between 20 and 40 years of hard core experience. As far as I can tell, they weren't really able to get anyone that we know in common to join. I certainly did not, and no one I know did either. Last I heard they'd had to hire third choice guys in Europe to do the work and it wasn't going well. There's a pretty good reason for it, and it comes down to a sociological problem. HW people don’t value SW people. It's just basically true and has been true everywhere I've looked. Maybe if you're doing a system (like a router or maybe a drone) then the HW people will begrudgingly admit that the SW is a major part of the delivery, but that isn't true for chip companies (including chips-on-reference-boards). You can rest assured that at a chip company, all of the high comp people in the company are going to be on the ASIC team and the SW team will never be on the same tier. The argument is always the same, no matter how many times it bites the companies on the ass and sends them careening into the dumpster: “yes, but the chip without SW is the chip! we can buy SW, if we have to. SW without the chip has zero value.” Almost every chip company ends up like that, and the kind of low level, experienced SW people that work in the space know to avoid them and work at systems companies instead. As far as I've been able to determine, with _maybe_ the exception of Cerebras - maybe - this is the situation that has played out at all of the 201x AI chip companies. They get founded by ASIC guys, most of whom have more than a small chip on their shoulder about the relative value of ASICs-vs-SW. These guys are all ex-SGI, ex-Sun, ex-Google, ex-Nvidia, ex-Intel HW guys who saw SW people making a lot more, not just in broader industry terms over the last few years, but at hardware-focused companies. In general, ASIC guys make less than SW guys unless they are the very narrow set of top level architects. IMHO from a value creation standpoint, that is _super unfair_ and I am not here to justify it, but it is how it is. The result poisons ASIC companies. SW people who know what needs to be don't won't go to them most of the time, for good reason, and so they fail. So I will say, given that, starting with SW first is brilliant.
- Mistletoe 3y agoCan anyone comment on the TinyBox they are taking preorders for? The tinybox 738 FP16 TFLOPS 144 GB GPU RAM 5.76 TB/s RAM bandwidth 30 GB/s model load bandwidth (big llama loads in around 4 seconds) AMD EPYC CPU 1600W (one 120V outlet) Runs 65B FP16 LLaMA out of the box (using tinygrad, subject to software development risks) $15,000
- 1letterunixname 3y agoThere's a reason no one uses ATI GPUs in datacenters. Their dev support is shit. Don't waste your money. Buy 6 RTX 4090's and a decent ECC-memory server, and call it a day.
- wmf 3y agoDidn't read the article.
- guraf 3y agoI thought you weren't allowed to use Nvidia's consumer GPUs in the datacenter?
- evanchisholm 3y agoYou aren't, but who's going to stop you?
- interlinked 3y agoTheir closed source driver?
- ilyt 3y agoAnd how exactly they are going to know you're running card in DC ?
- wmf 3y agoWe can't really comment much on it because a bunch of specs are lacking. Is it using 6x 7900 XTXs? Which Epyc CPU (Epycs vary in price from $1K to $11K)?
- nightowl_games 3y agoI wonder what Tenstorrent is doing. Didn't realize they are Canadian. https://tenstorrent.com/research/tenstorrent-raises-over-200-million-at-1-billion-valuation-to-create-programmable-high-performance-ai-computers/ https://tenstorrent.com/research/tenstorrent-raises-over-200...
- Havoc 3y agoGlad he is going ahead with this. Will make for many entertaining live streams no doubt
- j0hnyl 3y agowhat does he mean by "tape out" chips?
- bigdict 3y agohttps://en.m.wikipedia.org/wiki/Tape-out https://en.m.wikipedia.org/wiki/Tape-out
- detaro 3y agoIndustry term for a chip design being sent to manufacturing/being produced.
- bloggie 3y agoWhen transforming the logic gate design of a chip into the lithographic plates used for chip production, the plates were originally made by applying tape to create the photo masks. The name stuck and it now means to move a semiconductor project from the design phase to the manufacturing phase.
- bigdict 3y agoThe end goal is getting acquired by AMD of course.
- thatguyknows 3y agoI was always surprised at how AMD hasn't already thrown a bunch of money at this problem. Maybe they have and are just incompetent in this area. My prediction is AMD is already working on this internally, except more oriented around PyTorch not Hotz's Tinygrad, which I doubt will get much traction.
- tzhenghao 3y agoI think AMD is going down a different path, ie. ROCm then partnering with ML frameworks further up the stack for first class support. https://pytorch.org/blog/pytorch-for-amd-rocm-platform-now-available-as-python-package https://pytorch.org/blog/pytorch-for-amd-rocm-platform-now-a...
- lbhdc 3y agoHe mentioned ROCm, and apparently had lack luster experience with it. >The software is called ROCm, it’s open source, and supposedly it works with PyTorch. Though I’ve tried 3 times in the last couple years to build it, and every time it didn’t build out of the box, I struggled to fix it, got it built, and it either segfaulted or returned the wrong answer. In comparison, I have probably built CUDA PyTorch 10 times and never had a single issue.
- tzhenghao 3y agoNot surprising lol. This was also the experience I had while experimenting with MLIR approximately 3 years ago. You'd need to git checkout a very specific commit and then even change some flags in code to have a successful build. I'm sure things are better now but I haven't messed with it since then.
- radq 3y ago> I'm sure things are better now but I haven't messed with it since then. I had the same experience ~3 months ago. Gave up and switched to Nvidia 3090s for my workloads.
- 3y ago
- camdenlock 3y agoI’d love to see this succeed, because I own a 7900 XTX already, but there’s so much already built on top of PyTorch. Why would anyone port it all to tinygrad or whatever? Bummer.
- pizza 3y agoIn theory pytorch compiler can boil down to 50 or so fundamental functions and tinygrad IR to 12. So possibly you could just re-map a fairly limited set of base instructions. Devil’s in the details though..
- currymj 3y agopeople already ported a lot of stuff from pytorch to jax. if you're a research scientist or grad student, to a certain extent a lot of projects are "greenfield" so it's easy to jump on a new framework if it is nice to use and offers some advantage.
- quickthrower2 3y ago> I started tinygrad in Oct 2020. It started as a toy project to teach me about neural networks Shows you what is possible in 2.5 years. Keeps me motivated to learn.
- vippy 3y agoThe math isn't super difficult. Some books will try to throw a mess of differential equations at you, but some simple calculus is all you need for backpropagation.
- quickthrower2 3y agoI have been through the math thanks to the youtube videos by A. Karpathy. Deriving some of the differentials, e.g. for batchnorm seems fairly hard (hard as in slogging through something with many steps where you can't make a mistake at any step). But the principles are quite simple - I think by design. If they were hard to compute or reason about then the neural net wouldn't work very well!
- sva_ 3y agoDoing the compute efficiently, especially from Python, is the tricky part.
- whalesalad 3y agoDidn't realize the 7900 XTX was so good at this kind of work. Glad to have went with it over the 4080.
- jitl 3y agoThe whole point is that it’s not good at this work — and it’s a $Xx billion opportunity to make it work.
- csense 3y agoI respect Geohot's reputation and this company looks amazing. I might be in the market to work there... except "No Remote." For such a smart guy, locking yourself out of a ton of talent by requiring software developers to be on-site in 2023 seems...out of character, to put it politely. (Rephrased, my original post was a bit too ad hominem and accumulating downvotes rapidly. I wanted to delete this entire comment but apparently HN no longer allows comments to be deleted.)
- throwawayadvsec 3y agowhat about IP theft though?
- detaro 3y agowhat about it?
- duxup 3y agoHow big is Tiny Corp? I doubt they need mass volumes of employees at this stage and they maybe want to work closely with the people they choose?
- ftxbro 3y ago> For such a smart guy, locking yourself out of a ton of talent by requiring software developers to be on-site in 2023 seems...out of character, to put it politely. I mean a lot of smart people seem to do their hacking by themselves. I'm thinking like Fabrice Bellard. This is at least a step beyond that.
- paxys 3y agoSoftware is part of it, sure, but I doubt anyone can realistically work on this project/company without being around a bunch of specialized hardware and iterating on prototypes. Hard to contribute to any of that from home.
- foobiekr 3y agoA lot of people have now had the direct experience that new things which are highly technical or highly collaborative are not really compatible with the remote work thing. I know that's hard for a lot of people to hear, but the world is not web apps (which do remote well) and a lot of projects benefit hugely from being able to grab the two or three people and get into a room with a whiteboard.
- mhh__ 3y agoGood idea. I don't think George Hotz has the skill set to actually deliver on a lot of this stuff (specifically I suspect trying to replace the compiler for the GPU is something that he will probably make a song and dance about with some simple prototype but then quietly scrap it because even for AI workloads its still a very very tricky problem) but he has the strength of vision to get and direct other people to do it for him.
- wahnfrieden 3y agodo you think those talented workers will accept giving him ownership over that work and its value?
- mhh__ 3y agoTraditionally the way startups entice people is by giving them equity
- wahnfrieden 3y agowow ok
- dharma1 3y agoSo much untapped potential in AMD, and funny they keep failing at the software and geohot has to save them
- turnsout 3y agoI clicked through to the previous blog post, to read more about the unit of a "person" of compute [0]. Definitely worth a read, if only for this quote: > One Humanity is 20,000 Tampas. I'll never think of humanity the same way! [0]: https://geohot.github.io/blog/jekyll/update/2023/04/26/a-person-of-compute.html https://geohot.github.io/blog/jekyll/update/2023/04/26/a-per...
- kevmo314 3y agoNvidia or AMD, the real winner here is truly TSMC.
- vader777 3y ago[dead]
- godelski 3y ago> There’s a [Radeon RX 7900 XTX 24GB] already on the market. For $999, you get a 123 TFLOP card with 24 GB of 960 GB/s RAM. This is the best FLOPS per dollar today, and yet…nobody in ML uses it. > I promise it’s better than the chip you taped out! It has 58B transistors on TSMC N5, and it’s like the 20th generation chip made by the company, 3rd in this series. Why are you so arrogant that you think you can make a better chip? And then, if no one uses this one, why would they use yours? > So why does no one use it? The software is terrible! > Forget all that software. The RDNA3 Instruction Set is well documented. The hardware is great. We are going to write our own software. So why not just fix AMD accelerators in pytorch? Both ROCm and pytorch are open sourced. Isn't the point of the OSS community to use the community to solve problems? Shouldn't this be the killer advantage over CUDA? Making a new library doesn't democratize access to the 123 (fp16-)TFLOP accelerator. You fix pytorch and suddenly all the existing code has access to these accelerators. Millions of people now have This then puts significant pressure on Nvidia, as they can't corner the DL market. But it is a catch-22 because the DL market already is mostly Nvidia so it takes priority. Isn't this EXACTLY where OSS is supposed to help? I get Hotz wants to make money, and there's nothing wrong with that (it also complements his other company), but the arguments here seem more for fixing ROCm and specifically the pytorch implementation. The mission is great, but AMD is in a much better position to compete with AMD. They caught up in the gamer's market (mostly) but have a long way to go for scientific work (which is what Nvidia is shifting focus to). This is realistically the only way to drive GPU prices down. Intel tried their hand (including in supercomputers) but failed too. I have to think there's a reason that's not obvious to most of us as to why this is happening. Note 1: I will add that supercomputers like Frontier (current #1) do use AMDs and a lot of the hope has been that this will fund the optimization from two places: 1) DOE optimizing their own code because that's the machine that they have access to and 2) AMD using the contract money to hire more devs. But this doesn't seem to be happening fast enough (I know some grad students working on ROCm). Note 2: There's a clear difference in how AMD and Nvidia measure TFLOPS. techpowerup shows AMD at 2-3x Nvidia, but performance is similar. Either AMD is crazy underutilized or something is wrong. Does anyone know the answer?
- wmf 3y agoIt's often less work to start from scratch than to fix an extremely complex broken stack. Of course people also say this when they just want to start from scratch. RDNA 3 has dual-issue that basically isn't used by the compiler so half the FPUs are idle.
- deleted 3y ago[deleted]
- sheepscreek 3y agoThis is great news. I’ve oft wondered the same about AMD’s GPUs. NVIDIA’s got a clear monopoly. He made a very good point about how this isn’t general purpose computing. The tensors and the layers are static. There’s an opportunity for a new type of optimization at the hardware level. I don’t know much about Google’s TPUs, except that they use a fraction of the power used by a GPU. For this experiment though, my sincere hope is that all the bugs are software only. Supporting argument - if they were hardware bugs, the buggy instructions would not have worked during gameplay.
- nalzok 3y ago> The main advantage is in the tinygrad IR. It has 12 operations, all of which are ADD/MUL only. `x[3]` is supported, `x[y]` is not. Can someone educate me why that is the case? Does `x[y]` require a Turing-complete kernel to compute?
- brrrrrm 3y agolayer of indirection introduces scatter/gather and other dynamic loads, which is tricky to optimize
- impulser_ 3y agoWhy wouldn't AMD throw a few million at this? Worst case they lose a small amount of money, but best case they finally get good software for their hardware. The past decade or so, they haven't been able to create any good software for their hardware. They made small improvements but the competition, Nvidia, has also made improvements to their already good software. It too the point where their software is the reason why most people/companies don't use their products. Their drivers for their customer products are just as bad. They are very competitive in hardware, but Nvidia dominates them at software which make companies buy Nvidia. No one wants to deal with the pain of AMD software. AMD is a better company to work with than Nvidia, but it not worth it when it comes to dealing with their software lol.
- tormeh 3y agoAFAIK the "AMDs drivers are bad" meme is outdated. Sure, their AI/ML software is garbage, but the graphics drivers are fine.
- aiappreciator 3y agoThe cutting edge for graphics is all raytracing, and Nvidia still dominates. DLSS3, pathtracing, etc, these are for 'graphics', but heavily dependent on AI post processing, so Nvidia still rules. So in the gaming market, Nvidia still commands a huge premium. No AMD card can play Cyberpunk on overdrive 4k.
- dikei 3y ago> The cutting edge for graphics is all raytracing, and Nvidia still dominates. DLSS3, pathtracing, etc, these are for 'graphics', but heavily dependent on AI post processing, so Nvidia still rules. IMHO, unless you have a $1000+ GPU, RayTracing is still not worth the performance hit. I prefer playing at 100+ FPS with RayTracing off, than turning it on and have my frame rate cut in half.
- tgsovlerkhgsel 3y agoI just had my AMD-based machine crash - repeatedly - every time I tried to use Google Maps inside Firefox for longer than a few minutes. I haven't confirmed, but I strongly assume that either their graphics drivers or something Ubuntu does with Wayland are not fine.
- daveed 3y agoI don't want to cast any judgement, I just want to ask what the initial product is. The claim is they sell computers, and there's a link to the tinybox. There's a $100 preorder, for a 15k computer (I guess I'd have to pay 14.9k eventually?). And then we get a computer that... how do I interact with it? Will it have its own OS? Some flavor of linux? Is the intent to work on it directly, or use it as an inference server, and talk over a network?
- digitallyfree 3y agoI think the tinybox is meant to be a training/inference server meant for tinygrad and filled with those AMD cards. Very likely it will run Linux.
- asdfman123 3y agoIs this the guy who couldn’t add a feature to Twitter?
- thesausageking 3y agoFor background, "geohot", is George Hotz, who's a known hacker / tech personality[0] This project fits the pattern of his previous projects: he gets excited about the currently hot thing in tech, makes his own knockoff version, generates a ton of buzz in the tech press for it, and then it fizzles out because he doesn't have the resources or attention span to actually make something at that scale. In 2016, Tesla and self-driving cars led to his comma one project ("I could build a better vision system than Tesla autopilot in 3 months"). In 2020, Ethereum got hot and so he created "cheapETH". In 2022 it was Elon's Twitter, which led him to "fixing Twitter search". And in 2023 it's NVIDIA. I'd love to see an alternative to CUDA / NVIDIA so I hope this one breaks the pattern, but I'd be very, very careful before giving him a deposit. [0] https://en.wikipedia.org/wiki/George_Hotz https://en.wikipedia.org/wiki/George_Hotz
- normaldist 3y agoComma didn't fizzle out. https://twitter.com/comma_ai/status/1578517666632900608 https://twitter.com/comma_ai/status/1578517666632900608
- thesausageking 3y agoThey're slowing burning through their VC money trying to make a business out of the hobbyist market while Cruise and Waymo have fully autonomous cars deployed in SF and are scaling up.
- Retric 3y agoThe scale of investment is wildly different. $5.57M in actual revenue vs $18.1M raised isn’t that far from a viable product and positive ROI for their investors. Cruse and Waymo have invested billions, they need 10’s of billions in annual sales or their project is a failure.
- birken 3y agoAs a very happy user of Comma, I think it is reasonable to say the company is going to fail, but that ignores that the product they created is still awesome. Comma is light years better than any built-in driving assist in any non-Tesla car. And it's comparable to Tesla for far less money. The reason the company might fail is because their main thesis, being that car manufacturers would just license the self driving tech to somebody else (like Comma), never came about. Car manufacturers are just too conservative. It was a perfectly reasonable bet to make though. Unfortunately they ended up in the business of selling hardware and giving away software for free when they wanted to be in the business of selling software.
- agnosticmantis 3y ago> The human brain has about 20 PFLOPS of compute. Where is this number coming from? The number of spikes per second? Edit: doing a quick search, it doesn’t seem like there’s a consensus on the order of magnitude of this. Here’s a summary of various estimates: https://aiimpacts.org/brain-performance-in-flops/ https://aiimpacts.org/brain-performance-in-flops/
- reaperman 3y agoNo idea. I don't think there's any scientific consensus on even an upper limit of a human brain's FLOP equivalence.
- re-thc 3y ago> Where is this number coming from? Used 20 PFLOPS of compute to simulate it.
- agnosticmantis 3y ago[flagged]
- chubs 3y agoHe claims there's a $999 AMD card that gives 123 TFLOPS, and his tinybox will cost $15k for 738 TFLOPS. In other words, the tinybox will have 6 of these GPUs, eg $6000 cost price. It seems a steep markup from 6k to 15k, and if the software is open-source, i'm not sure why you wouldn't build your own? Or is it worth 9k for a custom motherboard that can fit so many GPUs. Or can you buy 6-GPU motherboards off the shelf? Just curious what people think. Not being disparaging, kudos to geohot :)
- mattnewton 3y agoAs someone who has built their own deep learning rigs in the $10k BoM range, frankly it's a pain in the ass and I would gladly pay that in the future. I probably will pay lambdalabs a much larger markup.
- sbrother 3y agoAs someone who bought a deep learning rig for about $12k from lambdalabs years ago, I can't recommend them strongly enough. The support (and not having to deal with building it out myself) was well worth the markup. They're also just really great to deal with.
- zdw 3y agoLogic boards and CPUs with the necessary PCIe lanes capable of feeding the GPUs, storage fast enough to do the 30GB/s that is in the spec sheet, power supplies and a case to hold such a power hungry system, etc. is not cheap. $15k seems like a kind of low - if you've tried to spec out a server with similar quantities of accelerators recently, you'd have trouble hitting that figure, even if using consumer grade GPUs.
- AlotOfReading 3y agoYou can buy 6+ gpu motherboards, but the consumer ones are solely for mining because no normal person has that many cards and no consumer CPU exposes enough lanes to properly support that many GPUs. A lot of enterprise vendors exist to sell you that kind of system in the enterprise space, but you should expect to still be paying 5 figures minimum. $15k seems like the low end to me.
- SkyMarshal 3y ago> I think the only way to start an AI chip company is to start with the software. The computing in ML is not general purpose computing. 95% of models in use today (including LLMs and image generation) have all their compute and memory accesses statically computable. > Unfortunately, this advantage is thrown away the minute you have something like CUDA in your stack. Once you are calling in to Turing complete kernels, you can no longer reason about their behavior. You fall back to caching, warp scheduling, and branch prediction. > tinygrad is a simple framework with a PyTorch like frontend that will take you all the way to the hardware, without allowing terrible Turing completeness to creep in. I like his thinking here, constraining the software to something less than Turing complete so as to minimize complexity and maximize performance. I hope this approach succeeds as he anticipates.
- nonethewiser 3y agoCan anyone elaborate on how or why Turing completeness requires these sub optimal patterns? I recall reading about avoiding Turing completeness for similar reasons to avoid the halting problem. > Other times these programmers apply the rule of least power—they deliberately use a computer language that is not quite fully Turing-complete. Frequently, these are languages that guarantee all subroutines finish, such as Coq. https://en.m.wikipedia.org/wiki/Halting_problem https://en.m.wikipedia.org/wiki/Halting_problem
- mlazos 3y agoIt isn’t that there are suboptimal patterns, it’s just the more expressive your language can be at runtime, the less you can reason about statically. An example is data-dependent control flow. If you can’t reason about what branch your code is going to take statically (without your runtime data) it is harder to generate fast code for it.
- mlazos 3y agoThis 95% of models are statically computable thing really shows how much he is trivializing this problem. I’d be interested to see his SW stack compile MaskRCNN. His ISA is massively under-defined and people will not change their model code to run on this accelerator unless his performance beats cuda significantly and even then they still won’t - usability matters more than performance every time. In the end you need a compiler, and it needs to be compatible with an existing framework which is not trivial at all, since they are written in python.
- Spikeysash 3y ago[dead]
- ipsum2 3y agoI have some experience in this area, having both worked on machine learning frameworks, trained large models on datacenters, and have my own personal machine for tinkering around with. This makes very little sense. Even if he was able to achieve his goals, consumer GPU hardware is bounded by network and memory, so it's a bad target to optimize. Fast device-to-device communication is only available on datacenter GPUs, and is essential for models training like LLaMA, Stable Diffusion, etc. Amdahl's law strikes again.
- sequoia-capital 3y ago[dead]
- Lapsa 3y agoI find it peculiar. just recently folks were tryharding to make everything and your kitchen sink Turing complete and now it creeps in menacingly on its own
- Lapsa 3y agoalso, I like the spirit of: "I don’t want to live in a world of closed AI running in a cloud you’ve never seen". shake those shills a bit
- smasher164 3y agoTuring-completeness != un-optimizable! Literally the areas of type systems and compilers exist to serve this endeavor. It's gotta be a meme at this point every time someone brings up the halting problem or rice's theorem.
- bmacho 3y agoCS education considered harmful!
- redox99 3y agoI don't think anybody claims it's unoptimizable. It's just a harder task, compared to a more constrained system.
- smasher164 3y agoRight, but the type system is the constraint. Nobody's gonna take the untyped lambda calculus and run it on an accelerator. Even something like turing-completeness can be a type annotation, for example, the totality effect provided by languages like Koka.
- fancyfredbot 3y agoAMD only enabled their ROCm stack on consumer cards last month. This finally corrects a huge mistake - Nvidia made cuda available on all their cards for free from the start and made it easy/cheap for people to get started. Of course once they'd started they stuck with it... I hope it's not too late to turn this around.
- russ 3y agoHope George pulls this off. Just sent this to my dad, who turned down starting ATI with the Lau family on account of me being born. =)
- deleted 3y ago[deleted]
- willzhang100 3y agoWhere does the figure that the human brain is 20 pflops come from?
- neilv 3y agoI don't know whether it's a factor in the alleged software quality issues he mentions, but it's not unusual for a company that thinks of itself as a hardware company to neither understand nor respect software enough. Even if adopting "hardware/software co-design", leadership might be hardware people, and they might not understand that there's tons more to systems software engineering than the class they had in school or the Web&app development that their 5 year-old can do. That misunderstanding can exhibit in product concepts, resource allocation, scheduling, decisions on technical arguments, etc. (Granted, the stereotypical software techbro image in popular culture probably doesn't help the respect situation.)
- sidcool 3y agoGeorge was in the list of my favourite programmers till a few months ago. Now there are Jeff Dean, John Carmack and Karpathy
- Tepix 3y agoWhen you click on the strip link to preorder the tinybox, it is advertised as a box running LLaMA 65B FP16 for $15000. To be fair, the previous page has a bit more details on the hardware. I can run LLaMA 65B GPTQ4b on my $2300 PC (built from used parts, 128GB RAM, Dual RTX 3090 @ PCIe 4.0x8 + NVLink), and according to the GPTQ paper(§) the quality of the model will not suffer much at all by the quantization. Just saying, open source is squeezing an amazing amount of LLM goodness out of commodity hardware. (§) https://arxiv.org/abs/2210.17323 https://arxiv.org/abs/2210.17323
- coolspot 3y agoWhat case/MB/GPUs do you use for your dual 3090 build? Liquid cooled cards?
- Tepix 3y agoI'm using air cooling. Gigabyte X570 Pro, RTX 3090 FE, be quiet! Pure Base 500DX mesh case with four 140mm fans currently. It's not quiet under heavy load! The GPUs have gotten improved thermal pads. The 8 core Ryzen 3700X is a 65W model. It appears to be fast enough not to be a bottleneck for this purpose. 1200W PSU. I may swap the fan in front of the GPUs for a high-rpm model. Also for longer runs i throttle the GPU power draw. It doesn't cost much performance.
- magixx 3y agoAre you able to memory pool two 3090s for 48gb and if so what's your setup? I looked into this previously[1] but wasn't super confident it's possible or what hardware is required (2x x8 pcie and official SLI support?). AFAICT still would look like two GPUs to the system. [1] https://discuss.pytorch.org/t/is-there-will-have-total-48g-memory-if-i-use-nvlink-to-connect-two-3090/100378/10 https://discuss.pytorch.org/t/is-there-will-have-total-48g-m...
- Tepix 3y agoYou can memory pool with the right software however pytorch supports spreading large models over multiple GPUs OOTB. Just pass the --gpu-memory parameter with two values (one per GPU) to oobabooga's text-generation-webui for example.
- paulus-saulus 3y agoI reached the limit of free articles. impossible to browse and read the page
- bgitarts 3y agoWhy isn't AMD and Intel working on an alternative?
- wahnfrieden 3y agogeohot gave me some good advice/feedback from using my Japanese reading app, Manabi Reader: https://reader.manabi.io https://reader.manabi.io I'm always grateful for his user feedback, it led to next-level improvements. Thanks geohot