11 ms·
Advice to Tenstorrent
- bigyabai 1y agoWith all due respect Mr. Geohot, you've got an awful lotta nerve to throw stones from your glass house. This guy needs to drop the middle school tough-guy tone and post numbers. The central thesis of this article is sound, lead with it: > You aren't going to get better deals on tapeouts/IP than NVIDIA/AMD. You need some advantage. > If you want a dataflow graph compiler, build a dataflow graph compiler. Now explain why. Clearly Tenstorrent is happy to build Yet Another Abstraction Layer, so instead of bullying them over it you should at least attempt to actively humiliate them for the approach. You know, produce some manner of evidence that vindicates your position instead of relying on your authority alone. Jim Keller has no reason to take this seriously, even if you're right. Without any numbers this feels like one cult of personality trying to bait another into a shit-flinging contest as a marketing scheme. We've seen this happen several times before on Hacker News, and it doesn't end up with either side making an Nvidia-killer. This is not a model for productive discourse.
- RealityVoid 1y agoI think you're right, mostly, but... While I'm sure Jim Keller will make great silicon, I'm not sure how good he's going to be at shaping the SW platform thing. I hope he will, I'm rooting for him, but I feel that this might be a novel challenge for him. Geohot is abrasive to say the least, and, no, this is not a model for productive discourse(I'll try not to bring in some of his hot takes on the stream because giving them stage is probably also not productive) But I do think he has good taste in SW and he might be right about the number of layers of abstraction. For context, geohot wrote this live on a twitch stream.
- solarpunk 1y agoprior to the elections here in the us he spent a few minutes of his streams every few days talking about how stable trump was gonna be for business and how democrats pitching capital gains taxes was a nonstarter.
- Havoc 1y ago>This guy needs to drop the middle school tough-guy tone and post numbers. Pretty sure comma is profitable? Not particularly, but for a hardware startup selling multiple iterations and not getting wrecked is a sound start
- IncreasePosts 1y agoI thought George was going to save AMD. Now he's saving tenstorrent? Busy guy!
- add-sub-mul-div 1y agoDon't forget he was also going to save Twitter but noped out after a few weeks.
- henning 1y agoMaintaining and improving existing software sure is boring and often thankless compared to starting flashy new projects where you get to make and understand all the major decisions up front.
- Aurornis 1y ago> Maintaining and improving existing software sure is boring and often thankless There was a big public episode where he appealed to Elon Musk for a job at Twitter, was given an “internship”, tried soliciting public submissions for the code he was tasked with, left a lot of people shocked that he was struggled with basic FE work, and then resigned 4 weeks in: https://news.ycombinator.com/item?id=34074344 https://news.ycombinator.com/item?id=34074344
- moralestapia 1y agogeohot has a loud mouth ... but he has earned the cred and walks the walk. I wish there was a thousand more geohots than all the mediocre middle-managers at AMD or tenstorrent; or people who have never done anything beyond posting snarky comments in online forums.
- Aurornis 1y ago> geohot has a loud mouth ... but he has earned the cred Sadly, I think geohot is an example of someone who earned some cred for impressive accomplishments in the past and then tried to cash in that cred over and over again in unrelated future domains. His brief and very public flame out at Twitter after mysteriously abandoning another project and the bold claims about his AMD work that never really translated to anything have really detracted from whatever past “cred” he built up. I really hope he can find a new niche and succeed, but until then it might be time to lie low on social media and avoid throwing more mud.
- htrp 1y agoHas geohot done anything since the original iphone jailbreak? The ventures he has started (I can think of tinygrad and comma ai) all seem like half finished tech demos.
- moralestapia 1y agoHe's the founder of comma.ai Edit: you edited your comment after I told you he made comma.ai. Bit dishonest, but whatever, I wouldn't describe comma.ai as a "half finished tech demo" but you're allowed to your own opinion about it.
- Nk26 1y agoComma.ai is the best tech I've bought in the last 3 or 4 years at least. Ive put thousands of miles into it and I've never had any issues.
- roenxi 1y agoHow many times does he have to make international news before he is qualified to let off steam in a github README? He owns a company that works in this space and his opinions usually offer insights into the ML hardware world. Although he is a little bombastic and his grammar is vulnerable to criticism. His AMD rants were a valuable warning about the quality of their hardware. I wish he'd done that maybe 10 years ago when I was buying AMD cards thinking that they might work with pytorch in a year or so. I knew they had problems but if I'd realised how bad the situation was I'd have held my nose and gone with Nvidia.
- throwawaythekey 1y agoI'm thinking about him every day lately due to a new set of AMD GPU drivers which has stopped my pc from being able to reliably wake from sleep.
- Aurornis 1y ago> His AMD rants were a valuable warning about the quality of their hardware. The rants weren’t breaking news to anyone who was familiar with PyTorch or adjacent communities. He seized upon a weak moment for AMD to try to launch his own company. Unfortunately he launched his effort with an attack on the company he was effectively trying to partner with, making the entire venture DOA. It’s too bad, too, because it would have been interesting to see if anything could have been accomplished with a more friendly offer of cooperation. He’s obviously talented as a developer, but effectively going on the attack for the company that forms the foundation of the business you’re trying to build is obviously not going to end well.
- fadfsdfaes 1y ago[flagged]
- throwaway314155 1y agoGuy would really benefit from learning some manners. Just comes across as painfully toxic no matter how correct he is. edit: For what it's worth, if you can't see that this language is rude or think it is somehow acceptable for people of a certain caliber to talk this way - you're also probably toxic.
- Havoc 1y agoWell he called sony fudge packers in a youtube rap dis track after they sued him so I wouldn't hold my breath on that one https://www.youtube.com/watch?v=9iUvuaChDEg https://www.youtube.com/watch?v=9iUvuaChDEg
- gsf_emergency 1y agoGeohot is the voodoo medication I (secretly) take everyday
- throwawaythekey 1y agoNot trying to flame or anything but what is your age? As someone in my early 30's who grew up on message boards/gaming the language he is using is fairly mild. I think we just have very different social norms.
- Aurornis 1y ago> As someone in my early 30's who grew up on message boards/gaming That’s an extremely low bar. Nobody would bat an eye about a random person speaking like this on a gaming Discord or a message board. Posting public messages as a public figure with an audience to a company is not the same as Call of Duty voice chat.
- throwawaythekey 1y agoYes and to me it is a fine tone for all of 'hacker' culture if you will. Perhaps not if you were giving a presentation to your boss, but as this is a message between peers it seems fine. If it were on x it would seem fine.
- samsartor 1y agoI'm doing my PhD in ML shit. Before that I was a systems programming guy, lots of C++, bit of CUDA, big fan of Rust. On the side I'm obsessed with RISC-V. Own a couple of boards. I made a stupid little cuda-like-compiler on top of the RISC-V vector extensions, just for fun. What I'm saying is, tensorrent couldn't find a more excitable third-party developer if they grew one in a lab. And you know what? I can't make heads or tails out of all their various abstractions. I've tried! I've read the docs, I've read the examples, I've gone to meetups. I think OP is right that "one more abstraction bro" probably doesn't solve the problem. At a guess, the problem isn't a technical one, it is an organizational one. They don't have anybody to stand in for me, or devs like me (eg dumb people). There is no product leadership on the API design. Just a lot of really brilliant engineers obsessively tuning for their own usecases, unwilling to ever trade-off a hit in performance or expressivity for readability or writeability.
- 1024bees 1y agoThere is a natural tension between developing an API that is nice to use and having a full fledged graph compiler. Most graph compilers, and the hardware that requires them will be complex and difficult to approach. The "original sin" was pytorch vs tensorflow -- tensorflow capturing the entire graph and then compiling it with XLA (or whatever it was before, I'm probably mixing up tf1 and tf2 here) was such an intractable mess to actually hack on (also the runtime had unapproachable complexity, from what I recall). This has probably changed, but pytorch won out because it was both nice to use and develop. There are clear reasons why a hardware company would use a graph compiler -- they think such an approach is higher performance, and makes tenstorrent look better on price per dollar when compared to competitors (read: nvda). There is some legitimate criticism of TT here, their hardware is composed or simple blocks that compose into a complex system (5 separate CPUs being programmed per tensix tile, many tiles per chip), and that complexity has to be wrangled in the software stack -- paying that complexity in hardware so there is less of a VLIW model in software might remove a few abstractions.
- liaopeiyuan 1y agoThis is my sentiment too after trying to get a Blackhole to run a recent VLM (like Pixtral) over the weekend. Not just unit tests, but actual training loops. I write a lot of JAX in my day job to train large models but I used to do a bit of ML compiler development, which I guess also puts me in the dumb people crowd. I'm equally impressed by how smooth the lower-level setup is and frustrated by how little progress I was able to make towards the seemingly last mile of "just rewrite the code a little bit more bro I just need to get rid of this one hlo op because it's not supported." I don't think anyone is seriously training an NN on TT hardware at the moment and I think that's an issue. I think tinygrad works not only because geohot is one hell of an engineer but also because comma dogfoods it. TT's engineers are absolutely brilliant (from reading their commits) but I think they are stretched too thin. Bounties are not gonna work - you can't expect an outsider with no internal access/bandwidth/knowledge to suddenly make e.g. Mixtral work as the issue spans at least across tt-xla/tt-mlir. And to agree with ^ training is a kind of artifact where good CX can only be derived from strong leadership and a leaner view of the stack. NVIDIA accumulated that over the decades and the rest are trying to catch up by aggressive hiring (not to say that hiring is necessary). e.g. Annapurna has a presence on the CMU campus when I was there and has the Anthropic team to test it out. I'm an incredibly excited third-party developer as I think the pitch appeals a lot to grad students (who do model research) who need to run small experiments within the 13B range and reasonably scale them up to draw the first half of the scaling curve. I lose too much productivity to abstractions and incomplete e2e support in TT's current shape. I'd love to give it another go in 6 months.
- coolThingsFirst 1y agoThis guy is insufferable. He failed his internship at Twitter and was asking questions about it publicly but has a strong opinion on everything and is an expert on everything tech related.
- arresin 1y agoWhat a bunch of absolutely limp wrist milquetoasts hackernews is / has become. So much pearl clutching over the “tone”. Oh dear. I thought this polemic was amusing and am sure comes from a place of genuine concern.
- yepyip 1y agoExactly. I do not remember this site being so delusional. Everyday it is becoming more an example of a failed institution. This is how mind viruses work.
- donperignon 1y ago[flagged]
- mlazos 1y agoI’m amazed this is even viewed as a “hot take” tbh most of what he said here is pretty high level of abstraction and standard practice for custom hardware. In essence I feel like he’s saying nothing really controversial other than publicly calling out TT for too many abstraction layers (and tbh it’s just in a readme). This is completely fine, he’s a user and this is his experience. I’m a dev working on torch.compile at meta (previously I worked on ML focused FPGAs) and the approach I would use is build a static graph compiler, use torch.compile (and probably JAX) as graph extraction front-ends and call it a day. I feel like hardware companies don’t know how to handle the flexibility of PyTorch and as a result develop their own APIs which is mistake #1 and virtually makes it impossible to get any market penetration once you head down that path because nobody will ever ever rewrite their models for your hardware when they don’t even know what perf they will get, the risk is just too high. As a result, hardware companies offer inference APIs which hide all of this behind a REST API to basically paper over the lack of generality of the software/hardware interface. This is convenient because then nobody actually knows the perf/$ and they can burn VC money for as long as they want. Whether this is a viable business model or not, we will have to wait until they go public to actually see what their true inference costs are. To sum it up, start from PyTorch and work your way down to your hardware, this is the only general way if you want to actually sell chips and not just constantly port the model of the day to your hardware.