8 ms·
Tinybox – A powerful computer for deep learning
- vlovich123 7mo agoSurprising to see this with AMD GPUs considering how George famously threw up his hands as AMD not being worth working with.
- embedding-shape 7mo agoYeah, and labeling AMD "Driver Quality" as "Good" (for comparison, they label nvidia's driver quality as "Great").
- lostmsu 7mo agoThings changed. On my new Ryzen Strix Halo laptop I was able to run training experiments with PyTorch on Windows day 1: https://news.ycombinator.com/item?id=46052535 https://news.ycombinator.com/item?id=46052535
- vlovich123 7mo agoYeah, installing random wheels from non official sources is an improvement. Not sure I’d characterize that as an unmitigated win. But also as soon as you try to do more involved things, at least personally, I ran into serious challenges getting things to work.
- wongarsu 7mo agoSound like solid prebuilt with well balanced components and a pretty case Not revolutionary in any way, but nice. Unless I'm missing something here?
- eurekin 7mo agoIt's pretty close to what people have been frankenbuilding on r/LocaLLaMa... It's nice to have a prebuild option.
- speedgoose 7mo agoYou could also order such configurations from a classic server reseller as far as I know. The case is a bit original there.
- nextlevelwizard 7mo agoTiny boxes are already several years old IIRC
- llbbdd 7mo agoIf you wanted a box built by geohot, most recently known for signing on to Elons Twitter and then bailing, it's for you
- asadm 7mo agoactually known for comma.ai
- heinternets 7mo agoexabox - 720x RDNA5 AT0 XL 25,920 GB VRAM 23,040 GB System RAM ~ $10 Million Who is the target market here?
- spiderfarmer 7mo agoVC funded startups
- orochimaaru 7mo agoAnd... what about 20k lbs and 1360 cubic feet screams "tiny" :)
- smoyer 7mo agoThat is very close to a half-length shipping container.
- mayukh 7mo agoA non-trivial share of this market won’t show up in public data. That makes most estimates unreliable by default
- LorenDB 7mo agoI can't find sources but I think they are building it for Comma.ai (geohot's other company) so that Comma can scale up their training datacenter.
- dist-epoch 7mo agoA company which doesn't want the big LLM providers to see it's prompts or data - military, health, finance, research
- vessenes 7mo agoThe exabox is interesting. I wonder who the customer is; after watching the Vera Rubin launch, I cannot imagine deciding I wanted to compete with NVIDIA for hyperscale business right now. Maybe it’s aiming at a value-conscious buyer? Maybe it’s a sensible buy for a (relatively) cash-strapped ML startup; actually I just checked prices, and it looks like Vera Rubin costs half for a similar amount of GPU RAM. I’m certain that the interconnect will not be as good as NV’s. I have no idea who would buy this. Maybe if you think Vera Rubin is three years out? But NV ships, man, they are shipping.
- zozbot234 7mo ago> The exabox is interesting. Can it run Crysis?
- bastawhiz 7mo agoProbably, the rdna5 can do graphics. But it would be a huge waste, since you could probably only use one of the 720 GPUs
- WithinReason 7mo agoOnly gamers understand that reference -- Jensen Huang
- dist-epoch 7mo agoYes, it can generate Crysis with diffusion models at 60 fps.
- 7mo ago
- orliesaurus 7mo agoI wonder if this is frontpage right now because of the other tiiny (the names are similar) video that went viral ... which turns out wasn't an actual product by the tinygrad linked in this post[1] [1]https://x.com/ShriKaranHanda/status/2035284883384553953 https://x.com/ShriKaranHanda/status/2035284883384553953
- comrade1234 7mo agoCool that you have a dual power supply model. It says rack mountable or free standing. Does that mean two form factors? $65K is more than we can afford right now but we are definitely eventually in the market for something we can run in our own colo. It's funny though... we're using deepseek now for features in our service and based on our customer-type we thought that they would be completely against sending their data to a third-party. We thought we'd have to do everything locally. But they seem ok with deepseek which is practically free. And the few customers that still worry about privacy may not justify such a high price point.
- hrmtst93837 7mo ago[flagged]
- zozbot234 7mo agoThe real case for private inference is not "organic", it's "slow food". Offering slow-but-cheap inference is an afterthought for the big model providers, e.g. OpenRouter doesn't support it, not even as a way of redirecting to existing "batched inference" offerings. This is a natural opening for local AI.
- selectodude 7mo agoBut how slow is too slow (faster than you’d think) and even then, you’re in for $25,000 for even the most basic on-premise slow LLM.
- aplomb1026 7mo ago[dead]
- jauntywundrkind 7mo agoMy interest in anything associated with geohot took a colossal nose dive today after seeing this post against democracy, quoting frelling M*ncius M*ldbug: Democracy is a Liability. https://news.ycombinator.com/item?id=47469543 https://news.ycombinator.com/item?id=47469543 https://geohot.github.io//blog/jekyll/update/2026/03/21/democracy-liability.html https://geohot.github.io//blog/jekyll/update/2026/03/21/demo... Theres a lot there that makes sense & I think needs to be considered. But a lot just seems to be out of the blue, included without connection, in my view. Feels like maybe are in-grouo messages, that I don't understand. How this is headered as against democracy is unclear to me, and revolting. I both think we must grapple with the world as it is, and this post is in that area, strongly, but to let fear be the dominant ruling emotion is one of the main definitions of conservativism, and it's use here to scare us sounds bad.
- tadfisher 7mo agoFor those unaware, Mencius Moldbug is the pen name of Curtis Yarvin, thought leader for the Silicon Valley branch of right-wing technofascist weirdos which includes Peter Thiel and apparently half of a16z.
- pencilheads 7mo agoGeohot has always been an arrogant cunt who thinks he's better than everyone else. That blog post is totally on brand.
- fragmede 7mo agoDamn, that's a take.
- stale2002 7mo agoGeohotz's politics are fairly straightforward once you understand his background. Geohotz is the prodigy child who, at the age of ~16 accomplished amazing technical feats on his own. And his politics are a derivative of Great Man Theory, and his positions on things like democracy follow from that. This idea, and those espoused by some of the VC/tech elite like Peter Theil are that singular hardworking genius individuals can change the world on their own, and everyone who not in this top 0.1% are borderline NPCs. They do this both because of their genius/hardwork, and also because they are willing to break the rules that are set forth by this bottom 99.9%. I'm starting to call this ideology Authoritarian techno-Libertarianism. Its a delibriately oxymoronic name that I use, because these "Great Men" are definitely trying to change the world. IE, they are trying to impose their goals and values on the world without getting the buyin of other people. Thats the "authoritarian" part. And then the "libertarian" part is that they are going about this imposition of their will on the world by doing it all themselves, through their own hard work. Think "Person invents a world changing technology, that some people thing is bad, and just releases it open source for anyone to use". AI models are a great example, in fact. Once that technology is out there the genie cannot be put back into the bottle and a ton of people are going to lose their jobs, ect. A distain for democracy follows directly from things like this. You dont wait for people to vote to allow you to change the world by inventing something new. You just do and watch the results.
- ivraatiems 7mo agoThere's some irony in the fact that this website reads as extremely NOT AI-generated, very human in the way it's designed and the tone of its writing. Still, this is a great idea, and one I hope takes off. I think there's a good argument that the future of AI is in locally-trained models for everyone, rather than relying on a big company's own model. One thought: The ability to conveniently get this onto a 240v circuit would be nice. Having to find two different 120v circuits to plug this into will be a pain for many folks.
- wat10000 7mo agoIf you’re spending $65,000 on this thing, needing two circuits seems like a minor problem
- ivraatiems 7mo agoThe $12,000 one also requires it.
- isatty 7mo agoSurprisingly affordable but I’m not really interested in the 9070XT. If it shipped with like 4090+ (for a higher price) it’d be more tempting.
- dmarcos 7mo agoThey offered a version a few months ago with 4x5090 for 25k https://x.com/__tinygrad__/status/1983917797781426511 https://x.com/__tinygrad__/status/1983917797781426511 Stopped due to raising GPU prices: https://x.com/__tinygrad__/status/2011263292753526978 https://x.com/__tinygrad__/status/2011263292753526978
- ycui1986 7mo ago9070XT provide roughly same inference performance at double the power, half the cost, as RTX PRO 4500. So this one is optimized for total BOM cost.
- droidjj 7mo agoAdding this to my list of ~beautifully~ designed things to buy when I win the lottery.
- pink_eye 7mo ago[flagged]
- throwatdem12311 7mo agoFinally, a computer that should be able to run Monster Hunter Wilds with decent performance. But let’s be real, 12k is kinda pushing it - what kind of people are gonna spend $65k or even $10M (lmao WTAF) on a boutique thing like this. I dont think these kinds of things go in datacenters (happy to be corrected) and they are way too expensive (and probably way too HOT) to just go in a home or even an office “closet”.
- oofbey 7mo agoIt’s not for people to buy. It’s for companies to buy. Compare to salary, and it’s cheap.
- throwatdem12311 7mo agoWhat companies are buying this instead of like a Dell server or whatever?
- flumpcakes 7mo agoThese specs look enormously cheaper than doing it with dell servers. The last quote I had for a bog standard dell server was $50k and only if bought in the next few days or so. The prices are going up weekly.
- throwatdem12311 7mo agoSo what’s the catch? If it seems too good to be true it probably is.
- wmf 7mo agoThese are "unsupported" configurations. Nvidia/AMD discourage running multiple gaming/workstation cards and encourage customers to buy $500K SXM/OAM servers.
- lostmsu 7mo ago
- mayukh 7mo agoWhat’s the most effective ~$5k setup today? Interested in what people are actually running.
- oofbey 7mo agoDGX Spark is a fantastic option at this price point. You get 128GB VRAM which is extremely difficult to get at this price point. Also it’s a fairly fast GPU. And stupidly fast networking - 200gbps or 400gbps mellanox if you find coin for another one.
- BobbyJo 7mo agoInternet seems to think the SW support for those is bad, and that strix halo boxes are better ROI.
- oofbey 7mo agoMeh. DGX is Arm and CUDA. Strix is X86 and ROCm. Cuda has better support than ROCm . And x86 has better support than Arm. Nowadays I find most things work fine on Arm. Sometimes something needs to be built from source which is genuinely annoying. But moving from CUDA to ROCm is often more like a rewrite than a recompile.
- BobbyJo 7mo agoCUDA != Driver support. Driver support seems to be what's spotty with DGX, and iirc Nvidia jas only committed to updates for 2 years or something.
- overfeed 7mo ago> But moving from CUDA to ROCm is often more like a rewrite than a recompile. Isn't everyone* in this segment just using PyTorch for training, or wrappers like Ollama/vllm/llama.cpp for inference? None have a strict dependency on Cuda. PyTorch's AMD backend is solid (for supported platforms, and Strix Halo is supported). * enthusiasts whose budget is in the $5k range. If you're vendor-locked to CUDA, Mac Mini and Strix Halo are immediately ruled out.
- operatingthetan 7mo agoThe incremental price increases between products is funny. $12,000, $65,000, $10,000,000.
- sudo_cowsay 7mo agoI mean the difference in performance is quite big too. However, the 10,000,000 is a little bit too much (imo).
- znpy 7mo agoI was more worried by the 600kW power requirement... that's 200 houses at full load (3kw) in southern europe... which likely means 400 houses at half load. the town near my hometown has 650 – 800 houses (according to chatgpt). crazy.
- dist-epoch 7mo agoYour hometown also has public lightning, water pumps, and probably some other stuff.
- nine_k 7mo agoOr it's two 300kW fast EV chargers working together. A typical home just consumes rather little energy, now that LED lighting and heat pump cooling / heating became the norm.
- znpy 7mo ago> now that LED lighting and heat pump cooling / heating became the norm. My brother in Christ, you vastly overestimate southern europe
- nine_k 7mo agoI noticed that Southern Europe often basically ignores both heating and cooling, especially close to the warm sea. But with heat pumps becoming normal in the North and in the US, they become mass-produced, and the prices fall. Same has happened to LED lamps.
- Heer_J 7mo ago[dead]
- sudo_cowsay 7mo agoI always wonder about these expensive products: Does the company make them once its ordered or do they just make them beforehand?
- bastawhiz 7mo agoThere's no way the red v2 is doing anything with a 120b parameter model. I just finished building a dual a100 ai homelab (80gb vram combined with nvlink). Similar stats otherwise. 120b only fits with very heavy quantization, enough to make the model schizophrenic in my experience. And there's no room for kv, so you'll OOM around 4k of context. I'm running a 70b model now that's okay, but it's still fairly tight. And I've got 16gb more vram then the red v2. I'm also confused why this is 12U. My whole rig is 4u. The green v2 has better GPUs. But for $65k, I'd expect a much better CPU and 256gb of RAM. It's not like a threadripper 7000 is going to break the bank. I'm glad this exists but it's... honestly pretty perplexing
- oceanplexian 7mo agoIt will work fine but it’s not necessarily insane performance. I can run a q4 of gpt-oss-120b on my Epyc Milan box that has similar specs and get something like 30-50 Tok/sec by splitting it across RAM and GPU. The thing that’s less useful is the 64G VRAM/128G System RAM config, even the large MoE models only need 20B for the router, the rest of the VRAM is essentially wasted (Mixing experts between VRAM and/System RAM has basically no performance benefit).
- syntaxing 7mo agoSplit RAM and GPU impacts it more than you think. I would be surprised if the red box doesn’t outperform you by 2-3X for both PP and TG
- androiddrew 7mo agoCould you share what you are using for inference and how you are running it? I have a 64G VRAM/128G system RAM setup.
- sosodev 7mo agoMost people are using something in the llama family for inference. Llama server is my go to. Unsloth guides describe how to configure inference for your model of choice.
- operatingthetan 7mo agoAre we at the point where 2x 9070XT's are a viable LLM platform? (I know this has 4, just wondering for myself).
- oceanplexian 7mo agoThese things don’t have Flash Attention or either have a really hacked together version of it. Is it viable for a hobby? Sure. Is it viable for a serious workload with all the optimizations, CUDA, etc.. Not really.
- cyanydeez 7mo agoI'd go with strix halo if you're looking at that old of tech. the latest AMD GPUs are RX 9070 XT w/32GB each
- ekropotin 7mo agoIDK, I feel it’s quite overpriced, even with the current component prices. I almost sure it’s possible to custom build a machine as powerful as their red v2 within 9k budget. And have a lot of fun along the way.
- lostmsu 7mo agoAMD now has 32 GiB Radeon AI Pro 9700. 4 of these (just under 2k each) would put you at 128 GiB VRAM
- ekropotin 7mo agoVRAM is not everything - GPU cores also matter (a lot) for inference
- lostmsu 7mo ago4x Radeon will have significantly more GPU power than say Mac Studio or DGX Spark.
- cyanydeez 7mo agoinference speed is like monitor Hz; sure, you go from 60 to 120Hz and thats noticeable, but unless your model is AGI, at some point you're just generating more code than you'll ever realistically be able to control, audit and rely on. So, context is probably more $/programming worth than inference speed.
- himata4113 7mo agoexabox reads as if it was making a joke of something or someone. if it's real then it's really interesting!
- andai 7mo agoCan someone explain the exabox? They say it "functions as a single GPU". Is there anything like that currently existing?
- baibai008989 7mo ago[dead]
- ppap3 7mo agoI thought there was a typo in the price
- mmoustafa 7mo agoI would love to see real-life tokens/sec values advertised for one or various specific open source models. I'm currently shopping for offline hardware and it is very hard to estimate the performance I will get before dropping $12K, and would love to have a baseline that I can at least always get e.g. 40 tok/s running GPT-OSS-120B using Ollama on Ubuntu out of the box.
- hpcjoe 7mo agoLook for llmfit on github. This will help with that analysis. I've found it reasonably accurate. If you have Ollama already installed, it can download the relevant models directly.
- atwrk 7mo agoFor reference, 12k gets you at least 4 Strix Halo boxes each running GPT-OSS-120B at ~50tok/s.
- aabaker99 7mo ago> Can I pay with something besides wire transfer? In order to keep prices low and quality high, we don't offer any customization to the box or ordering process. Wire transfer is the only accepted form of payment. Sorry, what? Is this just a scam?
- ejpir 7mo agoman, cmon. a little more effort.
- aabaker99 7mo agoSure thing. For those who don’t know, wiring money like this is a good way to lose your money. https://consumer.ftc.gov/articles/what-know-you-wire-money https://consumer.ftc.gov/articles/what-know-you-wire-money
- metadata 7mo agoWire transfer is a bank transfer, not money wire to Western Union and like.
- aabaker99 7mo agoYeah I agree the FTC article could be more clear here. I think they call out Western Union because those are tools that are commonly used by scammers. But let’s be clear: the risks are the same if you are wiring money through Western Union or wiring through any other bank. Once you wire the money you do not have the same protections as other payment mechanisms. And if you don’t get the product as described, you are likely out your money. This is compared to other forms of payment like credit cards where you are protected. With a credit card you can issue a charge back to the seller and get your money back in the case of fraud. With a wire transfer you cannot.
- 101008 7mo agoWire transfer has no comission or extra costs associated to it, so I find it very honest.
- ilaksh 7mo agoI thought the most interesting thing about tinygrad was that theoretically you could render a model all the way into hardware similar to Taalas (tinygrad might be where Taalas got the idea for all I know). I could swear I filed a GitHub issue asking about the plans for that but I don't see it. Anyway I think he mentioned it when explaining tinygrad at one point and I have wondered why that hasn't got more attention. As far as boxes, I wish that there were more MI355X available for normal hourly rental. Or any.
- flykespice 7mo ago"tiny" and it's 20k lbs and cost about 10k... Since when did our perception of tiny blow out of size in tech? Is it the influence of "hello world" eletron apps consuming 100mb of mem while idle setting the new standard? Anyway being an AI bro seems like an expensive hobby...
- zahirbmirza 7mo ago10 mil today... 1k in 10 years. Are OpenAI and Anthropic overvalued?
- Gigachad 7mo agoLooking at these prices I’m just thinking that as a user it makes no sense to buy this when you can just use the subsidised stuff from AI companies and then buy it a few years later at a tiny % of the cost.
- adrianwaj 7mo agoPerhaps this company should think about acting as a landlord for their hardware. You buy (or lease) but they also offer colocation hosting. They could partner with crypto miners who are transitioning to AI factories to find the space and power to do this. I wonder if the machines require added cooling, though, in what would otherwise be a crypto mining center. CoreWeave made the transition and also do colocation. The switchover is real. I think Tinygrad should think about recycling. Are they planning ahead in this regard? Is anyone? My thought is if there was a central database of who own what and where, at least when the recycling tech become available, people will know where to source their specific trash (and even pay for it.) Having a database like that in the first place could even fuel the industry.
- fhn 7mo ago"but if you haven't contributed to tinygrad your application won't be considered" this company expects people to work for free?
- paxys 7mo ago> See our bounty page to judge if you might be a good fit. Bounties pay you while judging that fit. Literally the line above that
- roarcher 7mo agoThey MIGHT pay you IF you're a fit. They're bounties, i.e. spec work. They also pay a max of $1000, most of them significantly less. You can see more info at the link in that line: > All bounties paid out at my (geohot) discretion. Code must be clean and maintainable without serious hacks. No thanks. If you want to try before you buy, have your candidates do a paid test project. Founders need to stop acting like it's a privilege to work for them. Any talent worth hiring has plenty of other options that will treat them with respect.
- deleted 7mo ago[deleted]
- paxys 7mo agoThe problem with all these "AI box" startups is that the product is too expensive for hobbyists, and companies that need to run workloads at scale can always build their own servers and racks and save on the markup (which is substantial). Unless someone can figure out how to get cheaper GPUs & RAM there is really no margin left to squeeze out.
- kkralev 7mo ago[flagged]
- wmf 7mo agojust want to run a 7-8b model locally This is already solved by running LM Studio on a normal computer.
- zozbot234 7mo agoOllama or llama.cpp are also common alternatives. But a 8B model isn't going to have much real-world knowledge or be highly reliable for agentic workloads, so it makes sense that people will want more than that.
- zach_vantio 7mo agothe compute density is insane. but giving a 70B model actual write access locally for agentic workloads is a massive liability. they still hallucinate too much. raw compute without strict state control is basically just a blast radius waiting to happen.
- nine_k 7mo agoWould a hedge fund that does not want to trust to a public AI cloud just buy chassis, mobos, GPUs, etc, and build an equivalent themselves? I suspect they value their time differently.
- paxys 7mo ago
- rpastuszak 7mo agoWho is this for?
- siliconc0w 7mo agoTinybox is cool but I think the market is maybe looking more for a turn-key explicit promise of some level of intelligence @ a certain Tok/s like "Kimi 2.5 at 50Tok/s".
- p0w3n3d 7mo agoQuite expensive little bastard. I wonder how much does it make sense to invest in a such device, if you can get $0.40/mtok from hyperbolic for example
- sowbug 7mo agoIf you're OK letting them train on, and maybe keep, your data, then it's hard to beat cloud prices vs. local.
- jeremie_strand 7mo ago[flagged]
- latchkey 7mo agoOld news. ROCm works a lot better now than it did a year ago.
- Gigachad 7mo agoYou are still really limited in what you can run. So much stuff is cuda only.
- latchkey 7mo agoLike what? Most of the good stuff is ported over already and anything else, tag Anush on X and see what you get. Also happy to help. The point is that they care now.
- Gigachad 7mo agoTbh my experience is in the non AI uses, recently I was looking at Gaussian splatting tools and it seemed the majority of it was CUDA only. I’m also still bothered AMD for ages claimed my card (5700xt) would be getting rocm but just abandoned it.
- latchkey 7mo ago>I was looking at Gaussian splatting tools and it seemed the majority of it was CUDA only. Not surprising. True, the ecosystem is like early OSX vs. Windows. Eventually it'll get ported over if there is demand.
- djsjajah 7mo agotrl. give me a uv command to get that working. But even in the amd stack things (like ck and aiter) consumer cards are not even second class citizens. They are a distance third at best. If you just want to run vllm with the latest model, if you can get it running at all there are going to be paper cuts all along the way and even then the performance won't be close to what you could be getting out of the hardware.
- hmokiguess 7mo agoIs this like the new equivalent of crypto mining? I remember the early days when they would sell hardware for farming crypto, now it’s AI?
- renewiltord 7mo agoI have 8x RTX 6000 Pro. Better to run the 300 W version of the cards. And it costs close to their 4x version. I get why they make it so big. So you can cool it at home. I prefer to just put in datacenter. Much cheaper power.
- SmartestUnknown 7mo agoRegarding 2x faster than pytorch being a condition for tinygrad to come out of alpha: Can they/someone else give more details as to what workloads pytorch is more than 2x slower than the hardware provides? Most of the papers use standard components and I assume pytorch is already pretty performant at implementing them at 50+% of extractable performance from typical GPUs. If they mean more esoteric stuff that requires writing custom kernels to get good performance out of the chips, then that's a different issue.
- aplomb1026 7mo ago[dead]
- mciancia 7mo agoNot sure why they stopped using 6 GPUs in thei builds - with 4 GPUs, both 9070 and rtx6000 come in 2 slot designs, so it easy to build it yourself using a bit more expensive, but still fairly regular motherboard. With 6 GPUs you have to deal with risers, pcie retimers, dual PSUs and custom case for so value proposition there was much better IMO
- ks2048 7mo ago"... and likely the best performance/$". "likely" doesn't inspire much confidence. Surely, they have those numbers, and if it was, they'd publicize the comparisons.
- deleted 7mo ago[deleted]
- deleted 7mo ago[deleted]
- alexfromapex 7mo ago$12,000 for the base model is insane. I have an Apple M3 Max with 128GB RAM that can run 120B parameter models using like 80 watts of electricity at about 15-20 tokens/sec. It's not amazing for 120B parameter models but it's also not 12 grand.
- segmondy 7mo agoit's for fools. i bought 160gb of vram for $1000 last year. 96gb of p40 VRAM can be had for under $1000. And it will run gpt-oss-120b Q8 at probably 30tk/sec
- timschmidt 7mo agoP40 is Tesla architecture which is no longer receiving driver or CUDA updates. And only available as used hardware. Fine for hobbyists, startups, and home labs, but there is likely a growing market of businesses too large to depend on used gear from ebay, but too small for a full rack solution from Nvidia. Seems like that's who they're targeting.
- segmondy 7mo ago99% of interest is in inference. If you want to fine-tune a model, just rent the best gpu in the cloud. It's often cheaper and faster.
- timschmidt 7mo agoGreat option if you don't mind sharing your data with the cloud. Some businesses want to own the hardware their data resides on.
- cootsnuck 7mo agoHow many businesses have the capabilities and expertise to train their own models?
- adi_kurian 7mo agohttps://en.wikipedia.org/wiki/Decoy_effect https://en.wikipedia.org/wiki/Decoy_effect
- raincole 7mo agoHow does this thing cool down?
- mememememememo 7mo agoGive me token/s for favourite models.
- Buttons840 7mo agoOh, this is geohots product? He's an interesting guy. Seems to be one who does things the way he thinks is right, regardless of corporate profits.
- caijia 7mo ago[flagged]
- alasdair_ 7mo agoI just don’t believe that this can run inference on a 120 billion parameter model at actually useful speeds. Obviously any Turing machine can run any size of model, so the “120B” claim doesn’t mean much - what actually matters is speed and I just don’t believe this can be speedy enough on models that my $5000 5090-based pc is too slow for and lacks enough vram for.
- mnkyprskbd 7mo agoLook at the GPU and RAM spec; 120b seems workable.
- Aurornis 7mo agoFor the red v2? 120B could run, but I wouldn't want to be the person who had to use it for anything. To be fair, the 120B claim doesn't appear on the webpage. I don't know where it came from, other than the person who submitted this to HN
- mnkyprskbd 7mo agoIt is more than fair, also, you're comparing your 5k devices to 12k and more importantly 65k and >10m devices.
- Aurornis 7mo agoThe "to be fair" part of my comment was saying that the tinygrad website doesn't claim 120B. Also nobody is comparing this box to an $10M nVidia rack scale deployment. They're comparing it to putting all of the same parts into their Newegg basket and putting it together themself.
- arunakt 7mo agoGreat idea, can you publish the power consumption units for this device
- jgarzik 7mo agoSkeptical of their engineering, with replies to questions like this: https://x.com/jgarzik/status/2031312666036146460?s=20 https://x.com/jgarzik/status/2031312666036146460?s=20
- _2d30 7mo agoThey answered your question with a pretty specific uptime target. Calling it a dodge and then moving the goalposts with a new question as your follow up doesn’t speak to you acting in good faith.
- scratchyone 7mo agotbh they really didn't, tinygrad's was clearly a joke response. they were not providing a real uptime target.
- potamic 7mo agoCan't see replies, what did they say?
- Moduke 7mo agohttps://xcancel.com/jgarzik/status/2031312666036146460?s=20 https://xcancel.com/jgarzik/status/2031312666036146460?s=20
- mellosouls 7mo agoWhere is the 120B documented? This seems to be an editorialized title. Edit: found a third party referencing the claim but it doesn't belong in the title here I think: Meet the World’s Smallest ‘Supercomputer’ from Tiiny AI; A Machine Bold Enough to Run 120B AI Models Right in the Palm of Your Hand https://wccftech.com/meet-the-worlds-smallest-supercomputer-a-machine-bold-enough-to-run-120b-ai-models/ https://wccftech.com/meet-the-worlds-smallest-supercomputer-...
- Aurornis 7mo agoThat third party link is from a different company (Tiiny with an extra i) Now I'm wondering if the HN title was submitted by some AI bot that couldn't tell the difference.
- deleted 7mo ago[deleted]
- mellosouls 7mo agoHa, good catch, I googled for Tinybox 120B and clearly didn't read the article beyond the seeming match.
- roarcher 7mo ago> In order to keep prices low and quality high, we don't offer any customization to the box or ordering process. If you aren't capable of ordering through the website, I'm sorry but we won't be able to help. Has this guy never worked on a B2B product before? Nobody is going to order a $10 million piece of infrastructure through your website's order form. And they are definitely going to want to negotiate something, even if it's just a warranty. And you'll do it because they're waving a $10 million check in your face. The tone of this website is arrogant to the point of being almost hostile. The guy behind this seems to think that his name carries enough weight to dictate terms like this, among other things like requiring candidates to have already contributed to his product to even be considered for a job. I would be extremely surprised if anyone except him thinks he's that important.
- jrflowers 7mo agoI imagine that the FAQ might get updated when there’s actually a $10M machine for sale
- roarcher 7mo agoMaybe. Frankly I'd be very surprised if any business ordered a $65k machine that way either.
- jrflowers 7mo agoYeah it’s a little odd. Maybe they are meant to be really really cool toys? People regularly spend more than $65k on things like cars to show off, so it could be like that. I have no use for these but I might buy one anyway if I won the lottery. ¯\_(ツ)_/¯
- wmf 7mo agoHe's not actually selling the exabox yet. It sounds like he put up a hypothetical config to see if anyone is interested.
- jen729w 7mo agoYour framing of this section is misleading. On the site it's preceded by a FAQ-style 'question': > Can you fill out this supplier onboarding form? That's very important context, as anyone who has been asked to fill out a supplier onboarding form (hi) will attest.
- kylehotchkiss 7mo agoMeanwhile M-series processors and Qwen are racing to do the same thing for a much more approachable price.
- WWilliam 7mo ago[flagged]
- jmspring 7mo agoTinygrad devices are interesting, I wish I have screen captures - but their prices have gone up and some specs like RAM have gone down. A single box with those specs without having to build/configure (the red and green) - I could see being useful if you had $ and not time to build/configure/etc yourself.
- deleted 7mo ago[deleted]
- jee599 7mo ago[flagged]
- insane_dreamer 7mo agoIs this real? Reads like a joke. They sell a $12K machine, a $60K machine, and a $10M machine???
- wmf 7mo agoNvidia has $4K DGX Spark, $120K DGX Station, $500K DGX, and $7M NVL72.
- agnishom 7mo agoWho is the intended customer for this product? I am genuinely curious.
- moscoe 7mo agoAnyone who wants to run/train/finetune a local llm. “Not your weights, not your brain.”
- jee599 7mo ago[flagged]
- gymbeaux 7mo ago$12,000 gets you 1Gb/s networking and vanilla Ubuntu 24.04. Napkin math on the hardware it looks like margins are around 50% which feels like a school fundraiser where everyone pays what is obviously way more than normal retail price for X because "it's for the children." I'm not sure what tinygrad is but I assume the markup is because the customer is making a conscious choice to support the tinygrad project. But what's unusual is there is apparently no reason whatsoever to buy this hardware, even if you plan on using tinygrad exclusively for your project. At least with System76 hardware I get (in theory) first class support for Pop!_OS.
- qubex 7mo agoI just backed their TINY on Kickstarter.
- rick_dalton 7mo agoThat thing is NOT related to tinybox or tinygrad in any way. It is basically copyright infringement. Unless you’re astroturfing here I suggest you get your money back.
- qubex 7mo agoWasn’t astroturfing, I’ll look into it, thanks.
- rick_dalton 7mo agoSorry for even mentioning astroturfing, haha. It’s just because the promotion of the device is based on trying to fool people it was made by tiny corp.
- qubex 7mo agoIn my case they apparently succeeded.
- saidnooneever 7mo agoits a bit weird to me ud need to be contributor to their software to work in operations or hardware, but I suppose its ok for tinycompany. in long term its likely better to have domain experts and not bias everything towards the same thing. the boxes look cool but how good are they really? the cheapest box seems pricey at 12 for a what is essentially a few gaming gpus. i dont see why you couldnt make that like half the price. u could do a PC/server build thats much much faster for way less. size doesnt matter if its more than twice the price i think... the more expensive box has atleast real processing gpus but afaik also not very popular ones, this one seems maybe more fair priced (there seems a big difference in bang for buck between these???). the third one suggested looks like a joke. dont get me wrong, this seems like a really cool idea. But i dont see it taking off as the prices are corporate but the product seems more home use. maybe in time they will find a better balance, i do respect the fact that the component market now is sour as hell and making good products with stable prices is pretty much i possible. id love one of these machines someday, maybe when i am less poor, or when they are xD. (love the styling of everything, this is the most critical i could be from a dumb consumer perspective, which i totally am btw.)
- chloecv 7mo ago[dead]
- Yanko_11 7mo ago[dead]
- algolint 7mo ago[flagged]
- EruditeCoder108 7mo ago[dead]
- DeathArrow 7mo agoWhy do I get the impression that I get more bang for the buck by going through OpenRouter? Of course, not anyone can do that and there are security and other concerns.
- DeathArrow 7mo agoI wonder how much has he sold.
- triwats 7mo agoThis is cool, I'll add these as desktops to https://flopper.io https://flopper.io! How do you test/generate these numbers?
- the_arun 7mo agoCurious to know who will spend this much money without external funding? Would you spend any VC invested money into this nameless brand? Are there any guardrails or clauses to protect the kind of expenses?
- pugchat 7mo ago[dead]
- h14h 7mo agoWould be very curious how RL benchmarks shake out vs M5 Pro/Max. Doubt local inference is the target use case near nearly as much as post-training. I could totally see something like this being super appealing for a startup looking to do some fine-tuning/distillation to tune a small open-weight model for a narrow use case.
- Aissen 6mo agoIt might be a bit CPU and RAM starved… Which in theory should be OK, but in practice you'll find production workloads that struggle because of this. Just make sure whatever you want to run on this is indeed extremely GPU-bound, or you might have bad surprises later.