6 ms·
Show HN: Run TRELLIS.2 Image-to-3D generation natively on Apple Silicon
I ported Microsoft's TRELLIS.2 (4B parameter image-to-3D model) to run on Apple Silicon via PyTorch MPS. The original requires CUDA with flash_attn, nvdiffrast, and custom sparse convolution kernels: none of which work on Mac.
I replaced the CUDA-specific ops with pure-PyTorch alternatives: a gather-scatter sparse 3D convolution, SDPA attention for sparse transformers, and a Python-based mesh extraction replacing CUDA hashmap operations. Total changes are a few hundred lines across 9 files.
Generates ~400K vertex meshes from single photos in about 3.5 minutes on M4 Pro (24GB). Not as fast as H100 (where it takes seconds), but it works offline with no cloud dependency.
https://github.com/shivampkumar/trellis-mac https://github.com/shivampkumar/trellis-mac
- villgax 6mo agoThat’s always been possible with MPS backend, the reason people choose to omit it in HF spaces/demos is that HF doesn’t offer an MPS backend. People would rather have the thing work at best speeds than 10x worse speeds just for compatibility.
- Reubend 6mo agoAre you saying the original one worked with MPS? Or are you just saying it was always theoretically possible to build what OP posted?
- villgax 6mo agoLatter
- refulgentis 6mo agoIt’s always been possible, but it’s not possible because there’s no backend, and no one wants to it to be possible because everyone needs it 10x the speed of running on a Mac? I’m missing something, I think.
- shivampkumar 6mo agoI thought it was cool and then I found the open issue mentioned above, that convinced me its def something more people want. It IS significantly slower, about 3.5 minutes on my MacBook vs seconds on an H100. That's partly the pure-PyTorch backend overhead and partly just the hardware difference. For my use case the tradeoff works -- iterate locally without paying for cloud GPUs or waiting in queues.
- shivampkumar 6mo agoIMO TRELLIS.2 is slightly different case from the HF models scenario. It depends on five compiled CUDA-only extensions -- flex_gemm for sparse convolution, flash_attn, o_voxel for CUDA hashmap ops, cumesh for mesh processing, and nvdiffrast for differentiable rasterization. These aren't PyTorch ops that fall back to MPS -- they're custom C++/CUDA kernels. The upstream setup.sh literally exits with "No supported GPU found" if nvidia-smi isn't present. The only reason I picked this up because I thought it was cool and no one was working on this open issue for Silicon back then (github.com/microsoft/TRELLIS.2/issues/74) requesting non-CUDA support.
- gondar 6mo agoNice work. Although this model is not very good, I tried a lot of different image-to-3d models, the one from meshy.ai is the best, trellis is in the useless tier, really hope there could be some good open source models in this domain.
- shivampkumar 6mo agoHey, thanks for sharing this. I'm sure TRELLIS.2 definitely has room to improve, especially on texturing. From what I've seen personally, and community benchmarks, it does fair on geometry and visual fidelity among open-source options, but I agree it's not perfect for every use case. Meshy is solid, I used it to print my girlfriend a mini 3d model of her on her birthday last year! Though worth noting it's a paid service, and free tier has usage limitations while TRELLIS.2 is MIT licensed with unlimited local generation. Different tradeoffs for different workflows. Hopefully the open-source side keeps improving.
- isoprophlex 6mo agoMeshy is indeed great but I am terminally put off by their alltogether terrible, sleazy, gamified, opaque web UI. It's like aliexpress and a lootbox game had a baby that's into mesh generation. Ugh.
- jseabra 6mo ago[dead]
- kennyloginz 6mo agoSo much effort, but no examples in the landing page.
- shivampkumar 6mo agoYou're right, thanks for flagging this, let me run something and push images
- shivampkumar 6mo agoadded! will add more, maybe even a GIF
- hank808 6mo ago[flagged]
- kennyloginz 6mo agoGood question.
- shivampkumar 6mo agoI mean I can see that it's niche. Did not expect so many upvotes, but ig it's less niche than I tought If you're not working with 3D on Apple Silicon this isn't relevant to you. For the subset of people who are, running this 4B parameter 3D generation model locally on a Mac was previously blocked by hard CUDA dependencies with no workaround.
- refulgentis 6mo agoSunday night, and its kinda cool idk man
- jmatthews 6mo agoWell done
- serf 6mo agorad. how long does output take? trellis is a fun model.
- shivampkumar 6mo agoi was able to get it in 3.5 mins from a single image on my 24gb m4 pro macbook I'm still working on this to try to replicate nvdiffrast better. Found an open source port, might look it tonight
- shivampkumar 6mo agothanks!
- post-it 6mo agoHow much RAM does this use? Only sitting on 8 GB right now, I'm trying to figure out if I should buy 24 GB when it's time for a replacement or spring for 32.
- Serhii-Set 6mo ago[dead]
- shivampkumar 6mo agoThe model needed about 15GB at peak during generation - the 4B model loads multiple sub-models (1.3B each for shape and texture flow). 8GB won't be enough, but both 24GB and 32GB both should be fine.
- post-it 6mo agoThanks! Could it conceivably load the sub-models in series rather than parallel? 8 still won't be enough but I wonder if those with 16 could eke something out.
- shivampkumar 5mo agoIn theory yes - the pipeline already does this to some extent with its low_vram mode, offloading models to CPU between stages. The challenge at 16GB is that even a single 1.3B sub-model at fp32 plus activations can push past what's available after macOS takes its share. Someone on an M1 iMac with 16GB did get geometry generation working tho (issue #5 on the repo), so 16GB is probably possible. 24GB gives comfortable headroom though.
- vrr044 6mo ago[dead]
- jiexiang 6mo ago[dead]
- petargyurov 6mo agoThis is fantastic, great work. I will attempt to run it on my 16GB M1 but I doubt it'll run. Out of curiosity, how did you go about replacing the CUDA specific ops? Any resources you relied on or just experience? Would love to learn more.
- sebakubisz 6mo agoThis is the kind of porting work I always hope for when I see a CUDA-only release. Have you thought about publishing the gather-scatter sparse 3D convolution and SDPA attention swaps as a standalone toolkit or writeup? A lot of folks running models locally on Apple Silicon hit the same wall with flash_attn, nvdiffrast, and custom sparse kernels and end up redoing the same work.
- shivampkumar 6mo agothat makes so much sense...I am exploring if I can find someone who has done this well...If not I'll try to do it myself.
- techpulselab 6mo ago[dead]
- sergiopreira 6mo agoMost 'runs on Mac' ports are a wrapper around a cloud call or a quantized shell of the original model. Going after the CUDA-specific kernels with pure-PyTorch alternatives is the kind of work that ages well, because the next CUDA-locked research release is three weeks away. One question: how much of the gather-scatter sparse conv is reusable for other TRELLIS-like architectures, or is it bespoke to this one?
- shivampkumar 5mo agoThe gather-scatter sparse conv should be fairly generic. Any model using 3x3x3 or 5x5x5 sparse convolutions on voxel grids could use it directly. The main thing that's TRELLIS-specific is the neighbor cache key format, but that's a few lines to adapt. The SDPA attention swap is even more reusable - it's just padding variable-length sequences into batches and calling torch.nn.functional.scaled_dot_product_attention.
- antirez 6mo agoGreat. Potentially can go much faster rewriting it in terms of Metal shaders.
- shivampkumar 5mo agoAgreed...I've been adding some like mtlgemm, mtldiffrast from other contributors already
- drbscl 6mo agoDoes it support multi-view input?
- shivampkumar 5mo agoNot currently - TRELLIS.2 is single-image input only AFAIK
- drbscl 5mo agoAh, it’s still very useful though, thanks for the port!
- takahitoyoneda 6mo ago[dead]
- strimoza 6mo agoCool project. I've been working on something similar in spirit — a personal video cloud (strimoza.com) — and the hardest part was also getting local playback to work reliably without internet. How are you handling memory pressure on the M-series chips with larger models?
- pcoyne 6mo agoI wonder how these outputs compare between Apple Sharp https://github.com/apple/ml-sharp https://github.com/apple/ml-sharp No matter what it is cool seeing so much them work on different devices
- vunderba 6mo agoThey’re completely different projects. ML Sharp is designed to create a 3D Gaussian representation of a depicted scene and output a splat file. Trellis generates model files (GLB I believe) for use with 3D modeling applications.
- paolatauru 6mo agosolid port. the sdpa swap for sparse attention — did you notice a meaningful quality difference, or is it basically equivalent to the cuda version? curious if the pure-pytorch path added any noticeable latency hit on the m3 max
- Olivia_Pan 6mo ago[dead]
- Kush-7575 5mo agoThis is very cool!
- deleted 5mo ago[deleted]
- egore911 5mo ago[dead]