Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
areddyyt
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
areddyyt
2y ago
Our CPU implementation for X86/AMD64 utilizes AVX-512 or AVX-2 instructions where possible. We're experimenting with support for ARM with NEON.
2.
▲
by
areddyyt
2y ago
100%. When performing performance optimization on CPUs, I was impressed with Intel's suite of tools (like VTUNE). NVIDIA has some unbelievable tools, like Nsys and, of course, its container registry (NGC), which I think surpasses even
3.
▲
by
areddyyt
2y ago
I should note that our linear layers are not the same as Microsoft's, in fact, we think Microsoft made a mistake in the code they uploaded. When I have time later today, I'll link to where I think they made a mistake. I've be
4.
▲
by
areddyyt
2y ago
I don't think I ever implied we started this for money. We started working on the technology because it was exciting and enabled us to run LLMs locally. We wouldn't have started this company if someone else came along and did it,
5.
▲
by
areddyyt
2y ago
I think there is a comment somewhere here where I comment on NVIDIA, but I think NVIDIA is the best hardware company for making good software. We had a very niche software issue for which NVIDIA maintained open-source repos. I don't th
6.
▲
by
areddyyt
2y ago
We were waiting for a Bitnet-based software and hardware stack, particularly from Microsoft, but it never did. We were essentially nerd-sniped into working on this problem, then we realized it was also monetizable. On a side note, I deeply
7.
▲
by
areddyyt
2y ago
We do quantization-aware training, so the model should minimize the loss w.r.t. the ternary weights, hence no degradation in performance.
8.
▲
by
areddyyt
2y ago
There was another founder that said this exact same thing. We'll definitely look into it especially as we train more ViTs.
9.
▲
by
areddyyt
2y ago
Funnily enough, our ML engineer, Eddy, did a hackathon project working with Procyon to make a neural network with a photonic chip. Unfortunately, I think Lightmatter beat us to the punch. Edit: I don't think the company exists in its c
10.
▲
by
areddyyt
2y ago
Have you sat in on my conversations with my cofounder? The end plan is to have a single chip and flush all weights onto the chip at initialization. Because we are a single line of code that is Torch compatible (hence HF compatible), every o
11.
▲
by
areddyyt
2y ago
This seems super cool. I'll have my cofounder look into it :)
12.
▲
by
areddyyt
2y ago
We don't achieve peak compression efficiency because more complex weight unpacking mechanisms kill throughput. To be more explicit, the weight matrix's values belong to the set of -1, 0, and 1. When using two bits to encode these
13.
▲
by
areddyyt
2y ago
Thank you, and good catch. We recently acquired deepsilicon.com, and it looks like the forwarding hasn't been registered yet. abhi@deepsilicon.net should work.
14.
▲
by
areddyyt
2y ago
We actually were thinking about this to flush the weights in at initialization
15.
▲
by
areddyyt
2y ago
It's always possible, but transformers have been around since 2017 and don't seem to be going anywhere. I was bullish on Mamba and researched extended context for structured state-space models at Dartmouth. However, no one cared.
16.
▲
by
areddyyt
2y ago
Video cropping issues should be fixed!
17.
▲
by
areddyyt
2y ago
We've spent a lot of time thinking about these things, in particular, the 3Ps. Part of making the one line of code work is addressing programmability. If you're on Jetson, we should load the CUDA kernels for Jetson's. If you&
18.
▲
by
areddyyt
2y ago
You're absolutely right about mobile devices (Apple, Google, etc.). However, most companies, with the exception of Tesla, do use Nvidia for edge computing capabilities. We know for a fact that most of the automotive industry uses autom
19.
▲
by
areddyyt
2y ago
Oops, good catch. Will re upload shortly.
20.
▲
by
areddyyt
2y ago
Great question. So a little bit of background about quantization (apologies if you are already familiar). There are two types of quantization (generally), post training quantization (PTQ) and quantization aware training (QAT). PTQ almost al
21.
▲
by
areddyyt
2y ago
Thank you! CUDA and Nvidia are practically impenetrable on the server side. To be very concrete, we did training for our models on AWS with parallel cluster. We used P5 instances (8xH100) that were scheduled with SLURM. A problem we ran int
22.
▲
by
areddyyt
2y ago
We are not under the illusion these markets are easy to enter. Still, we believe providing an effortless and compatible experience for edge ML computing is a strong competitive advantage. We have not met anyone who likes using Jetsons yet,
23.
▲
by
areddyyt
2y ago
The non-linear layers, particularly the softmax(QK^T), will be crucial to getting ultra-low latency and high throughput. We're considering some custom silicon just for that portion of every transformer block
24.
▲
by
areddyyt
2y ago
Agreed. We don't plan on making hardware until there is enough demand from customers to make it economically viable.
25.
▲
by
areddyyt
2y ago
In general, Jetson has quite a large market. Vehicle companies use automotive-rated Jetson Orins, and defense companies also use Jetson Orins to power ML applications on the edge (Anduril). Many of the companies we currently talk to are rob
26.
▲
by
areddyyt
2y ago
We're targeting the edge market first, such as NVIDIA's Jetson line, because it's far less supported/focussed on. In our experience, whenever we did training runs on H100 clusters with x86, any pip package would be easil
27.
▲
Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
189 points
by
areddyyt
2y ago
|
79 comments