Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
chillee
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
61.
▲
by
chillee
3y ago
120M in annual revenue for 100 employees is not that high. That's only about 1.2M in revenue per employee. Netflix's main business makes 2.4M per employee (31.6 billion for 12k employees), and given the nature of such businesses,
62.
▲
by
chillee
3y ago
ChatGPT-4 seems to do fine on this for me. https://chat.openai.com/share/c5cc8cb6-ebb5-45eb-9476-ef85a6...
63.
▲
by
chillee
3y ago
It's confusing, there's OpenAI Triton (what you're thinking of) and Nvidia Triton server (a different thing).
64.
▲
by
chillee
3y ago
1000 Mana costs 10$ USD, are you using a different currency?
65.
▲
by
chillee
3y ago
No, AIT is a runtime - it exports a .so that you can load from C++ and call how you like.
66.
▲
by
chillee
4y ago
> For simple stuff, we can compare JAX to PyTorch on a 4090, and JAX seems faster by 10-50%. It's way way way faster on TPU. What benchmarks are you looking at here?
67.
▲
by
chillee
4y ago
I retracted my specific criticism about not comparing to a regular matmul, but I still keep my criticism about having weak baselines :)
68.
▲
by
chillee
4y ago
Quoting myself on twitter: https://twitter.com/cHHillee/status/1577713102434361344 I'm quite suspicious about their hardware benchmarks. They're not writing custom kernels, they're relying on a grap
69.
▲
by
chillee
4y ago
To clarify, by "compilers" I mean "deep learning compilers". There's many different paths to optimizing compilers folks use with PyTorch. One with close integration is NVFuser (see https://www.reddit.com&
70.
▲
by
chillee
4y ago
Ok, I work on PyTorch, so probably should clear up some misconceptions in this thread. 1. In PyTorch (and other array programming libraries like Numpy), the operations being passed around are tensors/arrays (i.e. large chunks of memory
71.
▲
by
chillee
4y ago
> That app usage gets shared with Facebook's app. That's not really the primary mechanism, as far as I know. My understanding of the flow is this: Before: 1. You browse Facebook. 2. Companies who want to sell a guitar tuner app
72.
▲
by
chillee
4y ago
> Neither OpenAI nor FAIR ever has the top score on anything unless Google delays publication. This is ... very incorrect. I am very certain (95%+) that Google had nothing even close to GPT-3 at the time of its release. It's been 2
73.
▲
by
chillee
4y ago
Yeah, my main point is that I don't think the issue is memory movement. PyTorch/Tensorflow do care a lot about memory movement, as memory movement doesn't stop being an issue with larger networks.
74.
▲
by
chillee
4y ago
Anything that’s not an actual computational kernel (python interpreter, Pytorch dispatcher, etc.) See the flame graph here in the overhead section: https://horace.io/brrr_intro.html
75.
▲
by
chillee
4y ago
In "small" setups, you're actually more likely to be overhead-bound if anything (especially on CPU).
76.
▲
by
chillee
5y ago
Oh, wrapper overhead of Python is certainly annoying. There's a couple things that ameliorate that. For one, in many situations, it's possible to "trace" out the Python operations. The second is that, as in my blog post,
77.
▲
by
chillee
5y ago
Author of the original blog post here. I agree, the point isn't really about how Python is super slow compared to C++, or any language. The point is about how easy it is for overhead to creep in when you're working with these unbe
78.
▲
Making Deep Learning Go Brrrr from First Principles
(horace.io)
3 points
by
chillee
5y ago
|
0 comments
79.
▲
by
chillee
5y ago
If you’re massively dominated by overhead I can see it. I’ve definitely done microbenchmarks before where compilation gets you 1000x improvement.
80.
▲
by
chillee
5y ago
It's not quite a fair comparison, since Numpy is running with float64 while Jax is running with float32. If you fix the benchmarks then looks like this 5 loops, best of 5: 99.2 ms per loop 10 loops, best of 5: 114 ms per loop 10 loops,
81.
▲
by
chillee
5y ago
This is not comparable with ML frameworks (Pytorch or Jax) - its closest analogue is Halide (which they mention in the paper).
82.
▲
by
chillee
5y ago
I think it's just a subtle way of saying "Implementing these rules in software might take a while, but you can use a physical chess board to try these variants"
83.
▲
by
chillee
5y ago
Same :P I'm actually responsible for one of these ( https://github.com/pytorch/pytorch/issues/70607 ), but it's a typo in a list of tests to skip.
84.
▲
by
chillee
5y ago
You might be interested in the upcoming AMD Mi200 GPUs, which have 96 teraflops of fp64 performance.
85.
▲
by
chillee
5y ago
Tf.function is largely morally equivalent to Torchscript, which he does discuss.
86.
▲
by
chillee
5y ago
> As a researcher in RL & ML in a big industry lab Is that big industry lab Google or Deepmind? haha
87.
▲
by
chillee
5y ago
Cool, thanks! Yeah, stack-based PRs are awesome - excited to perhaps try it out.
88.
▲
by
chillee
5y ago
How is this different from ghstack? https://github.com/ezyang/ghstack (which is what Edward Yang for PyTorch developers to mimic the stacked workflow, although it works with other repos).
89.
▲
by
chillee
5y ago
> What is "conv and batch norm fusion"? How does FX help with any of this? Essentially, during inference, batch norm is simply a multiply and add operation. If this occurs after a convolution, then you can simply fold (i.e. &qu
90.
▲
by
chillee
5y ago
Built with torch.FX!
More ›