Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
marmaduke
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
25 ms
·
91.
▲
by
marmaduke
3y ago
Just curious, I had a 2020 SE daily and never had a OOM, how do you do it?
92.
▲
by
marmaduke
3y ago
> Why should I even bother existing? you should bother existing to push back against economic forces that push us to work too much and feel bad that we aren't working more. you should exist to teach your colleagues that taking the
93.
▲
by
marmaduke
3y ago
I multiply every estimate by at least 3x and often add an order of magnitude, especially when I haven't done it before. when it takes less than I guessed, I use the time to resolve tech debt (or maybe leave early to enjoy the big blue
94.
▲
by
marmaduke
3y ago
In another article about high speed rail, someone pointed out there's a big difference in France: once the govt votes the build, it becomes law and can't be preempted whereas the US, any entity can pause infra projects with a suit
95.
▲
by
marmaduke
3y ago
Very nice! I just installed, I hope to eventually contribute down the line, especially in terms of custom operators. They weren't even document until recently, and there's still quite some work to add them.
96.
▲
by
marmaduke
3y ago
there's a paper about it that I just found, enjoy https://arxiv.org/pdf/2301.13062
97.
▲
by
marmaduke
3y ago
Because they use a funny format (BCOO). I'm not mocking, it must be a solid choice for some reasons, like sparsification or other fancy stuff. But for large and even with batches (ie multiply with tall dense matrix), it doesn't
98.
▲
by
marmaduke
3y ago
I think this is up to XLA to handle not Jax. The whole selling point in TF of the tf.function decorator (which uses XLA underneath as well) is that it fuses arithmetic to lower launch count.
99.
▲
by
marmaduke
3y ago
Oops so sorry. But this is recent isn't it? I thought it was actually due to XLA/Bazel not supporting it?
100.
▲
by
marmaduke
3y ago
If you like functional languages, then Jax will fit better for you. It provides a bunch of function transformations to implement eg grad, JIT etc.
101.
▲
by
marmaduke
3y ago
They don't build on Windows at all, as well.
102.
▲
by
marmaduke
3y ago
None of these libraries unfortunately allows making good use of CPU vectorized units. Xla might produce some SIMD code but it pales in (performance) comparison to routines written explicitly for SIMD on GPU. ISPC is a good example of th
103.
▲
by
marmaduke
3y ago
Jax JIT of scan is fairly good, so loops aren't as slow as you'd expect.
104.
▲
by
marmaduke
3y ago
Nope it's super slow for large sparse matrices. It's even faster to use generic scatter/gather to implement some, instead of that built in thing.
105.
▲
by
marmaduke
3y ago
WSLg works really well, even for OpenGL 3/4, CUDA, and when connected via RDP as well.
106.
▲
by
marmaduke
3y ago
Was the wind reference a pun? The strongest winds in southern France are called mistral.
107.
▲
by
marmaduke
3y ago
I don't recall SQS taking anytime to set up queues, at least when using Celery (python task queue). IIRC you just name it in your code and use it.
108.
▲
by
marmaduke
3y ago
I had the first SE, and it initially handled mixed language text correction phenomenally, but went downhill with each update, which was really puzzling
109.
▲
by
marmaduke
3y ago
Web peasants will just have/find/create a different set of problems to solve, no?
110.
▲
by
marmaduke
3y ago
The Rwkv family of models qualifies, since it computes like a recurrent network at inference time.
111.
▲
by
marmaduke
5y ago
Computational neuroscience, but since your username is earth science I would mention one of the core algorithms which is quite a bit faster is the spherical harmonic transform.
112.
▲
by
marmaduke
5y ago
Memory bandwidth is still 76 GB/s. One reason I’m moving science workloads to M1 is for the cache line size and memory bandwidth, which bump by factor 2-3x the speed.
113.
▲
by
marmaduke
5y ago
Isn’t a court ruling an opinion and not fact? Is it necessary to oppose one opinion not with another but a fact?
114.
▲
by
marmaduke
5y ago
It’s not irrelevant for machine learning in general, e.g. Stan (Bayesian/MCMC) is fp64 only, so its OpenCL offloading is more effective where fp64 is better supported. It’s less relevant for the kinds of deep neural net work enabled by
115.
▲
by
marmaduke
5y ago
Wt is a similar idea for C++/Java https://www.webtoolkit.eu/wt/ though seemingly more complete.
116.
▲
by
marmaduke
5y ago
At least for France, there is less to recommend doing a startup than in the US. The social security regime is worse (than for others eg employees, govt workers), the tax schemes are complex and penalizing at the low end, and the amount a c
117.
▲
by
marmaduke
5y ago
Stand-alone is a very useful concept. I don’t like deploying Python stacks much. Wouldn’t that additionally mean you could target CL, CUDA or Sycl variants of C?
118.
▲
by
marmaduke
5y ago
ASM makes sense when the time spent in a specific routine exceeds the time it takes to write the ASM, which makes a lot of sense for Blas, less so for other HPC yet speculative or less fundamental projects. Cvodes for instance doesn’t need
119.
▲
by
marmaduke
5y ago
From (1) > Some services … share your WLAN access data with your facebook contacts Is that true? There’s no citation but seems a lot more flagrant a violation than I’d expected from Windows.
120.
▲
by
marmaduke
5y ago
Maybe you could explain why a non commercial license would not be open, for anyone not looking to turn a profit off of the author’s work?
More ›