Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
celrod
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
14 ms
·
151.
▲
by
celrod
4y ago
Reminds of the Geifer grizzly. It raided human dwellings for food, signing its death warrant. It wore a radio collar, so it seems like it shouldn't have been a problem to track down. Yet, it evaded hunters for over a year, while still
152.
▲
by
celrod
4y ago
Julia's 1.8's load times are dramatically better than Julia 1.0's. It keeps getting a bit better with every release. But there is a project to cache native code (without needing to build a sysimage) that may make a huge diffe
153.
▲
by
celrod
4y ago
I'm working on loop modeling and optimization software in C++: https://github.com/JuliaSIMD/LoopModels/ This project was started in November of last year, and is still undergoing rapid development. C++ is cur
154.
▲
by
celrod
4y ago
I think slow aging/longevity are "hard". You need excellent repair mechanisms, and for nothing to go wrong for a very long time. Random mutations can easily cause problems/chop some time off your lifespan. The question i
155.
▲
by
celrod
4y ago
Many of us in the Julia community (myself included) take it very seriously and spend a substantial amount of our time working to mitigate it.
156.
▲
by
celrod
4y ago
I find `ccall` and especially `@ccall` easy to use, but thankfully haven't spent much time wrapping C libraries so I'd consider myself far from an expert in the matter. Creating mutable structs and `GC.@preserve`ing them is an eff
157.
▲
by
celrod
4y ago
matmul
158.
▲
by
celrod
4y ago
// Compiler doesn't make independent sum* accumulators, so unroll manually. // We cannot use an array because V might be a sizeless type. For reasonable // code, we unroll 4x, but 8x might help (2
159.
▲
by
celrod
4y ago
The bigger the vectors, the better the performance per watt.
160.
▲
by
celrod
4y ago
I am a big fan of the Kinesis Advantage 2's dished layout. I think that's a bigger deal for me than being split. The 2 unfortunately is not split, while the 360 is both split and dished. The thumb keys are also extremely important
161.
▲
by
celrod
4y ago
> But none of those, except for dex, has a solution for fusing kernels that rely on loops. The LV rewrite will. Some day, I'd like to have it target accelerators, but unlike fusion, I've not actually put any research/engin
162.
▲
by
celrod
4y ago
One of my intentions with the rewrite is to let `@turbo` to change the semantics of "unreachable"s, allowing it to hoist them out of loops. This changes the observed behavior of code like for i = firstindex(x):lastindex(x)+1
163.
▲
by
celrod
4y ago
Even when FMA is implemented in hardware, LLVM will generally use the software version when the arguments are known at compile time.
164.
▲
by
celrod
4y ago
We probably should have released the blog post version as 0.3 instead of 0.2.1. There were not any breaking changes, but enough fixes and additions (e.g., the convolutional layer and threading) that 0.2.0 isn't really comparable. Odds
165.
▲
by
celrod
4y ago
If you want the -Ofast equivalent in Julia, there is @fastmath. LoopVectorization.jl is substantially more powerful. It also has tons of limitations, but I am rewriting it to fix most of them. Disclosure: I'm the primary author of both
166.
▲
by
celrod
4y ago
Blis is very bad at small sizes, last I checked. Their approach was based on having a single microkernel, and setting up calls to it. Last I checked, blasfeo did not support AVX512, and this performed poorly on CPUs supporting it. I'
167.
▲
by
celrod
4y ago
SimpleChains does not call out to gemm. All layers are currently implemented as naive for loops. It uses LoopVectorization.jl to compile them, which does a good job leveraging AVX512 -- much better than llvm or gcc. It doesn't pack, wh
168.
▲
by
celrod
4y ago
SimpleChains.jl/LoopVectorization.jl author here, and a co-author of the blog post. I love working on performance. It's fun and challenging. Trying to best high scores (times/benchmarks) is fun and gamifies it. Personally, I
169.
▲
by
celrod
4y ago
Enzyme.jl works quite well (but the possibility of using it across languages is appealing).
170.
▲
by
celrod
4y ago
The Flux.jl example did this. A PR to the PyTorch example to do this would be welcome: https://github.com/chriselrod/LeNetTorch
171.
▲
by
celrod
4y ago
The performance advantage was still in the ballpark of 4-8x faster for training MNIST on the CPU, which while smaller than most networks people are training on their GPUs, still has more than 40 thousand parameters. For someone with a stati
172.
▲
by
celrod
5y ago
PumasAI - Scientist / Machine Learning - Globally Remote (although you should be welcome in the Somerville Julia Computing office) https://pumas.ai/company/machine-learning-scientist/ https://disco
173.
▲
by
celrod
5y ago
Yeah, 2 p5 (two uops on port 5) for a compressed store vs 1 p05 (1 uop on either port 0 or port 5) for a masked move. For throughput's sake, shifting the pointer is better.
174.
▲
by
celrod
5y ago
AVX512 has a compressed store which might be a bit easier than a normal masked store?
175.
▲
by
celrod
5y ago
Skylake-X has 168 vector registers. Ice and Tiger lake have 224 The connection to the 32 micro architectural registers is quite loose.
176.
▲
by
celrod
5y ago
In that particular benchmark (3d particle movement), the 11700K performed about 4x better than the AMD 5900X. Performance/watt clearly wasn't suffering. Perhaps it could downlock more to address the wattage while still coming well
177.
▲
by
celrod
5y ago
I don't think latency is multiple dispatch's fault. Certainly, compilation caching failures aren't.
178.
▲
by
celrod
5y ago
With 128 bit vectors, the M1 is less of a vector computer than the x86 competition, which all have 256-bit or 512-bit vectors. As you note, the firestorm core's real strength is in its scalar execution (and when it comes to the vector
179.
▲
by
celrod
5y ago
I've had aphantasia all my life, although my dreams are still visual, as though they're real. The article claims it is more common among academics and computer scientists, so perhaps it's fairly common here. Perhaps it'
180.
▲
by
celrod
5y ago
Hmm, uiCA results: xlatb: https://bit.ly/3cyBNN5 sequence: https://bit.ly/3nCmVTX xlatb is looking better here. There are also some front end concerns that may favor xlatb, in particular if it's friend
More ›