Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
celrod
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
12 ms
·
121.
▲
by
celrod
4y ago
Regarding the uniformity... I started using tiling window managers a few years ago. These took a little configuring. I always had one monitor with an emacs session, and then another with an internet browser, and a terminal running tmux, whi
122.
▲
by
celrod
4y ago
I've noticed that R often defaults to much higher tolerances than Julia, even when it's wrappers to the same C library, like cubature. R cubature[0]: 1e-5 Cubature.jl [1]: 1e-8 The difference for NLopt in R vs Julia is smaller. `
123.
▲
by
celrod
4y ago
Perhaps tree sitter can help the syntax highlighting? https://github.com/latex-lsp/tree-sitter-latex
124.
▲
by
celrod
4y ago
I just tried Python, and it behaves like Julia. So I agree the two reasonable behaviors are: 1. Chain the operations like Julia/Python 2. Throw an error, ideally with an informative message.
125.
▲
by
celrod
4y ago
> I meant that the intent is obviously the chained comparison. Yes, "Is `b` between `a` and `c`?" is a natural and common question to ask. FWIW, Julia does lower `a < b < c` as a chained comparison: julia> Meta.@lo
126.
▲
by
celrod
4y ago
True. I have never gone down the Gentoo rabbit hole. Might be fun to try sometime, but I'd seriously doubt that the time spent compiling would be won back from better performance. Clear Linux is probably a more practical alternative. I
127.
▲
by
celrod
4y ago
You can use `-mprefer-vector-width=512` to use 512 bit vectors, or if you want a particular function to use 512, you could try the min-vector-width attribute: https://clang.llvm.org/docs/AttributeReference.html#min-vec
128.
▲
by
celrod
4y ago
How much code is compiled with `-march=native` or function multiversioning? I would guess the percentage is relatively small, at least when it comes to distributed binaries. Compiler autovectorizers also aren't very good at producing f
129.
▲
by
celrod
4y ago
FWIW, I also wrote a lot of C++ for the first time last year, and found the lack of > package management, one-liner built-in toolchains, built-in testing and build system To be by far the least pleasant things about the language, especia
130.
▲
by
celrod
4y ago
Note that `assert`s are disabled if you define the macro `NDEBUG`, e.g. https://godbolt.org/z/hMWo8KM7q CMake defines the macro in release builds: https://github.com/Kitware/CMake/blob/e1
131.
▲
by
celrod
4y ago
> yet I realized yesterday that I have no clue what the difference is between signed and unsigned subtraction on x86_64 It's two's complement. There is no difference between signed and unsigned for addition, subtraction, or mul
132.
▲
Optimizing a symmetric quadratic form in Julia
(spmd.org)
4 points
by
celrod
4y ago
|
0 comments
133.
▲
by
celrod
4y ago
I'd read before that many mammals lose more weight than bears during hibernation, but the weight they lose is a much more even distribution of different body tissues. There's a lot of research on how bears are capable of burning a
134.
▲
by
celrod
4y ago
Aside from the built in docs, mastering emacs is quite beginner friendly. It's aimed at people with prior programming experience, but not any emacs experience. You can get a few chapters for free.
135.
▲
by
celrod
4y ago
I'm young (early 30s) but have a strong family history of late onset Alzheimers (my father, and both of my mother's parents). I also get cold sores on occasion. I've brought this -- and the HSV-Alzheimers connection -- up to
136.
▲
by
celrod
4y ago
I use a treadmill desk and am a big fan. It's nice to walk for a few hours a day while on the computer. Stephen Wolfram takes this up another level by also walking outside for a while in nature with a laptop holder[0]. Anecdotally, he
137.
▲
by
celrod
4y ago
I do think that Genoa's approach is a reasonable one. I'd like to see one of Gracemont's successors doing the same. Maybe we'll even see quadruple pumping for AVX-512 some day? I'll be impressed if/when an Atom
138.
▲
by
celrod
4y ago
Good analysis. It's also worth pointing out that this is for 2x 512 bit FMA, which is more than client Ice/Tiger/Rocket lake or Zen4 have. Personally, I bought HEDT (Skylake-X and Cascadelake) because I wanted 2x 512 bit AVX5
139.
▲
by
celrod
4y ago
Smarter bump allocator implementations, like the one in llvm, will allocate a new blob instead of bumping past the end of an old one.
140.
▲
by
celrod
4y ago
I recently wanted to run some Python benchmarks, and was surprised to find that some major Python packages (i.e. Numba) didn't work on the latest Python release. In Julia, all open source packages that pass tests on the latest release
141.
▲
by
celrod
4y ago
That is a closed issue from a Julia package. I'm sure I could go to popular packages from any language and find bug reports.
142.
▲
by
celrod
4y ago
I'm guessing that 20% is still enough for your zen4 to be faster than raptor lake running the avx2 path, while also probably using less power.
143.
▲
by
celrod
4y ago
I use emacs. I have little desire to customize configs and spend hardly any time on it. Checking now, I haven't modified it since June, which means I haven't even installed any new packages (the equivalent of VSCode extensions) si
144.
▲
by
celrod
4y ago
LUTs at least do well in microbenchmarks, but I do worry that they may do comparatively much worse in real code. That said, that's another advantage of small tables using vpermi2pd. The Julia/base implementations of log and exp bo
145.
▲
by
celrod
4y ago
The Intel optimization manual has a fun example where they use vpconflict for vectorizing sparse dot products: https://github.com/intel/optimization-manual/blob/main/chap1... I benchmarked it on Intel, a
146.
▲
by
celrod
4y ago
Yeah, the claim was that this is why it hit higher clock speeds. The front end will be hard pressed to hit/maintian 4 IPC, while 2 IPC is much easier.
147.
▲
by
celrod
4y ago
Looks like SIMD implementations that use LUTs should favor small tables that fit in registers and use `vperm2ipd` as look ups over larger tables + gather. With 64 bits, you still get a LUT size of 16 (shuffle indexes into two 8xdouble vecto
148.
▲
by
celrod
4y ago
It's my first real foray into C++ (most of my experience is with Julia) so code quality/awareness of idioms, etc are probably lacking, but it's using C++20 and I'd appreciate pointers if anyone has any: https:/
149.
▲
by
celrod
4y ago
If you're using Julia and want something like this, you can try `LoopVectorization.vfilter`. It lowers to LLVM intrinsics that should do the right thing given AVX512, but performance is less than stellar without it IIRC. May be worth t
150.
▲
by
celrod
4y ago
You can configure throttling in the bios if you have Skylake-X. You can set it to whatever you want. The caveat being it won't necessarily be stable (depending on voltage) nor will your cooling necessarily be able to handle it. My 109
More ›