Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
celrod
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
20 ms
·
181.
▲
by
celrod
5y ago
Yeah, I really like that Julia and LLVM allow applying it on a per-operation basis. Because most of LLVM's backends don't allow for the same level of granularity, they do end up propagating some information more than I would like.
182.
▲
by
celrod
5y ago
> So yes, that means that you become responsible to either guarantee that erroneous results do not matter or that you will take care to always check the ranges of input operands, as "okl" has already posted, to ensure that no o
183.
▲
by
celrod
5y ago
Golden cove core highlights[1] include a 6 wide decoder (up from 4; M1 is 8), reorder buffer of 512 (up from 352 in Ice/Tiger/Rocket lake, or 224 in Skylake; M1 is estimated around 650), and 12 execution units (up from 10 in Ice&#
184.
▲
by
celrod
5y ago
From context, do you mean "It definitely does support 1 hand lid operation" instead of "It definitely does not"?
185.
▲
by
celrod
5y ago
More than 20% for many inland grizzly bears, e.g. Yellowstone: Approximately 45 ± 22% (equation image ± SD) of the assimilated nitrogen consumed by male grizzly bears, 38 ± 20% by female grizzly bears, and 23 ± 7% by male and female bla
186.
▲
by
celrod
5y ago
I assume they were responding to: And if you have fighter jets buzzing by, the least of your problems is noise. I.e., they meant their primary concern with pilots familiarizing themselves with the terrain is the noise, rather than an
187.
▲
by
celrod
5y ago
On paper, saphire rapids is a large step forward. Compute tiles to offer larger core counts (while maintaining a monolithic topology), 512 reorder buffer entries (up from 354), 6 decode (up from 4), 2 more execution units, chips with HBM me
188.
▲
by
celrod
5y ago
Assuming die area is the constraint, the choice is between 4 firestorm + 4 icestorm and 5 firestorm, because a firestorm core is 4x bigger than an icestorm core. Or between 8 firestorm and 7 fire + 4 ice (or 6 fire + 8 ice).
189.
▲
by
celrod
5y ago
By "compile time", I meant the first time you call the function with a given type signature. Also, that comment is saying "compile time type" does not exist. I don't know C++, so I cannot comment on it, but from the
190.
▲
by
celrod
5y ago
In practice, Julia's multiple dispatch is almost always devirtualized. That is, dispatch is resolved at compile time. Generic code relies on specialization instead of dynamic dispatches to be generic with respect to input types. That i
191.
▲
by
celrod
5y ago
https://godbolt.org/z/v67Tc3jEc gcc and icc will use fma under `-O3`, but Clang won't.
192.
▲
by
celrod
5y ago
-t6 would give 6 threads. FWIW, the Ryzen 2600 with 6 threads took 6 ms, so the 3600's floating point and vector improvement is quite evident from your timing. (The 2600 emulates 256-bit operations with 2x 128-bit, while the 3600 has f
193.
▲
by
celrod
5y ago
exp(x) = exp(x.real()) * sincos(x.imag()) Anyway, it is the same algorithm as long as `x` is real. The reason for manually inlining exp(::Complex) and manually decomposing the complex number into its real and imaginary parts is that
194.
▲
by
celrod
5y ago
Would be good to confirm the number of threads being used by both, but it could also be that pythran is using a better `sincos` implementation. On very old hardware (which it would have to be if multithreaded), the exp implementation might
195.
▲
by
celrod
5y ago
Also, note that Julia starts with only a single thread by default. You'd either need to add `-t4` when starting julia, e.g. `julia -t4`, or set the environmental variable `JULIA_NUM_THREADS=4` for 4 threads, for example.
196.
▲
by
celrod
5y ago
Did you start Julia with multiple threads? Julia unfortunately starts with only a single thread by default, and the slow time you reported (14 ms) makes it appear that this may have been the case. You can start `julia -t4` for 4 threads, fo
197.
▲
by
celrod
5y ago
> I'm too lazy to figure out how to install and benchmark Julia on my machine. Assuming you're eager to try once you've found out: You should be able to simply download and unpack a binary from: https://julialan
198.
▲
by
celrod
5y ago
Or much faster (at small sizes): https://github.com/JuliaLinearAlgebra/Octavian.jl
199.
▲
by
celrod
5y ago
`vsqrtpd` is an assembly instruction. `exp` and `sincos` are implemented in Julia. If you want to use a library to speed up the Julia code, checkout https://discourse.julialang.org/t/i-just-decided-to-migrate-... (but
200.
▲
by
celrod
5y ago
Depending on the scaling governor you're using, you should be able to adjust behavior from the command line: https://www.kernel.org/doc/html/v4.12/admin-guide/pm/cpufreq... https://w
201.
▲
by
celrod
5y ago
Growing up, what my late father probably wanted most from me is for me to find a project of my own. When I was in high school, he once threatened me with "get a life, or I will get you one". Engines, and especially motorcycles, we
202.
▲
by
celrod
5y ago
Try Julia nightly (1.7): julia> @code_warntype not_type_stable() MethodInstance for not_type_stable() from not_type_stable() in Main at REPL[2]:1 Arguments #self#::Core.Const(not_type_stable) Locals o::Any Body::I
203.
▲
by
celrod
5y ago
FWIW, llvm-mca estimates 448 clock cycles per 100 iterations of the AVX2 loop vs 528 cycles for the AVX512 loop with `-mcpu=cascadelake`. That suggests the AVX512 loop should be about 2*(448/528)=1.85 times faster.
204.
▲
by
celrod
5y ago
p5 can do 512 bit operations, but not 256 bit, e.g. look at Skylake-AVX512 and Cascadelake (Xeon benched in the blog post was Cascadelake) ports for vaddpd: https://uops.info/html-instr/VADDPD_YMM_YMM_YMM.html Here is
205.
▲
by
celrod
5y ago
ISPC also has first class support for "(array of) structure of arrays", see: https://ispc.github.io/ispc.html#structure-of-array-types For example: soa<8> Point pts[...]; declares an array of struct o
206.
▲
by
celrod
5y ago
Kinesis Advantage 2: Likes: The dished profile of the keys, which more naturally reflects the arch your finger tips travel, reducing the overall need for movement. While common among split keyboards, I also really like the thumb keys. I use
207.
▲
by
celrod
5y ago
I'd also take a look at Images.jl: https://juliaimages.org/latest/examples/ Noise removal example: https://juliaimages.org/ImageFiltering.jl/stable/democards/d...
208.
▲
by
celrod
6y ago
Yes, and 32 GB seems rather limited for a 48 core chip. While not a "general purpose" CPU, this HPC CPU does show IMO that it's at least possible of making a CPU that does compete in some aspects with what GPUs do. Although a
209.
▲
by
celrod
6y ago
FWIW, the A64FX has 1TB/s bandwidth because it has 32GiB of HBM2.
210.
▲
by
celrod
6y ago
I like having the familiar emacs keybindings to navigate the terminal, e.g. to be able to scroll quickly, search, and copy/paste without needing to take my hands off the keyboard. I'm sure there's some terminal emulator that
More ›