Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
cliffc
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
cliffc
2y ago
C2/HotSpot uses a MemBarrier / MemFence node type for these. Serializes the memory ops as needed. Used to implement the Java memory model on all hardware. Sometimes (x86) the ops encode as (empty), sometimes its a Real Fence op
2.
▲
Simple – Sea of Nodes
(github.com)
3 points
by
cliffc
2y ago
|
4 comments
3.
▲
by
cliffc
2y ago
Simple ( https://github.com/SeaOfNodes ) is a compiler tutorial featuring the Sea-of-Nodes IR, with ports in Java (@cliffclick), Go (@yardenlaif), Rust (@RobertObkircher), and C++ (@Hels15), with help from @XmiliaH, @ThaliaAr
4.
▲
by
cliffc
10y ago
It was cheaper to run (power, heat, space) than the equivalent pile of X86's of that era and for years to come. Each box cost between $250k and $750k depending on cores & memory. You didn't buy one unless you had a specific
5.
▲
by
cliffc
10y ago
I have board with 2 Vega chips in my living room - 24 cores each; each a full 64-bit RISC with 32 registers and ie754 FP math, read & write barriers in hardware; inline-cache virtual calls in hardware, and transactional memory. Later c
6.
▲
by
cliffc
10y ago
And the actual data is stored in giant byte[] (hidden behind the API), so the GC costs are near zero. Cliff
7.
▲
by
cliffc
10y ago
Fun stuff I've been doing with the H2O project is basically using nearly-pure Java (some Unsafe) to hold onto numbers with better efficiency than e.g. int[], and giving out an easy-enough-to-use API for writing parallel & distribut
8.
▲
by
cliffc
10y ago
Nearly all the work you allude to is done in other threads - which indeed consume machine resources (CPU cycles, memory bandwidth). If your application does not burn all cores/bandwidth then the GC work is all done on the idle/sp
9.
▲
by
cliffc
10y ago
Sorta kinda all of the above. C4 the algo has no pauses, but individual threads stop to crawl their own stacks. i.e., threads stop doing mutator work, but only 1-by-1, and only for a very short partial self-stack crawl. C4 the impl I belie
10.
▲
by
cliffc
10y ago
Azul hardware had a very large core count of low-performing cores. If you had enough parallelism then it was hard to beat - but most applications didn't have enough parallel work, so the market wasn't big enough
11.
▲
by
cliffc
10y ago
No change to the X86, instead user-mode TLB handler from RedHat allows ptr-swizzeling, that plus some careful planning and the read barrier fell to 2 X86 ops - with acceptable runtime costs. Cliff
12.
▲
by
cliffc
10y ago
Right when Java was taking off, there were a bunch of research proposals for better over-the-wire executable formats (mostly funded by microsoft). Designs where you could convert the bytecodes to machine code as fast as L1 could take it wi
13.
▲
by
cliffc
10y ago
The X86 memory model is "conservative" in the sense that it puts strong restrictions on the order of observed loads and stores between cores. Azul and IA64 did not. Hence "sloppy" code - but technically correct (at lea
14.
▲
by
cliffc
10y ago
It means what it says it means? "fragile" code is code that can't be touched because touching it causes it to "break" - exhibit bugs that can't easily be fixed. "fluffy" code is bulky code: code that
15.
▲
by
cliffc
10y ago
Azul's GC required a read-barrier (some instructions to be executed on every pointer read) - which cost something like 5% performance on an X86. In exchange max GC pause time is something in the low microsecond range (I helped imp
16.
▲
by
cliffc
10y ago
Thanks for the nod! Cliff