Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
PixelOfDeath
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
61.
▲
by
PixelOfDeath
6y ago
I hat very good experience with using buffer variables to copy-"prefetch" unpredictable costly fetches. E.g. from cache lines that get touched by several cores for communication. And I only actual use them after one iteration of w
62.
▲
by
PixelOfDeath
6y ago
Should electro bicycles not make such issues easier to handle?
63.
▲
by
PixelOfDeath
6y ago
x86_64 malloc already uses 16 byte alignment and portions to support SSE commands. So why not go all the way and allocate full cache lines? This also prevents accidental false sharing between different malloc pointers.
64.
▲
by
PixelOfDeath
6y ago
ME TO! Especially because of the x86 oligopoly I would think that Arm is so much more important as an ecosystem.
65.
▲
by
PixelOfDeath
6y ago
In Germany and surrounding countries it is very common to eat raw minced pork meat (Mett) on bread/buns.
66.
▲
by
PixelOfDeath
6y ago
> keeping the natural local time Fuck that! I want a single unified global time!
67.
▲
by
PixelOfDeath
6y ago
Isn't AVX512 basically cacheline-instructions?
68.
▲
by
PixelOfDeath
6y ago
The K5 was actually a RISC CPU with an x86 decoder on top. So we can count that one, too?!
69.
▲
by
PixelOfDeath
6y ago
x86 guarants that memory writes to different locations are seen in order by other cores. If you write memory first to address A and then to B. Another core never can see the B change without also seeing the A change at any moment in time. B
70.
▲
by
PixelOfDeath
6y ago
> 1. This would dramatically increase cache line size. I don't have data, but I assume this would generally be bad. Why would it change cache line size? GPUs also use cache lines in the range of 32-128 byte?! I think that is indepen
71.
▲
by
PixelOfDeath
6y ago
GPUs can hide memory latency very well, because they are basically SMT on steroids. (Imagine instead of 2 threads per "core" you have 32-64 threads) But they are starved of memory bandwidth! And the lower latency memory CPUs prefe
72.
▲
by
PixelOfDeath
6y ago
AI scales better to how many jets you want to use. AI improvements can also instantaneously be rolled out to every jet. There is no need for continues live communication in "hot" combat. Less video leaks when they kill a bunch of
73.
▲
by
PixelOfDeath
7y ago
I prefer Hitchens answer to the question: "We have free will, we have no choice."
74.
▲
by
PixelOfDeath
7y ago
You just set your chosen segment register(s) in protect mode, jump back to real mode, and then can use this segment register(s) in combination with the 32bit prefix to make a 32 bit memory access. All other segment registers are untouched a
75.
▲
by
PixelOfDeath
7y ago
Even if it may still take some years, I have really high hopes for chiplet GPUs.
76.
▲
by
PixelOfDeath
7y ago
The small vector class contains a static array of a few bytes. And as long as your data fits in it, no extra heap allocation is needed. A vector always does heap allocation, even if you only use a few bytes.
77.
▲
by
PixelOfDeath
7y ago
Isn't the tendency to use more and more large pages (1MiB) on x86 anyway? And then just use user space allocators to split them up for malloc.
78.
▲
by
PixelOfDeath
7y ago
I wonder if there is anything between the complexity of HPX and this stack-less task libraries.