Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
obl
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
61.
▲
French ISPs Ordered to Block Sci-Hub and LibGen
(torrentfreak.com)
696 points
by
obl
8y ago
|
368 comments
62.
▲
by
obl
8y ago
There is a reason why basically every attempt to make this kind of language fast has to support some form of on-stack replacement. For example, it's hard to optimize even local variable dataflow in python since it's part of the AP
63.
▲
New OpenGL Driver for Intel Gen8 GPUs Merged into Mesa
(lists.freedesktop.org)
91 points
by
obl
8y ago
|
21 comments
64.
▲
by
obl
8y ago
Very clever use of virtual memory. Unfortunately I would expect this technique not to scale well with increasing page size, and I sure hope that we will be slowly moving to >4k pages as years go by since the TLB is a very common bottlene
65.
▲
by
obl
8y ago
Of course ? Yes it often makes sense to produce inefficient code for many business reasons (development speed, users don't care, easier platform to deploy to, etc). However what Slack does would be nowhere near taxing for a modern comp
66.
▲
System76 Announces New Darter Pro Linux Laptop
(notebookcheck.net)
4 points
by
obl
8y ago
|
1 comments
67.
▲
by
obl
8y ago
The "core" of the trick is nice : amortizing interpreter dispatch over many items. (ignoring the column layout/SIMD stuff which basically helps in any case) Essentially it's turning : LOAD DISPATCH OP1 DI
68.
▲
Nvidia's PhysX engine now open-source
(github.com)
3 points
by
obl
8y ago
|
0 comments
69.
▲
by
obl
8y ago
What's the point of using this for native applications ? I mean yeah, it might be a little easier to distribute but it's kind of absurd to agree to throw away 20% performance for a bit of convenience and then sched tears on how th
70.
▲
by
obl
8y ago
If you have time : https://www.youtube.com/watch?v=NGFhc8R_uO4 It's great.
71.
▲
Design of LuaJIT 2.0 (2009)
(lua-users.org)
59 points
by
obl
8y ago
|
3 comments
72.
▲
by
obl
8y ago
Extrapolating up to one physics frame is still not a good solution, it will be noticeable. For example a falling object will clip in the ground for tens of ms of displacement which is definitely visually significant. A better solution is to
73.
▲
Verilog sources for Western Digital's open source RISC-V core
(github.com)
320 points
by
obl
8y ago
|
78 comments
74.
▲
An Introduction to Intel GPU Assembly
(youtube.com)
3 points
by
obl
8y ago
|
0 comments
75.
▲
by
obl
8y ago
Of course that's a choice for everyone to make, but I feel like in general the benefit of fitting in the platform is overrated. After all the whole webapp-everything somewhat shows that people don't care that much, they are happy
76.
▲
by
obl
8y ago
You should probably use mipmaps to avoid aliasing when rendering from far away. I guess you'll have to compute the LOD by hand since neighbor fragments are not necessarily performing the same texture fetches.
77.
▲
Android's arbitrary precision calculator (Hans Boehm)
(cacm.acm.org)
3 points
by
obl
8y ago
|
0 comments
78.
▲
by
obl
8y ago
Since this is intented to be educational I find it a bit unfortunate to implement function calls using the host (python) stack since it's an important part of an interpreter. It means that you can't really add things like stack tr
79.
▲
by
obl
8y ago
I wonder how they are going to pull that off without security implications. If you've played with compute shaders (or any of the modern "general purpose" shader stuff, ie arbitrary loads & stores etc) you probably know th
80.
▲
ATIC is looking for potential buyers for GlobalFoundries
(bitsandchips.it)
2 points
by
obl
8y ago
|
0 comments
81.
▲
by
obl
8y ago
It's not free though. If you can actually remove most of the power/area hungry ooo scheduling logic for only 10% less performance on real world workload, you can pack a lot more cores or a bigger GPU on the same chip without blowi
82.
▲
by
obl
8y ago
The memory mapping trick they use on x86 to avoid masking only works up to the maximum 48 bits of addressable virtual memory, so there is much less than 22 free bits in that case. It's also not quite free since it takes up TLB space.
83.
▲
by
obl
8y ago
It's a cute idea to cobble together a meta-tracing JIT for sqlite's query interpreter using source extraction and gcc. I think it's basically a local minimum for implementation complexity vs performance though. To push that f
84.
▲
Paths: Stroking and Offsetting
(tavmjong.free.fr)
1 points
by
obl
8y ago
|
0 comments
85.
▲
Corners Don't Look Like That: Regarding Screenspace Ambient Occlusion
(nothings.org)
2 points
by
obl
8y ago
|
0 comments
86.
▲
by
obl
8y ago
It's disappointing to have this whole business go inside the driver / dx runtime. I'm sure the vast majority of it could be done from user code (CPU code & compute shaders) and it would be much more profitable for everyon
87.
▲
by
obl
8y ago
Beyond the technical aspect, the pedagogical side of this is huge. An argument often made in favor of Rust's static analysis restrictions is that "worst case scenario, you just drop down to unsafe and you're back to C semanti
88.
▲
by
obl
8y ago
Half a gig of machine code... In the paper they mention trying to limit the code size to 10% of that and getting ~60% of the baseline perf. However, they got this number by simply dropping to the interpreter for the rest of the code. I'
89.
▲
Stacked Borrows: An Aliasing Model for Rust
(ralfj.de)
5 points
by
obl
8y ago
|
0 comments
90.
▲
by
obl
8y ago
For a simple baseline JIT like that, you can improve the quality of the generated code a lot at basically no compile-time cost (still one pass) by doing some flavor of "destination driven code generation". It also adds very little
More ›