Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
zwegner
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
61.
▲
by
zwegner
7y ago
(author here.) I started the x86-sat project about a week ago, which parses Intel's pseudocode for all of the intrinsic functions and converts it into a Z3 SAT model. I got a basic prototype working in around a day, and I've done
62.
▲
x86 SIMD superoptimizer in ~100 lines of Python
(github.com)
3 points
by
zwegner
7y ago
|
1 comments
63.
▲
by
zwegner
7y ago
Thanks for the reply! After playing more, and switching from a trial-and-error/backtrack approach to a more deliberate style where I try to mentally prove each extension of a line, I feel less of a need for the start-line-anywhere thin
64.
▲
by
zwegner
7y ago
This is really nice! Very clean design, and a great puzzle too. Two small interface suggestions: first, allow starting a line on any square, not just one with a number (the selection would be canceled if it doesn't connect to a square
65.
▲
by
zwegner
7y ago
This is very cool! I hope this gets released--even if it's a rough prototype, I would still use it in its current state. Like others (apparently), I started writing a system remarkably similar to this earlier this year. I just wanted t
66.
▲
by
zwegner
7y ago
That's pretty cool! It's really not too different from make.py. Using decorators and python functions is likely cleaner in a lot of cases, though I'm not sure if that would fit well with make.py's pseudo-declarative mode
67.
▲
by
zwegner
7y ago
Ah, thanks for the references. Now that you mention them, I realize I definitely knew about tup and fabricate before (and possibly waf?), but had forgotten about them. I haven't really thought much about trace-based build systems in ye
68.
▲
by
zwegner
7y ago
> I think bazels solution here is to just always make fat binaries. Or at least that's how it works if you're using blaze. Oh, I should've specified soft/hard filesystem links--they'd make handling a virtual fi
69.
▲
by
zwegner
7y ago
> My understanding is that bazel is moving away from this, so that you can define toolchains by saying "here is a binary that serves the job of linking/compiling stuff". How do they ensure determinism in that case? Is it j
70.
▲
by
zwegner
7y ago
Wow, that's kind of incredible. That bug has been open for years , and is only starting to see some progress. Lends some credence to my unsupported claim of Bazel being overengineered...
71.
▲
by
zwegner
7y ago
Yeah, I know about Bazel, but only at a high level--I haven't used it. I generally think the hermetic build concept is a very good one, but IMO Bazel goes about it the wrong way, and is overengineered. Rather than needing custom-built
72.
▲
by
zwegner
7y ago
Note that the author of that paper (a friend of mine) wrote another build system, the dead-simple-but-awesome make.py. I have a mirror/fork of it[0], since it's been unmaintained for a while (but it mostly doesn't need any ma
73.
▲
by
zwegner
7y ago
Oh wow, that is most definitely a bug. Thank you very much for reporting that, before this spreads too much. How embarrassing... I was in a bit of a rush to stick this up on HN before the weekend, as otherwise I'd probably never get ar
74.
▲
by
zwegner
7y ago
BTW, I opened an issue to track this topic: https://github.com/zwegner/faster-utf8-validator/issues/3
75.
▲
by
zwegner
7y ago
In case anyone sees this: I initially agreed with you, thinking that pure UTF-8 validation is a bit more of a niche than might be expected. But if I'm not mistaken, I think at least for JSON validation, the UTF-8 validation can happen
76.
▲
by
zwegner
7y ago
Thanks for the link, that's a good post! I hadn't investigated the throttling issues much in the past, not owning an AVX-512 machine, but that's much more precise than the standard "don't use AVX-512" meme that
77.
▲
by
zwegner
7y ago
Definitely a good idea. I had already downloaded a Mandarin Wikipedia page to look for test sequences, but hadn't benchmarked them. Getting some more pages there is a great place to start.
78.
▲
by
zwegner
7y ago
Ah, thanks for the reply. I understand you better now, and for the most part I agree. Taking out the early exit for ASCII can speed up the frequently-changing case to avoid the misprediction penalty, as long as sufficiently many chunks of i
79.
▲
by
zwegner
7y ago
It depends on the implementation. A good starting point would be looking at the various implementations benchmarked here: https://github.com/lemire/fastvalidate-utf-8 It looks like the main naive validator there (valid
80.
▲
by
zwegner
7y ago
I don't have any particular plans for that, beyond getting a bit of publicity for it here :). Or really any plans at all, besides adding AVX-512 support. It was just something I hacked together in a couple days as I began thinking abou
81.
▲
by
zwegner
7y ago
Well, to be fair, my claim was that it was the fastest in the world that I'm aware of , which is a much weaker statement. :) The random UTF-8 in the benchmark was generated from the code in Daniel Lemire's fastvalidate-utf-8 repo
82.
▲
by
zwegner
7y ago
At least my compiler (LLVM 10) is smart enough to know that "data" isn't modified, and thus fold the addresses of the two vector loads inside the main loop into single instructions: 1a0: c5 fe 6f 6c 31 ff vmovdqu ym
83.
▲
by
zwegner
7y ago
Well how else are you supposed to nest your preprocessor macros? :P Thanks though! I like taking the time to make my code as clean as I can. Glad others appreciate it!
84.
▲
by
zwegner
7y ago
Oh interesting, I hadn't seen that. It looks like it uses the same idea of shuffle-lookups on the first three nibbles. That's a fairly large patch though, I don't think I've fully grokked it. At the very least, my code d
85.
▲
by
zwegner
7y ago
> this might save megawatts of power world wide That'd sure be neat! I really have no idea, though. > A small nitpick: not sure whether those macros bought anything, I guess the optimizer could inline function calls to intrinsic
86.
▲
Show HN: Faster UTF-8 validator
(github.com)
122 points
by
zwegner
7y ago
|
55 comments
87.
▲
by
zwegner
7y ago
Minor correction: the graph of the PCM data is displaying the data as unsigned instead of signed, so it has lots of discontinuities between ~0 and ~64k. Though, for the sake of showing the data as "raw", I guess it doesn't ma
88.
▲
by
zwegner
7y ago
Again, I don't see that as cheating or violating a spec (or at least any guarantee made by Intel about x86(-64) behavior that I know of). I would assume that speculation can do just about anything (access protected memory, run illegal&
89.
▲
by
zwegner
7y ago
> both systems protected against Spectre, Meltdown etc. (i.e. no cheating by ignoring the ISA specifications) Does the x86 specification say "speculation will have no observable side effects on the memory subsytem"? I wasn'
90.
▲
by
zwegner
7y ago
I was recently looking into audio antialiasing/interpolation algorithms for a software synthesis project and stumbled across this. It's written by Miller Puckette, who created Max/MSP and PureData. It's one of the best r
More ›