4 ms·
> I'm not sure this is going to work fast enough for that. From reading the SIMD Everywhere description, it seems to me that SIMDe is a way to allow code that
by gioele 6y ago
> I'm not sure this is going to work fast enough for that.
From reading the SIMD Everywhere description, it seems to me that SIMDe is a way to allow code that targets only a platform to work on other platforms as well. As a nice byproduct, you get _some_ speed up if architecture targeted by the code is similar to the architecture that will run the code.
Portability is the main focus, not speed.
Obviously, once you have a good emulation of an architecture the first question is going to be: can I make it faster?
- mtgx 6y agoAll other chip architectures should adopt Arm's SVE2 or something similar. https://community.arm.com/developer/ip-products/processors/b/processors-ip-blog/posts/new-technologies-for-the-arm-a-profile-architecture https://community.arm.com/developer/ip-products/processors/b...
- fao_ 6y agoHi there! It looks like you've been shadowbanned, I vouched a few of your comments because they didn't seem terrible, but I didn't scroll very far. It might be worth sending a mail to dang to ask him to take a look at your account or something!
- chmod775 6y agoThat account has 33k karma and is 8 years old. I just can't imagine a set of circumstances where a shadowban of all things could be justified if they transgressed in some way. It's probably a mistake?
- fao_ 6y agoor they said something ridiculously heinous
- chmod775 6y agoIn that case one should point that out to them and warn them. Shadowbans are quite a heinous punishment in themselves, and better used to prevent spammers etc. from creating new accounts. But people who make thousands upon thousands of comments are bound to make an emotional comment or an error of judgment at some point.
- saagarjha 6y agoIt'd be nice, but they won't :( I'm not even sure if any consumer ARM chips ship with SVE2 at the moment.
- nemequ1729 6y agoI'm not aware of anything in the works for Intel, but RISC-V seems to be going in that direction. I'm hoping to start using SVE to implement AVX-512 and AVX in SIMDe soon.
- nemequ1729 6y agoI'm the lead developer of SIMDe. I wouldn't say that portability is the main focus. The first step is to get portable implementations up and running, but a huge number of functions have optimized implementations for NEON, AltiVec/VSX, and WASM SIMD 128, and we're working on adding more. We go to a lot of trouble to get good performance on multiple architectures, basically writing each implementation several times and using ifdefs to switch depending on what the fastest version available to a given architecture will be. Even just for the portable implementations, we use a lot of hints to help the compiler auto-vectorize. Almost every portable implementation has a loop which uses a pragma to try to get the compiler do the right thing (OpenMP SIMD, clang loop-specific pragmas, GCC ivdep, etc.). On top of that we take advantage of lots of compiler-sepecific features to speed things up where possible, including GCC-style vector extensions, __builtin_shuffle/__builtin_shufflevector, and __builtin_convertvector. SIMDe never going to be as fast as someone who knows what they're doing writing an optimized implementation for a given target. However, it should be as fast (or faster) than someone who is just trying to do a direct port where they just try to match the existing code as closely as possible.
- innocenat 6y agoI don't have experience with SIMD on platform other than x86/amd64, but a few of data-shuffling type functions [0] have SIMD version that is not that faster than scalar C implementation, and the overhead of translation might make then slower. [0]: http://web.archive.org/web/20140807014206/http://x264dev.multimedia.cx/archives/232 http://web.archive.org/web/20140807014206/http://x264dev.mul...
- nemequ1729 6y agoI hadn't seen that post before, thanks. This doesn't quite apply to SIMDe; the problem that post is talking about is really at a higher level… whether it is faster to do a bunch of shuffles or use some scalar code. Once SIMDe is called you've already made your decision, and at that level the hardware-based shuffles are much faster than scalar code. For example, see the decompression speed benchmarks for LZSSE-SIMDe (<https://github.com/nemequ/LZSSE-SIMDe> https://github.com/nemequ/LZSSE-SIMDe>) (they're in the README). It sounds like what that post really needs is a fast 16-bit gather operation. AVX2 has some 32-bit gather functions which you may be usable (2 gathers + a blend could emulate 16-bit gathers). For NEON, you could probably use one of the `vtbl` functions; they're all 8-bit, but that just means you have separate index entries for high and low bytes… it's a bit more code, but there shouldn't be any runtime overhead.