You can either have performance (=write manual ASM for each platform), or portability, but not both.
What so-called "portable SIMD" libraries give you is "portable auto-vectorization". "Portable performance" is a global property of the algorithm. Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice. The microbenchmarks will look great, though. ;)
louthy•37m ago
> there is no portable SIMD
Except in languages with a JIT compiler
Tanjreeve•27m ago
Starts to get a bit philosophical on what constitutes "portable" but JIT compilers would emit an opcode based off of whatever the frontend/IR is saying to do surely?
Scene_Cast2•35m ago
What about numpy, numba, and torch.compile?
Sharlin•30m ago
Getting 2x or 4x performance in your inner loops using a reasonable SIMD library is infinitely better than theoretically getting 8x performance with hand-coded nonportable intrinsics, because the latter is never going to happen in most programs, so the actual point of comparison is scalar code, or autovectorized code at best.
IshKebab•15m ago
It's a continuum. Some things basically all SIMD implementations support. Want to add 2 4xf32 vectors together? That's pretty easy to do portably.
But yeah to be fair if you are at that point, you probably want to go fully non-portable anyway. Especially with AI.
Has anyone even figured out how to do vector stuff (SVE/RVV) without assembly?
Archit3ch•48m ago
You can either have performance (=write manual ASM for each platform), or portability, but not both.
What so-called "portable SIMD" libraries give you is "portable auto-vectorization". "Portable performance" is a global property of the algorithm. Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice. The microbenchmarks will look great, though. ;)
louthy•37m ago
Except in languages with a JIT compiler
Tanjreeve•27m ago
Scene_Cast2•35m ago
Sharlin•30m ago
IshKebab•15m ago
But yeah to be fair if you are at that point, you probably want to go fully non-portable anyway. Especially with AI.
Has anyone even figured out how to do vector stuff (SVE/RVV) without assembly?