FFmpeg/libavutil/x86
Lynne 35080149ef
x86/tx_float: mark AVX2 functions as AVXSLOW
Makes Bulldozer prefer AVX functions rather than AVX2,
which are 64% slower:

AVX:  117653 decicycles in av_tx (fft), 1048535 runs,     41 skips
AVX2: 193385 decicycles in av_tx (fft), 1048561 runs,     15 skips

The only difference between both is that vgatherdpd is used in
the former. We don't want to mark them with the new SLOW_GATHER
flag however, since gathers are still faster on Haswell/Zen 2/3
than plain loads.
2022-01-29 03:08:16 +01:00
..
asm.h
bswap.h libavutil: x86: Include stdlib.h before using _byteswap_ulong 2020-01-23 18:30:26 +02:00
cpu.c avutil/cpu: move slow gather checks below in the function 2021-12-21 17:51:17 -03:00
cpu.h Merge commit '4cf84e254a' 2018-02-11 23:08:48 -03:00
cpuid.asm
emms.asm Merge commit '7911186ed6' 2017-03-23 18:28:56 -03:00
emms.h Merge commit '7911186ed6' 2017-03-23 18:28:56 -03:00
fixed_dsp.asm
fixed_dsp_init.c
float_dsp.asm x86/float_dsp: add ff_vector_dmul_{sse2,avx} 2018-09-14 12:54:42 -03:00
float_dsp_init.c x86/float_dsp: add ff_vector_dmul_{sse2,avx} 2018-09-14 12:54:42 -03:00
imgutils.asm Merge commit '6be7944ee2' 2017-03-23 18:05:27 -03:00
imgutils_init.c Merge commit 'd7bc52bf45' 2017-03-20 08:34:10 +01:00
intmath.h x86/intmath: add VEX encoded versions of av_clipf() and av_clipd() 2021-11-19 11:21:03 -03:00
intreadwrite.h
lls.asm
lls_init.c Include attributes.h directly 2021-04-19 14:34:10 +02:00
Makefile x86/tx_float: do not build tx_float_init.c if x86 assembly is disabled 2022-01-27 02:17:46 +01:00
pixelutils.asm x86/pixelutils: add missing preprocessor wrapper to the AVX2 functions 2018-07-31 22:14:42 -03:00
pixelutils.h
pixelutils_init.c x86/pixelutils: don't use the AVX2 functions on CPUs known to be slow with them 2018-07-31 22:14:53 -03:00
timer.h
tx_float.asm x86/tx_float: add permute-free FFT versions 2022-01-26 04:13:58 +01:00
tx_float_init.c x86/tx_float: mark AVX2 functions as AVXSLOW 2022-01-29 03:08:16 +01:00
w64xmmtest.h
x86inc.asm avutil/x86inc: fix warnings when assembling with Nasm 2.15 2020-07-12 11:30:23 -03:00
x86util.asm avutil/x86util : add macro for loading a 128 bits constants in an xmm or in each part of an ymm in order to simplify avx2 asm func 2017-12-02 18:25:15 +01:00