OpenBLAS

Commit Graph

Author	SHA1	Message	Date
yancheng	d4c96a35a8	loongarch64: Add optimizations for axpy and axpby.	1 year ago
yancheng	360acc0a41	loongarch64: Add optimizations for swap.	1 year ago
yancheng	174c25766b	loongarch64: Add optimizations for copy.	1 year ago
yancheng	49829b2b7d	loongarch64: Add optimizations for iamin.	1 year ago
yancheng	be83f5e4e0	loongarch64: Add optimizations for iamax.	1 year ago
yancheng	e3fb2b5afa	loongarch64: Add optimizations for imin.	1 year ago
yancheng	e46b48e372	loongarch64: Add optimizations for imax.	1 year ago
yancheng	702fc1d56d	loongarch64: Add optimization for min.	1 year ago
yancheng	346b384d1c	loongarch64: Add optimization for max.	1 year ago
yancheng	ff2ecc6cda	loongarch64: Add optimization for amin.	1 year ago
yancheng	265b5f2e80	loongarch64: Add optimizations for amax.	1 year ago
yancheng	993ede7c70	loongarch64: Add optimizations for scal.	1 year ago
Octavian Maghiar	4a12cf53ec	[RISC-V] Improve RVV kernel generator LMUL usage The RVV kernel generation script uses the provided LMUL to increase the number of accumulator registers. Since the effect of the LMUL is to group together the vector registers into larger ones, it actually should be used as a multiplier in the calculation of vlenmax. At the moment, no matter what LMUL is provided, the generated kernels would only set the maximum number of vector elements equal to VLEN/SEW. Commit changes the use of LMUL to properly adjust vlenmax. Note that an increase in LMUL results in a decrease in the number of effective vector registers.	1 year ago
Octavian Maghiar	e4586e81b8	[RISC-V] Add RISC-V Vector 128-bit target Current RVV x280 target depends on vlen=512-bits for Level 3 operations. Commit adds generic target that supports vlen=128-bits. New target uses the same scalable kernels as x280 for Level 1&2 operations, and autogenerated kernels for Level 3 operations. Functional correctness of Level 3 operations tested on vlen=128-bits using QEMU v8.1.1 for ctests and BLAS-Tester.	1 year ago
Martin Kroeker	39bf8ece20	Merge pull request #4340 from yinshiyou/la-dev Add some refines and optimizations for LoongArch.	1 year ago
Shiyou Yin	9fe07d82fd	loongarch: Add LSX optimization for dot.	1 year ago
Shiyou Yin	13b8c44b44	loongarch: Add optimization for dsdot kernel.	1 year ago
Shiyou Yin	3def6a8143	loongarch: Add LASX optimization for dot.	1 year ago
Bart Oldeman	c34e2cf380	Use _mm_set1_epi{32,64x} to init mask in x86-64 [cz]asum for skylake kernels. This is the same method as used in [sd]asum. _mm_set1_epi64x was commented out for zasum, but has the advantage of avoiding possible undefined behaviour (using an uninitialized variable), optimized out by NVHPC and icx. The new code works fine with those compilers. For GCC 12.3 the generated code is identical; no matter what method you use, the compiler optimizes the code into a compile-time constant, there is no performance benefit using mm_cmpeq_epi8 since the corresponding instruction (VPCMPEQB) isn't actually generated!	1 year ago
Martin Kroeker	22aa401656	Temporarily disable the AVX512 CASUM/ZASUM microkernels for any version of NVIDIA HPC (#4327 ) * Temporarily disable the C/ZASUM microkernels for any version of NVHPC	1 year ago
Bart Oldeman	f8ad5344c2	Fix casum fallback kernel. This kernel is only used on Skylake+ if the kernel with AVX512 intrinsics can't be used, but used the variable x1 incorrectly in the tail end of the loop, as it is still at the initial value instead of where x points to. This caused 55 "other error"s in the LAPACK tests (https://github.com/OpenMathLib/OpenBLAS/issues/4282) This change makes casum.c as similar as possible as zasum.c, because zasum.c does this correctly.	1 year ago
Martin Kroeker	04bc801999	(Re)apply fixes for supporting only a subset of precision types from PR 3915	1 year ago
Martin Kroeker	9019bc4945	Use SkylakeX ?ASUM microkernel for Cooperlake/Sapphirerapids as well	1 year ago
Martin Kroeker	3bfa4d4dcc	Fix outdated SVE kernel definitions for Cortex cpus by aliasing to ARMV8SVE	1 year ago
Rajalakshmi Srinivasaraghavan	980f702f72	POWER: AIX: Make use of power10 optimization POWER10 optimizations are disabled when using default AIX assembler. As we have fixed many issues recently, enabling optimization path for default assembler.	1 year ago
Rajalakshmi Srinivasaraghavan	9f42570e33	POWER: Increase macro size limit for AIX This patch increases the macro size limit from 4096 to 16384 to allow compiling larger assembly files in AIX. Tested with GCC and IBM Open XL C.	2 years ago
Martin Kroeker	9f49aef91b	Merge pull request #4255 from RajalakshmiSR/AIX-P10 POWER10: Fix compilation issues with Open XL C	2 years ago
Martin Kroeker	e7d05402e0	Fix up S/D GEMM copy function definitions after #4009	2 years ago
Rajalakshmi Srinivasaraghavan	71d733e5f7	POWER: Avoid m4 conversions for C files This patch removes intermediate m4 conversions used in sbgemm compilation as it is not needed for .c files. Tested on AIX with gcc and IBM Open XL C.	2 years ago
Rajalakshmi Srinivasaraghavan	82fc29a57a	POWER10: Fallback to POWER8 functions As cgemm and zgemm kernels are not optimized for big endian falling back to POWER8 versions. Tested on AIX using gcc and Open XL C.	2 years ago
Rajalakshmi Srinivasaraghavan	db0805906b	powerpc: Fix build errors with Open XL C This patch fixes errors when using Open XL C compiler on AIX. Tested with gcc/xlf and ibm-clang/xlf compiler combinations.	2 years ago
Martin Kroeker	675cd551da	fix improper function prototypes (empty parentheses)	2 years ago
gxw	d15e0a055c	LoongArch64: Fixed compilation issues when enable DYNAMIC_ARCH	2 years ago
gxw	4670eb1462	LoongArch64: Add dtrsm kernel	2 years ago
gxw	f2cf929374	LoongArch64: Add sgemv kernel	2 years ago
Martin Kroeker	8e6d93359d	Merge pull request #4196 from TiborGY/obsolete_inlines Modernize obsolete inline order	2 years ago
gxw	394a1fd1bf	LoongArch64: Compatible with early internal toolchain __loongarch_grlen and __loongarch_frlen were introduced in gcc version 8.3.0 (Loongnix 8.3.0-6.lnd.vec.31) internally within Loongson to standardize the general and floating-point register widths. However, previous versions did not have them, requiring additional checks to be added.	2 years ago
Martin Kroeker	9c4ae4d4fb	Merge pull request #4206 from martin-frbg/issue4201-2 Work around miscompilation of zdot_thunderx2t99 by the current NVIDIA HPC compiler	2 years ago
Martin Kroeker	88435104c8	Merge pull request #4204 from martin-frbg/llvm17-2 Work around LLVM17 miscompiling the AVX512 microkernels for CASUM/ZASUM	2 years ago
Martin Kroeker	fc8894dd98	Workaround miscompilation by NVIDIA nvc	2 years ago
Martin Kroeker	7a6203ffa1	restore default Neoverse SVE build instructions for non-NVIDIA compilers	2 years ago
Martin Kroeker	2c3034ff7f	Disable the C/ZASUM AVX512 microkernels when compiling with LLVM17 as well	2 years ago
Martin Kroeker	8794544b43	Add support for compiling the Neoverse SVE kernels with the NVIDIA HPC compiler	2 years ago
gxw	553cc1372f	LoongArch64: Add sgemm_kernel	2 years ago
Martin Kroeker	12ede72ab7	Merge pull request #4192 from imciner2/im/clangfix Fix cooperlake and sapphire rapids march flags on clang	2 years ago
Ian McInerney	79c15db348	Fix power10 gcc intrinsic check __builtin_vsx_assemble_pair was only in GCC 10-11.2 and was replaced by __builtin_vsx_build_pair thereafter.	2 years ago
TGY	b5ba95a6c0	Modernize obsolete inline order	2 years ago
Ian McInerney	8a8a8479be	Fix cooperlake and sapphire rapids march flags on clang The march=cooperlake and march=sapphirerapids flags were never getting added when building with Clang targetting those architectures. Instead it was falling back to the skylake AVX512 implementation. Clang added support for these two architectures in Clang 9 and Clang 12, so introduce new checks for those versions to enable the appropriate march flag, and fallback to skylake otherwise.	2 years ago
Martin Kroeker	34da1a067d	Allow negative INCX (API change from version 3.10 of the reference implementation)	2 years ago
Martin Kroeker	07e32c4cb8	Allow negative INCX (API change from version 3.10 of the reference implementation)	2 years ago

... 8 9 10 11 12 ...

2524 Commits (2c0dd2468e253ec7ecdabafcb15d5016a7218a12)