Ashwin Sekhar T K
4e1be0e481
ARM64: Add THUNDERX3T110 Target
5 years ago
Martin Kroeker
bd2498c886
Use POWER6 GEMM parameters on 32bit POWER8
5 years ago
Rajalakshmi Srinivasaraghavan
d23419accc
powerpc: Optimized SHGEMM kernel for POWER10
This patch introduces new optimized version of SHGEMM kernel
using power10 Matrix-Multiply Assist (MMA) feature introduced in
POWER ISA v3.1. This patch makes use of new POWER10 compute instructions
for matrix multiplication operation.
Tested on simulator and there are no new test failures.
5 years ago
Rajalakshmi Srinivasaraghavan
9fe930f205
powerpc: Add support for future processor
This is the initial patch to support build infrastructure
for POWER10 architecture.
5 years ago
Martin Kroeker
f16e39554d
Change PPCG4 CGEMM_M to match kernel change
5 years ago
张丹枫
ea5bdc3f72
split cortex-a53 param to match 8x8 kernel
5 years ago
Marius Hillenbrand
1b0b4349a1
s390x/Z14: Change register blocking for SGEMM to 16x4
Change register blocking for SGEMM (and STRMM) on z14 from 8x4 to 16x4
by adjusting SGEMM_DEFAULT_UNROLL_M and choosing the appropriate copy
implementations. Actually make KERNEL.Z14 more flexible, so that the
change in param.h suffices. As a result, performance for SGEMM improves
by around 30% on z15.
On z14, FP SIMD instructions can operate on float-sized scalars in
vector registers, while z13 could do that for double-sized scalars only.
Thus, we can double the amount of elements of C that are held in
registers in an SGEMM kernel.
Signed-off-by: Marius Hillenbrand <mhillen@linux.ibm.com>
5 years ago
Martin Kroeker
03ff213c51
Increase POWER8 ZGEMM_R and use same R values for POWER9
fixes lapack-test zger failures seen in #2299 after application of my PR #2551
5 years ago
Martin Kroeker
00172d440b
Typo fix in MIPS24K addition
5 years ago
Martin Kroeker
61bbae3ac1
Handle MIPS24K like P5600
and allow enforcing TARGET=1004K as well (omission from earlier 1004K merge and later introduction of TARGET check)
5 years ago
Martin Kroeker
a33d177430
Increase default BUFFER_SIZE on ARM, ZARCH and newer x86_64, add GEMM_R for POWER8/9
As shown in #2538 , default buffersizes on some platforms were smaller than required in memory.c
and the requirement could never be fulfilled for a calculated GEMM_R on PPC given the fomula used
5 years ago
Martin Kroeker
567d2760e6
Merge pull request #2520 from wjc404/develop
Fix avx512 sgemm performance bug when ldc is a multiple of 1024
5 years ago
wjc404
64daad4365
Update param.h
5 years ago
Martin Kroeker
ea8eec5d17
Merge pull request #2422 from wjc404/develop
Adjust SkylakeX GEMM3M parameters, add an AVX512 STRMM kernel and fix performance bugs in AVX2 s/c/z GEMM
5 years ago
Ali Saidi
c623a965f9
Add Neoverse-N1 core
The implementation is a hybird of the ARMV8 one with some of the
improved TX2 rountines along with specifying -march=v8.2-a
5 years ago
Martin Kroeker
8164fd1328
Always assume server-class cpu count for TSV110 and EMAG8180
5 years ago
Martin Kroeker
71e5669c3e
Add preliminary support for EMAG8180 ARMV8 processor
5 years ago
wjc404
b0558c11b9
Update param.h
5 years ago
wjc404
83b6be7976
Update param.h
5 years ago
wjc404
f3f969f681
Update param.h
5 years ago
Wang,Long
fbf4f48f4a
fix a few performance drop in some matrix size per data type
Signed-off-by: Wang,Long <long1.wang@intel.com>
5 years ago
wjc404
1c67567008
improve skylakex paralleled sgemm performance
5 years ago
wjc404
b7b408a120
optimize AVX2 SGEMM
5 years ago
wjc404
6362c34ee6
Update param.h
5 years ago
wjc404
64639f440f
Update param.h
5 years ago
wjc404
611445c7f8
Update param.h
6 years ago
wjc404
105e26e12a
Adjust Haswell ZGEMM blocking parameters
6 years ago
wjc404
e20709e976
Update param.h
6 years ago
Martin Kroeker
6082e556cd
Use "generic" S/CGEMM unroll M on big-endian PPC970
as the respective PPC970 "altivec" kernels give wrong results when compiled for big endian
6 years ago
Martin Kroeker
4c6a457358
Merge pull request #2300 from wjc404/develop
Optimize SGEMM on SKYLAKEX CPUs
6 years ago
wjc404
ae43b75a6a
Add files via upload
6 years ago
wjc404
274ff5cdb8
update sgemm_q on skylakex cpus
6 years ago
Martin Kroeker
df857551c0
Remove special parameter set for obsolete IOS/ARMV8 workaround
6 years ago
wjc404
5da9484d93
Add files via upload
6 years ago
Martin Kroeker
6b83079368
Count cpu cores on ARMV8 and use that to pick the GEMM_PQ parameters ( #2267 )
There is currently no simple way to query cache sizes on ARMV8, so this takes the number of cores as a trivial indication if the target is a server-class device with a big cache, or just a single-board toy or smartphone.
6 years ago
Martin Kroeker
6b6c9b1441
Merge pull request #2172 from quickwritereader/develop
power9 cgemm/ctrmm. new sgemm 8x16
6 years ago
AbdelRauf
a97b301aaa
cgemm/ctrmm power9
6 years ago
pkubaj
7c7505a778
Fix build for PPC970 on FreeBSD pt.2
FreeBSD needs those macros too.
6 years ago
AbdelRauf
cdbfb891da
new sgemm 8x16
6 years ago
AbdelRauf
d0c3543c3f
power9 zgemm ztrmm optimized
6 years ago
AbdelRauf
a469b32cf4
sgemm pipeline improved, zgemm rewritten without inner packs, ABI lxvx v20 fixed with vs52
6 years ago
AbdelRauf
8fe794f059
improved zgemm power9 based on power8
6 years ago
AbdelRauf
628b335e83
Merge branch 'develop' of https://github.com/quickwritereader/OpenBLAS into develop
6 years ago
AbdelRauf
0f105dd8a5
sgemm/strmm
6 years ago
Martin Kroeker
7c51cc8527
Merge branch 'develop' into develop
6 years ago
AbdelRauf
853a18bc17
power9 makefile. dgemm based on power8 kernel with following changes : 32x unrolled 16x4 kernel and 8x4 kernel using (lxv stxv butterfly rank1 update). improvement from 17 to 22-23gflops. dtrmm cases were added into dgemm itself
6 years ago
Martin Kroeker
03d7110900
Merge pull request #2042 from maomao194313/develop
add TARGET support for HiSilicon tsv110 CPUs
6 years ago
maomao194313
7e3eb9b25d
make DYNAMIC_ARCH=1 package work on TSV110
6 years ago
ken-cunningham-webuse
b0c714ef60
param.h : enable defines for PPC970 on DarwinOS
fixes:
gemm.c: In function 'sgemm_':
../common_param.h:981:18: error: 'SGEMM_DEFAULT_P' undeclared (first use in this function)
#define SGEMM_P SGEMM_DEFAULT_P
^
6 years ago
Martin Kroeker
bdc73a49e0
Add parameters for Z14
from patch provided by aarnez in #991
6 years ago