69 Commits (optimized_for_deeplearning)

Author SHA1 Message Date
  Zhang Xianyi d6cab3f37e Refs #113. Support AMD Bobcate using Barcelona kernel codes. Replace 3DNow! with MMX. 13 years ago
  Xianyi Zhang 19a48b82cf Init Sandybridge codes based on Nehalem. 13 years ago
  traz 7af0139a09 Modify P Q R size of Loongson3b. 13 years ago
  Wang Qian 66904fc4e8 BLAS3 used standard MIPS instructions without extensions on Loongson 3B. 14 years ago
  Wang Qian 8163ab7e55 Change the block size on Loongson 3B. 14 years ago
  Xianyi Zhang b95ad4cfaf Support detecting ICT Loongson-3B CPU. 14 years ago
  traz 831858b883 Modify aligned address of sa and sb to improve the performance of multi-threads. 14 years ago
  traz d238a768ab Use ps instructions in cgemm. 14 years ago
  Xianyi Zhang 4727fe8abf Refs #47. On Loongson 3A, set DGEMM_R parameter depending on different number of threads. It would improve double precision BLAS3 on multi-threads. 14 years ago
  traz 74a3f63489 Tuning mb, kb, nb size to get the best performance. 14 years ago
  traz cb0214787b Modify compile options. 14 years ago
  traz c8360e3ae5 Complete all the plura single precision functions of level3 on Loongson3a, the performance is 2.3GFlops. 14 years ago
  traz e72113f06a Add ztrmm and ztrsm part on loongson3a. The average performance is 2.2G. 14 years ago
  traz 1c96d345e2 Improve zgemm performance from 1G to 1.8G, change block size in param.h. 14 years ago
  traz 88d94d0ec8 Fixed #30 strmm computational error on Loongson3A. 14 years ago
  traz ab9e4ce351 Adjust kc size from 112 to 116 . 14 years ago
  traz 1aa9a298e1 Change BLOCK SIZE of LOONGSON3A TARGET. 14 years ago
  Xianyi Zhang 0597c1076f Added the configures of loongson 3a. refs #1 14 years ago
  Xianyi Zhang 342bbc3871 Import GotoBLAS2 1.13 BSD version codes. 14 years ago