luoyu-intel
|
11182fae34
|
fix scratch size of softmax (llama/8642)
|
2024-08-08 22:48:46 +03:00 |
|
luoyu-intel
|
29a2739d27
|
Fix WARP_SIZE=16 bug of Intel GPU (llama/8266)
* fix group_norm ut
* split softmax
* fix softmax
* add concat support condition
* revert debug code
* move QK_WARP_SIZE to presets.hpp
|
2024-07-08 14:53:55 +03:00 |
|
luoyu-intel
|
4a2ba1a065
|
Fix win build conflict of math library (llama/8230)
* fix win build conflict of math library
* fix the condition: !(win32 & SYCL)
* revert warp_size=16
|
2024-07-08 14:53:55 +03:00 |
|
luoyu-intel
|
f096cc6807
|
Fix the sub group size of Intel (llama/8106)
* use warp_size macro for all sycl kernels
* fix mask of permute_sub_group_by_xor
* fix rms_norm with correct warp number
* fix rms_norm_f32/group_norm_f32
* move norm to norm.cpp file
* fix quantize bug
* fix mmvq's batch size
|
2024-07-08 14:53:55 +03:00 |
|