Georgi Gerganov
b0aeef2d52
ci : fix windows builds to use 2019
2024-11-21 14:28:14 +02:00
Georgi Gerganov
37c88027e1
whisper : use backend registry ( #0 )
2024-11-20 21:00:08 +02:00
Georgi Gerganov
7fd8d9c220
whisper : adapt to new ggml (wip)
2024-11-20 21:00:08 +02:00
Georgi Gerganov
06e059b8f8
talk-llama : sync llama.cpp
2024-11-20 21:00:08 +02:00
Georgi Gerganov
c9f49d5f9d
sync : ggml
2024-11-20 21:00:08 +02:00
Georgi Gerganov
f4c1d7df39
ggml : sync resolve (skip) ( #0 )
2024-11-20 21:00:08 +02:00
Georgi Gerganov
2a444dc5bd
metal : refactor kernel args into structs (llama/10238)
...
* metal : add kernel arg structs (wip)
* metal : fattn args
ggml-ci
* metal : cont + avoid potential int overflow [no ci]
* metal : mul mat struct (wip)
* cont : mul mat vec
* cont : pass by reference
* cont : args is first argument
* cont : use char ptr
* cont : shmem style
* cont : thread counters style
* cont : mul mm id
ggml-ci
* cont : int safety + register optimizations
ggml-ci
* metal : GGML_OP_CONCAT
ggml-ci
* metal : GGML_OP_ADD, GGML_OP_SUB, GGML_OP_MUL, GGML_OP_DIV
* metal : GGML_OP_REPEAT
* metal : GGML_OP_CPY
* metal : GGML_OP_RMS_NORM
* metal : GGML_OP_NORM
* metal : add TODOs for rest of ops
* ggml : add ggml-metal-impl.h
ggml-ci
2024-11-20 21:00:08 +02:00
Georgi Gerganov
bd574b05af
ggml : inttypes.h -> cinttypes (llama/0)
...
ggml-ci
2024-11-20 21:00:08 +02:00
Georgi Gerganov
7e0eafcb1e
ggml : adapt AMX to tensor->grad removal (llama/0)
...
ggml-ci
2024-11-20 21:00:08 +02:00
Georgi Gerganov
75670ae673
ggml : fix compile warnings (llama/0)
...
ggml-ci
2024-11-20 21:00:08 +02:00
Georgi Gerganov
d4fcdf602b
llamafile : fix include path (llama/0)
...
ggml-ci
2024-11-20 21:00:08 +02:00
Georgi Gerganov
6a55015dc4
ggml : remove duplicated sources from the last sync (ggml/1017)
...
* ggml : remove duplicated sources from the last sync
ggml-ci
* cont : remove FindSIMD.cmake [no ci]
2024-11-20 21:00:08 +02:00
Georgi Gerganov
401fbea326
sync : leftovers (ggml/0)
...
ggml-ci
2024-11-20 21:00:08 +02:00
Georgi Gerganov
44d1cbdfe9
cmake : restore CMakeLists.txt (llama/10256)
...
ggml-ci
2024-11-20 21:00:08 +02:00
Georgi Gerganov
5f7e094ccb
scripts : update sync
2024-11-20 21:00:08 +02:00
Georgi Gerganov
6266a9f9e5
release : v1.7.2
2024-11-19 18:54:22 +02:00
Georgi Gerganov
01d3bd7d5c
ci : use local ggml in Android build ( #2567 )
2024-11-16 20:45:41 +02:00
Georgi Gerganov
bb12cd9b77
ggml : tmp workaround for whisper.cpp (skip) ( #2565 )
2024-11-16 20:21:24 +02:00
Georgi Gerganov
f02b40bcb4
update : readme
2024-11-15 16:00:10 +02:00
Georgi Gerganov
83ac2842bd
scripts : fix sync path
2024-11-15 15:24:09 +02:00
Georgi Gerganov
e23721f3fb
cmake : fix ppc64 check ( #0 )
2024-11-15 15:21:04 +02:00
Georgi Gerganov
c0a9f8ef85
whisper : include ggml-cpu.h ( #0 )
2024-11-15 15:21:04 +02:00
Georgi Gerganov
6477b84eb6
build : fixes
2024-11-15 15:21:04 +02:00
Georgi Gerganov
24d706774d
talk-llama : sync llama.cpp
2024-11-15 15:21:04 +02:00
Georgi Gerganov
5089ab2d6a
whisper : fix build ( #0 )
2024-11-15 15:21:04 +02:00
Georgi Gerganov
bdbb906817
sync : ggml
2024-11-15 15:21:04 +02:00
Georgi Gerganov
26a31b78e9
metal : more precise Q*K in FA vec kernel (llama/10247)
2024-11-15 15:21:04 +02:00
Georgi Gerganov
5e110c2eb5
metal : reorder write loop in mul mat kernel + style (llama/10231)
...
* metal : reorder write loop
* metal : int -> short, style
ggml-ci
2024-11-15 15:21:04 +02:00
Georgi Gerganov
4a9926d521
metal : fix build and some more comments (llama/10229)
2024-11-15 15:21:04 +02:00
Georgi Gerganov
ae3c5642d0
metal : fix F32 accumulation in FA vec kernel (llama/10232)
2024-11-15 15:21:04 +02:00
Georgi Gerganov
e287a3b627
metal : hide debug messages from normal log
2024-11-15 15:21:04 +02:00
Georgi Gerganov
9f67aab211
metal : opt-in compile flag for BF16 (llama/10218)
...
* metal : opt-in compile flag for BF16
ggml-ci
* ci : use BF16
ggml-ci
* swift : switch back to v12
* metal : has_float -> use_float
ggml-ci
* metal : fix BF16 check in MSL
ggml-ci
2024-11-15 15:21:04 +02:00
Georgi Gerganov
8f0f785d88
metal : improve clarity (minor) (llama/10171)
2024-11-15 15:21:04 +02:00
Georgi Gerganov
d0b8335789
metal : optimize FA kernels (llama/10171)
...
* ggml : add ggml_flash_attn_ext_get_prec
* metal : use F16 precision in FA kernels
ggml-ci
* metal : minor clean-up
* metal : compile-guard bf16 FA kernels
ggml-ci
* build : remove obsolete compile flag [no ci]
* metal : prevent int overflows [no ci]
* cuda : disable BF16 FA
ggml-ci
* metal : fix BF16 requirement for FA kernels
ggml-ci
* make : clean-up [no ci]
2024-11-15 15:21:04 +02:00
Georgi Gerganov
31c3482a4e
metal : add BF16 support (llama/8439)
...
* ggml : add initial BF16 support
ggml-ci
* metal : add mul_mat_id BF16 support
ggml-ci
* metal : check for bfloat support on the Metal device
ggml-ci
* metal : better var names [no ci]
* metal : do not build bfloat kernels when not supported
ggml-ci
* metal : try to fix BF16 support check
ggml-ci
* metal : this should correctly check bfloat support
2024-11-15 15:21:04 +02:00
Georgi Gerganov
d111a0987e
ggml : adjust is_first_call init value (llama/10193)
...
ggml-ci
2024-11-15 15:21:04 +02:00
Georgi Gerganov
915bcd2c63
metal : add quantized FA support (llama/10149)
...
* metal : add quantized FA (vec) support
ggml-ci
* metal : add quantized FA (non-vec) support
* metal : fix support check
ggml-ci
* metal : clean-up
* metal : clean-up (cont)
* metal : fix shared memory calc + reduce smem + comments
* metal : float-correctness
* metal : minor [no ci]
2024-11-15 15:21:04 +02:00
Georgi Gerganov
939d36fb4c
metal : simplify f16 and f32 dequant kernels (llama/0)
2024-11-15 15:21:04 +02:00
Georgi Gerganov
1471e41180
metal : move dequantize templates to beginning of MSL source (llama/0)
2024-11-15 15:21:04 +02:00
Georgi Gerganov
24a0feb5d9
metal : minor fixup in FA kernel (llama/10143)
...
* metal : minor fixup in FA kernel
ggml-ci
* metal : use the unrolled loop variable
* metal : remove unused var
2024-11-15 15:21:04 +02:00
Georgi Gerganov
0665168ef3
ggml : remove ggml_scratch (llama/10121)
...
ggml-ci
2024-11-15 15:21:04 +02:00
Georgi Gerganov
498ac0dc27
scripts : update sync
2024-11-15 15:21:04 +02:00
Georgi Gerganov
0377596b77
whisper : backend registry init before model load
2024-11-01 10:19:05 +02:00
Georgi Gerganov
c65d0fd3c8
talk-llama : sync llama.cpp
2024-11-01 10:19:05 +02:00
Georgi Gerganov
d9efb664ac
sync : ggml
2024-11-01 10:19:05 +02:00
Georgi Gerganov
ab36d02560
metal : support permuted matrix multiplicaions (llama/10033)
...
* metal : support permuted matrix multiplicaions
ggml-ci
* cont : use nb01 directly for row steps
ggml-ci
* cont : add comments [no ci]
* metal : minor refactor
* metal : minor
2024-11-01 10:19:05 +02:00
Georgi Gerganov
741c138aa1
ggml : add asserts for type conversion in fattn kernels (llama/9971)
...
ggml-ci
2024-11-01 10:19:05 +02:00
Georgi Gerganov and slaren
315364d7de
ggml : add metal backend registry / device (llama/9713)
...
* ggml : add metal backend registry / device
ggml-ci
* metal : fix names [no ci]
* metal : global registry and device instances
ggml-ci
* cont : alternative initialization of global objects
ggml-ci
* llama : adapt to backend changes
ggml-ci
* fixes
* metal : fix indent
* metal : fix build when MTLGPUFamilyApple3 is not available
ggml-ci
* fix merge
* metal : avoid unnecessary singleton accesses
ggml-ci
* metal : minor fix [no ci]
* metal : g_state -> g_ggml_ctx_dev_main [no ci]
* metal : avoid reference of device context in the backend context
ggml-ci
* metal : minor [no ci]
* metal : fix maxTransferRate check
* metal : remove transfer rate stuff
---------
Co-authored-by: slaren <slarengh@gmail.com >
2024-11-01 10:19:05 +02:00
Georgi Gerganov
4e10afb5a9
scripts : sync amx
2024-10-31 22:13:24 +02:00
Georgi Gerganov
aa037a60f3
ggml : alloc ggml_contexts on the heap ( #2525 )
...
* whisper : reduce ggml_context usage
* ggml : allocate contexts on the heap (v2)
* ggml : aligned malloc -> malloc
2024-10-31 22:00:09 +02:00
Georgi Gerganov and Tamotsu Takahashi
19dca2bb14
ci : fix openblas build ( #2511 )
...
* ci : fix openblas build
* cont : would this work?
* ci : I'm sorry, windows
* cont : disabled wrong build
* ci : fix openblas build with pkgconfiglite (#2517 )
- choco install pkgconfiglite (vcpkg-pkgconf doesn't contain pkg-config executable?)
- vcpkg install openblas (otherwise it is not detected now)
---------
Co-authored-by: Tamotsu Takahashi <ttakah+github@gmail.com >
2024-10-30 12:58:26 +02:00
Georgi Gerganov
55e422109b
scripts : add turbo-q8_0 to the benchmark
2024-10-29 19:37:24 +02:00
Georgi Gerganov
3f020fac9d
whisper : minor compile warning
2024-10-29 19:30:26 +02:00
Georgi Gerganov
1d5752fa42
make : fix GGML_VULKAN=1 build ( #2485 )
2024-10-16 18:42:47 +03:00
Georgi Gerganov
ebca09a3d1
release : v1.7.1
2024-10-07 13:06:48 +03:00
Georgi Gerganov
6a94163b91
release : v1.7.0
2024-10-05 16:43:26 +03:00
Georgi Gerganov
8a35b58c4f
scripts : bench v3-turbo
2024-10-05 16:22:53 +03:00
Georgi Gerganov
1789abca84
whisper : remove mel leftover constants ( 396089f)
2024-10-05 16:13:03 +03:00
Georgi Gerganov
847f94fdeb
whisper : zero-out the KV cache upon clear ( #2445 )
2024-10-05 15:23:51 +03:00
Georgi Gerganov
6e40108a59
objc : fix build
2024-10-05 15:23:51 +03:00
Georgi Gerganov
1ba185f4af
metal : zero-init buffer contexts ( #0 )
2024-10-05 15:23:51 +03:00
Georgi Gerganov
396089f3cf
whisper : revert mel-related changes ( #0 )
...
too much extra logic and complexity for small benefit
2024-10-05 15:23:51 +03:00
Georgi Gerganov
941912467d
whisper : adapt to latest ggml (skip) ( #0 )
2024-10-05 15:23:51 +03:00
Georgi Gerganov
f7d55e0614
scripts : sync ggml-backend.cpp
2024-10-05 15:23:51 +03:00
Georgi Gerganov
f62a546e03
whisper : fix excessive memory usage ( #2443 )
...
* whisper : fix KV cache allocation
* whisper : reduce memory overhead from unused input tensors
2024-10-05 12:36:40 +03:00
Georgi Gerganov
ccc2547210
talk-llama : sync llama.cpp
2024-10-03 12:22:17 +03:00
Georgi Gerganov
162a455402
metal : reduce command encoding overhead (llama/9698)
2024-10-03 12:22:17 +03:00
Georgi Gerganov
ff2cb0811f
sync : ggml
2024-10-03 12:22:17 +03:00
Georgi Gerganov and Willy Tarreau
6c91da80b8
ggml : define missing HWCAP flags (llama/9684)
...
ggml-ci
Co-authored-by: Willy Tarreau <w@1wt.eu >
2024-10-03 12:22:17 +03:00
Georgi Gerganov
5963004ff9
ggml : fix GGML_MAX_N_THREADS + improve formatting (ggml/969)
2024-10-03 12:22:17 +03:00
Georgi Gerganov
2ef717b293
whisper : add large-v3-turbo ( #2440 )
2024-10-01 15:57:06 +03:00
Georgi Gerganov
8feb375fbd
tests : remove test-backend-ops ( #2434 )
2024-09-27 11:49:01 +03:00
Georgi Gerganov
69339af2d1
ci : disable failing CUDA and Java builds
2024-09-25 10:05:04 +03:00
Georgi Gerganov
451e9ee92c
make : remove "talk" target until updated
2024-09-24 19:45:08 +03:00
Georgi Gerganov
1133ac98a8
ggml : add ggml-cpu-impl.h (skip) ( #0 )
2024-09-24 19:45:08 +03:00
Georgi Gerganov
76d27eec9a
sync : ggml
2024-09-24 19:45:08 +03:00
Georgi Gerganov
fe18c29ab8
talk-llama : sync llama.cpp
2024-09-24 19:45:08 +03:00
Georgi Gerganov
3b183cfae7
log : add CONT level for continuing previous log entry (llama/9610)
2024-09-24 19:45:08 +03:00
Georgi Gerganov
896c41ef30
metal : use F32 prec for K*Q in vec FA (llama/9595)
...
ggml-ci
2024-09-24 19:45:08 +03:00
Georgi Gerganov
54e5095765
examples : adapt to ggml.h changes (ggml/0)
...
ggml-ci
2024-09-24 19:45:08 +03:00
Georgi Gerganov
34291099fb
ggml : refactoring (llama/#0)
...
- d6a04f87
- 23e0d70b
2024-09-24 19:45:08 +03:00
Georgi Gerganov
d245d7aec7
ggml : fix builds (llama/0)
...
ggml-ci
2024-09-24 19:45:08 +03:00
Georgi Gerganov
d661283e68
ggml : fix trailing whitespace (llama/0)
...
ggml-ci
2024-09-24 19:45:08 +03:00
Georgi Gerganov
1fd78999e8
cmake : do not hide GGML options + rename option (llama/9465)
...
* cmake : do not hide GGML options
ggml-ci
* build : rename flag GGML_CUDA_USE_GRAPHS -> GGML_CUDA_GRAPHS
for consistency
ggml-ci
2024-09-24 19:45:08 +03:00
Georgi Gerganov
a2cb5b4183
metal : handle zero-sized allocs (llama/9466)
2024-09-24 19:45:08 +03:00
Georgi Gerganov
288ae5176e
common : reimplement logging (llama/9418)
...
https://github.com/ggerganov/llama.cpp/pull/9418
2024-09-24 19:45:08 +03:00
Georgi Gerganov and Michael Podvitskiy
66b00fad0d
cmake : use list(APPEND ...) instead of set() + dedup linker (llama/9463)
...
* cmake : use list(APPEND ...) instead of set() + dedup linker
ggml-ci
* cmake : try fix sycl
* cmake : try to fix sycl 2
* cmake : fix sycl build (llama/9469)
* try fix sycl build
* use CMAKE_CXX_FLAGS as a string variable
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
* one more CMAKE_CXX_FLAGS fix (llama/9471)
---------
Co-authored-by: Michael Podvitskiy <podvitskiymichael@gmail.com >
2024-09-24 19:45:08 +03:00
Georgi Gerganov
a785232bf9
metal : fix compile warning with GGML_METAL_NDEBUG (llama/0)
2024-09-24 19:45:08 +03:00
Georgi Gerganov
26225f1fb0
cuda : fix FA Q src index (1 -> 0) (llama/9374)
2024-09-24 19:45:08 +03:00
Georgi Gerganov
253ce30004
examples : add null threadpool args where needed (ggml/0)
...
ggml-ci
2024-09-24 19:45:08 +03:00
Georgi Gerganov
03a6fae484
metal : update support condition for im2col + fix warning (llama/0)
2024-09-24 19:45:08 +03:00
Georgi Gerganov
5236f02784
revert : cmake : set MSVC to use UTF-8 on source files ( #2346 )
...
This reverts commit c96906d84d .
2024-09-02 15:24:50 +03:00
Georgi Gerganov
2abaf19e0d
sync : ggml
2024-09-02 15:24:50 +03:00
Georgi Gerganov
e8f0f9b5f0
cann : fix doxy (ggml/0)
2024-09-02 15:24:50 +03:00
Georgi Gerganov
d8e24b877d
vulkan : fix build (llama/0)
...
ggml-ci
2024-09-02 15:24:50 +03:00
Georgi Gerganov
cc68f31577
cuda : mark BF16 CONT as unsupported
2024-09-02 15:24:50 +03:00
Georgi Gerganov
e2e55a6fed
readme : fix link ( #2394 )
2024-08-30 13:58:22 +03:00
Georgi Gerganov
da9809f243
talk-llama : sync llama.cpp
2024-08-28 13:22:20 +03:00
Georgi Gerganov
9d754a56cf
whisper : update FA call
2024-08-28 13:22:20 +03:00
Georgi Gerganov
8cc90a0e80
sync : ggml
2024-08-28 13:22:20 +03:00