1 #30

apicalshark · 2024-12-19T14:11:28Z

Make sure to read the contributing guidelines before submitting a PR

* ggml_pad_reflect_1d defined in header * implemented on CPU * called the forward pass * impl Metal kernel * added Metal kernel * added OP_PAD_REFLECT_1D in test-backend-ops.cpp * add test-pad-reflect-1d test case * test case support multiple backend

* implemented cpu kernel * add i32 test cases in test-backend-ops * typedef `ggml_metal_kargs_set` * implemented `kernel_set` * memcpy

* Support for Minerva 7B * Update convert_hf_to_gguf_update.py

…ttention (ggerganov#10206)

* server : (refactoring) reduce usage of json internally * move all response types to struct * wip [no ci] * many fixes * add virtual function * fix index * minor style fix * add std::move * refactor handle_completions_generic * add virtual functions * remove server.hpp * clarify server_sent_event RFC specs * apply review comments * fix model_alias and completion_probabilities * small clean up * remove virtual for to_json_oai_compat() * naming oai_compat --> oaicompat * fix unwanted recursive call * update docs

* metal : Extend how Llama.cpp locates metal resources (ggerganov#10675) * It searches the resource file in the directory where the current binary is located as well. * Resolves symbolic links. Rationale: When we plug this dependency into a Bazel build and run it in the context of Bazel (e.g. testing): * the execution directory is often very different from where the files are located and no direct control over this (Bazel sandboxing), * the Bazel sandbox often use symbolic links to make files available. With this patch, we can have the resource file added to the target, can build and run tests in the context of Bazel. * Update ggml/src/ggml-metal/ggml-metal.m Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> * Update ggml/src/ggml-metal/ggml-metal.m Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

…ng (ggerganov#10597) * Vulkan: Implement VK_KHR_cooperative_matrix support in the matrix matrix multiplication shader * Improve performance with better q4_k and q5_k dequant and store unrolling * Add Vulkan MUL_MAT and MUL_MAT_ID accumulator precision selection * Rework mulmat shader selection and compilation logic, avoid compiling shaders that won't get used by device * Vulkan: Implement accumulator switch for specific mul mat mat shaders * Vulkan: Unroll more loops for more mul mat mat performance * Vulkan: Add VK_AMD_shader_core_properties2 support to read Compute Unit count for split_k logic * Disable coopmat support on AMD proprietary driver * Remove redundant checks * Add environment variable GGML_VK_DISABLE_COOPMAT to disable VK_KHR_cooperative_matrix support * Fix rebase typo * Fix coopmat2 MUL_MAT_ID pipeline selection

ggml-ci

* rename ggml-cpu-aarch64.c to .cpp * reformat extra cpu backend. - clean Q4_0_N_M and IQ4_0_N_M - remove from "file" tensor type - allow only with dynamic repack - extract cpu extra bufts and convert to C++ - hbm - "aarch64" - more generic use of extra buffer - generalise extra_supports_op - new API for "cpu-accel": - amx - aarch64 * clang-format * Clean Q4_0_N_M ref Enable restrict on C++ * add op GGML_OP_MUL_MAT_ID for Q4_0_N_M with runtime repack * added/corrected control on tensor size for Q4 repacking. * Update ggml/src/ggml-cpu/ggml-cpu-aarch64.cpp Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> * Update ggml/src/ggml-cpu/ggml-cpu-aarch64.cpp Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> * add debug logs on repacks. --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* server : various fixes ggml-ci * server : show curent seed in slot_params ggml-ci * fix /slots endpoint * Update examples/server/server.cpp Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> * server : reflect endpoint response changes in the readme ggml-ci --------- Co-authored-by: Xuan Son Nguyen <son@huggingface.co> Co-authored-by: Xuan Son Nguyen <thichthat@gmail.com>

ggml-ci

* server : (refactor) no more json in server_task input * add test for slots endpoint * add tests for /props and /slots * remove task inf_type * fix CI by adding safe_json_to_str * add "model_path" to /props * update readme

* add 128k yarn context for Qwen * added property for model tensors * removing useless line

…gerganov#10713)

* llama : use cmake for swift build * swift : <> -> "" * ci : remove make * ci : disable ios build * Revert "swift : <> -> """ This reverts commit d39ffd9. * ci : try fix ios build * ci : cont * ci : cont --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

…gerganov#10723) * Vulkan: fix NaN in tanh.comp * Faster NaN-free tanh

* server : bring back into to final chunk in stream mode * clarify a bit * traling space

* server : fix format_infill * fix * rename * update test * use another model * update test * update test * test_invalid_input_extra_req

…erganov#10668) * Update cmakepreset.json to use clang with ninja by default * Update cmakepreset.json to add clang and ninja based configs * Updates to build.md file * Make updates to rename preset targets * Update with .cmake file * Remove additional whitespaces * Add .cmake file for x64-windows-llvm * Update docs/build.md * Update docs/build.md --------- Co-authored-by: Max Krasnyansky <max.krasnyansky@gmail.com>

There are some bugs in the 1.3.296 SDK, so disable this. It isn't strictly necessary anyway. Add missing dependency on vulkan-shaders-gen, so shaders get recompiled when it changes. Fix coopmat support reporting when glslc doesn't support NV_coopmat2.

…10751) Co-authored-by: eugenio.segala <esegala@deloitte.co.uk>

* Renames NVIDIA GPU-architecture flags to avoid name clashes with WinAPI. (e.g. CC_PASCAL, GPU architecture or WinAPI pascal compiler flag?) * Reverts erroneous rename in SYCL-code. * Renames GGML_CUDA_MIN_CC_DP4A to GGML_CUDA_CC_DP4A. * Renames the rest of the compute capability macros for consistency.

…erganov#10872) * server : (embeddings) using same format for "input" and "content" * fix test case * handle empty input case * fix test

* server : add "tokens" output ggml-ci * server : update readme ggml-ci * server : return tokens ids only if requested ggml-ci * tests : improve "tokens" type check Co-authored-by: Xuan Son Nguyen <thichthat@gmail.com> * server : remove "tokens" from the OAI endpoint ggml-ci --------- Co-authored-by: Xuan Son Nguyen <thichthat@gmail.com>

…nov#10861) * server : add "tokens" output ggml-ci * server : output embeddings for all tokens when pooling = none ggml-ci * server : update readme [no ci] * server : fix spacing [no ci] Co-authored-by: Xuan Son Nguyen <thichthat@gmail.com> * server : be explicit about the pooling type in the tests ggml-ci * server : update /embeddings and /v1/embeddings endpoints ggml-ci * server : do not normalize embeddings when there is no pooling ggml-ci * server : update readme ggml-ci * server : fixes * tests : update server tests ggml-ci * server : update readme [no ci] * server : remove rebase artifact --------- Co-authored-by: Xuan Son Nguyen <thichthat@gmail.com>

* server: avoid overwriting Authorization header If no API key is set, leave the Authorization header as is. It may be used by another part of the Web stack, such as an authenticating proxy. Fixes ggerganov#10854 * rebuild --------- Co-authored-by: Xuan Son Nguyen <son@huggingface.co>

* server : add "tokens" output ggml-ci * server : output embeddings for all tokens when pooling = none ggml-ci * server : be explicit about the pooling type in the tests ggml-ci * server : do not normalize embeddings when there is no pooling ggml-ci * llama : add OuteTTS support (wip) * wip * extract features * first conv * group norm * resnet conv * resnet * attn * pos net * layer norm * convnext * head * hann window * fix n_embd + remove llama.cpp hacks * compute hann window * fft * spectrum processing * clean-up * tts : receive input text and generate codes * clip : fix new conv name * tts : minor fix * tts : add header + minor fixes ggml-ci * tts : add matchematical constant ggml-ci * tts : fix sampling + cut initial noise * tts : fixes * tts : update default samplers ggml-ci * tts : text pre-processing * tts : outetts-voc -> wavtokenizer-dec * tts : remove hardcoded constants ggml-ci * tts : fix tensor shapes * llama : refactor wavtokenizer tensors ggml-ci * cont ggml-ci * cont [no ci] * llama : update WavTokenizer to non-causal attn * llama : handle no-vocab detokenization * tts : add Python example for OuteTTS (wip) * tts : extend python example to generate spectrogram ggml-ci * server : fix rebase artifacts * tts : enable "return_tokens" in Python example ggml-ci * tts : minor fixes * common : support HF download for vocoder

* ggml: GGML_NATIVE uses -mcpu=native on ARM Signed-off-by: Adrien Gallouët <angt@huggingface.co> * ggml: Show detected features with GGML_NATIVE Signed-off-by: Adrien Gallouët <angt@huggingface.co> * remove msvc support, add GGML_CPU_ARM_ARCH option * disable llamafile in android example * march -> mcpu, skip adding feature macros ggml-ci --------- Signed-off-by: Adrien Gallouët <angt@huggingface.co> Co-authored-by: Adrien Gallouët <angt@huggingface.co>

Set default width to whatever the terminal is. Also fixed a small bug around default n_gpu_layers value. Signed-off-by: Eric Curtin <ecurtin@redhat.com>

* convert : use GPT2 vocab for Phi-4 model * convert : use null value of sliding_window to distinguish Phi-4 from other PHI3-based models * llama : do not use sliding window attention mask for Phi-4 model --------- Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>

* fix: Use gpt2 tokenizer for roberta and add eos/bos tokens Branch: RobertaTokenizer Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fixes to position embeddings Signed-off-by: Sukriti-Sharma4 <sukriti.sharma4@ibm.com> * map roberta-bpe to gpt-2 Signed-off-by: Sukriti-Sharma4 <sukriti.sharma4@ibm.com> * fix linting Signed-off-by: Sukriti-Sharma4 <sukriti.sharma4@ibm.com> --------- Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> Signed-off-by: Sukriti-Sharma4 <sukriti.sharma4@ibm.com> Co-authored-by: Gabe Goodhart <ghart@us.ibm.com>

Signed-off-by: Adrien Gallouët <angt@huggingface.co>

PABannier and others added 30 commits December 5, 2024 13:27

ggml: add GGML_SET Metal kernel + i32 CPU kernel (ggml/1037)

a8cbab2

* implemented cpu kernel * add i32 test cases in test-backend-ops * typedef `ggml_metal_kargs_set` * implemented `kernel_set` * memcpy

sync : ggml

0cd182e

llama : add Minerva 7B model support (ggerganov#10673)

6fe6247

* Support for Minerva 7B * Update convert_hf_to_gguf_update.py

vulkan: Add VK_NV_cooperative_matrix2 support for mul_mat and flash a…

c9c6e01

…ttention (ggerganov#10206)

fix(server) : not show alert when DONE is received (ggerganov#10674)

7736837

common : bring back --no-warmup to server (ggerganov#10686)

f162d45

convert : add custom attention mapping

c5ede38

convert : add support for Roberta embeddings (ggerganov#10695)

784a14a

server : fix free of spec context and batch (ggerganov#10651)

c2a16c0

ggml-ci

ggml : disable iq4_nl interleave size 8 (ggerganov#10709)

d9c3ba2

ggml-ci

llama : add 128k yarn context for Qwen (ggerganov#10698)

62e84d9

* add 128k yarn context for Qwen * added property for model tensors * removing useless line

vulkan: compile a test shader in cmake to check for coopmat2 support (g…

ecc93d0

…gerganov#10713)

Vulkan: fix NaN in tanh.comp with AMD proprietary driver on Windows (g…

06d7014

…gerganov#10723) * Vulkan: fix NaN in tanh.comp * Faster NaN-free tanh

server : bring back info of final chunk in stream mode (ggerganov#10722)

e52522b

* server : bring back into to final chunk in stream mode * clarify a bit * traling space

server : fix format_infill (ggerganov#10724)

ce8784b

* server : fix format_infill * fix * rename * update test * use another model * update test * update test * test_invalid_input_extra_req

cmake : simplify msvc charsets (ggerganov#10672)

1a05004

vulkan: fix compile warnings (ggerganov#10731)

3d98b4c

CUDA: fix shared memory access condition for mmv (ggerganov#10740)

26a8406

server : add flag to disable the web-ui (ggerganov#10762) (ggerganov#…

a86ad84

…10751) Co-authored-by: eugenio.segala <esegala@deloitte.co.uk>

ngxson and others added 11 commits December 18, 2024 10:55

server : (embeddings) using same format for "input" and "content" (gg…

4682887

…erganov#10872) * server : (embeddings) using same format for "input" and "content" * fix test case * handle empty input case * fix test

llama-run : improve progress bar (ggerganov#10821)

7909e85

Set default width to whatever the terminal is. Also fixed a small bug around default n_gpu_layers value. Signed-off-by: Eric Curtin <ecurtin@redhat.com>

tests: disable GGUF test for bad value size (ggerganov#10886)

cd920d0

ggml: fix arm build with gcc (ggerganov#10895)

a3c33b1

Signed-off-by: Adrien Gallouët <angt@huggingface.co>

github-actions bot added documentation Improvements or additions to documentation examples server build devops nix testing python script ggml SYCL Nvidia GPU Vulkan Kompute Apple Metal android labels Dec 19, 2024

Merge branch 'master' into 1

4ec10e6

apicalshark merged commit aca6a17 into master Dec 19, 2024
5 of 8 checks passed

apicalshark deleted the 1 branch December 19, 2024 14:13

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

1 #30

1 #30

apicalshark commented Dec 19, 2024

1 #30

1 #30

Conversation

apicalshark commented Dec 19, 2024