Releases · ggerganov/llama.cpp

18 Oct 22:26

cda0e4b

llama : remove all_pos_0, all_pos_1, all_seq_id from llama_batch (#9745)

* refactor llama_batch_get_one

* adapt all examples

* fix simple.cpp

* fix llama_bench

* fix

* fix context shifting

* free batch before return

* use common_batch_add, reuse llama_batch in loop

* null terminated seq_id list

* fix save-load-state example

* fix perplexity

* correct token pos in llama_batch_allocr

Assets 22

cudart-llama-bin-win-cu11.7.1-x64.zip

293 MB 2024-10-18T22:26:05Z
cudart-llama-bin-win-cu12.2.0-x64.zip

413 MB 2024-10-18T22:26:16Z
llama-b1-bin-win-hip-x64-gfx1030.zip

236 MB 2024-10-18T22:26:28Z
llama-b1-bin-win-hip-x64-gfx1100.zip

238 MB 2024-10-18T22:26:35Z
llama-b1-bin-win-hip-x64-gfx1101.zip

237 MB 2024-10-18T22:26:43Z
llama-b3943-bin-macos-arm64.zip

52.1 MB 2024-10-18T22:26:51Z
llama-b3943-bin-macos-x64.zip

53 MB 2024-10-18T22:26:53Z
llama-b3943-bin-ubuntu-x64.zip

58.7 MB 2024-10-18T22:26:55Z
llama-b3943-bin-win-avx-x64.zip

7.81 MB 2024-10-18T22:26:57Z
llama-b3943-bin-win-avx2-x64.zip

7.81 MB 2024-10-18T22:26:58Z
Source code (zip)

2024-10-18T21:18:01Z
Source code (tar.gz)

2024-10-18T21:18:01Z

18 Oct 19:00

github-actions

b3942

afd9909

b3942

rpc : backend refactoring (#9912)

* rpc : refactor backend

Use structs for RPC request/response messages

* rpc : refactor server

Assets 22

18 Oct 07:03

github-actions

b3941

87421a2

b3941

[SYCL] Add SYCL Backend registry, device and Event Interfaces (#9705)

* implemented missing SYCL event APIs

* sycl : Added device and backend reg interfaces

* Restructured ggml-sycl.cpp

Assets 22

18 Oct 06:30

github-actions

b3940

60ce97c

b3940

add amx kernel for gemm (#8998)

add intel amx isa detection

add vnni kernel for gemv cases

add vnni and amx kernel support for block_q8_0

code cleanup

fix packing B issue

enable openmp

fine tune amx kernel

switch to aten parallel pattern

add error message for nested parallelism

code cleanup

add f16 support in ggml-amx

add amx kernels for QK_K quant formats: Q4_K, Q5_K, Q6_K and IQ4_XS

update CMakeList

update README

fix some compilation warning

fix compiler warning when amx is not enabled

minor change

ggml-ci

move ggml_amx_init from ggml.c to ggml-amx/mmq.cpp

ggml-ci

update CMakeLists with -mamx-tile, -mamx-int8 and -mamx-bf16

ggml-ci

add amx as an ggml-backend

update header file, the old path for immintrin.h has changed to ggml-cpu-impl.h

minor change

update CMakeLists.txt

minor change

apply weight prepacking in set_tensor method in ggml-backend

fix compile error

ggml-ci

minor change

ggml-ci

update CMakeLists.txt

ggml-ci

add march dependency

minor change

ggml-ci

change ggml_backend_buffer_is_host to return false for amx backend

ggml-ci

fix supports_op

use device reg for AMX backend

ggml-ci

minor change

ggml-ci

minor change

fix rebase

set .buffer_from_host_ptr to be false for AMX backend

Assets 22

18 Oct 05:20

github-actions

b3939

8901755

b3939

server : add n_indent parameter for line indentation requirement (#9929)

ggml-ci

Assets 22

18 Oct 00:32

github-actions

b3938

6f55bcc

b3938

llama : rename batch_all to batch (#8881)

This commit addresses the TODO in the code to rename the `batch_all`
parameter to `batch` in `llama_decode_internal`.

Assets 22

17 Oct 22:15

github-actions

b3936

9f45fc1

b3936

llama : change warning to debug log

Assets 22

17 Oct 21:50

github-actions

b3935

99bd4ac

b3935

llama : infill sampling handle very long tokens (#9924)

* llama : infill sampling handle very long tokens

ggml-ci

* cont : better indices

ggml-ci

Assets 22

17 Oct 01:44

github-actions

b3933

f010b77

b3933

vulkan : add backend registry / device interfaces (#9721)

* vulkan : add backend registry / device interfaces

* llama : print devices used on model load

Assets 22

17 Oct 00:46

github-actions

b3932

2194200

b3932

fix: allocating CPU buffer with size `0` (#9917)

Assets 22

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Releases: ggerganov/llama.cpp

b3943

b3942

b3941

b3940

b3939

b3938

b3936

b3935

b3933

b3932