Skip to content

Tags: wrongbutworks/llama.cpp

Tags

b10341

Toggle b10341's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
readme : remove dev branches (ggml-org#26832)

b10182

Toggle b10182's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
llama: move suppress_tokens handling to common/sampling (ggml-org#26276)

* llama: move suppress_tokens handling to common/sampling

* address security issues

* rm has_logit_bias

b10181

Toggle b10181's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
ggml-cuda : disable MMQ on devices with less than 48 KiB shared memory (

ggml-org#26141)

ggml_cuda_should_use_mmq() selects MMQ purely from the quantization
type. The current MMQ configurations are designed and maintained against
a minimum of 48 KiB per-block shared memory, the limit provided by
NVIDIA Pascal GPUs and later. On devices that report less, no supported
MMQ tile fits and mul_mat_q_switch_J() aborts when every tile size
exceeds the device's per-block shared memory budget.

Disable MMQ when smpbo < 48 KiB so the caller falls back to the BLAS
path instead of hitting GGML_ABORT. Some current MUSA QY1 devices
report only 28 KiB and are covered by this guard.

Reproduced on a Moore Threads MTT S70 (arch mp_21, 28 KiB shared memory
per block) with an RWKV-7 0.1B Q8_0 model:

  $ llama-bench -m rwkv7-g1d-0.1b-Q8_0.gguf -p 128 -n 0
  J_best=0
  ggml/src/ggml-cuda/template-instances/../mmq.cuh:1521: fatal error
  (core dumped)

Only prefill (batch > 1) is affected; token generation is fine. After
the fix the same device falls back to the BLAS path:

  Q8_0    pp128 1470.7 t/s, tg8 55.3 t/s   (was: abort)
  FP16    unchanged
  Q4_K_M  unchanged

This matches a -DGGML_CUDA_FORCE_CUBLAS=ON build (pp128 1464.2 t/s),
which confirms the fallback path is the one being taken.

This is not MUSA-specific: any device with less than 48 KiB per-block
shared memory is affected.

Co-authored-by: KakaruHayate <KakaruHayate@users.noreply.github.com>

b10180

Toggle b10180's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
sycl: contiguous fast path + 32-bit index math for unary elementwise …

…ops (ggml-org#25946)

* sycl: contiguous fast path + 32-bit index math for unary elementwise ops

* sycl: use fastdiv for elementwise index math

b10179

Toggle b10179's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
vendor: update BoringSSL to 0.20260728.0 (ggml-org#26241)

b10178

Toggle b10178's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
server : add trace logging for slot similarity checking (ggml-org#26271)

Adds trace logging in server-context.cpp for slot similarity checking
during prompt cache slot selection, including skip reasons and similarity
calculation details.

Assisted-by: llama.cpp:Qwen3.6-27B

b10176

Toggle b10176's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
RPC: add tensor_memset (ggml-org#25912)

b10175

Toggle b10175's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
add rdna3.5, and 3 to mmq configs so they can be tuned independently. (

…ggml-org#26199)

b10174

Toggle b10174's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
model: add NextN/MTP speculative decoding support for GLM_DSA (GLM-5.…

…2) (ggml-org#25980)

* model: add NextN/MTP speculative decoding support for GLM_DSA (GLM-5.2)

Adds GLM-5.2 NextN/MTP as a --spec-type draft-mtp target: nextn tensor
loading via the qwen35moe/step35-style presence probe, a graph_mtp
builder (enorm/hnorm/eh_proj + dense MLA + sigmoid-gated MoE with
shared expert + shared head with fallbacks, _s scale tensors passed
for NVFP4), t_h_nextn extraction in the trunk graph, and MTP-context
KV setup: the draft head runs dense MLA, so the MTP context uses a
plain attention KV cache holding only the nextn layer(s) (same
pattern as the hybrid Qwen3.5 MTP context) while the main context
keeps the DSA cache, now filtered to trunk layers only.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* convert : support --mtp/--no-mtp export for GlmMoeDsaForCausalLM (GLM-5.2)

Opt GLM-5.2 into the supports_mtp_export contract (post-ggml-org#25641 shape,
mirroring HYV3Model/Step35Model): --no-mtp drops the appended NextN
block (blk.78) and its nextn_predict_layers KV; --mtp keeps only the
NextN block plus shared embeddings/norm/lm_head. Default (bundled)
output is unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

b10173

Toggle b10173's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
model: Add Laguna-S-2.1 LLM_TYPE (ggml-org#26233)