CVE-2025-49847
17.06.2025, 20:15
llama.cpp is an inference of several LLM models in C/C++. Prior to version b5662, an attackersupplied GGUF model vocabulary can trigger a buffer overflow in llama.cpps vocabularyloading code. Specifically, the helper _try_copy in llama.cpp/src/vocab.cpp: llama_vocab::impl::token_to_piece() casts a very large size_t token length into an int32_t, causing the length check (if (length < (int32_t)size)) to be bypassed. As a result, memcpy is still called with that oversized size, letting a malicious model overwrite memory beyond the intended buffer. This can lead to arbitrary memory corruption and potential code execution. This issue has been patched in version b5662.Enginsight
Awaiting analysis
This vulnerability is currently awaiting analysis.
Common Weakness Enumeration