CVE-2026-43632Ggml · Llama.cpp
Vulnerability data via NVD (ingested)
llama.cpp builds b7492 through the latest b9060 contains a use-after-free vulnerability in llama-server affecting six tokenization endpoints (/tokenize, /detokenize, /infill, /apply-template, /rerank, and /anthropic/count_tokens) that bypass the task queue and access ctx_server.vocab directly on HTTP worker threads. Attackers can exploit a time-of-check-time-of-use race condition where the main thread destroys and frees vocab after the synchronization lock is released but before the handler finishes using it, causing a crash or potential code execution when --sleep-idle-seconds is configured.
External references
Search for exposed instances
Shodan + Censys queries derived from NVD's CPE data. The vuln tag catches assets Shodan has explicitly linked to this CVE; the product / banner fingerprints find exposed instances even when the vuln tag was never applied (which is common).
vuln:CVE-2026-43632product:"Ggml Llama.cpp"http.html:"Llama.cpp"More intel sources (5)
vuln:CVE-2026-43632vulnerabilities.cve_id: CVE-2026-43632CVE-2026-43632CVE-2026-43632"CVE-2026-43632" exploit -site:nvd.nist.gov