Source: The Hacker News
JFrog disclosed a critical vulnerability in LMCache on October 7, and any team running distributed LLM inference should read the details before the next deployment window. LMCache is open-source software that speeds up large language model servers such as vLLM, and the flaw, tracked as CVE-2026-105192 and rated 9.8 out of 10, lets an attacker run code on the cache server without any authentication. No fixed version is available.
How the flaw works
The problem is in LMCache’s multiprocess mode, where the cache runs as a standalone server that LLM workers reach over the ZeroMQ messaging library. That socket has no authentication, and one message type is unpacked with pickle, a Python format that can execute code while data is decoded. The server unpacks the payload before it even checks the message type, so a single crafted network message runs the sender’s code with the privileges of the LMCache process. On the project’s official container images, that process runs as root. The flaw affects versions from 0.3.9, released in October 2025, through 0.5.5, the latest stable release, and is also present in the 0.5.6 release candidates and the development branch.
The configuration detail that decides everything
By default the multiprocess server listens only on the local machine, which keeps the flaw contained. The risk appears when an operator starts the server bound to a routable address, which is exactly what multi-node deployments do when they share a cache across machines, and it is how LMCache’s own example Kubernetes deployment is configured, listening on every network interface. JFrog’s guidance while no patch exists is to keep the port on localhost or a trusted cluster network, and to firewall it as tightly as possible, while accepting that any host that can still open the connection can run code. LMCache has not published a security advisory, and there is no documented way to determine whether a server has already been attacked. Separately, six unverified security reports alleging multi-tenant data access were filed on GitHub the day before disclosure, and a related denial-of-service flaw in vLLM, CVE-2026-105756, was fixed in version 0.30.0 released September 22.
Why this matters for European AI platforms
European organizations building inference infrastructure face pressure to ship fast, and this flaw is a textbook example of how an optimization component becomes the weakest link. The same pattern, unauthenticated sockets handing data to pickle, was seen across other AI inference frameworks in the ShadowMQ research from November 2025, which suggests a systemic issue across the GenAI tooling ecosystem rather than a one-off bug. If your platform runs distributed inference, this is the week to inventory what listens on routable interfaces inside your cluster.
Contact Excello Digital at https://excello.digital/contact/ if you want a security review of your AI inference stack. We help teams audit network exposure of internal services, harden container configurations, and put guardrails around the fast-moving open-source components your AI platform depends on.
