Unpatched LMCache flaw allows unauthenticated code execution on LLM cache servers

JFrog disclosed CVE-2026-105192, a 9.8-rated flaw in LMCache's multiprocess mode that lets a single network message run code. No fixed version is available yet.

Unpatched LMCache flaw allows unauthenticated code execution on LLM cache servers

At a glance

  • CVE-2026-105192 affects LMCache 0.3.9 through 0.5.5, the 0.5.6 release candidates and the development branch.
  • The ZeroMQ socket used by workers has no authentication and deserializes one message type with Python pickle.
  • The default binds to localhost, but LMCache's example Kubernetes deployment listens on all interfaces.
  • No patch or vendor advisory exists; JFrog advises keeping the port off routable networks.

A critical, unpatched vulnerability in LMCache, an open-source caching layer used to speed up large language model servers such as vLLM, can let unauthenticated attackers execute code with a single network message. According to The Hacker News, the flaw, tracked as CVE-2026-105192, was disclosed on October 7 by JFrog security researcher Yuval Moravchick, and no fixed version is available.

What happened

JFrog rated the issue 9.8 out of 10, the score it assigns when the server is bound to a routable address, The Hacker News reported. The bug sits in LMCache's multiprocess mode, where the cache runs as a standalone server that LLM worker processes reach over ZeroMQ. LMCache has not published a security advisory, and JFrog's advisory gives operators no way to determine whether a server has already been attacked.

Technical details

According to the report, the ZeroMQ socket used by worker processes has no authentication. One message type is deserialized with Python's pickle module, which can execute code while decoding, and the server unpacks a message's arguments before checking its type. As a result, a crafted message can run code with the privileges of the LMCache process, which JFrog says runs as root on the project's official container images.

The Hacker News notes that the issue resembles the class of pickle-over-network flaws described in November 2025 under the name ShadowMQ in other AI inference frameworks, though a common code origin has not been established. Separately, a GitHub user filed six additional LMCache security reports on October 6 alleging cross-tenant cache access and command execution in other network services; those claims come from a single account and have not been confirmed by maintainers or assigned CVEs.

A related vLLM bug, CVE-2026-105756 (rated 6.5), lets a malformed cache_salt value crash the engine when the LMCache multiprocess connector is used. It is a denial-of-service issue fixed in vLLM 0.30.0.

Who is affected

The vulnerable code is present from version 0.3.9, released in October 2025, through 0.5.5, the latest stable release, as well as the 0.5.6 release candidates and the development branch. By default the server listens only on localhost, so exposure depends on an operator binding it to a routable address. However, The Hacker News reports that LMCache's example Kubernetes deployment does exactly that by listening on all network interfaces. A single-process LMCache running inside vLLM does not open the port. The report gives no figures for exposed or compromised servers.

What to do

Until a fix ships, JFrog recommends not assigning the multiprocess server a routable address and keeping the port on the local machine or a trusted cluster network. Teams that adapted the example Kubernetes manifests should review their bind addresses and network policies. JFrog cautions that a firewall reduces the risk but does not eliminate it, since any host that can connect to the port can run code. Operators using the LMCache connector with vLLM should also upgrade vLLM to 0.30.0 to address the related crash bug.

Related CVEs

Sources

This story is based on the sources listed above. Always check the vendor’s official advisory before acting on critical systems.