on this page
SGLang lets operators disable direct Pickle transport for its internal messages. I expected that choice to remove Pickle from the receive path. It changed the envelope and left the decoder where it was.
I ran a stock SGLang v0.5.18 service with a real model across two GPUs. I made one of its internal receivers reachable from a second machine, on purpose, and sent it a message from a peer that had not authenticated. Code chosen by that peer ran inside the receiver and wrote a small marker file containing its own process ID. The ID matched the receiver the service had launched. It did this with the setting on and with the setting off, on three fresh starts each.
Two questions fall out of that, and they need separate answers. Why did disabling direct Pickle transport still reach Pickle? And when can anyone outside the service send bytes to that decoder at all? The second matters as much as the first. In the default single-node layout, nobody on another machine can. The route I tested needed a particular configuration, and I will be precise about which one.
The switch changed the envelope
SGLang runs a model as several cooperating processes. They pass work and results to each other over ZeroMQ, a messaging library that can carry messages over local sockets or over TCP. This is inter-process communication, IPC for short. The word describes the relationship between the processes. It says nothing about who can reach the connection.
The receiver I followed was the detokenizer, the process that turns the model’s output token IDs back into text. In normal operation it receives messages from the scheduler. Before it can act on one, it has to turn the incoming bytes into a Python value.
Python Pickle does more than read fields from a data container. Its reconstruction instructions can identify a function and call it with supplied arguments. Those calls happen while loading the object. The Python documentation therefore warns against unpickling untrusted data.
In the tested release, SGLANG_USE_PICKLE_IPC defaults to true. The receive function calls PyZMQ’s recv_pyobj, which reconstructs the received Python object using Pickle. The PyZMQ documentation makes the same trust requirement explicit.
Setting the flag to false selects MessagePack instead. MessagePack represents data such as numbers, byte strings and structured records. Choosing it can remove the need to reconstruct arbitrary Python objects. Here, though, the accepted message types included PickleWrapper, a record carrying Pickle bytes.
The wrapper was accepted as a top-level message. Once MessagePack had decoded that record, SGLang immediately unwrapped it with Python’s Pickle decoder. The outer format had changed, but the sender could still reach the operation that invokes functions.
This was an explicit compatibility path in the decoder. The wrapper let the typed transport carry Python objects that its ordinary message types could not represent. The security consequence was that a peer could select that same path.
The check came after the effect
The relevant code is small enough to follow without a payload. The following is condensed pseudocode for the v0.5.18 receive and unwrap functions. It omits logging and other message types.
# Condensed pseudocode, not an exploit client.
def receive(socket):
if use_pickle_ipc:
return socket.recv_pyobj()
message = typed_msgpack_decode(socket.recv())
if isinstance(message, PickleWrapper):
return pickle.loads(message.data)
return message
# The receiver checks the reconstructed value afterward.
message = receive(socket)
dispatch(message)
There are two different checks here. The MessagePack decoder checks whether the outer record belongs to an accepted type. It accepts the wrapper. Later, the application dispatcher checks whether the reconstructed result is a message the detokenizer knows how to handle.
Neither check prevents a callable from running during the intervening Pickle reconstruction. The receiver’s event loop does not call its dispatcher until sock_recv has returned.
That order was visible in the experiment. Reconstruction ran the marker-writing action. Dispatch then received a value it did not recognise, rejected it, and the detokenizer exited. The error came after the code had already run.
Both transport choices could reach Pickle reconstruction. Compare the two formats.
An interactive cutaway switches between direct Pickle and a MessagePack envelope. The same highlighted Pickle core remains in both views. A shared receive sequence shows the marker written during reconstruction; the later rejection cannot undo that effect.
A typed envelope with a compatibility wrapper inside.
Changing the envelope did not remove this operation.
A check on the finished object cannot undo work done while building it. A regression test that only expects an invalid-message error would miss this. It also has to show that the rejection happened without the supplied action running.
When an internal receiver becomes reachable
Finding Pickle on an internal path does not tell you who can send to it. The next thing I followed was how the receiver gets its address.
The tested source has three branches. Without data-parallel attention, the component receivers use Unix-domain sockets, which never leave the host. Data-parallel attention is a serving mode that spreads attention work across parallel workers, and turning it on moves those receivers to TCP. Even then they bind to loopback, unless the launch also supplies a distribution address. Give it an address that is not loopback and the receivers bind there instead.
So a route from another machine needs all three. Data-parallel attention, a non-loopback distribution address, and a network path to the receiver’s port. That is a supported configuration, and it is the one I tested. It is not the default.
In the tested non-default setup, HTTP stayed on loopback while a second authorized host could reach the internal receiver.
A clearly bounded SGLang service host contains the HTTP API, the scheduler and the detokenizer receiver. A quiet return path at HTTP represents the successful local inference check. A second quiet path carries normal internal messages from the scheduler to the detokenizer. Outside the boundary, a separate authorized test host sends inward to the receiver. A small warm aperture marks the non-loopback TCP binding at the host boundary. Beside it, the tested setup is qualified as data-parallel attention plus a reachable distribution binding. The HTTP API stays inside the host. This schematic shows the tested non-default arrangement, not default remote exposure or delivery through HTTP.
I kept the public HTTP endpoint on loopback and exposed only the detokenizer’s receiver to a second authorized machine. The crafted message went to that internal receiver. I did not demonstrate delivery through the inference API, and I am not claiming there is one.
The operational point is narrow, and worth saying plainly. A loopback HTTP endpoint does not keep a separately bound component receiver local. If the service runs with data-parallel attention and a distribution address, its internal receivers are listening on that address, and the network rules around them are doing the work the serialization setting cannot.
From a code path to a running service
The first evidence was the source. Reading it gave me the shape of the problem, and a reconstruction of the receive logic confirmed how the decoder behaved. That is enough to know the code is unsafe as written. It is not enough to know that the process an ordinary launch creates will accept the same bytes, and that is the claim that matters.
So I moved up one tier at a time. A local campaign reached the launcher-created receiver over Unix-domain sockets. The step after that was a real service with a receiver reachable from a second physical machine, which tests the deployment condition and the decoder together.
Getting a listener to bind was the easy part. Several smaller model fixtures never produced a working inference baseline under the two-GPU layout, and I did not count any of them. A marker from an unhealthy launch would not demonstrate the effect in a working inference service. The final campaign used the Qwen3-30B-A3B model that SGLang’s own two-GPU test exercises, and every trial had to answer an ordinary request before the crafted message was sent.
The service was stock v0.5.18 at commit 71de97b264b04dcd514cf904003028aefe9775c8, on one host with two NVIDIA A100-SXM4-80GB GPUs, tensor and data parallelism both set to two, data-parallel attention on. The second machine only sent messages. It was not another model worker.
The witness was a marker write. Reconstruction selected a callable and its arguments, and that callable wrote the receiver’s process ID to a file. I wanted an effect I could inspect and tie to the listening process, and nothing beyond that.
That distinction carries the result. A crash would show the input disrupted the service. Here the selected action finished before the crash, inside the process that owned the receiver. That is code execution.
What six fresh starts established
I started the service three times with the setting on and three times with it off. Each trial began with a successful ordinary request. Then I connected from the second machine without sending anything. Connecting alone produced no marker.
Sending the crafted message produced exactly one marker in every trial, and every marker’s process ID matched the receiver the launcher had created. Three out of three with raw Pickle. Three out of three with MessagePack carrying a PickleWrapper.
Two aligned groups show three fresh starts with the setting true, the default Direct Pickle transport, and three with it false, the MessagePack transport. Each start has four literal outcomes: Yes, the ordinary request answered; zero markers after connecting without sending a message; one marker after the crafted message; and Match, the marker process ID equals the listener process ID. The marker-written row is highlighted. All connect-only outcomes are recorded in the operator transcript. For the first default-setting trial, the baseline and listener identity are available only in that transcript, indicated by daggers. The connection-only control is not a matched benign-message comparison.
| Observation | Setting true · default · Direct Pickle | Setting false · MessagePack | ||||
|---|---|---|---|---|---|---|
| Trial 1 | Trial 2 | Trial 3 | Trial 1 | Trial 2 | Trial 3 | |
| Baseline answered: ordinary request | Yes: ordinary request answered; baseline and listener identity are recorded only in the operator transcript for this trial | Yes: ordinary request answered | Yes: ordinary request answered | Yes: ordinary request answered | Yes: ordinary request answered | Yes: ordinary request answered |
| Connect-only markers: no message sent | Zero markers after connecting without sending a message; operator transcript | Zero markers after connecting without sending a message; operator transcript | Zero markers after connecting without sending a message; operator transcript | Zero markers after connecting without sending a message; operator transcript | Zero markers after connecting without sending a message; operator transcript | Zero markers after connecting without sending a message; operator transcript |
| Marker written: crafted message | One marker after the crafted message | One marker after the crafted message | One marker after the crafted message | One marker after the crafted message | One marker after the crafted message | One marker after the crafted message |
| Marker process: matches receiver | Marker process ID equals listener process ID; baseline and listener identity are recorded only in the operator transcript for this trial | Marker process ID equals listener process ID | Marker process ID equals listener process ID | Marker process ID equals listener process ID | Marker process ID equals listener process ID | Marker process ID equals listener process ID |
All six connect-only outcomes live in one hash-bound operator transcript. Five of the six trials also keep separate state, inference, marker and service records. The first default-mode trial keeps its baseline and listener identity only in that transcript. And a connect-only check is a control for the connection, not a matched benign message. It rules out the handshake as the cause of the marker. It does not compare a crafted message against a well-formed one.
Taken together, this is remote code execution in the reachable receiver under the tested configuration. The peer did not authenticate. The bytes it supplied selected an action, and that action ran with the receiver’s own permissions, in both receive modes.
Scope of the experimentThe witness was a deliberately small marker write. The receiver exited afterwards, so I did not establish continued service. Remote delivery needed access to the internal receiver in the configuration described above. I did not establish default remote exposure or a path through the HTTP API, and I have not observed this exploited in the wild.
What a repair has to change
Disabling direct Pickle transport did not mitigate this, because an accepted message type brought Pickle back after the outer decode. A repair has to follow the bytes through the whole receive path, not just the first function that touches them.
For any transport a peer can write to, I would remove both the raw recv_pyobj path and PickleWrapper from the accepted message types. Every field needs a data-only representation the decoder can validate without calling a function the sender picked. Once that validation has passed, peer-controlled bytes must not reach pickle.loads through an adapter further down.
The envelope can stay. The authority to rebuild arbitrary objects must go.
Matched cutaways compare the tested MessagePack path and a proposed data-only wire format. The tested envelope contains Pickle bytes and still invokes object reconstruction. The proposed format contains only typed data fields, without an executable reconstruction step. The proposed receiver authenticates the peer and validates the data before dispatch. This is a proposal, not a shipped or verified repair.
A later rejection cannot undo an action performed while decoding.
Unknown types and invalid fields are rejected before dispatch.
Peers also need to authenticate, and to be authorized, before the decoder accepts work from them. That narrows who can send. A non-executing format narrows what an accepted sender can make reconstruction do. Both belong in the design, and neither replaces the other.
Where cross-host traffic is unnecessary, keep the receivers local. Where TCP is required, restrict it at the host and network layers to the components that need it. That limits which peers can reach the decoder while the path is being repaired; the serialization setting did not prevent execution in these tests.
I would test a repair against both raw and wrapped inputs, and I would check for zero side effects as well as for rejection. Valid typed messages should still reach their handlers. That is the property the repair has to restore. An exception on its own does not show it.
Coordination and source status
CERT/CC received the report on July 14, 2026. On September 17 the coordinator confirmed it as arbitrary code execution under the non-loopback, data-parallel-attention condition. The case is VU#765030, and the identifier assigned to it is CVE-2026-93034.
The executed results above are for v0.5.18. A separate commit-pinned source inspection on October 1 found the same receive and unwrap primitives in v0.5.20 and in the main revision inspected that day. That is source evidence. It is not a runtime test of those versions, and it is not a claim about a continuous affected range.
A public release check on October 7 found v0.5.21, released after that source inspection. This post has not assessed that release’s receive path or runtime behavior, so the October 1 inspection does not establish its fix status.
There is an earlier public SGLang report involving Pickle in shm_broadcast.MessageQueue. It concerns a different receive path. This post follows the manager-message decoder and the launcher-created detokenizer.
What stayed with me is how reasonable the setting looked until I followed its other branch. A typed outer format can still carry an instruction to rebuild an object the executable way. The question to ask of any decoder is whether some accepted message can get back to that operation before the application decides what it is willing to handle.
This post documents CRUCIBLE-2026-131, coordinated through CERT/CC as VU#765030 and assigned CVE-2026-93034. The full-service experiments used stock SGLang v0.5.18.