DeepSeek recently published a system paper, disclosing for the first time the operational status of its internal sandbox platform, DSec. The paper states that from DeepSeek V3.2 to V4.1, all sandbox experiments involving reinforcement learning training and evaluation were conducted on this platform, allowing the outside world to see for the first time the core infrastructure of its Agent training environment.
Approximately 3 million sandboxes are served in a single day.
The paper shows that DSec is a sandbox platform designed for training with Agent, mainly used to provide an isolated, recoverable, and batch-createable execution environment. Unlike traditional large model training, Agent training requires the use of code repositories, compilation tools, browsers, and various external tools, and the execution process continuously changes the state of the environment, thus placing higher demands on the sandbox system.
According to the data disclosed in the paper, a production unit contains approximately 160 CPU nodes, 30,000 CPU cores, and 250 TB memories, and manages PB levels of image data. The platform can serve about 3 million sandbox instances per day, with a peak concurrency of over 380,000 instances, and a creation rate of more than 5,000 instances per second. A single training task can launch up to 32,000 sandboxes at a time.
Layered images reduce startup overhead.
The paper states that the main challenge with DSec is that different tasks require different codes, dependencies, and toolchains. If the image has to be downloaded in its entirety every time it is started, the cluster's I/O load will rise rapidly.
For this reason, DeepSeek divides the basic system, task workspaces, and toolkits into composable environmental layers that can be assembled as needed. Image distribution relies on its proprietary 3FS distributed file system, which reads EROFS images on demand rather than pulling the entire package at once. Experimental results presented in the paper show that in a scenario where 8,192 containers are started simultaneously, this approach can improve startup efficiency by about 42%, reduce disk write volume by about 57%, and increase task completion time by approximately 1.7 times.
In terms of resource utilization, the paper states that approximately 90% of the actual usage of CPU does not exceed 5% of the applied amount. Based on this characteristic, the platform adopts a high-density deployment and resource overselling strategy, with a single node capable of accommodating up to 3200 containers or 800 microVM. Combined with a memory sharing and recycling mechanism, this approach reduces peak memory usage by about 40%.
Behavior to bypass restrictions has been observed during training.
The paper also mentions that Agent attempts to bypass restrictions during the training process, including searching for residual files, forging RPC, evading access controls, and attempting to read protected content. DeepSeek uses AppArmor and eBPF domain-level network whitelists for protection within the system, but the paper also clearly states that no single mechanism can prevent all abnormal behaviors and system failures.
This paper has been submitted to arXiv. The list of authors exceeds 130 people, with Liang Wenfeng, the founder of DeepSeek, signing at the end. Public information indicates that the paper focuses on the field of distributed computing, with an emphasis not on the model itself, but on the underlying system that supports large-scale Agent reinforcement learning training.


Additional information:The original text is from a reprint by Wall Street Insights. The core data and system descriptions in the article are all cited from the paper "_DeepSeek Elastic Compute (DSec): _A Sandbox Infrastructure for Effective Agentic Training at Scale" submitted to _arXiv by _DeepSeek.












