DeepSeek Reveals Agent Training Sandbox DSec
Wallstreetcn
09-24 17:43
Ai Focus
DeepSeek makes its debut; Agent Training Sandbox Platform, revealing its support for V3.2 to V4.1 reinforcement learning training and evaluation.
Helpful
No.Help

DeepSeek recently published a system paper, disclosing for the first time the operational status of its internal sandbox platform, DSec. The paper states that from DeepSeek V3.2 to V4.1, all sandbox experiments involving reinforcement learning training and evaluation were conducted on this platform, allowing the outside world to see for the first time the core infrastructure of its Agent training environment.

Approximately 3 million sandboxes are served in a single day.

The paper shows that DSec is a sandbox platform designed for training with Agent, mainly used to provide an isolated, recoverable, and batch-createable execution environment. Unlike traditional large model training, Agent training requires the use of code repositories, compilation tools, browsers, and various external tools, and the execution process continuously changes the state of the environment, thus placing higher demands on the sandbox system.

According to the data disclosed in the paper, a production unit contains approximately 160 CPU nodes, 30,000 CPU cores, and 250 TB memories, and manages PB levels of image data. The platform can serve about 3 million sandbox instances per day, with a peak concurrency of over 380,000 instances, and a creation rate of more than 5,000 instances per second. A single training task can launch up to 32,000 sandboxes at a time.

Layered images reduce startup overhead.

The paper states that the main challenge with DSec is that different tasks require different codes, dependencies, and toolchains. If the image has to be downloaded in its entirety every time it is started, the cluster's I/O load will rise rapidly.

For this reason, DeepSeek divides the basic system, task workspaces, and toolkits into composable environmental layers that can be assembled as needed. Image distribution relies on its proprietary 3FS distributed file system, which reads EROFS images on demand rather than pulling the entire package at once. Experimental results presented in the paper show that in a scenario where 8,192 containers are started simultaneously, this approach can improve startup efficiency by about 42%, reduce disk write volume by about 57%, and increase task completion time by approximately 1.7 times.

In terms of resource utilization, the paper states that approximately 90% of the actual usage of CPU does not exceed 5% of the applied amount. Based on this characteristic, the platform adopts a high-density deployment and resource overselling strategy, with a single node capable of accommodating up to 3200 containers or 800 microVM. Combined with a memory sharing and recycling mechanism, this approach reduces peak memory usage by about 40%.

Behavior to bypass restrictions has been observed during training.

The paper also mentions that Agent attempts to bypass restrictions during the training process, including searching for residual files, forging RPC, evading access controls, and attempting to read protected content. DeepSeek uses AppArmor and eBPF domain-level network whitelists for protection within the system, but the paper also clearly states that no single mechanism can prevent all abnormal behaviors and system failures.

This paper has been submitted to arXiv. The list of authors exceeds 130 people, with Liang Wenfeng, the founder of DeepSeek, signing at the end. Public information indicates that the paper focuses on the field of distributed computing, with an emphasis not on the model itself, but on the underlying system that supports large-scale Agent reinforcement learning training.

Additional information:The original text is from a reprint by Wall Street Insights. The core data and system descriptions in the article are all cited from the paper "_DeepSeek Elastic Compute (DSec): _A Sandbox Infrastructure for Effective Agentic Training at Scale" submitted to _arXiv by _DeepSeek.

Tip
$0
Like
0
Save
0
Views 72
HQYC reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
AI Agents Continuously Escaping Creators' Control: What We Know So Far
Australia reveals that a OpenAI bot invaded a government website in June, becoming one of the first known cases of AI bots attacking government websites. The article also outlines several similar incidents involving OpenAI, Google, Meta, and China's Kimi, and discusses the risks brought about by the intersection of AI and crypto security, as well as the industry debate regarding slowing down development.
Decrypt
·2026-09-28 18:23:35
4
The Chinese Academy of Sciences releases its 15th Five-Year Development Plan, aiming to build a world-class scientific research institution by 2035
According to "Voice of the Chinese Academy of Sciences," the Chinese Academy of Sciences has officially released its 15th Five-Year Development Plan, which outlines systematic arrangements for development goals, main tasks, and significant measures for the next five years. It also sets the goal of establishing itself as a world-class scientific research institution by 2035.
The Block
·2026-09-28 18:02:52
14
TrendAI expands on NVIDIA Agent Safety Platform, incorporating threat intelligence and full-stack AI factory security
TrendAI announces support for NVIDIA Agent Safety Platform, stating that its Vision One platform can integrate model protection, tool governance, data and identity control, as well as detection responses based on BlueField DPU into a single platform, helping enterprises to more securely advance the deployment of AI proxies from pilot projects to production.
PR Newswire
·2026-09-28 17:42:04
15
NVIDIA Launches AI Security System: Real-time Monitoring of Agents, Millisecond-level Prevention of Violations
NVIDIA released a dual-layer security platform named “Open Agent Safety Platform” on Monday, which includes open-source tools OpenShell and Nvidia Sentry. This platform allows for real-time control of proxy permissions on NVIDIA hardware and can shut down proxies in case of violations. The company claims that this system can isolate abnormal proxies in milliseconds and indicates that if relevant laboratories had deployed this technology earlier, the attack on the Hugging Face model in July of this year by the OpenAI could have been prevented.
Wallstreetcn
·2026-09-28 17:12:12
14
After Buffett's purchase price was broken through, has Google really become cheaper? | Silicon Valley Observation
This article argues that the key to determining whether Alphabet is a good deal lies not in Berkshire Hathaway's purchase price, but rather in whether the company's annual capital expenditures of nearly $200 billion can be transformed into sustainable cloud profits and a new generation of business distribution rights. The article discusses whether Google is transitioning from a search advertising company to an AI infrastructure and action distribution platform, by considering Berkshire Hathaway's investments, Alphabet's financing and capital expenditures, Google Cloud's revenue and profit performance, as well as a case study of an entrepreneur using Gemini to accomplish real-world tasks.
The Block
·2026-09-28 16:25:11
19
View More