← Thinking Thinking

DeepSeek DSec: Training Infrastructure Serving About 3 Million Sandboxes a Day

DeepSeek’s DSec serves about 3 million sandboxes a day, letting agents run training tasks in real environments. This article unpacks how it uses independent environment layers and on-demand reads to cut preparation overhead, memory sharing and reclamation with CPU scheduling to carry large numbers of concurrent sessions, and decoupling from preemptible GPU training to preserve execution progress. Public experiments verify the gains of some mechanisms, but they are not yet enough to fix the ratio between environment and GPUs. Judging such platforms also depends on the supply of valid training trajectories, recovery overhead, and total cost.

2026-09-29Thinking49 min read
DeepSeek DSec: Training Infrastructure Serving About 3 Million Sandboxes a Day

Training an agent that writes code cannot stop at letting it read code and answers. The model has to enter a software repository, install dependencies, edit files, run tests, and decide the next step from the results. A single task often goes through many rounds of such interaction. While the model generates the next instruction, the computing environment has to stay in place: the packages just installed, the files half-edited, and the services still running can all affect what happens next.

When training runs thousands of such attempts at once, supplying these environments becomes an infrastructure job of its own. The platform must create clean environments quickly and keep the state of those mid-execution; it must hold large numbers of sandboxes that are waiting for the model, and it must keep one runaway operation from destroying other tasks.

The DSec paper that DeepSeek published on September 19 shows this work at production scale. The paper reports that one production unit has about 160 CPU nodes, 30,000 CPU cores, and roughly 250 TB of memory in total; a typical day serves about 3 million sandbox instances, peak concurrency is about 380,000, and the creation rate tops 5,000 per second. It carried every sandbox workload for the reinforcement-learning training and evaluation of DeepSeek V3.2 through V4.1. [1]

These numbers show that environments now need capacity planning, resource scheduling, and fault handling of their own, just as model compute does. They do not yet prove that the main bottleneck of training has moved from the GPU to the environment. What makes DSec worth a closer look is something else: it manages separately the state that model compute and environment execution each need to keep, and then rearranges storage, memory, and CPU around how environments are actually used.

First, the Whole Picture: What One Training Run Has to Coordinate

In agentic reinforcement learning, a training sample usually comes from many rounds of interaction between the model and an environment. The model decides the next action, the sandbox executes the command, and the result returns to the model until the task ends. This whole stretch of interaction is called a rollout. The training framework then scores the result and updates the model with the trajectory and the reward. Evaluation generates the same kind of trajectories, except that no parameters are updated from them.

Three kinds of work with different natures are involved: model inference generates actions, environments execute actions and hold state, and training updates parameters from the collected data. A real system can run them overlapped. DSec manages the environments and the execution attached to them; moving rollout execution into DSec should not be read as model inference no longer needing GPUs.

Figure 1: Arranged from the paper, §6.2. From V4.1 on, the worker container and the agent sandbox keep rollout state outside the preemptible GPU pool; model inference, reward computation, and parameter updates are distinct responsibilities. The figure omits reward computation and the concrete training pipeline.
Figure 1: Arranged from the paper, §6.2. From V4.1 on, the worker container and the agent sandbox keep rollout state outside the preemptible GPU pool; model inference, reward computation, and parameter updates are distinct responsibilities. The figure omits reward computation and the concrete training pipeline.

DSec exposes a Python library called libdsec. The caller specifies the environment, resources, and network permissions, creates a sandbox, runs commands and reads results, and finally releases the environment. A similar interface does not mean the ways of running are interchangeable. A short code execution need not carry the cost of a full virtual machine, and a security-sensitive task at the kernel boundary cannot chase startup speed alone.

Runtime Suited tasks Main trade-off
FnCall Short, stateless code execution, compilation, or operator tests Reuses pre-created containers, cutting per-call preparation overhead
Containers Repository edits, tests, and general tool calls Fast to start and dense to pack, but tasks on one container host share a kernel
Firecracker microVMs Linux tasks that need stronger isolation Each lightweight VM has its own kernel; higher memory and startup overhead
Full virtual machines Tasks that depend on full system capabilities, such as Android or graphical interfaces Runs the most complete system environments, at the highest resource cost

One should also be precise about where containers sit relative to the bare metal. DSec’s FnCall and container tiers run inside QEMU/libvirt virtual machines, and that VM adds an isolation boundary between untrusted code and the bare metal. Containers share the kernel of this VM, not the bare-metal kernel directly. This detail determines which tasks a failure can reach, and it should not be flattened away by a simple ladder diagram from light to heavy backends. [1, §2.2, §3.3]

Placing a creation request takes two steps. On the cluster side, the system first filters out unhealthy or mismatched nodes, then samples a few candidates at random and picks the one with the lowest load. This reduces the effect of a large burst of concurrent requests all converging on the same “least loaded” machine. The node-local edge component then decides admission from its immediate resource state; under pressure it can reject the request and have the system choose another node.

Both steps are necessary. The cluster view refreshes periodically; it suits fast filtering but cannot reflect every instantaneous change in resources. Leaving final admission to the node is what prevents a stale cluster estimate from overriding current memory or CPU pressure. The platform also counts placements too recent to appear in the monitoring snapshot into the node’s local estimate, which reduces the error from back-to-back allocations. [1, §3, §7]

Preparing Environments: Split Updates First, Then Move Less Data

Once sandboxes scale up, the first problem to surface is usually not slow command execution, but how much content has to be moved and unpacked before a command can even start.

In one production week, the container backend served 11,266 base images and 102,171 workspaces, and the platform offered 103 toolkits. Base images carry the operating system and language runtimes; workspaces carry task code and dependencies; toolkits include frequently updated components such as the agent execution framework. The three update at different frequencies, yet they must combine into the single filesystem a task sees.

If every task packed all of this into one complete image, a toolkit upgrade would force a rebuild of every image embedding it, even with the operating system and the repository untouched. Another option is to unpack workspaces and toolkits after startup, but then a batch of tasks repeatedly does the same unpacking and disk writes. When requests arrive in a burst, this preparation competes for resources with tasks already running.

DSec stores them as independently versioned read-only layers and composes them when a sandbox is created. The bottom is the base image; workspaces and toolkits stack on top; a writable directory sits at the very top. Linux’s overlayfs presents these layers as one directory tree: file conflicts resolve by layer priority, and runtime modifications land in the writable layer without changing the read-only content below that other sandboxes reuse.

A toolkit upgrade then only has to produce a new toolkit layer, and tasks keep using their existing base images and workspaces. The gain comes from not rebuilding components that did not change. Multiplying the number of base images by the number of workspaces gives more than a billion theoretical combinations, but that does not mean the platform ever maintained that many images; what actually needs shrinking is the repeated building and duplicated data inside the combinations in use.

Figure 2: The container path, arranged from the paper, §5.1 and §5.3. Independent layers decide what can be reused; on-demand reads decide when content is moved. The microVM writable disk follows a separate OverlayBD block-device path and should not be read as the same implementation.
Figure 2: The container path, arranged from the paper, §5.1 and §5.3. Independent layers decide what can be reused; on-demand reads decide when content is moved. The microVM writable disk follows a separate OverlayBD block-device path and should not be read as the same implementation.

Even with layers, there is still no reason to download every layer in full before startup. Across the container images of different languages sampled in the paper, the data actually accessed during a run is only 4.2% to 13.3% of the image. Many task images, meanwhile, are reused only a few times, so local caches can hardly pre-cover such scattered environments. Full pre-warming only moves the transfer cost earlier; it does not remove it.

DSec stores published layers in the read-only filesystem EROFS. Unlike an archive that must be decompressed in order, EROFS can read and decompress just the blocks at the accessed positions. On the container path, metadata such as directory structures is placed on local nodes in advance, file content lives in DeepSeek’s distributed filesystem 3FS and is fetched only as far as it is read, and runtime writes stay in the local writable layer.

This division of labor follows the shape of the storage system. 3FS is good at large reads and writes, while small, frequent random I/O performs poorly. So path lookups stay local, reads happen on demand and move in large blocks, and the logs and file modifications that cannot be predicted in advance stay local too. The value of on-demand loading is not just an earlier start; it is less data actually moved over the whole task.

The microVM tier needs a separate storage path. For example, when Docker itself runs inside a virtual machine, its data directory cannot simply sit on top of overlayfs, and the Firecracker used in the paper does not support virtio-fs. DSec therefore keeps the read-only base and toolkit layers on EROFS and serves the writable ext4 disk through OverlayBD and the user-space block framework ublk. This path fetches data from the remote side in 256 KiB blocks and keeps a local second-level cache. The cost is that ext4 metadata still lives inside the disk image, so some metadata access can also trigger remote reads. [1, §5]

This shows that a unified platform does not erase backend differences. The container tier can optimize metadata placement along the filesystem structure; the virtual-machine tier has to accommodate block devices and application compatibility. Both paths aim to move less and write less repeatedly, but a common interface alone does not mean identical startup and access performance.

Running at High Density: Idle CPU Does Not Mean Idle Memory

Once an environment is in place, the agent often spends its time waiting for the model to generate the next instruction. The paper reports that for about 90% of containers and microVMs, average CPU usage stays within 5% of the requested amount. If the platform reserved CPU by the requested amount, a large share of allocated compute would sit idle for long stretches. DSec therefore allows the CPU total allocated to sandboxes to exceed physical capacity, letting intermittent tasks share the machine.

Low average usage does not mean low at every moment. Installing dependencies, compiling, and testing can form compute peaks at the same time, and some tasks have strict response deadlines. DSec splits sandboxes into latency-sensitive and best-effort classes and has the latter yield CPU first through the Linux SCHED_IDLE scheduling policy.

Lowering scheduling priority alone is not enough. A physical CPU core usually hosts several logical threads that share execution resources. A background task can compete with a foreground task even without preempting its thread, by occupying the sibling thread of the same physical core. DSec therefore also enables core scheduling, which keeps unrelated background tasks off the sibling threads of latency-sensitive tasks.

In the paper’s chess-evaluation experiment, with background load at 50% of a node’s CPU capacity, single-step latency without protection rose 45.2% compared with no co-located load; with the two-layer control, the rise fell to 17.3%. The two layers did not eliminate all interference: heavy multi-core load pulls down turbo frequencies, and tasks still contend for memory bandwidth and the last-level cache. What this experiment supports is latency protection under a specific workload, not a proportional throughput gain for every kind of task. [1, §8.5]

Memory is the opposite problem: CPU time can be handed over while the model works, but task state cannot simply vanish. The microVM tier adds waste of its own. The same image content can sit simultaneously in the host page cache and in the page caches of several VMs; pages that have gone idle inside a VM may never be reclaimed by the host unless the VM reports them.

DSec treats the two kinds of waste separately. For the read-only base images and toolkit layers, it uses virtio-pmem with DAX, letting virtual machines access file data in host memory directly, so no second page cache is kept inside each VM. For the larger writable disks, Linux’s DAMON memory-access monitor finds file pages untouched for a long time and pushes the kernel to reclaim them; the virtual machine then reports free pages, handing releasable memory back to the host.

Figure 3: The paper’s §8.4 comparison covers Firecracker microVM configurations. 40.2% is the reduction in peak host memory, and 21.2% is the reduction in time-integrated memory occupancy; the two are different statistics, cannot be added, and do not explain container-backend capacity.
Figure 3: The paper’s §8.4 comparison covers Firecracker microVM configurations. 40.2% is the reduction in peak host memory, and 21.2% is the reduction in time-integrated memory occupancy; the two are different statistics, cannot be added, and do not explain container-backend capacity.

The two mechanisms improve different metrics. In the controlled experiment on a real agentic reinforcement-learning workload, virtio-pmem with DAX cut peak host memory by 40.2%; DAMON with free-page reporting, used alone, left the peak largely unchanged but reduced time-integrated memory consumption by 21.2%. The former mainly removes duplicate caches; the latter mainly shortens how long useless pages stay occupied.

The savings have costs. virtio-pmem requires the VM to keep page metadata for the mapped range: the paper’s example is that a 128 GB device needs about 2 GB of VM memory for this metadata. Cold accesses can also trigger synchronous page faults. In the experiment, enabling the mechanism raised the transient CPU peak from 26.5% to 41.4%, and the paper lists the cold-access path difference as one possible cause. A deployment already tight on CPU may therefore enable only free-page reclamation and keep the original block-device read path. [1, §5.2, §8.4]

What decides deployment density is exactly these exchange relations among resources. Seeing only the low CPU usage while a sandbox waits for the model does not mean more tasks can keep being added; seeing only the memory savings does not mean the extra CPU and storage pressure at first access and on recovery can be ignored.

Preemption and Recovery: Keep Task Progress, Release Idle Resources

Long tasks meet another kind of interruption: the GPU training job is preempted by the scheduler. The code and processes in the sandbox can stay where they are, but if the loop that drives the agent’s interaction exits with the training process, the system no longer knows which step to resume from.

In the early architecture, the agent loop ran inside the preemptible GPU training pod, together with model serving and the reinforcement-learning framework. After a training job was interrupted, the sandboxes still existed, but recovery had to rely on the command log to reconcile the state restored by the training framework with the sandbox’s current state. For commands already completed, the system reused the recorded results instead of re-executing them, to avoid repeating side effects. The DeepSeek V4 technical report also describes this recovery method. [2, §5.2.5]

From V4.1 on, DSec splits rollout execution into an agent sandbox and a worker container. The former hosts the agent execution framework and tools; the latter manages sandboxes and controls the rollout. Both sit outside the preemptible GPU pool and together hold the complete execution state. After the GPU job resumes, it reconnects and continues, without rebuilding this stretch of execution through the command log. [1, §6.2]

The point of this change is that execution progress now follows the lifecycle of the work itself. A training job interrupted to raise cluster utilization should not force a half-finished code edit to re-explain its own history. But this does not mean the model can keep generating actions without inference compute, and it does not remove all recovery overhead: the kept environments still occupy memory.

The training framework therefore actively pauses the sandboxes tied to a preempted job. On the container path it freezes the process, permits swapping through memory.swap.max, and then calls memory.reclaim to reclaim memory proactively. On resume, the platform issues asynchronous prefetch with MADV_WILLNEED and then unfreezes the process. For microVMs, memory and execution state are saved as a snapshot and the Firecracker process is terminated to release runtime memory; on resume a new process starts and loads the snapshot. [1, §6.3]

Two conditions must hold at once: task state must be kept, and physical memory cannot stay occupied by idle tasks forever. Pause and resume therefore become part of scheduling. The platform still pays I/O costs for swap-out, snapshots, and prefetch; whether it pays off depends on how long the pause is, how large the state is, and how long a wait recovery can tolerate. The paper gives no independent measurement of these operations’ net effect on total training time.

What the Public Experiments Show, and What They Cannot Compute

What deserves credit in DSec’s experiments is that each mechanism was verified against its own control. Two results speak most directly to why environment preparation is worth optimizing.

The first was run on a separate 10-node test cluster, launching 8,192 container tasks at once. On-demand loading finished the batch in about 35 minutes, close to the control with all images pre-placed locally; a full remote pull took over 60 minutes, which the paper reports as about 1.71 times the on-demand completion time. On-demand loading also cut cumulative disk writes from over 1,600 GB to about 700 GB per node, a reduction of about 57%. Those 35 minutes include task execution, so task count divided by elapsed time is not a sandbox creation rate. [1, §8.2]

The second compared ways of preparing workspaces and toolkits, replacing model generation with pre-recorded deterministic tool calls to confine the difference to environment preparation itself. With per-sandbox decompression of archives, the whole batch took 79 minutes; with direct mounting of EROFS layers, 45 minutes. Reducing repeated unpacking and writing at startup did improve end-to-end completion time in this experiment. It is still not a comparison of GPU utilization or total cost on a complete reinforcement-learning pipeline. [1, §8.3]

Production-scale numbers need a different set of boundaries. A node steadily running 3,200 containers or 800 microVMs is an operating point the paper observed, not a hard ceiling of the backend, and not proof that isolation costs exactly four times the density under comparable task conditions. The paper’s single-day, single-node peaks are load records of that sample; they cannot be copied to every node and used to refute the whole unit’s stated concurrency.

Likewise, the median sandbox lifetime cannot stand in for the mean. The paper reports a container lifetime median of 17.4 minutes, with p99 (the value at the 99th position when lifetimes are sorted from short to long) above three hours for both containers and microVMs. That shows long tasks exist, but it gives no mean residence time, and it does not establish that a sandbox’s lifetime equals one training trajectory’s duration. Estimating average concurrency from these numbers directly would mix different statistical objects.

So how many environment nodes does a 10,000-GPU training cluster actually need? Public materials are not yet enough to give a credible number. At least three things must be measured first: the rate of trajectories actually entering the system, the total time each trajectory occupies across its sandboxes, and how much of that load a node can carry safely. Training, evaluation, and environment building must also be told apart, rather than assumed to run at the same rhythm.

Within a stable observation window, Little’s Law estimates average occupancy: the average number in the system equals the actual admission rate times the average residence time. Both the rate and the residence time should be averaged over the same stable window, and neither side’s theoretical peak capacity should be substituted in. If one trajectory uses several sandboxes, each sandbox’s occupied time counts; if there is a reusable resident pool, it should be accounted separately.

A hypothetical, to illustrate the method only: suppose a task class admits 10 trajectories per second, each trajectory uses one sandbox, and the measured average residence is 20 minutes; then average concurrency is about 12,000. If load testing confirms the same workload can run stably at 1,000 per node, 12 nodes are only the equivalent requirement for average load; the actual configuration still needs headroom for fluctuation, bursts of backlog, and failure redundancy. None of these numbers are DSec’s measurements, and they cannot be used to back out its ratio with the GPUs.

To find the bottleneck, watch trajectory completion rate, training consumption rate, model-generation waiting, environment-creation waiting, and tool execution time together. Low GPU utilization only signals that waiting exists; on its own it cannot distinguish inference supply, environment supply, data preparation, or training scheduling. For burst requests, watch how queues build up and drain over time; an adequate average capacity does not mean a batch of evaluations can start without queuing.

For burst demand, DSec borrows cloud capacity, but the tasks it can divert are still constrained by their dependencies. It pre-syncs a set of high-coverage image content, and only tasks whose dependencies all fall within that set can move to the cloud. The paper reports that when local utilization exceeds 80%, the platform diverts part of the qualifying new requests, with 200 cloud virtual machines absorbing about 30% of the peak overflow. This is selective capacity top-up, not a promise that any environment can move to the cloud at any time. [1, §3.4]

Environment Correctness Is Itself Training-Data Quality

Sandboxes also carry a piece of work that resource numbers easily obscure: making sure task results really come from the model completing the task as prescribed.

The behaviors the paper records include searching platform logs for answers, attempting to forge internal requests, and fetching ready-made implementations from external mirrors or packages. In such cases, even a passing final test cannot be read directly as improved model capability. The platform must restrict which files the model can read and which services it can contact, and check that the scoring process follows the task’s rules. [1, §6.4]

This is a different problem from “whether the agent can escape the sandbox.” An agent can stay entirely inside the isolation boundary and still see reference answers it was never meant to see; an agent with no intent to cheat can trigger a kernel defect with one command traversing unusual files and wreck the execution environment. Virtual-machine isolation, file permissions, and network rules therefore address different risks and cannot substitute for one another.

DSec uses AppArmor to restrict file reads and writes and internal socket access, and these rules bind even when the agent runs as root inside the sandbox; on the network side, according to the services a task is allowed to reach, eBPF filters by IP, port, and protocol, and the rules can be updated as a task moves to a new stage. The paper is also explicit that these controls do not generally stop kernel defects; continuous observation and hardening are still required. [1, §6.5]

Agents building their own training environments need the same boundary. DSec provides pack_diff, an incremental disk snapshot that saves an environment built interactively and restores it into new sandboxes. Building, verification, and execution can then share one infrastructure. But the paper also states that environments must follow platform rules, that an internal platform performs quality checks before handing them to training and evaluation in a standard format, that builders and runtime agents use different accounts, and that reference answers possibly lingering in the writable layer are cleared before packaging. [1, §6.1]

This capability shows that environment building can be automated further; by itself it does not prove the model has formed a closed loop of continuous self-improvement. That conclusion would require showing that generated environments are effective enough, that training gains can be attributed to these environments, and that quality keeps improving across rounds. The DSec paper does not provide that set of experiments. Separating “able to generate environments” from “environments good enough to produce reliable training signal” shows what it has actually solved.

What Is Really Worth Borrowing: How Resources and State Are Divided

DSec did not invent a new runtime that owns every task. It keeps distinct isolation backends, uses independent read-only layers and on-demand reads to cut preparation work, supports high-density execution with memory reclamation and CPU scheduling, and separates long-lived rollout state from preemptible training jobs. Together these choices answer one question: in a task that runs over many rounds of interaction, what can be shared, what can be temporarily returned, and what state must be kept.

For anyone building a similar platform, the first things to measure are environment preparation time, actual file access volume, CPU activity, mean and long-tail residence times, and post-recovery waiting. Choosing layering, caching, CPU oversubscription, or pause strategies from those numbers is what tells you which stretch of waiting an optimization removes and where it moves the cost.

The paper provides concrete evidence on environment preparation, memory efficiency, and latency protection; what is still missing is the net benefit and total cost of these improvements on a complete training pipeline. The way to judge a training-environment platform, in the end, is whether it delivers valid trajectories more reliably, and at what overall cost, not how many sandboxes it runs at once.


Disclosure: This article is an analysis based on public materials; DSec experiments were not reproduced. Production-scale and experimental figures are values reported by the paper; the capacity example is an assumption set explicitly by this article. The figures are explanatory diagrams drawn from the mechanisms described in the paper.

[1] DeepSeek, DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale, arXiv:2609.22978v1, September 19, 2026.

[2] DeepSeek, DeepSeek V4 technical report, §5.2.5; used here only to verify the early preemption-recovery method.