TLDR⌗

Since I wrote up running Claude Code in a Docker container back in May, Anthropic have shipped their own sandbox for the CLI’s Bash tool. On Linux it wraps every command in bubblewrap, and an organisation can make it mandatory through managed settings. Turn that on inside a Docker container and it fails – and the failures lead you, one plausible step at a time, towards --cap-add SYS_ADMIN, which doesn’t fix it either. The short version of what does:

  • A seccomp profile: Docker’s default, plus five allow rules so an unprivileged process can create user, PID and mount namespaces and call mount.
  • An AppArmor profile: Docker’s default docker-default, with its deny mount, line replaced by mount,. Not apparmor=unconfined – on an Ubuntu host that is the thing that causes the final failure.
  • One Claude Code setting, sandbox.enableWeakerNestedSandbox: true, so the inner sandbox binds the container’s /proc instead of mounting a fresh one.

No --privileged, no --cap-add, no systempaths=unconfined. The container ends up more confined than the one I was running while debugging this.

What changed since May⌗

The May setup was about getting the agent off my workstation: a container with my source tree mounted, running as my own user, and a shell function that refuses to start it anywhere outside ~/development. The agent’s commands ran unconfined inside that container, and the container was the boundary. That is still my setup for my personal account, and it still works.

What Anthropic added is a second, inner boundary. With sandbox.enabled on, every Bash command Claude runs is wrapped in bwrap with a curated set of read-only and writable binds, no network except through an egress proxy the CLI runs, a fresh PID namespace, and a small helper called apply-seccomp that installs a BPF filter to stop commands sneaking out through Unix sockets. It’s good engineering, and for the developers on my company’s account it is a genuine step-up in the baseline security of AI tooling, and also something we can configure at enterprise level - so I made it mandatory in our managed settings. Unfortunately, it broke my own container - and I did not want to carve out an exception for myself, because exceptions however well-intentioned are the path to future vulnerabilities.

Which is how I came to be sitting in front of a Claude Code session whose Bash tool didn’t work.

Peeling the onion⌗

The failures come one layer at a time, and each one has an obvious fix that is also the wrong one. Here is the sequence, because you will hit them in this order.

Layer 1: seccomp. Docker’s default seccomp profile only allows clone and unshare with the CLONE_NEW* flags, and the whole mount family, when the container holds CAP_SYS_ADMIN. bubblewrap is unprivileged by design: it creates a user namespace and does its mounting inside that. Under the default profile the Bash tool simply never starts. The quick fix is --security-opt seccomp=unconfined, and it gets you to the next layer.

Layer 2: AppArmor. The docker-default profile contains a flat deny mount,. bubblewrap’s first act after creating its namespaces is mount(NULL, "/", NULL, MS_REC|MS_SLAVE), so you get:

bwrap: Failed to make / slave: Permission denied

The quick fix is --security-opt apparmor=unconfined. Remember this one; we’re coming back to it.

Layer 3: the masked /proc. Docker hides /proc/kcore, /proc/sysrq-trigger and a handful of other files by over-mounting them. bubblewrap wants a fresh procfs for its new PID namespace, and the kernel refuses to let a less-privileged user namespace mount one if doing so would un-hide paths the outer one has covered up. So:

bwrap: Can't mount proc on /proc: Operation not permitted

The quick fix is --security-opt systempaths=unconfined, which un-hides those paths for the whole container. The right fix is Claude Code’s own setting, sandbox.enableWeakerNestedSandbox, which tells it to bind the existing /proc rather than mount a new one. Anthropic document this in their sandboxing troubleshooting section, and it is exactly the case it exists for.

Layer 4: the one that looks fatal. With all three of those unconfined, bubblewrap now sets up fine, and the failure moves inside it:

apply-seccomp: write /proc/self/setgroups (nested userns is capability-restricted; caller must provide CAP_SYS_ADMIN): Permission denied

apply-seccomp runs inside the bubblewrap sandbox. To install its filter and hide the outer PIDs it creates a second, nested user namespace, maps its uid into it, and unshares PID and mount namespaces again. Writing deny to setgroups is the first step of configuring that namespace, and it wants CAP_SYS_ADMIN in the new namespace – which, on any normal Linux box, the process that created the namespace automatically has.

The error message says “caller must provide CAP_SYS_ADMIN”. So I added --cap-add SYS_ADMIN. Same error. At that point the reasoning in my terminal – and I was doing this with Claude’s help, more on which below – concluded that bubblewrap drops all capabilities before executing its payload, so nothing you grant the container can reach a helper running inside the sandbox, and the whole approach was a dead end. Plausible. Wrong.

What is actually going on⌗

The clue is the phrase nested userns is capability-restricted. That is a very specific thing to write in an error message, and it turns out Anthropic wrote it for a very specific reason.

Ubuntu, since 24.04, ships with kernel.apparmor_restrict_unprivileged_userns=1. The behaviour it enables is not “unprivileged user namespaces are disabled”. It is stranger than that: when an unconfined process without CAP_SYS_ADMIN creates a user namespace, AppArmor lets the unshare succeed, and then silently transitions the process into a special profile called unprivileged_userns, which denies every capability. The namespace exists, but you can’t do anything privileged in it. The setgroups write is the first privileged thing apply-seccomp tries, so that is where it dies.

You can watch it happen. Inside the container, read /proc/self/attr/current before and after an unshare(CLONE_NEWUSER):

[before]        apparmor label: unconfined
unshare(NEWUSER): ok
[after userns]  apparmor label: unprivileged_userns (enforce)
write /proc/self/setgroups: Permission denied

Inside bubblewrap it’s the same story with a different name – the host’s stock bwrap profile has a child called unpriv_bwrap that does the same job, so the label reads bwrap//&unpriv_bwrap. Those profiles come from the host’s /etc/apparmor.d; Docker didn’t put them there and docker run can’t take them away.

Now look back at layer 2. The rule applies to unconfined processes. docker-default is a confined profile – and, because it pins itself to AppArmor ABI 3.0, which predates user-namespace mediation, namespace creation isn’t even a question it gets asked. A process confined by docker-default creates a nested namespace and keeps its capabilities. The only thing docker-default was actually blocking was mount. By replacing it with apparmor=unconfined I had fixed one line and, without noticing, opted the whole container into Ubuntu’s userns restriction.

And --cap-add SYS_ADMIN genuinely can’t help, for the reason the original diagnosis gave: bubblewrap drops everything before apply-seccomp runs. The diagnosis was right about the capability and wrong about the conclusion. The fix isn’t more privilege. It’s being confined.

The fix⌗

Three pieces. Two are files on the host, one is a Claude Code setting.

The AppArmor profile⌗

Take Docker’s docker-default template – it lives in moby/profiles and is short – and change exactly one thing: deny mount, becomes mount,, plus pivot_root, which bubblewrap also needs. Everything else stays, including the denies on /proc/kcore and /proc/sysrq-trigger, which still apply inside the sandbox’s mount namespace.

abi <abi/3.0>,

#include <tunables/global>

profile "docker-claude" flags=(attach_disconnected,mediate_deleted) {
  network,
  deny network alg,
  deny network vsock,
  capability,
  file,
  umount,
  # --- the change: docker-default has `deny mount,` here ---
  mount,
  pivot_root,
  signal (receive) peer=unconfined,
  signal (receive) peer=runc,
  signal (receive) peer=crun,
  signal (send,receive) peer="docker-claude",

  deny @{PROC}/* w,
  deny @{PROC}/{[^1-9/],[^1-9/][^0-9/],[^1-9s/][^0-9y/][^0-9s/],[^1-9/][^0-9/][^0-9/][^0-9/]*}/** w,
  deny @{PROC}/sys/[^k]** w,
  deny @{PROC}/sys/kernel/{?,??,[^s][^h][^m]**} w,
  deny @{PROC}/sysrq-trigger rwklx,
  deny @{PROC}/kcore rwklx,

  deny /sys/[^f]*/** wklx,
  deny /sys/f[^s]*/** wklx,
  deny /sys/fs/[^c]*/** wklx,
  deny /sys/fs/c[^g]*/** wklx,
  deny /sys/fs/cg[^r]*/** wklx,
  deny /sys/firmware/** rwklx,
  deny /sys/devices/virtual/powercap/** rwklx,
  deny /sys/kernel/security/** rwklx,

  ptrace (trace,tracedby,read,readby) peer="docker-claude",
}

Load it once:

sudo cp docker-claude.apparmor /etc/apparmor.d/docker-claude
sudo apparmor_parser -r -W /etc/apparmor.d/docker-claude

Is mount, a real relaxation? Yes – it’s the reason bubblewrap works at all. But AppArmor sits on top of the kernel’s own rules, not instead of them, and the kernel still only honours mount(2) from a process holding CAP_SYS_ADMIN in the user namespace that owns the mount namespace. For an unprivileged container that means: only inside a user namespace you created yourself, on mounts you can already see. It’s the same permission every unprivileged user on a bare Ubuntu box has.

The seccomp profile⌗

Take Docker’s default default.json – also in moby/profiles – and append five rules to the syscalls array. They are all conditioned on the container not having CAP_SYS_ADMIN; if it does, the stock rules already allow everything.

{
  "names": ["clone"],
  "action": "SCMP_ACT_ALLOW",
  "args": [{ "index": 0, "value": 268435456, "valueTwo": 268435456, "op": "SCMP_CMP_MASKED_EQ" }],
  "excludes": { "caps": ["CAP_SYS_ADMIN"], "arches": ["s390", "s390x"] },
  "comment": "clone() with CLONE_NEWUSER: bwrap creates its user+pid+net namespaces in one clone"
},
{
  "names": ["unshare"],
  "action": "SCMP_ACT_ALLOW",
  "args": [{ "index": 0, "value": 268435456, "valueTwo": 268435456, "op": "SCMP_CMP_MASKED_EQ" }],
  "excludes": { "caps": ["CAP_SYS_ADMIN"] },
  "comment": "unshare(CLONE_NEWUSER): apply-seccomp's nested user namespace"
},
{
  "names": ["unshare"],
  "action": "SCMP_ACT_ALLOW",
  "args": [{ "index": 0, "value": 131072, "valueTwo": 131072, "op": "SCMP_CMP_MASKED_EQ" },
           { "index": 0, "value": 268435456, "valueTwo": 0, "op": "SCMP_CMP_MASKED_EQ" }],
  "excludes": { "caps": ["CAP_SYS_ADMIN"] },
  "comment": "unshare(CLONE_NEWNS[|CLONE_NEWPID]) from inside an owned user namespace"
},
{
  "names": ["unshare"],
  "action": "SCMP_ACT_ALLOW",
  "args": [{ "index": 0, "value": 536870912, "valueTwo": 536870912, "op": "SCMP_CMP_MASKED_EQ" },
           { "index": 0, "value": 268435456, "valueTwo": 0, "op": "SCMP_CMP_MASKED_EQ" }],
  "excludes": { "caps": ["CAP_SYS_ADMIN"] },
  "comment": "unshare(CLONE_NEWPID) from inside an owned user namespace (bwrap --unshare-pid)"
},
{
  "names": ["mount", "umount2", "pivot_root"],
  "action": "SCMP_ACT_ALLOW",
  "excludes": { "caps": ["CAP_SYS_ADMIN"] },
  "comment": "mount family bwrap uses to build its root; only effective inside an owned user namespace"
}

The magic numbers are CLONE_NEWUSER (0x10000000), CLONE_NEWNS (0x20000) and CLONE_NEWPID (0x20000000). The MASKED_EQ comparisons say “this bit set” or “this bit clear” without caring about the others. These are the syscalls bubblewrap 0.12 and apply-seccomp actually make; I checked the sources rather than guessing. If you’d rather not maintain a profile, seccomp=unconfined also works and is a much smaller sin than the AppArmor equivalent – but the file is cheap and it keeps the several hundred per-syscall decisions Docker’s default makes for you.

The Claude Code setting⌗

In the user settings for the profile that runs in the container – ~/.claude/settings.json from the container’s point of view:

{ "sandbox": { "enableWeakerNestedSandbox": true } }

What it does: bubblewrap gets --bind /proc /proc instead of --proc /proc. What you lose: commands inside the sandbox can see the container’s process list rather than only their own. Given that everything in the container is the same Claude session anyway, I’m content with that. Note that Claude Code honours this key from user settings unless your managed settings pin it, and there is an open issue where setting CLAUDE_CODE_SUBPROCESS_ENV_SCRUB silently forces it back off; if you hit Can't mount proc again with the setting in place, that is the first thing to check.

I did also read apply-seccomp’s source (Anthropic publish it in sandbox-runtime) to make sure its own /proc remount, inside the nested PID namespace, wouldn’t trip over the same masked-paths rule. It tolerates that specific EPERM on purpose, with a comment naming enableWeakerNestedSandbox as the environment it’s for. Somebody at Anthropic has been down this road.

The wrapper⌗

The safeclaude function from May gains two lines:

docker run -it --rm \
    -e "HOST_UID=$UID" \
    -e "HOST_GID=$GID" \
    --security-opt seccomp=/path/to/docker-claude.seccomp.json \
    --security-opt apparmor=docker-claude \
    ...

That’s it. The image didn’t change at all – bubblewrap and socat were already in it, because Claude Code’s dependency check nags you until they are.

Test it before you trust it⌗

The thing that turned this from guesswork into a fix was a probe that walks the same four layers in order, without Claude Code in the loop, and prints the AppArmor label as it goes. The bubblewrap steps are one-liners you can paste into a container:

# 1. user namespace at all (seccomp)
unshare -U true

# 2. bwrap with a fresh /proc  -- EXPECTED TO FAIL in Docker; enableWeakerNestedSandbox avoids it
bwrap --ro-bind / / --dev /dev --unshare-net --unshare-pid --unshare-user --cap-drop ALL --proc /proc -- true

# 3. bwrap with the container's /proc bound -- this is the form Claude Code uses with the setting on
bwrap --ro-bind / / --dev /dev --unshare-net --unshare-pid --unshare-user --cap-drop ALL --bind /proc /proc -- true

Step 4 is a twenty-line C program that does what apply-seccomp does – unshare(CLONE_NEWUSER), write setgroups/uid_map/gid_map, unshare(CLONE_NEWPID|CLONE_NEWNS), remount /proc – and reports each call, run inside the step-3 sandbox. With the profiles loaded, the run that mattered looked like this:

kernel: 7.0.0-31-generic   uid: 1000   apparmor label: docker-claude (enforce)
apparmor_restrict_unprivileged_userns: 1

== 1. create user namespace (unshare -U)
   PASS
== 2. bwrap, fresh /proc  (what Claude Code does by default)
bwrap: Can't mount proc on /proc: Operation not permitted
   FAIL (expected)
== 3. bwrap, bound /proc  (enableWeakerNestedSandbox form)
   PASS
== 4. apply-seccomp's nested namespace setup, inside bwrap
    [before] apparmor label: docker-claude (enforce)
    unshare(NEWUSER): ok
    [after userns] apparmor label: docker-claude (enforce)
    write /proc/self/setgroups <- "deny": ok
    write /proc/self/uid_map <- "1000 1000 1": ok
    write /proc/self/gid_map <- "1000 1000 1": ok
    unshare(NEWPID|NEWNS) after userns: ok
   PASS

The line to look at is the label after the namespace is created. Under apparmor=unconfined it flips to unprivileged_userns; under docker-claude it stays put, and everything after it just works. Then I started a real session on the company account, with its sandbox-mandating managed settings, and ran echo ok through the Bash tool. It said ok.

What you are and aren’t giving up⌗

Honesty section, because the whole point of the May post was a security boundary and I’ve just poked three holes in it.

  • mount, in AppArmor. Real, but kernel-bounded as above. The container still can’t remount anything it couldn’t already; the kcore and sysrq-trigger denies remain.
  • Namespace syscalls in seccomp. Docker blocked these mainly because user namespaces have had a run of kernel CVEs over the years, and Docker’s threat model is “you might be running untrusted images”. Mine is “I’m running an image I built, and I’d like the agent inside it to be more confined, not less”. The trade is an unprivileged-userns attack surface in exchange for a second sandbox around every command the agent runs. I think that’s the right trade; you may not.
  • enableWeakerNestedSandbox. Sandboxed commands can see the container’s process table. Not the host’s.
  • Nothing else. No capabilities added, /proc still masked, and the container is confined by a profile that is ninety-five percent docker-default. Compared with where I was mid-debugging – seccomp, AppArmor and system paths all unconfined and SYS_ADMIN, with the Bash tool still broken – this is a different league.

If you’ve gone further than I did and run the container with --cap-drop ALL and no-new-privileges, the seccomp rules above are written for exactly that case (they only fire when CAP_SYS_ADMIN is absent) and bubblewrap on Debian isn’t setuid, so I’d expect it to work unchanged. I haven’t tested it, so treat that as a prediction rather than a promise.

A note on who did the debugging⌗

I should be clear about how this was found, because it’s a bit funny. I was running Claude Code, on my personal account, inside the container, and asking it to work out why Claude Code, on my company account, couldn’t run commands inside the same container. It wrote the C probe, compiled it in the sandbox, read the AppArmor label off /proc/self/attr/current, and pulled the apply-seccomp source to confirm the EPERM tolerance – and it was the one that noticed the “nested userns is capability-restricted” wording was too specific to be generic, went looking, and found Ubuntu’s transition rule. The previous session, the one that concluded “dead end”, was also Claude. Same model family, same evidence available. The difference was that the second session tested the hypothesis instead of reasoning about it.

There’s a lesson in that for the humans too.

Drafted by Claude, edited by me. I write these things mostly for my own memory – and because the next person to see nested userns is capability-restricted inside a container deserves a better search result than “add SYS_ADMIN”.