Tech Digest

Guide

Should you run Docker in an LXC container or a VM on Proxmox?

Proxmox's own documentation and the forum consensus seem to disagree. They do not: they are answering different questions, and the one you are asking has a clear answer.

Last reviewed

Should you run Docker in an LXC container or a VM on Proxmox?

Put the Docker host that runs your stateful stacks (photos, passwords, databases, anything you would restore from backup) in a QEMU VM. Proxmox's own documentation recommends a VM where you need maximum isolation or live migration, running containers cannot be live-migrated, and Docker in LXC depends on the host's kernel, cgroup layout and AppArmor profile, so a host or guest update can break it. Use an unprivileged LXC with nesting and keyctl only on low-RAM hosts or for stacks you could rebuild in ten minutes.

Run the Docker host that holds your stateful stacks in a QEMU VM. Use an LXC container only when the host is short on RAM or the stacks inside it are ones you could rebuild from a Compose file in ten minutes. That is the rule, and it holds because of four specific properties of LXC on Proxmox VE, not because of a vague sense that VMs are "more secure".

The confusion exists because two true statements circulate. Proxmox's own Linux Container page says that "for use cases demanding maximum isolation and the ability to live-migrate, nesting containers inside a Proxmox QEMU VM remains a recommended practice." Forum threads and blog posts say Docker in LXC works fine. Both are right about what they describe. The forum posts describe the day of installation; the documentation describes the next three years.

The decision, in one table#

Your situationPut Docker inWhy
Stacks with data you would restore from backup: Immich, Vaultwarden, Nextcloud, anything with PostgresVMHost upgrades cannot break it, and it can live-migrate
A cluster where you want to move guests off a node for maintenance without downtimeVMRunning containers cannot be live-migrated
Anything reachable from the internetVMA container escape lands on the hypervisor's kernel, not a guest's
A host with 8 GB of RAM or less running several small stacksLXC, unprivilegedA VM reserves memory and runs a second kernel you may not have room for
Throwaway or test stacks, a dashboard, a Dozzle log viewerLXC, unprivilegedIf it breaks after an upgrade you redeploy it, not restore it
One self-contained image you want managed like any other guestProxmox OCI import (preview)Converted to a normal LXC container, no Docker at all

If a stack straddles two rows, the VM row wins. The cost of being wrong in the VM direction is some RAM. The cost of being wrong in the LXC direction is a morning spent on AppArmor while your password manager is down.

Reason one: containers cannot live-migrate#

This is stated plainly in the admin guide: "Running containers cannot live-migrated due to technical limitations." The alternative is a restart migration, which shuts the container down, moves it and starts it again on the target node. The documentation says this normally costs "some hundreds of milliseconds" of downtime.

That figure is for the container. It is not the figure for what runs inside it. A restart migration of an LXC holding a Docker host stops dockerd, which stops every Compose service, which means every database inside shuts down and replays its log on the other side. For a stack of three small services that is a few seconds. For a Postgres-backed photo library it is however long the application takes to come back and pass its health checks.

On a single node you never migrate anything, so this reason does not apply to you yet. It becomes decisive the day you add a second node and want to patch the first one without telling the household.

Reason two: the kernel belongs to the host#

The Proxmox wiki is direct about it: containers use the kernel of the host system, and that "exposes an attack surface for malicious users." A VM has its own kernel. A kernel bug reachable from inside a Docker container in a VM gets the attacker a guest kernel. The same bug reached from Docker nested in LXC gets them the kernel your hypervisor, your storage and every other guest are running on.

Unprivileged containers soften this. Root inside the container is mapped to an unprivileged user outside, using user namespaces, and that mapping is what makes Docker in LXC defensible at all. Which is why the answer to "should the container be privileged to make Docker work" is no: a privileged container running Docker has removed the one boundary the whole arrangement depended on.

Reason three: Docker needs two features switched on, and one has a price#

Docker will not run in a default unprivileged container. It needs two container features, both off by default:

bash
# on the Proxmox host, container stopped
pct set 110 --features nesting=1,keyctl=1
pct start 110

nesting exposes procfs and sysfs to the container so that nested containers can be created. keyctl is the more interesting one. The pct.conf reference describes it as "required to use docker inside a container", and explains that by default unprivileged containers see the keyctl() system call as non-existent. It then adds the line almost nobody quotes: "Essentially, you can choose between running systemd-networkd or docker." With keyctl on, a container using systemd-networkd can treat denied keyctl operations as a fatal error. Pick a container template that does not depend on systemd-networkd, or expect to debug networking.

Community answers have said the same thing for years. A 2022 thread on the Proxmox forum settled on nesting plus keyctl as the working configuration for an unprivileged Ubuntu container, and one of the first replies in it opened with the observation that Docker in LXC is not recommended because a VM would be more secure and more reliable.

Reason four: host upgrades change the ground under Docker#

This is the reason the forum advice underweights, and the one that actually costs people weekends. Everything Docker relies on for isolation (the kernel, the cgroup hierarchy, the AppArmor profile that confines the container) is supplied by the Proxmox host. You upgrade the host on its schedule and Docker on its own, and neither upgrade knows the other exists.

Three concrete instances:

  • cgroup v1 is gone. The Proxmox VE 9.0 release notes list "Dropped support for cgroupv1/hybrid hierarchy". The 8 to 9 upgrade guide spells out the consequence: containers running systemd version 230 or older, such as CentOS 7 or Ubuntu 16.04, are not supported on Proxmox VE 9. The admin guide puts the threshold as systemd 231 or newer for cgroup v2 support. A VM with any guest OS keeps working, because it has its own kernel and its own cgroup tree.
  • AppArmor changed. The same 9.0 release lists the move to AppArmor 4 among its known issues and breaking changes, and AppArmor is what confines every LXC container.
  • A Docker security fix broke Docker in LXC. In November 2025, runc released fixes for CVE-2025-52881 in versions 1.2.8 and 1.3.3. Inside Proxmox LXC containers, Docker then refused to start containers with open sysctl net.ipv4.ip_unprivileged_port_start file: reopen fd 8: permission denied. The runc issue explains why: AppArmor's path-based rules saw the process reaching for /proc/sys/net/... as a reach for /sys/net/..., which the LXC profile denies. The fix arrived on the host side, in lxc-pve 6.0.5-2.

That last one is the pattern in miniature. A routine apt upgrade inside the guest, applying a security fix, broke every container because of a profile on the host. The workaround that circulated on the Proxmox forum on 11 November 2025 was to set lxc.apparmor.profile: unconfined, and forum regulars objected immediately that unconfining AppArmor is a strange response to a security fix. They were right. And a Docker host in a VM never saw any of it, because its AppArmor was its own.

Why the "LXC is fine" answer keeps circulating#

It is not wrong about the costs it measures. An LXC container shares the host's memory rather than reserving it, starts in seconds, and reaches host storage through a bind mount with no virtual disk in between. On a mini PC with 8 GB of RAM those savings are real, and Docker Engine itself idles at about 120 MB either way, so the overhead you are avoiding is the VM's guest kernel and page cache, not Docker's.

What the advice does not measure is time. A blog post or forum answer is written when the setup works. It is almost never revisited after the next host major, and the failures in reason four all arrive at upgrade time, often months later, on a machine someone else configured from that post.

A backup trap specific to the LXC route#

Docker in LXC usually ends up with its data on a bind mount from the host, because that is the convenient way to give a container a big dataset. Check what your backups actually contain. The admin guide notes that the contents of bind mount points are not backed up when using vzdump, and neither are device mount points. Only mount points with the backup option enabled are included.

So a nightly Proxmox backup of an LXC Docker host can succeed every night while containing the container's root filesystem and none of the data. A VM with its data on its own virtual disk does not have this gap. Whichever you choose, restore it once and look, as Backups that actually restore argues for every backup job.

What the new OCI support is, and what it is not#

Proxmox VE 9.1, released on 19 November 2025, added the ability to create LXC containers from OCI images. The release notes split it in two: full system containers from suitable OCI images are supported, and application containers from OCI images are initial support, marked technology preview. The current 9.2 documentation still describes running application containers as a technology preview.

The mechanics matter for this decision. You pull an image with the "Pull from OCI registry" button on a storage's container template view, or upload one, then create a container from it with the wizard or pct create. The documentation says the image "is automatically converted to the LXC stack that Proxmox VE uses". There is no Docker daemon involved at any point.

That makes it a third option rather than a replacement for either. It fits a single self-contained service, such as a small web app with one image and one data directory, that you would like backed up, snapshotted and firewalled like any other guest. It does not fit a Compose stack with a database, a cache and a worker on a shared network, and it inherits every property of LXC described above, including the shared kernel and restart-only migration. Until the preview label comes off, keep it to services you would not mind rebuilding.

Where this answer stops applying#

  • No Proxmox. On a bare-metal Debian host there is no LXC layer to choose; Docker runs directly, which is the path Your first self-hosted server recommends for a first server.
  • TrueNAS as the base. TrueNAS runs Docker apps natively, and the trade-offs there are about the appliance model, covered in Proxmox vs TrueNAS.
  • Podman instead of Docker. Rootless Podman in LXC changes the privilege picture but not the shared kernel, the migration limit or the AppArmor coupling. Docker vs Podman covers the engine choice itself.
  • One node, no internet exposure, nothing irreplaceable. Then LXC is a legitimate choice and the RAM saving is the deciding factor. Just keep the container unprivileged.

What to do next#

If you already run Docker in LXC and it holds data you care about, you do not need to move it tonight. Before your next Proxmox upgrade, take a backup that includes the bind-mounted data, and plan the move to a VM for the stacks in the first row of the table. Moving a service to a new machine has the procedure, and it works the same between two guests as between two machines. Then read An update strategy that does not lose data, because the same rule applies one layer down: snapshot before you upgrade, and never let the host and the guest both change on the same night.

Questions#

Does Proxmox officially support Docker in LXC?

Not in the sense of recommending it. The container toolkit exposes the two features Docker needs, nesting and keyctl, and the keyctl description says outright that it is required to use Docker inside a container. But the Linux Container documentation says that for maximum isolation and the ability to live-migrate, nesting containers inside a QEMU VM remains a recommended practice. It works, it is exposed, and it is not the configuration Proxmox points you towards.

Can you live-migrate an LXC container in Proxmox?

No. The admin guide states that running containers cannot be live-migrated due to technical limitations. What you get instead is a restart migration: the container is shut down, moved and started on the target node, which the docs say normally costs some hundreds of milliseconds of downtime. For a container running Docker, the real cost is longer, because every Compose service inside it restarts too, including databases that replay their logs on start.

Should the LXC be privileged or unprivileged for Docker?

Unprivileged. In an unprivileged container, root inside is mapped to an unprivileged user outside, which is the only thing standing between a container escape and root on your hypervisor, since the kernel is shared. If a guide tells you to make the container privileged to get Docker working, you are removing the isolation that made LXC an acceptable choice.

Is the new Proxmox OCI image support a replacement for Docker?

Not for a Compose stack. Proxmox VE 9.1 added the ability to pull an OCI image from a registry and create a container from it, but each image is converted into an ordinary LXC container, and running application containers this way is still labelled a technology preview. It suits a single self-contained service you want managed like any other guest. A stack of four services with a shared network and volumes is still a job for Docker in a VM.

Why do so many people say Docker in LXC is fine?

Because on the day you set it up, it is. It uses less memory than a VM, starts in seconds, and shares host storage without a virtual disk. The problems are not in the setup, they are in the updates: the kernel, cgroup layout and AppArmor policy all belong to the host, so a host upgrade or a Docker security fix can stop containers from starting. Most forum advice is written on the day it worked.

Sources#

Published . Last reviewed . Found something out of date? Tell us and we will fix it and log the change.