Rotting Layers: The Hidden Price of an Unmanaged Container Registry
Photo: Federal Bureau of Investigation, Public domain, via Wikimedia Commons
There is a particular kind of technical debt that accumulates not through bad decisions but through no decisions at all. Container registries are one of its most reliable breeding grounds. A team ships a feature, pushes an image, and moves on. The CI pipeline runs nightly builds that tag images with commit hashes no one will ever type again. A developer tests a fix in a staging environment, pushes a snapshot, and forgets it entirely. Multiply these moments across a year, across a dozen engineers, across multiple services—and what you have is not a registry. It is a graveyard.
The problem is not dramatic. It does not surface in an incident postmortem or trigger an on-call alert. It simply grows, quietly, until one day a storage invoice arrives that no one can fully explain, a security scanner flags an image from eighteen months ago that still carries a critical CVE, or an engineer spends three hours trying to determine whether a particular image tag is safe to delete or silently depended upon by something in production.
Understanding how this situation develops—and what it actually costs—is the first step toward reversing it.
How Registries Become Ungoverned
Container image accumulation follows a predictable pattern. In the early stages of a project, teams push images manually or through loosely configured pipelines. Tags are informal: latest, dev, test-again-v2. As the team grows and CI/CD becomes more automated, build systems begin generating images on every commit, every pull request merge, and every scheduled pipeline run. These images pile up without ceremony.
The absence of a deletion policy is rarely intentional. Most teams assume they will establish governance later, once the product stabilizes. Later rarely arrives. What arrives instead is a registry containing thousands of images—some actively deployed, many orphaned, and a significant portion completely unknown in origin or purpose.
Untagged images present a particular challenge. When a new image is pushed under an existing tag, the previous image loses its tag reference but remains in the registry as a dangling layer. These untagged images are invisible to casual inspection, consume storage silently, and often carry older, unpatched software layers that no one has thought to address.
The Three Cost Vectors Teams Consistently Underestimate
Storage at scale is not trivial. Cloud-hosted registries—whether Amazon ECR, Google Artifact Registry, or Azure Container Registry—charge for storage by the gigabyte. A single uncompressed image layer for a Node.js or Java application can reach several hundred megabytes. Multiply that by hundreds or thousands of accumulated images and the monthly cost becomes material. Teams that have audited their registries for the first time frequently discover that a meaningful percentage of their infrastructure spend is attributable to images that have not been pulled in six months or more.
Security exposure compounds invisibly. Every layer in a container image represents a potential attack surface. An image built against a base that contained a known vulnerability at build time will carry that vulnerability indefinitely—even if the base image has since been patched. In a governed registry, outdated images are replaced or removed on a defined schedule. In an ungoverned one, they persist indefinitely. Security scanners can identify these vulnerabilities, but without a lifecycle policy that acts on the findings, the scan results accumulate alongside the images themselves: a catalog of known risks that no one has the mandate to address.
Cognitive load erodes engineering velocity. When engineers cannot determine which images are in active use, they become reluctant to delete anything. This is a rational response to an irrational situation, but it has a real cost. Decisions about what is safe to prune require investigation—checking deployment manifests, querying orchestration platforms, reviewing recent pipeline logs. Each of these investigations takes time that could be spent building. Over a team of ten engineers, even a modest recurring tax on registry-related confusion adds up to significant lost capacity over the course of a quarter.
What Lifecycle Governance Actually Looks Like
The good news is that registry discipline does not require sophisticated tooling or a dedicated platform team. It requires clear policies, consistent enforcement, and a shared vocabulary around what an image version means.
Define retention rules by purpose, not by age alone. A blanket rule that deletes images older than ninety days will eventually destroy something important. A more durable approach distinguishes between image categories: production release images, which should be retained for a defined number of versions; CI build images, which can be pruned aggressively after a short window; and staging or ephemeral images, which should be deleted immediately after the associated environment is torn down. Most modern registries support tag-based lifecycle policies that can enforce these distinctions automatically.
Establish a versioning convention and enforce it at the pipeline level. Images that carry meaningful, parseable version tags—semantic version strings, release identifiers tied to your deployment manifest—are far easier to reason about than images tagged with raw commit hashes or timestamps. A convention that makes the relationship between an image and a deployed artifact explicit reduces the investigation burden dramatically. When every engineer on the team can look at a tag and understand its provenance, the question of whether an image is safe to delete becomes answerable without an archaeological dig through pipeline logs.
Integrate provenance tracking early. Knowing which images are actually running in which environments is the prerequisite for any meaningful cleanup effort. Container orchestration platforms like Kubernetes expose this information through their APIs, and several open-source and commercial tools can aggregate it into a view that maps registry contents to live deployments. Teams that invest in this visibility early find that the ongoing cost of registry governance drops substantially—because the question "is this image in use?" has a reliable answer.
Treat untagged image cleanup as routine maintenance. Scheduling a weekly or monthly job to remove dangling, untagged images is a low-risk, high-return practice. These images are by definition unreferenced by any deployment manifest. Removing them reclaims storage immediately and reduces the surface area that security scanners must evaluate.
The Versioning Discipline That Prevents the Problem
Registry bloat is ultimately a versioning problem. Teams that treat container images as ephemeral build artifacts—rather than as versioned, governed software components—will inevitably accumulate more images than they can manage. The discipline that prevents this is the same discipline that governs any artifact: clear ownership, explicit lifecycle expectations, and a tagging scheme that communicates intent.
At ExVersion, we return often to the principle that what you do not version deliberately, you version accidentally—and accidental versioning always costs more in the end. A container registry that has never been audited is not a neutral resource. It is a growing liability, accumulating storage costs, security risk, and engineering confusion with every pipeline run.
The teams that manage this well are not the ones with the most sophisticated tooling. They are the ones that decided, early and explicitly, what their images mean, how long they should live, and who is responsible for cleaning up after them. That decision is available to every engineering organization, regardless of size or stack. The cost of not making it, as registries across the industry quietly demonstrate, is considerably higher than most teams expect.