Strengthen control of upstream Docker image repositories with pull through caching in Amazon ECR
What is pull through caching with Amazon ECR?
Pull through caching is a functionality of Amazon ECR (Elastic Container Registry) that allows you to configure a synchronization of contents of an upstream container image repository to your own Amazon ECR private registry. Essentially making your own ECR registry an intermediary between your container services and the upstream container image source.
Amazon ECR supports the creation of upstream registry caching from a range of public container registries, such as Docker Hub, Quay or GitHub Container Registry to name some.
See the official Amazon ECR documentation for more details on the functionality.
Why pull through caching?
Now why would you need to replicate an already publicly available container image repository to your own private repository? Seems a bit unneccessary you may think.
Pull through caching is definitely not a must-have functionality for most, but in certain scenarios it can be very powerful. Such as some of the following scenarios.
Public image distribution in restricted or no internet access environments
In environments with restricted or no internet access for workloads or supporting infrastructure, you can use ECR pull through cache to distribute publicly available container images to your workloads by using VPC service endpoints for communicating with ECR. This means that your workloads can fetch the images without traversing the internet themselves.
Public container image registry outages or End of Life
Caching container images in your own ECR private registry protects your from potential outages in the upstream source or even End of Life scenarios for certain development projects.
Do you remember when Docker announced the sunsetting of their Free Team plans back in 2023? Thankfully, Docker went back on their words due to massive complaints. This caused many developers to worry about their container image consumption from Docker Hub. With pull through caching in ECR, you could buy yourself valuable time in scenarios like this.
Governance over allowed container image sources
With pull through cache, you can effectively govern allowed container image sources in your environments by applying Service Control Policies (SCP) to restrict the use of container images from other sources than the ones you control, for example pull through cache repositories.
Protection from upstream changes to existing tags with tag immutability
Maybe the coolest functionality of pull through caching in ECR is in combination with the ECR feature of tag immutability. This feature ensures that once a tag is published, the container image for that tag cannot be overwritten while still retaining the same tag name. What this effectively means for pull through caching is that you can protect yourself from potentially malicious tag changes in the upstream image repositories. This could be a scenario where for instance a trusted supplier in Docker Hub, such as HashiCorp were to be under an attack where existing image tags were to be injected with malicious code while retaining their tag names.
Without tag immutability, you could in this scenario end up with completely different images in your environment without having changed anything in your code or container image tags.
Important to note: Tag immutability will also break the possibility to rely on tags such as latest, since the tag cannot be reused once set for an image.
Key components
The key components of a pull through cache in Amazon ECR are:
-
Pull through cache rule is a rule type where you define your private cache repository prefix, such as
mycache-docker-ioand the upstream registry URL, i.e.registry-1.docker.io. This is a caveat of pull through caching: You cannot scope down your rule to a specific repository, but must configure pull through for an entire registry. -
Repository creation template (optional) is a template where you define the configuration that the pull through cache created repositories will use, such as repository permissions policy, tag immutability and encryption. If you don’t define a repository creation template that matches your private cache repository prefix, the repositories will end up with a default configuration from AWS.
-
Custom IAM role (optional) is necessary if your repository creation template dictates the use of AWS Key Management Service (KMS) keys for encryption or you have other requirements that must be implemented. Since pull through caching is targeting an entire container image registry (such as registry-1.docker.io) as the upstream cache source, the IAM repository policy associated with the repositories created from the cache rule can utilize conditions to only allow use of specified container image repositories. Note however that the cache rule itself operates at the ECR registry level, and so anyone with permissions can still trigger a repository to be cached, even if the repository rule says it can’t be used. Repository lifecycle policies or custom function logic can help determine and clean up unwanted cached images and repositories.
-
Pull through cache container image address is the address your container services must use when specifying which container image to use. This adress is predictably built in the following format:
{account}.dkr.ecr.{region}.amazonaws.com/{cache-prefix}/{UpstreamRepositoryName}:{ImageTag}. Broken down:{account}is the AWS account ID where the Amazon ECR pull through cache rule is configured.{region}is the AWS region where the Amazon ECR pull through cache rule is configured.{cache-prefix}is the prefix set in the pull through cache rule. As the name implies, it is a prefix that is added to the upstream repository name.{UpstreamRepositoryName}is the repository or (container image if you like) in the upstream container image registry. For examplepython.{ImageTag}is the tag name of the image you wish to use. For examplelatestorv1.0.0.
Pull through cache flow
Using the diagram above as an example, the following flow is followed when pulling an initial container image with pull through caching using Amazon ECR.
- The container service of your choice (in this case Elastic Container Service) initiates a
docker pullcommand using the pull through cache container image address. - The pull through cache rule is matched with the prefix in the container image address used. and ECR is triggered to pull the image and specified tag from the upstream source.
- If a repository creation template is matched with the pull through cache prefix, an ECR repository is created for the image based on the template configuration. If a template is not defined, an ECR repository is created for the image with default ECR configuration.
- The container image is ultimately successfully pulled by the ECS service from the newly created ECR repository.
For consecutive docker pulls of the same container image and specified tag, Only step 1 and 4 is followed. If a new tag is specified that has not previously been cached, the entire flow is used again in order to cache the new tag.
Setting up pull through caching
Configuring pull through caching in Amazon ECR is fairly easy. Below is an example written in Terraform for setting up a simple pull through cache rule for registry-1.docker.io with a repository creation template. In order to keep it short, I have not included KMS key encryption of the repositories or with tag immutability configured.
The repository policy used by the template allows any principal in an AWS organization to pull images from the repository.
data "aws_region" "current" {}
resource "aws_ecr_pull_through_cache_rule" "registry" {
ecr_repository_prefix = "mycache-docker-io"
upstream_registry_url = "registry-1.docker.io"
}
resource "aws_ecr_repository_creation_template" "pull_through" {
prefix = "mycache-docker-io"
description = "ECR repsitory creation template for pull-through cache of registry registry-1.docker.io"
applied_for = [
"PULL_THROUGH_CACHE",
]
repository_policy = data.aws_iam_policy_document.pull_through_template.json
}
# IAM policy document for repository template
data "aws_iam_policy_document" "pull_through_template" {
statement {
sid = "AllowPullFromOrganization"
effect = "Allow"
principals {
type = "*"
identifiers = ["*"]
}
actions = [
"ecr:BatchGetImage",
"ecr:GetDownloadUrlForLayer",
"ecr:BatchImportUpstreamImage",
"ecr:CreateRepository"
]
condition {
test = "StringEquals"
variable = "aws:PrincipalOrgID"
values = ["your-organization-principalid"]
}
}
}
Important! Regarding ECS and Task execution role permissions with ECR pull through caching
If you intend to use Elastic Container Service (ECS) to pull images from a pull through cache repository in Amazon ECR, you must use a custom Task execution role policy that includes the permission ecr:BatchImportUpstreamImage.
The AWS managed role policy AmazonECSTaskExecutionRolePolicy does NOT contain this permission which is crucial for the initial pulling of a container image:tag combination if it has not already been cached by the pull through cache rule.
Not including this permission in your ECS Task Execution role will result in some strange behavior which is difficult to pin-point the root cause of. In CloudTrail you will see API calls of BatchImportUpstreamImage failing with access denied in the account you have your ECR pull through cache configured, but from the event it is not clear who the requester or source of the API call is.
I assume the same permission must be set for other AWS container services (such as App Runner) aiming to use ECR pull through cache container images as well.