Public Project or Nothing: The Proxy-Cache Setting That Is Not a Preference

The kubelet pulls with no imagePullSecrets at all — 32 of 32 pods pull from harbor/dockerhub and it works only because the project is Public. Copying the 'apps' project's Private setting took two workloads down.

Public Project or Nothing: The Proxy-Cache Setting That Is Not a Preference

The signature is unmistakable and the cause is always the same:

failed to resolve image: pull access denied, repository does not exist or may
require authorization: authorization failed: no basic auth credentials

It does not mean the image is missing. It does not mean the proxy cache is broken — a crane manifest from your laptop will happily succeed while every pod fails, because your CLI is authenticated and the nodes are not. It means the project is Private and the kubelet has no credentials.

The kubelet has no imagePullSecrets — by design

This cluster's pods carry no imagePullSecrets at all. That is verified, not assumed: 32 of 32 pods pulling from the dockerhub proxy cache have none, and they pull fine — because the project is Public. The entire registry-routing design rests on that fact. Every proxy cache project the cluster pulls from must be Public, or every pull fails.

The mistake that keeps threatening this is symmetry-thinking: the apps project is Private, so surely the others should be too. But apps is Private because CI pushes there with a robot account. A proxy cache the cluster pulls from is the opposite case.

Public is not optional — the pull path with no credentials
Public is not optional — the pull path with no credentials

The 2026-09-09 incident: a Private gitlabcom project took gitlab-runner1 and gitlab-gitlab-exporter into ImagePullBackOff until the project was flipped to Public.

Why the laptop test lies

The runbook's most repeated warning: do not verify with crane manifest from a workstation. Your CLI is authenticated; the nodes are not. The test that works is a pod:

kubectl create ns pulltest
kubectl -n pulltest run t --image=harbor.private.example.com/<proj>/<img> --restart=Never
kubectl -n pulltest describe pod t | grep -cE 'no basic auth|pull access denied'

The count must be zero. Confirm, then fix in Harbor: Projects → the project → Configuration → tick Public → SAVE. Stuck pods recover on their own via ImagePullBackOff retry — deleting them just hurries it along.

The corollary is worth writing on the wall: .private. in a hostname is a naming convention, not access control. Access control is the Envoy SecurityPolicy IP allowlist in front of it.

The exact failure, end to end

The 2026-09-09 incident is worth walking once, because every step of it was ordinary. A new proxy-cache project was created for gitlabcom — endpoint tested, proxy cache enabled, scan-on-push ticked. The Access Level was set to Private, copying the apps project's setting, because apps is the project everyone had looked at most recently. Nothing errored. The project worked from the UI.

Then the first pod whose image had been rewritten to the new path tried to pull. The kubelet — holding no credentials, as every kubelet in this cluster does — asked Harbor for the manifest and received 401 with authorization failed: no basic auth credentials. gitlab-runner1 and gitlab-gitlab-exporter went into ImagePullBackOff. The failure surfaced in two workloads first, looked like a Harbor problem, and was only recognised when the pull path itself was checked:

kubectl -n <ns> get pod <pod> -o jsonpath='{.spec.imagePullSecrets[*].name}{"\n"}'
# empty output — that is normal here
curl -sI https://harbor.private.example.com/v2/ | head -1

The fix was one checkbox: Projects → gitlabcom → Configuration → Public → SAVE. Pods recovered on their own through ImagePullBackOff retry.

Why Public is safe here — and where it would not be

"Public" in Harbor means unauthenticated pull within the instance; it does not mean the internet can reach Harbor at all, because the Envoy SecurityPolicy in front of harbor.private.example.com denies everything except the allowlisted CIDRs (the Scaleway public gateway that CI, Copacetic and the kubelets egress through, plus office and VPN ranges). The defense in depth is registry-level Public + network-level allowlist. Removing either half breaks the design: a Private project breaks the pull path; removing the IP policy would make "Public" mean something very different.

The cases where Public is wrong are equally clear and are documented:

  • apps — Private, because CI pushes there with a robot account, and pull-blocking is planned once it holds real images;
  • patched — pull-blocking off and effectively internal; the nightly job writes there with the robot account, and nothing should pull from it directly;
  • any project where "prevent vulnerable images from running" is enabled — pull-blocking is a deliberate feature for images you own, never for a proxy cache feeding cluster pulls.

The rule of thumb that survived the incident: the pull-path project's access level is a property of the kubelet's credentials, not of your threat model. The kubelet has none. Design accordingly.

The detection gap

Nothing alerts on "project flipped to Private." The failure surfaces as ImagePullBackOff on whatever restarts next, which could be minutes later (the two GitLab workloads) or days later (a node drain). The pull-test in the definition of done exists partly to catch it in the same session — run immediately after creating the project, it fails in one pod-creation instead of in production. It is the cheapest alarm in the whole procedure.

Next: why registry.gitlab.com went first — ordering the queue by CVE weight.