Caching PROJ and GDAL Layers in CI
This page removes the setup cost from a spatial CI job — the conda solve, the GDAL and PROJ install, and the transformation-grid download — by pinning and caching each of them with a key that changes only when the thing itself changes. On a typical PDAL and GDAL job that is six minutes of every run, and removing it also removes a source of non-reproducibility, because a cached PROJ release is a pinned PROJ release.
Why you hit this
A spatial CI job spends most of its wall clock before it touches any data. Resolving a conda environment containing GDAL, PROJ and PDAL takes minutes on its own, and projsync pulling transformation grids adds more. Teams notice the time and reach for a faster runner, which does not help because the cost is I/O and network rather than CPU. What does help is not doing the work twice, and the same change happens to pin the geodetic behaviour — which is the more valuable half.
The wider job structure is in GitHub Actions GDAL/PDAL pipeline jobs.
Prerequisites
- A CI system with a cache action — the examples use GitHub Actions, and the same keying logic applies to GitLab, Buildkite and CircleCI.
- A lockfile for the environment:
conda-lock,environment.ymlwith pinned versions, or a container image. - Somewhere to put a container image if you take that path — GHCR, ECR or Docker Hub.
Step-by-Step
1. Prefer a pinned container over installing at job time
The largest single win is not installing GDAL at all.
jobs:
process:
runs-on: ubuntu-latest
container:
# Pin by DIGEST, never by tag — a tag is mutable and silently changes the toolchain.
image: ghcr.io/example/spatial-base@sha256:9f2c1ab4e6d0c73a1b8e5f2d4a6c8e0b2d4f6a8c0e2f4a6c8e0b2d4f6a8c0e2f
steps:
- uses: actions/checkout@v4
- run: pdal --version && gdalinfo --version && projinfo --searchpaths
A digest is the only reference that guarantees the same bytes. A tag such as :3.8 moves whenever the publisher rebuilds, which means a job that passed last week can fail today with no change in your repository — and, worse, can silently produce different coordinates because PROJ moved.
2. Cache the PROJ data directory, keyed on the release
Transformation grids are large, static, and versioned. That combination is exactly what a cache is for.
- name: Resolve the PROJ data release
id: proj
run: echo "release=$(projinfo --searchpaths >/dev/null 2>&1; \
python -c 'import pyproj; print(pyproj.__proj_version__)')" >> "$GITHUB_OUTPUT"
- name: Cache PROJ grids
uses: actions/cache@v4
with:
path: ~/.local/share/proj
key: proj-data-${{ steps.proj.outputs.release }}-v2
restore-keys: proj-data-${{ steps.proj.outputs.release }}-
- name: Fetch any missing grids
run: projsync --system-directory --list-files >/dev/null && projsync --all --quiet
Keying on the PROJ release rather than on the workflow file is the important detail. The grids belong to PROJ, not to your pipeline, so a cache keyed on the workflow invalidates whenever anyone edits an unrelated step, and one keyed on latest never invalidates when it should.
3. Cache the environment itself when a container is not an option
Where the job must install at run time, cache the package layer and key it on the lockfile.
- uses: conda-incubator/setup-miniconda@v3
with:
miniforge-version: latest
use-mamba: true
environment-file: environment.lock.yml
activate-environment: spatial
- name: Cache conda packages
uses: actions/cache@v4
with:
path: ~/conda_pkgs_dir
key: conda-${{ runner.os }}-${{ hashFiles('environment.lock.yml') }}
restore-keys: conda-${{ runner.os }}-
hashFiles on the lockfile is what makes this correct. Hashing environment.yml with loose version specifiers produces a stable key across a solve that resolved differently, so the cache returns packages that do not match the environment the job then builds.
4. Cache the source data too, when it is stable
Test fixtures and reference tiles change far less often than code.
- name: Cache test fixtures
uses: actions/cache@v4
with:
path: fixtures/
key: fixtures-${{ hashFiles('fixtures/manifest.sha256') }}
- name: Fetch anything missing
run: |
test -f fixtures/tile_utm33n.laz || \
aws s3 cp s3://twin-fixtures/tile_utm33n.laz fixtures/
sha256sum -c fixtures/manifest.sha256
The checksum verification after restore is not optional. A partially restored cache is indistinguishable from a complete one to the job, and a truncated LAZ produces a confusing failure several steps later rather than at the point it was restored.
5. Measure what each cache actually saved
Otherwise a cache that stopped working goes unnoticed, because a slow job looks like a busy runner.
- name: Record step timings
if: always()
run: |
echo "::notice title=timings::proj=${PROJ_S}s conda=${CONDA_S}s process=${PROC_S}s"
import json
import subprocess
runs = json.loads(subprocess.run(
["gh", "run", "list", "--workflow", "process.yml", "--limit", "30",
"--json", "conclusion,createdAt,updatedAt"],
capture_output=True, text=True).stdout)
import datetime as dt
durations = []
for r in runs:
if r["conclusion"] != "success":
continue
a = dt.datetime.fromisoformat(r["createdAt"].replace("Z", "+00:00"))
b = dt.datetime.fromisoformat(r["updatedAt"].replace("Z", "+00:00"))
durations.append((b - a).total_seconds())
durations.sort()
print(f"median {durations[len(durations)//2]:.0f}s | "
f"p90 {durations[int(len(durations)*0.9)]:.0f}s | n={len(durations)}")
A median that creeps upward over weeks is a cache that stopped hitting — most often because a key started including something volatile, or because the cache exceeded the size limit and is being evicted between runs.
Expected Output & Verification
A representative before and after on a matrix of twenty shards:
median 578s | p90 641s | n=30 # before
median 249s | p90 288s | n=30 # after
proj=3s conda=0s process=211s
Two checks confirm the caches are real rather than apparently real. The PROJ step should report seconds rather than minutes on a hit, and projinfo --searchpaths should list the cached directory first. And the pipeline’s own numeric output — a transformed control point, a point count — must be identical to the pre-caching run, because a cache that changed the geodetic behaviour has broken something far more important than the runtime.
projinfo --searchpaths
python - <<'PY'
from pyproj import Transformer
t = Transformer.from_crs("EPSG:4326+5773", "EPSG:32618+5703", always_xy=True)
print(t.description)
print([round(v, 6) for v in t.transform(-73.985428, 40.748817, 12.30)])
PY
Common Errors
The cache never hits. The key includes something that changes every run — a timestamp, a run number, or hashFiles over a directory the job writes into. Print the resolved key and compare it across two runs.
Cache hits and the job still installs everything. The cached path is not the one the tool actually uses. conda respects CONDA_PKGS_DIRS, and PROJ looks at PROJ_DATA and its compiled-in search path — check with projinfo --searchpaths rather than assuming.
Results changed after adding caching. A different PROJ data release was restored than the one the previous runs used. This is the failure worth taking seriously: pin the release explicitly and re-verify a control point.
Cache uploads fail on large directories. Most CI caches have a size ceiling in the low gigabytes. Cache only the grids you use — projsync --area-of-use fetches a bounding box rather than the world.
Frequently Asked Questions
Container or cached environment?
Container, where you can. It pins the whole toolchain in one digest, it starts in seconds, and it is the same artifact developers can run locally. Cached environments are the fallback when the CI system cannot run containers.
Should the PROJ data cache be shared between repositories?
If the CI system allows it, yes — the grids are identical and large. Where caches are per-repository, publishing a container that already contains them is the equivalent.
Does this affect reproducibility?
It improves it, provided the keys are right. A pinned container digest plus a PROJ release in the cache key means the geodetic behaviour is fixed, which is one of the inputs an incremental rebuild’s content hash depends on.
One caveat about restore-keys. A prefix fallback is useful for the conda cache, where a partially matching package set still saves most of the download. It is actively harmful for the PROJ cache, because a fallback restores grids from a different PROJ release, which is precisely the silent change of geodetic behaviour the exact key was there to prevent. Use restore-keys where a partial hit is a saving and omit it where a partial hit is a correctness risk.
The second caveat is about verifying a restore rather than trusting it. Cache actions report a hit when they found and extracted an archive, not when the contents are complete — a truncated upload from a previous run restores cleanly and is missing files. A checksum pass over the restored directory costs a second and converts a confusing mid-job failure into an explicit one at the point of restore.
Finally, treat the cache as an optimisation and never as a dependency. A job that cannot run with every cache cold is not cached, it is broken in a way that happens not to show — and it will show on the first run in a new fork, a new runner pool, or after a cache eviction.
Related Guides
- CI/CD Automation for Spatial Pipelines — the pipeline this speeds up
- GitHub Actions GDAL/PDAL Pipeline Jobs — the job structure being cached
- Asserting CRS and Units with pyproj — the check that catches a changed PROJ release