# Photo Analyzer
Integrated, restart-safe photo analysis, duplicate review, metadata, upload, and
archive workflow. Planning lives in `INTEGRATED_PIPELINE_CONCEPT.md` and
`delivery_backlog/`.
## Application (`photo_pipeline`)
The target application lives in `photo_pipeline/` (FastAPI + SQLAlchemy + Alembic).
Install it into a virtualenv once:
```bash
python3.12 -m venv .venv
.venv/bin/pip install -e ".[vision]" # drop [vision] for a review-only install
```
Then run the two processes:
```bash
.venv/bin/python -m photo_pipeline migrate # apply database migrations
.venv/bin/python -m photo_pipeline serve # API + review UI at 127.0.0.1:8000/app/
.venv/bin/python -m photo_pipeline worker # second terminal: runs the jobs
```
The server enqueues work and serves the UI; nothing actually scans, scores,
analyses, uploads, or archives without a worker. `work_item/scripts/python` is the
*helper's* launcher — it prefers Conda base and falls back to a bare system
interpreter, so it is not how the application is run.
Configuration comes from `PHOTO_PIPELINE_*` environment variables (see
`photo_pipeline/config.py`); secrets are referenced, never logged.
### Configuration file
`.env` in the working directory is read at startup, or any path named by
`PHOTO_PIPELINE_ENV_FILE`. It is parsed, never executed: `KEY=value` lines,
`#` comments, optional quotes — no interpolation and no `export`. **Anything already
exported wins**, so the file is the standing configuration and the shell is the
override for one run.
The archived CLI's variable names still work, so an existing `photo_analyzer.env`
can be used as-is:
| in the file | applied as |
|---|---|
| `LLM_API_KEY` / `GEMINI_API_KEY` | `OPENAI_API_KEY` |
| `LLM_BASE_URL` | `OPENAI_BASE_URL` |
| `LIBRARY` | `PHOTO_PIPELINE_LIBRARY_ROOTS` |
`.env` and `*.env` are gitignored and denied by the work-item safety checks: the
file holds a real key and must never be committed.
### API access (US07-02)
The app listens on loopback, so its attacker is another page in the same browser.
Every `/api/v1` route except `health/live`, `health/ready`, and `session` requires
the application session, and every mutation requires its CSRF token as well:
```bash
BASE=http://127.0.0.1:8000
TOKEN=$(curl -sc /tmp/pp.jar $BASE/api/v1/session | python -c 'import json,sys; print(json.load(sys.stdin)["csrf_token"])')
curl -sb /tmp/pp.jar -H "X-CSRF-Token: $TOKEN" -X POST $BASE/api/v1/albums/proposals -d '{}' -H 'Content-Type: application/json'
```
The session is per server process — restarting `serve` invalidates it, and the
browser client re-bootstraps by itself. Requests are also refused when the `Host` is
not a loopback name (DNS rebinding), when `Origin` is any other origin, when
`Sec-Fetch-Site` says the request came from another site (an `
` pointed at a
thumbnail), or when the body exceeds `PHOTO_PIPELINE_MAX_REQUEST_BYTES`. There is no
CORS middleware at all, so no other origin can read a response.
### Reaching it through a hostname or proxy (US08-01)
| variable | meaning |
|---|---|
| `PHOTO_PIPELINE_ALLOWED_HOSTS` | comma-separated extra names the app answers to; empty means loopback only |
| `PHOTO_PIPELINE_ACCESS_SECRET` | traded for the session cookie at `GET /api/v1/session` via `X-Access-Secret` |
| `PHOTO_PIPELINE_TRUSTED_PROXIES` | comma-separated peer addresses whose `X-Forwarded-Proto`/`X-Forwarded-Host` are believed |
Being reachable *was* the authentication: whoever could open `127.0.0.1:8000` owned
the library. So naming any non-loopback host — or binding to one, `0.0.0.0` included
— makes the access secret mandatory, and `serve` refuses to start without it rather
than publishing the library. Loopback-only deployments need no secret and behave
exactly as before.
```bash
curl -sc /tmp/pp.jar -H "X-Access-Secret: $PHOTO_PIPELINE_ACCESS_SECRET" \
https://photos.example.com/api/v1/session
```
The browser asks for the secret once per tab and keeps it in `sessionStorage`.
Wrong secrets are rate-limited (5 per minute) and logged with the caller's address
only. `Host` and `Origin` are judged against the configured names; the *external*
scheme and host come from the forwarded headers only when the request arrived from a
`PHOTO_PIPELINE_TRUSTED_PROXIES` address, so a client cannot declare its own origin,
and the session cookie is marked `Secure` when that external scheme is HTTPS. Health
endpoints stay reachable without the secret so an orchestrator can restart the
container; nothing else does.
## Container image (US08-02)
One image runs either role. It is built from a clean checkout with no arguments:
```bash
docker build -t photo-pipeline:dev .
```
Everything external is pinned, and the build fails rather than drifting: the Python
base image by tag *and* digest, `exiftool` by its Debian package version (verified
against `exiftool -ver`), and `immich-go` by release version and per-architecture
SHA-256 of the release asset. The verified versions become image labels and
`/etc/photo-pipeline/versions.json`, which `diagnostics` reports as `tools[].pinned`
beside the version actually installed — so a replaced binary shows up as a
`tool_version_drift` warning instead of as a misparsed upload report.
| build argument | default | why change it |
|---|---|---|
| `UID` / `GID` | `1000` | must match the owner of the mounted photo library |
| `PYTHON_IMAGE` | pinned digest | upgrading the base image |
| `EXIFTOOL_VERSION` | Debian package version | upgrading exiftool |
| `IMMICH_GO_VERSION` + `IMMICH_GO_SHA256_AMD64`/`_ARM64` | pinned release | upgrading the uploader (take the digests from that release's `checksums.txt`) |
The first argument is the role, and every other management command still works:
```bash
docker run --rm -v /srv/photos:/srv/photos -v pp-data:/data \
-e PHOTO_PIPELINE_LIBRARY_ROOTS=/srv/photos photo-pipeline:dev migrate
docker run -d -p 127.0.0.1:8000:8000 -v /srv/photos:/srv/photos -v pp-data:/data \
-e PHOTO_PIPELINE_HOST=0.0.0.0 -e PHOTO_PIPELINE_ACCESS_SECRET=... \
-e PHOTO_PIPELINE_LIBRARY_ROOTS=/srv/photos photo-pipeline:dev serve
docker run -d -v /srv/photos:/srv/photos -v pp-data:/data \
-e PHOTO_PIPELINE_LIBRARY_ROOTS=/srv/photos photo-pipeline:dev worker
```
One role per container: `serve` and `worker` each take the library process lock for
their role (US07-05), so no supervisor starts both. The container refuses to run as
UID 0 — files it renames or writes must keep the ownership the host library expects —
and `/data` is the persistent volume holding the database, journals, backups, and
thumbnail cache. Binding to `0.0.0.0` makes the access secret mandatory
([above](#reaching-it-through-a-hostname-or-proxy-us08-01)); `serve` refuses to start
without it. The declared `HEALTHCHECK` polls `/api/v1/health/ready`, so a container
whose database is unmigrated or misconfigured is never reported healthy.
## Composed runtime (US08-03)
`docker-compose.yml` is the deployment: one `serve` container, one `worker`
container, one bind-mounted library, one data volume, and a one-shot `migrate` that
both roles wait for.
```bash
cp .env.example .env && $EDITOR .env # nothing has a value in it; fill in yours
docker compose up -d --build
```
`.env.example` lists every `PHOTO_PIPELINE_*` variable with its default in a comment
and no values at all. The three the composition cannot start without are
`PHOTO_PIPELINE_LIBRARY_HOST_PATH` (the library on this host),
`PHOTO_PIPELINE_LIBRARY_ROOTS` (where it is mounted *inside* the container), and
`PHOTO_PIPELINE_ACCESS_SECRET` — publishing the port means the app is reachable from
outside the container, so US08-01 makes the secret mandatory. Set
`PHOTO_PIPELINE_UID`/`GID` to the owner of the library: what the containers rename
and rewrite keeps that ownership.
| invariant | how the composition keeps it |
|---|---|
| one writer | `serve` and `worker` take their role's library lock (US07-05) in the shared `/data` volume, so `--scale worker=2` is refused by the lock, not by convention |
| migrations first | `migrate` runs the backup-then-migrate path and must exit 0 before `api` and `worker` start; a failed upgrade leaves the previous database and its pre-migration backup intact |
| container paths | `PHOTO_PIPELINE_LIBRARY_ROOTS` is both the mount target and the configured root; a root that is not mounted makes `serve`/`worker` exit 5 at startup instead of writing into the container's throwaway layer |
| loopback by default | the port is published to `127.0.0.1` unless `PHOTO_PIPELINE_PUBLISH_ADDRESS` says otherwise, and exposing it needs the hostname in `PHOTO_PIPELINE_ALLOWED_HOSTS` plus the access secret |
| restart safety | both roles are `restart: unless-stopped` with a 30 s stop grace period, and a job interrupted by a restart resumes exactly as it does on a host restart |
The `data` volume holds the database, its write-ahead log, the thumbnail cache,
journals, and backups. **It must stay on a local filesystem** — SQLite in WAL mode
needs real local locking, so NFS, SMB, and network volume drivers are unsupported
there and corrupt the database rather than slow it down. The library bind mount has
no such restriction.
Operating the deployment is operating the same CLI:
```bash
docker compose run --rm --no-deps api diagnostics
docker compose run --rm --no-deps api backup --reason pre-upgrade
docker compose run --rm --no-deps api verify-backup /data/backups/
docker compose run --rm --no-deps api restore /data/backups/ --into /data/restored
docker compose logs -f worker
```
`--no-deps` keeps a one-off command from starting a second stack; `api` is only the
service the command borrows the image and mounts from. `restore` is deliberately not
an API call: it replaces the state of an installation and belongs to a stopped one,
so stop `api` and `worker` first and restart them against the restored directory.
## Publishing and deploying (US08-04)
Two workflows in `.gitea/workflows/` make `main` the only path to the running stack:
| workflow | runs on | does |
|---|---|---|
| `test.yml` (**Test**) | every pull request, and every push to `main` | installs exiftool, the package, and Chromium, then runs the suites configured in `work_item/.work-item.yml` |
| `deploy.yml` (**Deploy**) | a **successful** `Test` run on `main`, or manual dispatch | builds the image, pushes it, triggers the Portainer webhook, prunes dangling layers |
Deploy waits for `Test` through `workflow_run`, so publishing is downstream of a green
suite rather than of a push. `completed` is not `success`: the publish job runs only on
`conclusion == 'success'`. Because an unmatched `workflow_run` filter does not fail but
simply never fires, `tests/integration/test_deploy_workflows.py` asserts the link from
both ends — renaming either workflow, or dropping one of the configured suites out of
the test job, fails that test rather than silently unhooking the gate.
The image is `git.domverse-berlin.eu/domverse/photoanalyzer`, tagged with the commit
SHA and `latest`. The commit tag is pushed first, so a `latest` that exists always has
a commit tag beside it and **a rollback is a tag change, not a rebuild**:
```bash
docker pull git.domverse-berlin.eu/domverse/photoanalyzer:
docker tag git.domverse-berlin.eu/domverse/photoanalyzer: \
git.domverse-berlin.eu/domverse/photoanalyzer:latest
docker push git.domverse-berlin.eu/domverse/photoanalyzer:latest
# then trigger the Portainer webhook, or redeploy the stack from Portainer
```
Manual dispatch defaults to **dry run**: it builds and pushes `scratch-` only, and
touches neither `latest` nor the running stack. Clear the `dry_run` input to publish by
hand. A workflow-level `concurrency: deploy` group with `cancel-in-progress: false`
queues deploys instead of overlapping them, so two commits can never race to `latest`.
### Secrets
Repository secrets — used only by the pipeline:
| secret | purpose |
|---|---|
| `REGISTRY_USER` | registry user with `write:package` |
| `REGISTRY_TOKEN` | that user's token, passed on stdin, never as an argument |
| `PORTAINER_WEBHOOK_URL` | the stack's auto-update POST URL; `curl --fail` makes a refused redeploy a failed workflow |
Runtime secrets — `PHOTO_PIPELINE_VISION_API_KEY`, `PHOTO_PIPELINE_IMMICH_API_KEY`,
`PHOTO_PIPELINE_ACCESS_SECRET` — are **not** repository secrets and are not in the
image. The stack is managed by Portainer from git (`docker-compose.yml`), and those
values live in the Portainer stack's environment, so rotation is one place and a
repository read discloses nothing. A test asserts the workflows never name them.
## Testing
One offline command runs the whole suite (unit, integration, and browser
end-to-end); it needs no network and uses only deterministic synthetic fixtures:
```bash
work_item/scripts/python -m pytest tests -q
```
Browser end-to-end tests require a one-time Playwright browser install:
```bash
python -m playwright install chromium
```
### Phase A acceptance gate
Phase A (Epic E01: shared identity, inventory, duplicates, thumbnails, review UI) is
gated by a reproducible end-to-end suite:
```bash
work_item/scripts/python -m pytest tests/e2e tests/integration -q
```
- `tests/e2e/test_phase_a_pipeline.py` launches the real API process against a fresh
database and a deterministic fixture library, then drives scan, `_IGNORE/`
exclusion, move reconciliation, exact/fuzzy duplicate review, thumbnail
orientation, canonical selection, browser reload, and a full process restart —
asserting durable API and database state after the restart.
- `tests/story_traceability.json` maps every delivered story to its tests;
`tests/e2e/test_traceability.py` fails if a Phase A story loses coverage or a test
file is left unexercised.
### Phase B acceptance gate
Phase B (Epic E02: durable jobs, workflow shell, safety/analysis views) is proven
through the real process and browser boundaries. One command runs the Phase B API,
worker-recovery, and browser (Playwright) suites:
```bash
work_item/scripts/python -m pytest tests/e2e -m phase_b -q
```
- `tests/e2e/test_phase_b_pipeline.py` launches the real server and durable worker as
child processes and drives them only over HTTP/SSE: analysis start/progress,
resumable SSE reconnect, the polling fallback, cancellation, per-asset error
inspection, one-mutating-job rejection during read-only browsing, the NSFW→vision
privacy gate, and durability across a full restart.
- `tests/e2e/test_worker_kill.py` kills a worker mid-item and proves a fresh worker
resumes the fenced job (the "resume" journey).
- `tests/e2e/test_analysis_browser.py` starts a job from the Analyze view and watches
live progress arrive over the browser's real SSE adapter.
The full Phase B regression, including the unchanged Phase A gate, is the whole
end-to-end suite:
```bash
work_item/scripts/python -m pytest tests/e2e -q
```
### Phase C acceptance gate
Phase C (Epic E03: album evidence, naming policy, versioned proposals, Albums view)
is proven through the real process and browser boundaries. One command runs the
Phase C API and browser (Playwright) suites with the deterministic naming provider:
```bash
work_item/scripts/python -m pytest tests/e2e -m phase_c -q
```
- `tests/e2e/test_phase_c_pipeline.py` drives a real server over HTTP: evidence
aggregation, generation through the deterministic naming fake (asserting the exact
provider inputs and that no file path or asset ID ever reaches it), provider
failure and retry, invalid names, path-separator sanitization, editing with
optimistic versions, stale-evidence approval refusal, valid approval, and
durability across a full restart.
- `tests/e2e/test_albums_ui.py` covers the browser journeys: evidence display,
editing, prompt validation, collision guidance, approval, stale conflict, and
keyboard operation.
- Both suites assert that **no fixture path changes** — Phase C proposes names and
never renames.
The deterministic naming provider is enabled only by test configuration
(`PHOTO_PIPELINE_FAKE_NAMING_LOG`); without it the application falls back to the
offline naming-policy name. Phase A and B suites remain green in the full run above.
### Phase D acceptance gate
Phase D (Epic E04: guarded renaming) is the first phase that changes the library on
disk, so its gate is the strictest. One command runs the rename API journeys, the
filesystem fault injection, and the browser suite:
```bash
work_item/scripts/python -m pytest tests/e2e -m phase_d -q
```
- `tests/e2e/test_phase_d_pipeline.py` drives a real server over HTTP: plan and
export, confirmation with the plan version and checksum (a stale token is refused
without touching disk), a valid apply, the case-only rename procedure, a collision
whose occupant survives, a source that changed after planning, and durability
across a full restart.
- **Fault injection is real.** `PHOTO_PIPELINE_FAULT_AFTER=` kills the
server process the instant that state is persisted. The suite crashes it at every
journal transition in turn (`moving`, `moved`, `database_updated`, `verified`),
starts a fresh process against the same database and library, and requires recovery
to converge from journal and disk evidence alone — with the asset set, the stable
IDs, and every content hash unchanged. Ambiguous evidence is never guessed: it stays
classified `manual` and keeps blocking. An unresolved rename is the cancellation
boundary — there is no cancel once a run starts, and unrelated mutations (album
proposal generation and approval) are refused with 409 `rename_recovery_required`
until it is resolved, while reads stay available.
- `tests/e2e/test_renames_ui.py` covers the browser journeys: preview of every
affected path, confirmation carrying the server-issued token, apply with progress
and terminal verification, stale confirmation, collision, interruption, recovery,
rollback, keyboard confirmation, and the view still matching the journal after a
server restart.
The fault barrier is test-only configuration; without `PHOTO_PIPELINE_FAULT_AFTER`
the apply path has no crash points. Phases A–C remain green in the full run above.
### Phase E acceptance gate
Phase E (Epic E05: Immich upload) is the one stage the application cannot take back,
so its gate runs the fake-uploader suites, the black-box upload API journeys, and the
browser suite as a single command:
```bash
work_item/scripts/python -m pytest -m phase_e -q
```
- `tests/integration/test_upload_*.py` drive a **real executable** standing in for
`immich-go` through the real adapter and `subprocess` — argument construction,
output bounding, report parsing, verification, and killing a running process.
- `tests/e2e/test_phase_e_pipeline.py` drives a real server and a real durable worker
over HTTP: credential failure and an unreachable server, preflight blockers and the
explicitly approved partial scope, a new album, an exact duplicate, an upgrade, a
retryable failure and its successful retry, a lost acceptance response, verification
against Immich, an inconclusive answer resolved by an operator with evidence, bytes
edited after upload, cancellation and resume, and an interrupted attempt recovered
across a restart.
- **EXIF precedes upload** is asserted, not assumed: an album without its verified
safety and analysis checkpoints cannot be approved, and the uploader's own argv log
proves it was never executed. Each finished upload re-hashes the files in the folder
the uploader was handed and requires the persisted SHA-256/SHA-1 to match.
- **No secret is retained.** The API key is a sentinel string; after a full upload and
verification it must appear in the uploader's argv and nowhere else — not in the
database, the retained report, or any response the browser can read.
- `tests/e2e/test_uploads_ui.py` covers the browser journeys (preflight preview,
confirmation, progress, stopping a run, verification, manual resolution, stale
bytes, and recovery after a restart).
Phases A–D remain green in the full run above.
### Phase F acceptance gate
Phase F (Epic E06: archive lifecycle) is the only stage that *removes* originals
from the library, and the only one whose storage can walk away in someone's bag.
One command runs the archive fault-injection suites, the black-box archive and
restore API journeys, and the browser suite:
```bash
work_item/scripts/python -m pytest -m phase_f -q
```
- `tests/integration/test_archive_*.py` and `tests/integration/test_restore.py`
drive real files on real filesystems: preflight against a mounted, missing,
swapped, read-only, or full medium; copy-verify-remove and the same-filesystem
move path; and a crash at **every** persisted journal transition in both transfer
modes, asserting that no source is ever removed without a durable, byte-identical
archive copy.
- `tests/e2e/test_phase_f_pipeline.py` drives a real server and a real durable
worker over HTTP: preflight blockers (offline medium, wrong volume, insufficient
capacity, bytes changed after upload), a verified archive whose manifest, hashes,
and path history are checked on the medium itself, a worker killed at each of
`transferring`, `verified`, `removing`, `source_removed`, and `complete`, the
evidence-based recovery that follows, offline deduplication of an exact and a
fuzzy copy while the medium is away, mount return, restore, and a collision that
restores beside its occupant.
- **Archived is not missing.** An unmounted medium leaves its photos
`archived_offline` — still hashed, still in the duplicate indexes, still
previewable through their protected thumbnails — and a rescan neither prunes nor
flags them.
- **Ambiguity is never guessed.** A journal state the medium contradicts stays
`manual`, offers no automatic action, and keeps blocking further archiving until
a human decides.
- `tests/e2e/test_archive_ui.py` covers the browser journeys (preview with
destination identity and reclaimable bytes, blockers and mount instructions,
progress split into transfer/verification/removal, interruption and recovery,
offline browsing, restore, collision, keyboard confirmation, and reload).
Phases A–E remain green in the full run above.
## Media and metadata hardening (US07-03)
Every pixel the application reads goes through `photo_pipeline/imaging.py`: the
declared dimensions are checked before anything is decoded, Pillow's
decompression-bomb warning is treated as a refusal, JPEG decodes near the requested
size, and each decoder failure becomes one of two typed errors. A damaged file is a
per-item error with a persisted code, never a failed scan or a dead worker.
Every metadata stage ends with an EXIF checkpoint (`services/exif_checkpoint.py`):
snapshot, write the owned keywords, read back, prove the owned fields landed and that
nothing else moved, refresh the file hash. A field the stage does not own that
changed anyway makes the checkpoint `divergent` — recorded in `exif_projections`,
shown in the review queue, never repaired behind the user's back, and not counted as
verified, so upload stays blocked.
The golden corpus that proves all of it is generated, not committed:
`tests/fixtures/media_corpus.py` declares every format, orientation, profile,
damage, and metadata case with its expected outcome, and the suite regenerates it
twice to prove it does not drift.
```bash
work_item/scripts/python -m pytest tests/integration/test_media_hardening.py tests/integration/test_exif_checkpoints.py -q
```
## Concurrency and crash recovery (US07-04)
Crash safety is proven by crashing. `photo_pipeline/faults.py` defines the control
points — the persisted transitions of the rename, archive, EXIF, upload, and job
lanes — and arms one only when `PHOTO_PIPELINE_FAULT_AFTER` names it, at which
point the process dies the way a `SIGKILL` does. There is no endpoint and no
configuration field that can reach a barrier; a deployment that never sets the
variable can never hit one.
The race suite runs each scenario several times with a seed recorded on the test
result (`race_seed`) and asserts invariants rather than schedules: work is never
claimed or executed twice, a stale fencing token never commits, no file body is
lost or overwritten, and the database still passes `PRAGMA integrity_check`.
```bash
work_item/scripts/python -m pytest tests/integration/test_concurrency_races.py \
tests/integration/test_fault_matrix.py tests/e2e/test_crash_recovery.py -q
# replay a failure, or soak for new interleavings
PHOTO_PIPELINE_RACE_SEED=1234 PHOTO_PIPELINE_RACE_REPEATS=50 \
work_item/scripts/python -m pytest tests/integration/test_concurrency_races.py -q
```
Any failing test keeps its evidence: the temporary database (with its write-ahead
log), the journals, the logs, the recorded seed, and a SHA-256 manifest of every
file in the temporary library are copied to `.artifacts//` before pytest
deletes the directory. Point `PHOTO_PIPELINE_TEST_ARTIFACTS` elsewhere to collect
them from CI.
## Release gate (US07-07)
One command runs every suite in an isolated stack and keeps the evidence:
```bash
work_item/scripts/python -m photo_pipeline release-gate --output data/release/$(date -u +%Y%m%dT%H%M%SZ)
```
It fails — and exits non-zero — when any stage fails, when a suite skips a test for
a reason that is not a documented environment limit (`exiftool not installed`,
`root ignores directory permissions`), or when the story matrix has a hole. The
evidence directory holds `release-report.json` (revision, per-stage result, timings,
summaries), `logs/.log`, and `CHECKSUMS.sha256` over both.
**The story matrix** lives in `tests/story_traceability.json`: every story under
`delivery_backlog/stories/` is either mapped to test files that exist, or listed in
`planned` as an accepted but unimplemented story. A story that is neither, or a
mapping to a file that has been deleted, fails the gate.
**The journey** (`tests/e2e/test_release_journey.py`) takes one fresh library through
discovery, duplicate review, safety, analysis, EXIF verification, album proposal,
guarded rename, rescan, upload with server-side verification, archive, offline
deduplication, and restore — over HTTP against real server and worker processes,
with a full restart in the middle and at the end.
### Real-library dry run and approval
Before the application is pointed at photos that cannot be replaced:
```bash
work_item/scripts/python -m photo_pipeline dry-run --output dry-run.json
work_item/scripts/python -m photo_pipeline approve-dry-run dry-run.json --approver "$(whoami)"
```
The dry run is strictly read-only: it opens no file for writing, writes no database
row, and reports what it found — file counts by extension, folders, bytes, unreadable
files, excluded directories, and a reconciliation against what the database already
knows (already registered, new, recorded but absent). Set
`PHOTO_PIPELINE_REQUIRE_DRY_RUN_APPROVAL=1` and **every mutating API request is
refused with `403 dry_run_not_approved`** until a report for exactly those library
roots has been approved. Reading stays open — you have to be able to see what was
found in order to approve it — and so does taking a backup. Change the library roots
and the approval no longer applies: it approves that reconciliation, not the idea of
mutating.
## Performance budgets (US07-06)
Budgets are measured, not asserted in prose. `python -m photo_pipeline benchmark`
builds a synthetic library of a stated size, runs the same scenarios every time,
writes a machine-readable report, and **exits non-zero when a budget is breached**.
```bash
work_item/scripts/python -m photo_pipeline benchmark --profile smoke # ~2 s, runs in CI
work_item/scripts/python -m photo_pipeline benchmark --profile short # 25k assets
work_item/scripts/python -m photo_pipeline benchmark --profile full \
--output data/benchmarks/full.json # 25k + 100k
work_item/scripts/python -m photo_pipeline benchmark --profile huge \
--soak-seconds 3600 --output data/benchmarks/soak.json # 500k + soak
```
| Metric | Budget | Why |
|---|---|---|
| `latency_p95_ms` | 250 ms | a list or search page must feel immediate |
| `latency_max_ms` | 2 000 ms | no single page may stall the review flow |
| `rss_growth_bytes` | 400 MB | a run must not leak the library |
| `open_files` | 256 | file descriptors are a hard operating-system limit |
| `wal_bytes` | 200 MB | a growing write-ahead log means checkpoints are starving |
| `queue_depth` | 1 000 | an unbounded queue is an out-of-memory in waiting |
| `cache_over_quota_bytes` | 0 | the thumbnail cache has to respect its quota |
Measured on the reference machine (Apple Silicon, SQLite WAL), p95 per scenario:
| Scenario | 25k | 100k |
|---|---|---|
| `inventory_page` | 0.5 ms | 0.6 ms |
| `library_search` | 4.8 ms | 17.1 ms |
| `library_stats` | 56.8 ms | 197.4 ms |
| `workflow_readiness` | 52.6 ms | 235.4 ms |
| `duplicate_cluster_list` | 0.6 ms | 0.5 ms |
| `duplicate_cluster_page` | 1.7 ms | 1.7 ms |
CI runs the `smoke` profile through `tests/integration/test_performance_budgets.py`;
the 25k/100k/500k matrix and the multi-hour soak belong to scheduled infrastructure,
because minutes of build time do not belong in the suite that runs on every change.
**Exceptions.** A budget that cannot be met is not a warning to ignore: it goes into
`APPROVED_EXCEPTIONS` in `photo_pipeline/services/benchmarks.py` with its raised
limit, who approved it, why, and a review date. Every report lists the exceptions it
applied, so a release review sees them.
Approved today, both for the 500k `huge` profile only, review by 2027-02-17:
| Scenario | Measured at 500k | Raised limit |
|---|---|---|
| `library_stats` | 1.08 s p95 · 3.2 s max | 1.5 s p95 · 4 s max |
| `workflow_readiness` | 1.40 s p95 · 3.3 s max | 1.8 s p95 · 4 s max |
Both are library-wide aggregates — the current safety decision of every asset, and
the album/tag/year breakdown of every analysis row — and both meet the 250 ms budget
at the 100k rows the concept sets it for. Beyond that they are linear against one
SQLite writer; the fix is denormalized totals or the planned PostgreSQL transition,
not a query tweak. Everything else at 500k is inside budget, and a soak at that size
grows neither resident memory nor the job queue.
## Backup and recovery (US07-05)
Backups go through SQLite's online backup API, never a file copy: with WAL enabled
the `.db` file alone is missing every committed page still in the write-ahead log.
Each backup is a directory under `data/backups/` holding the snapshot and a
`manifest.json` describing it — schema revision, SHA-256, row counts, the archive
media the library depends on, and which configuration was set. Secrets are recorded
as `configured`, never as values, so a manifest is safe to attach to a bug report.
```bash
work_item/scripts/python -m photo_pipeline backup --reason before-upgrade --keep 7
work_item/scripts/python -m photo_pipeline verify-backup data/backups/
work_item/scripts/python -m photo_pipeline diagnostics
```
The same is available at `GET /api/v1/diagnostics`, `GET|POST /api/v1/backups`,
`GET /api/v1/backups/{name}/verify`, and `POST /api/v1/backups/prune`. **Restore is
not an endpoint** — it replaces the state of an installation, so it belongs to a
stopped one and a person at a terminal.
### Integrity check
`verify-backup` runs `PRAGMA integrity_check` (structure) *and*
`PRAGMA foreign_key_check` (references), compares the snapshot's SHA-256 with the
manifest, and re-counts every table the manifest recorded. Any mismatch — bit rot, a
truncated copy, a "repaired" snapshot — fails the check, and `restore` refuses a
backup that does not verify.
### Restore drill
1. Stop the server and the worker.
2. `python -m photo_pipeline verify-backup data/backups/` — never restore an
unverified snapshot.
3. `python -m photo_pipeline restore data/backups/ --into /path/to/fresh-data`
(a target that already holds a database is refused; recovering in place means
moving the old data directory aside first).
4. Point `PHOTO_PIPELINE_DATA_DIR` at the restored directory and run
`python -m photo_pipeline migrate`.
5. Run an inventory scan so paths are reconciled against the real library.
6. Mount every archive location named in the manifest before archiving again — the
database records where archived originals are, but it does not contain them.
Practise this against a copy before you need it; the drill is exercised
automatically by `tests/integration/test_backup_recovery.py`.
### Failed migration
A pending schema upgrade is snapshotted first (`reason: pre-migration`), by both the
API startup and `python -m photo_pipeline migrate`. If a migration fails, the error
log names the backup directory: stop everything and run the restore drill against
it. An up-to-date database is not backed up again on every start.
### Archive media
Archived originals live on their medium, not in the backup. The manifest lists every
archive location with its `media_id` and whether it was mounted when the backup was
taken. Keep one copy of each medium off-site, and remount a location before
restoring assets from it.
### Retention and disk
`--keep N` (default 7) prunes the oldest backups and never the newest.
`diagnostics` reports the database, write-ahead log, thumbnail cache, uploader
reports, backups, and logs separately, with free space and warnings for low disk
(`disk_low`, `disk_critical`), a cache over its quota, a write-ahead log outgrowing
its database, and a legacy CLI writing the library.
### Process locking
`serve` and `worker` take a JSON lock in the data directory (`api.lock.json`,
`worker.lock.json`). A second worker exits `2` and names the holder; a lock whose
process is gone is taken over. If the frozen CLI's state files are being written,
both refuse with exit `3` — `--allow-legacy` overrides, and you own the outcome.
## Legacy CLI archive
The command-line tools this application was extracted from are frozen in
`legacy_cli_archive/` (US07-01): the original sources, their docs, the dependency
lock they were last verified against, schema notes, a redacted sample
configuration, the donor ledger, and a checksum for every file.
```bash
cd legacy_cli_archive && shasum -a 256 -c CHECKSUMS.sha256 # verify the archive
work_item/scripts/python -m pytest tests/unit/test_legacy_archive.py -q # lint it
```
They are reference material and rollback evidence only. No module under
`photo_pipeline/` imports or executes them, the archive is not on the application's
import path, and `tests/unit/test_legacy_archive.py` enforces that along with the
checksums and the redaction. Only the two suites that compare *against* the donors —
`tests/characterization/` and `tests/integration/test_safety_parity.py` — put the
archived sources on `sys.path`.
The last path-keyed state they owned, `nsfw_scores.csv`, is imported once and then
left alone:
```bash
work_item/scripts/python -m photo_pipeline import-legacy-scores /path/to/nsfw_scores.csv --dry-run
```
The import writes scored-but-unreviewed `safety_reviews` rows onto stable asset ids,
never invents an asset for an unknown path, never overwrites a human decision, and
writes a reconciliation report to the data directory saying exactly what it did.
`legacy_cli_archive/donor_ledger.yaml` records every migrated behavior with its
target, the tests that pin the donor, the tests that prove the replacement, and each
intentional delta; rows still marked `pending` name the story that will resolve them.