Photo Analyzer
Integrated, restart-safe photo analysis, duplicate review, metadata, upload, and
archive workflow. Planning lives in INTEGRATED_PIPELINE_CONCEPT.md and
delivery_backlog/.
Application (photo_pipeline)
The target application lives in photo_pipeline/ (FastAPI + SQLAlchemy + Alembic).
Run it with:
python -m photo_pipeline migrate # apply database migrations
python -m photo_pipeline serve # start the API + static review UI (127.0.0.1:8000)
Configuration comes from PHOTO_PIPELINE_* environment variables (see
photo_pipeline/config.py); secrets are referenced, never logged.
API access (US07-02)
The app listens on loopback, so its attacker is another page in the same browser.
Every /api/v1 route except health/live, health/ready, and session requires
the application session, and every mutation requires its CSRF token as well:
BASE=http://127.0.0.1:8000
TOKEN=$(curl -sc /tmp/pp.jar $BASE/api/v1/session | python -c 'import json,sys; print(json.load(sys.stdin)["csrf_token"])')
curl -sb /tmp/pp.jar -H "X-CSRF-Token: $TOKEN" -X POST $BASE/api/v1/albums/proposals -d '{}' -H 'Content-Type: application/json'
The session is per server process — restarting serve invalidates it, and the
browser client re-bootstraps by itself. Requests are also refused when the Host is
not a loopback name (DNS rebinding), when Origin is any other origin, when
Sec-Fetch-Site says the request came from another site (an <img> pointed at a
thumbnail), or when the body exceeds PHOTO_PIPELINE_MAX_REQUEST_BYTES. There is no
CORS middleware at all, so no other origin can read a response.
Testing
One offline command runs the whole suite (unit, integration, and browser end-to-end); it needs no network and uses only deterministic synthetic fixtures:
work_item/scripts/python -m pytest tests -q
Browser end-to-end tests require a one-time Playwright browser install:
python -m playwright install chromium
Phase A acceptance gate
Phase A (Epic E01: shared identity, inventory, duplicates, thumbnails, review UI) is gated by a reproducible end-to-end suite:
work_item/scripts/python -m pytest tests/e2e tests/integration -q
tests/e2e/test_phase_a_pipeline.pylaunches the real API process against a fresh database and a deterministic fixture library, then drives scan,_IGNORE/exclusion, move reconciliation, exact/fuzzy duplicate review, thumbnail orientation, canonical selection, browser reload, and a full process restart — asserting durable API and database state after the restart.tests/story_traceability.jsonmaps every delivered story to its tests;tests/e2e/test_traceability.pyfails if a Phase A story loses coverage or a test file is left unexercised.
Phase B acceptance gate
Phase B (Epic E02: durable jobs, workflow shell, safety/analysis views) is proven through the real process and browser boundaries. One command runs the Phase B API, worker-recovery, and browser (Playwright) suites:
work_item/scripts/python -m pytest tests/e2e -m phase_b -q
tests/e2e/test_phase_b_pipeline.pylaunches the real server and durable worker as child processes and drives them only over HTTP/SSE: analysis start/progress, resumable SSE reconnect, the polling fallback, cancellation, per-asset error inspection, one-mutating-job rejection during read-only browsing, the NSFW→vision privacy gate, and durability across a full restart.tests/e2e/test_worker_kill.pykills a worker mid-item and proves a fresh worker resumes the fenced job (the "resume" journey).tests/e2e/test_analysis_browser.pystarts a job from the Analyze view and watches live progress arrive over the browser's real SSE adapter.
The full Phase B regression, including the unchanged Phase A gate, is the whole end-to-end suite:
work_item/scripts/python -m pytest tests/e2e -q
Phase C acceptance gate
Phase C (Epic E03: album evidence, naming policy, versioned proposals, Albums view) is proven through the real process and browser boundaries. One command runs the Phase C API and browser (Playwright) suites with the deterministic naming provider:
work_item/scripts/python -m pytest tests/e2e -m phase_c -q
tests/e2e/test_phase_c_pipeline.pydrives a real server over HTTP: evidence aggregation, generation through the deterministic naming fake (asserting the exact provider inputs and that no file path or asset ID ever reaches it), provider failure and retry, invalid names, path-separator sanitization, editing with optimistic versions, stale-evidence approval refusal, valid approval, and durability across a full restart.tests/e2e/test_albums_ui.pycovers the browser journeys: evidence display, editing, prompt validation, collision guidance, approval, stale conflict, and keyboard operation.- Both suites assert that no fixture path changes — Phase C proposes names and never renames.
The deterministic naming provider is enabled only by test configuration
(PHOTO_PIPELINE_FAKE_NAMING_LOG); without it the application falls back to the
offline naming-policy name. Phase A and B suites remain green in the full run above.
Phase D acceptance gate
Phase D (Epic E04: guarded renaming) is the first phase that changes the library on disk, so its gate is the strictest. One command runs the rename API journeys, the filesystem fault injection, and the browser suite:
work_item/scripts/python -m pytest tests/e2e -m phase_d -q
tests/e2e/test_phase_d_pipeline.pydrives a real server over HTTP: plan and export, confirmation with the plan version and checksum (a stale token is refused without touching disk), a valid apply, the case-only rename procedure, a collision whose occupant survives, a source that changed after planning, and durability across a full restart.- Fault injection is real.
PHOTO_PIPELINE_FAULT_AFTER=<journal state>kills the server process the instant that state is persisted. The suite crashes it at every journal transition in turn (moving,moved,database_updated,verified), starts a fresh process against the same database and library, and requires recovery to converge from journal and disk evidence alone — with the asset set, the stable IDs, and every content hash unchanged. Ambiguous evidence is never guessed: it stays classifiedmanualand keeps blocking. An unresolved rename is the cancellation boundary — there is no cancel once a run starts, and unrelated mutations (album proposal generation and approval) are refused with 409rename_recovery_requireduntil it is resolved, while reads stay available. tests/e2e/test_renames_ui.pycovers the browser journeys: preview of every affected path, confirmation carrying the server-issued token, apply with progress and terminal verification, stale confirmation, collision, interruption, recovery, rollback, keyboard confirmation, and the view still matching the journal after a server restart.
The fault barrier is test-only configuration; without PHOTO_PIPELINE_FAULT_AFTER
the apply path has no crash points. Phases A–C remain green in the full run above.
Phase E acceptance gate
Phase E (Epic E05: Immich upload) is the one stage the application cannot take back, so its gate runs the fake-uploader suites, the black-box upload API journeys, and the browser suite as a single command:
work_item/scripts/python -m pytest -m phase_e -q
tests/integration/test_upload_*.pydrive a real executable standing in forimmich-gothrough the real adapter andsubprocess— argument construction, output bounding, report parsing, verification, and killing a running process.tests/e2e/test_phase_e_pipeline.pydrives a real server and a real durable worker over HTTP: credential failure and an unreachable server, preflight blockers and the explicitly approved partial scope, a new album, an exact duplicate, an upgrade, a retryable failure and its successful retry, a lost acceptance response, verification against Immich, an inconclusive answer resolved by an operator with evidence, bytes edited after upload, cancellation and resume, and an interrupted attempt recovered across a restart.- EXIF precedes upload is asserted, not assumed: an album without its verified safety and analysis checkpoints cannot be approved, and the uploader's own argv log proves it was never executed. Each finished upload re-hashes the files in the folder the uploader was handed and requires the persisted SHA-256/SHA-1 to match.
- No secret is retained. The API key is a sentinel string; after a full upload and verification it must appear in the uploader's argv and nowhere else — not in the database, the retained report, or any response the browser can read.
tests/e2e/test_uploads_ui.pycovers the browser journeys (preflight preview, confirmation, progress, stopping a run, verification, manual resolution, stale bytes, and recovery after a restart).
Phases A–D remain green in the full run above.
Phase F acceptance gate
Phase F (Epic E06: archive lifecycle) is the only stage that removes originals from the library, and the only one whose storage can walk away in someone's bag. One command runs the archive fault-injection suites, the black-box archive and restore API journeys, and the browser suite:
work_item/scripts/python -m pytest -m phase_f -q
tests/integration/test_archive_*.pyandtests/integration/test_restore.pydrive real files on real filesystems: preflight against a mounted, missing, swapped, read-only, or full medium; copy-verify-remove and the same-filesystem move path; and a crash at every persisted journal transition in both transfer modes, asserting that no source is ever removed without a durable, byte-identical archive copy.tests/e2e/test_phase_f_pipeline.pydrives a real server and a real durable worker over HTTP: preflight blockers (offline medium, wrong volume, insufficient capacity, bytes changed after upload), a verified archive whose manifest, hashes, and path history are checked on the medium itself, a worker killed at each oftransferring,verified,removing,source_removed, andcomplete, the evidence-based recovery that follows, offline deduplication of an exact and a fuzzy copy while the medium is away, mount return, restore, and a collision that restores beside its occupant.- Archived is not missing. An unmounted medium leaves its photos
archived_offline— still hashed, still in the duplicate indexes, still previewable through their protected thumbnails — and a rescan neither prunes nor flags them. - Ambiguity is never guessed. A journal state the medium contradicts stays
manual, offers no automatic action, and keeps blocking further archiving until a human decides. tests/e2e/test_archive_ui.pycovers the browser journeys (preview with destination identity and reclaimable bytes, blockers and mount instructions, progress split into transfer/verification/removal, interruption and recovery, offline browsing, restore, collision, keyboard confirmation, and reload).
Phases A–E remain green in the full run above.
Media and metadata hardening (US07-03)
Every pixel the application reads goes through photo_pipeline/imaging.py: the
declared dimensions are checked before anything is decoded, Pillow's
decompression-bomb warning is treated as a refusal, JPEG decodes near the requested
size, and each decoder failure becomes one of two typed errors. A damaged file is a
per-item error with a persisted code, never a failed scan or a dead worker.
Every metadata stage ends with an EXIF checkpoint (services/exif_checkpoint.py):
snapshot, write the owned keywords, read back, prove the owned fields landed and that
nothing else moved, refresh the file hash. A field the stage does not own that
changed anyway makes the checkpoint divergent — recorded in exif_projections,
shown in the review queue, never repaired behind the user's back, and not counted as
verified, so upload stays blocked.
The golden corpus that proves all of it is generated, not committed:
tests/fixtures/media_corpus.py declares every format, orientation, profile,
damage, and metadata case with its expected outcome, and the suite regenerates it
twice to prove it does not drift.
work_item/scripts/python -m pytest tests/integration/test_media_hardening.py tests/integration/test_exif_checkpoints.py -q
Concurrency and crash recovery (US07-04)
Crash safety is proven by crashing. photo_pipeline/faults.py defines the control
points — the persisted transitions of the rename, archive, EXIF, upload, and job
lanes — and arms one only when PHOTO_PIPELINE_FAULT_AFTER names it, at which
point the process dies the way a SIGKILL does. There is no endpoint and no
configuration field that can reach a barrier; a deployment that never sets the
variable can never hit one.
The race suite runs each scenario several times with a seed recorded on the test
result (race_seed) and asserts invariants rather than schedules: work is never
claimed or executed twice, a stale fencing token never commits, no file body is
lost or overwritten, and the database still passes PRAGMA integrity_check.
work_item/scripts/python -m pytest tests/integration/test_concurrency_races.py \
tests/integration/test_fault_matrix.py tests/e2e/test_crash_recovery.py -q
# replay a failure, or soak for new interleavings
PHOTO_PIPELINE_RACE_SEED=1234 PHOTO_PIPELINE_RACE_REPEATS=50 \
work_item/scripts/python -m pytest tests/integration/test_concurrency_races.py -q
Any failing test keeps its evidence: the temporary database (with its write-ahead
log), the journals, the logs, the recorded seed, and a SHA-256 manifest of every
file in the temporary library are copied to .artifacts/<test id>/ before pytest
deletes the directory. Point PHOTO_PIPELINE_TEST_ARTIFACTS elsewhere to collect
them from CI.
Release gate (US07-07)
One command runs every suite in an isolated stack and keeps the evidence:
work_item/scripts/python -m photo_pipeline release-gate --output data/release/$(date -u +%Y%m%dT%H%M%SZ)
It fails — and exits non-zero — when any stage fails, when a suite skips a test for
a reason that is not a documented environment limit (exiftool not installed,
root ignores directory permissions), or when the story matrix has a hole. The
evidence directory holds release-report.json (revision, per-stage result, timings,
summaries), logs/<stage>.log, and CHECKSUMS.sha256 over both.
The story matrix lives in tests/story_traceability.json: every story under
delivery_backlog/stories/ is either mapped to test files that exist, or listed in
planned as an accepted but unimplemented story. A story that is neither, or a
mapping to a file that has been deleted, fails the gate.
The journey (tests/e2e/test_release_journey.py) takes one fresh library through
discovery, duplicate review, safety, analysis, EXIF verification, album proposal,
guarded rename, rescan, upload with server-side verification, archive, offline
deduplication, and restore — over HTTP against real server and worker processes,
with a full restart in the middle and at the end.
Real-library dry run and approval
Before the application is pointed at photos that cannot be replaced:
work_item/scripts/python -m photo_pipeline dry-run --output dry-run.json
work_item/scripts/python -m photo_pipeline approve-dry-run dry-run.json --approver "$(whoami)"
The dry run is strictly read-only: it opens no file for writing, writes no database
row, and reports what it found — file counts by extension, folders, bytes, unreadable
files, excluded directories, and a reconciliation against what the database already
knows (already registered, new, recorded but absent). Set
PHOTO_PIPELINE_REQUIRE_DRY_RUN_APPROVAL=1 and every mutating API request is
refused with 403 dry_run_not_approved until a report for exactly those library
roots has been approved. Reading stays open — you have to be able to see what was
found in order to approve it — and so does taking a backup. Change the library roots
and the approval no longer applies: it approves that reconciliation, not the idea of
mutating.
Performance budgets (US07-06)
Budgets are measured, not asserted in prose. python -m photo_pipeline benchmark
builds a synthetic library of a stated size, runs the same scenarios every time,
writes a machine-readable report, and exits non-zero when a budget is breached.
work_item/scripts/python -m photo_pipeline benchmark --profile smoke # ~2 s, runs in CI
work_item/scripts/python -m photo_pipeline benchmark --profile short # 25k assets
work_item/scripts/python -m photo_pipeline benchmark --profile full \
--output data/benchmarks/full.json # 25k + 100k
work_item/scripts/python -m photo_pipeline benchmark --profile huge \
--soak-seconds 3600 --output data/benchmarks/soak.json # 500k + soak
| Metric | Budget | Why |
|---|---|---|
latency_p95_ms |
250 ms | a list or search page must feel immediate |
latency_max_ms |
2 000 ms | no single page may stall the review flow |
rss_growth_bytes |
400 MB | a run must not leak the library |
open_files |
256 | file descriptors are a hard operating-system limit |
wal_bytes |
200 MB | a growing write-ahead log means checkpoints are starving |
queue_depth |
1 000 | an unbounded queue is an out-of-memory in waiting |
cache_over_quota_bytes |
0 | the thumbnail cache has to respect its quota |
Measured on the reference machine (Apple Silicon, SQLite WAL), p95 per scenario:
| Scenario | 25k | 100k |
|---|---|---|
inventory_page |
0.5 ms | 0.6 ms |
library_search |
4.8 ms | 17.1 ms |
library_stats |
56.8 ms | 197.4 ms |
workflow_readiness |
52.6 ms | 235.4 ms |
duplicate_cluster_list |
0.6 ms | 0.5 ms |
duplicate_cluster_page |
1.7 ms | 1.7 ms |
CI runs the smoke profile through tests/integration/test_performance_budgets.py;
the 25k/100k/500k matrix and the multi-hour soak belong to scheduled infrastructure,
because minutes of build time do not belong in the suite that runs on every change.
Exceptions. A budget that cannot be met is not a warning to ignore: it goes into
APPROVED_EXCEPTIONS in photo_pipeline/services/benchmarks.py with its raised
limit, who approved it, why, and a review date. Every report lists the exceptions it
applied, so a release review sees them.
Approved today, both for the 500k huge profile only, review by 2027-02-17:
| Scenario | Measured at 500k | Raised limit |
|---|---|---|
library_stats |
1.08 s p95 · 3.2 s max | 1.5 s p95 · 4 s max |
workflow_readiness |
1.40 s p95 · 3.3 s max | 1.8 s p95 · 4 s max |
Both are library-wide aggregates — the current safety decision of every asset, and the album/tag/year breakdown of every analysis row — and both meet the 250 ms budget at the 100k rows the concept sets it for. Beyond that they are linear against one SQLite writer; the fix is denormalized totals or the planned PostgreSQL transition, not a query tweak. Everything else at 500k is inside budget, and a soak at that size grows neither resident memory nor the job queue.
Backup and recovery (US07-05)
Backups go through SQLite's online backup API, never a file copy: with WAL enabled
the .db file alone is missing every committed page still in the write-ahead log.
Each backup is a directory under data/backups/ holding the snapshot and a
manifest.json describing it — schema revision, SHA-256, row counts, the archive
media the library depends on, and which configuration was set. Secrets are recorded
as configured, never as values, so a manifest is safe to attach to a bug report.
work_item/scripts/python -m photo_pipeline backup --reason before-upgrade --keep 7
work_item/scripts/python -m photo_pipeline verify-backup data/backups/<name>
work_item/scripts/python -m photo_pipeline diagnostics
The same is available at GET /api/v1/diagnostics, GET|POST /api/v1/backups,
GET /api/v1/backups/{name}/verify, and POST /api/v1/backups/prune. Restore is
not an endpoint — it replaces the state of an installation, so it belongs to a
stopped one and a person at a terminal.
Integrity check
verify-backup runs PRAGMA integrity_check (structure) and
PRAGMA foreign_key_check (references), compares the snapshot's SHA-256 with the
manifest, and re-counts every table the manifest recorded. Any mismatch — bit rot, a
truncated copy, a "repaired" snapshot — fails the check, and restore refuses a
backup that does not verify.
Restore drill
- Stop the server and the worker.
python -m photo_pipeline verify-backup data/backups/<name>— never restore an unverified snapshot.python -m photo_pipeline restore data/backups/<name> --into /path/to/fresh-data(a target that already holds a database is refused; recovering in place means moving the old data directory aside first).- Point
PHOTO_PIPELINE_DATA_DIRat the restored directory and runpython -m photo_pipeline migrate. - Run an inventory scan so paths are reconciled against the real library.
- Mount every archive location named in the manifest before archiving again — the database records where archived originals are, but it does not contain them.
Practise this against a copy before you need it; the drill is exercised
automatically by tests/integration/test_backup_recovery.py.
Failed migration
A pending schema upgrade is snapshotted first (reason: pre-migration), by both the
API startup and python -m photo_pipeline migrate. If a migration fails, the error
log names the backup directory: stop everything and run the restore drill against
it. An up-to-date database is not backed up again on every start.
Archive media
Archived originals live on their medium, not in the backup. The manifest lists every
archive location with its media_id and whether it was mounted when the backup was
taken. Keep one copy of each medium off-site, and remount a location before
restoring assets from it.
Retention and disk
--keep N (default 7) prunes the oldest backups and never the newest.
diagnostics reports the database, write-ahead log, thumbnail cache, uploader
reports, backups, and logs separately, with free space and warnings for low disk
(disk_low, disk_critical), a cache over its quota, a write-ahead log outgrowing
its database, and a legacy CLI writing the library.
Process locking
serve and worker take a JSON lock in the data directory (api.lock.json,
worker.lock.json). A second worker exits 2 and names the holder; a lock whose
process is gone is taken over. If the frozen CLI's state files are being written,
both refuse with exit 3 — --allow-legacy overrides, and you own the outcome.
Legacy CLI archive
The command-line tools this application was extracted from are frozen in
legacy_cli_archive/ (US07-01): the original sources, their docs, the dependency
lock they were last verified against, schema notes, a redacted sample
configuration, the donor ledger, and a checksum for every file.
cd legacy_cli_archive && shasum -a 256 -c CHECKSUMS.sha256 # verify the archive
work_item/scripts/python -m pytest tests/unit/test_legacy_archive.py -q # lint it
They are reference material and rollback evidence only. No module under
photo_pipeline/ imports or executes them, the archive is not on the application's
import path, and tests/unit/test_legacy_archive.py enforces that along with the
checksums and the redaction. Only the two suites that compare against the donors —
tests/characterization/ and tests/integration/test_safety_parity.py — put the
archived sources on sys.path.
The last path-keyed state they owned, nsfw_scores.csv, is imported once and then
left alone:
work_item/scripts/python -m photo_pipeline import-legacy-scores /path/to/nsfw_scores.csv --dry-run
The import writes scored-but-unreviewed safety_reviews rows onto stable asset ids,
never invents an asset for an unknown path, never overwrites a human decision, and
writes a reconciliation report to the data directory saying exactly what it did.
legacy_cli_archive/donor_ledger.yaml records every migrated behavior with its
target, the tests that pin the donor, the tests that prove the replacement, and each
intentional delta; rows still marked pending name the story that will resolve them.