Photo Analyzer

Integrated, restart-safe photo analysis, duplicate review, metadata, upload, and archive workflow. Planning lives in INTEGRATED_PIPELINE_CONCEPT.md and delivery_backlog/.

Application (photo_pipeline)

The target application lives in photo_pipeline/ (FastAPI + SQLAlchemy + Alembic). Install it into a virtualenv once:

python3.12 -m venv .venv
.venv/bin/pip install -e ".[vision]"   # drop [vision] for a review-only install

Then run the two processes:

.venv/bin/python -m photo_pipeline migrate   # apply database migrations
.venv/bin/python -m photo_pipeline serve     # API + review UI at 127.0.0.1:8000/app/
.venv/bin/python -m photo_pipeline worker    # second terminal: runs the jobs

The server enqueues work and serves the UI; nothing actually scans, scores, analyses, uploads, or archives without a worker. work_item/scripts/python is the helper's launcher — it prefers Conda base and falls back to a bare system interpreter, so it is not how the application is run.

Configuration comes from PHOTO_PIPELINE_* environment variables (see photo_pipeline/config.py); secrets are referenced, never logged.

Configuration file

.env in the working directory is read at startup, or any path named by PHOTO_PIPELINE_ENV_FILE. It is parsed, never executed: KEY=value lines, # comments, optional quotes — no interpolation and no export. Anything already exported wins, so the file is the standing configuration and the shell is the override for one run.

The archived CLI's variable names still work, so an existing photo_analyzer.env can be used as-is:

in the file applied as
LLM_API_KEY / GEMINI_API_KEY OPENAI_API_KEY
LLM_BASE_URL OPENAI_BASE_URL
LIBRARY PHOTO_PIPELINE_LIBRARY_ROOTS

.env and *.env are gitignored and denied by the work-item safety checks: the file holds a real key and must never be committed.

API access (US07-02)

The app listens on loopback, so its attacker is another page in the same browser. Every /api/v1 route except health/live, health/ready, and session requires the application session, and every mutation requires its CSRF token as well:

BASE=http://127.0.0.1:8000
TOKEN=$(curl -sc /tmp/pp.jar $BASE/api/v1/session | python -c 'import json,sys; print(json.load(sys.stdin)["csrf_token"])')
curl -sb /tmp/pp.jar -H "X-CSRF-Token: $TOKEN" -X POST $BASE/api/v1/albums/proposals -d '{}' -H 'Content-Type: application/json'

The session is per server process — restarting serve invalidates it, and the browser client re-bootstraps by itself. Requests are also refused when the Host is not a loopback name (DNS rebinding), when Origin is any other origin, when Sec-Fetch-Site says the request came from another site (an <img> pointed at a thumbnail), or when the body exceeds PHOTO_PIPELINE_MAX_REQUEST_BYTES. There is no CORS middleware at all, so no other origin can read a response.

Reaching it through a hostname or proxy (US08-01)

variable meaning
PHOTO_PIPELINE_ALLOWED_HOSTS comma-separated extra names the app answers to; empty means loopback only
PHOTO_PIPELINE_ACCESS_SECRET traded for the session cookie at GET /api/v1/session via X-Access-Secret
PHOTO_PIPELINE_TRUSTED_PROXIES comma-separated peer addresses whose X-Forwarded-Proto/X-Forwarded-Host are believed

Being reachable was the authentication: whoever could open 127.0.0.1:8000 owned the library. So naming any non-loopback host — or binding to one, 0.0.0.0 included — makes the access secret mandatory, and serve refuses to start without it rather than publishing the library. Loopback-only deployments need no secret and behave exactly as before.

curl -sc /tmp/pp.jar -H "X-Access-Secret: $PHOTO_PIPELINE_ACCESS_SECRET" \
  https://photos.example.com/api/v1/session

The browser asks for the secret once per tab and keeps it in sessionStorage. Wrong secrets are rate-limited (5 per minute) and logged with the caller's address only. Host and Origin are judged against the configured names; the external scheme and host come from the forwarded headers only when the request arrived from a PHOTO_PIPELINE_TRUSTED_PROXIES address, so a client cannot declare its own origin, and the session cookie is marked Secure when that external scheme is HTTPS. Health endpoints stay reachable without the secret so an orchestrator can restart the container; nothing else does.

Testing

One offline command runs the whole suite (unit, integration, and browser end-to-end); it needs no network and uses only deterministic synthetic fixtures:

work_item/scripts/python -m pytest tests -q

Browser end-to-end tests require a one-time Playwright browser install:

python -m playwright install chromium

Phase A acceptance gate

Phase A (Epic E01: shared identity, inventory, duplicates, thumbnails, review UI) is gated by a reproducible end-to-end suite:

work_item/scripts/python -m pytest tests/e2e tests/integration -q
  • tests/e2e/test_phase_a_pipeline.py launches the real API process against a fresh database and a deterministic fixture library, then drives scan, _IGNORE/ exclusion, move reconciliation, exact/fuzzy duplicate review, thumbnail orientation, canonical selection, browser reload, and a full process restart — asserting durable API and database state after the restart.
  • tests/story_traceability.json maps every delivered story to its tests; tests/e2e/test_traceability.py fails if a Phase A story loses coverage or a test file is left unexercised.

Phase B acceptance gate

Phase B (Epic E02: durable jobs, workflow shell, safety/analysis views) is proven through the real process and browser boundaries. One command runs the Phase B API, worker-recovery, and browser (Playwright) suites:

work_item/scripts/python -m pytest tests/e2e -m phase_b -q
  • tests/e2e/test_phase_b_pipeline.py launches the real server and durable worker as child processes and drives them only over HTTP/SSE: analysis start/progress, resumable SSE reconnect, the polling fallback, cancellation, per-asset error inspection, one-mutating-job rejection during read-only browsing, the NSFW→vision privacy gate, and durability across a full restart.
  • tests/e2e/test_worker_kill.py kills a worker mid-item and proves a fresh worker resumes the fenced job (the "resume" journey).
  • tests/e2e/test_analysis_browser.py starts a job from the Analyze view and watches live progress arrive over the browser's real SSE adapter.

The full Phase B regression, including the unchanged Phase A gate, is the whole end-to-end suite:

work_item/scripts/python -m pytest tests/e2e -q

Phase C acceptance gate

Phase C (Epic E03: album evidence, naming policy, versioned proposals, Albums view) is proven through the real process and browser boundaries. One command runs the Phase C API and browser (Playwright) suites with the deterministic naming provider:

work_item/scripts/python -m pytest tests/e2e -m phase_c -q
  • tests/e2e/test_phase_c_pipeline.py drives a real server over HTTP: evidence aggregation, generation through the deterministic naming fake (asserting the exact provider inputs and that no file path or asset ID ever reaches it), provider failure and retry, invalid names, path-separator sanitization, editing with optimistic versions, stale-evidence approval refusal, valid approval, and durability across a full restart.
  • tests/e2e/test_albums_ui.py covers the browser journeys: evidence display, editing, prompt validation, collision guidance, approval, stale conflict, and keyboard operation.
  • Both suites assert that no fixture path changes — Phase C proposes names and never renames.

The deterministic naming provider is enabled only by test configuration (PHOTO_PIPELINE_FAKE_NAMING_LOG); without it the application falls back to the offline naming-policy name. Phase A and B suites remain green in the full run above.

Phase D acceptance gate

Phase D (Epic E04: guarded renaming) is the first phase that changes the library on disk, so its gate is the strictest. One command runs the rename API journeys, the filesystem fault injection, and the browser suite:

work_item/scripts/python -m pytest tests/e2e -m phase_d -q
  • tests/e2e/test_phase_d_pipeline.py drives a real server over HTTP: plan and export, confirmation with the plan version and checksum (a stale token is refused without touching disk), a valid apply, the case-only rename procedure, a collision whose occupant survives, a source that changed after planning, and durability across a full restart.
  • Fault injection is real. PHOTO_PIPELINE_FAULT_AFTER=<journal state> kills the server process the instant that state is persisted. The suite crashes it at every journal transition in turn (moving, moved, database_updated, verified), starts a fresh process against the same database and library, and requires recovery to converge from journal and disk evidence alone — with the asset set, the stable IDs, and every content hash unchanged. Ambiguous evidence is never guessed: it stays classified manual and keeps blocking. An unresolved rename is the cancellation boundary — there is no cancel once a run starts, and unrelated mutations (album proposal generation and approval) are refused with 409 rename_recovery_required until it is resolved, while reads stay available.
  • tests/e2e/test_renames_ui.py covers the browser journeys: preview of every affected path, confirmation carrying the server-issued token, apply with progress and terminal verification, stale confirmation, collision, interruption, recovery, rollback, keyboard confirmation, and the view still matching the journal after a server restart.

The fault barrier is test-only configuration; without PHOTO_PIPELINE_FAULT_AFTER the apply path has no crash points. Phases AC remain green in the full run above.

Phase E acceptance gate

Phase E (Epic E05: Immich upload) is the one stage the application cannot take back, so its gate runs the fake-uploader suites, the black-box upload API journeys, and the browser suite as a single command:

work_item/scripts/python -m pytest -m phase_e -q
  • tests/integration/test_upload_*.py drive a real executable standing in for immich-go through the real adapter and subprocess — argument construction, output bounding, report parsing, verification, and killing a running process.
  • tests/e2e/test_phase_e_pipeline.py drives a real server and a real durable worker over HTTP: credential failure and an unreachable server, preflight blockers and the explicitly approved partial scope, a new album, an exact duplicate, an upgrade, a retryable failure and its successful retry, a lost acceptance response, verification against Immich, an inconclusive answer resolved by an operator with evidence, bytes edited after upload, cancellation and resume, and an interrupted attempt recovered across a restart.
  • EXIF precedes upload is asserted, not assumed: an album without its verified safety and analysis checkpoints cannot be approved, and the uploader's own argv log proves it was never executed. Each finished upload re-hashes the files in the folder the uploader was handed and requires the persisted SHA-256/SHA-1 to match.
  • No secret is retained. The API key is a sentinel string; after a full upload and verification it must appear in the uploader's argv and nowhere else — not in the database, the retained report, or any response the browser can read.
  • tests/e2e/test_uploads_ui.py covers the browser journeys (preflight preview, confirmation, progress, stopping a run, verification, manual resolution, stale bytes, and recovery after a restart).

Phases AD remain green in the full run above.

Phase F acceptance gate

Phase F (Epic E06: archive lifecycle) is the only stage that removes originals from the library, and the only one whose storage can walk away in someone's bag. One command runs the archive fault-injection suites, the black-box archive and restore API journeys, and the browser suite:

work_item/scripts/python -m pytest -m phase_f -q
  • tests/integration/test_archive_*.py and tests/integration/test_restore.py drive real files on real filesystems: preflight against a mounted, missing, swapped, read-only, or full medium; copy-verify-remove and the same-filesystem move path; and a crash at every persisted journal transition in both transfer modes, asserting that no source is ever removed without a durable, byte-identical archive copy.
  • tests/e2e/test_phase_f_pipeline.py drives a real server and a real durable worker over HTTP: preflight blockers (offline medium, wrong volume, insufficient capacity, bytes changed after upload), a verified archive whose manifest, hashes, and path history are checked on the medium itself, a worker killed at each of transferring, verified, removing, source_removed, and complete, the evidence-based recovery that follows, offline deduplication of an exact and a fuzzy copy while the medium is away, mount return, restore, and a collision that restores beside its occupant.
  • Archived is not missing. An unmounted medium leaves its photos archived_offline — still hashed, still in the duplicate indexes, still previewable through their protected thumbnails — and a rescan neither prunes nor flags them.
  • Ambiguity is never guessed. A journal state the medium contradicts stays manual, offers no automatic action, and keeps blocking further archiving until a human decides.
  • tests/e2e/test_archive_ui.py covers the browser journeys (preview with destination identity and reclaimable bytes, blockers and mount instructions, progress split into transfer/verification/removal, interruption and recovery, offline browsing, restore, collision, keyboard confirmation, and reload).

Phases AE remain green in the full run above.

Media and metadata hardening (US07-03)

Every pixel the application reads goes through photo_pipeline/imaging.py: the declared dimensions are checked before anything is decoded, Pillow's decompression-bomb warning is treated as a refusal, JPEG decodes near the requested size, and each decoder failure becomes one of two typed errors. A damaged file is a per-item error with a persisted code, never a failed scan or a dead worker.

Every metadata stage ends with an EXIF checkpoint (services/exif_checkpoint.py): snapshot, write the owned keywords, read back, prove the owned fields landed and that nothing else moved, refresh the file hash. A field the stage does not own that changed anyway makes the checkpoint divergent — recorded in exif_projections, shown in the review queue, never repaired behind the user's back, and not counted as verified, so upload stays blocked.

The golden corpus that proves all of it is generated, not committed: tests/fixtures/media_corpus.py declares every format, orientation, profile, damage, and metadata case with its expected outcome, and the suite regenerates it twice to prove it does not drift.

work_item/scripts/python -m pytest tests/integration/test_media_hardening.py tests/integration/test_exif_checkpoints.py -q

Concurrency and crash recovery (US07-04)

Crash safety is proven by crashing. photo_pipeline/faults.py defines the control points — the persisted transitions of the rename, archive, EXIF, upload, and job lanes — and arms one only when PHOTO_PIPELINE_FAULT_AFTER names it, at which point the process dies the way a SIGKILL does. There is no endpoint and no configuration field that can reach a barrier; a deployment that never sets the variable can never hit one.

The race suite runs each scenario several times with a seed recorded on the test result (race_seed) and asserts invariants rather than schedules: work is never claimed or executed twice, a stale fencing token never commits, no file body is lost or overwritten, and the database still passes PRAGMA integrity_check.

work_item/scripts/python -m pytest tests/integration/test_concurrency_races.py \
    tests/integration/test_fault_matrix.py tests/e2e/test_crash_recovery.py -q

# replay a failure, or soak for new interleavings
PHOTO_PIPELINE_RACE_SEED=1234 PHOTO_PIPELINE_RACE_REPEATS=50 \
    work_item/scripts/python -m pytest tests/integration/test_concurrency_races.py -q

Any failing test keeps its evidence: the temporary database (with its write-ahead log), the journals, the logs, the recorded seed, and a SHA-256 manifest of every file in the temporary library are copied to .artifacts/<test id>/ before pytest deletes the directory. Point PHOTO_PIPELINE_TEST_ARTIFACTS elsewhere to collect them from CI.

Release gate (US07-07)

One command runs every suite in an isolated stack and keeps the evidence:

work_item/scripts/python -m photo_pipeline release-gate --output data/release/$(date -u +%Y%m%dT%H%M%SZ)

It fails — and exits non-zero — when any stage fails, when a suite skips a test for a reason that is not a documented environment limit (exiftool not installed, root ignores directory permissions), or when the story matrix has a hole. The evidence directory holds release-report.json (revision, per-stage result, timings, summaries), logs/<stage>.log, and CHECKSUMS.sha256 over both.

The story matrix lives in tests/story_traceability.json: every story under delivery_backlog/stories/ is either mapped to test files that exist, or listed in planned as an accepted but unimplemented story. A story that is neither, or a mapping to a file that has been deleted, fails the gate.

The journey (tests/e2e/test_release_journey.py) takes one fresh library through discovery, duplicate review, safety, analysis, EXIF verification, album proposal, guarded rename, rescan, upload with server-side verification, archive, offline deduplication, and restore — over HTTP against real server and worker processes, with a full restart in the middle and at the end.

Real-library dry run and approval

Before the application is pointed at photos that cannot be replaced:

work_item/scripts/python -m photo_pipeline dry-run --output dry-run.json
work_item/scripts/python -m photo_pipeline approve-dry-run dry-run.json --approver "$(whoami)"

The dry run is strictly read-only: it opens no file for writing, writes no database row, and reports what it found — file counts by extension, folders, bytes, unreadable files, excluded directories, and a reconciliation against what the database already knows (already registered, new, recorded but absent). Set PHOTO_PIPELINE_REQUIRE_DRY_RUN_APPROVAL=1 and every mutating API request is refused with 403 dry_run_not_approved until a report for exactly those library roots has been approved. Reading stays open — you have to be able to see what was found in order to approve it — and so does taking a backup. Change the library roots and the approval no longer applies: it approves that reconciliation, not the idea of mutating.

Performance budgets (US07-06)

Budgets are measured, not asserted in prose. python -m photo_pipeline benchmark builds a synthetic library of a stated size, runs the same scenarios every time, writes a machine-readable report, and exits non-zero when a budget is breached.

work_item/scripts/python -m photo_pipeline benchmark --profile smoke    # ~2 s, runs in CI
work_item/scripts/python -m photo_pipeline benchmark --profile short    # 25k assets
work_item/scripts/python -m photo_pipeline benchmark --profile full \
    --output data/benchmarks/full.json                                  # 25k + 100k
work_item/scripts/python -m photo_pipeline benchmark --profile huge \
    --soak-seconds 3600 --output data/benchmarks/soak.json              # 500k + soak
Metric Budget Why
latency_p95_ms 250 ms a list or search page must feel immediate
latency_max_ms 2 000 ms no single page may stall the review flow
rss_growth_bytes 400 MB a run must not leak the library
open_files 256 file descriptors are a hard operating-system limit
wal_bytes 200 MB a growing write-ahead log means checkpoints are starving
queue_depth 1 000 an unbounded queue is an out-of-memory in waiting
cache_over_quota_bytes 0 the thumbnail cache has to respect its quota

Measured on the reference machine (Apple Silicon, SQLite WAL), p95 per scenario:

Scenario 25k 100k
inventory_page 0.5 ms 0.6 ms
library_search 4.8 ms 17.1 ms
library_stats 56.8 ms 197.4 ms
workflow_readiness 52.6 ms 235.4 ms
duplicate_cluster_list 0.6 ms 0.5 ms
duplicate_cluster_page 1.7 ms 1.7 ms

CI runs the smoke profile through tests/integration/test_performance_budgets.py; the 25k/100k/500k matrix and the multi-hour soak belong to scheduled infrastructure, because minutes of build time do not belong in the suite that runs on every change.

Exceptions. A budget that cannot be met is not a warning to ignore: it goes into APPROVED_EXCEPTIONS in photo_pipeline/services/benchmarks.py with its raised limit, who approved it, why, and a review date. Every report lists the exceptions it applied, so a release review sees them.

Approved today, both for the 500k huge profile only, review by 2027-02-17:

Scenario Measured at 500k Raised limit
library_stats 1.08 s p95 · 3.2 s max 1.5 s p95 · 4 s max
workflow_readiness 1.40 s p95 · 3.3 s max 1.8 s p95 · 4 s max

Both are library-wide aggregates — the current safety decision of every asset, and the album/tag/year breakdown of every analysis row — and both meet the 250 ms budget at the 100k rows the concept sets it for. Beyond that they are linear against one SQLite writer; the fix is denormalized totals or the planned PostgreSQL transition, not a query tweak. Everything else at 500k is inside budget, and a soak at that size grows neither resident memory nor the job queue.

Backup and recovery (US07-05)

Backups go through SQLite's online backup API, never a file copy: with WAL enabled the .db file alone is missing every committed page still in the write-ahead log. Each backup is a directory under data/backups/ holding the snapshot and a manifest.json describing it — schema revision, SHA-256, row counts, the archive media the library depends on, and which configuration was set. Secrets are recorded as configured, never as values, so a manifest is safe to attach to a bug report.

work_item/scripts/python -m photo_pipeline backup --reason before-upgrade --keep 7
work_item/scripts/python -m photo_pipeline verify-backup data/backups/<name>
work_item/scripts/python -m photo_pipeline diagnostics

The same is available at GET /api/v1/diagnostics, GET|POST /api/v1/backups, GET /api/v1/backups/{name}/verify, and POST /api/v1/backups/prune. Restore is not an endpoint — it replaces the state of an installation, so it belongs to a stopped one and a person at a terminal.

Integrity check

verify-backup runs PRAGMA integrity_check (structure) and PRAGMA foreign_key_check (references), compares the snapshot's SHA-256 with the manifest, and re-counts every table the manifest recorded. Any mismatch — bit rot, a truncated copy, a "repaired" snapshot — fails the check, and restore refuses a backup that does not verify.

Restore drill

  1. Stop the server and the worker.
  2. python -m photo_pipeline verify-backup data/backups/<name> — never restore an unverified snapshot.
  3. python -m photo_pipeline restore data/backups/<name> --into /path/to/fresh-data (a target that already holds a database is refused; recovering in place means moving the old data directory aside first).
  4. Point PHOTO_PIPELINE_DATA_DIR at the restored directory and run python -m photo_pipeline migrate.
  5. Run an inventory scan so paths are reconciled against the real library.
  6. Mount every archive location named in the manifest before archiving again — the database records where archived originals are, but it does not contain them.

Practise this against a copy before you need it; the drill is exercised automatically by tests/integration/test_backup_recovery.py.

Failed migration

A pending schema upgrade is snapshotted first (reason: pre-migration), by both the API startup and python -m photo_pipeline migrate. If a migration fails, the error log names the backup directory: stop everything and run the restore drill against it. An up-to-date database is not backed up again on every start.

Archive media

Archived originals live on their medium, not in the backup. The manifest lists every archive location with its media_id and whether it was mounted when the backup was taken. Keep one copy of each medium off-site, and remount a location before restoring assets from it.

Retention and disk

--keep N (default 7) prunes the oldest backups and never the newest. diagnostics reports the database, write-ahead log, thumbnail cache, uploader reports, backups, and logs separately, with free space and warnings for low disk (disk_low, disk_critical), a cache over its quota, a write-ahead log outgrowing its database, and a legacy CLI writing the library.

Process locking

serve and worker take a JSON lock in the data directory (api.lock.json, worker.lock.json). A second worker exits 2 and names the holder; a lock whose process is gone is taken over. If the frozen CLI's state files are being written, both refuse with exit 3--allow-legacy overrides, and you own the outcome.

Legacy CLI archive

The command-line tools this application was extracted from are frozen in legacy_cli_archive/ (US07-01): the original sources, their docs, the dependency lock they were last verified against, schema notes, a redacted sample configuration, the donor ledger, and a checksum for every file.

cd legacy_cli_archive && shasum -a 256 -c CHECKSUMS.sha256   # verify the archive
work_item/scripts/python -m pytest tests/unit/test_legacy_archive.py -q   # lint it

They are reference material and rollback evidence only. No module under photo_pipeline/ imports or executes them, the archive is not on the application's import path, and tests/unit/test_legacy_archive.py enforces that along with the checksums and the redaction. Only the two suites that compare against the donors — tests/characterization/ and tests/integration/test_safety_parity.py — put the archived sources on sys.path.

The last path-keyed state they owned, nsfw_scores.csv, is imported once and then left alone:

work_item/scripts/python -m photo_pipeline import-legacy-scores /path/to/nsfw_scores.csv --dry-run

The import writes scored-but-unreviewed safety_reviews rows onto stable asset ids, never invents an asset for an unknown path, never overwrites a human decision, and writes a reconciliation report to the data directory saying exactly what it did. legacy_cli_archive/donor_ledger.yaml records every migrated behavior with its target, the tests that pin the donor, the tests that prove the replacement, and each intentional delta; rows still marked pending name the story that will resolve them.

Description
No description provided
Readme 1.8 MiB
Languages
Python 89.7%
JavaScript 6.2%
HTML 3.2%
CSS 0.4%
Dockerfile 0.3%
Other 0.2%