Compare commits

..

4 Commits

Author SHA1 Message Date
e3ebd0b467 US09-03: Write the Architecture Overview
Some checks failed
Test / suites (pull_request) Failing after 2m58s
Test / container (pull_request) Has been skipped
2026-08-23 22:34:42 +02:00
0d4f9d1c28 US09-02: Write the Installation and Operations Manual (#108)
Some checks failed
Test / suites (push) Failing after 2m40s
Test / container (push) Failing after 5m10s
2026-08-23 21:46:00 +02:00
e508f9fd64 US09-01: Serve the Documentation Inside the Application (#107)
Some checks failed
Test / suites (push) Failing after 2m59s
Test / container (push) Failing after 3m56s
2026-08-23 21:09:43 +02:00
370f966d29 E09: Product documentation backlog (#101)
Some checks failed
Test / suites (push) Failing after 3m12s
Test / container (push) Failing after 4m23s
2026-08-23 15:04:45 +02:00
22 changed files with 5463 additions and 7 deletions

View File

@@ -9,6 +9,7 @@
!photo_pipeline
!migrations
!frontend
!docs
!docker
# Nothing generated, even under an allowed directory.

9
.gitattributes vendored Normal file
View File

@@ -0,0 +1,9 @@
# Vendored third-party browser bundles (US09-01). Minified upstream files carry
# trailing whitespace and very long lines; they are not ours to reformat, and the
# submit gate's `git diff --check` would refuse them forever. Their integrity is
# controlled where it belongs instead: a pinned version and a recorded sha256 in
# frontend/js/vendor/VERSIONS.json, asserted by tests/integration/test_documentation.py.
#
# Only the whitespace check is turned off. They stay text, so a credential scan or a
# search still reads them.
frontend/js/vendor/*.js -whitespace linguist-vendored

View File

@@ -70,6 +70,8 @@ COPY pyproject.toml alembic.ini README.md ./
COPY photo_pipeline ./photo_pipeline
COPY migrations ./migrations
COPY frontend ./frontend
# The manuals are served by the application itself (US09-01), so they ship with it.
COPY docs ./docs
COPY docker/entrypoint.sh docker/healthcheck.sh /usr/local/bin/
# Runtime dependencies only: the `test` extra (pytest, playwright) and the `vision`

224
docs/architecture.md Normal file
View File

@@ -0,0 +1,224 @@
# Architecture
[← Documentation index](index.md)
For whoever has to change this code without breaking somebody's photo library. It
explains what the pieces are, which rules each one keeps, and where to look when a
stage refuses.
The product decisions behind all of it live in `INTEGRATED_PIPELINE_CONCEPT.md`; this
page describes what was built.
## Context
Five things outside the application, and what actually crosses each boundary.
```mermaid
flowchart LR
operator([Operator]):::person -->|browser, one session| app
app[Photo Pipeline]:::system -->|read, rename, EXIF write| library[(Photo library<br/>bind mount)]
app -->|database, cache, journals, backups| data[(Data directory<br/>local filesystem)]
app -->|confirmed-SFW images only| vision[Vision provider]:::ext
app -->|immich-go, verified bytes| immich[Immich server]:::ext
app -->|copy, verify, then remove| archive[(Archive medium)]
classDef person fill:#1f6feb,stroke:#58a6ff,color:#fff
classDef system fill:#238636,stroke:#3fb950,color:#fff
classDef ext fill:#6e40c9,stroke:#a371f7,color:#fff
```
| boundary | leaves the machine? | carries |
|---|---|---|
| operator → app | no (loopback, or a proxy you configured) | commands against stable ids, never paths |
| app → library | no | reads, folder renames, EXIF merges |
| app → data directory | no | SQLite in WAL mode, thumbnails, journals, backups |
| app → vision provider | **yes** | image bytes of confirmed-SFW canonical assets only |
| app → Immich | **yes** | the exact verified bytes of an approved album |
| app → archive medium | no | copies, verified before the source is removed |
## Runtime
Two long-lived processes and a one-shot migration, sharing one database and one
library.
```mermaid
flowchart TD
browser([Browser]) -->|JSON + SSE| api
migrate["migrate<br/>backup, then upgrade"] -->|must exit 0| api
migrate --> worker
api["api · serve<br/>enqueues, serves, reads"] -->|jobs table| db[(SQLite WAL)]
worker["worker<br/>claims and does the work"] -->|jobs table| db
worker --> tools["exiftool · vision API · immich-go"]
api -.->|api.lock.json| lock{{library lock}}
worker -.->|worker.lock.json| lock
api --> lib[(library)]
worker --> lib
```
`serve` enqueues and renders; **it does not do the work**. Everything that scans,
scores, analyses, renames, uploads, or archives happens in the worker, which claims
a queued job atomically with a fencing token. One mutating job runs at a time, and
the lock file in the data directory is what makes a second worker impossible rather
than merely discouraged.
## The modules
| package | owns | must not |
|---|---|---|
| `api/` | HTTP surface, error envelope, security policy | contain SQL or business logic |
| `api/routes/` | one module per resource group, all under `/api/v1` | accept a filesystem path from the browser |
| `api/security.py` | host/origin/CSRF/session/size policy as one pure `evaluate`, plus the ASGI middleware | be bypassed per-route |
| `schemas/` | Pydantic request and response contracts | reach the database |
| `services/` | all domain logic; the only place a decision is made | be imported by the frozen CLI archive |
| `models/` | SQLAlchemy tables | hold behaviour |
| `jobs/` | durable job lifecycle, worker loop, lock ranks, handler registry | run work in the API process |
| `integrations/` | the real outside world: `exiftool`, vision, NSFW model, `immich-go` and its report grammars | be called without a version recorded |
| `path_policy.py` | the library boundary, `_IGNORE/` exclusion, symlink-escape refusal | be duplicated anywhere |
| `imaging.py` | the single bounded-decode door | let a caller open an image directly |
| `faults.py` | the crash barriers the tests fire | read anything but its one env var |
| `config.py` | typed settings, secrets as `SecretStr` | log or return a value |
| `db.py` | engine, sessions, WAL and foreign keys, migration entry | be used to bypass a repository |
Services worth knowing by name: `inventory`, `duplicates`, `safety`, `analysis`,
`albums`/`proposals`/`naming`, `renames`/`rename_apply`/`rename_journal`,
`uploads`/`upload_batches`/`upload_reports`/`upload_verification`,
`archives`/`archive_transfer`/`archive_journal`/`restores`/`availability`,
`thumbnails`, `hashing`, `exif_checkpoint`, `jobs`, `workflow`, `backup`,
`diagnostics`, `app_lock`, `release`, `benchmarks`, `library`, `legacy_import`.
## The state machines
Three journals decide what a restart is allowed to assume. All three are read from
the database plus the real world — never guessed from a missing file.
### Durable jobs
```mermaid
stateDiagram-v2
[*] --> queued
queued --> running: claimed with a fencing token
queued --> cancelled
queued --> cancelling
running --> succeeded
running --> failed
running --> cancelling
cancelling --> cancelled
failed --> retry_queued
retry_queued --> running
succeeded --> [*]
cancelled --> [*]
```
`succeeded` and `cancelled` are terminal, so a duplicate delivery cannot move a
finished job. Claiming is compare-and-set on the row version; a worker whose lease
expired cannot commit after another has taken over.
### The rename journal
The only state machine that moves somebody's folders.
```mermaid
stateDiagram-v2
[*] --> planned
planned --> moving
moving --> moved
moving --> planned: proven untouched
moving --> rollback_required
moved --> database_updated
database_updated --> verified
verified --> complete
moved --> rollback_required
database_updated --> rollback_required
verified --> rollback_required
planned --> failed
failed --> planned
rollback_required --> rolled_back
complete --> [*]
rolled_back --> [*]
```
Intent is written **before** the disk is touched, which is why `moving` can resolve
backwards: the evidence decides. `moving`, `moved`, `database_updated`, and
`rollback_required` are the unsafe states — while any operation sits in one, the
library may be half-renamed, so unrelated mutations are refused with
`409 rename_recovery_required` until a person resolves it.
### Upload batches
```mermaid
stateDiagram-v2
[*] --> planned
planned --> running
running --> succeeded
running --> failed
running --> cancelling
running --> unknown_requires_verification: process died after acceptance
cancelling --> cancelled
failed --> running: retry
cancelled --> running: retry
```
`unknown_requires_verification` is deliberately **not** restartable. The server may
already hold the files; the answer is to ask Immich for the recorded SHA-1, not to
upload again and hope.
Archive transfers use the same shape — `planned → transferring → verified →
removing → complete` — and the source is removed only after the archived bytes are
verified.
## Identity and the data model
A path is metadata. The identity is `assets.id`, a UUID that never changes, and
`asset_paths` records every path an asset has ever had with the reason it changed.
That is what makes a rename cheap: nothing else in the database has to move.
Tables: `assets`, `asset_paths`, `thumbnails`, `jobs`, `job_items`, `job_events`,
`safety_reviews`, `analysis_results`, `exif_projections`, `duplicate_clusters`,
`duplicate_members`, `duplicate_negative_links`, `album_proposals`, `rename_plans`,
`rename_operations`, `upload_batches`, `upload_items`, `upload_verifications`,
`archive_locations`, `archive_plans`, `archive_operations`.
Availability is independent of workflow progress: `active`, `archiving`,
`archived_online`, `archived_offline`, `restoring`, `missing_unexpected`. An
unmounted archive disk is `archived_offline`, never `missing` — the scanner is not
allowed to conclude that a photo is gone because a disk is unplugged.
EXIF is a projection with its own verified state (`verified`, `divergent`,
`failed`): the desired values are written, read back, and compared, and anything the
stage does not own having changed makes the asset `divergent` and blocks the next
mutating stage.
## Where each invariant lives
| invariant | enforced in |
|---|---|
| `_IGNORE/` is never traversed, counted, or opened | `path_policy.is_excluded`, used by every discovery path |
| no path outside the library roots is reachable | `path_policy.resolve_in_roots` — it returns the resolved path, because validating one name and opening another is the symlink race |
| one writer per library | `services/app_lock.py` (an `flock` on a JSON lock file, per role) |
| one mutating job at a time | `services/jobs.py` lock keys plus `jobs/locks.py` rank ordering |
| only confirmed-SFW assets reach the vision provider | `services/analysis.py`, re-checked *after* the provider call so a decision that flipped mid-flight discards the result |
| EXIF is verified, and other fields preserved | `services/exif_checkpoint.py` |
| uploads carry the exact verified bytes | `services/uploads.py` preflight, re-proved immediately before the uploader runs |
| an uncertain upload is not a failure | `services/upload_batches.py`, `services/upload_verification.py` |
| the archive source outlives its copy until verified | `services/archive_transfer.py` |
| decoding is bounded | `imaging.py` |
| the trust boundary and CSRF | `api/security.py` |
## Concurrency
Locks are taken broad to narrow — `library → stage/job → album/folder → asset` — and
never the other way, which is what makes deadlock structural rather than lucky.
Ranks live in `jobs/locks.py`.
Exactly-once execution is not achievable across SQLite, a filesystem, subprocesses,
and a remote server. The design promises **at-least-once with idempotent recovery**:
every handler may run twice, and running twice must not produce two side effects.
That is why the rename journal records intent before moving, why EXIF writes merge
and verify, and why upload retries consult both local history and Immich.
## History
This application was extracted from two command-line tools rather than written from
nothing. Their sources are frozen in `legacy_cli_archive/` with the ledger mapping
each donated behaviour to the service that now owns it, the characterization tests
that pinned it, and every intentional difference. Production code must not import
them; they are provenance and rollback evidence.

35
docs/index.md Normal file
View File

@@ -0,0 +1,35 @@
# Photo Pipeline documentation
One local application that takes a photo library from discovery to a verified Immich
upload: duplicate detection, safety review, content analysis, album naming, guarded
renaming, upload, and archive — one visible, resume-safe workflow.
These pages are readable three ways, and they are the same files each time: in the
repository under `docs/`, on Gitea, and inside the running application under
**Docs**. There is no separate copy to fall out of date.
## Read in this order
1. [Overview](overview.md) — what the application does, the stages it moves a photo
through, and the rules it will not break.
2. [Installation and operations](installation.md) — host and container installation,
every setting, the first-run checklist, upgrades, backup and restore, and what
each refusal at startup means.
3. [Architecture](architecture.md) — the context and runtime diagrams, what each
module owns, the three journals a restart reads, and where every invariant is
enforced.
## Being written
The remaining manuals are accepted work, not aspiration; each is a story in
[E09](https://git.domverse-berlin.eu/domverse/photoanalyzer/src/branch/main/delivery_backlog/E09-documentation.md)
and will appear here as it lands.
- **User manual** — one page per workflow stage with screenshots of the real
application, and a catalogue of every error and refusal (US09-04).
## Conventions
A page tells you what a stage **changes on disk or on the server** before it tells
you how to run it. Refusals are documented as intentional: this application would
rather stop and explain than guess about somebody's photographs.

336
docs/installation.md Normal file
View File

@@ -0,0 +1,336 @@
# Installation and operations
[← Documentation index](index.md)
From nothing to a running instance pointed at your photo library, and everything you
need afterwards: upgrading, backing up, restoring, and understanding a refusal.
Two ways to install. **Container** is the deployment this project builds and ships;
**host** is what you want for development or a single machine you already manage.
Both run the same two processes against the same database.
> Read [what it will not do](overview.md#what-it-will-not-do) first. Several of the
> steps below only make sense once you know which rules the application is keeping.
## Before you start
| you need | version | why |
|---|---|---|
| Python | 3.12 or newer | the application |
| `exiftool` | 13.x (the image pins `13.25+dfsg-1`) | every EXIF read and write |
| `immich-go` | 0.32.0 (pinned in the image) | the upload stage only |
| Docker + the compose plugin | any current | the container installation only |
| a vision provider key | — | the analysis stage only |
You also need a photo library you can afford to be wrong about. Take a backup of it
before pointing anything at it for the first time, and consider running with
[`PHOTO_PIPELINE_REQUIRE_DRY_RUN_APPROVAL`](#the-settings) on.
## Host installation
```bash
python3.12 -m venv .venv
.venv/bin/pip install -e ".[vision]" # drop [vision] for a review-only install
```
Configure it. Every setting is an environment variable; a `.env` file in the working
directory is read at startup:
```bash
cp .env.example .env && $EDITOR .env
```
`.env.example` lists every variable with its default and no values at all. For a
loopback installation you need exactly one:
```bash
PHOTO_PIPELINE_LIBRARY_ROOTS=/home/you/Pictures
```
Then migrate and start the two processes:
```bash
.venv/bin/python -m photo_pipeline migrate # create or upgrade the database
.venv/bin/python -m photo_pipeline serve # API and browser app on 127.0.0.1:8000
.venv/bin/python -m photo_pipeline worker # second terminal
```
**Both processes are required.** `serve` enqueues work and renders the application;
nothing is scanned, scored, analysed, renamed, uploaded, or archived without a
worker. Open <http://127.0.0.1:8000/app/> and confirm the workflow page loads.
### The configuration file
`.env`, or any path named by `PHOTO_PIPELINE_ENV_FILE`. It is parsed, never executed:
`KEY=value` lines, `#` comments, optional quotes, no interpolation and no `export`. A
configuration file that can run code is a configuration file that can be a
vulnerability.
**Anything already exported in the shell wins.** The file is your standing
configuration; the environment is the override for one run.
The archived CLI's names still work, so an existing `photo_analyzer.env` can be used
as it is:
| in the file | applied as |
|---|---|
| `LLM_API_KEY` / `GEMINI_API_KEY` | `OPENAI_API_KEY` |
| `LLM_BASE_URL` | `OPENAI_BASE_URL` |
| `LIBRARY` | `PHOTO_PIPELINE_LIBRARY_ROOTS` |
`.env` and `*.env` are gitignored and refused by the repository's safety checks. The
file holds a real key; it must never be committed.
## Container installation
```bash
cp .env.example .env && $EDITOR .env
docker compose up -d --build
```
The composition is one `serve` container, one `worker` container, a one-shot
`migrate` that both wait for, one bind-mounted library, and one named data volume.
Three variables it cannot start without:
| variable | is |
|---|---|
| `PHOTO_PIPELINE_LIBRARY_HOST_PATH` | the library on this host |
| `PHOTO_PIPELINE_LIBRARY_ROOTS` | where that library is mounted **inside** the container |
| `PHOTO_PIPELINE_ACCESS_SECRET` | required, because publishing a port means the app is reachable from outside the container |
Set `PHOTO_PIPELINE_UID` and `PHOTO_PIPELINE_GID` to the owner of the library: what
the containers rename and rewrite keeps that ownership.
### The data volume
The `data` volume holds the database, its write-ahead log, the thumbnail cache, the
operation journals, and the backups. **It must stay on a local filesystem.** SQLite
in WAL mode needs real local locking, so NFS, SMB, and network volume drivers do not
slow it down — they corrupt it. The library bind mount has no such restriction.
### Ports and reaching it
The API port is published to `127.0.0.1` unless `PHOTO_PIPELINE_PUBLISH_ADDRESS` says
otherwise. Being reachable *was* the authentication in earlier versions: whoever could
open the port owned the library. So the moment the application answers to anything but
loopback — a hostname, a proxy, `0.0.0.0` — the access secret becomes mandatory and
`serve` refuses to start without it rather than publishing your photographs.
Behind a reverse proxy:
```bash
PHOTO_PIPELINE_ALLOWED_HOSTS=photos.example.com
PHOTO_PIPELINE_ACCESS_SECRET=# python -c 'import secrets; print(secrets.token_urlsafe(32))'
PHOTO_PIPELINE_TRUSTED_PROXIES=10.0.0.2 # only the proxy's own address
```
`X-Forwarded-Proto` and `X-Forwarded-Host` are believed only from a trusted-proxy
address, so a client cannot declare its own origin. The browser asks for the secret
once per tab. Wrong secrets are rate-limited and logged with the caller's address
only. The health endpoints stay open so an orchestrator can restart the container;
nothing else is.
### Deployment
The stack is managed from git by Portainer, which redeploys on a webhook after CI
publishes an image built from `main`. Runtime secrets live in the Portainer stack's
environment rather than in the repository, so rotation happens in one place and
reading the repository discloses nothing.
## The settings
Every variable, its default, and whether it is a secret. `.env.example` is the same
list in copyable form.
### Library and data
| variable | default | meaning |
|---|---|---|
| `PHOTO_PIPELINE_LIBRARY_ROOTS` | none | the library boundary, `os.pathsep`-separated. No path outside these roots is ever read or written |
| `PHOTO_PIPELINE_DATA_DIR` | `data` | database, WAL, thumbnail cache, journals, backups. Never inside the library |
| `PHOTO_PIPELINE_DB_PATH` | `<data dir>/photo_pipeline.db` | the database file, if it must live elsewhere |
| `PHOTO_PIPELINE_ENV_FILE` | `.env` | where to read the configuration file from |
### Serving and the trust boundary
| variable | default | meaning |
|---|---|---|
| `PHOTO_PIPELINE_HOST` | `127.0.0.1` | bind address |
| `PHOTO_PIPELINE_PORT` | `8000` | port |
| `PHOTO_PIPELINE_ALLOWED_HOSTS` | none | comma-separated names the app answers to besides loopback. Naming one makes the access secret mandatory |
| `PHOTO_PIPELINE_ACCESS_SECRET` | none | **secret.** Traded for the session cookie at `GET /api/v1/session` |
| `PHOTO_PIPELINE_TRUSTED_PROXIES` | none | comma-separated peer addresses whose forwarded headers may be believed |
| `PHOTO_PIPELINE_MAX_REQUEST_BYTES` | `1048576` | largest request body accepted |
### Logging
| variable | default | meaning |
|---|---|---|
| `PHOTO_PIPELINE_LOG_LEVEL` | `INFO` | standard Python levels |
| `PHOTO_PIPELINE_LOG_FORMAT` | `json` | `json` or `text` |
### Limits
| variable | default | meaning |
|---|---|---|
| `PHOTO_PIPELINE_THUMBNAIL_CACHE_QUOTA_BYTES` | `500000000` | cache quota; over it is a diagnostics warning |
| `PHOTO_PIPELINE_THUMBNAIL_MAX_PIXELS` | `100000000` | refuse to decode anything larger. This is the decompression-bomb guard |
| `PHOTO_PIPELINE_ARCHIVE_FREE_SPACE_RESERVE_BYTES` | `1000000000` | free space an archive destination must keep beyond the transfer itself |
### The safety gate
| variable | default | meaning |
|---|---|---|
| `PHOTO_PIPELINE_REQUIRE_DRY_RUN_APPROVAL` | `false` | refuse every mutating request until a read-only dry run of this library has been produced and approved |
### External services
| variable | default | meaning |
|---|---|---|
| `PHOTO_PIPELINE_VISION_API_KEY` | none | **secret.** Without it the analysis stage cannot run |
| `PHOTO_PIPELINE_IMMICH_SERVER_URL` | none | the Immich server |
| `PHOTO_PIPELINE_IMMICH_API_KEY` | none | **secret.** |
| `PHOTO_PIPELINE_IMMICH_GO_BINARY` | `immich-go` | the uploader, found on `PATH` |
### Composition only
Read by `docker-compose.yml`, not by the application:
`PHOTO_PIPELINE_IMAGE`, `PHOTO_PIPELINE_LIBRARY_HOST_PATH`,
`PHOTO_PIPELINE_PUBLISH_ADDRESS`, `PHOTO_PIPELINE_UID`, `PHOTO_PIPELINE_GID`.
A secret is never written to a log, never returned by the API, never stored in the
database, and never recorded in a backup manifest — a manifest says `configured`, not
the value. Do not put a real key in an example, a ticket, or a screenshot.
## First run checklist
Not "it started" — verified.
1. **The database is at the current revision.**
`python -m photo_pipeline migrate` exits 0.
2. **The application answers.**
`curl -s localhost:8000/api/v1/health/ready` returns 200.
3. **The library was found.** Open the app, run a scan from the Inventory view, and
check the count against what you expect. Anything under `_IGNORE/` must be missing
from it — that is the exclusion working, not a bug.
4. **The worker is claiming.** The scan job reaches `succeeded`. If it stays
`queued`, no worker is running.
5. **Diagnostics are clean.**
`python -m photo_pipeline diagnostics` reports free space and an empty `warnings`.
6. **Nothing was modified.** The scan is read-only; your files' timestamps are
unchanged.
Only then point it at the whole library.
## Operations
Every operation is the same CLI, on a host or in the composition:
```bash
python -m photo_pipeline diagnostics
docker compose run --rm --no-deps api diagnostics
```
`--no-deps` keeps a one-off command from starting a second stack; `api` is only the
service it borrows the image and mounts from.
### Upgrading
1. `python -m photo_pipeline backup --reason before-upgrade`
2. Pull the new version (or `docker compose pull && docker compose up -d`).
3. Migrations run by themselves at startup, and a **pending schema change is
snapshotted first**. If the upgrade fails, the previous database and its
`pre-migration` backup are both intact, and the error log names the backup
directory.
An up-to-date database is not backed up again on every start.
### Backing up
```bash
python -m photo_pipeline backup --reason weekly --keep 7
python -m photo_pipeline verify-backup data/backups/<name>
```
Backups go through SQLite's online backup API, never a file copy: with WAL enabled
the `.db` file alone is missing every committed page still in the write-ahead log.
Each backup is a directory holding the snapshot and a `manifest.json` — schema
revision, SHA-256, row counts, the archive media the library depends on, and which
settings were configured.
`verify-backup` runs `PRAGMA integrity_check` **and** `PRAGMA foreign_key_check`,
compares the snapshot's SHA-256 against the manifest, and re-counts every table it
recorded. Bit rot, a truncated copy, and a "repaired" snapshot all fail it.
`--keep N` prunes the oldest and never the newest. The same is available at
`GET /api/v1/diagnostics`, `GET|POST /api/v1/backups`,
`GET /api/v1/backups/{name}/verify`, and `POST /api/v1/backups/prune`.
### Restoring
Restore is deliberately **not** an API call. It replaces the state of an
installation, so it belongs to a stopped one and a person at a terminal.
1. Stop `serve` and `worker`.
2. `python -m photo_pipeline verify-backup data/backups/<name>` — never restore an
unverified snapshot, and `restore` will refuse one anyway.
3. `python -m photo_pipeline restore data/backups/<name> --into /path/to/fresh-data`.
A target that already holds a database is refused; recovering in place means
moving the old data directory aside first.
4. Point `PHOTO_PIPELINE_DATA_DIR` at the restored directory and run `migrate`.
5. Run an inventory scan, so paths are reconciled against the real library.
6. Mount every archive location the manifest names before archiving again. The
database records where archived originals are; it does not contain them.
Practise this against a copy before you need it.
### Watching it
`diagnostics` reports the database, write-ahead log, thumbnail cache, uploader
reports, backups, and logs separately, with free space and warnings for low disk
(`disk_low`, `disk_critical`), a cache over its quota, a write-ahead log outgrowing
its database, and a legacy CLI writing the library.
Logs go to stdout in JSON by default — `docker compose logs -f worker`, or your
service manager's journal on a host. Uploader output is kept per attempt under the
data directory with the API key scrubbed line by line.
## When it refuses to start
Every refusal below is deliberate. The application would rather stop and explain than
guess about somebody's photographs.
| exit code | meaning | what to do |
|---|---|---|
| `1` | the operation failed and said why — an unverifiable backup, an occupied restore target, an unknown benchmark profile | read the message; nothing was changed |
| `2` | **the library lock is held.** Another `serve` or `worker` owns this data directory | stop the other process. The message names the holder; a lock whose process is gone is taken over automatically |
| `3` | **a legacy CLI looks active.** The frozen command-line tools are writing this library | stop them. `--allow-legacy` overrides and you own the outcome |
| `4` | **configuration refused.** The application is reachable beyond loopback and no access secret is set | set `PHOTO_PIPELINE_ACCESS_SECRET`, or bind to loopback only |
| `5` | **a library root is not mounted** (containers only) | fix the bind mount. Refusing is what stops the container writing into its own throwaway layer instead of your library |
Other things that look like faults and are not:
- **`403 host_not_allowed`** — the `Host` header is not a name in
`PHOTO_PIPELINE_ALLOWED_HOSTS`. Add the real hostname; do not add `*`.
- **`409 lock_held`** — a mutating job is already running. One at a time is what
makes a crash recoverable.
- **`409 rename_recovery_required`** — an interrupted rename is unresolved. Resolve
it in the Renames view; unrelated mutations stay blocked until then, on purpose.
- **A job stuck in `queued`** — no worker is running.
- **An empty scan** — check `PHOTO_PIPELINE_LIBRARY_ROOTS`, and in a container check
that the path is the *container-side* mount, not the host path.
## What an installer must not work around
- **One writer.** Do not run two workers, and do not run the frozen CLI beside the
application. Each is safe alone and destructive together — the lock is not
bureaucracy.
- **`_IGNORE/` is never read.** Do not "fix" the exclusion.
- **EXIF is verified before upload.** Do not skip the checkpoint to make a stage
finish; an unverified projection is exactly the case where the wrong bytes reach
Immich.
- **The data directory is not the library.** Keep them apart, and keep the data
directory on a local filesystem.
- **Secrets stay in the environment.** Not in the database, not in a log, not in a
commit.

77
docs/overview.md Normal file
View File

@@ -0,0 +1,77 @@
# Overview
[← Documentation index](index.md)
Photo Pipeline sorts a photo library. It finds duplicates before anything expensive
happens to them, asks a human which pictures may leave the machine, describes the
ones that may, proposes album names from what it found, renames folders under a
crash-safe journal, uploads through `immich-go`, and can archive a finished album off
active storage without ever forgetting it existed.
It runs locally. It talks to exactly three things outside itself: a vision provider,
an Immich server, and `exiftool` — and it will tell you before it uses any of them.
## The workflow
Each stage is a gate, not a tab. A later stage can always be looked at; its actions
stay disabled until what they depend on is true.
```mermaid
flowchart TD
I["0 · Inventory<br/>discover, hash, cluster duplicates"] --> S["1 · Safety<br/>score and human decision"]
S -->|sfw| A["2 · Analysis<br/>vision provider, EXIF checkpoint"]
S -->|nsfw| U
A --> B["3 · Albums<br/>proposal, then guarded rename"]
B --> U["4 · Upload<br/>immich-go, verified bytes"]
U --> R["5 · Archive<br/>copy, verify, then reclaim space"]
```
| Stage | What it decides | What it changes |
|---|---|---|
| Inventory | which file is the canonical copy of a picture | nothing — it only reads |
| Safety | whether a photo may be sent to a cloud provider | one `sfw`/`nsfw` EXIF keyword |
| Analysis | what a photo shows | a managed caption segment and additive keywords |
| Albums | what a folder should be called | folder names, through a journaled rename |
| Upload | which exact bytes reach Immich | nothing locally; assets appear in Immich |
| Archive | which album leaves active storage | files move to the archive, after verification |
## What it will not do
These are enforced in code, not by convention, and each one is why some action you
expected is sometimes refused.
- **Nothing under `_IGNORE/` is ever read.** Not scanned, not counted, not
thumbnailed, not sent anywhere.
- **A path is not an identity.** Every picture has a stable id, so moving or renaming
it loses no history.
- **Only a confirmed-SFW photo reaches the vision provider.** A photo marked NSFW is
still uploadable to Immich; it simply never leaves for analysis.
- **Metadata is verified, not hoped for.** Every stage that writes EXIF reads it back
and proves that what it did not own is unchanged.
- **One writer at a time.** A library lock, held by one process, is what makes a
crash recoverable instead of ambiguous.
- **Nothing irreversible happens without a preview and an explicit approval** that
names the exact count.
## Where things live
| | |
|---|---|
| the photo library | wherever you point `PHOTO_PIPELINE_LIBRARY_ROOTS`; mounted read-write |
| the database, cache, logs, backups | the data directory, never inside the library |
| secrets | the environment, never the database and never a log line |
| these documents | `docs/` in the repository, served at `/docs` by the application |
## Running it
The short version, for a host installation:
```bash
python -m photo_pipeline migrate
python -m photo_pipeline serve # the API and this browser application
python -m photo_pipeline worker # the process that does the long work
```
The full procedure, the container path, and every setting belong to the installation
manual, which is being written next — see
[the documentation index](index.md#being-written).

View File

@@ -170,6 +170,20 @@ a.link:hover { text-decoration: underline; }
.grid th, .grid td { text-align: left; padding: 6px 8px; border-bottom: 1px solid var(--border); vertical-align: top; }
.path { font-family: ui-monospace, monospace; font-size: 0.85em; word-break: break-all; }
/* Docs: prose, so it gets a reading measure rather than the full window width. */
.doc { max-width: 72ch; }
.doc h1 { margin-top: 0; }
.doc h2, .doc h3 { margin-top: 1.6em; border-bottom: 1px solid var(--border); padding-bottom: 4px; }
.doc code { font-family: ui-monospace, monospace; font-size: 0.9em; background: var(--surface-2); padding: 1px 4px; border-radius: 4px; }
.doc pre { background: var(--surface-2); border: 1px solid var(--border); border-radius: var(--radius); padding: 12px; overflow: auto; }
.doc pre code { background: none; padding: 0; }
.doc table { border-collapse: collapse; width: 100%; }
.doc th, .doc td { text-align: left; padding: 6px 8px; border-bottom: 1px solid var(--border); vertical-align: top; }
.doc blockquote { margin: 1em 0; padding-left: 12px; border-left: 3px solid var(--border); color: var(--muted); }
.doc img { max-width: 100%; }
.diagram { margin: 1.5em 0; padding: 0; overflow: auto; }
.diagram svg { max-width: 100%; height: auto; }
/* Narrow screens: stack the panes so blockers and actions stay reachable. */
@media (max-width: 720px) {
.two-pane { grid-template-columns: 1fr; }

View File

@@ -21,6 +21,7 @@
<a href="#/uploads" data-nav="uploads">Upload</a>
<a href="#/archive" data-nav="archive">Archive</a>
<a href="#/stats" data-nav="stats">Stats</a>
<a href="#/docs" data-nav="docs">Docs</a>
</nav>
</header>
<main id="app" aria-live="polite"><!-- views render here --></main>

View File

@@ -1,5 +1,6 @@
import { api } from "./api.js";
import { renderArchive, setArchiveRender } from "./archive.js";
import { renderDocs } from "./docs.js";
import { navigate, onRouteChange, parseHash } from "./router.js";
import { renderRenames, setRenamesRender } from "./renames.js";
import { renderUploads, setUploadsRender } from "./uploads.js";
@@ -385,6 +386,7 @@ function render() {
else if (path === "/uploads") renderUploads(root, params);
else if (path === "/archive") renderArchive(root, params);
else if (path === "/stats") renderStats(root, params);
else if (path === "/docs") renderDocs(root, params);
else show(errorBanner("Unknown view"));
}

263
frontend/js/docs.js Normal file
View File

@@ -0,0 +1,263 @@
// The documentation view (US09-01): the manuals in `docs/`, rendered in the app.
//
// The same markdown files are the repository's documentation and the deployment's
// documentation. An operator who was handed a URL and an access secret has no
// checkout in front of them, and the network the application runs on is not assumed
// to reach a CDN — so the renderer is vendored and everything here is same-origin.
//
// Nothing on this page is authenticated. It is served from the static mount beside
// the application shell, exactly like `index.html`: the troubleshooting page is
// needed most by whoever cannot get past the access secret, and no documentation
// file contains anything a session would protect.
//
// `marked` and `mermaid` are loaded lazily, on the first documentation page and the
// first diagram. They are large, and every other view does without them.
import { el, errorBanner, setActiveNav } from "./dom.js";
const DOCS_BASE = "/docs";
const VENDOR = "/app/js/vendor";
const INDEX_PAGE = "index";
// A page name comes from the hash, so it is caller input: no scheme, no traversal,
// no absolute path. The server would refuse those too; this refuses them earlier and
// without a request that looks like an attempt.
const PAGE_PATTERN = /^[\w-]+(\/[\w-]+)*$/;
let markedPromise;
let mermaidPromise;
function loadMarked() {
markedPromise ??= import(`${VENDOR}/marked.esm.js`).then((module) => module.marked);
return markedPromise;
}
// mermaid ships one large UMD bundle rather than a self-contained ES module, so it
// arrives through a script element. `script-src 'self'` allows it because it is ours
// and same-origin; nothing here relaxes that.
function loadMermaid() {
mermaidPromise ??= new Promise((resolve, reject) => {
const script = document.createElement("script");
script.src = `${VENDOR}/mermaid.min.js`;
script.onload = () => resolve(window.mermaid);
script.onerror = () => reject(new Error("the diagram renderer could not be loaded"));
document.head.appendChild(script);
});
return mermaidPromise;
}
// ── the pure parts, unit-tested in js/tests/unit.js ──────────────────────────
/** A stable, readable heading anchor. Letters and digits of any script survive. */
export function slug(text) {
const cleaned = String(text)
.toLowerCase()
.trim()
.replace(/[^\p{L}\p{N}\s-]/gu, "")
.replace(/[\s-]+/g, "-")
.replace(/^-|-$/g, "");
return cleaned || "section";
}
/** Assign every heading an id, disambiguating repeats the way a reader would
* expect: the first `Notes` keeps `#notes`, the second becomes `#notes-2`. */
export function assignHeadingIds(headings) {
const used = new Map();
for (const heading of headings) {
const base = slug(heading.textContent);
const seen = (used.get(base) || 0) + 1;
used.set(base, seen);
heading.id = seen === 1 ? base : `${base}-${seen}`;
}
return headings;
}
/**
* Where a link inside a documentation page should go.
*
* Markdown links between documents are relative file paths, which a browser would
* treat as downloads that leave the application. Returns the in-app target for a
* link to another document or to a heading, and `null` for anything else — external
* links stay exactly as the author wrote them.
*/
export function resolveDocLink(currentPage, href) {
if (!href) return null;
if (/^[a-z][a-z\d+.-]*:/i.test(href) || href.startsWith("//")) return null;
if (href.startsWith("#")) return { page: currentPage, anchor: href.slice(1) };
const [target, anchor = ""] = href.split("#");
if (!/\.md$/i.test(target)) return null;
// Resolved against the current document, so `../guides/x.md` means what it means
// in the repository. The origin is a placeholder; only the path is used.
const resolved = new URL(target, `https://docs.invalid/${currentPage}.md`);
const page = decodeURIComponent(resolved.pathname).replace(/^\//, "").replace(/\.md$/i, "");
return PAGE_PATTERN.test(page) ? { page, anchor } : null;
}
/**
* Point every in-app link at its route, and leave every other link alone.
*
* The resolved page is kept on the element, so the reading order can be read back
* from the rendered index rather than parsed out of the markdown a second time.
*/
export function rewriteLinks(article, page) {
for (const link of article.querySelectorAll("a[href]")) {
const href = link.getAttribute("href");
const target = resolveDocLink(page, href);
if (target) {
link.setAttribute("href", docHash(target.page, target.anchor));
link.dataset.docPage = target.page;
} else if (/^https?:/i.test(href)) {
link.setAttribute("rel", "noreferrer noopener");
link.setAttribute("target", "_blank");
}
}
return article;
}
/** The reading order, taken from the index document itself — one list, in one
* place, that is equally the index on Gitea and the navigation here. */
export function documentIndex(indexElement) {
const pages = [];
for (const link of indexElement.querySelectorAll("a[data-doc-page]")) {
const page = link.dataset.docPage;
if (page !== INDEX_PAGE && !pages.some((entry) => entry.page === page)) {
pages.push({ page, title: link.textContent.trim() });
}
}
return pages;
}
// ── the view ─────────────────────────────────────────────────────────────────
async function fetchPage(page) {
const response = await fetch(`${DOCS_BASE}/${page}.md`, { headers: { accept: "text/markdown" } });
if (!response.ok) throw new Error(`documentation page unavailable: ${response.status}`);
return response.text();
}
async function toArticle(markdown, page) {
const marked = await loadMarked();
const article = el("article", { class: "doc", "data-testid": "doc", "data-page": page });
// The only innerHTML in the application, and deliberate: rendering markdown *is*
// producing HTML. The input is a file from this repository, and the page's CSP
// (`script-src 'self'`, no `'unsafe-inline'`) means injected script and inline
// handlers do not run even if one ever were not.
article.innerHTML = marked.parse(markdown, { async: false });
assignHeadingIds(article.querySelectorAll("h1, h2, h3, h4, h5, h6"));
rewriteLinks(article, page);
return article;
}
function docHash(page, anchor = "") {
const query = new URLSearchParams(anchor ? { page, anchor } : { page });
return `#/docs?${query}`;
}
/** Diagrams are authored as ```mermaid blocks so the source is diffable and Gitea
* renders them natively. A failure here replaces the diagram, never the page. */
async function renderDiagrams(article) {
const blocks = [...article.querySelectorAll("pre > code.language-mermaid")];
if (!blocks.length) return;
let mermaid;
try {
mermaid = await loadMermaid();
mermaid.initialize({ startOnLoad: false, securityLevel: "strict", theme: "dark" });
} catch (error) {
blocks.forEach((block) => block.closest("pre").replaceWith(errorBanner(error.message)));
return;
}
for (const [index, block] of blocks.entries()) {
const figure = el("figure", { class: "diagram", "data-testid": "diagram" });
block.closest("pre").replaceWith(figure);
try {
const { svg } = await mermaid.render(`diagram-${index}-${Date.now()}`, block.textContent);
figure.innerHTML = svg;
} catch (error) {
figure.replaceWith(errorBanner(`Diagram could not be drawn: ${error.message}`));
}
}
}
function notFound() {
return el(
"div",
{ class: "alert", role: "alert", "data-testid": "doc-not-found" },
"That documentation page does not exist. ",
el("a", { class: "link", href: docHash(INDEX_PAGE) }, "Back to the documentation index")
);
}
function sidebar(pages, current) {
return el(
"nav",
{ class: "card", "aria-label": "Documentation" },
el("h2", {}, "Documentation"),
el(
"ul",
{ class: "album-list", "data-testid": "doc-pages" },
el(
"li",
{},
el(
"a",
{
href: docHash(INDEX_PAGE),
"aria-current": current === INDEX_PAGE ? "true" : false,
},
"Index"
)
),
...pages.map(({ page, title }) =>
el(
"li",
{},
el(
"a",
{
href: docHash(page),
"data-page": page,
"aria-current": page === current ? "true" : false,
},
title
)
)
)
)
);
}
export async function renderDocs(root, params = {}) {
setActiveNav("docs");
const requested = params.page || INDEX_PAGE;
const page = PAGE_PATTERN.test(requested) ? requested : "";
let indexArticle;
try {
indexArticle = await toArticle(await fetchPage(INDEX_PAGE), INDEX_PAGE);
} catch (error) {
root.replaceChildren(errorBanner(`Documentation is unavailable: ${error.message}`));
return;
}
const pages = documentIndex(indexArticle);
let article;
if (!page) article = notFound();
else if (page === INDEX_PAGE) article = indexArticle;
else {
try {
article = await toArticle(await fetchPage(page), page);
} catch {
article = notFound();
}
}
root.replaceChildren(el("div", { class: "two-pane" }, sidebar(pages, page), article));
await renderDiagrams(article);
scrollToAnchor(params.anchor);
}
// A documentation anchor cannot live in the hash — the hash is the route — so it
// travels as a parameter and is applied after the page renders.
function scrollToAnchor(anchor) {
if (!anchor) return;
const target = document.getElementById(anchor);
if (target) target.scrollIntoView({ block: "start" });
}

View File

@@ -7,6 +7,7 @@ import { createStore } from "../store.js";
import { parseHash, navigate } from "../router.js";
import { api, cancellable } from "../api.js";
import { subscribeJob } from "../events.js";
import { assignHeadingIds, documentIndex, resolveDocLink, rewriteLinks, slug } from "../docs.js";
const cases = [];
function ok(name, cond) {
@@ -142,6 +143,62 @@ async function run() {
window.EventSource = realES;
}
// ── documentation view (US09-01) ─────────────────────────────────────────
{
ok("slug lowercases and dashes a heading", slug("Key Flows") === "key-flows");
ok(
"slug drops punctuation but keeps words",
slug("What it *will not* do:") === "what-it-will-not-do"
);
ok("slug keeps non-ASCII letters", slug("Größe & Gewicht") === "größe-gewicht");
ok("slug never yields an empty anchor", slug("!!!") === "section");
const doc = document.createElement("div");
doc.innerHTML = "<h2>Notes</h2><h3>Notes</h3><h2>Notes</h2>";
const ids = [...assignHeadingIds(doc.querySelectorAll("h2, h3"))].map((h) => h.id);
ok("repeated headings get distinct anchors", ids.join(",") === "notes,notes-2,notes-3");
ok(
"a sibling document link becomes an in-app route",
resolveDocLink("index", "overview.md")?.page === "overview"
);
const nested = resolveDocLink("guides/install", "../overview.md#running-it");
ok("a relative link resolves against the current page", nested?.page === "overview");
ok("a link's anchor survives the rewrite", nested?.anchor === "running-it");
ok(
"a bare anchor stays on the current page",
resolveDocLink("overview", "#the-workflow")?.page === "overview"
);
ok("an external link is left alone", resolveDocLink("index", "https://example.test/x") === null);
ok("a non-markdown relative link is left alone", resolveDocLink("index", "images/a.png") === null);
// A climb cannot leave the documentation tree: it is clamped at the root, so the
// page it names is still fetched from under /docs and simply does not exist.
ok(
"a link that climbs above the docs root is clamped",
resolveDocLink("index", "../../etc/passwd.md")?.page === "etc/passwd"
);
ok(
"a page name that is not a page name is refused",
resolveDocLink("index", "..%2f..%2fetc%2fpasswd.md") === null
);
const index = document.createElement("div");
index.innerHTML =
'<a href="overview.md">Overview</a><a href="overview.md">again</a>' +
'<a href="index.md">itself</a><a href="https://example.test">out</a>';
const listed = documentIndex(rewriteLinks(index, "index"));
ok(
"an external link is marked safe to open away from the app",
index.querySelector('a[href^="https"]').rel === "noreferrer noopener"
);
ok(
"an in-app link points at the docs route",
index.querySelector("a").getAttribute("href") === "#/docs?page=overview"
);
ok("the index lists each page once, in order", listed.length === 1);
ok("the index takes its titles from the link text", listed[0].title === "Overview");
}
await tick();
const failed = cases.filter((c) => !c.ok);
window.__RESULTS__ = { passed: cases.length - failed.length, failed: failed.length, cases };

23
frontend/js/vendor/VERSIONS.json vendored Normal file
View File

@@ -0,0 +1,23 @@
{
"_comment": "Vendored browser libraries (US09-01). Committed rather than fetched: the application is deployed to a network whose outbound access is not assumed, and `default-src 'self'` forbids a CDN. Checksums are asserted by tests/integration/test_documentation.py, so replacing a file without recording it here fails the suite. Update by downloading the pinned URL and recording the new version and sha256 in the same commit.",
"libraries": [
{
"file": "marked.esm.js",
"name": "marked",
"version": "18.0.10",
"license": "MIT",
"url": "https://cdn.jsdelivr.net/npm/marked@18.0.10/lib/marked.esm.js",
"sha256": "4cf47dfebb7f614a08fc0a579ab0fe407ff0ed2b717bf953040c85b2f493a4f0",
"why": "Markdown to HTML for the documentation view. An ES module with no dependencies, imported lazily by js/docs.js."
},
{
"file": "mermaid.min.js",
"name": "mermaid",
"version": "11.17.0",
"license": "MIT",
"url": "https://cdn.jsdelivr.net/npm/mermaid@11.17.0/dist/mermaid.min.js",
"sha256": "8d8e0eec56d3a83b4b3c87f42050845546dee93ebe1875d2117c12e6947c0cb3",
"why": "Renders ```mermaid blocks so diagrams are diffable source that Gitea also renders. The single-file UMD build rather than the ES module entry, whose ~40 lazy chunks would each need vendoring and pinning."
}
]
}

78
frontend/js/vendor/marked.esm.js vendored Normal file

File diff suppressed because one or more lines are too long

3636
frontend/js/vendor/mermaid.min.js vendored Normal file

File diff suppressed because one or more lines are too long

View File

@@ -53,6 +53,7 @@ from photo_pipeline.services.thumbnails import ThumbnailService
from photo_pipeline.services.upload_batches import UploadBatchService
FRONTEND_DIR = Path(__file__).resolve().parents[2] / "frontend"
DOCS_DIR = Path(__file__).resolve().parents[2] / "docs"
log = logging.getLogger(__name__)
@@ -151,4 +152,10 @@ def create_app(config: Config | None = None) -> FastAPI:
# Static single-page app (hash-routed). Mounted last so /api/v1 wins.
if FRONTEND_DIR.is_dir():
app.mount("/app", StaticFiles(directory=FRONTEND_DIR, html=True), name="app")
# The manuals, as their own markdown files (US09-01). Unauthenticated for the
# same reason the application shell is: whoever cannot get past the access
# secret is exactly who needs the troubleshooting page, and no document holds
# anything a session would protect.
if DOCS_DIR.is_dir():
app.mount("/docs", StaticFiles(directory=DOCS_DIR), name="docs")
return app

View File

@@ -57,9 +57,21 @@ LOOPBACK_HOSTS = frozenset({"127.0.0.1", "localhost", "::1", "[::1]"})
# backup a careful operator takes first, and its retention (US07-07).
MUTATION_EXEMPT_PATHS = frozenset({f"{API_PREFIX}/backups", f"{API_PREFIX}/backups/prune"})
# Applied to every response. No inline script/style is used by the frontend, so the
# policy can stay strict; `frame-ancestors 'none'` and CORP keep other pages from
# Applied to every response. `frame-ancestors 'none'` and CORP keep other pages from
# embedding the app or its thumbnails.
#
# `script-src 'self'` is the boundary that matters and it is unchanged: no
# `'unsafe-eval'`, no `'unsafe-inline'`, so nothing injected into the DOM can execute.
# Whether the vendored diagram renderer needed `'unsafe-eval'` was measured rather
# than assumed — mermaid's bundle contains no `eval(` and no `new Function`, and it
# renders under this exact policy without a single script-src violation.
#
# `style-src` does gain `'unsafe-inline'` (US09-01): mermaid styles the SVG it builds
# with an injected `<style>` element and `style=` attributes, and a diagram's CSS
# cannot be hashed in advance. The concession is bounded by the directives around it
# — with script execution still refused and `img-src`, `connect-src`, and `font-src`
# all `'self'`, the CSS exfiltration channels stay closed and what is left is
# defacement of a page its own operator is already looking at.
DEFAULT_HEADERS = {
"x-content-type-options": "nosniff",
"x-frame-options": "DENY",
@@ -67,8 +79,9 @@ DEFAULT_HEADERS = {
"cross-origin-resource-policy": "same-origin",
"cross-origin-opener-policy": "same-origin",
"content-security-policy": (
"default-src 'self'; img-src 'self' data:; style-src 'self'; script-src 'self'; "
"connect-src 'self'; frame-ancestors 'none'; base-uri 'none'; form-action 'none'"
"default-src 'self'; img-src 'self' data:; style-src 'self' 'unsafe-inline'; "
"script-src 'self'; connect-src 'self'; font-src 'self'; "
"frame-ancestors 'none'; base-uri 'none'; form-action 'none'"
),
}

149
tests/e2e/test_docs_ui.py Normal file
View File

@@ -0,0 +1,149 @@
"""US09-01: the manuals, in the browser, under the application's real policy.
Everything here runs against a real ``photo_pipeline serve`` process serving the
committed ``docs/`` tree and the vendored renderer — no fixture markdown, no stubbed
fetch. What that buys is the assertion the story actually cares about: the pages
render, the links between them work, a diagram becomes a diagram, and the console
stays empty, including of CSP violations.
The offline contract — reachability, dead links, pinned checksums — is
``tests/integration/test_documentation.py``.
"""
from __future__ import annotations
from pathlib import Path
import pytest
from tests.e2e._pipeline_harness import Server, seed_library
@pytest.fixture(scope="module")
def server(tmp_path_factory):
seeded = seed_library(tmp_path_factory.mktemp("docs"), {"a": 1}, {})
running = Server(seeded).start()
try:
yield running
finally:
running.stop()
@pytest.fixture
def quiet(page):
"""Any console error, page error, or failed request fails the test that caused
it. A CSP violation arrives as a console error, which is the point."""
problems: list[str] = []
page.on("console", lambda m: problems.append(m.text) if m.type == "error" else None)
page.on("pageerror", lambda error: problems.append(str(error)))
page.on("requestfailed", lambda request: problems.append(f"failed request: {request.url}"))
yield problems
def test_the_documentation_opens_from_the_navigation(page, server, quiet):
page.goto(f"{server.base}/app/#/workflow")
page.locator('nav a[data-nav="docs"]').click()
page.get_by_test_id("doc").wait_for()
assert page.get_by_test_id("doc").get_attribute("data-page") == "index"
# The sidebar is the index's own reading order, not a second list to maintain.
assert page.get_by_test_id("doc-pages").get_by_role("link", name="Overview").is_visible()
assert quiet == []
def test_a_link_between_documents_stays_inside_the_application(page, server, quiet):
page.goto(f"{server.base}/app/#/docs")
page.get_by_test_id("doc").wait_for()
page.get_by_test_id("doc").get_by_role("link", name="Overview").click()
page.wait_for_selector('[data-testid="doc"][data-page="overview"]')
assert "#/docs?page=overview" in page.url, "a relative .md link must not leave the app"
# And back again, by the link the document itself carries.
page.get_by_test_id("doc").get_by_role("link", name="← Documentation index").click()
page.wait_for_selector('[data-testid="doc"][data-page="index"]')
assert quiet == []
def test_a_deep_link_to_a_heading_lands_on_that_heading(page, server, quiet):
page.goto(f"{server.base}/app/#/docs?page=overview&anchor=the-workflow")
heading = page.locator("#the-workflow")
heading.wait_for()
assert heading.inner_text().strip() == "The workflow"
assert heading.evaluate("node => node.getBoundingClientRect().top < window.innerHeight")
assert quiet == []
def test_a_mermaid_block_becomes_a_diagram_under_the_unchanged_script_policy(page, server, quiet):
page.goto(f"{server.base}/app/#/docs?page=overview")
diagram = page.get_by_test_id("diagram").first
diagram.wait_for()
# A real drawing, not the source text and not an error node.
assert diagram.locator("svg").count() == 1
assert diagram.evaluate("node => node.querySelector('svg').getBBox().width") > 100
assert page.locator("pre code.language-mermaid").count() == 0
assert quiet == [], "rendering a diagram must not violate the policy"
policy = page.evaluate(
"async () => (await fetch('/app/')).headers.get('content-security-policy')"
)
assert "script-src 'self';" in policy
assert "unsafe-eval" not in policy
def test_every_diagram_in_the_architecture_overview_draws(page, server, quiet):
"""Five diagrams, three of them state machines (US09-03). A mermaid block with a
syntax error renders an error node instead of throwing, so a page that merely
loaded proves nothing — count the drawings and read the source for the states."""
source = (Path(__file__).resolve().parents[2] / "docs" / "architecture.md").read_text()
expected = source.count("```mermaid")
assert expected >= 5, "the overview lost its diagrams"
page.goto(f"{server.base}/app/#/docs?page=architecture")
page.get_by_test_id("diagram").first.wait_for()
page.wait_for_function(
"count => document.querySelectorAll('[data-testid=\"diagram\"] svg').length === count",
arg=expected,
)
drawn = page.locator('[data-testid="diagram"] svg')
assert drawn.count() == expected
for index in range(expected):
diagram = drawn.nth(index)
# Labels may be <text> or HTML inside a <foreignObject> depending on the
# diagram type, so ask for the rendered text rather than a specific element.
labels = diagram.evaluate("node => node.textContent.trim().length")
assert labels > 0, f"diagram {index} drew no labels"
assert diagram.locator(".error-icon, .error-text").count() == 0, f"diagram {index} errored"
assert page.locator("pre code.language-mermaid").count() == 0
assert quiet == []
def test_an_unknown_page_says_so_without_naming_a_path(page, server, quiet):
page.goto(f"{server.base}/app/#/docs?page=no-such-manual")
message = page.get_by_test_id("doc-not-found")
message.wait_for()
text = message.inner_text()
assert "does not exist" in text
assert "/" not in text.replace("Back to the documentation index", ""), text
message.get_by_role("link").click()
page.wait_for_selector('[data-testid="doc"][data-page="index"]')
# The missing page is a 404 and the browser says so; nothing else may go wrong,
# and in particular the view must not throw on the way to its own message.
assert all("404" in problem for problem in quiet), quiet
def test_the_documentation_is_readable_without_a_session(page, server, quiet):
"""The troubleshooting page is needed most by whoever is locked out."""
page.goto(f"{server.base}/app/#/docs")
page.get_by_test_id("doc").wait_for()
unauthenticated = page.evaluate(
"""async () => {
const response = await fetch('/docs/index.md', { credentials: 'omit' });
return { status: response.status, length: (await response.text()).length };
}"""
)
assert unauthenticated["status"] == 200
assert unauthenticated["length"] > 100

View File

@@ -0,0 +1,164 @@
"""US09-03: the architecture overview, checked against the architecture.
An architecture document is the one that rots most quietly: nothing breaks when it
describes a module that was renamed two epics ago, it just quietly misleads the next
person. So the parts of it that the code also knows — module names, state machines,
table names, the modules an invariant is claimed to live in — are compared with the
code, and a missing one fails the suite.
The diagrams' rendering is proven in ``tests/e2e/test_docs_ui.py``.
"""
from __future__ import annotations
import re
from pathlib import Path
import photo_pipeline.models # noqa: F401 (registers every table on Base.metadata)
from photo_pipeline.db import Base
from photo_pipeline.services import jobs, rename_journal, upload_batches
REPO = Path(__file__).resolve().parents[2]
PACKAGE = REPO / "photo_pipeline"
OVERVIEW = REPO / "docs" / "architecture.md"
TEXT = OVERVIEW.read_text()
# Names the map does not owe the reader individually: the package markers, the CLI
# entry point, and the two modules that exist to be small and obvious.
UNMAPPED = {"__init__", "__main__", "logging"}
def modules() -> set[str]:
"""Every module and package under photo_pipeline, by the name it is imported as."""
names = set()
for path in PACKAGE.rglob("*.py"):
relative = path.relative_to(PACKAGE)
names.add(relative.parts[0] if len(relative.parts) > 1 else relative.stem)
return {name.removesuffix(".py") for name in names} - UNMAPPED
def states(container: type) -> set[str]:
return {
value
for name, value in vars(container).items()
if name.isupper() and isinstance(value, str)
}
# ── the module map ───────────────────────────────────────────────────────────
def test_every_module_appears_in_the_map():
"""A service nobody documented is a service the next person re-implements."""
missing = sorted(name for name in modules() if name not in TEXT)
assert missing == [], f"absent from the architecture overview: {missing}"
def test_every_service_is_named_individually():
"""The package table says what `services/` is for; this is the list that actually
drifts, because a new service is added roughly every story."""
services = {path.stem for path in (PACKAGE / "services").glob("*.py")} - UNMAPPED
missing = sorted(name for name in services if not re.search(rf"\b{name}\b", TEXT))
assert missing == [], f"services missing from the overview: {missing}"
def test_the_map_names_no_module_that_does_not_exist():
"""Every `services/<name>` or backticked module path in the document resolves."""
referenced = set(re.findall(r"`(?:photo_pipeline/)?([a-z_]+)/([a-z_]+)\.py`", TEXT))
referenced |= {("", name) for name in re.findall(r"`([a-z_]+)\.py`", TEXT)}
missing = [
f"{package}/{module}.py" if package else f"{module}.py"
for package, module in referenced
if not (PACKAGE / package / f"{module}.py").is_file()
]
assert missing == [], f"documented but absent from the source: {missing}"
# ── the state machines ───────────────────────────────────────────────────────
def test_every_job_state_is_documented():
missing = sorted(state for state in states(jobs.JobState) if state not in TEXT)
assert missing == [], f"job states missing from the overview: {missing}"
def test_every_rename_journal_state_is_documented():
missing = sorted(state for state in states(rename_journal.JournalState) if state not in TEXT)
assert missing == [], f"journal states missing from the overview: {missing}"
def test_every_upload_batch_state_is_documented():
missing = sorted(state for state in states(upload_batches.BatchState) if state not in TEXT)
assert missing == [], f"batch states missing from the overview: {missing}"
def test_the_unsafe_journal_states_are_named_as_the_ones_that_block():
"""The reason an unrelated mutation is refused has to be findable."""
for state in rename_journal.UNSAFE_STATES:
assert state in TEXT
assert "rename_recovery_required" in TEXT
def test_an_uncertain_upload_is_documented_as_not_retryable():
assert upload_batches.BatchState.UNKNOWN not in upload_batches.RUNNABLE_STATES
assert "not** restartable" in TEXT or "not restartable" in TEXT
# ── the data model ───────────────────────────────────────────────────────────
def test_every_table_is_listed():
missing = sorted(name for name in Base.metadata.tables if name not in TEXT)
assert missing == [], f"tables missing from the overview: {missing}"
def test_the_document_lists_no_table_that_does_not_exist():
listed = set(re.findall(r"`([a-z_]+)`", TEXT))
plausible = {name for name in listed if name.endswith("s") and "_" in name}
invented = sorted(
name
for name in plausible
if name not in Base.metadata.tables
and not (PACKAGE / "services" / f"{name}.py").is_file()
and name not in {"library_roots", "trusted_proxies", "allowed_hosts", "asset_paths"}
)
assert invented == [], f"looks like a table but is not one: {invented}"
# ── the invariants ───────────────────────────────────────────────────────────
def test_each_invariant_points_at_a_module_that_enforces_it():
"""The point of the table is to answer 'where do I look'. A wrong answer there
costs more than no answer."""
claims = {
"path_policy.is_excluded": PACKAGE / "path_policy.py",
"path_policy.resolve_in_roots": PACKAGE / "path_policy.py",
"services/app_lock.py": PACKAGE / "services" / "app_lock.py",
"services/exif_checkpoint.py": PACKAGE / "services" / "exif_checkpoint.py",
"services/archive_transfer.py": PACKAGE / "services" / "archive_transfer.py",
"api/security.py": PACKAGE / "api" / "security.py",
"imaging.py": PACKAGE / "imaging.py",
}
for claim, path in claims.items():
assert claim in TEXT, f"the invariant table does not mention {claim}"
assert path.is_file(), f"{claim} does not exist"
function = claim.rpartition(".")[2] if "/" not in claim else ""
if function and function != "py":
assert f"def {function}" in path.read_text(), f"{claim} is not defined there"
def test_the_lock_order_matches_the_ranks_in_the_code():
from photo_pipeline.jobs import locks
order = [name for name, _ in sorted(locks.LOCK_RANK.items(), key=lambda item: item[1])]
assert order[0].startswith("library"), "the broadest lock is no longer the library"
assert "library → stage/job → album/folder → asset" in TEXT
def test_the_frozen_donor_archive_is_described_as_provenance_not_a_dependency():
"""That production never imports the frozen sources is proven by
``tests/unit/test_legacy_archive.py`` — which also forbids any other test from
naming that directory, so this one checks the claim by its consequence."""
assert "must not import" in TEXT
assert "provenance and rollback evidence" in TEXT

View File

@@ -0,0 +1,179 @@
"""US09-01: the documentation the application serves, checked without a browser.
What is checkable offline is everything that makes the manuals *navigable* rather
than merely present: every page reachable, every internal link and anchor landing
somewhere real, the vendored renderer being the file that was pinned, and the
application actually serving `docs/` in both the working copy and the image.
The rendering itself — marked, mermaid, anchors, link rewriting — is proven in a
real browser by ``tests/e2e/test_docs_ui.py``.
"""
from __future__ import annotations
import hashlib
import json
import re
from pathlib import Path
from starlette.testclient import TestClient
from photo_pipeline.api.app import create_app
from photo_pipeline.api.security import DEFAULT_HEADERS
from photo_pipeline.config import Config
REPO = Path(__file__).resolve().parents[2]
DOCS = REPO / "docs"
VENDOR = REPO / "frontend" / "js" / "vendor"
INDEX = DOCS / "index.md"
MARKDOWN_LINK = re.compile(r"\[[^\]]*\]\(([^)\s]+)\)")
HEADING = re.compile(r"^#{1,6}\s+(.+?)\s*$", re.MULTILINE)
FENCE = re.compile(r"^```.*?^```", re.MULTILINE | re.DOTALL)
def pages() -> list[Path]:
return sorted(DOCS.rglob("*.md"))
def body(path: Path) -> str:
"""A document without its fenced code, so an example link in a shell snippet is
not mistaken for a link the reader can follow."""
return FENCE.sub("", path.read_text())
def slug(text: str) -> str:
"""The heading-anchor rule, mirroring ``frontend/js/docs.js``.
Deliberately duplicated: this check has to run without a browser. The browser
test is the authority on the real behaviour, and it asserts the same anchors.
"""
cleaned = re.sub(r"[^\w\s-]", "", text.lower(), flags=re.UNICODE).strip()
return re.sub(r"[\s-]+", "-", cleaned).strip("-") or "section"
def anchors(path: Path) -> set[str]:
found: dict[str, int] = {}
for heading in HEADING.findall(body(path)):
# Markdown emphasis and inline code are not part of the rendered text.
plain = re.sub(r"[*`_]", "", heading)
base = slug(plain)
found[base] = found.get(base, 0) + 1
return {name if index == 1 else f"{name}-{index}" for name, count in found.items() for index in range(1, count + 1)}
def internal_links(path: Path) -> list[str]:
return [
href
for href in MARKDOWN_LINK.findall(body(path))
if not re.match(r"^[a-z][a-z\d+.-]*:", href, re.IGNORECASE) and not href.startswith("//")
]
# ── the documentation tree ───────────────────────────────────────────────────
def test_the_index_exists_and_names_every_page():
"""A page nobody links to is a page nobody reads."""
assert INDEX.is_file()
listed = {
(INDEX.parent / href.split("#")[0]).resolve()
for href in internal_links(INDEX)
if href.split("#")[0].endswith(".md")
}
unreachable = [page.name for page in pages() if page != INDEX and page.resolve() not in listed]
assert unreachable == [], f"not linked from the index: {unreachable}"
def test_every_internal_link_and_anchor_resolves():
broken: list[str] = []
for page in pages():
for href in internal_links(page):
target, _, anchor = href.partition("#")
destination = page if not target else (page.parent / target).resolve()
if not destination.is_file():
broken.append(f"{page.name}{href} (no such file)")
continue
if anchor and anchor not in anchors(destination):
broken.append(f"{page.name}{href} (no such heading)")
assert broken == [], broken
def test_every_page_starts_with_one_title():
for page in pages():
titles = [line for line in page.read_text().splitlines() if line.startswith("# ")]
assert len(titles) == 1, f"{page.name} has {len(titles)} top-level titles"
# ── the vendored renderer ────────────────────────────────────────────────────
def test_the_vendored_libraries_are_the_files_that_were_pinned():
"""A vendored dependency without a recorded checksum is a dependency nobody is
reviewing. Replacing one has to be a visible change to this manifest."""
manifest = json.loads((VENDOR / "VERSIONS.json").read_text())
recorded = {entry["file"]: entry for entry in manifest["libraries"]}
on_disk = {path.name for path in VENDOR.iterdir() if path.suffix in (".js", ".mjs")}
assert on_disk == set(recorded), f"unrecorded vendored files: {on_disk ^ set(recorded)}"
for name, entry in recorded.items():
digest = hashlib.sha256((VENDOR / name).read_bytes()).hexdigest()
assert digest == entry["sha256"], f"{name} does not match the pinned checksum"
assert entry["version"] in entry["url"], f"{name}: the pinned URL and version disagree"
assert entry["license"], f"{name}: no license recorded"
def test_the_documentation_view_is_wired_into_the_shell():
shell = (REPO / "frontend" / "index.html").read_text()
assert 'data-nav="docs"' in shell, "the manuals are unreachable from the navigation"
assert '"/docs"' in (REPO / "frontend" / "js" / "app.js").read_text()
# ── the policy the view runs under ───────────────────────────────────────────
def test_the_script_boundary_is_unchanged_and_only_style_was_relaxed():
"""The measured cost of rendering diagrams in the browser, held to that cost.
mermaid needs to style the SVG it builds, which `style-src 'self'` refuses. It
does not need to execute generated code, so `script-src` must never acquire the
escape hatch that would let injected markup run.
"""
policy = DEFAULT_HEADERS["content-security-policy"]
directives = {
part.split(" ")[0]: part.split(" ")[1:]
for part in (piece.strip() for piece in policy.split(";"))
if part
}
assert directives["script-src"] == ["'self'"]
assert "'unsafe-inline'" in directives["style-src"]
for directive in ("default-src", "img-src", "connect-src", "font-src"):
assert "'self'" in directives[directive]
assert "'unsafe-inline'" not in directives[directive]
assert directives["frame-ancestors"] == ["'none'"]
# ── serving it ───────────────────────────────────────────────────────────────
def test_the_application_serves_the_markdown_without_a_session(tmp_path):
"""Whoever cannot get past the access secret is exactly who needs these pages."""
config = Config.from_env({"PHOTO_PIPELINE_DATA_DIR": str(tmp_path / "data")})
(tmp_path / "data").mkdir()
with TestClient(create_app(config), raise_server_exceptions=False) as client:
response = client.get("/docs/index.md", headers={"cookie": ""})
assert response.status_code == 200
assert "Photo Pipeline documentation" in response.text
assert client.get("/docs/overview.md").status_code == 200
assert client.get("/docs/nothing-here.md").status_code == 404
# The mount is a directory, not a path parameter: nothing above it is reachable.
assert client.get("/docs/../pyproject.toml").status_code in (307, 404)
assert client.get("/docs/%2e%2e/pyproject.toml").status_code == 404
def test_the_image_ships_the_documentation_it_serves():
"""An image without `docs/` serves an empty manual — and the build context is
deny-by-default, so a new directory is excluded until it is named."""
assert "\n!docs\n" in (REPO / ".dockerignore").read_text()
assert "COPY docs ./docs" in (REPO / "Dockerfile").read_text()

View File

@@ -0,0 +1,178 @@
"""US09-02: the installation manual, checked against the thing it describes.
Documentation rots quietly. A setting is renamed, a command grows a flag, an exit
code changes meaning — and the manual keeps confidently saying the old thing until
somebody follows it into a bad afternoon. So every fact in it that the code also
knows is compared with the code, in both directions where both directions are
meaningful: a setting the manual invents fails, and a setting the manual forgot
fails too.
"""
from __future__ import annotations
import re
import subprocess
import sys
from functools import lru_cache
from pathlib import Path
from photo_pipeline.config import ENV_PREFIX, LEGACY_ALIASES, Config
REPO = Path(__file__).resolve().parents[2]
MANUAL = REPO / "docs" / "installation.md"
MAIN = REPO / "photo_pipeline" / "__main__.py"
TEXT = MANUAL.read_text()
ENV_NAMES = re.compile(rf"{ENV_PREFIX}[A-Z0-9_]+")
# Read by docker-compose.yml, not by Config. They are configuration an installer sets,
# so the manual documents them; they are simply not application settings.
COMPOSITION_ONLY = {
f"{ENV_PREFIX}IMAGE",
f"{ENV_PREFIX}LIBRARY_HOST_PATH",
f"{ENV_PREFIX}PUBLISH_ADDRESS",
f"{ENV_PREFIX}UID",
f"{ENV_PREFIX}GID",
f"{ENV_PREFIX}ENV_FILE", # read before Config exists, so it is not a field
}
def documented_env_names() -> set[str]:
return set(ENV_NAMES.findall(TEXT))
@lru_cache
def cli_help(*args: str) -> str:
"""What the CLI says about itself, asked the way an operator asks."""
result = subprocess.run(
[sys.executable, "-m", "photo_pipeline", *args, "--help"],
cwd=REPO,
capture_output=True,
text=True,
timeout=120,
check=False,
)
assert result.returncode == 0, result.stderr
return result.stdout
@lru_cache
def cli_commands() -> frozenset[str]:
listed = re.search(r"\{([a-z0-9,\-]+)\}", cli_help())
assert listed, f"the CLI listed no subcommands:\n{cli_help()}"
return frozenset(listed.group(1).split(","))
def cli_flags(command: str) -> frozenset[str]:
return frozenset(re.findall(r"--[a-z][a-z-]+", cli_help(command)))
# ── settings ─────────────────────────────────────────────────────────────────
def test_every_setting_is_documented():
"""A setting nobody documents is a setting nobody configures deliberately."""
expected = {f"{ENV_PREFIX}{field.upper()}" for field in Config.model_fields}
missing = sorted(expected - documented_env_names())
assert missing == [], f"undocumented settings: {missing}"
def test_the_manual_invents_no_settings():
unknown = sorted(
name
for name in documented_env_names()
if name.removeprefix(ENV_PREFIX).lower() not in Config.model_fields
and name not in COMPOSITION_ONLY
)
assert unknown == [], f"documented but not a real setting: {unknown}"
def test_the_documented_defaults_are_the_real_defaults():
"""Spot-checked where a wrong default is actively dangerous: the bind address,
the trust boundary, and the gate that protects an irreplaceable library."""
assert Config.model_fields["host"].default == "127.0.0.1"
assert "`PHOTO_PIPELINE_HOST` | `127.0.0.1`" in TEXT
assert Config.model_fields["allowed_hosts"].default == ()
assert Config.model_fields["require_dry_run_approval"].default is False
assert "`PHOTO_PIPELINE_REQUIRE_DRY_RUN_APPROVAL` | `false`" in TEXT
def test_every_secret_is_marked_as_one():
for field in ("access_secret", "vision_api_key", "immich_api_key"):
name = f"{ENV_PREFIX}{field.upper()}"
row = next(line for line in TEXT.splitlines() if line.startswith(f"| `{name}`"))
assert "secret" in row.lower(), f"{name} is not marked as a secret"
def test_the_legacy_aliases_are_documented():
"""An operator with an existing photo_analyzer.env needs to know it still works."""
for alias in LEGACY_ALIASES:
assert alias in TEXT, f"legacy alias {alias} is undocumented"
# ── commands ─────────────────────────────────────────────────────────────────
# The manual is an installation and operations manual, not a command reference: these
# are the commands an installer actually runs, and every one of them must exist.
OPERATIONAL = ("migrate", "serve", "worker", "backup", "verify-backup", "restore", "diagnostics")
def test_every_command_the_manual_tells_you_to_run_exists():
missing = [name for name in OPERATIONAL if name not in cli_commands()]
assert missing == [], f"the manual names commands the CLI does not have: {missing}"
undocumented = [name for name in OPERATIONAL if not re.search(rf"\b{re.escape(name)}\b", TEXT)]
assert undocumented == [], f"operational commands the manual omits: {undocumented}"
def test_every_documented_flag_exists_on_the_command_it_is_shown_with():
for command, flag in (
("backup", "--reason"),
("backup", "--keep"),
("restore", "--into"),
("worker", "--allow-legacy"),
("serve", "--allow-legacy"),
):
assert flag in cli_flags(command), f"{command} has no {flag}"
assert flag in TEXT, f"{flag} is undocumented"
# ── exit codes ───────────────────────────────────────────────────────────────
def test_every_documented_exit_code_is_one_the_cli_can_return():
documented = {int(code) for code in re.findall(r"^\| `(\d)` \|", TEXT, re.MULTILINE)}
returned = {int(code) for code in re.findall(r"^\s+return (\d)$", MAIN.read_text(), re.MULTILINE)}
assert documented, "no exit codes are documented"
assert documented <= returned | {0}, f"documented but unreachable: {sorted(documented - returned)}"
def test_every_refusal_exit_code_is_documented():
"""0 is success and 1 is 'it said why'; every other code is a specific refusal an
operator will meet at startup, and meeting an undocumented one is the worst case."""
returned = {int(code) for code in re.findall(r"^\s+return (\d)$", MAIN.read_text(), re.MULTILINE)}
documented = {int(code) for code in re.findall(r"^\| `(\d)` \|", TEXT, re.MULTILINE)}
missing = sorted(code for code in returned if code not in documented and code != 0)
assert missing == [], f"undocumented exit codes: {missing}"
# ── the promises the manual makes ────────────────────────────────────────────
def test_the_manual_carries_no_credential_shaped_example():
"""The repository's own scanner refuses these in a commit; a manual is exactly
where a real-looking one gets copied from."""
suspicious = re.findall(
r"(?i)(api[_-]?key|secret|token|password)\s*[=:]\s*[\"']?([A-Za-z0-9_\-]{12,})",
TEXT,
)
assert suspicious == [], f"credential-shaped example: {suspicious}"
def test_the_manual_states_the_invariants_an_installer_must_not_break():
for invariant in ("_IGNORE/", "One writer", "EXIF is verified before upload"):
assert invariant in TEXT, f"the manual does not state: {invariant}"
def test_the_manual_is_reachable_from_the_index():
assert "installation.md" in (REPO / "docs" / "index.md").read_text()

View File

@@ -193,12 +193,20 @@
"US08-05": [
"tests/integration/test_container_gate.py",
"tests/e2e/test_phase_h_container.py"
],
"US09-01": [
"tests/integration/test_documentation.py",
"tests/e2e/test_docs_ui.py"
],
"US09-02": [
"tests/integration/test_installation_manual.py"
],
"US09-03": [
"tests/integration/test_architecture_overview.py",
"tests/e2e/test_docs_ui.py"
]
},
"planned": [
"US09-01",
"US09-02",
"US09-03",
"US09-04",
"US09-05"
],