Compare commits

..

3 Commits

Author SHA1 Message Date
eae7466bb8 US09-02: Write the Installation and Operations Manual
Some checks failed
Test / suites (pull_request) Failing after 2m40s
Test / container (pull_request) Has been skipped
2026-08-23 21:45:09 +02:00
e508f9fd64 US09-01: Serve the Documentation Inside the Application (#107)
Some checks failed
Test / suites (push) Failing after 2m59s
Test / container (push) Failing after 3m56s
2026-08-23 21:09:43 +02:00
370f966d29 E09: Product documentation backlog (#101)
Some checks failed
Test / suites (push) Failing after 3m12s
Test / container (push) Failing after 4m23s
2026-08-23 15:04:45 +02:00
20 changed files with 5040 additions and 6 deletions

View File

@@ -9,6 +9,7 @@
!photo_pipeline
!migrations
!frontend
!docs
!docker
# Nothing generated, even under an allowed directory.

9
.gitattributes vendored Normal file
View File

@@ -0,0 +1,9 @@
# Vendored third-party browser bundles (US09-01). Minified upstream files carry
# trailing whitespace and very long lines; they are not ours to reformat, and the
# submit gate's `git diff --check` would refuse them forever. Their integrity is
# controlled where it belongs instead: a pinned version and a recorded sha256 in
# frontend/js/vendor/VERSIONS.json, asserted by tests/integration/test_documentation.py.
#
# Only the whitespace check is turned off. They stay text, so a credential scan or a
# search still reads them.
frontend/js/vendor/*.js -whitespace linguist-vendored

View File

@@ -70,6 +70,8 @@ COPY pyproject.toml alembic.ini README.md ./
COPY photo_pipeline ./photo_pipeline
COPY migrations ./migrations
COPY frontend ./frontend
# The manuals are served by the application itself (US09-01), so they ship with it.
COPY docs ./docs
COPY docker/entrypoint.sh docker/healthcheck.sh /usr/local/bin/
# Runtime dependencies only: the `test` extra (pytest, playwright) and the `vision`

34
docs/index.md Normal file
View File

@@ -0,0 +1,34 @@
# Photo Pipeline documentation
One local application that takes a photo library from discovery to a verified Immich
upload: duplicate detection, safety review, content analysis, album naming, guarded
renaming, upload, and archive — one visible, resume-safe workflow.
These pages are readable three ways, and they are the same files each time: in the
repository under `docs/`, on Gitea, and inside the running application under
**Docs**. There is no separate copy to fall out of date.
## Read in this order
1. [Overview](overview.md) — what the application does, the stages it moves a photo
through, and the rules it will not break.
2. [Installation and operations](installation.md) — host and container installation,
every setting, the first-run checklist, upgrades, backup and restore, and what
each refusal at startup means.
## Being written
The remaining manuals are accepted work, not aspiration; each is a story in
[E09](https://git.domverse-berlin.eu/domverse/photoanalyzer/src/branch/main/delivery_backlog/E09-documentation.md)
and will appear here as it lands.
- **Architecture** — context and runtime diagrams, the module map, the job and
journal state machines, and where each invariant is enforced (US09-03).
- **User manual** — one page per workflow stage with screenshots of the real
application, and a catalogue of every error and refusal (US09-04).
## Conventions
A page tells you what a stage **changes on disk or on the server** before it tells
you how to run it. Refusals are documented as intentional: this application would
rather stop and explain than guess about somebody's photographs.

336
docs/installation.md Normal file
View File

@@ -0,0 +1,336 @@
# Installation and operations
[← Documentation index](index.md)
From nothing to a running instance pointed at your photo library, and everything you
need afterwards: upgrading, backing up, restoring, and understanding a refusal.
Two ways to install. **Container** is the deployment this project builds and ships;
**host** is what you want for development or a single machine you already manage.
Both run the same two processes against the same database.
> Read [what it will not do](overview.md#what-it-will-not-do) first. Several of the
> steps below only make sense once you know which rules the application is keeping.
## Before you start
| you need | version | why |
|---|---|---|
| Python | 3.12 or newer | the application |
| `exiftool` | 13.x (the image pins `13.25+dfsg-1`) | every EXIF read and write |
| `immich-go` | 0.32.0 (pinned in the image) | the upload stage only |
| Docker + the compose plugin | any current | the container installation only |
| a vision provider key | — | the analysis stage only |
You also need a photo library you can afford to be wrong about. Take a backup of it
before pointing anything at it for the first time, and consider running with
[`PHOTO_PIPELINE_REQUIRE_DRY_RUN_APPROVAL`](#the-settings) on.
## Host installation
```bash
python3.12 -m venv .venv
.venv/bin/pip install -e ".[vision]" # drop [vision] for a review-only install
```
Configure it. Every setting is an environment variable; a `.env` file in the working
directory is read at startup:
```bash
cp .env.example .env && $EDITOR .env
```
`.env.example` lists every variable with its default and no values at all. For a
loopback installation you need exactly one:
```bash
PHOTO_PIPELINE_LIBRARY_ROOTS=/home/you/Pictures
```
Then migrate and start the two processes:
```bash
.venv/bin/python -m photo_pipeline migrate # create or upgrade the database
.venv/bin/python -m photo_pipeline serve # API and browser app on 127.0.0.1:8000
.venv/bin/python -m photo_pipeline worker # second terminal
```
**Both processes are required.** `serve` enqueues work and renders the application;
nothing is scanned, scored, analysed, renamed, uploaded, or archived without a
worker. Open <http://127.0.0.1:8000/app/> and confirm the workflow page loads.
### The configuration file
`.env`, or any path named by `PHOTO_PIPELINE_ENV_FILE`. It is parsed, never executed:
`KEY=value` lines, `#` comments, optional quotes, no interpolation and no `export`. A
configuration file that can run code is a configuration file that can be a
vulnerability.
**Anything already exported in the shell wins.** The file is your standing
configuration; the environment is the override for one run.
The archived CLI's names still work, so an existing `photo_analyzer.env` can be used
as it is:
| in the file | applied as |
|---|---|
| `LLM_API_KEY` / `GEMINI_API_KEY` | `OPENAI_API_KEY` |
| `LLM_BASE_URL` | `OPENAI_BASE_URL` |
| `LIBRARY` | `PHOTO_PIPELINE_LIBRARY_ROOTS` |
`.env` and `*.env` are gitignored and refused by the repository's safety checks. The
file holds a real key; it must never be committed.
## Container installation
```bash
cp .env.example .env && $EDITOR .env
docker compose up -d --build
```
The composition is one `serve` container, one `worker` container, a one-shot
`migrate` that both wait for, one bind-mounted library, and one named data volume.
Three variables it cannot start without:
| variable | is |
|---|---|
| `PHOTO_PIPELINE_LIBRARY_HOST_PATH` | the library on this host |
| `PHOTO_PIPELINE_LIBRARY_ROOTS` | where that library is mounted **inside** the container |
| `PHOTO_PIPELINE_ACCESS_SECRET` | required, because publishing a port means the app is reachable from outside the container |
Set `PHOTO_PIPELINE_UID` and `PHOTO_PIPELINE_GID` to the owner of the library: what
the containers rename and rewrite keeps that ownership.
### The data volume
The `data` volume holds the database, its write-ahead log, the thumbnail cache, the
operation journals, and the backups. **It must stay on a local filesystem.** SQLite
in WAL mode needs real local locking, so NFS, SMB, and network volume drivers do not
slow it down — they corrupt it. The library bind mount has no such restriction.
### Ports and reaching it
The API port is published to `127.0.0.1` unless `PHOTO_PIPELINE_PUBLISH_ADDRESS` says
otherwise. Being reachable *was* the authentication in earlier versions: whoever could
open the port owned the library. So the moment the application answers to anything but
loopback — a hostname, a proxy, `0.0.0.0` — the access secret becomes mandatory and
`serve` refuses to start without it rather than publishing your photographs.
Behind a reverse proxy:
```bash
PHOTO_PIPELINE_ALLOWED_HOSTS=photos.example.com
PHOTO_PIPELINE_ACCESS_SECRET=# python -c 'import secrets; print(secrets.token_urlsafe(32))'
PHOTO_PIPELINE_TRUSTED_PROXIES=10.0.0.2 # only the proxy's own address
```
`X-Forwarded-Proto` and `X-Forwarded-Host` are believed only from a trusted-proxy
address, so a client cannot declare its own origin. The browser asks for the secret
once per tab. Wrong secrets are rate-limited and logged with the caller's address
only. The health endpoints stay open so an orchestrator can restart the container;
nothing else is.
### Deployment
The stack is managed from git by Portainer, which redeploys on a webhook after CI
publishes an image built from `main`. Runtime secrets live in the Portainer stack's
environment rather than in the repository, so rotation happens in one place and
reading the repository discloses nothing.
## The settings
Every variable, its default, and whether it is a secret. `.env.example` is the same
list in copyable form.
### Library and data
| variable | default | meaning |
|---|---|---|
| `PHOTO_PIPELINE_LIBRARY_ROOTS` | none | the library boundary, `os.pathsep`-separated. No path outside these roots is ever read or written |
| `PHOTO_PIPELINE_DATA_DIR` | `data` | database, WAL, thumbnail cache, journals, backups. Never inside the library |
| `PHOTO_PIPELINE_DB_PATH` | `<data dir>/photo_pipeline.db` | the database file, if it must live elsewhere |
| `PHOTO_PIPELINE_ENV_FILE` | `.env` | where to read the configuration file from |
### Serving and the trust boundary
| variable | default | meaning |
|---|---|---|
| `PHOTO_PIPELINE_HOST` | `127.0.0.1` | bind address |
| `PHOTO_PIPELINE_PORT` | `8000` | port |
| `PHOTO_PIPELINE_ALLOWED_HOSTS` | none | comma-separated names the app answers to besides loopback. Naming one makes the access secret mandatory |
| `PHOTO_PIPELINE_ACCESS_SECRET` | none | **secret.** Traded for the session cookie at `GET /api/v1/session` |
| `PHOTO_PIPELINE_TRUSTED_PROXIES` | none | comma-separated peer addresses whose forwarded headers may be believed |
| `PHOTO_PIPELINE_MAX_REQUEST_BYTES` | `1048576` | largest request body accepted |
### Logging
| variable | default | meaning |
|---|---|---|
| `PHOTO_PIPELINE_LOG_LEVEL` | `INFO` | standard Python levels |
| `PHOTO_PIPELINE_LOG_FORMAT` | `json` | `json` or `text` |
### Limits
| variable | default | meaning |
|---|---|---|
| `PHOTO_PIPELINE_THUMBNAIL_CACHE_QUOTA_BYTES` | `500000000` | cache quota; over it is a diagnostics warning |
| `PHOTO_PIPELINE_THUMBNAIL_MAX_PIXELS` | `100000000` | refuse to decode anything larger. This is the decompression-bomb guard |
| `PHOTO_PIPELINE_ARCHIVE_FREE_SPACE_RESERVE_BYTES` | `1000000000` | free space an archive destination must keep beyond the transfer itself |
### The safety gate
| variable | default | meaning |
|---|---|---|
| `PHOTO_PIPELINE_REQUIRE_DRY_RUN_APPROVAL` | `false` | refuse every mutating request until a read-only dry run of this library has been produced and approved |
### External services
| variable | default | meaning |
|---|---|---|
| `PHOTO_PIPELINE_VISION_API_KEY` | none | **secret.** Without it the analysis stage cannot run |
| `PHOTO_PIPELINE_IMMICH_SERVER_URL` | none | the Immich server |
| `PHOTO_PIPELINE_IMMICH_API_KEY` | none | **secret.** |
| `PHOTO_PIPELINE_IMMICH_GO_BINARY` | `immich-go` | the uploader, found on `PATH` |
### Composition only
Read by `docker-compose.yml`, not by the application:
`PHOTO_PIPELINE_IMAGE`, `PHOTO_PIPELINE_LIBRARY_HOST_PATH`,
`PHOTO_PIPELINE_PUBLISH_ADDRESS`, `PHOTO_PIPELINE_UID`, `PHOTO_PIPELINE_GID`.
A secret is never written to a log, never returned by the API, never stored in the
database, and never recorded in a backup manifest — a manifest says `configured`, not
the value. Do not put a real key in an example, a ticket, or a screenshot.
## First run checklist
Not "it started" — verified.
1. **The database is at the current revision.**
`python -m photo_pipeline migrate` exits 0.
2. **The application answers.**
`curl -s localhost:8000/api/v1/health/ready` returns 200.
3. **The library was found.** Open the app, run a scan from the Inventory view, and
check the count against what you expect. Anything under `_IGNORE/` must be missing
from it — that is the exclusion working, not a bug.
4. **The worker is claiming.** The scan job reaches `succeeded`. If it stays
`queued`, no worker is running.
5. **Diagnostics are clean.**
`python -m photo_pipeline diagnostics` reports free space and an empty `warnings`.
6. **Nothing was modified.** The scan is read-only; your files' timestamps are
unchanged.
Only then point it at the whole library.
## Operations
Every operation is the same CLI, on a host or in the composition:
```bash
python -m photo_pipeline diagnostics
docker compose run --rm --no-deps api diagnostics
```
`--no-deps` keeps a one-off command from starting a second stack; `api` is only the
service it borrows the image and mounts from.
### Upgrading
1. `python -m photo_pipeline backup --reason before-upgrade`
2. Pull the new version (or `docker compose pull && docker compose up -d`).
3. Migrations run by themselves at startup, and a **pending schema change is
snapshotted first**. If the upgrade fails, the previous database and its
`pre-migration` backup are both intact, and the error log names the backup
directory.
An up-to-date database is not backed up again on every start.
### Backing up
```bash
python -m photo_pipeline backup --reason weekly --keep 7
python -m photo_pipeline verify-backup data/backups/<name>
```
Backups go through SQLite's online backup API, never a file copy: with WAL enabled
the `.db` file alone is missing every committed page still in the write-ahead log.
Each backup is a directory holding the snapshot and a `manifest.json` — schema
revision, SHA-256, row counts, the archive media the library depends on, and which
settings were configured.
`verify-backup` runs `PRAGMA integrity_check` **and** `PRAGMA foreign_key_check`,
compares the snapshot's SHA-256 against the manifest, and re-counts every table it
recorded. Bit rot, a truncated copy, and a "repaired" snapshot all fail it.
`--keep N` prunes the oldest and never the newest. The same is available at
`GET /api/v1/diagnostics`, `GET|POST /api/v1/backups`,
`GET /api/v1/backups/{name}/verify`, and `POST /api/v1/backups/prune`.
### Restoring
Restore is deliberately **not** an API call. It replaces the state of an
installation, so it belongs to a stopped one and a person at a terminal.
1. Stop `serve` and `worker`.
2. `python -m photo_pipeline verify-backup data/backups/<name>` — never restore an
unverified snapshot, and `restore` will refuse one anyway.
3. `python -m photo_pipeline restore data/backups/<name> --into /path/to/fresh-data`.
A target that already holds a database is refused; recovering in place means
moving the old data directory aside first.
4. Point `PHOTO_PIPELINE_DATA_DIR` at the restored directory and run `migrate`.
5. Run an inventory scan, so paths are reconciled against the real library.
6. Mount every archive location the manifest names before archiving again. The
database records where archived originals are; it does not contain them.
Practise this against a copy before you need it.
### Watching it
`diagnostics` reports the database, write-ahead log, thumbnail cache, uploader
reports, backups, and logs separately, with free space and warnings for low disk
(`disk_low`, `disk_critical`), a cache over its quota, a write-ahead log outgrowing
its database, and a legacy CLI writing the library.
Logs go to stdout in JSON by default — `docker compose logs -f worker`, or your
service manager's journal on a host. Uploader output is kept per attempt under the
data directory with the API key scrubbed line by line.
## When it refuses to start
Every refusal below is deliberate. The application would rather stop and explain than
guess about somebody's photographs.
| exit code | meaning | what to do |
|---|---|---|
| `1` | the operation failed and said why — an unverifiable backup, an occupied restore target, an unknown benchmark profile | read the message; nothing was changed |
| `2` | **the library lock is held.** Another `serve` or `worker` owns this data directory | stop the other process. The message names the holder; a lock whose process is gone is taken over automatically |
| `3` | **a legacy CLI looks active.** The frozen command-line tools are writing this library | stop them. `--allow-legacy` overrides and you own the outcome |
| `4` | **configuration refused.** The application is reachable beyond loopback and no access secret is set | set `PHOTO_PIPELINE_ACCESS_SECRET`, or bind to loopback only |
| `5` | **a library root is not mounted** (containers only) | fix the bind mount. Refusing is what stops the container writing into its own throwaway layer instead of your library |
Other things that look like faults and are not:
- **`403 host_not_allowed`** — the `Host` header is not a name in
`PHOTO_PIPELINE_ALLOWED_HOSTS`. Add the real hostname; do not add `*`.
- **`409 lock_held`** — a mutating job is already running. One at a time is what
makes a crash recoverable.
- **`409 rename_recovery_required`** — an interrupted rename is unresolved. Resolve
it in the Renames view; unrelated mutations stay blocked until then, on purpose.
- **A job stuck in `queued`** — no worker is running.
- **An empty scan** — check `PHOTO_PIPELINE_LIBRARY_ROOTS`, and in a container check
that the path is the *container-side* mount, not the host path.
## What an installer must not work around
- **One writer.** Do not run two workers, and do not run the frozen CLI beside the
application. Each is safe alone and destructive together — the lock is not
bureaucracy.
- **`_IGNORE/` is never read.** Do not "fix" the exclusion.
- **EXIF is verified before upload.** Do not skip the checkpoint to make a stage
finish; an unverified projection is exactly the case where the wrong bytes reach
Immich.
- **The data directory is not the library.** Keep them apart, and keep the data
directory on a local filesystem.
- **Secrets stay in the environment.** Not in the database, not in a log, not in a
commit.

77
docs/overview.md Normal file
View File

@@ -0,0 +1,77 @@
# Overview
[← Documentation index](index.md)
Photo Pipeline sorts a photo library. It finds duplicates before anything expensive
happens to them, asks a human which pictures may leave the machine, describes the
ones that may, proposes album names from what it found, renames folders under a
crash-safe journal, uploads through `immich-go`, and can archive a finished album off
active storage without ever forgetting it existed.
It runs locally. It talks to exactly three things outside itself: a vision provider,
an Immich server, and `exiftool` — and it will tell you before it uses any of them.
## The workflow
Each stage is a gate, not a tab. A later stage can always be looked at; its actions
stay disabled until what they depend on is true.
```mermaid
flowchart TD
I["0 · Inventory<br/>discover, hash, cluster duplicates"] --> S["1 · Safety<br/>score and human decision"]
S -->|sfw| A["2 · Analysis<br/>vision provider, EXIF checkpoint"]
S -->|nsfw| U
A --> B["3 · Albums<br/>proposal, then guarded rename"]
B --> U["4 · Upload<br/>immich-go, verified bytes"]
U --> R["5 · Archive<br/>copy, verify, then reclaim space"]
```
| Stage | What it decides | What it changes |
|---|---|---|
| Inventory | which file is the canonical copy of a picture | nothing — it only reads |
| Safety | whether a photo may be sent to a cloud provider | one `sfw`/`nsfw` EXIF keyword |
| Analysis | what a photo shows | a managed caption segment and additive keywords |
| Albums | what a folder should be called | folder names, through a journaled rename |
| Upload | which exact bytes reach Immich | nothing locally; assets appear in Immich |
| Archive | which album leaves active storage | files move to the archive, after verification |
## What it will not do
These are enforced in code, not by convention, and each one is why some action you
expected is sometimes refused.
- **Nothing under `_IGNORE/` is ever read.** Not scanned, not counted, not
thumbnailed, not sent anywhere.
- **A path is not an identity.** Every picture has a stable id, so moving or renaming
it loses no history.
- **Only a confirmed-SFW photo reaches the vision provider.** A photo marked NSFW is
still uploadable to Immich; it simply never leaves for analysis.
- **Metadata is verified, not hoped for.** Every stage that writes EXIF reads it back
and proves that what it did not own is unchanged.
- **One writer at a time.** A library lock, held by one process, is what makes a
crash recoverable instead of ambiguous.
- **Nothing irreversible happens without a preview and an explicit approval** that
names the exact count.
## Where things live
| | |
|---|---|
| the photo library | wherever you point `PHOTO_PIPELINE_LIBRARY_ROOTS`; mounted read-write |
| the database, cache, logs, backups | the data directory, never inside the library |
| secrets | the environment, never the database and never a log line |
| these documents | `docs/` in the repository, served at `/docs` by the application |
## Running it
The short version, for a host installation:
```bash
python -m photo_pipeline migrate
python -m photo_pipeline serve # the API and this browser application
python -m photo_pipeline worker # the process that does the long work
```
The full procedure, the container path, and every setting belong to the installation
manual, which is being written next — see
[the documentation index](index.md#being-written).

View File

@@ -170,6 +170,20 @@ a.link:hover { text-decoration: underline; }
.grid th, .grid td { text-align: left; padding: 6px 8px; border-bottom: 1px solid var(--border); vertical-align: top; }
.path { font-family: ui-monospace, monospace; font-size: 0.85em; word-break: break-all; }
/* Docs: prose, so it gets a reading measure rather than the full window width. */
.doc { max-width: 72ch; }
.doc h1 { margin-top: 0; }
.doc h2, .doc h3 { margin-top: 1.6em; border-bottom: 1px solid var(--border); padding-bottom: 4px; }
.doc code { font-family: ui-monospace, monospace; font-size: 0.9em; background: var(--surface-2); padding: 1px 4px; border-radius: 4px; }
.doc pre { background: var(--surface-2); border: 1px solid var(--border); border-radius: var(--radius); padding: 12px; overflow: auto; }
.doc pre code { background: none; padding: 0; }
.doc table { border-collapse: collapse; width: 100%; }
.doc th, .doc td { text-align: left; padding: 6px 8px; border-bottom: 1px solid var(--border); vertical-align: top; }
.doc blockquote { margin: 1em 0; padding-left: 12px; border-left: 3px solid var(--border); color: var(--muted); }
.doc img { max-width: 100%; }
.diagram { margin: 1.5em 0; padding: 0; overflow: auto; }
.diagram svg { max-width: 100%; height: auto; }
/* Narrow screens: stack the panes so blockers and actions stay reachable. */
@media (max-width: 720px) {
.two-pane { grid-template-columns: 1fr; }

View File

@@ -21,6 +21,7 @@
<a href="#/uploads" data-nav="uploads">Upload</a>
<a href="#/archive" data-nav="archive">Archive</a>
<a href="#/stats" data-nav="stats">Stats</a>
<a href="#/docs" data-nav="docs">Docs</a>
</nav>
</header>
<main id="app" aria-live="polite"><!-- views render here --></main>

View File

@@ -1,5 +1,6 @@
import { api } from "./api.js";
import { renderArchive, setArchiveRender } from "./archive.js";
import { renderDocs } from "./docs.js";
import { navigate, onRouteChange, parseHash } from "./router.js";
import { renderRenames, setRenamesRender } from "./renames.js";
import { renderUploads, setUploadsRender } from "./uploads.js";
@@ -385,6 +386,7 @@ function render() {
else if (path === "/uploads") renderUploads(root, params);
else if (path === "/archive") renderArchive(root, params);
else if (path === "/stats") renderStats(root, params);
else if (path === "/docs") renderDocs(root, params);
else show(errorBanner("Unknown view"));
}

263
frontend/js/docs.js Normal file
View File

@@ -0,0 +1,263 @@
// The documentation view (US09-01): the manuals in `docs/`, rendered in the app.
//
// The same markdown files are the repository's documentation and the deployment's
// documentation. An operator who was handed a URL and an access secret has no
// checkout in front of them, and the network the application runs on is not assumed
// to reach a CDN — so the renderer is vendored and everything here is same-origin.
//
// Nothing on this page is authenticated. It is served from the static mount beside
// the application shell, exactly like `index.html`: the troubleshooting page is
// needed most by whoever cannot get past the access secret, and no documentation
// file contains anything a session would protect.
//
// `marked` and `mermaid` are loaded lazily, on the first documentation page and the
// first diagram. They are large, and every other view does without them.
import { el, errorBanner, setActiveNav } from "./dom.js";
const DOCS_BASE = "/docs";
const VENDOR = "/app/js/vendor";
const INDEX_PAGE = "index";
// A page name comes from the hash, so it is caller input: no scheme, no traversal,
// no absolute path. The server would refuse those too; this refuses them earlier and
// without a request that looks like an attempt.
const PAGE_PATTERN = /^[\w-]+(\/[\w-]+)*$/;
let markedPromise;
let mermaidPromise;
function loadMarked() {
markedPromise ??= import(`${VENDOR}/marked.esm.js`).then((module) => module.marked);
return markedPromise;
}
// mermaid ships one large UMD bundle rather than a self-contained ES module, so it
// arrives through a script element. `script-src 'self'` allows it because it is ours
// and same-origin; nothing here relaxes that.
function loadMermaid() {
mermaidPromise ??= new Promise((resolve, reject) => {
const script = document.createElement("script");
script.src = `${VENDOR}/mermaid.min.js`;
script.onload = () => resolve(window.mermaid);
script.onerror = () => reject(new Error("the diagram renderer could not be loaded"));
document.head.appendChild(script);
});
return mermaidPromise;
}
// ── the pure parts, unit-tested in js/tests/unit.js ──────────────────────────
/** A stable, readable heading anchor. Letters and digits of any script survive. */
export function slug(text) {
const cleaned = String(text)
.toLowerCase()
.trim()
.replace(/[^\p{L}\p{N}\s-]/gu, "")
.replace(/[\s-]+/g, "-")
.replace(/^-|-$/g, "");
return cleaned || "section";
}
/** Assign every heading an id, disambiguating repeats the way a reader would
* expect: the first `Notes` keeps `#notes`, the second becomes `#notes-2`. */
export function assignHeadingIds(headings) {
const used = new Map();
for (const heading of headings) {
const base = slug(heading.textContent);
const seen = (used.get(base) || 0) + 1;
used.set(base, seen);
heading.id = seen === 1 ? base : `${base}-${seen}`;
}
return headings;
}
/**
* Where a link inside a documentation page should go.
*
* Markdown links between documents are relative file paths, which a browser would
* treat as downloads that leave the application. Returns the in-app target for a
* link to another document or to a heading, and `null` for anything else — external
* links stay exactly as the author wrote them.
*/
export function resolveDocLink(currentPage, href) {
if (!href) return null;
if (/^[a-z][a-z\d+.-]*:/i.test(href) || href.startsWith("//")) return null;
if (href.startsWith("#")) return { page: currentPage, anchor: href.slice(1) };
const [target, anchor = ""] = href.split("#");
if (!/\.md$/i.test(target)) return null;
// Resolved against the current document, so `../guides/x.md` means what it means
// in the repository. The origin is a placeholder; only the path is used.
const resolved = new URL(target, `https://docs.invalid/${currentPage}.md`);
const page = decodeURIComponent(resolved.pathname).replace(/^\//, "").replace(/\.md$/i, "");
return PAGE_PATTERN.test(page) ? { page, anchor } : null;
}
/**
* Point every in-app link at its route, and leave every other link alone.
*
* The resolved page is kept on the element, so the reading order can be read back
* from the rendered index rather than parsed out of the markdown a second time.
*/
export function rewriteLinks(article, page) {
for (const link of article.querySelectorAll("a[href]")) {
const href = link.getAttribute("href");
const target = resolveDocLink(page, href);
if (target) {
link.setAttribute("href", docHash(target.page, target.anchor));
link.dataset.docPage = target.page;
} else if (/^https?:/i.test(href)) {
link.setAttribute("rel", "noreferrer noopener");
link.setAttribute("target", "_blank");
}
}
return article;
}
/** The reading order, taken from the index document itself — one list, in one
* place, that is equally the index on Gitea and the navigation here. */
export function documentIndex(indexElement) {
const pages = [];
for (const link of indexElement.querySelectorAll("a[data-doc-page]")) {
const page = link.dataset.docPage;
if (page !== INDEX_PAGE && !pages.some((entry) => entry.page === page)) {
pages.push({ page, title: link.textContent.trim() });
}
}
return pages;
}
// ── the view ─────────────────────────────────────────────────────────────────
async function fetchPage(page) {
const response = await fetch(`${DOCS_BASE}/${page}.md`, { headers: { accept: "text/markdown" } });
if (!response.ok) throw new Error(`documentation page unavailable: ${response.status}`);
return response.text();
}
async function toArticle(markdown, page) {
const marked = await loadMarked();
const article = el("article", { class: "doc", "data-testid": "doc", "data-page": page });
// The only innerHTML in the application, and deliberate: rendering markdown *is*
// producing HTML. The input is a file from this repository, and the page's CSP
// (`script-src 'self'`, no `'unsafe-inline'`) means injected script and inline
// handlers do not run even if one ever were not.
article.innerHTML = marked.parse(markdown, { async: false });
assignHeadingIds(article.querySelectorAll("h1, h2, h3, h4, h5, h6"));
rewriteLinks(article, page);
return article;
}
function docHash(page, anchor = "") {
const query = new URLSearchParams(anchor ? { page, anchor } : { page });
return `#/docs?${query}`;
}
/** Diagrams are authored as ```mermaid blocks so the source is diffable and Gitea
* renders them natively. A failure here replaces the diagram, never the page. */
async function renderDiagrams(article) {
const blocks = [...article.querySelectorAll("pre > code.language-mermaid")];
if (!blocks.length) return;
let mermaid;
try {
mermaid = await loadMermaid();
mermaid.initialize({ startOnLoad: false, securityLevel: "strict", theme: "dark" });
} catch (error) {
blocks.forEach((block) => block.closest("pre").replaceWith(errorBanner(error.message)));
return;
}
for (const [index, block] of blocks.entries()) {
const figure = el("figure", { class: "diagram", "data-testid": "diagram" });
block.closest("pre").replaceWith(figure);
try {
const { svg } = await mermaid.render(`diagram-${index}-${Date.now()}`, block.textContent);
figure.innerHTML = svg;
} catch (error) {
figure.replaceWith(errorBanner(`Diagram could not be drawn: ${error.message}`));
}
}
}
function notFound() {
return el(
"div",
{ class: "alert", role: "alert", "data-testid": "doc-not-found" },
"That documentation page does not exist. ",
el("a", { class: "link", href: docHash(INDEX_PAGE) }, "Back to the documentation index")
);
}
function sidebar(pages, current) {
return el(
"nav",
{ class: "card", "aria-label": "Documentation" },
el("h2", {}, "Documentation"),
el(
"ul",
{ class: "album-list", "data-testid": "doc-pages" },
el(
"li",
{},
el(
"a",
{
href: docHash(INDEX_PAGE),
"aria-current": current === INDEX_PAGE ? "true" : false,
},
"Index"
)
),
...pages.map(({ page, title }) =>
el(
"li",
{},
el(
"a",
{
href: docHash(page),
"data-page": page,
"aria-current": page === current ? "true" : false,
},
title
)
)
)
)
);
}
export async function renderDocs(root, params = {}) {
setActiveNav("docs");
const requested = params.page || INDEX_PAGE;
const page = PAGE_PATTERN.test(requested) ? requested : "";
let indexArticle;
try {
indexArticle = await toArticle(await fetchPage(INDEX_PAGE), INDEX_PAGE);
} catch (error) {
root.replaceChildren(errorBanner(`Documentation is unavailable: ${error.message}`));
return;
}
const pages = documentIndex(indexArticle);
let article;
if (!page) article = notFound();
else if (page === INDEX_PAGE) article = indexArticle;
else {
try {
article = await toArticle(await fetchPage(page), page);
} catch {
article = notFound();
}
}
root.replaceChildren(el("div", { class: "two-pane" }, sidebar(pages, page), article));
await renderDiagrams(article);
scrollToAnchor(params.anchor);
}
// A documentation anchor cannot live in the hash — the hash is the route — so it
// travels as a parameter and is applied after the page renders.
function scrollToAnchor(anchor) {
if (!anchor) return;
const target = document.getElementById(anchor);
if (target) target.scrollIntoView({ block: "start" });
}

View File

@@ -7,6 +7,7 @@ import { createStore } from "../store.js";
import { parseHash, navigate } from "../router.js";
import { api, cancellable } from "../api.js";
import { subscribeJob } from "../events.js";
import { assignHeadingIds, documentIndex, resolveDocLink, rewriteLinks, slug } from "../docs.js";
const cases = [];
function ok(name, cond) {
@@ -142,6 +143,62 @@ async function run() {
window.EventSource = realES;
}
// ── documentation view (US09-01) ─────────────────────────────────────────
{
ok("slug lowercases and dashes a heading", slug("Key Flows") === "key-flows");
ok(
"slug drops punctuation but keeps words",
slug("What it *will not* do:") === "what-it-will-not-do"
);
ok("slug keeps non-ASCII letters", slug("Größe & Gewicht") === "größe-gewicht");
ok("slug never yields an empty anchor", slug("!!!") === "section");
const doc = document.createElement("div");
doc.innerHTML = "<h2>Notes</h2><h3>Notes</h3><h2>Notes</h2>";
const ids = [...assignHeadingIds(doc.querySelectorAll("h2, h3"))].map((h) => h.id);
ok("repeated headings get distinct anchors", ids.join(",") === "notes,notes-2,notes-3");
ok(
"a sibling document link becomes an in-app route",
resolveDocLink("index", "overview.md")?.page === "overview"
);
const nested = resolveDocLink("guides/install", "../overview.md#running-it");
ok("a relative link resolves against the current page", nested?.page === "overview");
ok("a link's anchor survives the rewrite", nested?.anchor === "running-it");
ok(
"a bare anchor stays on the current page",
resolveDocLink("overview", "#the-workflow")?.page === "overview"
);
ok("an external link is left alone", resolveDocLink("index", "https://example.test/x") === null);
ok("a non-markdown relative link is left alone", resolveDocLink("index", "images/a.png") === null);
// A climb cannot leave the documentation tree: it is clamped at the root, so the
// page it names is still fetched from under /docs and simply does not exist.
ok(
"a link that climbs above the docs root is clamped",
resolveDocLink("index", "../../etc/passwd.md")?.page === "etc/passwd"
);
ok(
"a page name that is not a page name is refused",
resolveDocLink("index", "..%2f..%2fetc%2fpasswd.md") === null
);
const index = document.createElement("div");
index.innerHTML =
'<a href="overview.md">Overview</a><a href="overview.md">again</a>' +
'<a href="index.md">itself</a><a href="https://example.test">out</a>';
const listed = documentIndex(rewriteLinks(index, "index"));
ok(
"an external link is marked safe to open away from the app",
index.querySelector('a[href^="https"]').rel === "noreferrer noopener"
);
ok(
"an in-app link points at the docs route",
index.querySelector("a").getAttribute("href") === "#/docs?page=overview"
);
ok("the index lists each page once, in order", listed.length === 1);
ok("the index takes its titles from the link text", listed[0].title === "Overview");
}
await tick();
const failed = cases.filter((c) => !c.ok);
window.__RESULTS__ = { passed: cases.length - failed.length, failed: failed.length, cases };

23
frontend/js/vendor/VERSIONS.json vendored Normal file
View File

@@ -0,0 +1,23 @@
{
"_comment": "Vendored browser libraries (US09-01). Committed rather than fetched: the application is deployed to a network whose outbound access is not assumed, and `default-src 'self'` forbids a CDN. Checksums are asserted by tests/integration/test_documentation.py, so replacing a file without recording it here fails the suite. Update by downloading the pinned URL and recording the new version and sha256 in the same commit.",
"libraries": [
{
"file": "marked.esm.js",
"name": "marked",
"version": "18.0.10",
"license": "MIT",
"url": "https://cdn.jsdelivr.net/npm/marked@18.0.10/lib/marked.esm.js",
"sha256": "4cf47dfebb7f614a08fc0a579ab0fe407ff0ed2b717bf953040c85b2f493a4f0",
"why": "Markdown to HTML for the documentation view. An ES module with no dependencies, imported lazily by js/docs.js."
},
{
"file": "mermaid.min.js",
"name": "mermaid",
"version": "11.17.0",
"license": "MIT",
"url": "https://cdn.jsdelivr.net/npm/mermaid@11.17.0/dist/mermaid.min.js",
"sha256": "8d8e0eec56d3a83b4b3c87f42050845546dee93ebe1875d2117c12e6947c0cb3",
"why": "Renders ```mermaid blocks so diagrams are diffable source that Gitea also renders. The single-file UMD build rather than the ES module entry, whose ~40 lazy chunks would each need vendoring and pinning."
}
]
}

78
frontend/js/vendor/marked.esm.js vendored Normal file

File diff suppressed because one or more lines are too long

3636
frontend/js/vendor/mermaid.min.js vendored Normal file

File diff suppressed because one or more lines are too long

View File

@@ -53,6 +53,7 @@ from photo_pipeline.services.thumbnails import ThumbnailService
from photo_pipeline.services.upload_batches import UploadBatchService
FRONTEND_DIR = Path(__file__).resolve().parents[2] / "frontend"
DOCS_DIR = Path(__file__).resolve().parents[2] / "docs"
log = logging.getLogger(__name__)
@@ -151,4 +152,10 @@ def create_app(config: Config | None = None) -> FastAPI:
# Static single-page app (hash-routed). Mounted last so /api/v1 wins.
if FRONTEND_DIR.is_dir():
app.mount("/app", StaticFiles(directory=FRONTEND_DIR, html=True), name="app")
# The manuals, as their own markdown files (US09-01). Unauthenticated for the
# same reason the application shell is: whoever cannot get past the access
# secret is exactly who needs the troubleshooting page, and no document holds
# anything a session would protect.
if DOCS_DIR.is_dir():
app.mount("/docs", StaticFiles(directory=DOCS_DIR), name="docs")
return app

View File

@@ -57,9 +57,21 @@ LOOPBACK_HOSTS = frozenset({"127.0.0.1", "localhost", "::1", "[::1]"})
# backup a careful operator takes first, and its retention (US07-07).
MUTATION_EXEMPT_PATHS = frozenset({f"{API_PREFIX}/backups", f"{API_PREFIX}/backups/prune"})
# Applied to every response. No inline script/style is used by the frontend, so the
# policy can stay strict; `frame-ancestors 'none'` and CORP keep other pages from
# Applied to every response. `frame-ancestors 'none'` and CORP keep other pages from
# embedding the app or its thumbnails.
#
# `script-src 'self'` is the boundary that matters and it is unchanged: no
# `'unsafe-eval'`, no `'unsafe-inline'`, so nothing injected into the DOM can execute.
# Whether the vendored diagram renderer needed `'unsafe-eval'` was measured rather
# than assumed — mermaid's bundle contains no `eval(` and no `new Function`, and it
# renders under this exact policy without a single script-src violation.
#
# `style-src` does gain `'unsafe-inline'` (US09-01): mermaid styles the SVG it builds
# with an injected `<style>` element and `style=` attributes, and a diagram's CSS
# cannot be hashed in advance. The concession is bounded by the directives around it
# — with script execution still refused and `img-src`, `connect-src`, and `font-src`
# all `'self'`, the CSS exfiltration channels stay closed and what is left is
# defacement of a page its own operator is already looking at.
DEFAULT_HEADERS = {
"x-content-type-options": "nosniff",
"x-frame-options": "DENY",
@@ -67,8 +79,9 @@ DEFAULT_HEADERS = {
"cross-origin-resource-policy": "same-origin",
"cross-origin-opener-policy": "same-origin",
"content-security-policy": (
"default-src 'self'; img-src 'self' data:; style-src 'self'; script-src 'self'; "
"connect-src 'self'; frame-ancestors 'none'; base-uri 'none'; form-action 'none'"
"default-src 'self'; img-src 'self' data:; style-src 'self' 'unsafe-inline'; "
"script-src 'self'; connect-src 'self'; font-src 'self'; "
"frame-ancestors 'none'; base-uri 'none'; form-action 'none'"
),
}

119
tests/e2e/test_docs_ui.py Normal file
View File

@@ -0,0 +1,119 @@
"""US09-01: the manuals, in the browser, under the application's real policy.
Everything here runs against a real ``photo_pipeline serve`` process serving the
committed ``docs/`` tree and the vendored renderer — no fixture markdown, no stubbed
fetch. What that buys is the assertion the story actually cares about: the pages
render, the links between them work, a diagram becomes a diagram, and the console
stays empty, including of CSP violations.
The offline contract — reachability, dead links, pinned checksums — is
``tests/integration/test_documentation.py``.
"""
from __future__ import annotations
import pytest
from tests.e2e._pipeline_harness import Server, seed_library
@pytest.fixture(scope="module")
def server(tmp_path_factory):
seeded = seed_library(tmp_path_factory.mktemp("docs"), {"a": 1}, {})
running = Server(seeded).start()
try:
yield running
finally:
running.stop()
@pytest.fixture
def quiet(page):
"""Any console error, page error, or failed request fails the test that caused
it. A CSP violation arrives as a console error, which is the point."""
problems: list[str] = []
page.on("console", lambda m: problems.append(m.text) if m.type == "error" else None)
page.on("pageerror", lambda error: problems.append(str(error)))
page.on("requestfailed", lambda request: problems.append(f"failed request: {request.url}"))
yield problems
def test_the_documentation_opens_from_the_navigation(page, server, quiet):
page.goto(f"{server.base}/app/#/workflow")
page.locator('nav a[data-nav="docs"]').click()
page.get_by_test_id("doc").wait_for()
assert page.get_by_test_id("doc").get_attribute("data-page") == "index"
# The sidebar is the index's own reading order, not a second list to maintain.
assert page.get_by_test_id("doc-pages").get_by_role("link", name="Overview").is_visible()
assert quiet == []
def test_a_link_between_documents_stays_inside_the_application(page, server, quiet):
page.goto(f"{server.base}/app/#/docs")
page.get_by_test_id("doc").wait_for()
page.get_by_test_id("doc").get_by_role("link", name="Overview").click()
page.wait_for_selector('[data-testid="doc"][data-page="overview"]')
assert "#/docs?page=overview" in page.url, "a relative .md link must not leave the app"
# And back again, by the link the document itself carries.
page.get_by_test_id("doc").get_by_role("link", name="← Documentation index").click()
page.wait_for_selector('[data-testid="doc"][data-page="index"]')
assert quiet == []
def test_a_deep_link_to_a_heading_lands_on_that_heading(page, server, quiet):
page.goto(f"{server.base}/app/#/docs?page=overview&anchor=the-workflow")
heading = page.locator("#the-workflow")
heading.wait_for()
assert heading.inner_text().strip() == "The workflow"
assert heading.evaluate("node => node.getBoundingClientRect().top < window.innerHeight")
assert quiet == []
def test_a_mermaid_block_becomes_a_diagram_under_the_unchanged_script_policy(page, server, quiet):
page.goto(f"{server.base}/app/#/docs?page=overview")
diagram = page.get_by_test_id("diagram").first
diagram.wait_for()
# A real drawing, not the source text and not an error node.
assert diagram.locator("svg").count() == 1
assert diagram.evaluate("node => node.querySelector('svg').getBBox().width") > 100
assert page.locator("pre code.language-mermaid").count() == 0
assert quiet == [], "rendering a diagram must not violate the policy"
policy = page.evaluate(
"async () => (await fetch('/app/')).headers.get('content-security-policy')"
)
assert "script-src 'self';" in policy
assert "unsafe-eval" not in policy
def test_an_unknown_page_says_so_without_naming_a_path(page, server, quiet):
page.goto(f"{server.base}/app/#/docs?page=no-such-manual")
message = page.get_by_test_id("doc-not-found")
message.wait_for()
text = message.inner_text()
assert "does not exist" in text
assert "/" not in text.replace("Back to the documentation index", ""), text
message.get_by_role("link").click()
page.wait_for_selector('[data-testid="doc"][data-page="index"]')
# The missing page is a 404 and the browser says so; nothing else may go wrong,
# and in particular the view must not throw on the way to its own message.
assert all("404" in problem for problem in quiet), quiet
def test_the_documentation_is_readable_without_a_session(page, server, quiet):
"""The troubleshooting page is needed most by whoever is locked out."""
page.goto(f"{server.base}/app/#/docs")
page.get_by_test_id("doc").wait_for()
unauthenticated = page.evaluate(
"""async () => {
const response = await fetch('/docs/index.md', { credentials: 'omit' });
return { status: response.status, length: (await response.text()).length };
}"""
)
assert unauthenticated["status"] == 200
assert unauthenticated["length"] > 100

View File

@@ -0,0 +1,179 @@
"""US09-01: the documentation the application serves, checked without a browser.
What is checkable offline is everything that makes the manuals *navigable* rather
than merely present: every page reachable, every internal link and anchor landing
somewhere real, the vendored renderer being the file that was pinned, and the
application actually serving `docs/` in both the working copy and the image.
The rendering itself — marked, mermaid, anchors, link rewriting — is proven in a
real browser by ``tests/e2e/test_docs_ui.py``.
"""
from __future__ import annotations
import hashlib
import json
import re
from pathlib import Path
from starlette.testclient import TestClient
from photo_pipeline.api.app import create_app
from photo_pipeline.api.security import DEFAULT_HEADERS
from photo_pipeline.config import Config
REPO = Path(__file__).resolve().parents[2]
DOCS = REPO / "docs"
VENDOR = REPO / "frontend" / "js" / "vendor"
INDEX = DOCS / "index.md"
MARKDOWN_LINK = re.compile(r"\[[^\]]*\]\(([^)\s]+)\)")
HEADING = re.compile(r"^#{1,6}\s+(.+?)\s*$", re.MULTILINE)
FENCE = re.compile(r"^```.*?^```", re.MULTILINE | re.DOTALL)
def pages() -> list[Path]:
return sorted(DOCS.rglob("*.md"))
def body(path: Path) -> str:
"""A document without its fenced code, so an example link in a shell snippet is
not mistaken for a link the reader can follow."""
return FENCE.sub("", path.read_text())
def slug(text: str) -> str:
"""The heading-anchor rule, mirroring ``frontend/js/docs.js``.
Deliberately duplicated: this check has to run without a browser. The browser
test is the authority on the real behaviour, and it asserts the same anchors.
"""
cleaned = re.sub(r"[^\w\s-]", "", text.lower(), flags=re.UNICODE).strip()
return re.sub(r"[\s-]+", "-", cleaned).strip("-") or "section"
def anchors(path: Path) -> set[str]:
found: dict[str, int] = {}
for heading in HEADING.findall(body(path)):
# Markdown emphasis and inline code are not part of the rendered text.
plain = re.sub(r"[*`_]", "", heading)
base = slug(plain)
found[base] = found.get(base, 0) + 1
return {name if index == 1 else f"{name}-{index}" for name, count in found.items() for index in range(1, count + 1)}
def internal_links(path: Path) -> list[str]:
return [
href
for href in MARKDOWN_LINK.findall(body(path))
if not re.match(r"^[a-z][a-z\d+.-]*:", href, re.IGNORECASE) and not href.startswith("//")
]
# ── the documentation tree ───────────────────────────────────────────────────
def test_the_index_exists_and_names_every_page():
"""A page nobody links to is a page nobody reads."""
assert INDEX.is_file()
listed = {
(INDEX.parent / href.split("#")[0]).resolve()
for href in internal_links(INDEX)
if href.split("#")[0].endswith(".md")
}
unreachable = [page.name for page in pages() if page != INDEX and page.resolve() not in listed]
assert unreachable == [], f"not linked from the index: {unreachable}"
def test_every_internal_link_and_anchor_resolves():
broken: list[str] = []
for page in pages():
for href in internal_links(page):
target, _, anchor = href.partition("#")
destination = page if not target else (page.parent / target).resolve()
if not destination.is_file():
broken.append(f"{page.name}{href} (no such file)")
continue
if anchor and anchor not in anchors(destination):
broken.append(f"{page.name}{href} (no such heading)")
assert broken == [], broken
def test_every_page_starts_with_one_title():
for page in pages():
titles = [line for line in page.read_text().splitlines() if line.startswith("# ")]
assert len(titles) == 1, f"{page.name} has {len(titles)} top-level titles"
# ── the vendored renderer ────────────────────────────────────────────────────
def test_the_vendored_libraries_are_the_files_that_were_pinned():
"""A vendored dependency without a recorded checksum is a dependency nobody is
reviewing. Replacing one has to be a visible change to this manifest."""
manifest = json.loads((VENDOR / "VERSIONS.json").read_text())
recorded = {entry["file"]: entry for entry in manifest["libraries"]}
on_disk = {path.name for path in VENDOR.iterdir() if path.suffix in (".js", ".mjs")}
assert on_disk == set(recorded), f"unrecorded vendored files: {on_disk ^ set(recorded)}"
for name, entry in recorded.items():
digest = hashlib.sha256((VENDOR / name).read_bytes()).hexdigest()
assert digest == entry["sha256"], f"{name} does not match the pinned checksum"
assert entry["version"] in entry["url"], f"{name}: the pinned URL and version disagree"
assert entry["license"], f"{name}: no license recorded"
def test_the_documentation_view_is_wired_into_the_shell():
shell = (REPO / "frontend" / "index.html").read_text()
assert 'data-nav="docs"' in shell, "the manuals are unreachable from the navigation"
assert '"/docs"' in (REPO / "frontend" / "js" / "app.js").read_text()
# ── the policy the view runs under ───────────────────────────────────────────
def test_the_script_boundary_is_unchanged_and_only_style_was_relaxed():
"""The measured cost of rendering diagrams in the browser, held to that cost.
mermaid needs to style the SVG it builds, which `style-src 'self'` refuses. It
does not need to execute generated code, so `script-src` must never acquire the
escape hatch that would let injected markup run.
"""
policy = DEFAULT_HEADERS["content-security-policy"]
directives = {
part.split(" ")[0]: part.split(" ")[1:]
for part in (piece.strip() for piece in policy.split(";"))
if part
}
assert directives["script-src"] == ["'self'"]
assert "'unsafe-inline'" in directives["style-src"]
for directive in ("default-src", "img-src", "connect-src", "font-src"):
assert "'self'" in directives[directive]
assert "'unsafe-inline'" not in directives[directive]
assert directives["frame-ancestors"] == ["'none'"]
# ── serving it ───────────────────────────────────────────────────────────────
def test_the_application_serves_the_markdown_without_a_session(tmp_path):
"""Whoever cannot get past the access secret is exactly who needs these pages."""
config = Config.from_env({"PHOTO_PIPELINE_DATA_DIR": str(tmp_path / "data")})
(tmp_path / "data").mkdir()
with TestClient(create_app(config), raise_server_exceptions=False) as client:
response = client.get("/docs/index.md", headers={"cookie": ""})
assert response.status_code == 200
assert "Photo Pipeline documentation" in response.text
assert client.get("/docs/overview.md").status_code == 200
assert client.get("/docs/nothing-here.md").status_code == 404
# The mount is a directory, not a path parameter: nothing above it is reachable.
assert client.get("/docs/../pyproject.toml").status_code in (307, 404)
assert client.get("/docs/%2e%2e/pyproject.toml").status_code == 404
def test_the_image_ships_the_documentation_it_serves():
"""An image without `docs/` serves an empty manual — and the build context is
deny-by-default, so a new directory is excluded until it is named."""
assert "\n!docs\n" in (REPO / ".dockerignore").read_text()
assert "COPY docs ./docs" in (REPO / "Dockerfile").read_text()

View File

@@ -0,0 +1,178 @@
"""US09-02: the installation manual, checked against the thing it describes.
Documentation rots quietly. A setting is renamed, a command grows a flag, an exit
code changes meaning — and the manual keeps confidently saying the old thing until
somebody follows it into a bad afternoon. So every fact in it that the code also
knows is compared with the code, in both directions where both directions are
meaningful: a setting the manual invents fails, and a setting the manual forgot
fails too.
"""
from __future__ import annotations
import re
import subprocess
import sys
from functools import lru_cache
from pathlib import Path
from photo_pipeline.config import ENV_PREFIX, LEGACY_ALIASES, Config
REPO = Path(__file__).resolve().parents[2]
MANUAL = REPO / "docs" / "installation.md"
MAIN = REPO / "photo_pipeline" / "__main__.py"
TEXT = MANUAL.read_text()
ENV_NAMES = re.compile(rf"{ENV_PREFIX}[A-Z0-9_]+")
# Read by docker-compose.yml, not by Config. They are configuration an installer sets,
# so the manual documents them; they are simply not application settings.
COMPOSITION_ONLY = {
f"{ENV_PREFIX}IMAGE",
f"{ENV_PREFIX}LIBRARY_HOST_PATH",
f"{ENV_PREFIX}PUBLISH_ADDRESS",
f"{ENV_PREFIX}UID",
f"{ENV_PREFIX}GID",
f"{ENV_PREFIX}ENV_FILE", # read before Config exists, so it is not a field
}
def documented_env_names() -> set[str]:
return set(ENV_NAMES.findall(TEXT))
@lru_cache
def cli_help(*args: str) -> str:
"""What the CLI says about itself, asked the way an operator asks."""
result = subprocess.run(
[sys.executable, "-m", "photo_pipeline", *args, "--help"],
cwd=REPO,
capture_output=True,
text=True,
timeout=120,
check=False,
)
assert result.returncode == 0, result.stderr
return result.stdout
@lru_cache
def cli_commands() -> frozenset[str]:
listed = re.search(r"\{([a-z0-9,\-]+)\}", cli_help())
assert listed, f"the CLI listed no subcommands:\n{cli_help()}"
return frozenset(listed.group(1).split(","))
def cli_flags(command: str) -> frozenset[str]:
return frozenset(re.findall(r"--[a-z][a-z-]+", cli_help(command)))
# ── settings ─────────────────────────────────────────────────────────────────
def test_every_setting_is_documented():
"""A setting nobody documents is a setting nobody configures deliberately."""
expected = {f"{ENV_PREFIX}{field.upper()}" for field in Config.model_fields}
missing = sorted(expected - documented_env_names())
assert missing == [], f"undocumented settings: {missing}"
def test_the_manual_invents_no_settings():
unknown = sorted(
name
for name in documented_env_names()
if name.removeprefix(ENV_PREFIX).lower() not in Config.model_fields
and name not in COMPOSITION_ONLY
)
assert unknown == [], f"documented but not a real setting: {unknown}"
def test_the_documented_defaults_are_the_real_defaults():
"""Spot-checked where a wrong default is actively dangerous: the bind address,
the trust boundary, and the gate that protects an irreplaceable library."""
assert Config.model_fields["host"].default == "127.0.0.1"
assert "`PHOTO_PIPELINE_HOST` | `127.0.0.1`" in TEXT
assert Config.model_fields["allowed_hosts"].default == ()
assert Config.model_fields["require_dry_run_approval"].default is False
assert "`PHOTO_PIPELINE_REQUIRE_DRY_RUN_APPROVAL` | `false`" in TEXT
def test_every_secret_is_marked_as_one():
for field in ("access_secret", "vision_api_key", "immich_api_key"):
name = f"{ENV_PREFIX}{field.upper()}"
row = next(line for line in TEXT.splitlines() if line.startswith(f"| `{name}`"))
assert "secret" in row.lower(), f"{name} is not marked as a secret"
def test_the_legacy_aliases_are_documented():
"""An operator with an existing photo_analyzer.env needs to know it still works."""
for alias in LEGACY_ALIASES:
assert alias in TEXT, f"legacy alias {alias} is undocumented"
# ── commands ─────────────────────────────────────────────────────────────────
# The manual is an installation and operations manual, not a command reference: these
# are the commands an installer actually runs, and every one of them must exist.
OPERATIONAL = ("migrate", "serve", "worker", "backup", "verify-backup", "restore", "diagnostics")
def test_every_command_the_manual_tells_you_to_run_exists():
missing = [name for name in OPERATIONAL if name not in cli_commands()]
assert missing == [], f"the manual names commands the CLI does not have: {missing}"
undocumented = [name for name in OPERATIONAL if not re.search(rf"\b{re.escape(name)}\b", TEXT)]
assert undocumented == [], f"operational commands the manual omits: {undocumented}"
def test_every_documented_flag_exists_on_the_command_it_is_shown_with():
for command, flag in (
("backup", "--reason"),
("backup", "--keep"),
("restore", "--into"),
("worker", "--allow-legacy"),
("serve", "--allow-legacy"),
):
assert flag in cli_flags(command), f"{command} has no {flag}"
assert flag in TEXT, f"{flag} is undocumented"
# ── exit codes ───────────────────────────────────────────────────────────────
def test_every_documented_exit_code_is_one_the_cli_can_return():
documented = {int(code) for code in re.findall(r"^\| `(\d)` \|", TEXT, re.MULTILINE)}
returned = {int(code) for code in re.findall(r"^\s+return (\d)$", MAIN.read_text(), re.MULTILINE)}
assert documented, "no exit codes are documented"
assert documented <= returned | {0}, f"documented but unreachable: {sorted(documented - returned)}"
def test_every_refusal_exit_code_is_documented():
"""0 is success and 1 is 'it said why'; every other code is a specific refusal an
operator will meet at startup, and meeting an undocumented one is the worst case."""
returned = {int(code) for code in re.findall(r"^\s+return (\d)$", MAIN.read_text(), re.MULTILINE)}
documented = {int(code) for code in re.findall(r"^\| `(\d)` \|", TEXT, re.MULTILINE)}
missing = sorted(code for code in returned if code not in documented and code != 0)
assert missing == [], f"undocumented exit codes: {missing}"
# ── the promises the manual makes ────────────────────────────────────────────
def test_the_manual_carries_no_credential_shaped_example():
"""The repository's own scanner refuses these in a commit; a manual is exactly
where a real-looking one gets copied from."""
suspicious = re.findall(
r"(?i)(api[_-]?key|secret|token|password)\s*[=:]\s*[\"']?([A-Za-z0-9_\-]{12,})",
TEXT,
)
assert suspicious == [], f"credential-shaped example: {suspicious}"
def test_the_manual_states_the_invariants_an_installer_must_not_break():
for invariant in ("_IGNORE/", "One writer", "EXIF is verified before upload"):
assert invariant in TEXT, f"the manual does not state: {invariant}"
def test_the_manual_is_reachable_from_the_index():
assert "installation.md" in (REPO / "docs" / "index.md").read_text()

View File

@@ -193,11 +193,16 @@
"US08-05": [
"tests/integration/test_container_gate.py",
"tests/e2e/test_phase_h_container.py"
],
"US09-01": [
"tests/integration/test_documentation.py",
"tests/e2e/test_docs_ui.py"
],
"US09-02": [
"tests/integration/test_installation_manual.py"
]
},
"planned": [
"US09-01",
"US09-02",
"US09-03",
"US09-04",
"US09-05"