Compare commits

...

2 Commits

Author SHA1 Message Date
adfdf1dbdb US09-01: Serve the Documentation Inside the Application
Some checks failed
Test / suites (pull_request) Failing after 3m10s
Test / container (pull_request) Has been skipped
2026-08-23 18:45:43 +02:00
370f966d29 E09: Product documentation backlog (#101)
Some checks failed
Test / suites (push) Failing after 3m12s
Test / container (push) Failing after 4m23s
2026-08-23 15:04:45 +02:00
25 changed files with 4844 additions and 9 deletions

View File

@@ -9,6 +9,7 @@
!photo_pipeline
!migrations
!frontend
!docs
!docker
# Nothing generated, even under an allowed directory.

9
.gitattributes vendored Normal file
View File

@@ -0,0 +1,9 @@
# Vendored third-party browser bundles (US09-01). Minified upstream files carry
# trailing whitespace and very long lines; they are not ours to reformat, and the
# submit gate's `git diff --check` would refuse them forever. Their integrity is
# controlled where it belongs instead: a pinned version and a recorded sha256 in
# frontend/js/vendor/VERSIONS.json, asserted by tests/integration/test_documentation.py.
#
# Only the whitespace check is turned off. They stay text, so a credential scan or a
# search still reads them.
frontend/js/vendor/*.js -whitespace linguist-vendored

View File

@@ -70,6 +70,8 @@ COPY pyproject.toml alembic.ini README.md ./
COPY photo_pipeline ./photo_pipeline
COPY migrations ./migrations
COPY frontend ./frontend
# The manuals are served by the application itself (US09-01), so they ship with it.
COPY docs ./docs
COPY docker/entrypoint.sh docker/healthcheck.sh /usr/local/bin/
# Runtime dependencies only: the `test` extra (pytest, playwright) and the `vision`

View File

@@ -0,0 +1,77 @@
# E09 — Product Documentation
Concept phase: none. Like [E08](E08-container-deployment.md), this epic is a delivery
addition rather than a product-scope change: the same application, documented well
enough that somebody who did not build it can install it, understand it, and operate it
without reading the source.
It does not change the product scope in
[`INTEGRATED_PIPELINE_CONCEPT.md`](../INTEGRATED_PIPELINE_CONCEPT.md). No safety
invariant moves, no schema changes, no new runtime capability. The one change to the
running application is a documentation view and the static assets it needs.
## Why the application serves its own documentation
The manuals describe an application that is reached over HTTP behind an access secret.
Documentation that lives only in the repository is unreachable from the deployment it
describes: an operator who has just been handed a URL and a secret has no Gitea account
in front of them. So the same markdown files are both the repository's documentation and
the deployment's `/docs` view, and neither is a copy of the other.
## Decisions made in this epic
**Markdown is the source.** Everything is written as markdown under `docs/`, so it is
reviewable in a diff, readable on Gitea, and renderable in the app. No documentation
format that only a tool can read.
**The renderer is vendored, not written and not fetched.** `marked` (MIT, no
dependencies, ships an ES module) is pinned and committed under
`frontend/js/vendor/`. A CDN is not an option: the application is deployed to a network
whose outbound access is not assumed, and `default-src 'self'` forbids it.
**Diagrams are mermaid, rendered client-side, with the script boundary intact.** The
claim that mermaid requires `'unsafe-eval'` was tested rather than believed: mermaid 11's
bundle contains no `eval(` and no `new Function` — only a lodash `Function("return this")`
global-detection fallback that short-circuits on `globalThis` and never executes.
Rendered under this application's exact CSP, a flowchart produced a 15.8 KB SVG and
raised **no `script-src` violation**. What it does raise is `style-src`: mermaid styles
its output with an injected `<style>` element and `style=` attributes.
So `script-src 'self'` stays exactly as US07-02 and US08-01 left it, and `style-src`
gains `'unsafe-inline'`. That relaxation is bounded on purpose: with `script-src` intact
no injected markup can execute, and with `img-src`, `connect-src`, and `font-src` all
still `'self'`, the CSS-based exfiltration channels stay closed. What remains is
defacement of a page the operator is already authenticated on.
The rejected alternative is recorded because it may become the better trade later:
pre-rendering each diagram to a committed SVG with the already-installed Playwright
(no Node toolchain needed) and serving that, which would leave CSP untouched entirely at
the cost of a generator and a freshness check. If `style-src 'unsafe-inline'` is ever
judged too much, that is the migration, and it does not change a single markdown file.
**Screenshots are produced, not pasted.** Every screenshot in the user manual is captured
by Playwright from the real application against a temporary fixture library, by the same
kind of code that already drives the browser suites. A screenshot nobody can regenerate
is a screenshot that silently stops being true.
**The documentation is held to the code by tests.** Configuration keys, CLI commands,
exit codes, API error codes, job and journal states, and module names all appear in both
the code and the manuals. Each of those is cross-checked, so the failure mode of stale
documentation is a red test rather than a misled operator.
## Stories
1. [US09-01 — Serve the documentation inside the application](stories/US09-01-docs-in-app.md)
2. [US09-02 — Write the installation and operations manual](stories/US09-02-installation-manual.md)
3. [US09-03 — Write the architecture overview](stories/US09-03-architecture-overview.md)
4. [US09-04 — Write the user manual with generated screenshots](stories/US09-04-user-manual.md)
5. [US09-05 — Automate documentation acceptance](stories/US09-05-docs-gate.md)
## Epic outcome
A deployed instance serves its own installation manual, architecture overview, and
illustrated user manual at `/app/#/docs`, rendered from the same markdown files that are
readable in the repository. Screenshots are regenerated from the running application,
every documented command, setting, state, and error code is cross-checked against the
code that implements it, and one gate fails the build when documentation and application
disagree.

View File

@@ -2,13 +2,14 @@
This backlog decomposes the phases in
[`INTEGRATED_PIPELINE_CONCEPT.md`](../INTEGRATED_PIPELINE_CONCEPT.md) into seven
epics and small, independently verifiable user stories, plus one delivery-format
epic (E08) that packages the released application as a deployable container.
epics and small, independently verifiable user stories, plus two delivery-format
epics: E08 packages the released application as a deployable container, and E09
documents it for the people who install, operate, and use it.
## Numbering and file naming
- Epics: `E01` through `E07`, matching concept Phases A through G; `E08` has no
concept phase and must not change product scope.
- Epics: `E01` through `E07`, matching concept Phases A through G; `E08` and `E09`
have no concept phase and must not change product scope.
- Stories: `US<epic>-<sequence>`, for example `US03-02`.
- Epic files: `E01-<slug>.md`.
- Story files: `stories/US01-01-<slug>.md`.
@@ -39,6 +40,7 @@ epic (E08) that packages the released application as a deployable container.
6. [E06 — Archive lifecycle](E06-archive-lifecycle.md)
7. [E07 — Hardening and release](E07-hardening-release.md)
8. [E08 — Container deployment](E08-container-deployment.md)
9. [E09 — Product documentation](E09-documentation.md)
## Shared definition of done

View File

@@ -0,0 +1,49 @@
# US09-01 — Serve the Documentation Inside the Application
Epic: [E09](../E09-documentation.md)
As an operator who was handed a URL and an access secret, I want the manuals inside the
application I am looking at, so that understanding it does not require a repository
checkout or an internet connection.
## Acceptance criteria
- Documentation lives as markdown under `docs/`, with an index that names every page and
the order it is meant to be read in. The same files render on Gitea without
modification.
- The application serves a `/docs` view reachable from the main navigation, listing the
pages and rendering the selected one.
- Rendering uses a vendored, version-pinned markdown library committed under
`frontend/js/vendor/`. It is not fetched from a CDN, not bundled by a build step, and
not written by hand. The build records the version and a checksum of the vendored file.
- Diagrams written as ```mermaid``` fenced blocks render as diagrams. Mermaid is vendored
the same way.
- `script-src` in the Content-Security-Policy is unchanged: no `'unsafe-eval'`, no
`'unsafe-inline'`. Only `style-src` gains `'unsafe-inline'`, and the reason is recorded
where the policy is defined.
- Links between documents work in the app: a relative `../foo.md#section` link navigates
to that page and heading rather than downloading a file or leaving the application.
Every heading has a stable anchor, and a deep link to an anchor scrolls to it.
- An unknown page renders a documentation-specific not-found message with a link back to
the index, and never exposes a filesystem path.
- Documentation is readable without an access secret **or** behind the same session as
the rest of the application — whichever is chosen is stated explicitly in the story's
implementation notes and covered by a test, because an operator locked out by a
configuration mistake is exactly who needs the troubleshooting page.
- The view works with no outbound network access at all.
## Automated tests
- Unit: heading-anchor slugs (duplicates, punctuation, non-ASCII), and the rewriting of
relative markdown links into in-app routes.
- Integration: every markdown file under `docs/` is reachable from the index; every
internal link in every document resolves to a file and, where an anchor is given, to a
heading that exists; the vendored library files match their recorded checksums.
- Browser: the documentation view renders a page, follows an internal link, deep-links to
an anchor, renders a mermaid diagram to an SVG, and shows the not-found message for an
unknown page — all with **zero console errors and zero CSP violations**, asserted, not
eyeballed.
## Dependencies
- None beyond the delivered application.

View File

@@ -0,0 +1,46 @@
# US09-02 — Write the Installation and Operations Manual
Epic: [E09](../E09-documentation.md)
As somebody installing this application for the first time, I want one manual that takes
me from nothing to a running instance pointed at my photo library, so that I do not have
to reconstruct the procedure from the README, the compose file, and the test suite.
## Acceptance criteria
- A **host installation** path: prerequisites and their versions (Python, exiftool,
immich-go, Playwright's browser for the test suite), virtual environment, dependency
install, `.env`, database migration, starting `serve` and `worker`, and how to verify
the install succeeded.
- A **container installation** path: the published image, the compose file, the data
volume, the library bind mount and why it is mounted the way it is, published ports,
running behind a reverse proxy with `PHOTO_PIPELINE_ALLOWED_HOSTS` and
`PHOTO_PIPELINE_TRUSTED_PROXIES`, and the Portainer/webhook deployment already in use.
- A **configuration reference**: every `PHOTO_PIPELINE_*` setting with its meaning,
default, accepted values, and whether it is a secret. Secrets are described, never
exemplified with a real-looking value.
- A **first run checklist** that ends in a verified state: library discovered, worker
claiming jobs, readiness endpoint green, diagnostics clean.
- **Operations**: upgrading (including what the migration backup does), backup, verify,
restore, pruning, diagnostics, and the log locations for each role.
- **Troubleshooting**: each refusal an operator can hit at startup — every non-zero exit
code of the CLI, the trust-boundary refusal, the unmounted library root refusal, the
library lock being held, and a legacy CLI still running — with the cause and the fix.
- The manual states the safety invariants an installer must not work around: one worker
writes, `_IGNORE/` is never read, EXIF is verified before upload, and the library lock
is authoritative.
## Automated tests
- Every `PHOTO_PIPELINE_*` setting documented exists as a field on `Config`, and every
field on `Config` is documented — both directions, so a new setting cannot ship
undocumented.
- Every CLI command and flag named in the manual exists in the argument parser, and every
subcommand the parser accepts appears in the manual.
- Every exit code documented is one the CLI can actually return.
- No documented example value collides with a real secret pattern (the repository's own
credential scanner runs over `docs/`).
## Dependencies
- US09-01

View File

@@ -0,0 +1,47 @@
# US09-03 — Write the Architecture Overview
Epic: [E09](../E09-documentation.md)
As a developer or reviewer new to this repository, I want an architecture overview that
explains what the pieces are and which rules they enforce, so that I can find the right
module and not violate an invariant I never knew existed.
## Acceptance criteria
- A **context diagram**: the operator, the browser application, the API and worker, the
photo library, the Immich server, the vision provider, and the archive destination —
with the direction and nature of every interaction, including which ones leave the
machine.
- A **runtime diagram**: the migrate/api/worker composition, the data volume, the library
bind mount, the SQLite database, and the process lock — showing what is shared and what
is exclusive.
- A **module map**: every package under `photo_pipeline/` with its responsibility and the
boundary it must not cross (which modules may touch the filesystem, which may call an
external provider, which own schema).
- **Key flows**, each as a diagram plus prose: the durable job lifecycle from enqueue to
recovery; the rename journal state machine including the rollback and manual-resolution
paths; the upload lifecycle through preflight, run, report ingestion, and verification.
- **The invariants and where they are enforced**, each pointing at the module that owns
it: one writer at a time, the library process lock, the path policy and library roots,
`_IGNORE/` exclusion, verified EXIF before upload, safety decisions gating the vision
provider, and restart-safety of every mutation.
- **The data model**: the tables, what identity means for an asset, and how a moved file
keeps its identity.
- A short **history and donor** section: what was migrated out of the legacy CLIs, where
the archive and its ledger live, and why they are kept.
- Diagrams are ```mermaid``` blocks so the source is diffable and Gitea renders them
natively.
## Automated tests
- Every package and top-level module under `photo_pipeline/` appears in the module map,
and every module named in the map exists — a new service cannot appear on the map's
blind side.
- Every job state, journal state, and upload batch state named in the document exists in
the code, and every state the code defines is named in the document.
- Every mermaid block parses (rendered in the browser test without an error node).
- Every code path referenced by file name exists.
## Dependencies
- US09-01

View File

@@ -0,0 +1,52 @@
# US09-04 — Write the User Manual with Generated Screenshots
Epic: [E09](../E09-documentation.md)
As the person actually sorting a photo library, I want a manual that walks the workflow
screen by screen and tells me what each refusal means, so that I can use the application
confidently and know what it is about to do to my files before it does it.
## Acceptance criteria
- One page per workflow stage, in the order the application presents them: discovery and
inventory, duplicate review, safety review, analysis, album proposals, renames, upload,
archive, and diagnostics. Each page answers the same four questions: what this stage is
for, what I have to decide, **what it changes on disk or on the server**, and what it
refuses to do.
- A **guided first pass** that takes a new library from scan to a verified upload, naming
the point of no return in each stage and what is reversible after it.
- Every stage page carries at least one screenshot of the real application showing the
state being described.
- Screenshots are **generated** by a committed Playwright script that seeds a temporary
fixture library, drives the application, and writes the images. Regenerating them is one
documented command. No screenshot is captured by hand.
- Screenshots contain no real library paths, no personal photos, and no secrets — the
fixture library is synthetic, and this is asserted rather than assumed.
- An **error and refusal catalogue**: every `error.code` the API can return, with what
causes it, what the application did or refused to do, and what the operator should do
next. Refusals that protect data (`lock_held`, `rename_recovery_required`,
`stale_preflight`, `path_not_allowed`, `host_not_allowed`, `dry_run_not_approved`,
conflict codes) are explained as intentional, not as faults.
- A **recovery** page: an interrupted rename, an uncertain upload, a failed migration, a
restored backup — what the application does by itself and what needs a decision.
- The work-item safety allow list is extended to permit `docs/images/**`, since the
repository denies image files by default. This is an explicit, reviewed change, not a
quiet one, and it stays narrow enough that a real photo still cannot be committed.
## Automated tests
- Every `error.code` the application can emit is documented, and every code documented
exists in the code — both directions.
- Every image referenced by a documentation page exists, and every image under
`docs/images/` is referenced by a page.
- The screenshot generator runs end to end in the test environment and produces every
image the manual references, against a temporary fixture library that is destroyed
afterwards.
- Generated screenshots are checked for library paths outside the fixture root and for
the configured secrets.
- Browser: each stage page renders in the documentation view with its screenshot loaded
and no console error.
## Dependencies
- US09-01

View File

@@ -0,0 +1,38 @@
# US09-05 — Automate Documentation Acceptance
Epic: [E09](../E09-documentation.md)
As a release owner, I want one gate that proves the documentation still describes the
application, so that a change to the code cannot quietly make the manuals wrong.
## Acceptance criteria
- One documented command runs every documentation check and retains its evidence, in the
shape the release and container gates already use.
- The gate fails when: an internal link or anchor is dead; a page is unreachable from the
index; an image is referenced but missing or present but unreferenced; a documented
setting, command, exit code, error code, or state does not exist in the code; a code
path exists that the documentation is required to cover and does not.
- The gate regenerates the screenshots and fails when a regenerated image no longer
matches the committed one beyond a stated tolerance, so a UI change that invalidates the
manual is a red build rather than a discovery months later.
- The gate renders every documentation page in a real browser and fails on any console
error or CSP violation, and asserts that `script-src` contains neither `'unsafe-eval'`
nor `'unsafe-inline'`.
- The checks run on a `phase_i` marker; CI runs the gate, and the earlier epic suites keep
running unchanged.
- The README and the application's documentation index point at each other, so neither is
the forgotten copy.
- Documentation stories are mapped in the story traceability matrix like every other
story.
## Automated tests
- The gate's own contract is testable without a browser: a seeded broken link, a missing
image, an undocumented error code, and a stale screenshot each fail it, and a clean tree
passes.
- The full documentation suite runs on `phase_i` in CI.
## Dependencies
- US09-01 through US09-04

34
docs/index.md Normal file
View File

@@ -0,0 +1,34 @@
# Photo Pipeline documentation
One local application that takes a photo library from discovery to a verified Immich
upload: duplicate detection, safety review, content analysis, album naming, guarded
renaming, upload, and archive — one visible, resume-safe workflow.
These pages are readable three ways, and they are the same files each time: in the
repository under `docs/`, on Gitea, and inside the running application under
**Docs**. There is no separate copy to fall out of date.
## Read in this order
1. [Overview](overview.md) — what the application does, the stages it moves a photo
through, and the rules it will not break.
## Being written
The remaining manuals are accepted work, not aspiration; each is a story in
[E09](https://git.domverse-berlin.eu/domverse/photoanalyzer/src/branch/main/delivery_backlog/E09-documentation.md)
and will appear here as it lands.
- **Installation and operations** — host and container installation, every setting,
the first-run checklist, upgrades, backup and restore, and what each refusal at
startup means (US09-02).
- **Architecture** — context and runtime diagrams, the module map, the job and
journal state machines, and where each invariant is enforced (US09-03).
- **User manual** — one page per workflow stage with screenshots of the real
application, and a catalogue of every error and refusal (US09-04).
## Conventions
A page tells you what a stage **changes on disk or on the server** before it tells
you how to run it. Refusals are documented as intentional: this application would
rather stop and explain than guess about somebody's photographs.

77
docs/overview.md Normal file
View File

@@ -0,0 +1,77 @@
# Overview
[← Documentation index](index.md)
Photo Pipeline sorts a photo library. It finds duplicates before anything expensive
happens to them, asks a human which pictures may leave the machine, describes the
ones that may, proposes album names from what it found, renames folders under a
crash-safe journal, uploads through `immich-go`, and can archive a finished album off
active storage without ever forgetting it existed.
It runs locally. It talks to exactly three things outside itself: a vision provider,
an Immich server, and `exiftool` — and it will tell you before it uses any of them.
## The workflow
Each stage is a gate, not a tab. A later stage can always be looked at; its actions
stay disabled until what they depend on is true.
```mermaid
flowchart TD
I["0 · Inventory<br/>discover, hash, cluster duplicates"] --> S["1 · Safety<br/>score and human decision"]
S -->|sfw| A["2 · Analysis<br/>vision provider, EXIF checkpoint"]
S -->|nsfw| U
A --> B["3 · Albums<br/>proposal, then guarded rename"]
B --> U["4 · Upload<br/>immich-go, verified bytes"]
U --> R["5 · Archive<br/>copy, verify, then reclaim space"]
```
| Stage | What it decides | What it changes |
|---|---|---|
| Inventory | which file is the canonical copy of a picture | nothing — it only reads |
| Safety | whether a photo may be sent to a cloud provider | one `sfw`/`nsfw` EXIF keyword |
| Analysis | what a photo shows | a managed caption segment and additive keywords |
| Albums | what a folder should be called | folder names, through a journaled rename |
| Upload | which exact bytes reach Immich | nothing locally; assets appear in Immich |
| Archive | which album leaves active storage | files move to the archive, after verification |
## What it will not do
These are enforced in code, not by convention, and each one is why some action you
expected is sometimes refused.
- **Nothing under `_IGNORE/` is ever read.** Not scanned, not counted, not
thumbnailed, not sent anywhere.
- **A path is not an identity.** Every picture has a stable id, so moving or renaming
it loses no history.
- **Only a confirmed-SFW photo reaches the vision provider.** A photo marked NSFW is
still uploadable to Immich; it simply never leaves for analysis.
- **Metadata is verified, not hoped for.** Every stage that writes EXIF reads it back
and proves that what it did not own is unchanged.
- **One writer at a time.** A library lock, held by one process, is what makes a
crash recoverable instead of ambiguous.
- **Nothing irreversible happens without a preview and an explicit approval** that
names the exact count.
## Where things live
| | |
|---|---|
| the photo library | wherever you point `PHOTO_PIPELINE_LIBRARY_ROOTS`; mounted read-write |
| the database, cache, logs, backups | the data directory, never inside the library |
| secrets | the environment, never the database and never a log line |
| these documents | `docs/` in the repository, served at `/docs` by the application |
## Running it
The short version, for a host installation:
```bash
python -m photo_pipeline migrate
python -m photo_pipeline serve # the API and this browser application
python -m photo_pipeline worker # the process that does the long work
```
The full procedure, the container path, and every setting belong to the installation
manual, which is being written next — see
[the documentation index](index.md#being-written).

View File

@@ -170,6 +170,20 @@ a.link:hover { text-decoration: underline; }
.grid th, .grid td { text-align: left; padding: 6px 8px; border-bottom: 1px solid var(--border); vertical-align: top; }
.path { font-family: ui-monospace, monospace; font-size: 0.85em; word-break: break-all; }
/* Docs: prose, so it gets a reading measure rather than the full window width. */
.doc { max-width: 72ch; }
.doc h1 { margin-top: 0; }
.doc h2, .doc h3 { margin-top: 1.6em; border-bottom: 1px solid var(--border); padding-bottom: 4px; }
.doc code { font-family: ui-monospace, monospace; font-size: 0.9em; background: var(--surface-2); padding: 1px 4px; border-radius: 4px; }
.doc pre { background: var(--surface-2); border: 1px solid var(--border); border-radius: var(--radius); padding: 12px; overflow: auto; }
.doc pre code { background: none; padding: 0; }
.doc table { border-collapse: collapse; width: 100%; }
.doc th, .doc td { text-align: left; padding: 6px 8px; border-bottom: 1px solid var(--border); vertical-align: top; }
.doc blockquote { margin: 1em 0; padding-left: 12px; border-left: 3px solid var(--border); color: var(--muted); }
.doc img { max-width: 100%; }
.diagram { margin: 1.5em 0; padding: 0; overflow: auto; }
.diagram svg { max-width: 100%; height: auto; }
/* Narrow screens: stack the panes so blockers and actions stay reachable. */
@media (max-width: 720px) {
.two-pane { grid-template-columns: 1fr; }

View File

@@ -21,6 +21,7 @@
<a href="#/uploads" data-nav="uploads">Upload</a>
<a href="#/archive" data-nav="archive">Archive</a>
<a href="#/stats" data-nav="stats">Stats</a>
<a href="#/docs" data-nav="docs">Docs</a>
</nav>
</header>
<main id="app" aria-live="polite"><!-- views render here --></main>

View File

@@ -1,5 +1,6 @@
import { api } from "./api.js";
import { renderArchive, setArchiveRender } from "./archive.js";
import { renderDocs } from "./docs.js";
import { navigate, onRouteChange, parseHash } from "./router.js";
import { renderRenames, setRenamesRender } from "./renames.js";
import { renderUploads, setUploadsRender } from "./uploads.js";
@@ -385,6 +386,7 @@ function render() {
else if (path === "/uploads") renderUploads(root, params);
else if (path === "/archive") renderArchive(root, params);
else if (path === "/stats") renderStats(root, params);
else if (path === "/docs") renderDocs(root, params);
else show(errorBanner("Unknown view"));
}

263
frontend/js/docs.js Normal file
View File

@@ -0,0 +1,263 @@
// The documentation view (US09-01): the manuals in `docs/`, rendered in the app.
//
// The same markdown files are the repository's documentation and the deployment's
// documentation. An operator who was handed a URL and an access secret has no
// checkout in front of them, and the network the application runs on is not assumed
// to reach a CDN — so the renderer is vendored and everything here is same-origin.
//
// Nothing on this page is authenticated. It is served from the static mount beside
// the application shell, exactly like `index.html`: the troubleshooting page is
// needed most by whoever cannot get past the access secret, and no documentation
// file contains anything a session would protect.
//
// `marked` and `mermaid` are loaded lazily, on the first documentation page and the
// first diagram. They are large, and every other view does without them.
import { el, errorBanner, setActiveNav } from "./dom.js";
const DOCS_BASE = "/docs";
const VENDOR = "/app/js/vendor";
const INDEX_PAGE = "index";
// A page name comes from the hash, so it is caller input: no scheme, no traversal,
// no absolute path. The server would refuse those too; this refuses them earlier and
// without a request that looks like an attempt.
const PAGE_PATTERN = /^[\w-]+(\/[\w-]+)*$/;
let markedPromise;
let mermaidPromise;
function loadMarked() {
markedPromise ??= import(`${VENDOR}/marked.esm.js`).then((module) => module.marked);
return markedPromise;
}
// mermaid ships one large UMD bundle rather than a self-contained ES module, so it
// arrives through a script element. `script-src 'self'` allows it because it is ours
// and same-origin; nothing here relaxes that.
function loadMermaid() {
mermaidPromise ??= new Promise((resolve, reject) => {
const script = document.createElement("script");
script.src = `${VENDOR}/mermaid.min.js`;
script.onload = () => resolve(window.mermaid);
script.onerror = () => reject(new Error("the diagram renderer could not be loaded"));
document.head.appendChild(script);
});
return mermaidPromise;
}
// ── the pure parts, unit-tested in js/tests/unit.js ──────────────────────────
/** A stable, readable heading anchor. Letters and digits of any script survive. */
export function slug(text) {
const cleaned = String(text)
.toLowerCase()
.trim()
.replace(/[^\p{L}\p{N}\s-]/gu, "")
.replace(/[\s-]+/g, "-")
.replace(/^-|-$/g, "");
return cleaned || "section";
}
/** Assign every heading an id, disambiguating repeats the way a reader would
* expect: the first `Notes` keeps `#notes`, the second becomes `#notes-2`. */
export function assignHeadingIds(headings) {
const used = new Map();
for (const heading of headings) {
const base = slug(heading.textContent);
const seen = (used.get(base) || 0) + 1;
used.set(base, seen);
heading.id = seen === 1 ? base : `${base}-${seen}`;
}
return headings;
}
/**
* Where a link inside a documentation page should go.
*
* Markdown links between documents are relative file paths, which a browser would
* treat as downloads that leave the application. Returns the in-app target for a
* link to another document or to a heading, and `null` for anything else — external
* links stay exactly as the author wrote them.
*/
export function resolveDocLink(currentPage, href) {
if (!href) return null;
if (/^[a-z][a-z\d+.-]*:/i.test(href) || href.startsWith("//")) return null;
if (href.startsWith("#")) return { page: currentPage, anchor: href.slice(1) };
const [target, anchor = ""] = href.split("#");
if (!/\.md$/i.test(target)) return null;
// Resolved against the current document, so `../guides/x.md` means what it means
// in the repository. The origin is a placeholder; only the path is used.
const resolved = new URL(target, `https://docs.invalid/${currentPage}.md`);
const page = decodeURIComponent(resolved.pathname).replace(/^\//, "").replace(/\.md$/i, "");
return PAGE_PATTERN.test(page) ? { page, anchor } : null;
}
/**
* Point every in-app link at its route, and leave every other link alone.
*
* The resolved page is kept on the element, so the reading order can be read back
* from the rendered index rather than parsed out of the markdown a second time.
*/
export function rewriteLinks(article, page) {
for (const link of article.querySelectorAll("a[href]")) {
const href = link.getAttribute("href");
const target = resolveDocLink(page, href);
if (target) {
link.setAttribute("href", docHash(target.page, target.anchor));
link.dataset.docPage = target.page;
} else if (/^https?:/i.test(href)) {
link.setAttribute("rel", "noreferrer noopener");
link.setAttribute("target", "_blank");
}
}
return article;
}
/** The reading order, taken from the index document itself — one list, in one
* place, that is equally the index on Gitea and the navigation here. */
export function documentIndex(indexElement) {
const pages = [];
for (const link of indexElement.querySelectorAll("a[data-doc-page]")) {
const page = link.dataset.docPage;
if (page !== INDEX_PAGE && !pages.some((entry) => entry.page === page)) {
pages.push({ page, title: link.textContent.trim() });
}
}
return pages;
}
// ── the view ─────────────────────────────────────────────────────────────────
async function fetchPage(page) {
const response = await fetch(`${DOCS_BASE}/${page}.md`, { headers: { accept: "text/markdown" } });
if (!response.ok) throw new Error(`documentation page unavailable: ${response.status}`);
return response.text();
}
async function toArticle(markdown, page) {
const marked = await loadMarked();
const article = el("article", { class: "doc", "data-testid": "doc", "data-page": page });
// The only innerHTML in the application, and deliberate: rendering markdown *is*
// producing HTML. The input is a file from this repository, and the page's CSP
// (`script-src 'self'`, no `'unsafe-inline'`) means injected script and inline
// handlers do not run even if one ever were not.
article.innerHTML = marked.parse(markdown, { async: false });
assignHeadingIds(article.querySelectorAll("h1, h2, h3, h4, h5, h6"));
rewriteLinks(article, page);
return article;
}
function docHash(page, anchor = "") {
const query = new URLSearchParams(anchor ? { page, anchor } : { page });
return `#/docs?${query}`;
}
/** Diagrams are authored as ```mermaid blocks so the source is diffable and Gitea
* renders them natively. A failure here replaces the diagram, never the page. */
async function renderDiagrams(article) {
const blocks = [...article.querySelectorAll("pre > code.language-mermaid")];
if (!blocks.length) return;
let mermaid;
try {
mermaid = await loadMermaid();
mermaid.initialize({ startOnLoad: false, securityLevel: "strict", theme: "dark" });
} catch (error) {
blocks.forEach((block) => block.closest("pre").replaceWith(errorBanner(error.message)));
return;
}
for (const [index, block] of blocks.entries()) {
const figure = el("figure", { class: "diagram", "data-testid": "diagram" });
block.closest("pre").replaceWith(figure);
try {
const { svg } = await mermaid.render(`diagram-${index}-${Date.now()}`, block.textContent);
figure.innerHTML = svg;
} catch (error) {
figure.replaceWith(errorBanner(`Diagram could not be drawn: ${error.message}`));
}
}
}
function notFound() {
return el(
"div",
{ class: "alert", role: "alert", "data-testid": "doc-not-found" },
"That documentation page does not exist. ",
el("a", { class: "link", href: docHash(INDEX_PAGE) }, "Back to the documentation index")
);
}
function sidebar(pages, current) {
return el(
"nav",
{ class: "card", "aria-label": "Documentation" },
el("h2", {}, "Documentation"),
el(
"ul",
{ class: "album-list", "data-testid": "doc-pages" },
el(
"li",
{},
el(
"a",
{
href: docHash(INDEX_PAGE),
"aria-current": current === INDEX_PAGE ? "true" : false,
},
"Index"
)
),
...pages.map(({ page, title }) =>
el(
"li",
{},
el(
"a",
{
href: docHash(page),
"data-page": page,
"aria-current": page === current ? "true" : false,
},
title
)
)
)
)
);
}
export async function renderDocs(root, params = {}) {
setActiveNav("docs");
const requested = params.page || INDEX_PAGE;
const page = PAGE_PATTERN.test(requested) ? requested : "";
let indexArticle;
try {
indexArticle = await toArticle(await fetchPage(INDEX_PAGE), INDEX_PAGE);
} catch (error) {
root.replaceChildren(errorBanner(`Documentation is unavailable: ${error.message}`));
return;
}
const pages = documentIndex(indexArticle);
let article;
if (!page) article = notFound();
else if (page === INDEX_PAGE) article = indexArticle;
else {
try {
article = await toArticle(await fetchPage(page), page);
} catch {
article = notFound();
}
}
root.replaceChildren(el("div", { class: "two-pane" }, sidebar(pages, page), article));
await renderDiagrams(article);
scrollToAnchor(params.anchor);
}
// A documentation anchor cannot live in the hash — the hash is the route — so it
// travels as a parameter and is applied after the page renders.
function scrollToAnchor(anchor) {
if (!anchor) return;
const target = document.getElementById(anchor);
if (target) target.scrollIntoView({ block: "start" });
}

View File

@@ -7,6 +7,7 @@ import { createStore } from "../store.js";
import { parseHash, navigate } from "../router.js";
import { api, cancellable } from "../api.js";
import { subscribeJob } from "../events.js";
import { assignHeadingIds, documentIndex, resolveDocLink, rewriteLinks, slug } from "../docs.js";
const cases = [];
function ok(name, cond) {
@@ -142,6 +143,62 @@ async function run() {
window.EventSource = realES;
}
// ── documentation view (US09-01) ─────────────────────────────────────────
{
ok("slug lowercases and dashes a heading", slug("Key Flows") === "key-flows");
ok(
"slug drops punctuation but keeps words",
slug("What it *will not* do:") === "what-it-will-not-do"
);
ok("slug keeps non-ASCII letters", slug("Größe & Gewicht") === "größe-gewicht");
ok("slug never yields an empty anchor", slug("!!!") === "section");
const doc = document.createElement("div");
doc.innerHTML = "<h2>Notes</h2><h3>Notes</h3><h2>Notes</h2>";
const ids = [...assignHeadingIds(doc.querySelectorAll("h2, h3"))].map((h) => h.id);
ok("repeated headings get distinct anchors", ids.join(",") === "notes,notes-2,notes-3");
ok(
"a sibling document link becomes an in-app route",
resolveDocLink("index", "overview.md")?.page === "overview"
);
const nested = resolveDocLink("guides/install", "../overview.md#running-it");
ok("a relative link resolves against the current page", nested?.page === "overview");
ok("a link's anchor survives the rewrite", nested?.anchor === "running-it");
ok(
"a bare anchor stays on the current page",
resolveDocLink("overview", "#the-workflow")?.page === "overview"
);
ok("an external link is left alone", resolveDocLink("index", "https://example.test/x") === null);
ok("a non-markdown relative link is left alone", resolveDocLink("index", "images/a.png") === null);
// A climb cannot leave the documentation tree: it is clamped at the root, so the
// page it names is still fetched from under /docs and simply does not exist.
ok(
"a link that climbs above the docs root is clamped",
resolveDocLink("index", "../../etc/passwd.md")?.page === "etc/passwd"
);
ok(
"a page name that is not a page name is refused",
resolveDocLink("index", "..%2f..%2fetc%2fpasswd.md") === null
);
const index = document.createElement("div");
index.innerHTML =
'<a href="overview.md">Overview</a><a href="overview.md">again</a>' +
'<a href="index.md">itself</a><a href="https://example.test">out</a>';
const listed = documentIndex(rewriteLinks(index, "index"));
ok(
"an external link is marked safe to open away from the app",
index.querySelector('a[href^="https"]').rel === "noreferrer noopener"
);
ok(
"an in-app link points at the docs route",
index.querySelector("a").getAttribute("href") === "#/docs?page=overview"
);
ok("the index lists each page once, in order", listed.length === 1);
ok("the index takes its titles from the link text", listed[0].title === "Overview");
}
await tick();
const failed = cases.filter((c) => !c.ok);
window.__RESULTS__ = { passed: cases.length - failed.length, failed: failed.length, cases };

23
frontend/js/vendor/VERSIONS.json vendored Normal file
View File

@@ -0,0 +1,23 @@
{
"_comment": "Vendored browser libraries (US09-01). Committed rather than fetched: the application is deployed to a network whose outbound access is not assumed, and `default-src 'self'` forbids a CDN. Checksums are asserted by tests/integration/test_documentation.py, so replacing a file without recording it here fails the suite. Update by downloading the pinned URL and recording the new version and sha256 in the same commit.",
"libraries": [
{
"file": "marked.esm.js",
"name": "marked",
"version": "18.0.10",
"license": "MIT",
"url": "https://cdn.jsdelivr.net/npm/marked@18.0.10/lib/marked.esm.js",
"sha256": "4cf47dfebb7f614a08fc0a579ab0fe407ff0ed2b717bf953040c85b2f493a4f0",
"why": "Markdown to HTML for the documentation view. An ES module with no dependencies, imported lazily by js/docs.js."
},
{
"file": "mermaid.min.js",
"name": "mermaid",
"version": "11.17.0",
"license": "MIT",
"url": "https://cdn.jsdelivr.net/npm/mermaid@11.17.0/dist/mermaid.min.js",
"sha256": "8d8e0eec56d3a83b4b3c87f42050845546dee93ebe1875d2117c12e6947c0cb3",
"why": "Renders ```mermaid blocks so diagrams are diffable source that Gitea also renders. The single-file UMD build rather than the ES module entry, whose ~40 lazy chunks would each need vendoring and pinning."
}
]
}

78
frontend/js/vendor/marked.esm.js vendored Normal file

File diff suppressed because one or more lines are too long

3636
frontend/js/vendor/mermaid.min.js vendored Normal file

File diff suppressed because one or more lines are too long

View File

@@ -53,6 +53,7 @@ from photo_pipeline.services.thumbnails import ThumbnailService
from photo_pipeline.services.upload_batches import UploadBatchService
FRONTEND_DIR = Path(__file__).resolve().parents[2] / "frontend"
DOCS_DIR = Path(__file__).resolve().parents[2] / "docs"
log = logging.getLogger(__name__)
@@ -151,4 +152,10 @@ def create_app(config: Config | None = None) -> FastAPI:
# Static single-page app (hash-routed). Mounted last so /api/v1 wins.
if FRONTEND_DIR.is_dir():
app.mount("/app", StaticFiles(directory=FRONTEND_DIR, html=True), name="app")
# The manuals, as their own markdown files (US09-01). Unauthenticated for the
# same reason the application shell is: whoever cannot get past the access
# secret is exactly who needs the troubleshooting page, and no document holds
# anything a session would protect.
if DOCS_DIR.is_dir():
app.mount("/docs", StaticFiles(directory=DOCS_DIR), name="docs")
return app

View File

@@ -57,9 +57,21 @@ LOOPBACK_HOSTS = frozenset({"127.0.0.1", "localhost", "::1", "[::1]"})
# backup a careful operator takes first, and its retention (US07-07).
MUTATION_EXEMPT_PATHS = frozenset({f"{API_PREFIX}/backups", f"{API_PREFIX}/backups/prune"})
# Applied to every response. No inline script/style is used by the frontend, so the
# policy can stay strict; `frame-ancestors 'none'` and CORP keep other pages from
# Applied to every response. `frame-ancestors 'none'` and CORP keep other pages from
# embedding the app or its thumbnails.
#
# `script-src 'self'` is the boundary that matters and it is unchanged: no
# `'unsafe-eval'`, no `'unsafe-inline'`, so nothing injected into the DOM can execute.
# Whether the vendored diagram renderer needed `'unsafe-eval'` was measured rather
# than assumed — mermaid's bundle contains no `eval(` and no `new Function`, and it
# renders under this exact policy without a single script-src violation.
#
# `style-src` does gain `'unsafe-inline'` (US09-01): mermaid styles the SVG it builds
# with an injected `<style>` element and `style=` attributes, and a diagram's CSS
# cannot be hashed in advance. The concession is bounded by the directives around it
# — with script execution still refused and `img-src`, `connect-src`, and `font-src`
# all `'self'`, the CSS exfiltration channels stay closed and what is left is
# defacement of a page its own operator is already looking at.
DEFAULT_HEADERS = {
"x-content-type-options": "nosniff",
"x-frame-options": "DENY",
@@ -67,8 +79,9 @@ DEFAULT_HEADERS = {
"cross-origin-resource-policy": "same-origin",
"cross-origin-opener-policy": "same-origin",
"content-security-policy": (
"default-src 'self'; img-src 'self' data:; style-src 'self'; script-src 'self'; "
"connect-src 'self'; frame-ancestors 'none'; base-uri 'none'; form-action 'none'"
"default-src 'self'; img-src 'self' data:; style-src 'self' 'unsafe-inline'; "
"script-src 'self'; connect-src 'self'; font-src 'self'; "
"frame-ancestors 'none'; base-uri 'none'; form-action 'none'"
),
}

119
tests/e2e/test_docs_ui.py Normal file
View File

@@ -0,0 +1,119 @@
"""US09-01: the manuals, in the browser, under the application's real policy.
Everything here runs against a real ``photo_pipeline serve`` process serving the
committed ``docs/`` tree and the vendored renderer — no fixture markdown, no stubbed
fetch. What that buys is the assertion the story actually cares about: the pages
render, the links between them work, a diagram becomes a diagram, and the console
stays empty, including of CSP violations.
The offline contract — reachability, dead links, pinned checksums — is
``tests/integration/test_documentation.py``.
"""
from __future__ import annotations
import pytest
from tests.e2e._pipeline_harness import Server, seed_library
@pytest.fixture(scope="module")
def server(tmp_path_factory):
seeded = seed_library(tmp_path_factory.mktemp("docs"), {"a": 1}, {})
running = Server(seeded).start()
try:
yield running
finally:
running.stop()
@pytest.fixture
def quiet(page):
"""Any console error, page error, or failed request fails the test that caused
it. A CSP violation arrives as a console error, which is the point."""
problems: list[str] = []
page.on("console", lambda m: problems.append(m.text) if m.type == "error" else None)
page.on("pageerror", lambda error: problems.append(str(error)))
page.on("requestfailed", lambda request: problems.append(f"failed request: {request.url}"))
yield problems
def test_the_documentation_opens_from_the_navigation(page, server, quiet):
page.goto(f"{server.base}/app/#/workflow")
page.locator('nav a[data-nav="docs"]').click()
page.get_by_test_id("doc").wait_for()
assert page.get_by_test_id("doc").get_attribute("data-page") == "index"
# The sidebar is the index's own reading order, not a second list to maintain.
assert page.get_by_test_id("doc-pages").get_by_role("link", name="Overview").is_visible()
assert quiet == []
def test_a_link_between_documents_stays_inside_the_application(page, server, quiet):
page.goto(f"{server.base}/app/#/docs")
page.get_by_test_id("doc").wait_for()
page.get_by_test_id("doc").get_by_role("link", name="Overview").click()
page.wait_for_selector('[data-testid="doc"][data-page="overview"]')
assert "#/docs?page=overview" in page.url, "a relative .md link must not leave the app"
# And back again, by the link the document itself carries.
page.get_by_test_id("doc").get_by_role("link", name="← Documentation index").click()
page.wait_for_selector('[data-testid="doc"][data-page="index"]')
assert quiet == []
def test_a_deep_link_to_a_heading_lands_on_that_heading(page, server, quiet):
page.goto(f"{server.base}/app/#/docs?page=overview&anchor=the-workflow")
heading = page.locator("#the-workflow")
heading.wait_for()
assert heading.inner_text().strip() == "The workflow"
assert heading.evaluate("node => node.getBoundingClientRect().top < window.innerHeight")
assert quiet == []
def test_a_mermaid_block_becomes_a_diagram_under_the_unchanged_script_policy(page, server, quiet):
page.goto(f"{server.base}/app/#/docs?page=overview")
diagram = page.get_by_test_id("diagram").first
diagram.wait_for()
# A real drawing, not the source text and not an error node.
assert diagram.locator("svg").count() == 1
assert diagram.evaluate("node => node.querySelector('svg').getBBox().width") > 100
assert page.locator("pre code.language-mermaid").count() == 0
assert quiet == [], "rendering a diagram must not violate the policy"
policy = page.evaluate(
"async () => (await fetch('/app/')).headers.get('content-security-policy')"
)
assert "script-src 'self';" in policy
assert "unsafe-eval" not in policy
def test_an_unknown_page_says_so_without_naming_a_path(page, server, quiet):
page.goto(f"{server.base}/app/#/docs?page=no-such-manual")
message = page.get_by_test_id("doc-not-found")
message.wait_for()
text = message.inner_text()
assert "does not exist" in text
assert "/" not in text.replace("Back to the documentation index", ""), text
message.get_by_role("link").click()
page.wait_for_selector('[data-testid="doc"][data-page="index"]')
# The missing page is a 404 and the browser says so; nothing else may go wrong,
# and in particular the view must not throw on the way to its own message.
assert all("404" in problem for problem in quiet), quiet
def test_the_documentation_is_readable_without_a_session(page, server, quiet):
"""The troubleshooting page is needed most by whoever is locked out."""
page.goto(f"{server.base}/app/#/docs")
page.get_by_test_id("doc").wait_for()
unauthenticated = page.evaluate(
"""async () => {
const response = await fetch('/docs/index.md', { credentials: 'omit' });
return { status: response.status, length: (await response.text()).length };
}"""
)
assert unauthenticated["status"] == 200
assert unauthenticated["length"] > 100

View File

@@ -0,0 +1,179 @@
"""US09-01: the documentation the application serves, checked without a browser.
What is checkable offline is everything that makes the manuals *navigable* rather
than merely present: every page reachable, every internal link and anchor landing
somewhere real, the vendored renderer being the file that was pinned, and the
application actually serving `docs/` in both the working copy and the image.
The rendering itself — marked, mermaid, anchors, link rewriting — is proven in a
real browser by ``tests/e2e/test_docs_ui.py``.
"""
from __future__ import annotations
import hashlib
import json
import re
from pathlib import Path
from starlette.testclient import TestClient
from photo_pipeline.api.app import create_app
from photo_pipeline.api.security import DEFAULT_HEADERS
from photo_pipeline.config import Config
REPO = Path(__file__).resolve().parents[2]
DOCS = REPO / "docs"
VENDOR = REPO / "frontend" / "js" / "vendor"
INDEX = DOCS / "index.md"
MARKDOWN_LINK = re.compile(r"\[[^\]]*\]\(([^)\s]+)\)")
HEADING = re.compile(r"^#{1,6}\s+(.+?)\s*$", re.MULTILINE)
FENCE = re.compile(r"^```.*?^```", re.MULTILINE | re.DOTALL)
def pages() -> list[Path]:
return sorted(DOCS.rglob("*.md"))
def body(path: Path) -> str:
"""A document without its fenced code, so an example link in a shell snippet is
not mistaken for a link the reader can follow."""
return FENCE.sub("", path.read_text())
def slug(text: str) -> str:
"""The heading-anchor rule, mirroring ``frontend/js/docs.js``.
Deliberately duplicated: this check has to run without a browser. The browser
test is the authority on the real behaviour, and it asserts the same anchors.
"""
cleaned = re.sub(r"[^\w\s-]", "", text.lower(), flags=re.UNICODE).strip()
return re.sub(r"[\s-]+", "-", cleaned).strip("-") or "section"
def anchors(path: Path) -> set[str]:
found: dict[str, int] = {}
for heading in HEADING.findall(body(path)):
# Markdown emphasis and inline code are not part of the rendered text.
plain = re.sub(r"[*`_]", "", heading)
base = slug(plain)
found[base] = found.get(base, 0) + 1
return {name if index == 1 else f"{name}-{index}" for name, count in found.items() for index in range(1, count + 1)}
def internal_links(path: Path) -> list[str]:
return [
href
for href in MARKDOWN_LINK.findall(body(path))
if not re.match(r"^[a-z][a-z\d+.-]*:", href, re.IGNORECASE) and not href.startswith("//")
]
# ── the documentation tree ───────────────────────────────────────────────────
def test_the_index_exists_and_names_every_page():
"""A page nobody links to is a page nobody reads."""
assert INDEX.is_file()
listed = {
(INDEX.parent / href.split("#")[0]).resolve()
for href in internal_links(INDEX)
if href.split("#")[0].endswith(".md")
}
unreachable = [page.name for page in pages() if page != INDEX and page.resolve() not in listed]
assert unreachable == [], f"not linked from the index: {unreachable}"
def test_every_internal_link_and_anchor_resolves():
broken: list[str] = []
for page in pages():
for href in internal_links(page):
target, _, anchor = href.partition("#")
destination = page if not target else (page.parent / target).resolve()
if not destination.is_file():
broken.append(f"{page.name}{href} (no such file)")
continue
if anchor and anchor not in anchors(destination):
broken.append(f"{page.name}{href} (no such heading)")
assert broken == [], broken
def test_every_page_starts_with_one_title():
for page in pages():
titles = [line for line in page.read_text().splitlines() if line.startswith("# ")]
assert len(titles) == 1, f"{page.name} has {len(titles)} top-level titles"
# ── the vendored renderer ────────────────────────────────────────────────────
def test_the_vendored_libraries_are_the_files_that_were_pinned():
"""A vendored dependency without a recorded checksum is a dependency nobody is
reviewing. Replacing one has to be a visible change to this manifest."""
manifest = json.loads((VENDOR / "VERSIONS.json").read_text())
recorded = {entry["file"]: entry for entry in manifest["libraries"]}
on_disk = {path.name for path in VENDOR.iterdir() if path.suffix in (".js", ".mjs")}
assert on_disk == set(recorded), f"unrecorded vendored files: {on_disk ^ set(recorded)}"
for name, entry in recorded.items():
digest = hashlib.sha256((VENDOR / name).read_bytes()).hexdigest()
assert digest == entry["sha256"], f"{name} does not match the pinned checksum"
assert entry["version"] in entry["url"], f"{name}: the pinned URL and version disagree"
assert entry["license"], f"{name}: no license recorded"
def test_the_documentation_view_is_wired_into_the_shell():
shell = (REPO / "frontend" / "index.html").read_text()
assert 'data-nav="docs"' in shell, "the manuals are unreachable from the navigation"
assert '"/docs"' in (REPO / "frontend" / "js" / "app.js").read_text()
# ── the policy the view runs under ───────────────────────────────────────────
def test_the_script_boundary_is_unchanged_and_only_style_was_relaxed():
"""The measured cost of rendering diagrams in the browser, held to that cost.
mermaid needs to style the SVG it builds, which `style-src 'self'` refuses. It
does not need to execute generated code, so `script-src` must never acquire the
escape hatch that would let injected markup run.
"""
policy = DEFAULT_HEADERS["content-security-policy"]
directives = {
part.split(" ")[0]: part.split(" ")[1:]
for part in (piece.strip() for piece in policy.split(";"))
if part
}
assert directives["script-src"] == ["'self'"]
assert "'unsafe-inline'" in directives["style-src"]
for directive in ("default-src", "img-src", "connect-src", "font-src"):
assert "'self'" in directives[directive]
assert "'unsafe-inline'" not in directives[directive]
assert directives["frame-ancestors"] == ["'none'"]
# ── serving it ───────────────────────────────────────────────────────────────
def test_the_application_serves_the_markdown_without_a_session(tmp_path):
"""Whoever cannot get past the access secret is exactly who needs these pages."""
config = Config.from_env({"PHOTO_PIPELINE_DATA_DIR": str(tmp_path / "data")})
(tmp_path / "data").mkdir()
with TestClient(create_app(config), raise_server_exceptions=False) as client:
response = client.get("/docs/index.md", headers={"cookie": ""})
assert response.status_code == 200
assert "Photo Pipeline documentation" in response.text
assert client.get("/docs/overview.md").status_code == 200
assert client.get("/docs/nothing-here.md").status_code == 404
# The mount is a directory, not a path parameter: nothing above it is reachable.
assert client.get("/docs/../pyproject.toml").status_code in (307, 404)
assert client.get("/docs/%2e%2e/pyproject.toml").status_code == 404
def test_the_image_ships_the_documentation_it_serves():
"""An image without `docs/` serves an empty manual — and the build context is
deny-by-default, so a new directory is excluded until it is named."""
assert "\n!docs\n" in (REPO / ".dockerignore").read_text()
assert "COPY docs ./docs" in (REPO / "Dockerfile").read_text()

View File

@@ -193,8 +193,17 @@
"US08-05": [
"tests/integration/test_container_gate.py",
"tests/e2e/test_phase_h_container.py"
],
"US09-01": [
"tests/integration/test_documentation.py",
"tests/e2e/test_docs_ui.py"
]
},
"planned": [],
"planned": [
"US09-02",
"US09-03",
"US09-04",
"US09-05"
],
"_planned_comment": "Accepted backlog stories that are not implemented yet. The release gate (US07-07) requires every story file to be either mapped to tests or listed here, so an unimplemented story is a visible decision rather than a hole in the matrix."
}