Compare commits

..

2 Commits

Author SHA1 Message Date
adfdf1dbdb US09-01: Serve the Documentation Inside the Application
Some checks failed
Test / suites (pull_request) Failing after 3m10s
Test / container (pull_request) Has been skipped
2026-08-23 18:45:43 +02:00
370f966d29 E09: Product documentation backlog (#101)
Some checks failed
Test / suites (push) Failing after 3m12s
Test / container (push) Failing after 4m23s
2026-08-23 15:04:45 +02:00
18 changed files with 4523 additions and 5 deletions

View File

@@ -9,6 +9,7 @@
!photo_pipeline
!migrations
!frontend
!docs
!docker
# Nothing generated, even under an allowed directory.

9
.gitattributes vendored Normal file
View File

@@ -0,0 +1,9 @@
# Vendored third-party browser bundles (US09-01). Minified upstream files carry
# trailing whitespace and very long lines; they are not ours to reformat, and the
# submit gate's `git diff --check` would refuse them forever. Their integrity is
# controlled where it belongs instead: a pinned version and a recorded sha256 in
# frontend/js/vendor/VERSIONS.json, asserted by tests/integration/test_documentation.py.
#
# Only the whitespace check is turned off. They stay text, so a credential scan or a
# search still reads them.
frontend/js/vendor/*.js -whitespace linguist-vendored

View File

@@ -70,6 +70,8 @@ COPY pyproject.toml alembic.ini README.md ./
COPY photo_pipeline ./photo_pipeline
COPY migrations ./migrations
COPY frontend ./frontend
# The manuals are served by the application itself (US09-01), so they ship with it.
COPY docs ./docs
COPY docker/entrypoint.sh docker/healthcheck.sh /usr/local/bin/
# Runtime dependencies only: the `test` extra (pytest, playwright) and the `vision`

34
docs/index.md Normal file
View File

@@ -0,0 +1,34 @@
# Photo Pipeline documentation
One local application that takes a photo library from discovery to a verified Immich
upload: duplicate detection, safety review, content analysis, album naming, guarded
renaming, upload, and archive — one visible, resume-safe workflow.
These pages are readable three ways, and they are the same files each time: in the
repository under `docs/`, on Gitea, and inside the running application under
**Docs**. There is no separate copy to fall out of date.
## Read in this order
1. [Overview](overview.md) — what the application does, the stages it moves a photo
through, and the rules it will not break.
## Being written
The remaining manuals are accepted work, not aspiration; each is a story in
[E09](https://git.domverse-berlin.eu/domverse/photoanalyzer/src/branch/main/delivery_backlog/E09-documentation.md)
and will appear here as it lands.
- **Installation and operations** — host and container installation, every setting,
the first-run checklist, upgrades, backup and restore, and what each refusal at
startup means (US09-02).
- **Architecture** — context and runtime diagrams, the module map, the job and
journal state machines, and where each invariant is enforced (US09-03).
- **User manual** — one page per workflow stage with screenshots of the real
application, and a catalogue of every error and refusal (US09-04).
## Conventions
A page tells you what a stage **changes on disk or on the server** before it tells
you how to run it. Refusals are documented as intentional: this application would
rather stop and explain than guess about somebody's photographs.

77
docs/overview.md Normal file
View File

@@ -0,0 +1,77 @@
# Overview
[← Documentation index](index.md)
Photo Pipeline sorts a photo library. It finds duplicates before anything expensive
happens to them, asks a human which pictures may leave the machine, describes the
ones that may, proposes album names from what it found, renames folders under a
crash-safe journal, uploads through `immich-go`, and can archive a finished album off
active storage without ever forgetting it existed.
It runs locally. It talks to exactly three things outside itself: a vision provider,
an Immich server, and `exiftool` — and it will tell you before it uses any of them.
## The workflow
Each stage is a gate, not a tab. A later stage can always be looked at; its actions
stay disabled until what they depend on is true.
```mermaid
flowchart TD
I["0 · Inventory<br/>discover, hash, cluster duplicates"] --> S["1 · Safety<br/>score and human decision"]
S -->|sfw| A["2 · Analysis<br/>vision provider, EXIF checkpoint"]
S -->|nsfw| U
A --> B["3 · Albums<br/>proposal, then guarded rename"]
B --> U["4 · Upload<br/>immich-go, verified bytes"]
U --> R["5 · Archive<br/>copy, verify, then reclaim space"]
```
| Stage | What it decides | What it changes |
|---|---|---|
| Inventory | which file is the canonical copy of a picture | nothing — it only reads |
| Safety | whether a photo may be sent to a cloud provider | one `sfw`/`nsfw` EXIF keyword |
| Analysis | what a photo shows | a managed caption segment and additive keywords |
| Albums | what a folder should be called | folder names, through a journaled rename |
| Upload | which exact bytes reach Immich | nothing locally; assets appear in Immich |
| Archive | which album leaves active storage | files move to the archive, after verification |
## What it will not do
These are enforced in code, not by convention, and each one is why some action you
expected is sometimes refused.
- **Nothing under `_IGNORE/` is ever read.** Not scanned, not counted, not
thumbnailed, not sent anywhere.
- **A path is not an identity.** Every picture has a stable id, so moving or renaming
it loses no history.
- **Only a confirmed-SFW photo reaches the vision provider.** A photo marked NSFW is
still uploadable to Immich; it simply never leaves for analysis.
- **Metadata is verified, not hoped for.** Every stage that writes EXIF reads it back
and proves that what it did not own is unchanged.
- **One writer at a time.** A library lock, held by one process, is what makes a
crash recoverable instead of ambiguous.
- **Nothing irreversible happens without a preview and an explicit approval** that
names the exact count.
## Where things live
| | |
|---|---|
| the photo library | wherever you point `PHOTO_PIPELINE_LIBRARY_ROOTS`; mounted read-write |
| the database, cache, logs, backups | the data directory, never inside the library |
| secrets | the environment, never the database and never a log line |
| these documents | `docs/` in the repository, served at `/docs` by the application |
## Running it
The short version, for a host installation:
```bash
python -m photo_pipeline migrate
python -m photo_pipeline serve # the API and this browser application
python -m photo_pipeline worker # the process that does the long work
```
The full procedure, the container path, and every setting belong to the installation
manual, which is being written next — see
[the documentation index](index.md#being-written).

View File

@@ -170,6 +170,20 @@ a.link:hover { text-decoration: underline; }
.grid th, .grid td { text-align: left; padding: 6px 8px; border-bottom: 1px solid var(--border); vertical-align: top; }
.path { font-family: ui-monospace, monospace; font-size: 0.85em; word-break: break-all; }
/* Docs: prose, so it gets a reading measure rather than the full window width. */
.doc { max-width: 72ch; }
.doc h1 { margin-top: 0; }
.doc h2, .doc h3 { margin-top: 1.6em; border-bottom: 1px solid var(--border); padding-bottom: 4px; }
.doc code { font-family: ui-monospace, monospace; font-size: 0.9em; background: var(--surface-2); padding: 1px 4px; border-radius: 4px; }
.doc pre { background: var(--surface-2); border: 1px solid var(--border); border-radius: var(--radius); padding: 12px; overflow: auto; }
.doc pre code { background: none; padding: 0; }
.doc table { border-collapse: collapse; width: 100%; }
.doc th, .doc td { text-align: left; padding: 6px 8px; border-bottom: 1px solid var(--border); vertical-align: top; }
.doc blockquote { margin: 1em 0; padding-left: 12px; border-left: 3px solid var(--border); color: var(--muted); }
.doc img { max-width: 100%; }
.diagram { margin: 1.5em 0; padding: 0; overflow: auto; }
.diagram svg { max-width: 100%; height: auto; }
/* Narrow screens: stack the panes so blockers and actions stay reachable. */
@media (max-width: 720px) {
.two-pane { grid-template-columns: 1fr; }

View File

@@ -21,6 +21,7 @@
<a href="#/uploads" data-nav="uploads">Upload</a>
<a href="#/archive" data-nav="archive">Archive</a>
<a href="#/stats" data-nav="stats">Stats</a>
<a href="#/docs" data-nav="docs">Docs</a>
</nav>
</header>
<main id="app" aria-live="polite"><!-- views render here --></main>

View File

@@ -1,5 +1,6 @@
import { api } from "./api.js";
import { renderArchive, setArchiveRender } from "./archive.js";
import { renderDocs } from "./docs.js";
import { navigate, onRouteChange, parseHash } from "./router.js";
import { renderRenames, setRenamesRender } from "./renames.js";
import { renderUploads, setUploadsRender } from "./uploads.js";
@@ -385,6 +386,7 @@ function render() {
else if (path === "/uploads") renderUploads(root, params);
else if (path === "/archive") renderArchive(root, params);
else if (path === "/stats") renderStats(root, params);
else if (path === "/docs") renderDocs(root, params);
else show(errorBanner("Unknown view"));
}

263
frontend/js/docs.js Normal file
View File

@@ -0,0 +1,263 @@
// The documentation view (US09-01): the manuals in `docs/`, rendered in the app.
//
// The same markdown files are the repository's documentation and the deployment's
// documentation. An operator who was handed a URL and an access secret has no
// checkout in front of them, and the network the application runs on is not assumed
// to reach a CDN — so the renderer is vendored and everything here is same-origin.
//
// Nothing on this page is authenticated. It is served from the static mount beside
// the application shell, exactly like `index.html`: the troubleshooting page is
// needed most by whoever cannot get past the access secret, and no documentation
// file contains anything a session would protect.
//
// `marked` and `mermaid` are loaded lazily, on the first documentation page and the
// first diagram. They are large, and every other view does without them.
import { el, errorBanner, setActiveNav } from "./dom.js";
const DOCS_BASE = "/docs";
const VENDOR = "/app/js/vendor";
const INDEX_PAGE = "index";
// A page name comes from the hash, so it is caller input: no scheme, no traversal,
// no absolute path. The server would refuse those too; this refuses them earlier and
// without a request that looks like an attempt.
const PAGE_PATTERN = /^[\w-]+(\/[\w-]+)*$/;
let markedPromise;
let mermaidPromise;
function loadMarked() {
markedPromise ??= import(`${VENDOR}/marked.esm.js`).then((module) => module.marked);
return markedPromise;
}
// mermaid ships one large UMD bundle rather than a self-contained ES module, so it
// arrives through a script element. `script-src 'self'` allows it because it is ours
// and same-origin; nothing here relaxes that.
function loadMermaid() {
mermaidPromise ??= new Promise((resolve, reject) => {
const script = document.createElement("script");
script.src = `${VENDOR}/mermaid.min.js`;
script.onload = () => resolve(window.mermaid);
script.onerror = () => reject(new Error("the diagram renderer could not be loaded"));
document.head.appendChild(script);
});
return mermaidPromise;
}
// ── the pure parts, unit-tested in js/tests/unit.js ──────────────────────────
/** A stable, readable heading anchor. Letters and digits of any script survive. */
export function slug(text) {
const cleaned = String(text)
.toLowerCase()
.trim()
.replace(/[^\p{L}\p{N}\s-]/gu, "")
.replace(/[\s-]+/g, "-")
.replace(/^-|-$/g, "");
return cleaned || "section";
}
/** Assign every heading an id, disambiguating repeats the way a reader would
* expect: the first `Notes` keeps `#notes`, the second becomes `#notes-2`. */
export function assignHeadingIds(headings) {
const used = new Map();
for (const heading of headings) {
const base = slug(heading.textContent);
const seen = (used.get(base) || 0) + 1;
used.set(base, seen);
heading.id = seen === 1 ? base : `${base}-${seen}`;
}
return headings;
}
/**
* Where a link inside a documentation page should go.
*
* Markdown links between documents are relative file paths, which a browser would
* treat as downloads that leave the application. Returns the in-app target for a
* link to another document or to a heading, and `null` for anything else — external
* links stay exactly as the author wrote them.
*/
export function resolveDocLink(currentPage, href) {
if (!href) return null;
if (/^[a-z][a-z\d+.-]*:/i.test(href) || href.startsWith("//")) return null;
if (href.startsWith("#")) return { page: currentPage, anchor: href.slice(1) };
const [target, anchor = ""] = href.split("#");
if (!/\.md$/i.test(target)) return null;
// Resolved against the current document, so `../guides/x.md` means what it means
// in the repository. The origin is a placeholder; only the path is used.
const resolved = new URL(target, `https://docs.invalid/${currentPage}.md`);
const page = decodeURIComponent(resolved.pathname).replace(/^\//, "").replace(/\.md$/i, "");
return PAGE_PATTERN.test(page) ? { page, anchor } : null;
}
/**
* Point every in-app link at its route, and leave every other link alone.
*
* The resolved page is kept on the element, so the reading order can be read back
* from the rendered index rather than parsed out of the markdown a second time.
*/
export function rewriteLinks(article, page) {
for (const link of article.querySelectorAll("a[href]")) {
const href = link.getAttribute("href");
const target = resolveDocLink(page, href);
if (target) {
link.setAttribute("href", docHash(target.page, target.anchor));
link.dataset.docPage = target.page;
} else if (/^https?:/i.test(href)) {
link.setAttribute("rel", "noreferrer noopener");
link.setAttribute("target", "_blank");
}
}
return article;
}
/** The reading order, taken from the index document itself — one list, in one
* place, that is equally the index on Gitea and the navigation here. */
export function documentIndex(indexElement) {
const pages = [];
for (const link of indexElement.querySelectorAll("a[data-doc-page]")) {
const page = link.dataset.docPage;
if (page !== INDEX_PAGE && !pages.some((entry) => entry.page === page)) {
pages.push({ page, title: link.textContent.trim() });
}
}
return pages;
}
// ── the view ─────────────────────────────────────────────────────────────────
async function fetchPage(page) {
const response = await fetch(`${DOCS_BASE}/${page}.md`, { headers: { accept: "text/markdown" } });
if (!response.ok) throw new Error(`documentation page unavailable: ${response.status}`);
return response.text();
}
async function toArticle(markdown, page) {
const marked = await loadMarked();
const article = el("article", { class: "doc", "data-testid": "doc", "data-page": page });
// The only innerHTML in the application, and deliberate: rendering markdown *is*
// producing HTML. The input is a file from this repository, and the page's CSP
// (`script-src 'self'`, no `'unsafe-inline'`) means injected script and inline
// handlers do not run even if one ever were not.
article.innerHTML = marked.parse(markdown, { async: false });
assignHeadingIds(article.querySelectorAll("h1, h2, h3, h4, h5, h6"));
rewriteLinks(article, page);
return article;
}
function docHash(page, anchor = "") {
const query = new URLSearchParams(anchor ? { page, anchor } : { page });
return `#/docs?${query}`;
}
/** Diagrams are authored as ```mermaid blocks so the source is diffable and Gitea
* renders them natively. A failure here replaces the diagram, never the page. */
async function renderDiagrams(article) {
const blocks = [...article.querySelectorAll("pre > code.language-mermaid")];
if (!blocks.length) return;
let mermaid;
try {
mermaid = await loadMermaid();
mermaid.initialize({ startOnLoad: false, securityLevel: "strict", theme: "dark" });
} catch (error) {
blocks.forEach((block) => block.closest("pre").replaceWith(errorBanner(error.message)));
return;
}
for (const [index, block] of blocks.entries()) {
const figure = el("figure", { class: "diagram", "data-testid": "diagram" });
block.closest("pre").replaceWith(figure);
try {
const { svg } = await mermaid.render(`diagram-${index}-${Date.now()}`, block.textContent);
figure.innerHTML = svg;
} catch (error) {
figure.replaceWith(errorBanner(`Diagram could not be drawn: ${error.message}`));
}
}
}
function notFound() {
return el(
"div",
{ class: "alert", role: "alert", "data-testid": "doc-not-found" },
"That documentation page does not exist. ",
el("a", { class: "link", href: docHash(INDEX_PAGE) }, "Back to the documentation index")
);
}
function sidebar(pages, current) {
return el(
"nav",
{ class: "card", "aria-label": "Documentation" },
el("h2", {}, "Documentation"),
el(
"ul",
{ class: "album-list", "data-testid": "doc-pages" },
el(
"li",
{},
el(
"a",
{
href: docHash(INDEX_PAGE),
"aria-current": current === INDEX_PAGE ? "true" : false,
},
"Index"
)
),
...pages.map(({ page, title }) =>
el(
"li",
{},
el(
"a",
{
href: docHash(page),
"data-page": page,
"aria-current": page === current ? "true" : false,
},
title
)
)
)
)
);
}
export async function renderDocs(root, params = {}) {
setActiveNav("docs");
const requested = params.page || INDEX_PAGE;
const page = PAGE_PATTERN.test(requested) ? requested : "";
let indexArticle;
try {
indexArticle = await toArticle(await fetchPage(INDEX_PAGE), INDEX_PAGE);
} catch (error) {
root.replaceChildren(errorBanner(`Documentation is unavailable: ${error.message}`));
return;
}
const pages = documentIndex(indexArticle);
let article;
if (!page) article = notFound();
else if (page === INDEX_PAGE) article = indexArticle;
else {
try {
article = await toArticle(await fetchPage(page), page);
} catch {
article = notFound();
}
}
root.replaceChildren(el("div", { class: "two-pane" }, sidebar(pages, page), article));
await renderDiagrams(article);
scrollToAnchor(params.anchor);
}
// A documentation anchor cannot live in the hash — the hash is the route — so it
// travels as a parameter and is applied after the page renders.
function scrollToAnchor(anchor) {
if (!anchor) return;
const target = document.getElementById(anchor);
if (target) target.scrollIntoView({ block: "start" });
}

View File

@@ -7,6 +7,7 @@ import { createStore } from "../store.js";
import { parseHash, navigate } from "../router.js";
import { api, cancellable } from "../api.js";
import { subscribeJob } from "../events.js";
import { assignHeadingIds, documentIndex, resolveDocLink, rewriteLinks, slug } from "../docs.js";
const cases = [];
function ok(name, cond) {
@@ -142,6 +143,62 @@ async function run() {
window.EventSource = realES;
}
// ── documentation view (US09-01) ─────────────────────────────────────────
{
ok("slug lowercases and dashes a heading", slug("Key Flows") === "key-flows");
ok(
"slug drops punctuation but keeps words",
slug("What it *will not* do:") === "what-it-will-not-do"
);
ok("slug keeps non-ASCII letters", slug("Größe & Gewicht") === "größe-gewicht");
ok("slug never yields an empty anchor", slug("!!!") === "section");
const doc = document.createElement("div");
doc.innerHTML = "<h2>Notes</h2><h3>Notes</h3><h2>Notes</h2>";
const ids = [...assignHeadingIds(doc.querySelectorAll("h2, h3"))].map((h) => h.id);
ok("repeated headings get distinct anchors", ids.join(",") === "notes,notes-2,notes-3");
ok(
"a sibling document link becomes an in-app route",
resolveDocLink("index", "overview.md")?.page === "overview"
);
const nested = resolveDocLink("guides/install", "../overview.md#running-it");
ok("a relative link resolves against the current page", nested?.page === "overview");
ok("a link's anchor survives the rewrite", nested?.anchor === "running-it");
ok(
"a bare anchor stays on the current page",
resolveDocLink("overview", "#the-workflow")?.page === "overview"
);
ok("an external link is left alone", resolveDocLink("index", "https://example.test/x") === null);
ok("a non-markdown relative link is left alone", resolveDocLink("index", "images/a.png") === null);
// A climb cannot leave the documentation tree: it is clamped at the root, so the
// page it names is still fetched from under /docs and simply does not exist.
ok(
"a link that climbs above the docs root is clamped",
resolveDocLink("index", "../../etc/passwd.md")?.page === "etc/passwd"
);
ok(
"a page name that is not a page name is refused",
resolveDocLink("index", "..%2f..%2fetc%2fpasswd.md") === null
);
const index = document.createElement("div");
index.innerHTML =
'<a href="overview.md">Overview</a><a href="overview.md">again</a>' +
'<a href="index.md">itself</a><a href="https://example.test">out</a>';
const listed = documentIndex(rewriteLinks(index, "index"));
ok(
"an external link is marked safe to open away from the app",
index.querySelector('a[href^="https"]').rel === "noreferrer noopener"
);
ok(
"an in-app link points at the docs route",
index.querySelector("a").getAttribute("href") === "#/docs?page=overview"
);
ok("the index lists each page once, in order", listed.length === 1);
ok("the index takes its titles from the link text", listed[0].title === "Overview");
}
await tick();
const failed = cases.filter((c) => !c.ok);
window.__RESULTS__ = { passed: cases.length - failed.length, failed: failed.length, cases };

23
frontend/js/vendor/VERSIONS.json vendored Normal file
View File

@@ -0,0 +1,23 @@
{
"_comment": "Vendored browser libraries (US09-01). Committed rather than fetched: the application is deployed to a network whose outbound access is not assumed, and `default-src 'self'` forbids a CDN. Checksums are asserted by tests/integration/test_documentation.py, so replacing a file without recording it here fails the suite. Update by downloading the pinned URL and recording the new version and sha256 in the same commit.",
"libraries": [
{
"file": "marked.esm.js",
"name": "marked",
"version": "18.0.10",
"license": "MIT",
"url": "https://cdn.jsdelivr.net/npm/marked@18.0.10/lib/marked.esm.js",
"sha256": "4cf47dfebb7f614a08fc0a579ab0fe407ff0ed2b717bf953040c85b2f493a4f0",
"why": "Markdown to HTML for the documentation view. An ES module with no dependencies, imported lazily by js/docs.js."
},
{
"file": "mermaid.min.js",
"name": "mermaid",
"version": "11.17.0",
"license": "MIT",
"url": "https://cdn.jsdelivr.net/npm/mermaid@11.17.0/dist/mermaid.min.js",
"sha256": "8d8e0eec56d3a83b4b3c87f42050845546dee93ebe1875d2117c12e6947c0cb3",
"why": "Renders ```mermaid blocks so diagrams are diffable source that Gitea also renders. The single-file UMD build rather than the ES module entry, whose ~40 lazy chunks would each need vendoring and pinning."
}
]
}

78
frontend/js/vendor/marked.esm.js vendored Normal file

File diff suppressed because one or more lines are too long

3636
frontend/js/vendor/mermaid.min.js vendored Normal file

File diff suppressed because one or more lines are too long

View File

@@ -53,6 +53,7 @@ from photo_pipeline.services.thumbnails import ThumbnailService
from photo_pipeline.services.upload_batches import UploadBatchService
FRONTEND_DIR = Path(__file__).resolve().parents[2] / "frontend"
DOCS_DIR = Path(__file__).resolve().parents[2] / "docs"
log = logging.getLogger(__name__)
@@ -151,4 +152,10 @@ def create_app(config: Config | None = None) -> FastAPI:
# Static single-page app (hash-routed). Mounted last so /api/v1 wins.
if FRONTEND_DIR.is_dir():
app.mount("/app", StaticFiles(directory=FRONTEND_DIR, html=True), name="app")
# The manuals, as their own markdown files (US09-01). Unauthenticated for the
# same reason the application shell is: whoever cannot get past the access
# secret is exactly who needs the troubleshooting page, and no document holds
# anything a session would protect.
if DOCS_DIR.is_dir():
app.mount("/docs", StaticFiles(directory=DOCS_DIR), name="docs")
return app

View File

@@ -57,9 +57,21 @@ LOOPBACK_HOSTS = frozenset({"127.0.0.1", "localhost", "::1", "[::1]"})
# backup a careful operator takes first, and its retention (US07-07).
MUTATION_EXEMPT_PATHS = frozenset({f"{API_PREFIX}/backups", f"{API_PREFIX}/backups/prune"})
# Applied to every response. No inline script/style is used by the frontend, so the
# policy can stay strict; `frame-ancestors 'none'` and CORP keep other pages from
# Applied to every response. `frame-ancestors 'none'` and CORP keep other pages from
# embedding the app or its thumbnails.
#
# `script-src 'self'` is the boundary that matters and it is unchanged: no
# `'unsafe-eval'`, no `'unsafe-inline'`, so nothing injected into the DOM can execute.
# Whether the vendored diagram renderer needed `'unsafe-eval'` was measured rather
# than assumed — mermaid's bundle contains no `eval(` and no `new Function`, and it
# renders under this exact policy without a single script-src violation.
#
# `style-src` does gain `'unsafe-inline'` (US09-01): mermaid styles the SVG it builds
# with an injected `<style>` element and `style=` attributes, and a diagram's CSS
# cannot be hashed in advance. The concession is bounded by the directives around it
# — with script execution still refused and `img-src`, `connect-src`, and `font-src`
# all `'self'`, the CSS exfiltration channels stay closed and what is left is
# defacement of a page its own operator is already looking at.
DEFAULT_HEADERS = {
"x-content-type-options": "nosniff",
"x-frame-options": "DENY",
@@ -67,8 +79,9 @@ DEFAULT_HEADERS = {
"cross-origin-resource-policy": "same-origin",
"cross-origin-opener-policy": "same-origin",
"content-security-policy": (
"default-src 'self'; img-src 'self' data:; style-src 'self'; script-src 'self'; "
"connect-src 'self'; frame-ancestors 'none'; base-uri 'none'; form-action 'none'"
"default-src 'self'; img-src 'self' data:; style-src 'self' 'unsafe-inline'; "
"script-src 'self'; connect-src 'self'; font-src 'self'; "
"frame-ancestors 'none'; base-uri 'none'; form-action 'none'"
),
}

119
tests/e2e/test_docs_ui.py Normal file
View File

@@ -0,0 +1,119 @@
"""US09-01: the manuals, in the browser, under the application's real policy.
Everything here runs against a real ``photo_pipeline serve`` process serving the
committed ``docs/`` tree and the vendored renderer — no fixture markdown, no stubbed
fetch. What that buys is the assertion the story actually cares about: the pages
render, the links between them work, a diagram becomes a diagram, and the console
stays empty, including of CSP violations.
The offline contract — reachability, dead links, pinned checksums — is
``tests/integration/test_documentation.py``.
"""
from __future__ import annotations
import pytest
from tests.e2e._pipeline_harness import Server, seed_library
@pytest.fixture(scope="module")
def server(tmp_path_factory):
seeded = seed_library(tmp_path_factory.mktemp("docs"), {"a": 1}, {})
running = Server(seeded).start()
try:
yield running
finally:
running.stop()
@pytest.fixture
def quiet(page):
"""Any console error, page error, or failed request fails the test that caused
it. A CSP violation arrives as a console error, which is the point."""
problems: list[str] = []
page.on("console", lambda m: problems.append(m.text) if m.type == "error" else None)
page.on("pageerror", lambda error: problems.append(str(error)))
page.on("requestfailed", lambda request: problems.append(f"failed request: {request.url}"))
yield problems
def test_the_documentation_opens_from_the_navigation(page, server, quiet):
page.goto(f"{server.base}/app/#/workflow")
page.locator('nav a[data-nav="docs"]').click()
page.get_by_test_id("doc").wait_for()
assert page.get_by_test_id("doc").get_attribute("data-page") == "index"
# The sidebar is the index's own reading order, not a second list to maintain.
assert page.get_by_test_id("doc-pages").get_by_role("link", name="Overview").is_visible()
assert quiet == []
def test_a_link_between_documents_stays_inside_the_application(page, server, quiet):
page.goto(f"{server.base}/app/#/docs")
page.get_by_test_id("doc").wait_for()
page.get_by_test_id("doc").get_by_role("link", name="Overview").click()
page.wait_for_selector('[data-testid="doc"][data-page="overview"]')
assert "#/docs?page=overview" in page.url, "a relative .md link must not leave the app"
# And back again, by the link the document itself carries.
page.get_by_test_id("doc").get_by_role("link", name="← Documentation index").click()
page.wait_for_selector('[data-testid="doc"][data-page="index"]')
assert quiet == []
def test_a_deep_link_to_a_heading_lands_on_that_heading(page, server, quiet):
page.goto(f"{server.base}/app/#/docs?page=overview&anchor=the-workflow")
heading = page.locator("#the-workflow")
heading.wait_for()
assert heading.inner_text().strip() == "The workflow"
assert heading.evaluate("node => node.getBoundingClientRect().top < window.innerHeight")
assert quiet == []
def test_a_mermaid_block_becomes_a_diagram_under_the_unchanged_script_policy(page, server, quiet):
page.goto(f"{server.base}/app/#/docs?page=overview")
diagram = page.get_by_test_id("diagram").first
diagram.wait_for()
# A real drawing, not the source text and not an error node.
assert diagram.locator("svg").count() == 1
assert diagram.evaluate("node => node.querySelector('svg').getBBox().width") > 100
assert page.locator("pre code.language-mermaid").count() == 0
assert quiet == [], "rendering a diagram must not violate the policy"
policy = page.evaluate(
"async () => (await fetch('/app/')).headers.get('content-security-policy')"
)
assert "script-src 'self';" in policy
assert "unsafe-eval" not in policy
def test_an_unknown_page_says_so_without_naming_a_path(page, server, quiet):
page.goto(f"{server.base}/app/#/docs?page=no-such-manual")
message = page.get_by_test_id("doc-not-found")
message.wait_for()
text = message.inner_text()
assert "does not exist" in text
assert "/" not in text.replace("Back to the documentation index", ""), text
message.get_by_role("link").click()
page.wait_for_selector('[data-testid="doc"][data-page="index"]')
# The missing page is a 404 and the browser says so; nothing else may go wrong,
# and in particular the view must not throw on the way to its own message.
assert all("404" in problem for problem in quiet), quiet
def test_the_documentation_is_readable_without_a_session(page, server, quiet):
"""The troubleshooting page is needed most by whoever is locked out."""
page.goto(f"{server.base}/app/#/docs")
page.get_by_test_id("doc").wait_for()
unauthenticated = page.evaluate(
"""async () => {
const response = await fetch('/docs/index.md', { credentials: 'omit' });
return { status: response.status, length: (await response.text()).length };
}"""
)
assert unauthenticated["status"] == 200
assert unauthenticated["length"] > 100

View File

@@ -0,0 +1,179 @@
"""US09-01: the documentation the application serves, checked without a browser.
What is checkable offline is everything that makes the manuals *navigable* rather
than merely present: every page reachable, every internal link and anchor landing
somewhere real, the vendored renderer being the file that was pinned, and the
application actually serving `docs/` in both the working copy and the image.
The rendering itself — marked, mermaid, anchors, link rewriting — is proven in a
real browser by ``tests/e2e/test_docs_ui.py``.
"""
from __future__ import annotations
import hashlib
import json
import re
from pathlib import Path
from starlette.testclient import TestClient
from photo_pipeline.api.app import create_app
from photo_pipeline.api.security import DEFAULT_HEADERS
from photo_pipeline.config import Config
REPO = Path(__file__).resolve().parents[2]
DOCS = REPO / "docs"
VENDOR = REPO / "frontend" / "js" / "vendor"
INDEX = DOCS / "index.md"
MARKDOWN_LINK = re.compile(r"\[[^\]]*\]\(([^)\s]+)\)")
HEADING = re.compile(r"^#{1,6}\s+(.+?)\s*$", re.MULTILINE)
FENCE = re.compile(r"^```.*?^```", re.MULTILINE | re.DOTALL)
def pages() -> list[Path]:
return sorted(DOCS.rglob("*.md"))
def body(path: Path) -> str:
"""A document without its fenced code, so an example link in a shell snippet is
not mistaken for a link the reader can follow."""
return FENCE.sub("", path.read_text())
def slug(text: str) -> str:
"""The heading-anchor rule, mirroring ``frontend/js/docs.js``.
Deliberately duplicated: this check has to run without a browser. The browser
test is the authority on the real behaviour, and it asserts the same anchors.
"""
cleaned = re.sub(r"[^\w\s-]", "", text.lower(), flags=re.UNICODE).strip()
return re.sub(r"[\s-]+", "-", cleaned).strip("-") or "section"
def anchors(path: Path) -> set[str]:
found: dict[str, int] = {}
for heading in HEADING.findall(body(path)):
# Markdown emphasis and inline code are not part of the rendered text.
plain = re.sub(r"[*`_]", "", heading)
base = slug(plain)
found[base] = found.get(base, 0) + 1
return {name if index == 1 else f"{name}-{index}" for name, count in found.items() for index in range(1, count + 1)}
def internal_links(path: Path) -> list[str]:
return [
href
for href in MARKDOWN_LINK.findall(body(path))
if not re.match(r"^[a-z][a-z\d+.-]*:", href, re.IGNORECASE) and not href.startswith("//")
]
# ── the documentation tree ───────────────────────────────────────────────────
def test_the_index_exists_and_names_every_page():
"""A page nobody links to is a page nobody reads."""
assert INDEX.is_file()
listed = {
(INDEX.parent / href.split("#")[0]).resolve()
for href in internal_links(INDEX)
if href.split("#")[0].endswith(".md")
}
unreachable = [page.name for page in pages() if page != INDEX and page.resolve() not in listed]
assert unreachable == [], f"not linked from the index: {unreachable}"
def test_every_internal_link_and_anchor_resolves():
broken: list[str] = []
for page in pages():
for href in internal_links(page):
target, _, anchor = href.partition("#")
destination = page if not target else (page.parent / target).resolve()
if not destination.is_file():
broken.append(f"{page.name}{href} (no such file)")
continue
if anchor and anchor not in anchors(destination):
broken.append(f"{page.name}{href} (no such heading)")
assert broken == [], broken
def test_every_page_starts_with_one_title():
for page in pages():
titles = [line for line in page.read_text().splitlines() if line.startswith("# ")]
assert len(titles) == 1, f"{page.name} has {len(titles)} top-level titles"
# ── the vendored renderer ────────────────────────────────────────────────────
def test_the_vendored_libraries_are_the_files_that_were_pinned():
"""A vendored dependency without a recorded checksum is a dependency nobody is
reviewing. Replacing one has to be a visible change to this manifest."""
manifest = json.loads((VENDOR / "VERSIONS.json").read_text())
recorded = {entry["file"]: entry for entry in manifest["libraries"]}
on_disk = {path.name for path in VENDOR.iterdir() if path.suffix in (".js", ".mjs")}
assert on_disk == set(recorded), f"unrecorded vendored files: {on_disk ^ set(recorded)}"
for name, entry in recorded.items():
digest = hashlib.sha256((VENDOR / name).read_bytes()).hexdigest()
assert digest == entry["sha256"], f"{name} does not match the pinned checksum"
assert entry["version"] in entry["url"], f"{name}: the pinned URL and version disagree"
assert entry["license"], f"{name}: no license recorded"
def test_the_documentation_view_is_wired_into_the_shell():
shell = (REPO / "frontend" / "index.html").read_text()
assert 'data-nav="docs"' in shell, "the manuals are unreachable from the navigation"
assert '"/docs"' in (REPO / "frontend" / "js" / "app.js").read_text()
# ── the policy the view runs under ───────────────────────────────────────────
def test_the_script_boundary_is_unchanged_and_only_style_was_relaxed():
"""The measured cost of rendering diagrams in the browser, held to that cost.
mermaid needs to style the SVG it builds, which `style-src 'self'` refuses. It
does not need to execute generated code, so `script-src` must never acquire the
escape hatch that would let injected markup run.
"""
policy = DEFAULT_HEADERS["content-security-policy"]
directives = {
part.split(" ")[0]: part.split(" ")[1:]
for part in (piece.strip() for piece in policy.split(";"))
if part
}
assert directives["script-src"] == ["'self'"]
assert "'unsafe-inline'" in directives["style-src"]
for directive in ("default-src", "img-src", "connect-src", "font-src"):
assert "'self'" in directives[directive]
assert "'unsafe-inline'" not in directives[directive]
assert directives["frame-ancestors"] == ["'none'"]
# ── serving it ───────────────────────────────────────────────────────────────
def test_the_application_serves_the_markdown_without_a_session(tmp_path):
"""Whoever cannot get past the access secret is exactly who needs these pages."""
config = Config.from_env({"PHOTO_PIPELINE_DATA_DIR": str(tmp_path / "data")})
(tmp_path / "data").mkdir()
with TestClient(create_app(config), raise_server_exceptions=False) as client:
response = client.get("/docs/index.md", headers={"cookie": ""})
assert response.status_code == 200
assert "Photo Pipeline documentation" in response.text
assert client.get("/docs/overview.md").status_code == 200
assert client.get("/docs/nothing-here.md").status_code == 404
# The mount is a directory, not a path parameter: nothing above it is reachable.
assert client.get("/docs/../pyproject.toml").status_code in (307, 404)
assert client.get("/docs/%2e%2e/pyproject.toml").status_code == 404
def test_the_image_ships_the_documentation_it_serves():
"""An image without `docs/` serves an empty manual — and the build context is
deny-by-default, so a new directory is excluded until it is named."""
assert "\n!docs\n" in (REPO / ".dockerignore").read_text()
assert "COPY docs ./docs" in (REPO / "Dockerfile").read_text()

View File

@@ -193,10 +193,13 @@
"US08-05": [
"tests/integration/test_container_gate.py",
"tests/e2e/test_phase_h_container.py"
],
"US09-01": [
"tests/integration/test_documentation.py",
"tests/e2e/test_docs_ui.py"
]
},
"planned": [
"US09-01",
"US09-02",
"US09-03",
"US09-04",