Compare commits
1 Commits
us/US09-01
...
chore/E09-
| Author | SHA1 | Date | |
|---|---|---|---|
| 792134c72d |
@@ -9,7 +9,6 @@
|
||||
!photo_pipeline
|
||||
!migrations
|
||||
!frontend
|
||||
!docs
|
||||
!docker
|
||||
|
||||
# Nothing generated, even under an allowed directory.
|
||||
|
||||
9
.gitattributes
vendored
9
.gitattributes
vendored
@@ -1,9 +0,0 @@
|
||||
# Vendored third-party browser bundles (US09-01). Minified upstream files carry
|
||||
# trailing whitespace and very long lines; they are not ours to reformat, and the
|
||||
# submit gate's `git diff --check` would refuse them forever. Their integrity is
|
||||
# controlled where it belongs instead: a pinned version and a recorded sha256 in
|
||||
# frontend/js/vendor/VERSIONS.json, asserted by tests/integration/test_documentation.py.
|
||||
#
|
||||
# Only the whitespace check is turned off. They stay text, so a credential scan or a
|
||||
# search still reads them.
|
||||
frontend/js/vendor/*.js -whitespace linguist-vendored
|
||||
@@ -70,8 +70,6 @@ COPY pyproject.toml alembic.ini README.md ./
|
||||
COPY photo_pipeline ./photo_pipeline
|
||||
COPY migrations ./migrations
|
||||
COPY frontend ./frontend
|
||||
# The manuals are served by the application itself (US09-01), so they ship with it.
|
||||
COPY docs ./docs
|
||||
COPY docker/entrypoint.sh docker/healthcheck.sh /usr/local/bin/
|
||||
|
||||
# Runtime dependencies only: the `test` extra (pytest, playwright) and the `vision`
|
||||
|
||||
@@ -1,34 +0,0 @@
|
||||
# Photo Pipeline documentation
|
||||
|
||||
One local application that takes a photo library from discovery to a verified Immich
|
||||
upload: duplicate detection, safety review, content analysis, album naming, guarded
|
||||
renaming, upload, and archive — one visible, resume-safe workflow.
|
||||
|
||||
These pages are readable three ways, and they are the same files each time: in the
|
||||
repository under `docs/`, on Gitea, and inside the running application under
|
||||
**Docs**. There is no separate copy to fall out of date.
|
||||
|
||||
## Read in this order
|
||||
|
||||
1. [Overview](overview.md) — what the application does, the stages it moves a photo
|
||||
through, and the rules it will not break.
|
||||
|
||||
## Being written
|
||||
|
||||
The remaining manuals are accepted work, not aspiration; each is a story in
|
||||
[E09](https://git.domverse-berlin.eu/domverse/photoanalyzer/src/branch/main/delivery_backlog/E09-documentation.md)
|
||||
and will appear here as it lands.
|
||||
|
||||
- **Installation and operations** — host and container installation, every setting,
|
||||
the first-run checklist, upgrades, backup and restore, and what each refusal at
|
||||
startup means (US09-02).
|
||||
- **Architecture** — context and runtime diagrams, the module map, the job and
|
||||
journal state machines, and where each invariant is enforced (US09-03).
|
||||
- **User manual** — one page per workflow stage with screenshots of the real
|
||||
application, and a catalogue of every error and refusal (US09-04).
|
||||
|
||||
## Conventions
|
||||
|
||||
A page tells you what a stage **changes on disk or on the server** before it tells
|
||||
you how to run it. Refusals are documented as intentional: this application would
|
||||
rather stop and explain than guess about somebody's photographs.
|
||||
@@ -1,77 +0,0 @@
|
||||
# Overview
|
||||
|
||||
[← Documentation index](index.md)
|
||||
|
||||
Photo Pipeline sorts a photo library. It finds duplicates before anything expensive
|
||||
happens to them, asks a human which pictures may leave the machine, describes the
|
||||
ones that may, proposes album names from what it found, renames folders under a
|
||||
crash-safe journal, uploads through `immich-go`, and can archive a finished album off
|
||||
active storage without ever forgetting it existed.
|
||||
|
||||
It runs locally. It talks to exactly three things outside itself: a vision provider,
|
||||
an Immich server, and `exiftool` — and it will tell you before it uses any of them.
|
||||
|
||||
## The workflow
|
||||
|
||||
Each stage is a gate, not a tab. A later stage can always be looked at; its actions
|
||||
stay disabled until what they depend on is true.
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
I["0 · Inventory<br/>discover, hash, cluster duplicates"] --> S["1 · Safety<br/>score and human decision"]
|
||||
S -->|sfw| A["2 · Analysis<br/>vision provider, EXIF checkpoint"]
|
||||
S -->|nsfw| U
|
||||
A --> B["3 · Albums<br/>proposal, then guarded rename"]
|
||||
B --> U["4 · Upload<br/>immich-go, verified bytes"]
|
||||
U --> R["5 · Archive<br/>copy, verify, then reclaim space"]
|
||||
```
|
||||
|
||||
| Stage | What it decides | What it changes |
|
||||
|---|---|---|
|
||||
| Inventory | which file is the canonical copy of a picture | nothing — it only reads |
|
||||
| Safety | whether a photo may be sent to a cloud provider | one `sfw`/`nsfw` EXIF keyword |
|
||||
| Analysis | what a photo shows | a managed caption segment and additive keywords |
|
||||
| Albums | what a folder should be called | folder names, through a journaled rename |
|
||||
| Upload | which exact bytes reach Immich | nothing locally; assets appear in Immich |
|
||||
| Archive | which album leaves active storage | files move to the archive, after verification |
|
||||
|
||||
## What it will not do
|
||||
|
||||
These are enforced in code, not by convention, and each one is why some action you
|
||||
expected is sometimes refused.
|
||||
|
||||
- **Nothing under `_IGNORE/` is ever read.** Not scanned, not counted, not
|
||||
thumbnailed, not sent anywhere.
|
||||
- **A path is not an identity.** Every picture has a stable id, so moving or renaming
|
||||
it loses no history.
|
||||
- **Only a confirmed-SFW photo reaches the vision provider.** A photo marked NSFW is
|
||||
still uploadable to Immich; it simply never leaves for analysis.
|
||||
- **Metadata is verified, not hoped for.** Every stage that writes EXIF reads it back
|
||||
and proves that what it did not own is unchanged.
|
||||
- **One writer at a time.** A library lock, held by one process, is what makes a
|
||||
crash recoverable instead of ambiguous.
|
||||
- **Nothing irreversible happens without a preview and an explicit approval** that
|
||||
names the exact count.
|
||||
|
||||
## Where things live
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| the photo library | wherever you point `PHOTO_PIPELINE_LIBRARY_ROOTS`; mounted read-write |
|
||||
| the database, cache, logs, backups | the data directory, never inside the library |
|
||||
| secrets | the environment, never the database and never a log line |
|
||||
| these documents | `docs/` in the repository, served at `/docs` by the application |
|
||||
|
||||
## Running it
|
||||
|
||||
The short version, for a host installation:
|
||||
|
||||
```bash
|
||||
python -m photo_pipeline migrate
|
||||
python -m photo_pipeline serve # the API and this browser application
|
||||
python -m photo_pipeline worker # the process that does the long work
|
||||
```
|
||||
|
||||
The full procedure, the container path, and every setting belong to the installation
|
||||
manual, which is being written next — see
|
||||
[the documentation index](index.md#being-written).
|
||||
@@ -170,20 +170,6 @@ a.link:hover { text-decoration: underline; }
|
||||
.grid th, .grid td { text-align: left; padding: 6px 8px; border-bottom: 1px solid var(--border); vertical-align: top; }
|
||||
.path { font-family: ui-monospace, monospace; font-size: 0.85em; word-break: break-all; }
|
||||
|
||||
/* Docs: prose, so it gets a reading measure rather than the full window width. */
|
||||
.doc { max-width: 72ch; }
|
||||
.doc h1 { margin-top: 0; }
|
||||
.doc h2, .doc h3 { margin-top: 1.6em; border-bottom: 1px solid var(--border); padding-bottom: 4px; }
|
||||
.doc code { font-family: ui-monospace, monospace; font-size: 0.9em; background: var(--surface-2); padding: 1px 4px; border-radius: 4px; }
|
||||
.doc pre { background: var(--surface-2); border: 1px solid var(--border); border-radius: var(--radius); padding: 12px; overflow: auto; }
|
||||
.doc pre code { background: none; padding: 0; }
|
||||
.doc table { border-collapse: collapse; width: 100%; }
|
||||
.doc th, .doc td { text-align: left; padding: 6px 8px; border-bottom: 1px solid var(--border); vertical-align: top; }
|
||||
.doc blockquote { margin: 1em 0; padding-left: 12px; border-left: 3px solid var(--border); color: var(--muted); }
|
||||
.doc img { max-width: 100%; }
|
||||
.diagram { margin: 1.5em 0; padding: 0; overflow: auto; }
|
||||
.diagram svg { max-width: 100%; height: auto; }
|
||||
|
||||
/* Narrow screens: stack the panes so blockers and actions stay reachable. */
|
||||
@media (max-width: 720px) {
|
||||
.two-pane { grid-template-columns: 1fr; }
|
||||
|
||||
@@ -21,7 +21,6 @@
|
||||
<a href="#/uploads" data-nav="uploads">Upload</a>
|
||||
<a href="#/archive" data-nav="archive">Archive</a>
|
||||
<a href="#/stats" data-nav="stats">Stats</a>
|
||||
<a href="#/docs" data-nav="docs">Docs</a>
|
||||
</nav>
|
||||
</header>
|
||||
<main id="app" aria-live="polite"><!-- views render here --></main>
|
||||
|
||||
@@ -1,6 +1,5 @@
|
||||
import { api } from "./api.js";
|
||||
import { renderArchive, setArchiveRender } from "./archive.js";
|
||||
import { renderDocs } from "./docs.js";
|
||||
import { navigate, onRouteChange, parseHash } from "./router.js";
|
||||
import { renderRenames, setRenamesRender } from "./renames.js";
|
||||
import { renderUploads, setUploadsRender } from "./uploads.js";
|
||||
@@ -386,7 +385,6 @@ function render() {
|
||||
else if (path === "/uploads") renderUploads(root, params);
|
||||
else if (path === "/archive") renderArchive(root, params);
|
||||
else if (path === "/stats") renderStats(root, params);
|
||||
else if (path === "/docs") renderDocs(root, params);
|
||||
else show(errorBanner("Unknown view"));
|
||||
}
|
||||
|
||||
|
||||
@@ -1,263 +0,0 @@
|
||||
// The documentation view (US09-01): the manuals in `docs/`, rendered in the app.
|
||||
//
|
||||
// The same markdown files are the repository's documentation and the deployment's
|
||||
// documentation. An operator who was handed a URL and an access secret has no
|
||||
// checkout in front of them, and the network the application runs on is not assumed
|
||||
// to reach a CDN — so the renderer is vendored and everything here is same-origin.
|
||||
//
|
||||
// Nothing on this page is authenticated. It is served from the static mount beside
|
||||
// the application shell, exactly like `index.html`: the troubleshooting page is
|
||||
// needed most by whoever cannot get past the access secret, and no documentation
|
||||
// file contains anything a session would protect.
|
||||
//
|
||||
// `marked` and `mermaid` are loaded lazily, on the first documentation page and the
|
||||
// first diagram. They are large, and every other view does without them.
|
||||
import { el, errorBanner, setActiveNav } from "./dom.js";
|
||||
|
||||
const DOCS_BASE = "/docs";
|
||||
const VENDOR = "/app/js/vendor";
|
||||
const INDEX_PAGE = "index";
|
||||
// A page name comes from the hash, so it is caller input: no scheme, no traversal,
|
||||
// no absolute path. The server would refuse those too; this refuses them earlier and
|
||||
// without a request that looks like an attempt.
|
||||
const PAGE_PATTERN = /^[\w-]+(\/[\w-]+)*$/;
|
||||
|
||||
let markedPromise;
|
||||
let mermaidPromise;
|
||||
|
||||
function loadMarked() {
|
||||
markedPromise ??= import(`${VENDOR}/marked.esm.js`).then((module) => module.marked);
|
||||
return markedPromise;
|
||||
}
|
||||
|
||||
// mermaid ships one large UMD bundle rather than a self-contained ES module, so it
|
||||
// arrives through a script element. `script-src 'self'` allows it because it is ours
|
||||
// and same-origin; nothing here relaxes that.
|
||||
function loadMermaid() {
|
||||
mermaidPromise ??= new Promise((resolve, reject) => {
|
||||
const script = document.createElement("script");
|
||||
script.src = `${VENDOR}/mermaid.min.js`;
|
||||
script.onload = () => resolve(window.mermaid);
|
||||
script.onerror = () => reject(new Error("the diagram renderer could not be loaded"));
|
||||
document.head.appendChild(script);
|
||||
});
|
||||
return mermaidPromise;
|
||||
}
|
||||
|
||||
// ── the pure parts, unit-tested in js/tests/unit.js ──────────────────────────
|
||||
|
||||
/** A stable, readable heading anchor. Letters and digits of any script survive. */
|
||||
export function slug(text) {
|
||||
const cleaned = String(text)
|
||||
.toLowerCase()
|
||||
.trim()
|
||||
.replace(/[^\p{L}\p{N}\s-]/gu, "")
|
||||
.replace(/[\s-]+/g, "-")
|
||||
.replace(/^-|-$/g, "");
|
||||
return cleaned || "section";
|
||||
}
|
||||
|
||||
/** Assign every heading an id, disambiguating repeats the way a reader would
|
||||
* expect: the first `Notes` keeps `#notes`, the second becomes `#notes-2`. */
|
||||
export function assignHeadingIds(headings) {
|
||||
const used = new Map();
|
||||
for (const heading of headings) {
|
||||
const base = slug(heading.textContent);
|
||||
const seen = (used.get(base) || 0) + 1;
|
||||
used.set(base, seen);
|
||||
heading.id = seen === 1 ? base : `${base}-${seen}`;
|
||||
}
|
||||
return headings;
|
||||
}
|
||||
|
||||
/**
|
||||
* Where a link inside a documentation page should go.
|
||||
*
|
||||
* Markdown links between documents are relative file paths, which a browser would
|
||||
* treat as downloads that leave the application. Returns the in-app target for a
|
||||
* link to another document or to a heading, and `null` for anything else — external
|
||||
* links stay exactly as the author wrote them.
|
||||
*/
|
||||
export function resolveDocLink(currentPage, href) {
|
||||
if (!href) return null;
|
||||
if (/^[a-z][a-z\d+.-]*:/i.test(href) || href.startsWith("//")) return null;
|
||||
if (href.startsWith("#")) return { page: currentPage, anchor: href.slice(1) };
|
||||
const [target, anchor = ""] = href.split("#");
|
||||
if (!/\.md$/i.test(target)) return null;
|
||||
// Resolved against the current document, so `../guides/x.md` means what it means
|
||||
// in the repository. The origin is a placeholder; only the path is used.
|
||||
const resolved = new URL(target, `https://docs.invalid/${currentPage}.md`);
|
||||
const page = decodeURIComponent(resolved.pathname).replace(/^\//, "").replace(/\.md$/i, "");
|
||||
return PAGE_PATTERN.test(page) ? { page, anchor } : null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Point every in-app link at its route, and leave every other link alone.
|
||||
*
|
||||
* The resolved page is kept on the element, so the reading order can be read back
|
||||
* from the rendered index rather than parsed out of the markdown a second time.
|
||||
*/
|
||||
export function rewriteLinks(article, page) {
|
||||
for (const link of article.querySelectorAll("a[href]")) {
|
||||
const href = link.getAttribute("href");
|
||||
const target = resolveDocLink(page, href);
|
||||
if (target) {
|
||||
link.setAttribute("href", docHash(target.page, target.anchor));
|
||||
link.dataset.docPage = target.page;
|
||||
} else if (/^https?:/i.test(href)) {
|
||||
link.setAttribute("rel", "noreferrer noopener");
|
||||
link.setAttribute("target", "_blank");
|
||||
}
|
||||
}
|
||||
return article;
|
||||
}
|
||||
|
||||
/** The reading order, taken from the index document itself — one list, in one
|
||||
* place, that is equally the index on Gitea and the navigation here. */
|
||||
export function documentIndex(indexElement) {
|
||||
const pages = [];
|
||||
for (const link of indexElement.querySelectorAll("a[data-doc-page]")) {
|
||||
const page = link.dataset.docPage;
|
||||
if (page !== INDEX_PAGE && !pages.some((entry) => entry.page === page)) {
|
||||
pages.push({ page, title: link.textContent.trim() });
|
||||
}
|
||||
}
|
||||
return pages;
|
||||
}
|
||||
|
||||
// ── the view ─────────────────────────────────────────────────────────────────
|
||||
|
||||
async function fetchPage(page) {
|
||||
const response = await fetch(`${DOCS_BASE}/${page}.md`, { headers: { accept: "text/markdown" } });
|
||||
if (!response.ok) throw new Error(`documentation page unavailable: ${response.status}`);
|
||||
return response.text();
|
||||
}
|
||||
|
||||
async function toArticle(markdown, page) {
|
||||
const marked = await loadMarked();
|
||||
const article = el("article", { class: "doc", "data-testid": "doc", "data-page": page });
|
||||
// The only innerHTML in the application, and deliberate: rendering markdown *is*
|
||||
// producing HTML. The input is a file from this repository, and the page's CSP
|
||||
// (`script-src 'self'`, no `'unsafe-inline'`) means injected script and inline
|
||||
// handlers do not run even if one ever were not.
|
||||
article.innerHTML = marked.parse(markdown, { async: false });
|
||||
assignHeadingIds(article.querySelectorAll("h1, h2, h3, h4, h5, h6"));
|
||||
rewriteLinks(article, page);
|
||||
return article;
|
||||
}
|
||||
|
||||
function docHash(page, anchor = "") {
|
||||
const query = new URLSearchParams(anchor ? { page, anchor } : { page });
|
||||
return `#/docs?${query}`;
|
||||
}
|
||||
|
||||
/** Diagrams are authored as ```mermaid blocks so the source is diffable and Gitea
|
||||
* renders them natively. A failure here replaces the diagram, never the page. */
|
||||
async function renderDiagrams(article) {
|
||||
const blocks = [...article.querySelectorAll("pre > code.language-mermaid")];
|
||||
if (!blocks.length) return;
|
||||
let mermaid;
|
||||
try {
|
||||
mermaid = await loadMermaid();
|
||||
mermaid.initialize({ startOnLoad: false, securityLevel: "strict", theme: "dark" });
|
||||
} catch (error) {
|
||||
blocks.forEach((block) => block.closest("pre").replaceWith(errorBanner(error.message)));
|
||||
return;
|
||||
}
|
||||
for (const [index, block] of blocks.entries()) {
|
||||
const figure = el("figure", { class: "diagram", "data-testid": "diagram" });
|
||||
block.closest("pre").replaceWith(figure);
|
||||
try {
|
||||
const { svg } = await mermaid.render(`diagram-${index}-${Date.now()}`, block.textContent);
|
||||
figure.innerHTML = svg;
|
||||
} catch (error) {
|
||||
figure.replaceWith(errorBanner(`Diagram could not be drawn: ${error.message}`));
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
function notFound() {
|
||||
return el(
|
||||
"div",
|
||||
{ class: "alert", role: "alert", "data-testid": "doc-not-found" },
|
||||
"That documentation page does not exist. ",
|
||||
el("a", { class: "link", href: docHash(INDEX_PAGE) }, "Back to the documentation index")
|
||||
);
|
||||
}
|
||||
|
||||
function sidebar(pages, current) {
|
||||
return el(
|
||||
"nav",
|
||||
{ class: "card", "aria-label": "Documentation" },
|
||||
el("h2", {}, "Documentation"),
|
||||
el(
|
||||
"ul",
|
||||
{ class: "album-list", "data-testid": "doc-pages" },
|
||||
el(
|
||||
"li",
|
||||
{},
|
||||
el(
|
||||
"a",
|
||||
{
|
||||
href: docHash(INDEX_PAGE),
|
||||
"aria-current": current === INDEX_PAGE ? "true" : false,
|
||||
},
|
||||
"Index"
|
||||
)
|
||||
),
|
||||
...pages.map(({ page, title }) =>
|
||||
el(
|
||||
"li",
|
||||
{},
|
||||
el(
|
||||
"a",
|
||||
{
|
||||
href: docHash(page),
|
||||
"data-page": page,
|
||||
"aria-current": page === current ? "true" : false,
|
||||
},
|
||||
title
|
||||
)
|
||||
)
|
||||
)
|
||||
)
|
||||
);
|
||||
}
|
||||
|
||||
export async function renderDocs(root, params = {}) {
|
||||
setActiveNav("docs");
|
||||
const requested = params.page || INDEX_PAGE;
|
||||
const page = PAGE_PATTERN.test(requested) ? requested : "";
|
||||
|
||||
let indexArticle;
|
||||
try {
|
||||
indexArticle = await toArticle(await fetchPage(INDEX_PAGE), INDEX_PAGE);
|
||||
} catch (error) {
|
||||
root.replaceChildren(errorBanner(`Documentation is unavailable: ${error.message}`));
|
||||
return;
|
||||
}
|
||||
const pages = documentIndex(indexArticle);
|
||||
|
||||
let article;
|
||||
if (!page) article = notFound();
|
||||
else if (page === INDEX_PAGE) article = indexArticle;
|
||||
else {
|
||||
try {
|
||||
article = await toArticle(await fetchPage(page), page);
|
||||
} catch {
|
||||
article = notFound();
|
||||
}
|
||||
}
|
||||
|
||||
root.replaceChildren(el("div", { class: "two-pane" }, sidebar(pages, page), article));
|
||||
await renderDiagrams(article);
|
||||
scrollToAnchor(params.anchor);
|
||||
}
|
||||
|
||||
// A documentation anchor cannot live in the hash — the hash is the route — so it
|
||||
// travels as a parameter and is applied after the page renders.
|
||||
function scrollToAnchor(anchor) {
|
||||
if (!anchor) return;
|
||||
const target = document.getElementById(anchor);
|
||||
if (target) target.scrollIntoView({ block: "start" });
|
||||
}
|
||||
@@ -7,7 +7,6 @@ import { createStore } from "../store.js";
|
||||
import { parseHash, navigate } from "../router.js";
|
||||
import { api, cancellable } from "../api.js";
|
||||
import { subscribeJob } from "../events.js";
|
||||
import { assignHeadingIds, documentIndex, resolveDocLink, rewriteLinks, slug } from "../docs.js";
|
||||
|
||||
const cases = [];
|
||||
function ok(name, cond) {
|
||||
@@ -143,62 +142,6 @@ async function run() {
|
||||
window.EventSource = realES;
|
||||
}
|
||||
|
||||
// ── documentation view (US09-01) ─────────────────────────────────────────
|
||||
{
|
||||
ok("slug lowercases and dashes a heading", slug("Key Flows") === "key-flows");
|
||||
ok(
|
||||
"slug drops punctuation but keeps words",
|
||||
slug("What it *will not* do:") === "what-it-will-not-do"
|
||||
);
|
||||
ok("slug keeps non-ASCII letters", slug("Größe & Gewicht") === "größe-gewicht");
|
||||
ok("slug never yields an empty anchor", slug("!!!") === "section");
|
||||
|
||||
const doc = document.createElement("div");
|
||||
doc.innerHTML = "<h2>Notes</h2><h3>Notes</h3><h2>Notes</h2>";
|
||||
const ids = [...assignHeadingIds(doc.querySelectorAll("h2, h3"))].map((h) => h.id);
|
||||
ok("repeated headings get distinct anchors", ids.join(",") === "notes,notes-2,notes-3");
|
||||
|
||||
ok(
|
||||
"a sibling document link becomes an in-app route",
|
||||
resolveDocLink("index", "overview.md")?.page === "overview"
|
||||
);
|
||||
const nested = resolveDocLink("guides/install", "../overview.md#running-it");
|
||||
ok("a relative link resolves against the current page", nested?.page === "overview");
|
||||
ok("a link's anchor survives the rewrite", nested?.anchor === "running-it");
|
||||
ok(
|
||||
"a bare anchor stays on the current page",
|
||||
resolveDocLink("overview", "#the-workflow")?.page === "overview"
|
||||
);
|
||||
ok("an external link is left alone", resolveDocLink("index", "https://example.test/x") === null);
|
||||
ok("a non-markdown relative link is left alone", resolveDocLink("index", "images/a.png") === null);
|
||||
// A climb cannot leave the documentation tree: it is clamped at the root, so the
|
||||
// page it names is still fetched from under /docs and simply does not exist.
|
||||
ok(
|
||||
"a link that climbs above the docs root is clamped",
|
||||
resolveDocLink("index", "../../etc/passwd.md")?.page === "etc/passwd"
|
||||
);
|
||||
ok(
|
||||
"a page name that is not a page name is refused",
|
||||
resolveDocLink("index", "..%2f..%2fetc%2fpasswd.md") === null
|
||||
);
|
||||
|
||||
const index = document.createElement("div");
|
||||
index.innerHTML =
|
||||
'<a href="overview.md">Overview</a><a href="overview.md">again</a>' +
|
||||
'<a href="index.md">itself</a><a href="https://example.test">out</a>';
|
||||
const listed = documentIndex(rewriteLinks(index, "index"));
|
||||
ok(
|
||||
"an external link is marked safe to open away from the app",
|
||||
index.querySelector('a[href^="https"]').rel === "noreferrer noopener"
|
||||
);
|
||||
ok(
|
||||
"an in-app link points at the docs route",
|
||||
index.querySelector("a").getAttribute("href") === "#/docs?page=overview"
|
||||
);
|
||||
ok("the index lists each page once, in order", listed.length === 1);
|
||||
ok("the index takes its titles from the link text", listed[0].title === "Overview");
|
||||
}
|
||||
|
||||
await tick();
|
||||
const failed = cases.filter((c) => !c.ok);
|
||||
window.__RESULTS__ = { passed: cases.length - failed.length, failed: failed.length, cases };
|
||||
|
||||
23
frontend/js/vendor/VERSIONS.json
vendored
23
frontend/js/vendor/VERSIONS.json
vendored
@@ -1,23 +0,0 @@
|
||||
{
|
||||
"_comment": "Vendored browser libraries (US09-01). Committed rather than fetched: the application is deployed to a network whose outbound access is not assumed, and `default-src 'self'` forbids a CDN. Checksums are asserted by tests/integration/test_documentation.py, so replacing a file without recording it here fails the suite. Update by downloading the pinned URL and recording the new version and sha256 in the same commit.",
|
||||
"libraries": [
|
||||
{
|
||||
"file": "marked.esm.js",
|
||||
"name": "marked",
|
||||
"version": "18.0.10",
|
||||
"license": "MIT",
|
||||
"url": "https://cdn.jsdelivr.net/npm/marked@18.0.10/lib/marked.esm.js",
|
||||
"sha256": "4cf47dfebb7f614a08fc0a579ab0fe407ff0ed2b717bf953040c85b2f493a4f0",
|
||||
"why": "Markdown to HTML for the documentation view. An ES module with no dependencies, imported lazily by js/docs.js."
|
||||
},
|
||||
{
|
||||
"file": "mermaid.min.js",
|
||||
"name": "mermaid",
|
||||
"version": "11.17.0",
|
||||
"license": "MIT",
|
||||
"url": "https://cdn.jsdelivr.net/npm/mermaid@11.17.0/dist/mermaid.min.js",
|
||||
"sha256": "8d8e0eec56d3a83b4b3c87f42050845546dee93ebe1875d2117c12e6947c0cb3",
|
||||
"why": "Renders ```mermaid blocks so diagrams are diffable source that Gitea also renders. The single-file UMD build rather than the ES module entry, whose ~40 lazy chunks would each need vendoring and pinning."
|
||||
}
|
||||
]
|
||||
}
|
||||
78
frontend/js/vendor/marked.esm.js
vendored
78
frontend/js/vendor/marked.esm.js
vendored
File diff suppressed because one or more lines are too long
3636
frontend/js/vendor/mermaid.min.js
vendored
3636
frontend/js/vendor/mermaid.min.js
vendored
File diff suppressed because one or more lines are too long
@@ -53,7 +53,6 @@ from photo_pipeline.services.thumbnails import ThumbnailService
|
||||
from photo_pipeline.services.upload_batches import UploadBatchService
|
||||
|
||||
FRONTEND_DIR = Path(__file__).resolve().parents[2] / "frontend"
|
||||
DOCS_DIR = Path(__file__).resolve().parents[2] / "docs"
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
@@ -152,10 +151,4 @@ def create_app(config: Config | None = None) -> FastAPI:
|
||||
# Static single-page app (hash-routed). Mounted last so /api/v1 wins.
|
||||
if FRONTEND_DIR.is_dir():
|
||||
app.mount("/app", StaticFiles(directory=FRONTEND_DIR, html=True), name="app")
|
||||
# The manuals, as their own markdown files (US09-01). Unauthenticated for the
|
||||
# same reason the application shell is: whoever cannot get past the access
|
||||
# secret is exactly who needs the troubleshooting page, and no document holds
|
||||
# anything a session would protect.
|
||||
if DOCS_DIR.is_dir():
|
||||
app.mount("/docs", StaticFiles(directory=DOCS_DIR), name="docs")
|
||||
return app
|
||||
|
||||
@@ -57,21 +57,9 @@ LOOPBACK_HOSTS = frozenset({"127.0.0.1", "localhost", "::1", "[::1]"})
|
||||
# backup a careful operator takes first, and its retention (US07-07).
|
||||
MUTATION_EXEMPT_PATHS = frozenset({f"{API_PREFIX}/backups", f"{API_PREFIX}/backups/prune"})
|
||||
|
||||
# Applied to every response. `frame-ancestors 'none'` and CORP keep other pages from
|
||||
# Applied to every response. No inline script/style is used by the frontend, so the
|
||||
# policy can stay strict; `frame-ancestors 'none'` and CORP keep other pages from
|
||||
# embedding the app or its thumbnails.
|
||||
#
|
||||
# `script-src 'self'` is the boundary that matters and it is unchanged: no
|
||||
# `'unsafe-eval'`, no `'unsafe-inline'`, so nothing injected into the DOM can execute.
|
||||
# Whether the vendored diagram renderer needed `'unsafe-eval'` was measured rather
|
||||
# than assumed — mermaid's bundle contains no `eval(` and no `new Function`, and it
|
||||
# renders under this exact policy without a single script-src violation.
|
||||
#
|
||||
# `style-src` does gain `'unsafe-inline'` (US09-01): mermaid styles the SVG it builds
|
||||
# with an injected `<style>` element and `style=` attributes, and a diagram's CSS
|
||||
# cannot be hashed in advance. The concession is bounded by the directives around it
|
||||
# — with script execution still refused and `img-src`, `connect-src`, and `font-src`
|
||||
# all `'self'`, the CSS exfiltration channels stay closed and what is left is
|
||||
# defacement of a page its own operator is already looking at.
|
||||
DEFAULT_HEADERS = {
|
||||
"x-content-type-options": "nosniff",
|
||||
"x-frame-options": "DENY",
|
||||
@@ -79,9 +67,8 @@ DEFAULT_HEADERS = {
|
||||
"cross-origin-resource-policy": "same-origin",
|
||||
"cross-origin-opener-policy": "same-origin",
|
||||
"content-security-policy": (
|
||||
"default-src 'self'; img-src 'self' data:; style-src 'self' 'unsafe-inline'; "
|
||||
"script-src 'self'; connect-src 'self'; font-src 'self'; "
|
||||
"frame-ancestors 'none'; base-uri 'none'; form-action 'none'"
|
||||
"default-src 'self'; img-src 'self' data:; style-src 'self'; script-src 'self'; "
|
||||
"connect-src 'self'; frame-ancestors 'none'; base-uri 'none'; form-action 'none'"
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
@@ -1,119 +0,0 @@
|
||||
"""US09-01: the manuals, in the browser, under the application's real policy.
|
||||
|
||||
Everything here runs against a real ``photo_pipeline serve`` process serving the
|
||||
committed ``docs/`` tree and the vendored renderer — no fixture markdown, no stubbed
|
||||
fetch. What that buys is the assertion the story actually cares about: the pages
|
||||
render, the links between them work, a diagram becomes a diagram, and the console
|
||||
stays empty, including of CSP violations.
|
||||
|
||||
The offline contract — reachability, dead links, pinned checksums — is
|
||||
``tests/integration/test_documentation.py``.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import pytest
|
||||
|
||||
from tests.e2e._pipeline_harness import Server, seed_library
|
||||
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def server(tmp_path_factory):
|
||||
seeded = seed_library(tmp_path_factory.mktemp("docs"), {"a": 1}, {})
|
||||
running = Server(seeded).start()
|
||||
try:
|
||||
yield running
|
||||
finally:
|
||||
running.stop()
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def quiet(page):
|
||||
"""Any console error, page error, or failed request fails the test that caused
|
||||
it. A CSP violation arrives as a console error, which is the point."""
|
||||
problems: list[str] = []
|
||||
page.on("console", lambda m: problems.append(m.text) if m.type == "error" else None)
|
||||
page.on("pageerror", lambda error: problems.append(str(error)))
|
||||
page.on("requestfailed", lambda request: problems.append(f"failed request: {request.url}"))
|
||||
yield problems
|
||||
|
||||
|
||||
def test_the_documentation_opens_from_the_navigation(page, server, quiet):
|
||||
page.goto(f"{server.base}/app/#/workflow")
|
||||
page.locator('nav a[data-nav="docs"]').click()
|
||||
|
||||
page.get_by_test_id("doc").wait_for()
|
||||
assert page.get_by_test_id("doc").get_attribute("data-page") == "index"
|
||||
# The sidebar is the index's own reading order, not a second list to maintain.
|
||||
assert page.get_by_test_id("doc-pages").get_by_role("link", name="Overview").is_visible()
|
||||
assert quiet == []
|
||||
|
||||
|
||||
def test_a_link_between_documents_stays_inside_the_application(page, server, quiet):
|
||||
page.goto(f"{server.base}/app/#/docs")
|
||||
page.get_by_test_id("doc").wait_for()
|
||||
page.get_by_test_id("doc").get_by_role("link", name="Overview").click()
|
||||
|
||||
page.wait_for_selector('[data-testid="doc"][data-page="overview"]')
|
||||
assert "#/docs?page=overview" in page.url, "a relative .md link must not leave the app"
|
||||
# And back again, by the link the document itself carries.
|
||||
page.get_by_test_id("doc").get_by_role("link", name="← Documentation index").click()
|
||||
page.wait_for_selector('[data-testid="doc"][data-page="index"]')
|
||||
assert quiet == []
|
||||
|
||||
|
||||
def test_a_deep_link_to_a_heading_lands_on_that_heading(page, server, quiet):
|
||||
page.goto(f"{server.base}/app/#/docs?page=overview&anchor=the-workflow")
|
||||
heading = page.locator("#the-workflow")
|
||||
heading.wait_for()
|
||||
|
||||
assert heading.inner_text().strip() == "The workflow"
|
||||
assert heading.evaluate("node => node.getBoundingClientRect().top < window.innerHeight")
|
||||
assert quiet == []
|
||||
|
||||
|
||||
def test_a_mermaid_block_becomes_a_diagram_under_the_unchanged_script_policy(page, server, quiet):
|
||||
page.goto(f"{server.base}/app/#/docs?page=overview")
|
||||
diagram = page.get_by_test_id("diagram").first
|
||||
diagram.wait_for()
|
||||
|
||||
# A real drawing, not the source text and not an error node.
|
||||
assert diagram.locator("svg").count() == 1
|
||||
assert diagram.evaluate("node => node.querySelector('svg').getBBox().width") > 100
|
||||
assert page.locator("pre code.language-mermaid").count() == 0
|
||||
assert quiet == [], "rendering a diagram must not violate the policy"
|
||||
|
||||
policy = page.evaluate(
|
||||
"async () => (await fetch('/app/')).headers.get('content-security-policy')"
|
||||
)
|
||||
assert "script-src 'self';" in policy
|
||||
assert "unsafe-eval" not in policy
|
||||
|
||||
|
||||
def test_an_unknown_page_says_so_without_naming_a_path(page, server, quiet):
|
||||
page.goto(f"{server.base}/app/#/docs?page=no-such-manual")
|
||||
message = page.get_by_test_id("doc-not-found")
|
||||
message.wait_for()
|
||||
|
||||
text = message.inner_text()
|
||||
assert "does not exist" in text
|
||||
assert "/" not in text.replace("Back to the documentation index", ""), text
|
||||
message.get_by_role("link").click()
|
||||
page.wait_for_selector('[data-testid="doc"][data-page="index"]')
|
||||
# The missing page is a 404 and the browser says so; nothing else may go wrong,
|
||||
# and in particular the view must not throw on the way to its own message.
|
||||
assert all("404" in problem for problem in quiet), quiet
|
||||
|
||||
|
||||
def test_the_documentation_is_readable_without_a_session(page, server, quiet):
|
||||
"""The troubleshooting page is needed most by whoever is locked out."""
|
||||
page.goto(f"{server.base}/app/#/docs")
|
||||
page.get_by_test_id("doc").wait_for()
|
||||
unauthenticated = page.evaluate(
|
||||
"""async () => {
|
||||
const response = await fetch('/docs/index.md', { credentials: 'omit' });
|
||||
return { status: response.status, length: (await response.text()).length };
|
||||
}"""
|
||||
)
|
||||
assert unauthenticated["status"] == 200
|
||||
assert unauthenticated["length"] > 100
|
||||
@@ -1,179 +0,0 @@
|
||||
"""US09-01: the documentation the application serves, checked without a browser.
|
||||
|
||||
What is checkable offline is everything that makes the manuals *navigable* rather
|
||||
than merely present: every page reachable, every internal link and anchor landing
|
||||
somewhere real, the vendored renderer being the file that was pinned, and the
|
||||
application actually serving `docs/` in both the working copy and the image.
|
||||
|
||||
The rendering itself — marked, mermaid, anchors, link rewriting — is proven in a
|
||||
real browser by ``tests/e2e/test_docs_ui.py``.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
import re
|
||||
from pathlib import Path
|
||||
|
||||
from starlette.testclient import TestClient
|
||||
|
||||
from photo_pipeline.api.app import create_app
|
||||
from photo_pipeline.api.security import DEFAULT_HEADERS
|
||||
from photo_pipeline.config import Config
|
||||
|
||||
REPO = Path(__file__).resolve().parents[2]
|
||||
DOCS = REPO / "docs"
|
||||
VENDOR = REPO / "frontend" / "js" / "vendor"
|
||||
INDEX = DOCS / "index.md"
|
||||
|
||||
MARKDOWN_LINK = re.compile(r"\[[^\]]*\]\(([^)\s]+)\)")
|
||||
HEADING = re.compile(r"^#{1,6}\s+(.+?)\s*$", re.MULTILINE)
|
||||
FENCE = re.compile(r"^```.*?^```", re.MULTILINE | re.DOTALL)
|
||||
|
||||
|
||||
def pages() -> list[Path]:
|
||||
return sorted(DOCS.rglob("*.md"))
|
||||
|
||||
|
||||
def body(path: Path) -> str:
|
||||
"""A document without its fenced code, so an example link in a shell snippet is
|
||||
not mistaken for a link the reader can follow."""
|
||||
return FENCE.sub("", path.read_text())
|
||||
|
||||
|
||||
def slug(text: str) -> str:
|
||||
"""The heading-anchor rule, mirroring ``frontend/js/docs.js``.
|
||||
|
||||
Deliberately duplicated: this check has to run without a browser. The browser
|
||||
test is the authority on the real behaviour, and it asserts the same anchors.
|
||||
"""
|
||||
cleaned = re.sub(r"[^\w\s-]", "", text.lower(), flags=re.UNICODE).strip()
|
||||
return re.sub(r"[\s-]+", "-", cleaned).strip("-") or "section"
|
||||
|
||||
|
||||
def anchors(path: Path) -> set[str]:
|
||||
found: dict[str, int] = {}
|
||||
for heading in HEADING.findall(body(path)):
|
||||
# Markdown emphasis and inline code are not part of the rendered text.
|
||||
plain = re.sub(r"[*`_]", "", heading)
|
||||
base = slug(plain)
|
||||
found[base] = found.get(base, 0) + 1
|
||||
return {name if index == 1 else f"{name}-{index}" for name, count in found.items() for index in range(1, count + 1)}
|
||||
|
||||
|
||||
def internal_links(path: Path) -> list[str]:
|
||||
return [
|
||||
href
|
||||
for href in MARKDOWN_LINK.findall(body(path))
|
||||
if not re.match(r"^[a-z][a-z\d+.-]*:", href, re.IGNORECASE) and not href.startswith("//")
|
||||
]
|
||||
|
||||
|
||||
# ── the documentation tree ───────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_the_index_exists_and_names_every_page():
|
||||
"""A page nobody links to is a page nobody reads."""
|
||||
assert INDEX.is_file()
|
||||
listed = {
|
||||
(INDEX.parent / href.split("#")[0]).resolve()
|
||||
for href in internal_links(INDEX)
|
||||
if href.split("#")[0].endswith(".md")
|
||||
}
|
||||
unreachable = [page.name for page in pages() if page != INDEX and page.resolve() not in listed]
|
||||
assert unreachable == [], f"not linked from the index: {unreachable}"
|
||||
|
||||
|
||||
def test_every_internal_link_and_anchor_resolves():
|
||||
broken: list[str] = []
|
||||
for page in pages():
|
||||
for href in internal_links(page):
|
||||
target, _, anchor = href.partition("#")
|
||||
destination = page if not target else (page.parent / target).resolve()
|
||||
if not destination.is_file():
|
||||
broken.append(f"{page.name} → {href} (no such file)")
|
||||
continue
|
||||
if anchor and anchor not in anchors(destination):
|
||||
broken.append(f"{page.name} → {href} (no such heading)")
|
||||
assert broken == [], broken
|
||||
|
||||
|
||||
def test_every_page_starts_with_one_title():
|
||||
for page in pages():
|
||||
titles = [line for line in page.read_text().splitlines() if line.startswith("# ")]
|
||||
assert len(titles) == 1, f"{page.name} has {len(titles)} top-level titles"
|
||||
|
||||
|
||||
# ── the vendored renderer ────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_the_vendored_libraries_are_the_files_that_were_pinned():
|
||||
"""A vendored dependency without a recorded checksum is a dependency nobody is
|
||||
reviewing. Replacing one has to be a visible change to this manifest."""
|
||||
manifest = json.loads((VENDOR / "VERSIONS.json").read_text())
|
||||
recorded = {entry["file"]: entry for entry in manifest["libraries"]}
|
||||
|
||||
on_disk = {path.name for path in VENDOR.iterdir() if path.suffix in (".js", ".mjs")}
|
||||
assert on_disk == set(recorded), f"unrecorded vendored files: {on_disk ^ set(recorded)}"
|
||||
|
||||
for name, entry in recorded.items():
|
||||
digest = hashlib.sha256((VENDOR / name).read_bytes()).hexdigest()
|
||||
assert digest == entry["sha256"], f"{name} does not match the pinned checksum"
|
||||
assert entry["version"] in entry["url"], f"{name}: the pinned URL and version disagree"
|
||||
assert entry["license"], f"{name}: no license recorded"
|
||||
|
||||
|
||||
def test_the_documentation_view_is_wired_into_the_shell():
|
||||
shell = (REPO / "frontend" / "index.html").read_text()
|
||||
assert 'data-nav="docs"' in shell, "the manuals are unreachable from the navigation"
|
||||
assert '"/docs"' in (REPO / "frontend" / "js" / "app.js").read_text()
|
||||
|
||||
|
||||
# ── the policy the view runs under ───────────────────────────────────────────
|
||||
|
||||
|
||||
def test_the_script_boundary_is_unchanged_and_only_style_was_relaxed():
|
||||
"""The measured cost of rendering diagrams in the browser, held to that cost.
|
||||
|
||||
mermaid needs to style the SVG it builds, which `style-src 'self'` refuses. It
|
||||
does not need to execute generated code, so `script-src` must never acquire the
|
||||
escape hatch that would let injected markup run.
|
||||
"""
|
||||
policy = DEFAULT_HEADERS["content-security-policy"]
|
||||
directives = {
|
||||
part.split(" ")[0]: part.split(" ")[1:]
|
||||
for part in (piece.strip() for piece in policy.split(";"))
|
||||
if part
|
||||
}
|
||||
assert directives["script-src"] == ["'self'"]
|
||||
assert "'unsafe-inline'" in directives["style-src"]
|
||||
for directive in ("default-src", "img-src", "connect-src", "font-src"):
|
||||
assert "'self'" in directives[directive]
|
||||
assert "'unsafe-inline'" not in directives[directive]
|
||||
assert directives["frame-ancestors"] == ["'none'"]
|
||||
|
||||
|
||||
# ── serving it ───────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_the_application_serves_the_markdown_without_a_session(tmp_path):
|
||||
"""Whoever cannot get past the access secret is exactly who needs these pages."""
|
||||
config = Config.from_env({"PHOTO_PIPELINE_DATA_DIR": str(tmp_path / "data")})
|
||||
(tmp_path / "data").mkdir()
|
||||
with TestClient(create_app(config), raise_server_exceptions=False) as client:
|
||||
response = client.get("/docs/index.md", headers={"cookie": ""})
|
||||
assert response.status_code == 200
|
||||
assert "Photo Pipeline documentation" in response.text
|
||||
assert client.get("/docs/overview.md").status_code == 200
|
||||
assert client.get("/docs/nothing-here.md").status_code == 404
|
||||
# The mount is a directory, not a path parameter: nothing above it is reachable.
|
||||
assert client.get("/docs/../pyproject.toml").status_code in (307, 404)
|
||||
assert client.get("/docs/%2e%2e/pyproject.toml").status_code == 404
|
||||
|
||||
|
||||
def test_the_image_ships_the_documentation_it_serves():
|
||||
"""An image without `docs/` serves an empty manual — and the build context is
|
||||
deny-by-default, so a new directory is excluded until it is named."""
|
||||
assert "\n!docs\n" in (REPO / ".dockerignore").read_text()
|
||||
assert "COPY docs ./docs" in (REPO / "Dockerfile").read_text()
|
||||
@@ -193,13 +193,10 @@
|
||||
"US08-05": [
|
||||
"tests/integration/test_container_gate.py",
|
||||
"tests/e2e/test_phase_h_container.py"
|
||||
],
|
||||
"US09-01": [
|
||||
"tests/integration/test_documentation.py",
|
||||
"tests/e2e/test_docs_ui.py"
|
||||
]
|
||||
},
|
||||
"planned": [
|
||||
"US09-01",
|
||||
"US09-02",
|
||||
"US09-03",
|
||||
"US09-04",
|
||||
|
||||
Reference in New Issue
Block a user