--- id: ESRC-WP-0002 type: workplan title: "HTML and Markdown source representations" domain: infotech repo: evidence-source status: proposed owner: codex topic_slug: citation_evidence_mvp created: "2026-07-08" updated: "2026-07-08" spec_refs: - README.md - docs/ADR-0001-extraction-boundary.md - ../citation-engine/src/shared/document.ts state_hub_workstream_id: "c0d28886-e02f-4df2-abb6-ebe179171c94" --- # ESRC-WP-0002 — HTML and Markdown source representations The PDF slice (ESRC-WP-0001) established the ingest boundary and contract shape. This workplan extends ingest to HTML and Markdown sources, producing the same `{ document, representation }` pair over engine-owned contracts. ## Open questions - Which `representationType` values do HTML/MD map to, and do they need new `DocumentRepresentation` fields in `citation-engine`? - How is canonical text + offset map derived for reflowable formats that have no fixed pages? (Page map may be empty or synthetic.) ## Sketch ```task id: ESRC-WP-0002-T01 status: todo priority: medium state_hub_task_id: "bf1d2a2f-41b4-44b5-b1d1-4f0eec5fbf59" ``` Define the HTML/MD representation contract with `citation-engine`; agree `representationType` and any offset/page semantics for pageless formats. ```task id: ESRC-WP-0002-T02 status: todo priority: medium state_hub_task_id: "66320911-a4b2-404f-b025-8b23dc0c9680" ``` Implement `ingestHtml` / `ingestMarkdown` in `src/html/` and `src/markdown/`, reusing `fingerprintBytes` and the canonical-text normalizer. ```task id: ESRC-WP-0002-T03 status: todo priority: medium state_hub_task_id: "266d168c-d93e-4387-9527-70aa0d5b2a6b" ``` Add fixture-driven contract tests mirroring the PDF suite; decide fixture ownership per ADR-0002.