evidence-source/README.md
tegwick 97437dfa18
Some checks failed
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Failing after 15m18s
Implement ESRC-WP-0002/0003/0004: HTML/MD ingest, metadata, recovery
Add ingestHtml and ingestMarkdown with ADR-0003 pageless offset semantics,
PDF intrinsic metadata extraction with caller-wins merge (WP-0003), and
citation recovery primitives including re-ingest reconcile, local quote
search, and pluggable discovery hooks (ADR-0004). Mark all three workplans
finished with contract tests (80 passing).
2026-07-09 01:48:28 +02:00

2.8 KiB

evidence-source

Headless document ingest for the citation-evidence ecosystem. Turns raw PDF, HTML, and Markdown bytes into engine-owned evidence contracts — a Document (media type, SHA-256 fingerprint, optional title/uri/metadata) and a DocumentRepresentation (canonical text, page/offset maps where applicable). Ingest is pure over bytes: no persistence, no viewer state, no React.

Status

Implemented: PDF ingest (ESRC-WP-0001), HTML/Markdown ingest (ESRC-WP-0002), PDF metadata enrichment (ESRC-WP-0003), and citation recovery primitives (ESRC-WP-0004). See workplans/ for history.

Install

Requires a sibling checkout of citation-engine (domain contracts) and, for tests, citation-evidence (fixture corpus — see docs/ADR-0002-fixture-ownership.md).

pnpm install   # resolves @citation-evidence/engine via link:../citation-engine

Usage

import {
  ingestPdf,
  ingestHtml,
  ingestMarkdown,
} from "@citation-evidence/evidence-source";

const pdf = await ingestPdf(pdfBytes, { filename: "contract.pdf" });
const html = await ingestHtml(htmlBytes, { filename: "brief.html" });
const md = await ingestMarkdown("# Title\n\nBody text.", { filename: "notes.md" });
// document.fingerprint -> SHA-256 hex
// representation.canonicalText / pageMap / offsetMap (format-dependent)

Browser upload helpers (in-memory blob: byte store + ingestPdfFromFile) live behind a separate, clearly-isolated entry point so headless consumers do not pull in browser machinery:

import { createPdfByteStore, ingestPdfFromFile }
  from "@citation-evidence/evidence-source/browser";

Extraction requires the host to configure the PDF.js worker (GlobalWorkerOptions.workerSrc) before calling extractPdf/ingestPdf; the module itself does no worker setup so it loads cleanly in Node and browsers.

Dev commands

pnpm test        # vitest (fixture suites skip if the corpus is absent)
pnpm typecheck   # tsc --noEmit
pnpm lint        # eslint

Point EVIDENCE_SOURCE_FIXTURE_DIR at the PDF corpus when the sibling citation-evidence checkout is not at ../citation-evidence.

Architecture

  • src/pdf/ — PDF ingest, extraction, fingerprinting, metadata enrichment.
  • src/html/, src/markdown/ — reflowable-format ingest (ADR-0003).
  • src/recovery/ — re-ingest/reconcile, local quote search, discovery hooks (ADR-0004).
  • src/browser/ — browser upload surface: byte-store, upload. Separated by an ESLint boundary so the core never imports it.
  • Domain contracts come from @citation-evidence/engine/shared; none are copied locally.
  • Boundary and fixture decisions: docs/ADR-0001 through docs/ADR-0004.

viewer-url resolution stays in the consuming app (citation-evidence), which imports this package's ingest core through a thin façade.