evidence-source/README.md
tegwick cb93c322c0
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 2s
feat(pdf): extract standalone PDF ingest package (ESRC-WP-0001)
Bootstrap evidence-source from the citation-evidence src/source PDF slice.

- Headless core (src/pdf): ingest/extract/fingerprint, importing domain
  contracts from @citation-evidence/engine/shared (no local copies)
- Browser upload helpers isolated under src/browser behind a ./browser
  entry point, with an eslint boundary keeping the core browser-free
- pnpm/TS/vitest/eslint scaffold; 52 tests (contract + determinism)
- Fixtures resolved from the sibling citation-evidence checkout, not
  duplicated (real PII) — see docs/ADR-0002; suites skip when absent
- Boundary + fixture decisions recorded as docs/ADR-0001 / ADR-0002
- README/SCOPE rewritten; capability.infotech.pdf-evidence-ingest
  registered, NO_CAPABILITIES removed
- Follow-on workplans ESRC-WP-0002..0004 queued; ESRC-WP-0001 finished

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 20:55:28 +02:00

2.5 KiB

evidence-source

Headless document ingest for the citation-evidence ecosystem. Turns raw PDF bytes into engine-owned evidence contracts — a Document (media type, SHA-256 fingerprint, optional title/uri/metadata) and a DocumentRepresentation (pdf-text: canonical text, page map, gap-free offset map). Ingest is pure over bytes: no persistence, no viewer state, no React.

Status

Implemented: the PDF slice. As of ESRC-WP-0001 this repo hosts the extracted PDF ingest core that previously lived in citation-evidence/src/source/. HTML/Markdown representations, richer metadata enrichment, and citation recovery are deferred to follow-on workplans (see workplans/).

Install

Requires a sibling checkout of citation-engine (domain contracts) and, for tests, citation-evidence (fixture corpus — see docs/ADR-0002-fixture-ownership.md).

pnpm install   # resolves @citation-evidence/engine via link:../citation-engine

Usage

import { ingestPdf } from "@citation-evidence/evidence-source";

const { document, representation } = await ingestPdf(pdfBytes, {
  filename: "contract.pdf",
});
// document.fingerprint      -> SHA-256 hex
// representation.canonicalText / pageMap / offsetMap

Browser upload helpers (in-memory blob: byte store + ingestPdfFromFile) live behind a separate, clearly-isolated entry point so headless consumers do not pull in browser machinery:

import { createPdfByteStore, ingestPdfFromFile }
  from "@citation-evidence/evidence-source/browser";

Extraction requires the host to configure the PDF.js worker (GlobalWorkerOptions.workerSrc) before calling extractPdf/ingestPdf; the module itself does no worker setup so it loads cleanly in Node and browsers.

Dev commands

pnpm test        # vitest (fixture suites skip if the corpus is absent)
pnpm typecheck   # tsc --noEmit
pnpm lint        # eslint

Point EVIDENCE_SOURCE_FIXTURE_DIR at the PDF corpus when the sibling citation-evidence checkout is not at ../citation-evidence.

Architecture

  • src/pdf/ — headless core: ingest, extract, fingerprint. Runtime-agnostic.
  • src/browser/ — browser upload surface: byte-store, upload. Separated by an ESLint boundary so the core never imports it.
  • Domain contracts come from @citation-evidence/engine/shared; none are copied locally.
  • Boundary and fixture decisions: docs/ADR-0001, docs/ADR-0002.

viewer-url resolution stays in the consuming app (citation-evidence), which imports this package's ingest core through a thin façade.