Bootstrap evidence-source from the citation-evidence src/source PDF slice. - Headless core (src/pdf): ingest/extract/fingerprint, importing domain contracts from @citation-evidence/engine/shared (no local copies) - Browser upload helpers isolated under src/browser behind a ./browser entry point, with an eslint boundary keeping the core browser-free - pnpm/TS/vitest/eslint scaffold; 52 tests (contract + determinism) - Fixtures resolved from the sibling citation-evidence checkout, not duplicated (real PII) — see docs/ADR-0002; suites skip when absent - Boundary + fixture decisions recorded as docs/ADR-0001 / ADR-0002 - README/SCOPE rewritten; capability.infotech.pdf-evidence-ingest registered, NO_CAPABILITIES removed - Follow-on workplans ESRC-WP-0002..0004 queued; ESRC-WP-0001 finished Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2.5 KiB
evidence-source
Headless document ingest for the citation-evidence ecosystem. Turns raw PDF
bytes into engine-owned evidence contracts — a Document (media type,
SHA-256 fingerprint, optional title/uri/metadata) and a
DocumentRepresentation (pdf-text: canonical text, page map, gap-free
offset map). Ingest is pure over bytes: no persistence, no viewer state, no
React.
Status
Implemented: the PDF slice. As of ESRC-WP-0001 this repo hosts the
extracted PDF ingest core that previously lived in citation-evidence/src/source/.
HTML/Markdown representations, richer metadata enrichment, and citation
recovery are deferred to follow-on workplans (see workplans/).
Install
Requires a sibling checkout of citation-engine (domain contracts) and, for
tests, citation-evidence (fixture corpus — see docs/ADR-0002-fixture-ownership.md).
pnpm install # resolves @citation-evidence/engine via link:../citation-engine
Usage
import { ingestPdf } from "@citation-evidence/evidence-source";
const { document, representation } = await ingestPdf(pdfBytes, {
filename: "contract.pdf",
});
// document.fingerprint -> SHA-256 hex
// representation.canonicalText / pageMap / offsetMap
Browser upload helpers (in-memory blob: byte store + ingestPdfFromFile)
live behind a separate, clearly-isolated entry point so headless consumers do
not pull in browser machinery:
import { createPdfByteStore, ingestPdfFromFile }
from "@citation-evidence/evidence-source/browser";
Extraction requires the host to configure the PDF.js worker
(GlobalWorkerOptions.workerSrc) before calling extractPdf/ingestPdf; the
module itself does no worker setup so it loads cleanly in Node and browsers.
Dev commands
pnpm test # vitest (fixture suites skip if the corpus is absent)
pnpm typecheck # tsc --noEmit
pnpm lint # eslint
Point EVIDENCE_SOURCE_FIXTURE_DIR at the PDF corpus when the sibling
citation-evidence checkout is not at ../citation-evidence.
Architecture
src/pdf/— headless core:ingest,extract,fingerprint. Runtime-agnostic.src/browser/— browser upload surface:byte-store,upload. Separated by an ESLint boundary so the core never imports it.- Domain contracts come from
@citation-evidence/engine/shared; none are copied locally. - Boundary and fixture decisions:
docs/ADR-0001,docs/ADR-0002.
viewer-url resolution stays in the consuming app (citation-evidence), which
imports this package's ingest core through a thin façade.