|
|
||
|---|---|---|
| .claude/rules | ||
| .forgejo/workflows | ||
| docs | ||
| registry | ||
| src | ||
| tests | ||
| workplans | ||
| .custodian-brief.md | ||
| .gitignore | ||
| .repo-classification.yaml | ||
| AGENTS.md | ||
| CLAUDE.md | ||
| eslint.config.js | ||
| INTENT.md | ||
| LICENSE | ||
| package.json | ||
| pnpm-lock.yaml | ||
| README.md | ||
| SCOPE.md | ||
| tsconfig.json | ||
| vitest.config.ts | ||
evidence-source
Headless document ingest for the citation-evidence ecosystem. Turns raw PDF,
HTML, and Markdown bytes into engine-owned evidence contracts — a Document
(media type, SHA-256 fingerprint, optional title/uri/metadata) and a
DocumentRepresentation (canonical text, page/offset maps where applicable).
Ingest is pure over bytes: no persistence, no viewer state, no React.
Status
Implemented: PDF ingest (ESRC-WP-0001), HTML/Markdown ingest
(ESRC-WP-0002), PDF metadata enrichment (ESRC-WP-0003), and citation recovery
primitives (ESRC-WP-0004). See workplans/ for history.
Install
Requires a sibling checkout of citation-engine (domain contracts) and, for
tests, citation-evidence (fixture corpus — see docs/ADR-0002-fixture-ownership.md).
pnpm install # resolves @citation-evidence/engine via link:../citation-engine
Usage
import {
ingestPdf,
ingestHtml,
ingestMarkdown,
} from "@citation-evidence/evidence-source";
const pdf = await ingestPdf(pdfBytes, { filename: "contract.pdf" });
const html = await ingestHtml(htmlBytes, { filename: "brief.html" });
const md = await ingestMarkdown("# Title\n\nBody text.", { filename: "notes.md" });
// document.fingerprint -> SHA-256 hex
// representation.canonicalText / pageMap / offsetMap (format-dependent)
Browser upload helpers (in-memory blob: byte store + ingestPdfFromFile)
live behind a separate, clearly-isolated entry point so headless consumers do
not pull in browser machinery:
import { createPdfByteStore, ingestPdfFromFile }
from "@citation-evidence/evidence-source/browser";
Extraction requires the host to configure the PDF.js worker
(GlobalWorkerOptions.workerSrc) before calling extractPdf/ingestPdf; the
module itself does no worker setup so it loads cleanly in Node and browsers.
Dev commands
pnpm test # vitest (fixture suites skip if the corpus is absent)
pnpm typecheck # tsc --noEmit
pnpm lint # eslint
Point EVIDENCE_SOURCE_FIXTURE_DIR at the PDF corpus when the sibling
citation-evidence checkout is not at ../citation-evidence.
Architecture
src/pdf/— PDF ingest, extraction, fingerprinting, metadata enrichment.src/html/,src/markdown/— reflowable-format ingest (ADR-0003).src/recovery/— re-ingest/reconcile, local quote search, discovery hooks (ADR-0004).src/browser/— browser upload surface:byte-store,upload. Separated by an ESLint boundary so the core never imports it.- Domain contracts come from
@citation-evidence/engine/shared; none are copied locally. - Boundary and fixture decisions:
docs/ADR-0001throughdocs/ADR-0004.
viewer-url resolution stays in the consuming app (citation-evidence), which
imports this package's ingest core through a thin façade.