Document source, ingestion, extraction, metadata, and citation recovery.
Find a file
tegwick df6acad828
All checks were successful
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Successful in 3s
chore(consistency): write back state-hub workplan/task IDs [auto]
fix-consistency C-06/C-11 registered ESRC-WP-0002..0004 and the
ESRC-WP-0001 tasks; IDs written back into the workplan files.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 20:56:39 +02:00
.claude/rules feat(pdf): extract standalone PDF ingest package (ESRC-WP-0001) 2026-07-08 20:55:28 +02:00
.forgejo/workflows Add Forgejo CI smoke workflow (enablement template) 2026-07-08 12:32:09 +02:00
docs feat(pdf): extract standalone PDF ingest package (ESRC-WP-0001) 2026-07-08 20:55:28 +02:00
registry feat(pdf): extract standalone PDF ingest package (ESRC-WP-0001) 2026-07-08 20:55:28 +02:00
src feat(pdf): extract standalone PDF ingest package (ESRC-WP-0001) 2026-07-08 20:55:28 +02:00
tests feat(pdf): extract standalone PDF ingest package (ESRC-WP-0001) 2026-07-08 20:55:28 +02:00
workplans chore(consistency): write back state-hub workplan/task IDs [auto] 2026-07-08 20:56:39 +02:00
.custodian-brief.md chore(consistency): sync task status from DB [auto] 2026-07-08 20:56:16 +02:00
.gitignore feat(pdf): extract standalone PDF ingest package (ESRC-WP-0001) 2026-07-08 20:55:28 +02:00
.repo-classification.yaml Add .repo-classification.yaml (CUST-WP-0050 T11 agent first-pass) 2026-06-22 17:47:36 +02:00
AGENTS.md Regenerate agent instructions from state-hub templates (CUST-WP-0055 T01) 2026-07-08 14:50:21 +02:00
CLAUDE.md Normalize agent instructions and workplan frontmatter (STATE-WP-0067) 2026-06-22 23:16:24 +02:00
eslint.config.js feat(pdf): extract standalone PDF ingest package (ESRC-WP-0001) 2026-07-08 20:55:28 +02:00
INTENT.md Add MVP Coordination section: code lives in citation-evidence umbrella during MVP 2026-05-24 16:51:06 +02:00
LICENSE Initial commit 2026-05-24 13:49:33 +00:00
package.json feat(pdf): extract standalone PDF ingest package (ESRC-WP-0001) 2026-07-08 20:55:28 +02:00
pnpm-lock.yaml feat(pdf): extract standalone PDF ingest package (ESRC-WP-0001) 2026-07-08 20:55:28 +02:00
README.md feat(pdf): extract standalone PDF ingest package (ESRC-WP-0001) 2026-07-08 20:55:28 +02:00
SCOPE.md feat(pdf): extract standalone PDF ingest package (ESRC-WP-0001) 2026-07-08 20:55:28 +02:00
tsconfig.json feat(pdf): extract standalone PDF ingest package (ESRC-WP-0001) 2026-07-08 20:55:28 +02:00
vitest.config.ts feat(pdf): extract standalone PDF ingest package (ESRC-WP-0001) 2026-07-08 20:55:28 +02:00

evidence-source

Headless document ingest for the citation-evidence ecosystem. Turns raw PDF bytes into engine-owned evidence contracts — a Document (media type, SHA-256 fingerprint, optional title/uri/metadata) and a DocumentRepresentation (pdf-text: canonical text, page map, gap-free offset map). Ingest is pure over bytes: no persistence, no viewer state, no React.

Status

Implemented: the PDF slice. As of ESRC-WP-0001 this repo hosts the extracted PDF ingest core that previously lived in citation-evidence/src/source/. HTML/Markdown representations, richer metadata enrichment, and citation recovery are deferred to follow-on workplans (see workplans/).

Install

Requires a sibling checkout of citation-engine (domain contracts) and, for tests, citation-evidence (fixture corpus — see docs/ADR-0002-fixture-ownership.md).

pnpm install   # resolves @citation-evidence/engine via link:../citation-engine

Usage

import { ingestPdf } from "@citation-evidence/evidence-source";

const { document, representation } = await ingestPdf(pdfBytes, {
  filename: "contract.pdf",
});
// document.fingerprint      -> SHA-256 hex
// representation.canonicalText / pageMap / offsetMap

Browser upload helpers (in-memory blob: byte store + ingestPdfFromFile) live behind a separate, clearly-isolated entry point so headless consumers do not pull in browser machinery:

import { createPdfByteStore, ingestPdfFromFile }
  from "@citation-evidence/evidence-source/browser";

Extraction requires the host to configure the PDF.js worker (GlobalWorkerOptions.workerSrc) before calling extractPdf/ingestPdf; the module itself does no worker setup so it loads cleanly in Node and browsers.

Dev commands

pnpm test        # vitest (fixture suites skip if the corpus is absent)
pnpm typecheck   # tsc --noEmit
pnpm lint        # eslint

Point EVIDENCE_SOURCE_FIXTURE_DIR at the PDF corpus when the sibling citation-evidence checkout is not at ../citation-evidence.

Architecture

  • src/pdf/ — headless core: ingest, extract, fingerprint. Runtime-agnostic.
  • src/browser/ — browser upload surface: byte-store, upload. Separated by an ESLint boundary so the core never imports it.
  • Domain contracts come from @citation-evidence/engine/shared; none are copied locally.
  • Boundary and fixture decisions: docs/ADR-0001, docs/ADR-0002.

viewer-url resolution stays in the consuming app (citation-evidence), which imports this package's ingest core through a thin façade.