feat(pdf): extract standalone PDF ingest package (ESRC-WP-0001)
Bootstrap evidence-source from the citation-evidence src/source PDF slice. - Headless core (src/pdf): ingest/extract/fingerprint, importing domain contracts from @citation-evidence/engine/shared (no local copies) - Browser upload helpers isolated under src/browser behind a ./browser entry point, with an eslint boundary keeping the core browser-free - pnpm/TS/vitest/eslint scaffold; 52 tests (contract + determinism) - Fixtures resolved from the sibling citation-evidence checkout, not duplicated (real PII) — see docs/ADR-0002; suites skip when absent - Boundary + fixture decisions recorded as docs/ADR-0001 / ADR-0002 - README/SCOPE rewritten; capability.infotech.pdf-evidence-ingest registered, NO_CAPABILITIES removed - Follow-on workplans ESRC-WP-0002..0004 queued; ESRC-WP-0001 finished Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
parent
2fd715ba45
commit
cb93c322c0
31 changed files with 5066 additions and 105 deletions
179
registry/capabilities/capability.infotech.pdf-evidence-ingest.md
Normal file
179
registry/capabilities/capability.infotech.pdf-evidence-ingest.md
Normal file
|
|
@ -0,0 +1,179 @@
|
|||
---
|
||||
id: capability.infotech.pdf-evidence-ingest
|
||||
name: Headless PDF Evidence Ingest
|
||||
summary: Turns raw PDF bytes into an engine-shaped Document + DocumentRepresentation — SHA-256 fingerprint, canonical text, page map, and gap-free offset map — as a runtime-agnostic library with no viewer, persistence, or React coupling.
|
||||
owner: evidence-source
|
||||
status: draft
|
||||
domain: infotech
|
||||
tags:
|
||||
- pdf
|
||||
- ingest
|
||||
- fingerprint
|
||||
- canonical-text
|
||||
- evidence
|
||||
|
||||
maturity:
|
||||
discovery:
|
||||
current: D2
|
||||
target: D5
|
||||
confidence: medium
|
||||
rationale: >
|
||||
Intent, scope, and boundary are documented (ADR-0001/0002) and the
|
||||
surface is exercised by a fixture-driven contract suite plus a
|
||||
cross-repo consumer round-trip. HTML/MD ingest and citation recovery
|
||||
are declared out of scope for now, so discovery is bounded to the PDF
|
||||
slice.
|
||||
availability:
|
||||
current: A1
|
||||
target: A3
|
||||
confidence: medium
|
||||
rationale: >
|
||||
Consumable today as a linked library (`link:../evidence-source`) via the
|
||||
package root (`@citation-evidence/evidence-source`) and browser subpath.
|
||||
Reaching A3 needs a published/versioned artifact rather than a
|
||||
sibling-checkout link.
|
||||
|
||||
external_evidence:
|
||||
completeness:
|
||||
level: C1
|
||||
confidence: medium
|
||||
basis: scope_vs_intent_and_consumer_expectations
|
||||
satisfied_expectations:
|
||||
- PDF bytes -> Document + DocumentRepresentation
|
||||
- deterministic SHA-256 fingerprint
|
||||
- canonical text with gap-free page offset map
|
||||
broken_expectations: []
|
||||
out_of_scope_expectations:
|
||||
- HTML / Markdown ingest
|
||||
- metadata enrichment beyond pass-through options
|
||||
- citation recovery / external source discovery
|
||||
reliability:
|
||||
level: R1
|
||||
confidence: medium
|
||||
basis: consumer_quality_signals
|
||||
known_reliability_risks:
|
||||
- PDF.js worker must be configured by the host before extraction
|
||||
- text-extraction fidelity varies on scanned/OCR-noisy PDFs
|
||||
|
||||
discovery:
|
||||
intent: >
|
||||
Let any consumer turn a PDF into engine-owned evidence contracts without
|
||||
inheriting a viewer, a store, or a UI framework.
|
||||
includes:
|
||||
- PDF byte ingest (ingestPdf)
|
||||
- PDF text extraction to canonical text + page/offset maps (extractPdf)
|
||||
- SHA-256 byte fingerprint (fingerprintBytes)
|
||||
- browser upload helpers behind a separate entry point (createPdfByteStore, ingestPdfFromFile)
|
||||
excludes:
|
||||
- HTML / Markdown ingest
|
||||
- persistence of documents/representations
|
||||
- viewer URL resolution (stays in the consuming app)
|
||||
- citation recovery
|
||||
assumptions:
|
||||
- domain contracts come from '@citation-evidence/engine/shared'
|
||||
- host configures the PDF.js worker before extraction
|
||||
use_cases:
|
||||
- citation-evidence umbrella consumes ingest across the repo boundary
|
||||
research_memos: []
|
||||
|
||||
availability:
|
||||
current_level: A1
|
||||
target_level: A3
|
||||
current_artifacts:
|
||||
- '@citation-evidence/evidence-source (package root, headless core)'
|
||||
- '@citation-evidence/evidence-source/browser (upload helpers)'
|
||||
target_artifacts:
|
||||
- published versioned package
|
||||
consumption_modes:
|
||||
- library import
|
||||
|
||||
relations:
|
||||
depends_on: []
|
||||
supports: []
|
||||
related_to: []
|
||||
|
||||
evidence:
|
||||
documentation:
|
||||
- README.md
|
||||
- docs/ADR-0001-extraction-boundary.md
|
||||
- docs/ADR-0002-fixture-ownership.md
|
||||
tests:
|
||||
- src/pdf/ingest.test.ts
|
||||
- src/pdf/fingerprint.test.ts
|
||||
- src/browser/byte-store.test.ts
|
||||
- src/browser/upload.test.ts
|
||||
consumer_feedback: []
|
||||
bug_reports: []
|
||||
incidents: []
|
||||
|
||||
consumer_guidance:
|
||||
recommended_for:
|
||||
- headless PDF -> evidence-contract ingest in Node or the browser
|
||||
- deterministic content fingerprinting of document bytes
|
||||
not_recommended_for:
|
||||
- non-PDF formats (not yet implemented)
|
||||
- high-fidelity extraction from scanned/OCR-only PDFs
|
||||
known_limitations:
|
||||
- PDF only for now
|
||||
- requires host-configured PDF.js worker
|
||||
- fixtures resolved from sibling citation-evidence checkout (ADR-0002)
|
||||
|
||||
promotion_history:
|
||||
- date: "2026-07-08"
|
||||
dimension: discovery
|
||||
from: D0
|
||||
to: D2
|
||||
rationale: PDF ingest slice extracted, documented (ADR-0001/0002), and covered by contract tests plus a cross-repo consumer round-trip (ESRC-WP-0001).
|
||||
author: claude-code
|
||||
---
|
||||
|
||||
# Headless PDF Evidence Ingest
|
||||
|
||||
## Overview
|
||||
|
||||
`evidence-source` turns raw PDF bytes into the two engine-owned contracts that
|
||||
the rest of the citation-evidence ecosystem builds on — a `Document` (media
|
||||
type, SHA-256 fingerprint, optional title/uri/metadata) and a
|
||||
`DocumentRepresentation` (`pdf-text`: canonical text, page map, and a gap-free
|
||||
offset map). Ingest is pure over bytes: no persistence, no viewer state, no
|
||||
React.
|
||||
|
||||
The YAML front matter above is the machine-readable source of truth for
|
||||
maturity, evidence, relations, and consumer guidance.
|
||||
|
||||
## Assessment notes
|
||||
|
||||
### Discovery
|
||||
|
||||
The boundary is fixed in ADR-0001 (which files are headless core vs. browser
|
||||
helpers vs. app-only viewer concerns) and ADR-0002 (fixture ownership). The
|
||||
public surface is small and stable: `ingestPdf`, `extractPdf`,
|
||||
`fingerprintBytes`, plus browser upload helpers behind a separate entry point.
|
||||
|
||||
### Availability
|
||||
|
||||
Consumed today as a linked library via `link:../evidence-source`. The umbrella
|
||||
app (`citation-evidence`) imports the headless core and browser helpers through
|
||||
a thin façade and runs its source ↔ anchor round-trip across the boundary.
|
||||
Target A3 requires a published, versioned artifact.
|
||||
|
||||
### Completeness
|
||||
|
||||
PDF only. HTML/Markdown representations, metadata enrichment, and citation
|
||||
recovery are explicitly deferred to follow-on workplans (see README) and are
|
||||
recorded as out-of-scope expectations rather than gaps.
|
||||
|
||||
### Reliability
|
||||
|
||||
Backed by a fixture-driven contract suite (fingerprint determinism, page/offset
|
||||
invariants, known-good quote presence) and a downstream consumer round-trip.
|
||||
The main reliability caveat is extraction fidelity on scanned/OCR-noisy PDFs,
|
||||
which is inherent to text-layer extraction.
|
||||
|
||||
## Promotion checklist
|
||||
|
||||
- [x] ID follows `capability.<domain>.<name>` pattern
|
||||
- [x] Maturity enums match `specs/CapabilityMaturityStandard.md`
|
||||
- [x] `external_evidence` is populated separately from `maturity`
|
||||
- [x] Relations reference valid capability IDs
|
||||
- [x] Index entry added or updated in `registry/indexes/capabilities.yaml`
|
||||
Loading…
Add table
Add a link
Reference in a new issue