Bootstrap evidence-source from the citation-evidence src/source PDF slice. - Headless core (src/pdf): ingest/extract/fingerprint, importing domain contracts from @citation-evidence/engine/shared (no local copies) - Browser upload helpers isolated under src/browser behind a ./browser entry point, with an eslint boundary keeping the core browser-free - pnpm/TS/vitest/eslint scaffold; 52 tests (contract + determinism) - Fixtures resolved from the sibling citation-evidence checkout, not duplicated (real PII) — see docs/ADR-0002; suites skip when absent - Boundary + fixture decisions recorded as docs/ADR-0001 / ADR-0002 - README/SCOPE rewritten; capability.infotech.pdf-evidence-ingest registered, NO_CAPABILITIES removed - Follow-on workplans ESRC-WP-0002..0004 queued; ESRC-WP-0001 finished Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
1.2 KiB
1.2 KiB
| id | type | title | domain | repo | status | owner | topic_slug | created | updated | spec_refs | ||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ESRC-WP-0003 | workplan | Document metadata enrichment beyond pass-through options | infotech | evidence-source | proposed | codex | citation_evidence_mvp | 2026-07-08 | 2026-07-08 |
|
ESRC-WP-0003 — Metadata enrichment
Today ingestPdf only passes through caller-supplied title/uri/metadata.
This workplan extracts intrinsic document metadata (PDF info dictionary,
embedded XMP, page-derived signals like author/creation date/title) and folds
it into the Document record without breaking the pure-over-bytes contract.
Sketch
id: ESRC-WP-0003-T01
status: todo
priority: low
Enumerate available PDF metadata sources (info dict, XMP) via PDF.js and decide precedence vs. caller-supplied options (caller wins).
id: ESRC-WP-0003-T02
status: todo
priority: low
Implement extraction into a normalized metadata shape; keep it optional so minimal-metadata PDFs still ingest cleanly.
id: ESRC-WP-0003-T03
status: todo
priority: low
Contract tests over the fixture corpus asserting extracted vs. overridden metadata precedence.