49 lines
1.2 KiB
Markdown
49 lines
1.2 KiB
Markdown
|
|
---
|
||
|
|
id: ESRC-WP-0003
|
||
|
|
type: workplan
|
||
|
|
title: "Document metadata enrichment beyond pass-through options"
|
||
|
|
domain: infotech
|
||
|
|
repo: evidence-source
|
||
|
|
status: proposed
|
||
|
|
owner: codex
|
||
|
|
topic_slug: citation_evidence_mvp
|
||
|
|
created: "2026-07-08"
|
||
|
|
updated: "2026-07-08"
|
||
|
|
spec_refs:
|
||
|
|
- README.md
|
||
|
|
- src/pdf/ingest.ts
|
||
|
|
---
|
||
|
|
|
||
|
|
# ESRC-WP-0003 — Metadata enrichment
|
||
|
|
|
||
|
|
Today `ingestPdf` only passes through caller-supplied `title`/`uri`/`metadata`.
|
||
|
|
This workplan extracts intrinsic document metadata (PDF info dictionary,
|
||
|
|
embedded XMP, page-derived signals like author/creation date/title) and folds
|
||
|
|
it into the `Document` record without breaking the pure-over-bytes contract.
|
||
|
|
|
||
|
|
## Sketch
|
||
|
|
|
||
|
|
```task
|
||
|
|
id: ESRC-WP-0003-T01
|
||
|
|
status: todo
|
||
|
|
priority: low
|
||
|
|
```
|
||
|
|
Enumerate available PDF metadata sources (info dict, XMP) via PDF.js and decide
|
||
|
|
precedence vs. caller-supplied options (caller wins).
|
||
|
|
|
||
|
|
```task
|
||
|
|
id: ESRC-WP-0003-T02
|
||
|
|
status: todo
|
||
|
|
priority: low
|
||
|
|
```
|
||
|
|
Implement extraction into a normalized metadata shape; keep it optional so
|
||
|
|
minimal-metadata PDFs still ingest cleanly.
|
||
|
|
|
||
|
|
```task
|
||
|
|
id: ESRC-WP-0003-T03
|
||
|
|
status: todo
|
||
|
|
priority: low
|
||
|
|
```
|
||
|
|
Contract tests over the fixture corpus asserting extracted vs. overridden
|
||
|
|
metadata precedence.
|