Implement ESRC-WP-0002/0003/0004: HTML/MD ingest, metadata, recovery
Some checks failed
CI Smoke / host-smoke (push) Successful in 0s
CI Smoke / container-smoke (push) Failing after 15m18s

Add ingestHtml and ingestMarkdown with ADR-0003 pageless offset semantics,
PDF intrinsic metadata extraction with caller-wins merge (WP-0003), and
citation recovery primitives including re-ingest reconcile, local quote
search, and pluggable discovery hooks (ADR-0004). Mark all three workplans
finished with contract tests (80 passing).
This commit is contained in:
tegwick 2026-07-09 01:48:28 +02:00
parent df6acad828
commit 97437dfa18
28 changed files with 1295 additions and 69 deletions

View file

@ -1,18 +1,16 @@
# evidence-source # evidence-source
Headless document ingest for the citation-evidence ecosystem. Turns raw PDF Headless document ingest for the citation-evidence ecosystem. Turns raw PDF,
bytes into engine-owned evidence contracts — a `Document` (media type, HTML, and Markdown bytes into engine-owned evidence contracts — a `Document`
SHA-256 fingerprint, optional title/uri/metadata) and a (media type, SHA-256 fingerprint, optional title/uri/metadata) and a
`DocumentRepresentation` (`pdf-text`: canonical text, page map, gap-free `DocumentRepresentation` (canonical text, page/offset maps where applicable).
offset map). Ingest is pure over bytes: no persistence, no viewer state, no Ingest is pure over bytes: no persistence, no viewer state, no React.
React.
## Status ## Status
**Implemented: the PDF slice.** As of ESRC-WP-0001 this repo hosts the **Implemented:** PDF ingest (ESRC-WP-0001), HTML/Markdown ingest
extracted PDF ingest core that previously lived in `citation-evidence/src/source/`. (ESRC-WP-0002), PDF metadata enrichment (ESRC-WP-0003), and citation recovery
HTML/Markdown representations, richer metadata enrichment, and citation primitives (ESRC-WP-0004). See `workplans/` for history.
recovery are deferred to follow-on workplans (see `workplans/`).
## Install ## Install
@ -26,13 +24,17 @@ pnpm install # resolves @citation-evidence/engine via link:../citation-engine
## Usage ## Usage
```ts ```ts
import { ingestPdf } from "@citation-evidence/evidence-source"; import {
ingestPdf,
ingestHtml,
ingestMarkdown,
} from "@citation-evidence/evidence-source";
const { document, representation } = await ingestPdf(pdfBytes, { const pdf = await ingestPdf(pdfBytes, { filename: "contract.pdf" });
filename: "contract.pdf", const html = await ingestHtml(htmlBytes, { filename: "brief.html" });
}); const md = await ingestMarkdown("# Title\n\nBody text.", { filename: "notes.md" });
// document.fingerprint -> SHA-256 hex // document.fingerprint -> SHA-256 hex
// representation.canonicalText / pageMap / offsetMap // representation.canonicalText / pageMap / offsetMap (format-dependent)
``` ```
Browser upload helpers (in-memory `blob:` byte store + `ingestPdfFromFile`) Browser upload helpers (in-memory `blob:` byte store + `ingestPdfFromFile`)
@ -61,12 +63,14 @@ Point `EVIDENCE_SOURCE_FIXTURE_DIR` at the PDF corpus when the sibling
## Architecture ## Architecture
- `src/pdf/` — headless core: `ingest`, `extract`, `fingerprint`. Runtime-agnostic. - `src/pdf/` — PDF ingest, extraction, fingerprinting, metadata enrichment.
- `src/html/`, `src/markdown/` — reflowable-format ingest (ADR-0003).
- `src/recovery/` — re-ingest/reconcile, local quote search, discovery hooks (ADR-0004).
- `src/browser/` — browser upload surface: `byte-store`, `upload`. Separated by - `src/browser/` — browser upload surface: `byte-store`, `upload`. Separated by
an ESLint boundary so the core never imports it. an ESLint boundary so the core never imports it.
- Domain contracts come from `@citation-evidence/engine/shared`; none are - Domain contracts come from `@citation-evidence/engine/shared`; none are
copied locally. copied locally.
- Boundary and fixture decisions: `docs/ADR-0001`, `docs/ADR-0002`. - Boundary and fixture decisions: `docs/ADR-0001` through `docs/ADR-0004`.
`viewer-url` resolution stays in the consuming app (`citation-evidence`), which `viewer-url` resolution stays in the consuming app (`citation-evidence`), which
imports this package's ingest core through a thin façade. imports this package's ingest core through a thin façade.

View file

@ -8,8 +8,8 @@
## One-liner ## One-liner
Headless PDF ingest that turns document bytes into engine-owned evidence Headless document ingest that turns PDF, HTML, and Markdown bytes into
contracts (fingerprint, canonical text, page/offset maps). engine-owned evidence contracts (fingerprint, canonical text, page/offset maps).
--- ---
@ -18,25 +18,30 @@ contracts (fingerprint, canonical text, page/offset maps).
The rest of the citation-evidence ecosystem (anchoring, evidence linking, The rest of the citation-evidence ecosystem (anchoring, evidence linking,
binders) needs a stable, runtime-agnostic way to go from *raw document bytes* binders) needs a stable, runtime-agnostic way to go from *raw document bytes*
to a `Document` + `DocumentRepresentation`. This repo owns that transformation to a `Document` + `DocumentRepresentation`. This repo owns that transformation
for PDFs — pure over bytes, no persistence, no viewer, no UI framework. for PDF, HTML, and Markdown — pure over bytes, no persistence, no viewer,
no UI framework.
--- ---
## In Scope ## In Scope
- PDF byte ingest → `{ document, representation }` (`ingestPdf`) - PDF byte ingest → `{ document, representation }` (`ingestPdf`)
- HTML byte ingest → `{ document, representation }` (`ingestHtml`)
- Markdown byte ingest → `{ document, representation }` (`ingestMarkdown`)
- PDF text extraction → canonical text + page map + gap-free offset map (`extractPdf`) - PDF text extraction → canonical text + page map + gap-free offset map (`extractPdf`)
- PDF intrinsic metadata extraction with caller-wins merge (`extractPdfMetadata`)
- SHA-256 byte fingerprinting (`fingerprintBytes`) - SHA-256 byte fingerprinting (`fingerprintBytes`)
- Citation recovery primitives: re-ingest/reconcile, local quote search, discovery hooks
- Browser upload helpers behind a separate entry point (`createPdfByteStore`, `ingestPdfFromFile`) - Browser upload helpers behind a separate entry point (`createPdfByteStore`, `ingestPdfFromFile`)
--- ---
## Out of Scope ## Out of Scope
- HTML / Markdown representations (deferred — see `workplans/`)
- Persisting documents or representations (caller's job) - Persisting documents or representations (caller's job)
- Viewer URL resolution and blob-vs-fixture policy (stays in `citation-evidence`) - Viewer URL resolution and blob-vs-fixture policy (stays in `citation-evidence`)
- Citation recovery / external source discovery (deferred) - Stale selector detection and anchor resolution (stays in `evidence-anchor`)
- Network-backed external source discovery implementations (host registers hooks)
- Defining domain contracts — those are owned by `citation-engine` - Defining domain contracts — those are owned by `citation-engine`
--- ---
@ -51,7 +56,7 @@ for PDFs — pure over bytes, no persistence, no viewer, no UI framework.
## Not Relevant When ## Not Relevant When
- You need HTML/Markdown ingest (not yet implemented) - You need formats beyond PDF/HTML/Markdown (not yet implemented)
- You need viewer/session/UI behavior (see `citation-evidence`) - You need viewer/session/UI behavior (see `citation-evidence`)
- You are changing the `Document`/`DocumentRepresentation` contract (see `citation-engine`) - You are changing the `Document`/`DocumentRepresentation` contract (see `citation-engine`)
@ -60,7 +65,7 @@ for PDFs — pure over bytes, no persistence, no viewer, no UI framework.
## Current State ## Current State
- Status: active - Status: active
- Implementation: partial (PDF slice implemented; HTML/MD and recovery deferred) - Implementation: partial (PDF/HTML/MD ingest, metadata enrichment, recovery primitives)
- Stability: evolving - Stability: evolving
- Usage: internal (consumed by `citation-evidence` across the repo boundary) - Usage: internal (consumed by `citation-evidence` across the repo boundary)
@ -92,7 +97,7 @@ for PDFs — pure over bytes, no persistence, no viewer, no UI framework.
## Getting Oriented ## Getting Oriented
- Start with: `README.md`, then `docs/ADR-0001-extraction-boundary.md` - Start with: `README.md`, then `docs/ADR-0001-extraction-boundary.md`
- Key files / directories: `src/pdf/` (headless core), `src/browser/` (upload helpers) - Key files / directories: `src/pdf/`, `src/html/`, `src/markdown/`, `src/recovery/`, `src/browser/`
- Entry points: `src/index.ts` (headless), `src/browser/index.ts` (browser) - Entry points: `src/index.ts` (headless), `src/browser/index.ts` (browser)
--- ---

View file

@ -0,0 +1,58 @@
# ADR-0003 — HTML and Markdown representation contract
- **Status:** accepted
- **Date:** 2026-07-09
- **Workplan:** ESRC-WP-0002 (T01)
## Context
ESRC-WP-0001 established the PDF ingest boundary. ESRC-WP-0002 extends ingest to
HTML and Markdown while reusing engine-owned contracts from
`@citation-evidence/engine/shared`.
`RepresentationType` already reserves `html-dom` and `markdown-rendered` in
`citation-engine`; no engine schema change is required for this slice.
## Decision
### Representation types
| Source format | `mediaType` | `representationType` |
| ------------- | ------------------ | ---------------------- |
| HTML | `text/html` | `html-dom` |
| Markdown | `text/markdown` | `markdown-rendered` |
### Pageless offset semantics
Reflowable formats have no fixed pages:
- `pageMap` is **omitted** (undefined).
- `offsetMap` is a **single synthetic range** on page 1 covering
`[0, canonicalText.length)` with no gaps. This keeps
`TextPositionSelector` usable without inventing physical page geometry.
- `structureMap` remains `never` in the engine contract and is not populated.
### Canonical text
Both pipelines decode source bytes as UTF-8, derive plain text from the format,
then apply `normalize()` from `@citation-evidence/engine/shared`. HTML strips
active content (`script`, `style`, `iframe`, `object`, `embed`, `noscript`)
before text extraction.
### Content hash
`contentHash` is the SHA-256 fingerprint of the **source bytes** (same rule as
PDF ingest), not a hash of canonical text.
### Fixtures
HTML/Markdown contract tests use small inline fixtures owned by this repo. They
contain no PII and do not duplicate the upstream PDF corpus (ADR-0002).
## Consequences
- `ingestHtml` and `ingestMarkdown` mirror the PDF `{ document, representation }`
shape.
- Anchoring layers may add DOM/structural selectors later; this slice delivers
canonical text and position offsets only.
- Sanitization is extraction-time stripping, not a sandboxed renderer.

View file

@ -0,0 +1,54 @@
# ADR-0004 — Citation recovery boundary
- **Status:** accepted
- **Date:** 2026-07-09
- **Workplan:** ESRC-WP-0004 (T01)
## Context
`evidence-source` owns ingestion and representation generation. Citation
recovery spans stale-selector detection, local quote search, external source
discovery, and human confirmation — but selector resolution algorithms live in
`evidence-anchor`, and the `CitationRecoveryAttempt` type vocabulary lives in
`citation-engine`.
## Decision
### Owned by `evidence-source`
| Capability | Module | Notes |
| ---------- | ------ | ----- |
| Re-ingest + re-fingerprint | `src/recovery/re-ingest.ts` | Compare prior vs. fresh `Document.fingerprint`; flag `requiresReanchor` |
| Local canonical quote search | `src/recovery/local-search.ts` | Search `DocumentRepresentation.canonicalText` |
| Recovery attempt scaffold | `src/recovery/attempt.ts` | Local record using SharedContracts state vocabulary |
| Discovery hook interface | `src/recovery/discovery.ts` | Pluggable providers; **no network in core** |
### Owned elsewhere
| Capability | Owner |
| ---------- | ----- |
| Stale selector detection / resolution order | `evidence-anchor` |
| `CitationRecoveryAttempt` canonical type | `citation-engine` (future) |
| Human confirmation UX | `citation-work` / umbrella |
| External HTTP/API lookup implementations | Deployment-specific providers registered on the hook |
### Recovery flow
```text
CitationClue
→ createRecoveryAttempt (evidence-source)
→ local library search via searchCanonicalQuote (evidence-source)
→ optional SourceDiscoveryHook providers (registered by host)
→ reIngestAndCompare when fresh bytes arrive (evidence-source)
→ anchor layer re-resolves selectors when requiresReanchor is true
```
`evidence-source` **does not** decide whether a selector is stale. It supplies
primitives the anchor layer calls after it detects mismatch.
## Consequences
- No network coupling in the recovery core.
- Host apps register discovery providers explicitly (local-first default).
- Recovery state enum mirrors `wiki/SharedContracts.md` §2.6 locally until the
engine exports a shared `CitationRecoveryAttempt` type.

58
src/html/extract.ts Normal file
View file

@ -0,0 +1,58 @@
/**
* HTML text extraction canonical text + gap-free offset map.
*
* Implements ADR-0003 for the `html-dom` representation. Active content is
* stripped before text extraction; canonical text is produced via `normalize()`.
*/
import { normalize, type OffsetMap } from "@citation-evidence/engine/shared";
import { buildWholeTextOffsetMap } from "../shared/offset-map";
const ACTIVE_CONTENT =
/<(script|style|iframe|object|embed|noscript)\b[^>]*>[\s\S]*?<\/\1>/gi;
const BLOCK_BREAK = /<\/(p|div|h[1-6]|li|tr|section|article|blockquote|pre)>/gi;
const BR_TAG = /<br\s*\/?>/gi;
const TAG = /<[^>]+>/g;
export interface HtmlExtractionResult {
readonly canonicalText: string;
readonly sanitizedHtml: string;
readonly offsetMap: OffsetMap;
}
export function extractHtml(source: string): HtmlExtractionResult {
const sanitizedHtml = sanitizeHtml(source);
const rawText = htmlToPlainText(sanitizedHtml);
const canonicalText = normalize(rawText).text;
return {
canonicalText,
sanitizedHtml,
offsetMap: buildWholeTextOffsetMap(canonicalText),
};
}
export function sanitizeHtml(source: string): string {
return source.replace(ACTIVE_CONTENT, "");
}
function htmlToPlainText(html: string): string {
let text = html.replace(BR_TAG, "\n").replace(BLOCK_BREAK, "\n\n");
text = text.replace(TAG, "");
return decodeHtmlEntities(text);
}
function decodeHtmlEntities(text: string): string {
return text
.replace(/&nbsp;/gi, " ")
.replace(/&amp;/gi, "&")
.replace(/&lt;/gi, "<")
.replace(/&gt;/gi, ">")
.replace(/&quot;/gi, '"')
.replace(/&#39;/g, "'")
.replace(/&#x([0-9a-f]+);/gi, (_, hex: string) =>
String.fromCodePoint(Number.parseInt(hex, 16)),
)
.replace(/&#(\d+);/g, (_, dec: string) =>
String.fromCodePoint(Number.parseInt(dec, 10)),
);
}

54
src/html/ingest.test.ts Normal file
View file

@ -0,0 +1,54 @@
import { describe, expect, it } from "vitest";
import { fingerprintBytes } from "../pdf/fingerprint";
import { HTML_FIXTURE } from "../../tests/fixtures-text";
import { ingestHtml } from "./ingest";
import { extractHtml, sanitizeHtml } from "./extract";
describe("extractHtml", () => {
it("strips active content before text extraction", () => {
const { sanitizedHtml } = extractHtml(HTML_FIXTURE);
expect(sanitizeHtml(HTML_FIXTURE)).not.toMatch(/<script/i);
expect(sanitizedHtml).not.toMatch(/<script/i);
});
it("produces normalized canonical text with the known-good quote", () => {
const { canonicalText } = extractHtml(HTML_FIXTURE);
expect(canonicalText).toContain("Known good quote for HTML ingest testing.");
expect(canonicalText).toContain("Hello World");
});
});
describe("ingestHtml", () => {
const bytes = new TextEncoder().encode(HTML_FIXTURE);
it("produces a Document with text/html media type and SHA-256 fingerprint", async () => {
const { document } = await ingestHtml(bytes, { filename: "sample.html" });
expect(document.mediaType).toBe("text/html");
expect(document.fingerprint).toMatch(/^[0-9a-f]{64}$/);
expect(document.title).toBe("sample.html");
expect(document.fingerprint).toBe(await fingerprintBytes(bytes));
});
it("produces an html-dom representation with a gap-free offset map", async () => {
const { representation } = await ingestHtml(HTML_FIXTURE);
const text = representation.canonicalText ?? "";
const offsets = representation.offsetMap ?? [];
expect(representation.representationType).toBe("html-dom");
expect(representation.pageMap).toBeUndefined();
expect(offsets).toHaveLength(1);
expect(offsets[0]!.page).toBe(1);
expect(offsets[0]!.globalStart).toBe(0);
expect(offsets[0]!.globalEnd).toBe(text.length);
expect(text).toContain("Known good quote for HTML ingest testing.");
});
it("accepts string input and propagates uri/metadata", async () => {
const { document } = await ingestHtml(HTML_FIXTURE, {
uri: "file:///sample.html",
metadata: { source: "test" },
});
expect(document.uri).toBe("file:///sample.html");
expect(document.metadata).toEqual({ source: "test" });
});
});

81
src/html/ingest.ts Normal file
View file

@ -0,0 +1,81 @@
/**
* HTML ingest pipeline `{ document, representation }`.
*
* Implements ADR-0003 for the `html-dom` representation type.
*/
import {
newId,
type Document,
type DocumentRepresentation,
} from "@citation-evidence/engine/shared";
import { fingerprintBytes } from "../pdf/fingerprint";
import { toBytes, type ByteInput } from "../shared/to-bytes";
import { extractHtml } from "./extract";
const HTML_MEDIA_TYPE = "text/html";
export interface IngestHtmlOptions {
readonly filename?: string;
readonly title?: string;
readonly uri?: string;
readonly metadata?: Readonly<Record<string, unknown>>;
}
export interface IngestHtmlResult {
readonly document: Document;
readonly representation: DocumentRepresentation;
}
export type IngestHtmlInput = ByteInput | string;
export async function ingestHtml(
input: IngestHtmlInput,
options: IngestHtmlOptions = {},
): Promise<IngestHtmlResult> {
const { bytes, source } = await resolveInput(input);
const [fingerprint, extraction] = await Promise.all([
fingerprintBytes(bytes),
Promise.resolve(extractHtml(source)),
]);
const now = new Date().toISOString();
const documentId = newId("document");
const representationId = newId("representation");
const title = options.title ?? options.filename;
const document: Document = {
id: documentId,
mediaType: HTML_MEDIA_TYPE,
fingerprint,
createdAt: now,
updatedAt: now,
...(title !== undefined ? { title } : {}),
...(options.uri !== undefined ? { uri: options.uri } : {}),
...(options.metadata !== undefined ? { metadata: options.metadata } : {}),
};
const representation: DocumentRepresentation = {
id: representationId,
documentId,
representationType: "html-dom",
contentHash: fingerprint,
canonicalText: extraction.canonicalText,
offsetMap: extraction.offsetMap,
generatedAt: now,
};
return { document, representation };
}
async function resolveInput(
input: IngestHtmlInput,
): Promise<{ bytes: Uint8Array; source: string }> {
if (typeof input === "string") {
const bytes = new TextEncoder().encode(input);
return { bytes, source: input };
}
const bytes = await toBytes(input);
const source = new TextDecoder("utf-8", { fatal: false }).decode(bytes);
return { bytes, source };
}

View file

@ -1,7 +1,7 @@
/** /**
* `evidence-source` headless document ingest core (PDF slice). * `evidence-source` headless document ingest core.
* *
* This entry point exposes only runtime-agnostic ingest: pass PDF bytes, * This entry point exposes only runtime-agnostic ingest: pass document bytes,
* receive a `{ document, representation }` pair built entirely from * receive a `{ document, representation }` pair built entirely from
* engine-owned contracts (`@citation-evidence/engine/shared`). No browser, * engine-owned contracts (`@citation-evidence/engine/shared`). No browser,
* viewer, or persistence concerns leak through here. * viewer, or persistence concerns leak through here.
@ -17,4 +17,50 @@ export {
type IngestPdfResult, type IngestPdfResult,
} from "./pdf/ingest"; } from "./pdf/ingest";
export { extractPdf, type PdfExtractionResult } from "./pdf/extract"; export { extractPdf, type PdfExtractionResult } from "./pdf/extract";
export {
extractPdfMetadata,
mergeDocumentFields,
type MergeDocumentFieldOptions,
type MergedDocumentFields,
type PdfIntrinsicMetadata,
} from "./pdf/metadata";
export { fingerprintBytes } from "./pdf/fingerprint"; export { fingerprintBytes } from "./pdf/fingerprint";
export {
ingestHtml,
type IngestHtmlInput,
type IngestHtmlOptions,
type IngestHtmlResult,
} from "./html/ingest";
export { extractHtml, sanitizeHtml, type HtmlExtractionResult } from "./html/extract";
export {
ingestMarkdown,
type IngestMarkdownInput,
type IngestMarkdownOptions,
type IngestMarkdownResult,
} from "./markdown/ingest";
export { extractMarkdown, type MarkdownExtractionResult } from "./markdown/extract";
export {
applyLocalQuoteSearch,
createRecoveryAttempt,
createSourceDiscoveryRegistry,
reconcileFingerprint,
reIngestAndCompare,
recoveryStateAfterLocalSearch,
searchCanonicalQuote,
withRecoveryState,
type CitationClue,
type CitationRecoveryState,
type DiscoveryCandidate,
type IngestFn,
type IngestOutcome,
type QuoteMatch,
type ReIngestInput,
type ReIngestResult,
type ReconcileResult,
type RecoveryAttempt,
type SourceDiscoveryHook,
type SourceDiscoveryRegistry,
} from "./recovery";

53
src/markdown/extract.ts Normal file
View file

@ -0,0 +1,53 @@
/**
* Markdown text extraction canonical text + gap-free offset map.
*
* Implements ADR-0003 for the `markdown-rendered` representation. Markup is
* stripped to plain text before `normalize()` is applied.
*/
import { normalize, type OffsetMap } from "@citation-evidence/engine/shared";
import { buildWholeTextOffsetMap } from "../shared/offset-map";
const FENCED_CODE = /```[\s\S]*?```/g;
const INLINE_CODE = /`([^`]+)`/g;
const IMAGE = /!\[([^\]]*)\]\([^)]+\)/g;
const LINK = /\[([^\]]+)\]\([^)]+\)/g;
const HEADING = /^#{1,6}\s+/gm;
const BLOCKQUOTE = /^>\s?/gm;
const LIST_MARKER = /^[\s]*[-*+]\s+/gm;
const ORDERED_LIST = /^[\s]*\d+\.\s+/gm;
const EMPHASIS = /(\*\*|__)(.*?)\1/g;
const ITALIC = /(\*|_)([^*_]+)\1/g;
const STRIKETHROUGH = /~~(.*?)~~/g;
const HORIZONTAL_RULE = /^[-*_]{3,}\s*$/gm;
export interface MarkdownExtractionResult {
readonly canonicalText: string;
readonly offsetMap: OffsetMap;
}
export function extractMarkdown(source: string): MarkdownExtractionResult {
const rawText = markdownToPlainText(source);
const canonicalText = normalize(rawText).text;
return {
canonicalText,
offsetMap: buildWholeTextOffsetMap(canonicalText),
};
}
function markdownToPlainText(source: string): string {
let text = source;
text = text.replace(FENCED_CODE, "\n\n");
text = text.replace(INLINE_CODE, "$1");
text = text.replace(IMAGE, "$1");
text = text.replace(LINK, "$1");
text = text.replace(HEADING, "");
text = text.replace(BLOCKQUOTE, "");
text = text.replace(LIST_MARKER, "");
text = text.replace(ORDERED_LIST, "");
text = text.replace(EMPHASIS, "$2");
text = text.replace(ITALIC, "$2");
text = text.replace(STRIKETHROUGH, "$1");
text = text.replace(HORIZONTAL_RULE, "\n\n");
return text;
}

View file

@ -0,0 +1,49 @@
import { describe, expect, it } from "vitest";
import { fingerprintBytes } from "../pdf/fingerprint";
import { MARKDOWN_FIXTURE } from "../../tests/fixtures-text";
import { ingestMarkdown } from "./ingest";
import { extractMarkdown } from "./extract";
describe("extractMarkdown", () => {
it("strips markup and preserves the known-good quote", () => {
const { canonicalText } = extractMarkdown(MARKDOWN_FIXTURE);
expect(canonicalText).toContain("Known good quote for Markdown ingest testing.");
expect(canonicalText).toContain("bold");
expect(canonicalText).toContain("italic");
expect(canonicalText).not.toContain("**");
});
});
describe("ingestMarkdown", () => {
const bytes = new TextEncoder().encode(MARKDOWN_FIXTURE);
it("produces a Document with text/markdown media type and SHA-256 fingerprint", async () => {
const { document } = await ingestMarkdown(bytes, { filename: "sample.md" });
expect(document.mediaType).toBe("text/markdown");
expect(document.fingerprint).toMatch(/^[0-9a-f]{64}$/);
expect(document.title).toBe("sample.md");
expect(document.fingerprint).toBe(await fingerprintBytes(bytes));
});
it("produces a markdown-rendered representation with a gap-free offset map", async () => {
const { representation } = await ingestMarkdown(MARKDOWN_FIXTURE);
const text = representation.canonicalText ?? "";
const offsets = representation.offsetMap ?? [];
expect(representation.representationType).toBe("markdown-rendered");
expect(representation.pageMap).toBeUndefined();
expect(offsets).toHaveLength(1);
expect(offsets[0]!.page).toBe(1);
expect(offsets[0]!.globalStart).toBe(0);
expect(offsets[0]!.globalEnd).toBe(text.length);
expect(text).toContain("Known good quote for Markdown ingest testing.");
});
it("uses explicit title over filename", async () => {
const { document } = await ingestMarkdown(bytes, {
filename: "sample.md",
title: "Custom Title",
});
expect(document.title).toBe("Custom Title");
});
});

81
src/markdown/ingest.ts Normal file
View file

@ -0,0 +1,81 @@
/**
* Markdown ingest pipeline `{ document, representation }`.
*
* Implements ADR-0003 for the `markdown-rendered` representation type.
*/
import {
newId,
type Document,
type DocumentRepresentation,
} from "@citation-evidence/engine/shared";
import { fingerprintBytes } from "../pdf/fingerprint";
import { toBytes, type ByteInput } from "../shared/to-bytes";
import { extractMarkdown } from "./extract";
const MARKDOWN_MEDIA_TYPE = "text/markdown";
export interface IngestMarkdownOptions {
readonly filename?: string;
readonly title?: string;
readonly uri?: string;
readonly metadata?: Readonly<Record<string, unknown>>;
}
export interface IngestMarkdownResult {
readonly document: Document;
readonly representation: DocumentRepresentation;
}
export type IngestMarkdownInput = ByteInput | string;
export async function ingestMarkdown(
input: IngestMarkdownInput,
options: IngestMarkdownOptions = {},
): Promise<IngestMarkdownResult> {
const { bytes, source } = await resolveInput(input);
const [fingerprint, extraction] = await Promise.all([
fingerprintBytes(bytes),
Promise.resolve(extractMarkdown(source)),
]);
const now = new Date().toISOString();
const documentId = newId("document");
const representationId = newId("representation");
const title = options.title ?? options.filename;
const document: Document = {
id: documentId,
mediaType: MARKDOWN_MEDIA_TYPE,
fingerprint,
createdAt: now,
updatedAt: now,
...(title !== undefined ? { title } : {}),
...(options.uri !== undefined ? { uri: options.uri } : {}),
...(options.metadata !== undefined ? { metadata: options.metadata } : {}),
};
const representation: DocumentRepresentation = {
id: representationId,
documentId,
representationType: "markdown-rendered",
contentHash: fingerprint,
canonicalText: extraction.canonicalText,
offsetMap: extraction.offsetMap,
generatedAt: now,
};
return { document, representation };
}
async function resolveInput(
input: IngestMarkdownInput,
): Promise<{ bytes: Uint8Array; source: string }> {
if (typeof input === "string") {
const bytes = new TextEncoder().encode(input);
return { bytes, source: input };
}
const bytes = await toBytes(input);
const source = new TextDecoder("utf-8", { fatal: false }).decode(bytes);
return { bytes, source };
}

View file

@ -18,6 +18,8 @@ import {
} from "@citation-evidence/engine/shared"; } from "@citation-evidence/engine/shared";
import { extractPdf } from "./extract"; import { extractPdf } from "./extract";
import { fingerprintBytes } from "./fingerprint"; import { fingerprintBytes } from "./fingerprint";
import { extractPdfMetadata, mergeDocumentFields } from "./metadata";
import { toBytes, type ByteInput } from "../shared/to-bytes";
const PDF_MEDIA_TYPE = "application/pdf"; const PDF_MEDIA_TYPE = "application/pdf";
@ -37,22 +39,23 @@ export interface IngestPdfResult {
readonly representation: DocumentRepresentation; readonly representation: DocumentRepresentation;
} }
export type IngestPdfInput = Uint8Array | ArrayBuffer | Blob; export type IngestPdfInput = ByteInput;
export async function ingestPdf( export async function ingestPdf(
input: IngestPdfInput, input: IngestPdfInput,
options: IngestPdfOptions = {}, options: IngestPdfOptions = {},
): Promise<IngestPdfResult> { ): Promise<IngestPdfResult> {
const bytes = await toBytes(input); const bytes = await toBytes(input);
const [fingerprint, extraction] = await Promise.all([ const [fingerprint, extraction, intrinsicMetadata] = await Promise.all([
fingerprintBytes(bytes), fingerprintBytes(bytes),
extractPdf(bytes), extractPdf(bytes),
extractPdfMetadata(bytes),
]); ]);
const now = new Date().toISOString(); const now = new Date().toISOString();
const documentId = newId("document"); const documentId = newId("document");
const representationId = newId("representation"); const representationId = newId("representation");
const title = options.title ?? options.filename; const merged = mergeDocumentFields(intrinsicMetadata, options);
const document: Document = { const document: Document = {
id: documentId, id: documentId,
@ -60,9 +63,7 @@ export async function ingestPdf(
fingerprint, fingerprint,
createdAt: now, createdAt: now,
updatedAt: now, updatedAt: now,
...(title !== undefined ? { title } : {}), ...merged,
...(options.uri !== undefined ? { uri: options.uri } : {}),
...(options.metadata !== undefined ? { metadata: options.metadata } : {}),
}; };
const representation: DocumentRepresentation = { const representation: DocumentRepresentation = {
@ -78,11 +79,3 @@ export async function ingestPdf(
return { document, representation }; return { document, representation };
} }
async function toBytes(input: IngestPdfInput): Promise<Uint8Array> {
if (input instanceof Uint8Array) return input;
if (input instanceof ArrayBuffer) return new Uint8Array(input);
// Blob (covers `File` in browsers — File extends Blob).
const buf = await input.arrayBuffer();
return new Uint8Array(buf);
}

77
src/pdf/metadata.test.ts Normal file
View file

@ -0,0 +1,77 @@
import { describe, expect, it } from "vitest";
import {
extractPdfMetadata,
mergeDocumentFields,
type PdfIntrinsicMetadata,
} from "./metadata";
import {
fixtureBytes,
fixtureCorpusAvailable,
loadFixtures,
} from "../../tests/fixtures";
describe("mergeDocumentFields", () => {
const intrinsic: PdfIntrinsicMetadata = {
title: "Intrinsic Title",
author: "Intrinsic Author",
subject: "Intrinsic Subject",
creationDate: "2024-01-01T00:00:00.000Z",
xmp: { "dc:title": "XMP Title" },
};
it("prefers caller title over intrinsic title and filename", () => {
const merged = mergeDocumentFields(intrinsic, {
filename: "file.pdf",
title: "Caller Title",
});
expect(merged.title).toBe("Caller Title");
});
it("falls back to filename when caller title is absent", () => {
const merged = mergeDocumentFields(intrinsic, { filename: "file.pdf" });
expect(merged.title).toBe("file.pdf");
});
it("falls back to intrinsic title when caller title and filename are absent", () => {
const merged = mergeDocumentFields(intrinsic);
expect(merged.title).toBe("Intrinsic Title");
});
it("merges intrinsic metadata under caller metadata with caller winning", () => {
const merged = mergeDocumentFields(intrinsic, {
metadata: { author: "Caller Author", source: "test" },
});
expect(merged.metadata).toEqual({
author: "Caller Author",
subject: "Intrinsic Subject",
creationDate: "2024-01-01T00:00:00.000Z",
xmp: { "dc:title": "XMP Title" },
source: "test",
});
});
it("omits metadata when neither intrinsic nor caller metadata is present", () => {
const merged = mergeDocumentFields({});
expect(merged.metadata).toBeUndefined();
});
});
const hasCorpus = fixtureCorpusAvailable();
const FIXTURES = hasCorpus ? loadFixtures() : [];
describe.skipIf(!hasCorpus)("extractPdfMetadata — fixture corpus", { timeout: 30_000 }, () => {
for (const fixture of FIXTURES) {
it(`${fixture.id} extracts without throwing and caller overrides intrinsic title`, async () => {
const bytes = fixtureBytes(fixture.filename);
const intrinsic = await extractPdfMetadata(bytes);
expect(intrinsic).toBeTypeOf("object");
const merged = mergeDocumentFields(intrinsic, {
filename: fixture.filename,
title: "Override Title",
});
expect(merged.title).toBe("Override Title");
});
}
});

185
src/pdf/metadata.ts Normal file
View file

@ -0,0 +1,185 @@
/**
* PDF intrinsic metadata extraction and merge precedence.
*
* Implements ESRC-WP-0003: extract info-dictionary and XMP fields via PDF.js,
* then merge into the Document record with caller-supplied values winning.
*/
import { getDocument } from "pdfjs-dist";
import type { Metadata } from "pdfjs-dist";
export interface PdfIntrinsicMetadata {
readonly title?: string;
readonly author?: string;
readonly subject?: string;
readonly keywords?: string;
readonly creator?: string;
readonly producer?: string;
readonly creationDate?: string;
readonly modificationDate?: string;
readonly xmp?: Readonly<Record<string, string>>;
}
export interface MergedDocumentFields {
readonly title?: string;
readonly uri?: string;
readonly metadata?: Readonly<Record<string, unknown>>;
}
export interface MergeDocumentFieldOptions {
readonly filename?: string;
readonly title?: string;
readonly uri?: string;
readonly metadata?: Readonly<Record<string, unknown>>;
}
export async function extractPdfMetadata(
bytes: Uint8Array,
): Promise<PdfIntrinsicMetadata> {
const data = new Uint8Array(bytes);
const loadingTask = getDocument({ data });
const doc = await loadingTask.promise;
try {
const { info, metadata } = await doc.getMetadata();
return normalizePdfMetadata(info as Record<string, unknown>, metadata);
} finally {
await doc.destroy();
}
}
export function mergeDocumentFields(
intrinsic: PdfIntrinsicMetadata,
options: MergeDocumentFieldOptions = {},
): MergedDocumentFields {
const intrinsicRecord = intrinsicMetadataRecord(intrinsic);
const mergedMetadata =
Object.keys(intrinsicRecord).length > 0 || options.metadata !== undefined
? { ...intrinsicRecord, ...options.metadata }
: undefined;
const title =
options.title ??
options.filename ??
intrinsic.title;
return {
...(title !== undefined ? { title } : {}),
...(options.uri !== undefined ? { uri: options.uri } : {}),
...(mergedMetadata !== undefined ? { metadata: mergedMetadata } : {}),
};
}
function normalizePdfMetadata(
info: Record<string, unknown>,
metadata: Metadata | null,
): PdfIntrinsicMetadata {
const result: {
title?: string;
author?: string;
subject?: string;
keywords?: string;
creator?: string;
producer?: string;
creationDate?: string;
modificationDate?: string;
xmp?: Record<string, string>;
} = {};
const title = stringField(info, "Title");
if (title !== undefined) result.title = title;
const author = stringField(info, "Author");
if (author !== undefined) result.author = author;
const subject = stringField(info, "Subject");
if (subject !== undefined) result.subject = subject;
const keywords = keywordsField(info.Keywords);
if (keywords !== undefined) result.keywords = keywords;
const creator = stringField(info, "Creator");
if (creator !== undefined) result.creator = creator;
const producer = stringField(info, "Producer");
if (producer !== undefined) result.producer = producer;
const creationDate = pdfDateField(info.CreationDate);
if (creationDate !== undefined) result.creationDate = creationDate;
const modificationDate = pdfDateField(info.ModDate);
if (modificationDate !== undefined) result.modificationDate = modificationDate;
const xmp = xmpFields(metadata);
if (xmp !== undefined) result.xmp = xmp;
return result;
}
function intrinsicMetadataRecord(
intrinsic: PdfIntrinsicMetadata,
): Record<string, unknown> {
const record: Record<string, unknown> = {};
if (intrinsic.author !== undefined) record.author = intrinsic.author;
if (intrinsic.subject !== undefined) record.subject = intrinsic.subject;
if (intrinsic.keywords !== undefined) record.keywords = intrinsic.keywords;
if (intrinsic.creator !== undefined) record.creator = intrinsic.creator;
if (intrinsic.producer !== undefined) record.producer = intrinsic.producer;
if (intrinsic.creationDate !== undefined) {
record.creationDate = intrinsic.creationDate;
}
if (intrinsic.modificationDate !== undefined) {
record.modificationDate = intrinsic.modificationDate;
}
if (intrinsic.xmp !== undefined) record.xmp = intrinsic.xmp;
return record;
}
function stringField(
info: Record<string, unknown>,
key: string,
): string | undefined {
const value = info[key];
if (typeof value !== "string") return undefined;
const trimmed = value.trim();
return trimmed.length > 0 ? trimmed : undefined;
}
function keywordsField(value: unknown): string | undefined {
if (typeof value === "string") {
const trimmed = value.trim();
return trimmed.length > 0 ? trimmed : undefined;
}
if (Array.isArray(value)) {
const parts = value
.filter((entry): entry is string => typeof entry === "string")
.map((entry) => entry.trim())
.filter((entry) => entry.length > 0);
return parts.length > 0 ? parts.join(", ") : undefined;
}
return undefined;
}
function pdfDateField(value: unknown): string | undefined {
if (typeof value !== "string") return undefined;
const trimmed = value.trim();
if (trimmed.length === 0) return undefined;
const match = /^D:(\d{4})(\d{2})(\d{2})(\d{2})?(\d{2})?(\d{2})?/.exec(trimmed);
if (!match) return trimmed;
const [, year, month, day, hour = "00", minute = "00", second = "00"] = match;
const iso = `${year}-${month}-${day}T${hour}:${minute}:${second}Z`;
const parsed = new Date(iso);
return Number.isNaN(parsed.getTime()) ? trimmed : parsed.toISOString();
}
function xmpFields(metadata: Metadata | null): Record<string, string> | undefined {
if (metadata === null) return undefined;
const all = metadata.getAll() as Record<string, unknown>;
const xmp: Record<string, string> = {};
for (const [key, value] of Object.entries(all)) {
if (typeof value === "string" && value.trim().length > 0) {
xmp[key] = value.trim();
}
}
return Object.keys(xmp).length > 0 ? xmp : undefined;
}

39
src/recovery/attempt.ts Normal file
View file

@ -0,0 +1,39 @@
import { newId } from "@citation-evidence/engine/shared";
import type { CitationClue, CitationRecoveryState } from "./types";
export interface RecoveryAttempt {
readonly id: string;
readonly state: CitationRecoveryState;
readonly clue: CitationClue;
readonly createdAt: string;
readonly updatedAt: string;
readonly documentId?: string;
readonly notes: readonly string[];
}
export function createRecoveryAttempt(clue: CitationClue): RecoveryAttempt {
const now = new Date().toISOString();
return {
id: newId("citation-recovery"),
state: "created",
clue,
createdAt: now,
updatedAt: now,
notes: [],
};
}
export function withRecoveryState(
attempt: RecoveryAttempt,
state: CitationRecoveryState,
note?: string,
): RecoveryAttempt {
const now = new Date().toISOString();
return {
...attempt,
state,
updatedAt: now,
notes: note === undefined ? attempt.notes : [...attempt.notes, note],
};
}

37
src/recovery/discovery.ts Normal file
View file

@ -0,0 +1,37 @@
import type { CitationClue } from "./types";
export interface DiscoveryCandidate {
readonly uri: string;
readonly title?: string;
readonly confidence: number;
readonly source: "local" | "external";
readonly metadata?: Readonly<Record<string, unknown>>;
}
/** Pluggable discovery provider — implementations may perform network I/O. */
export interface SourceDiscoveryHook {
readonly name: string;
discover(clue: CitationClue): Promise<readonly DiscoveryCandidate[]>;
}
export interface SourceDiscoveryRegistry {
register(hook: SourceDiscoveryHook): void;
discover(clue: CitationClue): Promise<readonly DiscoveryCandidate[]>;
}
export function createSourceDiscoveryRegistry(): SourceDiscoveryRegistry {
const hooks: SourceDiscoveryHook[] = [];
return {
register(hook: SourceDiscoveryHook) {
hooks.push(hook);
},
async discover(clue: CitationClue): Promise<readonly DiscoveryCandidate[]> {
const results = await Promise.all(hooks.map((hook) => hook.discover(clue)));
return results
.flat()
.sort((a, b) => b.confidence - a.confidence);
},
};
}

26
src/recovery/index.ts Normal file
View file

@ -0,0 +1,26 @@
export type { CitationClue, CitationRecoveryState } from "./types";
export {
createRecoveryAttempt,
withRecoveryState,
type RecoveryAttempt,
} from "./attempt";
export { reconcileFingerprint, type ReconcileResult } from "./reconcile";
export {
reIngestAndCompare,
type IngestFn,
type IngestOutcome,
type ReIngestInput,
type ReIngestResult,
} from "./re-ingest";
export {
applyLocalQuoteSearch,
recoveryStateAfterLocalSearch,
searchCanonicalQuote,
type QuoteMatch,
} from "./local-search";
export {
createSourceDiscoveryRegistry,
type DiscoveryCandidate,
type SourceDiscoveryHook,
type SourceDiscoveryRegistry,
} from "./discovery";

View file

@ -0,0 +1,65 @@
import { normalize, type DocumentRepresentation } from "@citation-evidence/engine/shared";
import type { CitationRecoveryState } from "./types";
import { withRecoveryState, type RecoveryAttempt } from "./attempt";
export interface QuoteMatch {
readonly start: number;
readonly end: number;
readonly matchedText: string;
}
export function searchCanonicalQuote(
representation: Pick<DocumentRepresentation, "canonicalText">,
quote: string,
): readonly QuoteMatch[] {
const canonicalText = representation.canonicalText ?? "";
const normalizedQuote = normalize(quote).text;
if (normalizedQuote.length === 0 || canonicalText.length === 0) return [];
const normalizedCanonical = normalize(canonicalText).text;
const matches: QuoteMatch[] = [];
let from = 0;
while (from <= normalizedCanonical.length) {
const index = normalizedCanonical.indexOf(normalizedQuote, from);
if (index === -1) break;
matches.push({
start: index,
end: index + normalizedQuote.length,
matchedText: normalizedCanonical.slice(index, index + normalizedQuote.length),
});
from = index + 1;
}
return matches;
}
export function applyLocalQuoteSearch(
attempt: RecoveryAttempt,
representation: Pick<DocumentRepresentation, "canonicalText">,
): RecoveryAttempt {
const quote = attempt.clue.quote;
if (quote === undefined || quote.trim().length === 0) {
return withRecoveryState(attempt, "quote-not-found", "No quote in clue");
}
const matches = searchCanonicalQuote(representation, quote);
if (matches.length === 1) {
return withRecoveryState(attempt, "quote-found", "Exact quote match");
}
if (matches.length > 1) {
return withRecoveryState(
attempt,
"candidate-passages-found",
`${matches.length} quote matches`,
);
}
return withRecoveryState(attempt, "quote-not-found", "Quote not found locally");
}
export function recoveryStateAfterLocalSearch(
matchCount: number,
): CitationRecoveryState {
if (matchCount === 1) return "quote-found";
if (matchCount > 1) return "candidate-passages-found";
return "quote-not-found";
}

33
src/recovery/re-ingest.ts Normal file
View file

@ -0,0 +1,33 @@
import type { Document, DocumentRepresentation } from "@citation-evidence/engine/shared";
import { reconcileFingerprint, type ReconcileResult } from "./reconcile";
export interface IngestOutcome {
readonly document: Document;
readonly representation: DocumentRepresentation;
}
export type IngestFn<TOptions = void> = (
bytes: Uint8Array,
options?: TOptions,
) => Promise<IngestOutcome>;
export interface ReIngestInput<TOptions = void> {
readonly previous: Document;
readonly bytes: Uint8Array;
readonly options?: TOptions;
}
export interface ReIngestResult {
readonly outcome: IngestOutcome;
readonly reconcile: ReconcileResult;
}
export async function reIngestAndCompare<TOptions>(
input: ReIngestInput<TOptions>,
ingest: IngestFn<TOptions>,
): Promise<ReIngestResult> {
const outcome = await ingest(input.bytes, input.options);
const reconcile = reconcileFingerprint(input.previous, outcome.document);
return { outcome, reconcile };
}

27
src/recovery/reconcile.ts Normal file
View file

@ -0,0 +1,27 @@
import type { Document } from "@citation-evidence/engine/shared";
export interface ReconcileResult {
readonly unchanged: boolean;
readonly previousFingerprint?: string;
readonly currentFingerprint: string;
readonly requiresReanchor: boolean;
}
export function reconcileFingerprint(
previous: Pick<Document, "fingerprint">,
current: Pick<Document, "fingerprint">,
): ReconcileResult {
const previousFingerprint = previous.fingerprint;
const currentFingerprint = current.fingerprint ?? "";
const unchanged =
previousFingerprint !== undefined &&
currentFingerprint.length > 0 &&
previousFingerprint === currentFingerprint;
return {
unchanged,
previousFingerprint,
currentFingerprint,
requiresReanchor: !unchanged,
};
}

View file

@ -0,0 +1,97 @@
import { describe, expect, it, vi } from "vitest";
import { HTML_FIXTURE } from "../../tests/fixtures-text";
import { ingestHtml } from "../html/ingest";
import {
applyLocalQuoteSearch,
createRecoveryAttempt,
createSourceDiscoveryRegistry,
reconcileFingerprint,
reIngestAndCompare,
recoveryStateAfterLocalSearch,
searchCanonicalQuote,
} from "./index";
describe("reconcileFingerprint", () => {
it("flags re-anchor when fingerprints differ", () => {
const result = reconcileFingerprint(
{ fingerprint: "aaa" },
{ fingerprint: "bbb" },
);
expect(result.unchanged).toBe(false);
expect(result.requiresReanchor).toBe(true);
});
it("reports unchanged when fingerprints match", () => {
const fp = "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855";
const result = reconcileFingerprint({ fingerprint: fp }, { fingerprint: fp });
expect(result.unchanged).toBe(true);
expect(result.requiresReanchor).toBe(false);
});
});
describe("searchCanonicalQuote", () => {
it("finds normalized quote matches in canonical text", async () => {
const { representation } = await ingestHtml(HTML_FIXTURE);
const matches = searchCanonicalQuote(
representation,
"Known good quote for HTML ingest testing.",
);
expect(matches).toHaveLength(1);
expect(matches[0]!.matchedText).toContain("Known good quote");
});
it("maps match counts to recovery states", () => {
expect(recoveryStateAfterLocalSearch(1)).toBe("quote-found");
expect(recoveryStateAfterLocalSearch(2)).toBe("candidate-passages-found");
expect(recoveryStateAfterLocalSearch(0)).toBe("quote-not-found");
});
});
describe("applyLocalQuoteSearch", () => {
it("updates attempt state from local quote search", async () => {
const { representation } = await ingestHtml(HTML_FIXTURE);
const attempt = createRecoveryAttempt({
quote: "Known good quote for HTML ingest testing.",
});
const updated = applyLocalQuoteSearch(attempt, representation);
expect(updated.state).toBe("quote-found");
});
});
describe("reIngestAndCompare", () => {
it("detects unchanged document on identical bytes", async () => {
const bytes = new TextEncoder().encode(HTML_FIXTURE);
const first = await ingestHtml(bytes);
const { reconcile } = await reIngestAndCompare(
{ previous: first.document, bytes },
ingestHtml,
);
expect(reconcile.unchanged).toBe(true);
expect(reconcile.requiresReanchor).toBe(false);
});
});
describe("createSourceDiscoveryRegistry", () => {
it("aggregates provider results sorted by confidence", async () => {
const registry = createSourceDiscoveryRegistry();
registry.register({
name: "low",
discover: vi.fn(async () => [
{ uri: "file:///a", confidence: 0.2, source: "local" },
]),
});
registry.register({
name: "high",
discover: vi.fn(async () => [
{ uri: "https://example.com/doc", confidence: 0.9, source: "external" },
]),
});
const candidates = await registry.discover({ title: "Example" });
expect(candidates).toHaveLength(2);
expect(candidates[0]!.confidence).toBe(0.9);
expect(candidates[1]!.confidence).toBe(0.2);
});
});

31
src/recovery/types.ts Normal file
View file

@ -0,0 +1,31 @@
/**
* Citation recovery vocabulary aligned with citation-engine SharedContracts §2.6.
*
* The canonical `CitationRecoveryAttempt` type will live in citation-engine;
* this repo uses the same state strings for local attempt records.
*/
export type CitationRecoveryState =
| "created"
| "source-found-fulltext"
| "source-found-preview-only"
| "source-found-metadata-only"
| "source-not-found"
| "quote-found"
| "quote-not-found"
| "candidate-passages-found"
| "manual-confirmation-needed"
| "confirmed"
| "annotation-created"
| "failed";
/** Clue supplied when starting or continuing a recovery attempt. */
export interface CitationClue {
readonly quote?: string;
readonly prefix?: string;
readonly suffix?: string;
readonly title?: string;
readonly uri?: string;
readonly doi?: string;
readonly metadata?: Readonly<Record<string, unknown>>;
}

19
src/shared/offset-map.ts Normal file
View file

@ -0,0 +1,19 @@
import type { OffsetMap } from "@citation-evidence/engine/shared";
/**
* Gap-free offset map for pageless representations (HTML, Markdown).
*
* Uses a single synthetic page (page 1) covering `[0, canonicalText.length)`.
* `pageMap` stays unset there are no physical pages.
*/
export function buildWholeTextOffsetMap(canonicalText: string): OffsetMap {
const length = canonicalText.length;
return [
{
page: 1,
globalStart: 0,
globalEnd: length,
pageLength: length,
},
];
}

9
src/shared/to-bytes.ts Normal file
View file

@ -0,0 +1,9 @@
/** Runtime-agnostic byte input accepted by ingest pipelines. */
export type ByteInput = Uint8Array | ArrayBuffer | Blob;
export async function toBytes(input: ByteInput): Promise<Uint8Array> {
if (input instanceof Uint8Array) return input;
if (input instanceof ArrayBuffer) return new Uint8Array(input);
const buf = await input.arrayBuffer();
return new Uint8Array(buf);
}

42
tests/fixtures-text.ts Normal file
View file

@ -0,0 +1,42 @@
/**
* Inline HTML/Markdown fixtures for contract tests (ADR-0003).
*
* Unlike the PDF corpus (ADR-0002), these are small synthetic documents with
* no PII and are owned by this repo.
*/
export interface TextFixture {
readonly id: string;
readonly knownGoodQuote: string;
}
export const HTML_FIXTURE = `<!DOCTYPE html>
<html>
<head>
<title>Fixture Title</title>
<style>.hidden { display: none; }</style>
</head>
<body>
<h1>Hello World</h1>
<p>Known good quote for HTML ingest testing.</p>
<script>alert("removed")</script>
</body>
</html>`;
export const MARKDOWN_FIXTURE = `# Hello World
Known good quote for Markdown ingest testing.
**bold** and _italic_ text.
`;
export const TEXT_FIXTURES: readonly TextFixture[] = [
{
id: "html-basic",
knownGoodQuote: "Known good quote for HTML ingest testing.",
},
{
id: "markdown-basic",
knownGoodQuote: "Known good quote for Markdown ingest testing.",
},
];

View file

@ -4,14 +4,15 @@ type: workplan
title: "HTML and Markdown source representations" title: "HTML and Markdown source representations"
domain: infotech domain: infotech
repo: evidence-source repo: evidence-source
status: proposed status: finished
owner: codex owner: codex
topic_slug: citation_evidence_mvp topic_slug: citation_evidence_mvp
created: "2026-07-08" created: "2026-07-08"
updated: "2026-07-08" updated: "2026-07-09"
spec_refs: spec_refs:
- README.md - README.md
- docs/ADR-0001-extraction-boundary.md - docs/ADR-0001-extraction-boundary.md
- docs/ADR-0003-html-markdown-representation.md
- ../citation-engine/src/shared/document.ts - ../citation-engine/src/shared/document.ts
state_hub_workstream_id: "c0d28886-e02f-4df2-abb6-ebe179171c94" state_hub_workstream_id: "c0d28886-e02f-4df2-abb6-ebe179171c94"
--- ---
@ -22,18 +23,19 @@ The PDF slice (ESRC-WP-0001) established the ingest boundary and contract shape.
This workplan extends ingest to HTML and Markdown sources, producing the same This workplan extends ingest to HTML and Markdown sources, producing the same
`{ document, representation }` pair over engine-owned contracts. `{ document, representation }` pair over engine-owned contracts.
## Open questions ## Decisions (ADR-0003)
- Which `representationType` values do HTML/MD map to, and do they need new - HTML maps to `representationType: "html-dom"`; Markdown to
`DocumentRepresentation` fields in `citation-engine`? `"markdown-rendered"`.
- How is canonical text + offset map derived for reflowable formats that have - Pageless formats omit `pageMap`; `offsetMap` is a single synthetic page-1
no fixed pages? (Page map may be empty or synthetic.) range covering `[0, canonicalText.length)`.
- Inline fixtures owned by this repo (no PII); PDF corpus rules unchanged.
## Sketch ## Sketch
```task ```task
id: ESRC-WP-0002-T01 id: ESRC-WP-0002-T01
status: todo status: done
priority: medium priority: medium
state_hub_task_id: "bf1d2a2f-41b4-44b5-b1d1-4f0eec5fbf59" state_hub_task_id: "bf1d2a2f-41b4-44b5-b1d1-4f0eec5fbf59"
``` ```
@ -42,7 +44,7 @@ Define the HTML/MD representation contract with `citation-engine`; agree
```task ```task
id: ESRC-WP-0002-T02 id: ESRC-WP-0002-T02
status: todo status: done
priority: medium priority: medium
state_hub_task_id: "66320911-a4b2-404f-b025-8b23dc0c9680" state_hub_task_id: "66320911-a4b2-404f-b025-8b23dc0c9680"
``` ```
@ -51,9 +53,9 @@ reusing `fingerprintBytes` and the canonical-text normalizer.
```task ```task
id: ESRC-WP-0002-T03 id: ESRC-WP-0002-T03
status: todo status: done
priority: medium priority: medium
state_hub_task_id: "266d168c-d93e-4387-9527-70aa0d5b2a6b" state_hub_task_id: "266d168c-d93e-4387-9527-70aa0d5b2a6b"
``` ```
Add fixture-driven contract tests mirroring the PDF suite; decide fixture Add fixture-driven contract tests mirroring the PDF suite; decide fixture
ownership per ADR-0002. ownership per ADR-0002.

View file

@ -4,14 +4,15 @@ type: workplan
title: "Document metadata enrichment beyond pass-through options" title: "Document metadata enrichment beyond pass-through options"
domain: infotech domain: infotech
repo: evidence-source repo: evidence-source
status: proposed status: finished
owner: codex owner: codex
topic_slug: citation_evidence_mvp topic_slug: citation_evidence_mvp
created: "2026-07-08" created: "2026-07-08"
updated: "2026-07-08" updated: "2026-07-09"
spec_refs: spec_refs:
- README.md - README.md
- src/pdf/ingest.ts - src/pdf/ingest.ts
- src/pdf/metadata.ts
state_hub_workstream_id: "831267a1-fc93-4872-a25b-73793d9c5df4" state_hub_workstream_id: "831267a1-fc93-4872-a25b-73793d9c5df4"
--- ---
@ -26,7 +27,7 @@ it into the `Document` record without breaking the pure-over-bytes contract.
```task ```task
id: ESRC-WP-0003-T01 id: ESRC-WP-0003-T01
status: todo status: done
priority: low priority: low
state_hub_task_id: "45eeef15-41a9-4abd-9436-ed77e2c6b3a4" state_hub_task_id: "45eeef15-41a9-4abd-9436-ed77e2c6b3a4"
``` ```
@ -35,7 +36,7 @@ precedence vs. caller-supplied options (caller wins).
```task ```task
id: ESRC-WP-0003-T02 id: ESRC-WP-0003-T02
status: todo status: done
priority: low priority: low
state_hub_task_id: "41ca090f-c709-420a-b7c5-8332e4e82bd1" state_hub_task_id: "41ca090f-c709-420a-b7c5-8332e4e82bd1"
``` ```
@ -44,9 +45,9 @@ minimal-metadata PDFs still ingest cleanly.
```task ```task
id: ESRC-WP-0003-T03 id: ESRC-WP-0003-T03
status: todo status: done
priority: low priority: low
state_hub_task_id: "0004881b-6577-4a62-a2ff-cbb5c4706b1b" state_hub_task_id: "0004881b-6577-4a62-a2ff-cbb5c4706b1b"
``` ```
Contract tests over the fixture corpus asserting extracted vs. overridden Contract tests over the fixture corpus asserting extracted vs. overridden
metadata precedence. metadata precedence.

View file

@ -4,14 +4,15 @@ type: workplan
title: "Local citation recovery and external source discovery hooks" title: "Local citation recovery and external source discovery hooks"
domain: infotech domain: infotech
repo: evidence-source repo: evidence-source
status: proposed status: finished
owner: codex owner: codex
topic_slug: citation_evidence_mvp topic_slug: citation_evidence_mvp
created: "2026-07-08" created: "2026-07-08"
updated: "2026-07-08" updated: "2026-07-09"
spec_refs: spec_refs:
- INTENT.md - INTENT.md
- README.md - README.md
- docs/ADR-0004-citation-recovery-boundary.md
state_hub_workstream_id: "2ce8b799-d67c-4dbf-aac5-0ff1f0363d62" state_hub_workstream_id: "2ce8b799-d67c-4dbf-aac5-0ff1f0363d62"
--- ---
@ -22,18 +23,17 @@ selector goes stale and discovering the external source a document was drawn
from. This is the largest deferred strand and depends on the anchoring layer's from. This is the largest deferred strand and depends on the anchoring layer's
selector model. selector model.
## Open questions ## Decisions (ADR-0004)
- What is the recovery contract — does this repo *detect* stale selectors, or - `evidence-source` owns re-ingest/re-fingerprint, local quote search, and a
only provide the re-ingest/re-fingerprint primitives the anchor layer calls? pluggable discovery hook registry (no network in core).
- Where does external source discovery (URL/DOI resolution) belong relative to - Stale selector detection and resolution order stay in `evidence-anchor`.
`citation-engine` and the umbrella?
## Sketch ## Sketch
```task ```task
id: ESRC-WP-0004-T01 id: ESRC-WP-0004-T01
status: todo status: done
priority: low priority: low
state_hub_task_id: "e7ff29b9-f04f-44f5-a23a-b028fc20a80b" state_hub_task_id: "e7ff29b9-f04f-44f5-a23a-b028fc20a80b"
``` ```
@ -42,9 +42,9 @@ Design the recovery boundary with the anchor layer: define what
```task ```task
id: ESRC-WP-0004-T02 id: ESRC-WP-0004-T02
status: todo status: done
priority: low priority: low
state_hub_task_id: "88a512d1-70e0-4bfa-b6b0-c1e4e31abb97" state_hub_task_id: "88a512d1-70e0-4bfa-b6b0-c1e4e31abb97"
``` ```
Prototype re-ingest/re-fingerprint recovery primitives and a discovery hook Prototype re-ingest/re-fingerprint recovery primitives and a discovery hook
interface (no network coupling baked into the core). interface (no network coupling baked into the core).