1691 lines
34 KiB
Markdown
Executable file
1691 lines
34 KiB
Markdown
Executable file
# SearchCapabilitiesResearchProgram.md
|
||
|
||
**Project:** DocStorePG
|
||
**Research Track:** PostgreSQL-native Search Capabilities
|
||
**Working label:** SearchPG capability family
|
||
**Status:** Proposed
|
||
**Date:** 2026-08-23
|
||
|
||
## 1. Purpose
|
||
|
||
DocStorePG started from the question of how far PostgreSQL can be extended toward a document-database role normally associated with MongoDB while retaining PostgreSQL as the underlying transactional substrate.
|
||
|
||
This research track extends that question:
|
||
|
||
> **How far can PostgreSQL also absorb the search responsibilities commonly delegated to Elasticsearch or OpenSearch, such that document storage, transactional access, structured filtering, lexical search, semantic retrieval, and tenant-aware application search can operate over one coherent data substrate?**
|
||
|
||
The goal is not to prove that PostgreSQL should replace Elasticsearch everywhere. The goal is to establish a measured capability envelope, identify architectures that minimize unnecessary data-system proliferation, and determine where a dedicated distributed search engine remains justified.
|
||
|
||
The research should produce both practical components for DocStorePG and reusable evidence for architectural decisions.
|
||
|
||
---
|
||
|
||
## 2. Problem Statement
|
||
|
||
A common modern application architecture uses several specialized data systems:
|
||
|
||
```text
|
||
Application
|
||
|
|
||
+--> PostgreSQL / MongoDB -- system of record
|
||
|
|
||
+--> Elasticsearch -- search projection
|
||
^
|
||
|
|
||
CDC / ETL / queues
|
||
```
|
||
|
||
This creates a second representation of application reality.
|
||
|
||
That duplication introduces:
|
||
|
||
- synchronization delay;
|
||
- failed or incomplete index updates;
|
||
- schema and mapping duplication;
|
||
- reindexing workflows;
|
||
- CDC, queue, or ETL infrastructure;
|
||
- duplicated authorization and tenant filters;
|
||
- reconciliation requirements;
|
||
- separate backup and recovery procedures;
|
||
- separate observability and operational expertise;
|
||
- additional infrastructure cost;
|
||
- two query languages and often two application data-access paths;
|
||
- uncertainty about which representation is authoritative at a particular instant.
|
||
|
||
For large-scale distributed search, these costs can be justified. For many application-search workloads, they may be accidental complexity.
|
||
|
||
PostgreSQL already provides a substantial subset of search primitives through:
|
||
|
||
- generalized inverted indexes (GIN);
|
||
- full-text search using `tsvector` / `tsquery`;
|
||
- `jsonb`;
|
||
- trigram similarity through `pg_trgm`;
|
||
- relational and structured filtering;
|
||
- Row Level Security;
|
||
- PostGIS for spatial search;
|
||
- extension mechanisms for alternative index access methods.
|
||
|
||
The current extension ecosystem adds substantially more:
|
||
|
||
- `pg_search` / ParadeDB for BM25, advanced lexical search, faceting, highlighting and search-oriented execution;
|
||
- `pgvector` for exact and approximate vector similarity search;
|
||
- PGroonga as an alternative full-text/search implementation with strong multilingual and JSONB capabilities;
|
||
- ZomboDB as a useful comparison architecture in which Elasticsearch is managed as a PostgreSQL index.
|
||
|
||
The research problem is therefore no longer simply:
|
||
|
||
> “Can PostgreSQL do full-text search?”
|
||
|
||
It is:
|
||
|
||
> **Can PostgreSQL provide a sufficiently complete, performant, operable and secure application-search substrate to eliminate a separate Elasticsearch/OpenSearch tier for a defined class of workloads?**
|
||
|
||
---
|
||
|
||
## 3. Research Thesis
|
||
|
||
The working thesis is:
|
||
|
||
> **A PostgreSQL-centered architecture can cover most document-oriented application-search workloads when lexical search, structured filtering, vector retrieval and authorization operate over the same transactional objects. A separate distributed search engine should be introduced only when measured scale, ingestion, analytical, or distribution requirements exceed that capability envelope.**
|
||
|
||
A stronger version of the thesis is relevant to DocStorePG:
|
||
|
||
> **PostgreSQL may serve as a universal application-information substrate whose relational, document, search, vector and spatial capabilities are exposed through specialized access surfaces rather than implemented as separate authoritative databases.**
|
||
|
||
This yields the conceptual target:
|
||
|
||
```text
|
||
PostgreSQL
|
||
|
|
||
+------------------+------------------+
|
||
| | |
|
||
SQL API Document API Search API
|
||
| | |
|
||
tables JSONB Query IR
|
||
| | |
|
||
+------------------+------------------+
|
||
|
|
||
Common transactional objects
|
||
|
|
||
+------------------+------------------+
|
||
| | |
|
||
B-tree/GIN lexical index vector index
|
||
BM25 / FTS HNSW / IVF
|
||
```
|
||
|
||
SearchPG is a working label for this capability family inside DocStorePG. It does not initially imply a separate repository or product.
|
||
|
||
---
|
||
|
||
## 4. Research Goals
|
||
|
||
The research program shall answer six principal questions.
|
||
|
||
### G1 — Capability
|
||
|
||
Determine which Elasticsearch/OpenSearch search capabilities can be provided by PostgreSQL itself or by mature PostgreSQL extensions.
|
||
|
||
### G2 — Quality
|
||
|
||
Determine whether result relevance and search behavior are competitive with Elasticsearch/OpenSearch for representative application workloads.
|
||
|
||
### G3 — Performance and Scale
|
||
|
||
Determine the performance envelope for:
|
||
|
||
- ingestion;
|
||
- updates;
|
||
- top-K search;
|
||
- filtered search;
|
||
- faceting;
|
||
- fuzzy and prefix search;
|
||
- hybrid lexical/vector search;
|
||
- multi-tenant workloads;
|
||
- index construction and rebuilding.
|
||
|
||
### G4 — Architecture and Operations
|
||
|
||
Measure whether consolidating search into PostgreSQL materially reduces operational complexity and total cost without creating unacceptable OLTP interference or recovery risk.
|
||
|
||
### G5 — Compatibility
|
||
|
||
Determine whether a useful Elasticsearch/OpenSearch-compatible API subset can be implemented over PostgreSQL without attempting impractical full emulation.
|
||
|
||
### G6 — DocStorePG Integration
|
||
|
||
Determine how search becomes a first-class capability of the DocStorePG document model, API, multitenancy model and deployment architecture.
|
||
|
||
---
|
||
|
||
## 5. Non-Goals
|
||
|
||
The initial program will not attempt to:
|
||
|
||
- prove that PostgreSQL replaces Elasticsearch for every workload;
|
||
- reproduce the complete Elasticsearch REST API;
|
||
- reproduce every Lucene analyzer or plugin;
|
||
- optimize initially for petabyte-scale observability;
|
||
- replace SIEM platforms;
|
||
- reproduce cross-region Elasticsearch cluster behavior;
|
||
- implement a new search engine from scratch before existing PostgreSQL extensions have been evaluated;
|
||
- conflate document storage compatibility with search API compatibility.
|
||
|
||
The initial target is **application search**, not unrestricted distributed log analytics.
|
||
|
||
---
|
||
|
||
## 6. Search Capability Canon
|
||
|
||
The project should define a canonical capability model independent of any implementation.
|
||
|
||
### 6.1 Structured Search
|
||
|
||
- equality filters;
|
||
- ranges;
|
||
- boolean combinations;
|
||
- nested/document-field predicates;
|
||
- arrays;
|
||
- sorting;
|
||
- pagination;
|
||
- aggregation.
|
||
|
||
### 6.2 Lexical Search
|
||
|
||
- tokenization;
|
||
- stemming;
|
||
- stop words;
|
||
- language analyzers;
|
||
- field weighting;
|
||
- BM25-style ranking;
|
||
- phrase search;
|
||
- proximity search;
|
||
- fuzzy matching;
|
||
- prefix search;
|
||
- typo tolerance;
|
||
- synonyms;
|
||
- highlighting.
|
||
|
||
### 6.3 Faceting and Analytics
|
||
|
||
- term facets;
|
||
- range facets;
|
||
- counts;
|
||
- grouped aggregations;
|
||
- filtered aggregations;
|
||
- top-K plus facet evaluation.
|
||
|
||
### 6.4 Semantic Search
|
||
|
||
- vector storage;
|
||
- exact nearest-neighbor search;
|
||
- approximate nearest-neighbor search;
|
||
- cosine / inner-product / L2 distance;
|
||
- filtered vector search;
|
||
- vector index lifecycle.
|
||
|
||
### 6.5 Hybrid Search
|
||
|
||
- lexical + semantic retrieval;
|
||
- rank fusion;
|
||
- reranking;
|
||
- structured business signals;
|
||
- recency weighting;
|
||
- domain-specific ranking signals.
|
||
|
||
### 6.6 Document Search
|
||
|
||
- dynamic JSON fields;
|
||
- nested structures;
|
||
- field-specific indexing;
|
||
- indexing all textual JSON values;
|
||
- schema evolution;
|
||
- partial indexes;
|
||
- index configuration per document type.
|
||
|
||
### 6.7 Security and Multitenancy
|
||
|
||
- tenant filtering;
|
||
- Row Level Security interaction;
|
||
- tenant-specific indexes where appropriate;
|
||
- shared-index isolation;
|
||
- authorization-aware search;
|
||
- noisy-neighbor behavior;
|
||
- tenant-aware ranking and facets.
|
||
|
||
### 6.8 Operational Search Capabilities
|
||
|
||
- online index creation;
|
||
- rebuild/reindex;
|
||
- backup;
|
||
- restore;
|
||
- PITR;
|
||
- physical replication;
|
||
- logical replication;
|
||
- failover;
|
||
- rolling upgrade;
|
||
- extension upgrade;
|
||
- Kubernetes deployment;
|
||
- monitoring;
|
||
- resource governance.
|
||
|
||
### 6.9 Compatibility Surface
|
||
|
||
- search request;
|
||
- query clauses;
|
||
- filters;
|
||
- sort;
|
||
- pagination;
|
||
- highlighting;
|
||
- facets/aggregations;
|
||
- mappings/index definitions;
|
||
- bulk indexing semantics;
|
||
- selected Elasticsearch/OpenSearch response structures.
|
||
|
||
---
|
||
|
||
## 7. Candidate Solution Families
|
||
|
||
The program should evaluate implementations as competing or composable solution families.
|
||
|
||
### A. Native PostgreSQL
|
||
|
||
Components:
|
||
|
||
```text
|
||
PostgreSQL
|
||
+ tsvector / tsquery
|
||
+ GIN
|
||
+ pg_trgm
|
||
+ JSONB
|
||
+ ordinary SQL indexes
|
||
```
|
||
|
||
Purpose:
|
||
|
||
- establish the minimum-dependency baseline;
|
||
- determine how much search functionality requires no specialized engine at all;
|
||
- identify the point at which native ranking or performance becomes inadequate.
|
||
|
||
This is the architectural simplicity baseline.
|
||
|
||
---
|
||
|
||
### B. PostgreSQL + `pg_search` / ParadeDB
|
||
|
||
Components:
|
||
|
||
```text
|
||
PostgreSQL
|
||
+ pg_search
|
||
+ Tantivy-backed BM25 indexes
|
||
+ search-oriented aggregations
|
||
```
|
||
|
||
Investigate:
|
||
|
||
- BM25 ranking;
|
||
- fuzzy search;
|
||
- phrase/proximity behavior;
|
||
- highlighting;
|
||
- facets;
|
||
- top-K performance;
|
||
- structured filters;
|
||
- joins;
|
||
- transaction semantics;
|
||
- index build and update behavior;
|
||
- resource consumption;
|
||
- Community vs Enterprise deployment characteristics;
|
||
- licensing implications;
|
||
- replication and recovery.
|
||
|
||
This is currently the primary candidate for an Elasticsearch-class lexical search capability inside PostgreSQL.
|
||
|
||
---
|
||
|
||
### C. PostgreSQL + PGroonga
|
||
|
||
Investigate as both an alternative and a specialized capability.
|
||
|
||
Particular areas of interest:
|
||
|
||
- multilingual search;
|
||
- CJK language behavior;
|
||
- JSONB full-text search;
|
||
- similar-search functionality;
|
||
- Row Level Security interaction;
|
||
- performance;
|
||
- replication and recovery;
|
||
- permissive licensing.
|
||
|
||
PGroonga should prevent the research program from prematurely equating “PostgreSQL search” with a single extension.
|
||
|
||
---
|
||
|
||
### D. PostgreSQL + `pgvector`
|
||
|
||
Components:
|
||
|
||
```text
|
||
PostgreSQL
|
||
+ pgvector
|
||
+ HNSW
|
||
+ IVFFlat
|
||
```
|
||
|
||
Investigate:
|
||
|
||
- exact retrieval;
|
||
- HNSW;
|
||
- IVFFlat;
|
||
- filtered vector search;
|
||
- recall/latency trade-offs;
|
||
- index size;
|
||
- write behavior;
|
||
- tenant filters;
|
||
- partitioning;
|
||
- hybrid integration.
|
||
|
||
This represents the semantic-search layer.
|
||
|
||
---
|
||
|
||
### E. Combined Search PostgreSQL
|
||
|
||
The most important practical candidate is:
|
||
|
||
```text
|
||
PostgreSQL
|
||
+ JSONB
|
||
+ pg_search
|
||
+ pgvector
|
||
+ pg_trgm where useful
|
||
+ PostGIS where useful
|
||
```
|
||
|
||
This configuration should be treated as a single integrated search substrate rather than a bag of extensions.
|
||
|
||
Research must determine:
|
||
|
||
- planner interaction;
|
||
- extension interaction;
|
||
- index selection;
|
||
- write amplification;
|
||
- memory pressure;
|
||
- vacuum behavior;
|
||
- transaction semantics;
|
||
- operational compatibility.
|
||
|
||
---
|
||
|
||
### F. PostgreSQL Search Replica
|
||
|
||
Search need not execute on the transactional primary.
|
||
|
||
Candidate:
|
||
|
||
```text
|
||
DocStorePG Primary
|
||
|
|
||
| logical replication
|
||
v
|
||
Search PostgreSQL
|
||
+ pg_search
|
||
+ pgvector
|
||
```
|
||
|
||
This topology attempts to obtain:
|
||
|
||
- workload isolation;
|
||
- independent search scaling;
|
||
- PostgreSQL-compatible data model;
|
||
- reduced coupling between search and OLTP resource consumption.
|
||
|
||
Measure the consistency/lag trade-off explicitly.
|
||
|
||
---
|
||
|
||
### G. Distributed PostgreSQL Search
|
||
|
||
Candidate:
|
||
|
||
```text
|
||
PostgreSQL
|
||
+ Citus or comparable sharding layer
|
||
+ pg_search
|
||
```
|
||
|
||
Research questions:
|
||
|
||
- tenant-key distribution;
|
||
- distributed top-K;
|
||
- distributed facets;
|
||
- cross-shard joins;
|
||
- query planning;
|
||
- rebalancing;
|
||
- replica placement;
|
||
- operational complexity relative to Elasticsearch/OpenSearch.
|
||
|
||
This is important for finding the point at which “Postgres for search” remains viable after a single node is exhausted.
|
||
|
||
---
|
||
|
||
### H. ZomboDB Control Architecture
|
||
|
||
ZomboDB deliberately retains Elasticsearch as the physical search engine while exposing it through PostgreSQL index semantics.
|
||
|
||
It should be included as a control because it answers a different question:
|
||
|
||
> Can PostgreSQL own search consistency and query integration even when Elasticsearch remains the index engine?
|
||
|
||
This helps separate three benefits:
|
||
|
||
1. one application query surface;
|
||
2. transactional/index consistency;
|
||
3. elimination of Elasticsearch itself.
|
||
|
||
---
|
||
|
||
### I. Elasticsearch and OpenSearch Baselines
|
||
|
||
Both should be benchmark controls.
|
||
|
||
The research is invalid if PostgreSQL configurations are measured only against one another.
|
||
|
||
Reference configurations should include:
|
||
|
||
```text
|
||
Application -> Elasticsearch
|
||
Application -> OpenSearch
|
||
```
|
||
|
||
and, where relevant:
|
||
|
||
```text
|
||
PostgreSQL -> CDC -> Elasticsearch/OpenSearch
|
||
```
|
||
|
||
---
|
||
|
||
## 8. Proposed Search API Architecture
|
||
|
||
Full Elasticsearch compatibility should not be the first implementation target.
|
||
|
||
Instead, introduce a canonical **Search Query Intermediate Representation**.
|
||
|
||
```text
|
||
Elasticsearch-like request
|
||
|
|
||
Mongo/DocStore search request
|
||
|
|
||
Native DocStorePG Search API
|
||
|
|
||
v
|
||
Search Query IR
|
||
|
|
||
+-----+------+---------+----------+
|
||
| | | |
|
||
filters lexical vector facets
|
||
| | | |
|
||
+------------+---------+----------+
|
||
|
|
||
PostgreSQL planner
|
||
```
|
||
|
||
The Search Query IR should model:
|
||
|
||
```text
|
||
SearchRequest
|
||
query
|
||
lexical
|
||
semantic
|
||
filters
|
||
ranking
|
||
sort
|
||
pagination
|
||
facets
|
||
highlighting
|
||
projection
|
||
```
|
||
|
||
Backends can then compile the same request into:
|
||
|
||
- native PostgreSQL FTS;
|
||
- `pg_search`;
|
||
- PGroonga;
|
||
- `pgvector`;
|
||
- hybrid SQL;
|
||
- optionally Elasticsearch/OpenSearch for comparison.
|
||
|
||
This provides two strategic benefits.
|
||
|
||
First, DocStorePG does not become coupled to one extension.
|
||
|
||
Second, API compatibility can be implemented as adapters over the IR.
|
||
|
||
---
|
||
|
||
## 9. Elasticsearch/OpenSearch Compatibility Strategy
|
||
|
||
Compatibility should be **profile-based**, not claimed as universal.
|
||
|
||
Suggested profiles:
|
||
|
||
### ES-SEARCH-0 — Basic Search
|
||
|
||
- `_search`;
|
||
- `match`;
|
||
- `term`;
|
||
- `terms`;
|
||
- `range`;
|
||
- `bool`;
|
||
- sorting;
|
||
- pagination.
|
||
|
||
### ES-SEARCH-1 — Application Search
|
||
|
||
Adds:
|
||
|
||
- `multi_match`;
|
||
- phrase search;
|
||
- fuzzy;
|
||
- highlighting;
|
||
- selected aggregations;
|
||
- field boosting.
|
||
|
||
### ES-SEARCH-2 — Hybrid Search
|
||
|
||
Adds:
|
||
|
||
- vector query primitives;
|
||
- lexical/vector rank fusion;
|
||
- selected reranking.
|
||
|
||
### ES-SEARCH-X — Unsupported / Native Required
|
||
|
||
Explicitly mark unsupported functionality rather than silently approximate it.
|
||
|
||
Compatibility testing should be done against pinned Elasticsearch/OpenSearch versions because their APIs have diverged over time.
|
||
|
||
The project should prefer **useful compatibility** over cosmetic wire compatibility.
|
||
|
||
---
|
||
|
||
## 10. Research Hypotheses
|
||
|
||
### H1 — Native PostgreSQL covers a large basic-search envelope
|
||
|
||
For many application workloads, `tsvector`, GIN, `pg_trgm`, JSONB and SQL filters will be sufficient.
|
||
|
||
### H2 — BM25 is the main threshold between database FTS and search-engine-quality lexical retrieval
|
||
|
||
`pg_search` or another BM25-capable index should materially improve ranking quality and top-K performance over native PostgreSQL ranking.
|
||
|
||
### H3 — Search over authoritative rows is an architectural advantage
|
||
|
||
Removing an asynchronous external search projection will eliminate measurable classes of consistency and authorization defects.
|
||
|
||
### H4 — Hybrid search becomes simpler inside one database
|
||
|
||
Lexical relevance, vector similarity, structured filters and business ranking signals should be easier to compose when they operate over the same row identifiers.
|
||
|
||
### H5 — Tenant-aware filtering benefits disproportionately from PostgreSQL integration
|
||
|
||
Shared-table multitenancy, RLS, joins and application authorization should require less duplicated security logic than an external search index.
|
||
|
||
### H6 — Search workload isolation remains necessary above moderate load
|
||
|
||
A search replica or dedicated Search PostgreSQL node will likely become preferable before PostgreSQL's raw search capability is exhausted.
|
||
|
||
### H7 — A PostgreSQL search topology can scale farther than the common “single database” mental model suggests
|
||
|
||
Sharding and replication may extend the useful envelope considerably, but operational complexity may approach that of Elasticsearch.
|
||
|
||
### H8 — A small Elasticsearch-compatible API subset provides most migration value
|
||
|
||
A carefully selected 20–30% of API primitives may cover the majority of application-search integration use cases.
|
||
|
||
### H9 — Full Elasticsearch compatibility is economically unattractive
|
||
|
||
The long tail of mappings, analyzers, cluster APIs, plugins and Query DSL behavior is unlikely to justify reproduction.
|
||
|
||
### H10 — Elasticsearch/OpenSearch will retain a clear advantage for some workloads
|
||
|
||
Likely examples:
|
||
|
||
- extremely high-volume event ingestion;
|
||
- very large append-only observability datasets;
|
||
- massive horizontal fan-out;
|
||
- mature Lucene-specific custom analysis;
|
||
- large distributed aggregation workloads.
|
||
|
||
The research should locate this boundary rather than argue it away.
|
||
|
||
---
|
||
|
||
## 11. Workstreams
|
||
|
||
## WS0 — Capability Inventory and Baseline
|
||
|
||
Deliver:
|
||
|
||
- `SearchCapabilityCanon.md`;
|
||
- implementation capability matrix;
|
||
- version and licensing matrix;
|
||
- known deployment constraints.
|
||
|
||
Candidates:
|
||
|
||
- PostgreSQL native FTS;
|
||
- `pg_trgm`;
|
||
- `pg_search`;
|
||
- PGroonga;
|
||
- `pgvector`;
|
||
- ZomboDB;
|
||
- Elasticsearch;
|
||
- OpenSearch.
|
||
|
||
---
|
||
|
||
## WS1 — Lexical Search Quality
|
||
|
||
Evaluate:
|
||
|
||
- tokenization;
|
||
- stemming;
|
||
- phrase search;
|
||
- fuzzy search;
|
||
- prefix behavior;
|
||
- field weighting;
|
||
- BM25;
|
||
- language support;
|
||
- synonyms;
|
||
- highlighting.
|
||
|
||
Metrics:
|
||
|
||
- NDCG@10;
|
||
- MRR;
|
||
- Recall@K;
|
||
- judged relevance;
|
||
- latency.
|
||
|
||
Use at least one public relevance dataset plus one DocStorePG-style document corpus.
|
||
|
||
---
|
||
|
||
## WS2 — Structured and JSON Document Search
|
||
|
||
Evaluate:
|
||
|
||
- dynamic fields;
|
||
- JSON path filters;
|
||
- nested structures;
|
||
- arrays;
|
||
- numeric/date ranges;
|
||
- combinations of full-text and JSON predicates;
|
||
- schema evolution;
|
||
- index configuration changes.
|
||
|
||
This workstream directly connects search to the DocStorePG document model.
|
||
|
||
---
|
||
|
||
## WS3 — Facets and Aggregations
|
||
|
||
Evaluate:
|
||
|
||
- term facets;
|
||
- range facets;
|
||
- top-K + facets;
|
||
- high-cardinality facets;
|
||
- filtered facets;
|
||
- multi-facet requests.
|
||
|
||
Measure CPU, memory and latency separately from ordinary top-K search.
|
||
|
||
---
|
||
|
||
## WS4 — Semantic and Hybrid Retrieval
|
||
|
||
Compare:
|
||
|
||
```text
|
||
BM25
|
||
vector-only
|
||
BM25 + vector
|
||
BM25 + vector + structured business signals
|
||
```
|
||
|
||
Fusion methods should include at least:
|
||
|
||
- Reciprocal Rank Fusion;
|
||
- weighted normalized scores;
|
||
- reranking of lexical candidates;
|
||
- reranking of vector candidates.
|
||
|
||
Measure:
|
||
|
||
- relevance;
|
||
- vector recall;
|
||
- latency;
|
||
- filter interaction;
|
||
- index memory;
|
||
- ingestion/update cost.
|
||
|
||
---
|
||
|
||
## WS5 — Multitenancy and Security
|
||
|
||
Test at least three tenant models.
|
||
|
||
### Model A — Shared Table
|
||
|
||
```text
|
||
documents(
|
||
tenant_id,
|
||
...
|
||
)
|
||
```
|
||
|
||
### Model B — Partitioned Shared Database
|
||
|
||
Partitions by tenant or tenant group.
|
||
|
||
### Model C — Database / Cluster Isolation
|
||
|
||
Dedicated PostgreSQL databases or clusters for high-isolation tenants.
|
||
|
||
For each model measure:
|
||
|
||
- tenant filter cost;
|
||
- RLS correctness;
|
||
- search ranking correctness;
|
||
- facet isolation;
|
||
- vector filtering;
|
||
- noisy-neighbor behavior;
|
||
- cross-tenant leakage attempts;
|
||
- operational cost.
|
||
|
||
Search authorization tests should be adversarial, not merely functional.
|
||
|
||
---
|
||
|
||
## WS6 — API Compatibility
|
||
|
||
Implement a thin prototype:
|
||
|
||
```text
|
||
HTTP Search Gateway
|
||
|
|
||
Elasticsearch/OpenSearch subset
|
||
|
|
||
Search Query IR
|
||
|
|
||
PostgreSQL
|
||
```
|
||
|
||
Test existing client libraries where practical.
|
||
|
||
Record each feature as:
|
||
|
||
```text
|
||
compatible
|
||
compatible-with-difference
|
||
translated
|
||
unsupported
|
||
```
|
||
|
||
Do not hide semantic differences.
|
||
|
||
---
|
||
|
||
## WS7 — Consistency and Lifecycle
|
||
|
||
Test:
|
||
|
||
- insert visibility;
|
||
- update visibility;
|
||
- delete visibility;
|
||
- transaction rollback;
|
||
- concurrent updates;
|
||
- long-running transactions;
|
||
- index rebuild;
|
||
- schema migration;
|
||
- extension upgrade;
|
||
- backup/restore;
|
||
- failover;
|
||
- logical replication lag;
|
||
- crash recovery.
|
||
|
||
This should explicitly compare synchronous in-database indexing with CDC-backed Elasticsearch/OpenSearch.
|
||
|
||
---
|
||
|
||
## WS8 — Performance and Scale
|
||
|
||
Test dataset sizes such as:
|
||
|
||
```text
|
||
100 K
|
||
1 M
|
||
10 M
|
||
100 M
|
||
```
|
||
|
||
and larger where the environment makes the experiment meaningful.
|
||
|
||
Measure:
|
||
|
||
### Writes
|
||
|
||
- documents/sec;
|
||
- updates/sec;
|
||
- deletes/sec;
|
||
- bulk load;
|
||
- WAL volume;
|
||
- write amplification.
|
||
|
||
### Reads
|
||
|
||
- p50;
|
||
- p95;
|
||
- p99;
|
||
- QPS;
|
||
- concurrent users.
|
||
|
||
### Search Workloads
|
||
|
||
- top-K lexical;
|
||
- phrase;
|
||
- fuzzy;
|
||
- filter + search;
|
||
- faceted;
|
||
- vector;
|
||
- hybrid;
|
||
- tenant-filtered;
|
||
- deep pagination where supported.
|
||
|
||
### Resources
|
||
|
||
- CPU;
|
||
- RAM;
|
||
- disk;
|
||
- index size;
|
||
- cache hit behavior;
|
||
- temporary I/O.
|
||
|
||
---
|
||
|
||
## WS9 — Deployment and Operations
|
||
|
||
Evaluate:
|
||
|
||
```text
|
||
Single PostgreSQL
|
||
Primary + search replica
|
||
HA PostgreSQL cluster
|
||
CloudNativePG deployment
|
||
Sharded PostgreSQL
|
||
Elasticsearch/OpenSearch cluster
|
||
```
|
||
|
||
Measure:
|
||
|
||
- installation;
|
||
- upgrade;
|
||
- backup;
|
||
- restore;
|
||
- failover;
|
||
- rebuild;
|
||
- monitoring;
|
||
- day-2 operations;
|
||
- operator support;
|
||
- extension packaging;
|
||
- Kubernetes compatibility.
|
||
|
||
The benchmark should capture operator effort, not merely query latency.
|
||
|
||
---
|
||
|
||
## WS10 — Economics and Complexity
|
||
|
||
Create a normalized cost model.
|
||
|
||
### Infrastructure
|
||
|
||
- compute;
|
||
- memory;
|
||
- storage;
|
||
- network;
|
||
- backup;
|
||
- replicas.
|
||
|
||
### Operational Complexity
|
||
|
||
Score:
|
||
|
||
- systems operated;
|
||
- data copies;
|
||
- synchronization mechanisms;
|
||
- schemas/mappings;
|
||
- recovery procedures;
|
||
- monitoring surfaces;
|
||
- specialist knowledge;
|
||
- failure modes.
|
||
|
||
### Development Complexity
|
||
|
||
Measure:
|
||
|
||
- LOC for integration;
|
||
- number of APIs;
|
||
- authorization duplication;
|
||
- migration effort;
|
||
- test surface.
|
||
|
||
The result should produce both:
|
||
|
||
```text
|
||
technical capability envelope
|
||
```
|
||
|
||
and:
|
||
|
||
```text
|
||
economic/operational capability envelope
|
||
```
|
||
|
||
---
|
||
|
||
## 12. Benchmark Dataset Families
|
||
|
||
Use multiple dataset families rather than one synthetic benchmark.
|
||
|
||
### D1 — Relevance Corpus
|
||
|
||
Purpose:
|
||
|
||
- lexical ranking;
|
||
- semantic ranking;
|
||
- hybrid search.
|
||
|
||
Use a public IR/relevance dataset such as an appropriate BEIR/TREC-derived corpus.
|
||
|
||
### D2 — Product Catalog
|
||
|
||
Documents with:
|
||
|
||
- title;
|
||
- description;
|
||
- brand;
|
||
- categories;
|
||
- attributes;
|
||
- price;
|
||
- stock;
|
||
- ratings.
|
||
|
||
Purpose:
|
||
|
||
- faceting;
|
||
- filtering;
|
||
- fuzzy search;
|
||
- ranking;
|
||
- high-cardinality attributes.
|
||
|
||
### D3 — Document Management Corpus
|
||
|
||
Documents with:
|
||
|
||
- metadata;
|
||
- extracted text;
|
||
- participants;
|
||
- dates;
|
||
- document type;
|
||
- tags;
|
||
- tenant ACLs.
|
||
|
||
Purpose:
|
||
|
||
- realistic DocStorePG application search.
|
||
|
||
### D4 — Multi-Tenant SaaS Corpus
|
||
|
||
Generate tenants with a skewed distribution:
|
||
|
||
```text
|
||
many small
|
||
some medium
|
||
few very large
|
||
```
|
||
|
||
Purpose:
|
||
|
||
- tenant isolation;
|
||
- noisy-neighbor effects;
|
||
- partitioning;
|
||
- RLS;
|
||
- filtered vector search.
|
||
|
||
### D5 — Event/Log Corpus
|
||
|
||
Purpose:
|
||
|
||
- intentionally stress the likely boundary;
|
||
- compare ingestion and time-range search against Elasticsearch/OpenSearch.
|
||
|
||
This dataset is important because a good research program must contain workloads expected to falsify the PostgreSQL thesis.
|
||
|
||
---
|
||
|
||
## 13. Benchmark Harness
|
||
|
||
The benchmark harness should be reusable across backends.
|
||
|
||
Suggested structure:
|
||
|
||
```text
|
||
bench/search/
|
||
datasets/
|
||
workloads/
|
||
adapters/
|
||
postgres_native/
|
||
pg_search/
|
||
pgroonga/
|
||
pgvector/
|
||
postgres_hybrid/
|
||
zombodb/
|
||
elasticsearch/
|
||
opensearch/
|
||
runner/
|
||
metrics/
|
||
reports/
|
||
```
|
||
|
||
Each workload should produce a machine-readable result record containing:
|
||
|
||
```text
|
||
backend
|
||
version
|
||
dataset
|
||
dataset_size
|
||
query_class
|
||
concurrency
|
||
latency_p50
|
||
latency_p95
|
||
latency_p99
|
||
throughput
|
||
cpu
|
||
memory
|
||
io
|
||
index_size
|
||
relevance_metrics
|
||
consistency_metrics
|
||
```
|
||
|
||
Benchmark definitions should be immutable/versioned so results remain comparable over time.
|
||
|
||
---
|
||
|
||
## 14. Failure Injection
|
||
|
||
Search systems are often selected on benchmark speed and regretted during failure.
|
||
|
||
Inject:
|
||
|
||
- database restart;
|
||
- hard process kill;
|
||
- search-node loss;
|
||
- network interruption;
|
||
- replication interruption;
|
||
- disk pressure;
|
||
- index corruption where safely reproducible;
|
||
- stalled replication;
|
||
- partially completed bulk load;
|
||
- extension upgrade failure.
|
||
|
||
Measure:
|
||
|
||
```text
|
||
RPO
|
||
RTO
|
||
search availability
|
||
data correctness
|
||
repair procedure
|
||
operator actions
|
||
```
|
||
|
||
---
|
||
|
||
## 15. Decision Metrics
|
||
|
||
A configuration should not be declared successful merely because it is faster.
|
||
|
||
Evaluate across six dimensions.
|
||
|
||
### Q — Search Quality
|
||
|
||
- relevance;
|
||
- recall;
|
||
- query expressiveness.
|
||
|
||
### P — Performance
|
||
|
||
- latency;
|
||
- throughput;
|
||
- ingestion.
|
||
|
||
### C — Consistency
|
||
|
||
- transactional visibility;
|
||
- lag;
|
||
- correctness.
|
||
|
||
### O — Operability
|
||
|
||
- deployment;
|
||
- recovery;
|
||
- upgrades;
|
||
- monitoring.
|
||
|
||
### S — Security
|
||
|
||
- tenant isolation;
|
||
- authorization integrity.
|
||
|
||
### E — Economics
|
||
|
||
- infrastructure cost;
|
||
- engineering cost;
|
||
- operational cost.
|
||
|
||
The eventual decision model can be expressed as:
|
||
|
||
```text
|
||
fitness =
|
||
f(Q, P, C, O, S, E)
|
||
```
|
||
|
||
with workload-specific weights rather than one universal winner.
|
||
|
||
---
|
||
|
||
## 16. Milestones
|
||
|
||
## M0 — Research Foundation
|
||
|
||
Deliver:
|
||
|
||
- research program;
|
||
- capability canon;
|
||
- candidate matrix;
|
||
- benchmark harness skeleton;
|
||
- pinned component versions.
|
||
|
||
Exit criterion:
|
||
|
||
> Every later experiment can be described through the same capability and metric vocabulary.
|
||
|
||
---
|
||
|
||
## M1 — Native PostgreSQL Baseline
|
||
|
||
Implement and benchmark:
|
||
|
||
- `tsvector`;
|
||
- GIN;
|
||
- JSONB;
|
||
- `pg_trgm`.
|
||
|
||
Exit criterion:
|
||
|
||
> Native PostgreSQL search envelope is documented with measured limits.
|
||
|
||
---
|
||
|
||
## M2 — Search Extension Shootout
|
||
|
||
Add:
|
||
|
||
- `pg_search`;
|
||
- PGroonga;
|
||
- ZomboDB control;
|
||
- Elasticsearch;
|
||
- OpenSearch.
|
||
|
||
Exit criterion:
|
||
|
||
> Lexical search quality, functionality, performance and operational differences are quantified.
|
||
|
||
---
|
||
|
||
## M3 — Semantic and Hybrid Search
|
||
|
||
Add:
|
||
|
||
- `pgvector`;
|
||
- HNSW;
|
||
- IVFFlat;
|
||
- lexical/vector fusion.
|
||
|
||
Exit criterion:
|
||
|
||
> A recommended PostgreSQL-native hybrid search architecture is identified or rejected with evidence.
|
||
|
||
---
|
||
|
||
## M4 — Multitenancy and Security
|
||
|
||
Add:
|
||
|
||
- RLS;
|
||
- tenant filters;
|
||
- partitions;
|
||
- authorization joins;
|
||
- adversarial leakage tests.
|
||
|
||
Exit criterion:
|
||
|
||
> At least one secure shared-database search pattern is validated.
|
||
|
||
---
|
||
|
||
## M5 — Search API Prototype
|
||
|
||
Implement:
|
||
|
||
- Search Query IR;
|
||
- native DocStorePG search endpoint;
|
||
- ES-SEARCH-0 compatibility profile.
|
||
|
||
Exit criterion:
|
||
|
||
> A representative Elasticsearch-style application can query DocStorePG through the compatibility layer.
|
||
|
||
---
|
||
|
||
## M6 — Isolation and Scale
|
||
|
||
Evaluate:
|
||
|
||
- dedicated search PostgreSQL;
|
||
- logical replication;
|
||
- sharded PostgreSQL where justified;
|
||
- OLTP/search interference.
|
||
|
||
Exit criterion:
|
||
|
||
> The transition point from embedded search to dedicated search nodes is characterized.
|
||
|
||
---
|
||
|
||
## M7 — Failure, Recovery and Operations
|
||
|
||
Evaluate:
|
||
|
||
- backup;
|
||
- restore;
|
||
- crash;
|
||
- failover;
|
||
- reindex;
|
||
- upgrades;
|
||
- Kubernetes operations.
|
||
|
||
Exit criterion:
|
||
|
||
> Production-operability constraints and extension-specific risks are explicitly known.
|
||
|
||
---
|
||
|
||
## M8 — Architecture Decision
|
||
|
||
Produce:
|
||
|
||
- benchmark report;
|
||
- capability envelope;
|
||
- cost model;
|
||
- architecture patterns;
|
||
- adoption guide;
|
||
- boundary conditions.
|
||
|
||
Exit criterion:
|
||
|
||
> DocStorePG has an evidence-backed answer to when PostgreSQL should replace, complement, or retain Elasticsearch/OpenSearch.
|
||
|
||
---
|
||
|
||
## 17. Proposed Architecture Patterns to Validate
|
||
|
||
### Pattern P1 — Embedded Search
|
||
|
||
```text
|
||
Application
|
||
|
|
||
DocStorePG PostgreSQL
|
||
|
|
||
JSONB + native FTS
|
||
```
|
||
|
||
Target:
|
||
|
||
- small/medium applications;
|
||
- minimal operations.
|
||
|
||
---
|
||
|
||
### Pattern P2 — Enhanced Embedded Search
|
||
|
||
```text
|
||
Application
|
||
|
|
||
DocStorePG PostgreSQL
|
||
|
|
||
JSONB + pg_search + pgvector
|
||
```
|
||
|
||
Target:
|
||
|
||
- application search;
|
||
- knowledge systems;
|
||
- product search;
|
||
- document search.
|
||
|
||
---
|
||
|
||
### Pattern P3 — Search-Isolated PostgreSQL
|
||
|
||
```text
|
||
+--> OLTP reads
|
||
|
|
||
DocStorePG Primary
|
||
|
|
||
+--> logical replication --> Search PostgreSQL
|
||
pg_search
|
||
pgvector
|
||
```
|
||
|
||
Target:
|
||
|
||
- higher search load;
|
||
- workload isolation.
|
||
|
||
---
|
||
|
||
### Pattern P4 — Distributed PostgreSQL Search
|
||
|
||
```text
|
||
Application
|
||
|
|
||
Search Gateway
|
||
|
|
||
Distributed PostgreSQL
|
||
Citus + pg_search
|
||
```
|
||
|
||
Target:
|
||
|
||
- large tenant sets;
|
||
- large searchable corpora;
|
||
- horizontal scaling experiments.
|
||
|
||
---
|
||
|
||
### Pattern P5 — External Search Remains Correct
|
||
|
||
```text
|
||
DocStorePG
|
||
|
|
||
CDC / managed integration
|
||
|
|
||
Elasticsearch/OpenSearch
|
||
```
|
||
|
||
Target:
|
||
|
||
- workloads outside the PostgreSQL capability envelope.
|
||
|
||
This pattern must remain an acceptable research conclusion.
|
||
|
||
---
|
||
|
||
## 18. Integration into DocStorePG Scope
|
||
|
||
The DocStorePG project should expand from:
|
||
|
||
```text
|
||
MongoDB-like document capabilities
|
||
on PostgreSQL
|
||
```
|
||
|
||
to:
|
||
|
||
```text
|
||
Document-oriented application data capabilities
|
||
on PostgreSQL
|
||
including search and retrieval
|
||
```
|
||
|
||
The capability layers become:
|
||
|
||
```text
|
||
DocStorePG
|
||
|
|
||
+-- Document Model
|
||
| +-- JSONB
|
||
| +-- schema/index management
|
||
|
|
||
+-- Document API
|
||
| +-- Mongo-compatible primitives
|
||
|
|
||
+-- Search Model
|
||
| +-- structured
|
||
| +-- lexical
|
||
| +-- semantic
|
||
| +-- hybrid
|
||
| +-- facets
|
||
|
|
||
+-- Search API
|
||
| +-- native DocStorePG
|
||
| +-- optional Elasticsearch/OpenSearch profile
|
||
|
|
||
+-- Multitenancy
|
||
| +-- tenant routing
|
||
| +-- RLS
|
||
| +-- partitioning
|
||
|
|
||
+-- Operations
|
||
+-- CloudNativePG
|
||
+-- replication
|
||
+-- backup/recovery
|
||
+-- observability
|
||
```
|
||
|
||
This keeps the research coherent: document storage and search are not independent features. Search is an alternative projection and access mode over the same information model.
|
||
|
||
---
|
||
|
||
## 19. Proposed Repository Additions
|
||
|
||
```text
|
||
docs/
|
||
research/
|
||
search/
|
||
SearchCapabilitiesResearchProgram.md
|
||
SearchCapabilityCanon.md
|
||
SearchArchitectureOptions.md
|
||
SearchCompatibilityProfiles.md
|
||
SearchMultitenancyModel.md
|
||
SearchBenchmarkSpecification.md
|
||
SearchOperationalModel.md
|
||
SearchDecisionRecord.md
|
||
|
||
bench/
|
||
search/
|
||
datasets/
|
||
workloads/
|
||
adapters/
|
||
runner/
|
||
reports/
|
||
|
||
src/
|
||
search/
|
||
ir/
|
||
planner/
|
||
api/
|
||
adapters/
|
||
|
||
experiments/
|
||
search/
|
||
postgres-native/
|
||
pg-search/
|
||
pgroonga/
|
||
pgvector/
|
||
hybrid/
|
||
replication/
|
||
citus/
|
||
zombodb/
|
||
elasticsearch/
|
||
opensearch/
|
||
```
|
||
|
||
If the repository is still deliberately small, begin only with:
|
||
|
||
```text
|
||
docs/research/search/SearchCapabilitiesResearchProgram.md
|
||
bench/search/
|
||
experiments/search/
|
||
```
|
||
|
||
and allow the remaining structure to emerge from evidence.
|
||
|
||
---
|
||
|
||
## 20. Initial Priority
|
||
|
||
The first practical experiment should be intentionally narrow.
|
||
|
||
Build the same document corpus on:
|
||
|
||
```text
|
||
A. PostgreSQL native FTS
|
||
B. PostgreSQL + pg_search
|
||
C. Elasticsearch
|
||
```
|
||
|
||
Run:
|
||
|
||
```text
|
||
1. keyword top-K
|
||
2. phrase query
|
||
3. fuzzy query
|
||
4. structured filter + search
|
||
5. facet query
|
||
6. insert then immediate search
|
||
7. update then immediate search
|
||
8. delete then immediate search
|
||
```
|
||
|
||
Measure:
|
||
|
||
```text
|
||
relevance
|
||
p50/p95/p99 latency
|
||
index size
|
||
ingest speed
|
||
freshness/consistency
|
||
CPU/RAM
|
||
implementation complexity
|
||
```
|
||
|
||
Then add `pgvector` and hybrid search.
|
||
|
||
This sequence gives the project an early falsifiable answer before substantial compatibility work is attempted.
|
||
|
||
---
|
||
|
||
## 21. Success Criteria
|
||
|
||
The program succeeds even if PostgreSQL does **not** replace Elasticsearch in every tested scenario.
|
||
|
||
Success means producing a dependable decision map such as:
|
||
|
||
```text
|
||
Native PostgreSQL
|
||
sufficient here
|
||
|
|
||
v
|
||
PostgreSQL + search extensions
|
||
sufficient here
|
||
|
|
||
v
|
||
Dedicated PostgreSQL search nodes
|
||
sufficient here
|
||
|
|
||
v
|
||
Distributed PostgreSQL search
|
||
investigate here
|
||
|
|
||
v
|
||
Elasticsearch/OpenSearch
|
||
justified beyond here
|
||
```
|
||
|
||
The most valuable result is the location of the boundaries.
|
||
|
||
---
|
||
|
||
## 22. Expected Strategic Outcome
|
||
|
||
If the central hypotheses survive testing, DocStorePG becomes more than a MongoDB compatibility experiment.
|
||
|
||
It becomes an exploration of:
|
||
|
||
> **PostgreSQL as a unified application information substrate.**
|
||
|
||
That substrate would combine:
|
||
|
||
```text
|
||
relational transactions
|
||
+
|
||
document storage
|
||
+
|
||
structured retrieval
|
||
+
|
||
lexical search
|
||
+
|
||
semantic search
|
||
+
|
||
hybrid ranking
|
||
+
|
||
authorization
|
||
+
|
||
multitenancy
|
||
```
|
||
|
||
while allowing specialized external systems to remain available when scale or workload characteristics genuinely require them.
|
||
|
||
This is a stronger and more general research direction than “PostgreSQL as a MongoDB replacement,” because it tests whether several common application data-system boundaries are architectural necessities or historical artifacts.
|
||
|
||
---
|
||
|
||
## 23. Current Technology Baseline to Track
|
||
|
||
The following projects should be version-pinned in the research repository and periodically reevaluated:
|
||
|
||
- PostgreSQL 18 documentation: GIN and Full Text Search;
|
||
- ParadeDB / `pg_search`: BM25, full-text search, facets, hybrid-search integration;
|
||
- `pgvector`: HNSW and IVFFlat vector indexes;
|
||
- PGroonga: multilingual full-text and JSONB search;
|
||
- ZomboDB: PostgreSQL-managed Elasticsearch indexes;
|
||
- Elasticsearch;
|
||
- OpenSearch.
|
||
|
||
Important current research considerations include:
|
||
|
||
- `pg_search` is actively evolving;
|
||
- ParadeDB Community uses AGPL-3.0 while commercial licensing is available separately;
|
||
- some recovery/replication characteristics differ between ParadeDB Community and Enterprise;
|
||
- PGroonga uses the permissive PostgreSQL license;
|
||
- OpenSearch and Elasticsearch APIs have diverged enough that compatibility must be version/profile specific;
|
||
- `pgvector` approximate indexes trade recall for speed and therefore require explicit quality measurement.
|
||
|
||
These constraints belong in the benchmark and architecture decision process rather than being treated as implementation footnotes.
|