doc-store-pg/ResearchProgram.md

51 KiB
Executable file
Raw Blame History

ResearchProgram.md

Repository

doc-store-pg

Status

Initial research program / execution roadmap.

Purpose

This document turns the hypotheses and open questions defined in INTENT.md and ArchitectureBlueprint.md into an ordered, reproducible research program.

The goal is not to accumulate benchmarks.

The goal is to produce enough evidence to make progressively stronger architectural decisions about whether, where, and how a PostgreSQL / CloudNativePG platform can provide document-store capabilities that are competitive with MongoDB while retaining PostgreSQL-native strengths.

The research program should answer three classes of questions:

  1. Can it work?
    Compatibility, correctness, integration, and operational feasibility.

  2. How well does it work?
    Performance, efficiency, resilience, multitenancy, and cost.

  3. Where should it be used?
    Architectural fit, workload boundaries, tenant placement, and long-term product direction.

The program is deliberately staged so that expensive performance and scale research only begins after correctness and experimental reproducibility have been established.


1. Research North Star

The central research hypothesis is:

Document storage can be implemented as a composable PostgreSQL capability that is operationally native to CloudNativePG, sufficiently compatible with MongoDB for meaningful application reuse, and superior in selected hybrid or multitenant scenarios without requiring MongoDB's architecture to be reproduced wholesale.

The program must be capable of disproving this hypothesis.

A successful research outcome may therefore be any of the following:

  • PostgreSQL-based document storage is broadly competitive with MongoDB.
  • PostgreSQL-based document storage is attractive only for specific workload classes.
  • MongoDB remains clearly preferable for important document workloads.
  • Different tenant classes should use different storage architectures.
  • DocumentDB or FerretDB is useful only as a migration/compatibility surface.
  • Native PostgreSQL JSONB is sufficient for many workloads and the MongoDB layer adds little value.
  • A hybrid architecture is more compelling than a replacement architecture.
  • The operational simplicity of one PostgreSQL platform outweighs moderate performance disadvantages.
  • The complexity introduced by MongoDB compatibility is not justified.

The research program exists to distinguish among these outcomes.


2. Research Principles

2.1 Correctness Before Performance

A system that returns different results more quickly is not faster at the same task.

Compatibility and semantic behavior must therefore be understood before comparative performance results are treated as meaningful.

2.2 Reproducibility Before Optimization

Every important result must be reproducible from:

  • a repository commit
  • a machine-readable environment definition
  • pinned software versions
  • a workload definition
  • a dataset definition
  • a random seed where relevant
  • benchmark configuration
  • raw result references
  • analysis code

2.3 Compare Architectures, Not Brands

The primary objects of comparison are deployment and access profiles.

Stable profile identifiers from ArchitectureBlueprint.md are used:

M0  Native MongoDB reference

P0  PostgreSQL + native JSONB
P1  PostgreSQL + DocumentDB, native PostgreSQL access
P2  PostgreSQL + DocumentDB + DocumentDB Gateway
P3  PostgreSQL + DocumentDB + FerretDB

Tenant placement identifiers:

T0  POOLED
T1  NAMESPACE
T2  DATABASE
T3  DEDICATED

Workload identifiers:

W0  Basic CRUD
W1  Flexible Schema
W2  Nested Documents
W3  Indexing
W4  Aggregation
W5  Transactions
W6  Change/Event Workloads
W7  Hybrid Relational/Document
W8  Multitenancy
W9  Operations

Additional profiles may be added, but existing identifiers should not be silently redefined.

2.4 One Question Per Experiment Where Possible

An experiment should isolate one major uncertainty.

For example:

P1 vs P2

is useful for estimating gateway/protocol overhead.

By contrast:

different hardware
+ different storage
+ different gateway
+ different indexes
+ different dataset

does not provide a useful causal comparison.

2.5 Negative Results Are Results

Unsupported operations, poor performance, operational complexity, failed upgrades, and undesirable tenant behavior must be documented rather than optimized out of the published evidence.

2.6 Benchmark Realistic Shapes

The benchmark program should include both:

  • controlled synthetic workloads for causal isolation
  • application-shaped workloads representing plausible production usage

Neither is sufficient alone.

2.7 Measure Tails, Not Only Means

At minimum, performance work should consider:

  • median
  • p95
  • p99
  • error rate
  • saturation behavior

Multitenant systems should additionally consider per-tenant tail latency.

2.8 Separate Cold and Warm Behavior

Relevant experiments should distinguish:

  • cold start
  • cold cache
  • warm cache
  • steady state
  • post-failover recovery

2.9 Cost Is an Architectural Metric

CPU, RAM, storage, IOPS, network, operational effort, and idle capacity are part of the outcome.

A system that performs similarly but consumes twice the infrastructure is not equivalent.

2.10 Upstream Changes Are Expected

CloudNativePG, PostgreSQL, DocumentDB, FerretDB, MongoDB, Kubernetes, drivers, and related tooling will evolve.

The research program must therefore support periodic regression runs against newer versions while retaining historical result comparability.


3. Decision Questions

The research program is organized around the following decision questions.

DQ-01 — Functional Viability

Can a CloudNativePG PostgreSQL cluster run the selected document engine and expose a stable MongoDB-compatible interface?

DQ-02 — Semantic Compatibility

Which MongoDB behaviors are identical, equivalent, different, unsupported, or incorrect?

DQ-03 — Application Compatibility

Can representative MongoDB applications and official drivers operate without modification?

DQ-04 — Document Performance

How does PostgreSQL-backed document access compare with native MongoDB under common document workloads?

DQ-05 — Abstraction Cost

What performance and resource costs are introduced by:

  • DocumentDB compared with JSONB?
  • Mongo-compatible gateways compared with native DocumentDB access?
  • DocumentDB Gateway compared with FerretDB?

DQ-06 — PostgreSQL Advantage

Which workloads benefit materially from retaining SQL, relational modeling, PostgreSQL transactions, RLS, and extension compatibility?

DQ-07 — Multitenancy

How do T0T3 compare in:

  • density
  • isolation
  • noisy-neighbour behavior
  • security
  • recoverability
  • operational complexity
  • cost?

DQ-08 — Tenant Mobility

Can a tenant move between placement models without forcing application-level architectural changes?

DQ-09 — Operational Resilience

How does the PostgreSQL-backed system behave under:

  • failover
  • restart
  • backup
  • restore
  • upgrade
  • storage pressure
  • node loss?

DQ-10 — Scaling Boundary

At what point does MongoDB sharding or another distributed architecture become materially superior to DocStorePG placement-based scaling?

DQ-11 — Economic Boundary

For which tenant and workload profiles does a shared PostgreSQL data plane offer lower total cost than MongoDB or dedicated placements?

DQ-12 — Product Boundary

What should DocStorePG ultimately provide itself, and what should remain upstream configuration?


4. Research Program Structure

The research program is divided into ten phases.

R0  Research Foundation
R1  Functional Baseline
R2  Compatibility & Conformance
R3  Performance Anatomy
R4  PostgreSQL Differentiation
R5  Multitenancy
R6  Tenant Mobility
R7  Operations & Resilience
R8  Scaling & Economics
R9  Long-Term Regression Program

The phases are ordered.

Not every experiment inside a phase must complete before later exploratory work begins, but each phase has a research gate that should be satisfied before major effort is committed to the next phase.


5. R0 — Research Foundation

Goal

Create a trustworthy experimental environment.

Primary Question

Can two independently executed benchmark runs be meaningfully compared?

Hypotheses

H-R0-01

A benchmark environment can be fully described from version-controlled configuration plus a small amount of machine-discovered metadata.

H-R0-02

Repeated runs of the same workload on the same environment will exhibit bounded variance sufficient for architectural comparison.

H-R0-03

Results can be traced to exact software versions, infrastructure profiles, workload definitions, and repository commits.

Required Outputs

BenchmarkSpecification.md
WorkloadCatalog.md
CompatibilityModel.md
results schema
environment manifest schema
benchmark runner skeleton
MongoDB reference deployment
P0 baseline deployment
P1/P2 deployment skeleton

Experiments

E-R0-001 — Environment Capture

Capture:

  • Kubernetes version
  • node type
  • CPU
  • memory
  • kernel
  • storage class
  • volume type
  • filesystem
  • network configuration
  • CNPG version
  • PostgreSQL version
  • MongoDB version
  • DocumentDB version
  • gateway version
  • driver version
  • repository commit

Pass condition: a result can be reproduced without undocumented manual configuration.

E-R0-002 — Run Repeatability

Execute a simple fixed CRUD workload at least 10 times under unchanged conditions.

Measure:

  • throughput
  • p50
  • p95
  • p99
  • CPU
  • memory

Determine variance.

Pass condition: observed variance is characterized and acceptable experimental confidence rules can be defined.

E-R0-003 — Resource Contention Baseline

Run the same benchmark:

  • on an idle cluster
  • with unrelated CPU load
  • with unrelated storage load

Purpose:

Identify how much environmental noise can distort results.

E-R0-004 — Result Provenance

Given any benchmark summary, verify that a reviewer can identify:

  • exact workload
  • exact dataset
  • exact implementation profile
  • exact tenant profile
  • exact software versions
  • raw result location
  • analysis code

Research Gate G0

Proceed when:

  • environments are reproducible
  • results are machine-readable
  • variance is understood
  • version pinning exists
  • M0 and at least P0 can execute the same basic workload

6. R1 — Functional Baseline

Goal

Demonstrate that the proposed architecture works end to end before deep compatibility analysis.

Primary Question

Can a MongoDB client successfully use a CloudNativePG-backed document store?

Hypotheses

H-R1-01

DocumentDB can run reliably within the selected CloudNativePG deployment model.

H-R1-02

DocumentDB Gateway can expose basic MongoDB CRUD through the CloudNativePG backend.

H-R1-03

FerretDB can operate against the same DocumentDB-backed PostgreSQL substrate as an alternative gateway.

H-R1-04

PostgreSQL-native access remains usable alongside the Mongo-compatible surface.

Experiments

E-R1-001 — P0 Native JSONB Baseline

Deploy:

CloudNativePG
+ PostgreSQL
+ JSONB workload

Verify:

  • create
  • insert
  • read
  • update
  • delete
  • index
  • backup
  • replica startup

E-R1-002 — P1 DocumentDB Native

Deploy:

CloudNativePG
+ PostgreSQL
+ DocumentDB extensions

Verify document operations through PostgreSQL-native interfaces.

E-R1-003 — P2 Gateway Path

Deploy DocumentDB Gateway.

Test with a standard MongoDB client:

  • connect
  • create collection
  • insert
  • find
  • update
  • delete
  • index
  • basic aggregation

E-R1-004 — P3 FerretDB Path

Repeat E-R1-003 through FerretDB.

E-R1-005 — Concurrent SQL and Mongo Access

Write through Mongo-compatible API and inspect/query through PostgreSQL.

Then reverse where safely supported.

Purpose:

Confirm the document layer remains part of a visible PostgreSQL platform rather than an opaque silo.

E-R1-006 — Restart Survival

Restart:

  • gateway
  • PostgreSQL pod
  • replica
  • client

Verify basic recovery.

Required Evidence

For each profile:

deployment manifest
successful smoke test
known failures
startup time
idle CPU
idle memory
storage footprint

Research Gate G1

Proceed when:

  • P1 and at least one Mongo-compatible gateway work end to end
  • M0 and candidate profiles can execute a shared basic workload
  • no unexplained data-corruption or fundamental consistency issue exists
  • PostgreSQL-native access has been verified

7. R2 — Compatibility & Conformance

Goal

Determine how compatible the PostgreSQL-backed document surface actually is.

Primary Question

What does "MongoDB compatible" mean for DocStorePG in practice?

Research Model

Every test case should produce one of:

IDENTICAL
EQUIVALENT
DOCUMENTED_DIFFERENCE
UNSUPPORTED
INCORRECT
UNKNOWN

The reference behavior is M0 for MongoDB-specific semantics.

Research Clusters

C-R2-A — BSON and Type Semantics

Test:

  • numbers
  • strings
  • booleans
  • null
  • missing fields
  • arrays
  • nested documents
  • dates
  • ObjectId
  • binary
  • decimal
  • regex
  • timestamps
  • min/max key where relevant

C-R2-B — CRUD

Test:

  • insertOne
  • insertMany
  • find
  • findOne
  • updateOne
  • updateMany
  • replaceOne
  • deleteOne
  • deleteMany
  • findAndModify family
  • bulk operations

C-R2-C — Query Operators

Cover representative families:

  • equality
  • comparison
  • logical
  • array
  • element
  • evaluation
  • text
  • expression

C-R2-D — Aggregation

Cover:

  • $match
  • $project
  • $group
  • $sort
  • $limit
  • $skip
  • $unwind
  • $lookup
  • $facet
  • expressions
  • nested pipelines

C-R2-E — Index Semantics

Test:

  • unique
  • compound
  • sparse behavior
  • partial behavior
  • nested fields
  • arrays
  • sort use
  • index errors
  • index lifecycle

C-R2-F — Transactions and Sessions

Test:

  • single-document atomicity
  • multi-document transactions
  • abort
  • commit
  • retry
  • concurrent writers
  • session behavior

C-R2-G — Errors

Compare:

  • duplicate key
  • malformed query
  • invalid update
  • nonexistent namespace
  • authorization failure
  • transaction conflict
  • unsupported operation

Error compatibility should be evaluated independently from successful-response compatibility.

C-R2-H — Drivers

At minimum investigate current official MongoDB drivers for:

  • Python
  • Node.js
  • Go
  • Java

Additional languages may be added based on real application demand.

C-R2-I — Tooling

Where practical:

  • mongosh
  • migration tools
  • common ODM/ORM libraries
  • backup/export tools
  • schema tools

C-R2-J — Change Streams

Determine:

  • availability
  • semantics
  • resume behavior
  • ordering
  • failover behavior
  • unsupported edge cases

Experiments

E-R2-001 — Differential Primitive Suite

Run identical request sets against:

M0
P2
P3

Compare structured outcomes.

E-R2-002 — Driver Matrix

Run the same functional suite through each target driver.

E-R2-003 — Error Matrix

Generate controlled invalid operations and compare behavior.

E-R2-004 — Aggregation Corpus

Create a growing corpus of aggregation pipelines ranging from trivial to complex.

E-R2-005 — Real Application Smoke Tests

Select at least three open-source MongoDB-backed applications with distinct characteristics.

Prefer applications that exercise:

  • normal CRUD
  • nested data
  • indexes
  • aggregation
  • transactions or change streams where possible

Run with no code changes first.

Classify required changes.

Outputs

CompatibilityModel.md
compatibility matrix
unsupported-operation registry
behavioral-difference registry
driver compatibility matrix
application compatibility reports

Research Gate G2

Proceed when:

  • basic CRUD/query semantics are characterized
  • known incompatibilities are explicit
  • at least one realistic application has been exercised
  • performance benchmarks can avoid comparing semantically different operations unknowingly

No specific compatibility percentage is required to proceed.

The purpose of G2 is understanding, not declaring victory.


8. R3 — Performance Anatomy

Goal

Understand where performance and resource costs originate.

Primary Question

What is the cost of each architectural layer?

Core Comparisons

P0 vs P1
    document-engine cost/value

P1 vs P2
    DocumentDB Gateway cost

P1 vs P3
    FerretDB gateway cost

P2 vs P3
    gateway implementation difference

M0 vs P2
    MongoDB vs DocumentDB gateway architecture

M0 vs P3
    MongoDB vs FerretDB architecture

Hypotheses

H-R3-01

Gateway translation overhead is measurable but not dominant for typical networked application workloads.

H-R3-02

DocumentDB's storage/index model will outperform generic JSONB for some Mongo-shaped queries but may underperform native JSONB for other PostgreSQL-native access patterns.

H-R3-03

The relative performance ranking changes with document shape, index count, concurrency, and query complexity.

H-R3-04

Resource efficiency may differ substantially even when latency is similar.

Benchmark Dimensions

Dataset Size

Example tiers:

S1  100k documents
S2  1m documents
S3  10m documents
S4  larger, where infrastructure permits

Document Size

Example distributions:

small       ~1 KB
medium      ~10 KB
large       ~100 KB
mixed

Concurrency

Example:

1
8
32
128
saturation search

Index Count

0
1
5
20

Cache State

cold
warm
steady-state

Experiments

E-R3-001 — CRUD Latency Curve

For W0:

  • insert
  • point read
  • indexed read
  • update
  • delete

Across concurrency levels.

E-R3-002 — Document Size Curve

Measure throughput and tail latency as document size increases.

E-R3-003 — Nested Query Cost

Use W2 with increasing nesting and array sizes.

E-R3-004 — Index Write Amplification

Measure:

  • insert throughput
  • update throughput
  • index size
  • WAL/write volume
  • CPU

as index count rises.

E-R3-005 — Aggregation Complexity

Create graded pipelines:

A1 simple match/project
A2 match/group
A3 unwind/group
A4 lookup
A5 facet
A6 complex mixed pipeline

E-R3-006 — Saturation Point

Increase concurrency until:

  • latency rises nonlinearly
  • throughput plateaus
  • errors increase

Record saturation behavior rather than only peak throughput.

E-R3-007 — Gateway Resource Isolation

Measure gateway CPU/memory separately from PostgreSQL.

E-R3-008 — Storage Efficiency

Compare:

  • raw dataset size
  • table/collection size
  • index size
  • WAL/oplog-related write volume where comparable
  • backup size

E-R3-009 — Connection Scaling

Measure:

  • idle connection cost
  • active connection cost
  • pool behavior
  • reconnect storms

Analysis Rule

Do not publish one global "DocStorePG vs MongoDB" performance score.

Results must remain workload-specific.

Research Gate G3

Proceed when:

  • performance anatomy is understood for basic workload classes
  • layer overhead can be attributed reasonably
  • at least one strength and one weakness of the PostgreSQL-backed approach are documented
  • test variance remains controlled

9. R4 — PostgreSQL Differentiation

Goal

Test whether retaining PostgreSQL provides capabilities that justify the architecture beyond MongoDB emulation.

Primary Question

What can DocStorePG do better because it is PostgreSQL?

Research Tracks

C-R4-A — Hybrid Relational/Document Queries

Example model:

customers
contracts
billing
permissions
        |
        +---- document collections

Test queries that combine relational predicates with document predicates.

C-R4-B — Cross-Model Transactions

Test transactions touching:

  • relational row
  • document
  • related audit record

C-R4-C — Row-Level Security

Evaluate whether pooled document access can benefit from PostgreSQL-enforced tenant policy.

C-R4-D — SQL Analytics

Compare operational document access with downstream SQL analysis.

C-R4-E — Extension Composition

Potential future experiments may include compatibility with relevant extensions such as:

  • PostGIS
  • pgvector
  • full-text/search extensions
  • observability extensions

Only test extensions for which there is a plausible workload.

Hypotheses

H-R4-01

Hybrid relational/document queries can reduce duplication or synchronization between separate relational and document systems.

H-R4-02

Cross-model transactions provide a meaningful correctness advantage for some applications.

H-R4-03

Database-enforced tenant policy can provide a stronger pooled-security boundary than application-only tenant filtering.

H-R4-04

The operational value of one PostgreSQL platform may exceed raw document-query performance differences.

Experiments

E-R4-001 — Customer/Document Hybrid Query

Create relational customer/contract data and document-oriented device/application data.

Query:

Find documents for customers satisfying relational commercial conditions and document capability predicates.

Compare implementation complexity and runtime behavior with M0-based alternatives.

E-R4-002 — Cross-Model Atomicity

Update:

contract state
+ document state
+ audit row

in one PostgreSQL transaction.

Document equivalent patterns required under M0.

E-R4-003 — SQL Reporting

Run analytical queries over mixed relational/document data without ETL.

E-R4-004 — RLS Isolation

Attempt cross-tenant access through:

  • direct SQL
  • pooled application connection
  • gateway path
  • malicious tenant identifier manipulation

E-R4-005 — Operational Stack Reduction

Compare required operational components for:

PostgreSQL + MongoDB

versus:

DocStorePG on PostgreSQL

Measure:

  • operators
  • backup systems
  • monitoring systems
  • secrets
  • upgrade procedures
  • runbooks
  • expertise requirements

This is partly qualitative but should use explicit inventories.

Research Gate G4

Proceed when:

  • at least one PostgreSQL-specific benefit has been demonstrated or disproven
  • hybrid architecture value can be discussed with evidence rather than assumption

If no meaningful PostgreSQL-specific advantage exists, reconsider whether Mongo compatibility on PostgreSQL is worth pursuing beyond migration use cases.


10. R5 — Multitenancy

Goal

Compare tenant placement strategies and MongoDB multitenancy/sharding approaches.

Primary Question

What is the best continuum between tenant density and tenant sovereignty?

Placement Profiles

T0 POOLED
T1 NAMESPACE
T2 DATABASE
T3 DEDICATED

MongoDB comparison profiles should distinguish at least:

MT0 shared collection with tenant key
MT1 database per tenant
MT2 sharded shared deployment
MT3 dedicated deployment

Exact MongoDB topology definitions belong in BenchmarkSpecification.md.

Research Dimensions

Isolation

  • data
  • identity
  • compute
  • memory
  • I/O
  • failure
  • backup
  • maintenance
  • configuration

Density

  • tenants per database
  • tenants per cluster
  • idle resource cost
  • object/catalog overhead
  • connection overhead

Operability

  • provisioning
  • upgrade
  • backup
  • restore
  • tenant deletion
  • credential rotation
  • migration

Security

  • database-enforced boundary
  • gateway enforcement
  • application enforcement
  • privilege escape risk

Performance

  • average latency
  • p95/p99 per tenant
  • noisy-neighbour effects

Experiments

E-R5-001 — Tenant Density Curve

Test:

10 tenants
100 tenants
1,000 tenants
10,000 tenants

where technically practical.

Not every placement mode must reach every tenant count.

Record failure/scaling boundaries.

E-R5-002 — Uniform Tenant Load

All tenants generate similar load.

Purpose:

Establish baseline fairness.

E-R5-003 — Whale Tenant

Canonical distribution:

many small tenants
some medium tenants
one tenant consuming ~50% of workload

Measure:

  • small-tenant p99
  • medium-tenant p99
  • whale latency
  • CPU share
  • I/O share
  • connection share

E-R5-004 — Bursty Tenant

One tenant periodically produces extreme short bursts.

Measure recovery of other tenants after burst completion.

E-R5-005 — Tenant Security Boundary

Attempt:

  • missing tenant filter
  • manipulated tenant identifier
  • reused connection context
  • role escalation
  • direct database access
  • gateway bypass

E-R5-006 — Provisioning Cost

Measure time and resources to provision:

new pooled tenant
new schema tenant
new database tenant
new dedicated tenant

E-R5-007 — Idle Cost

Measure per-tenant marginal cost when inactive.

E-R5-008 — Backup/Restore Granularity

Attempt restoration of a single tenant under T0T3.

Measure:

  • time
  • complexity
  • collateral impact

E-R5-009 — Failure Blast Radius

Inject failures and observe affected tenants.

SWOT Output

Each placement model should receive a structured SWOT based on measured results.

Example categories:

Strengths
Weaknesses
Opportunities
Threats
Recommended tenant profile
Known scale boundary

Research Gate G5

Proceed when:

  • T0, T2, and T3 have measured characteristics
  • T1 has either been validated or explicitly rejected
  • noisy-neighbour behavior is known
  • tenant-security enforcement has been exercised
  • a first tenant-placement decision matrix can be written

11. R6 — Tenant Mobility

Goal

Determine whether tenant placement can evolve without redesigning the application.

Primary Question

Can tenant isolation become a runtime/platform decision instead of a permanent application architecture decision?

Target Transitions

Priority order:

T0 -> T2
T2 -> T3
T0 -> T3

Optional later:

T1 -> T2
T3 -> T2
T2 -> T0

Reverse transitions may be operationally useful but have lower initial priority.

Hypotheses

H-R6-01

A tenant can move from pooled storage to an isolated database while retaining the same logical document API.

H-R6-02

A high-volume tenant can move to dedicated infrastructure with bounded interruption.

H-R6-03

Migration complexity depends strongly on change capture and gateway routing capabilities.

Migration Correctness Model

A migration must account for:

snapshot point
delta changes
validation
cutover
rollback
identity
indexes
metadata
permissions

Experiments

E-R6-001 — Offline T0 -> T2

Stop tenant writes.

Export.

Import.

Validate.

Switch routing.

Measure:

  • downtime
  • data correctness
  • operational steps

This establishes the simplest baseline.

E-R6-002 — Online-ish T0 -> T2

Investigate:

  • snapshot
  • delta capture
  • short cutover

E-R6-003 — T2 -> T3

Move isolated database tenant to dedicated CNPG deployment.

E-R6-004 — Whale Evacuation

Start E-R5-003 whale scenario.

Trigger tenant migration when policy threshold is exceeded.

Measure effect on smaller tenants before, during, and after migration.

E-R6-005 — Rollback

Introduce a target validation failure.

Return traffic to source safely.

E-R6-006 — Endpoint Stability

Determine whether application connection details can remain unchanged through a stable routing abstraction.

Research Gate G6

Proceed when:

  • at least one tenant migration path is reliable and documented
  • migration correctness can be validated automatically
  • downtime and operational cost are measured
  • routing requirements are understood

If tenant mobility is prohibitively complex, treat placement as an initial provisioning decision rather than a dynamic property.


12. R7 — Operations & Resilience

Goal

Compare operational behavior under realistic failure and lifecycle events.

Primary Question

Does DocStorePG inherit enough operational strength from CloudNativePG to justify the composed architecture?

Research Tracks

C-R7-A — Failover

C-R7-B — Backup and Recovery

C-R7-C — Upgrade

C-R7-D — Extension Lifecycle

C-R7-E — Gateway Lifecycle

C-R7-F — Disaster Scenarios

Experiments

E-R7-001 — Primary Failure

Terminate PostgreSQL primary.

Measure:

  • detection time
  • promotion time
  • Mongo client-visible errors
  • SQL client-visible errors
  • recovery time
  • lost/ambiguous operations
  • gateway reconnection behavior

E-R7-002 — Gateway Failure

Terminate gateway instances progressively.

Measure:

  • availability
  • reconnect behavior
  • load redistribution

E-R7-003 — Node Failure

Remove Kubernetes node hosting database or gateway workload.

E-R7-004 — Planned PostgreSQL Restart

Measure application-visible disruption.

E-R7-005 — Backup

Measure:

  • duration
  • CPU
  • I/O
  • storage size
  • impact on workload latency

E-R7-006 — Point-in-Time Recovery

Create known writes.

Restore to a target point.

Validate.

E-R7-007 — PostgreSQL Minor Upgrade

Evaluate normal lifecycle.

E-R7-008 — PostgreSQL Major Upgrade

Evaluate selected supported CNPG path and extension compatibility.

E-R7-009 — DocumentDB Extension Upgrade

Measure:

  • downtime
  • compatibility
  • rollback options
  • replica behavior

E-R7-010 — Gateway Upgrade

Perform rolling upgrade where supported.

E-R7-011 — Storage Pressure

Approach:

  • high disk usage
  • reduced IOPS
  • delayed storage

Measure behavior.

E-R7-012 — Connection Storm

Restart many clients or gateway pods simultaneously.

E-R7-013 — Recovery Runbook Test

Have a clean environment execute recovery using only repository documentation.

Purpose:

Test operational documentation itself.

Comparative Operations Scorecard

Do not collapse into one score unless weights are explicit.

Compare:

availability
RTO
RPO
backup complexity
restore granularity
upgrade complexity
observability
failure transparency
staff/operator burden

Research Gate G7

Proceed when:

  • common failures have reproducible outcomes
  • backup/restore works
  • upgrade path is understood
  • client-visible behavior is documented
  • operational weaknesses are explicit

13. R8 — Scaling & Economics

Goal

Define where DocStorePG is economically and architecturally preferable, and where MongoDB or another system is better.

Primary Question

Where are the practical boundaries of the architecture?

Research Tracks

C-R8-A — Vertical Scaling

Measure performance across compute sizes.

C-R8-B — Read Scaling

Evaluate replica-based reads where semantics permit.

C-R8-C — Gateway Scaling

Increase gateway replicas.

C-R8-D — Placement Scaling

Increase independent DocStorePG deployments and distribute tenants.

C-R8-E — MongoDB Sharding Comparison

Create MongoDB sharded reference scenarios.

C-R8-F — Cost Modeling

Estimate total infrastructure and operational cost.

Experiments

E-R8-001 — Vertical Efficiency Curve

For each system/profile:

small
medium
large

infrastructure sizes.

Measure throughput per:

  • CPU
  • GB RAM
  • storage cost unit

E-R8-002 — Read Replica Experiment

Test read-heavy workload with replicas.

Validate consistency expectations explicitly.

E-R8-003 — Gateway Horizontal Scale

Increase gateway replicas until database becomes bottleneck.

E-R8-004 — Multi-Placement Tenant Fleet

Distribute tenants across independent CNPG clusters.

Measure:

  • fleet overhead
  • operational complexity
  • routing
  • cost

E-R8-005 — MongoDB Shard Scaling

Compare a workload where MongoDB's native sharding should be advantaged.

Purpose:

Identify a genuine architectural boundary rather than avoiding it.

E-R8-006 — Whale Scaling Response

Compare:

MongoDB shard placement
vs
DocStorePG tenant evacuation to T3

Evaluate:

  • time to isolation
  • client transparency
  • operational complexity
  • steady-state cost

E-R8-007 — Cost Per Tenant

For each tenancy mode:

monthly infra cost
idle cost
active cost
backup cost
estimated operational cost

across tenant sizes.

E-R8-008 — Cost Per Million Operations

Normalize infrastructure consumption where meaningful.

E-R8-009 — Dual-Database Avoidance Value

Estimate architecture where an organization would otherwise operate:

PostgreSQL
+ MongoDB

versus:

PostgreSQL / DocStorePG only

Include:

  • compute
  • storage
  • backup
  • monitoring
  • engineering/operator effort
  • failure modes
  • data synchronization

This is a model, not a benchmark; assumptions must be explicit.

Research Gate G8

Complete when a defensible architecture decision guide can answer:

  • use P0
  • use P1
  • use P2/P3
  • use native MongoDB
  • use MongoDB sharding
  • use dedicated placement
  • use another architecture

for representative workload profiles.


14. R9 — Long-Term Regression Program

Goal

Keep conclusions valid as upstream systems evolve.

Trigger Events

A regression cycle should be considered when there is a meaningful release of:

  • PostgreSQL
  • CloudNativePG
  • DocumentDB
  • FerretDB
  • MongoDB
  • Kubernetes
  • major MongoDB drivers

Not every patch release requires full reruns.

Regression Tiers

L0 — Smoke

Run:

  • deployment
  • connect
  • W0 basic CRUD
  • backup smoke test

L1 — Compatibility

Run:

  • conformance core
  • driver core
  • known-difference regression

L2 — Performance

Run:

  • representative CRUD
  • aggregation
  • index
  • hybrid workload
  • tenant benchmark

L3 — Full Program

Run major benchmark suite including operations and multitenancy.

Regression Goals

Track:

compatibility gained
compatibility lost
performance gained
performance lost
resource efficiency changes
operational behavior changes
new capabilities
deprecated assumptions

15. Research Workstream Map

Research can proceed through semi-independent workstreams once foundational gates are satisfied.

flowchart TD
    R0[R0 Foundation]
    R1[R1 Functional Baseline]
    R2[R2 Compatibility]
    R3[R3 Performance]
    R4[R4 PostgreSQL Differentiation]
    R5[R5 Multitenancy]
    R6[R6 Tenant Mobility]
    R7[R7 Operations]
    R8[R8 Scaling & Economics]
    R9[R9 Regression]

    R0 --> R1
    R1 --> R2
    R2 --> R3
    R2 --> R4
    R2 --> R5

    R5 --> R6

    R3 --> R7
    R5 --> R7

    R3 --> R8
    R4 --> R8
    R5 --> R8
    R6 --> R8
    R7 --> R8

    R8 --> R9

This means R3, R4, and R5 may overlap after compatibility is sufficiently understood.


16. Priority Order

The default priority is:

P0  Reproducibility
P1  Functional viability
P2  Compatibility
P3  Multitenancy correctness
P4  Performance
P5  Hybrid PostgreSQL advantages
P6  Operations
P7  Tenant mobility
P8  Scaling/economics
P9  long-term optimization

A key reason multitenancy appears early is security:

A pooled architecture should not be performance-optimized before its tenant boundary is understood.


17. Minimal First Research Cycle

The first practical research cycle should deliberately be small.

Cycle 1 — "Can It Work?"

Implement:

M0
P0
P1
P2

Run:

W0 basic CRUD
small W2 nested data
basic W4 aggregation

Outputs:

deployment manifests
environment manifest
first compatibility table
first performance table
known issues

Cycle 2 — "Is It Really Compatible?"

Add:

P3
driver matrix
error behavior
index semantics
aggregation corpus

Cycle 3 — "Does PostgreSQL Buy Us Anything?"

Add:

W7 hybrid workload
RLS experiments
cross-model transactions

Cycle 4 — "Can We Host Tenants Safely?"

Add:

T0
T2
T3

Run:

uniform tenants
whale tenant
security tests
backup/restore granularity

Cycle 5 — "Can Placement Evolve?"

Implement:

T0 -> T2
T2 -> T3

Cycle 6 — "Where Are the Boundaries?"

Add:

HA
upgrade
large datasets
MongoDB sharding
cost models

This gives the repository a practical path without requiring the whole program to exist before useful findings emerge.


18. Experiment Definition Standard

Each experiment should have a Markdown specification.

Suggested format:

# E-R5-003 Whale Tenant

## Question

What effect does one dominant tenant have on other pooled tenants?

## Hypothesis

...

## Profiles

- P2/T0
- P3/T0
- M0/MT0
- M0/MT2

## Environment

...

## Dataset

...

## Workload

...

## Controlled Variables

...

## Independent Variable

...

## Measurements

...

## Procedure

...

## Acceptance / Interpretation Rules

...

## Results

...

## Conclusion

...

## Follow-Up
...

Experiment definitions should be versioned before results are generated when practical.


19. Evidence Hierarchy

Not all evidence has equal strength.

Use this hierarchy:

E0  Assumption
E1  Documentation evidence
E2  Manual observation
E3  Reproducible functional test
E4  Repeated controlled experiment
E5  Cross-environment replication
E6  Longitudinal production evidence

Architecture claims should state the strongest evidence level supporting them.

Example:

Claim:
P2 supports operation X.

Evidence:
E3 reproducible conformance test.

A benchmark finding based on one local run should not be described with production-level certainty.


20. Result Classification

Every significant finding should be classified.

CONFIRMED

Evidence supports the hypothesis within the tested boundary.

REJECTED

Evidence contradicts the hypothesis.

PARTIAL

Hypothesis holds only under defined conditions.

INCONCLUSIVE

Evidence is insufficient or contradictory.

BLOCKED

Experiment cannot currently be run because of missing functionality or infrastructure.

OBSOLETE

Upstream changes invalidate the experiment design or finding.


21. Performance Statistics

At minimum report:

operations
duration
throughput
p50
p95
p99
maximum
error count
error rate
CPU
memory
storage growth

Where relevant:

WAL volume
network volume
index size
cache hit ratio
replication lag
gateway CPU
gateway memory

Confidence

For important comparisons:

  • repeat runs
  • report dispersion
  • avoid interpreting very small differences as meaningful
  • establish a minimum practical effect size

The exact statistical method should be defined in BenchmarkSpecification.md.


22. Multitenant Metrics

Per-tenant metrics are required for W8.

At minimum:

tenant throughput
tenant p50
tenant p95
tenant p99
tenant errors
tenant storage

System-level averages must not hide tenant starvation.

Useful derived metrics may include:

latency fairness
throughput fairness
whale impact factor
isolation recovery time
marginal tenant cost

These should be formally defined before use.


23. Compatibility Metrics

A compatibility summary may report counts such as:

IDENTICAL
EQUIVALENT
DOCUMENTED_DIFFERENCE
UNSUPPORTED
INCORRECT
UNKNOWN

A percentage may be calculated for navigation, but must never replace the detailed matrix.

Weighted compatibility scores are allowed only when:

  • weighting is documented
  • use case is specified
  • unsupported high-impact features remain visible

24. Operational Metrics

Operational experiments should record:

event start
failure detection
service interruption start
recovery start
service usable
full steady state

Derived:

detection time
failover time
client interruption
recovery time
RTO
data loss / RPO
ambiguous operation count

25. Cost Model

Cost analysis should distinguish:

Infrastructure

  • compute
  • memory
  • block storage
  • object storage
  • network
  • backup retention

Platform Overhead

  • Kubernetes
  • gateways
  • replicas
  • idle capacity

Operational Overhead

Estimate separately:

  • deployment
  • upgrades
  • monitoring
  • backup
  • recovery
  • incident response
  • expertise

Operational estimates must be labeled as modeled assumptions unless measured from actual practice.


26. Research Environments

At least three environment classes are desirable over time.

DEV

Purpose:

  • functional work
  • rapid iteration
  • local/small Kubernetes

Not authoritative for performance.

LAB

Purpose:

  • controlled benchmarks
  • known hardware
  • low environmental noise
  • repeatability

Primary environment for published benchmark results.

FIELD

Purpose:

  • realistic long-running workload
  • operational evidence
  • future production-derived findings

Not required initially.


27. Dataset Strategy

Use multiple dataset families.

Synthetic Canonical

Purpose:

  • controlled variation
  • deterministic generation

Examples:

  • flat documents
  • nested documents
  • large arrays
  • heterogeneous schemas

Application-Shaped

Examples:

  • product catalog
  • device/configuration model
  • event metadata
  • user/profile model
  • content objects

Adversarial

Purpose:

Expose boundaries.

Examples:

  • extreme nesting
  • huge arrays
  • large documents
  • high-cardinality indexes
  • low-cardinality indexes
  • heavy update churn

Synthetic datasets should be generated from seedable code rather than committed as large files.


28. Workload Strategy

Each workload should describe:

read/write ratio
operation distribution
document distribution
query distribution
index assumptions
concurrency
tenant distribution
duration
warmup
think time
retry behavior

Avoid hidden client behavior that changes results.

MongoDB driver retries and PostgreSQL client retries should be understood and reported.


29. Benchmark Anti-Patterns

The research program should reject the following.

AP-01 — Peak TPS Marketing

Publishing only maximum throughput.

AP-02 — Single Run Results

Treating one benchmark run as definitive.

AP-03 — Different Hardware

Comparing systems on materially different infrastructure without normalization.

AP-04 — Semantic Mismatch

Comparing operations that produce different correctness guarantees.

AP-05 — Hidden Index Differences

One system having more or different indexes than the other.

AP-06 — Cache Ambiguity

Mixing cold- and warm-cache results.

AP-07 — Client Bottleneck

Benchmark runner saturates before database.

AP-08 — Average-Only Multitenancy

Hiding small-tenant degradation behind aggregate throughput.

AP-09 — Ignoring Errors

Counting failed operations as successful throughput or omitting error rate.

AP-10 — Unpinned Versions

Results that cannot be recreated.

AP-11 — Hand-Tuned Winner

Aggressively tuning only one system.

Tuning stages should be explicit:

default
reasonable production
expert optimized

AP-12 — Benchmark-Only Architecture

Optimizing implementation specifically for benchmark patterns that are not representative.


30. Architecture Decision Integration

Research outcomes should update architecture through ADRs.

Suggested relationship:

Experiment
    |
Finding
    |
Architecture Decision Record
    |
ArchitectureBlueprint.md update

Example:

E-R5-005
    |
RLS through gateway is unsafe under current connection model
    |
ADR-012 Reject T0 gateway RLS strategy B
    |
MultitenancyModel.md revised

This keeps the blueprint evidence-driven.


31. Research Artifact Map

The program should gradually produce:

INTENT.md
ArchitectureBlueprint.md
ResearchProgram.md
BenchmarkSpecification.md
CompatibilityModel.md
MultitenancyModel.md
WorkloadCatalog.md

docs/
  research/
  findings/
  decisions/
  references/

experiments/
  R0/
  R1/
  R2/
  R3/
  R4/
  R5/
  R6/
  R7/
  R8/

results/
  manifests/
  summaries/
  reports/

Possible future files:

TenantPlacementDecisionGuide.md
MongoCompatibilityReport.md
OperationsComparison.md
PerformanceReport.md
CostModel.md
ArchitectureDecisionGuide.md

32. Proposed Experiment Naming

Use:

E-R<phase>-<number>

Examples:

E-R2-001
E-R5-003
E-R7-001

Sub-experiments may use:

E-R5-003-A
E-R5-003-B

Names must remain stable once results are published.


33. Research Issue Labels

Suggested repository labels:

research
experiment
benchmark
compatibility
performance
multitenancy
security
migration
operations
resilience
cost
architecture
upstream
blocked
finding
regression

Experiment lifecycle labels:

experiment:proposed
experiment:designed
experiment:ready
experiment:running
experiment:analyzed
experiment:closed

34. Initial Backlog

Foundation

  • Create BenchmarkSpecification.md
  • Create WorkloadCatalog.md
  • Create CompatibilityModel.md
  • Define result JSON/YAML schema
  • Define environment manifest
  • Implement benchmark runner skeleton

Deployment

  • Deploy M0 MongoDB reference
  • Deploy P0 PostgreSQL JSONB
  • Deploy P1 DocumentDB native
  • Deploy P2 DocumentDB Gateway
  • Deploy P3 FerretDB
  • Pin versions
  • Capture deployment metadata

R0

  • E-R0-001 Environment Capture
  • E-R0-002 Run Repeatability
  • E-R0-003 Resource Contention Baseline
  • E-R0-004 Result Provenance

R1

  • E-R1-001 P0 baseline
  • E-R1-002 P1 DocumentDB native
  • E-R1-003 P2 gateway
  • E-R1-004 P3 FerretDB
  • E-R1-005 SQL + Mongo coexistence
  • E-R1-006 restart survival

R2

  • BSON/type corpus
  • CRUD corpus
  • query corpus
  • aggregation corpus
  • index corpus
  • transaction corpus
  • error corpus
  • driver matrix
  • first real application test

R3

  • CRUD latency curve
  • document size curve
  • nested query benchmark
  • index amplification benchmark
  • aggregation complexity benchmark
  • saturation benchmark
  • gateway isolation benchmark
  • storage efficiency benchmark

R4

  • hybrid customer/document workload
  • cross-model transaction test
  • SQL reporting test
  • RLS tenant-security test
  • operational stack inventory

R5

  • pooled T0 deployment
  • database T2 deployment
  • dedicated T3 deployment
  • assess T1 viability
  • tenant density curve
  • whale benchmark
  • bursty tenant benchmark
  • tenant isolation attack tests
  • provisioning benchmark
  • idle-cost benchmark
  • tenant restore benchmark

R6

  • T0 -> T2 offline migration
  • T0 -> T2 delta migration
  • T2 -> T3 migration
  • whale evacuation test
  • rollback test
  • stable endpoint experiment

R7

  • primary failure
  • gateway failure
  • node failure
  • restart
  • backup
  • PITR
  • minor upgrade
  • major upgrade
  • DocumentDB extension upgrade
  • gateway upgrade
  • storage pressure
  • connection storm
  • runbook validation

R8

  • vertical scale curve
  • read-replica study
  • gateway horizontal scale
  • multi-placement fleet
  • MongoDB sharding reference
  • whale scale response comparison
  • cost-per-tenant model
  • cost-per-million-operations model
  • dual-database avoidance model

35. First Milestone Definition

Milestone M1 — Document Store Baseline

M1 is achieved when:

  • M0, P0, P1, and P2 deploy reproducibly
  • basic W0 workload runs on all applicable profiles
  • first differential compatibility results exist
  • benchmark results are machine-readable
  • environment provenance is complete
  • repeated-run variance has been characterized

M1 should not require performance optimization.


36. Second Milestone Definition

Milestone M2 — Compatibility Characterized

M2 is achieved when:

  • CRUD compatibility is well understood
  • major query operators are classified
  • index behavior is classified
  • aggregation corpus exists
  • at least three official drivers have been tested
  • at least one real application has been tested
  • P3 has been evaluated
  • unsupported behaviors are explicit

37. Third Milestone Definition

Milestone M3 — Architectural Value Demonstrated

M3 is achieved when:

  • performance anatomy exists
  • at least one hybrid relational/document workload has been evaluated
  • T0, T2, and T3 exist
  • first whale-tenant benchmark exists
  • PostgreSQL-specific strengths and weaknesses have evidence

At M3 the project should explicitly decide whether deeper investment is justified.


38. Fourth Milestone Definition

Milestone M4 — Operationally Credible

M4 is achieved when:

  • HA failover is tested
  • backups are tested
  • recovery is tested
  • upgrades are tested
  • tenant security has been exercised
  • one tenant migration path works
  • operational runbooks exist

39. Fifth Milestone Definition

Milestone M5 — Architecture Decision Guide

M5 is achieved when the evidence supports a practical guide answering:

For this workload and tenant profile, which architecture should be used?

The guide should cover at least:

P0 JSONB
P1 DocumentDB native
P2/P3 Mongo-compatible PostgreSQL
M0 native MongoDB
MongoDB sharded
DocStorePG dedicated placement

M5 marks transition from exploratory research into long-term engineering/productization decisions.


40. Stop / Pivot Criteria

The research program should deliberately support stopping or changing direction.

Pivot P1 — Compatibility Layer Only

Trigger:

MongoDB compatibility is useful for migration but too incomplete for general runtime compatibility.

Outcome:

DocStorePG may become primarily:

migration
interop
document import/export
PostgreSQL convergence

Pivot P2 — JSONB Wins

Trigger:

P0 provides equivalent application value with materially lower complexity than P1P3 for target workloads.

Outcome:

Focus on PostgreSQL document patterns rather than MongoDB emulation.

Pivot P3 — MongoDB Remains Superior

Trigger:

MongoDB demonstrates material advantages in the workloads that actually matter and PostgreSQL-specific benefits do not offset them.

Outcome:

DocStorePG becomes a decision/benchmark framework rather than production alternative.

Pivot P4 — Multitenancy Becomes Core

Trigger:

Tenant placement and isolation prove substantially more valuable than MongoDB compatibility itself.

Outcome:

DocStorePG evolves toward a PostgreSQL tenant-placement/control-plane project.

Pivot P5 — Upstream Absorbs the Product

Trigger:

DocumentDB Operator or another upstream project provides nearly all runtime functionality needed.

Outcome:

DocStorePG concentrates on:

profiles
benchmarks
multitenancy research
conformance
decision support

rather than duplicating runtime orchestration.


41. Long-Term Success Criteria

doc-store-pg has long-term research value if it can maintain:

  1. A reproducible reference implementation.
  2. A versioned compatibility corpus.
  3. A versioned workload catalog.
  4. Comparative benchmark history.
  5. Multitenancy evidence.
  6. Operational resilience evidence.
  7. Architecture decision records tied to experiments.
  8. A clear account of MongoDB's remaining strengths.
  9. A clear account of PostgreSQL-specific strengths.
  10. Regression tests that detect upstream improvements or regressions.

The repository should become more useful as the ecosystem evolves, not obsolete after one benchmark campaign.


42. Immediate Next Actions

The recommended execution order from this document is:

1. BenchmarkSpecification.md
2. WorkloadCatalog.md
3. CompatibilityModel.md
4. M0/P0/P1/P2 reproducible deployments
5. R0 reproducibility experiments
6. R1 smoke tests
7. first R2 compatibility corpus
8. first T0 tenant-security experiment
9. only then broader performance benchmarking

This order intentionally puts experimental validity and correctness ahead of speed.


43. Guiding Research Principle

Do not ask whether PostgreSQL can replace MongoDB in the abstract.

Ask instead:

For a defined workload, tenant model, operational environment, and compatibility requirement, which architecture produces the best combination of correctness, performance, isolation, flexibility, and cost?

That is the question doc-store-pg should become increasingly good at answering.