51 KiB
Executable file
ResearchProgram.md
Repository
doc-store-pg
Status
Initial research program / execution roadmap.
Purpose
This document turns the hypotheses and open questions defined in INTENT.md and ArchitectureBlueprint.md into an ordered, reproducible research program.
The goal is not to accumulate benchmarks.
The goal is to produce enough evidence to make progressively stronger architectural decisions about whether, where, and how a PostgreSQL / CloudNativePG platform can provide document-store capabilities that are competitive with MongoDB while retaining PostgreSQL-native strengths.
The research program should answer three classes of questions:
-
Can it work?
Compatibility, correctness, integration, and operational feasibility. -
How well does it work?
Performance, efficiency, resilience, multitenancy, and cost. -
Where should it be used?
Architectural fit, workload boundaries, tenant placement, and long-term product direction.
The program is deliberately staged so that expensive performance and scale research only begins after correctness and experimental reproducibility have been established.
1. Research North Star
The central research hypothesis is:
Document storage can be implemented as a composable PostgreSQL capability that is operationally native to CloudNativePG, sufficiently compatible with MongoDB for meaningful application reuse, and superior in selected hybrid or multitenant scenarios without requiring MongoDB's architecture to be reproduced wholesale.
The program must be capable of disproving this hypothesis.
A successful research outcome may therefore be any of the following:
- PostgreSQL-based document storage is broadly competitive with MongoDB.
- PostgreSQL-based document storage is attractive only for specific workload classes.
- MongoDB remains clearly preferable for important document workloads.
- Different tenant classes should use different storage architectures.
- DocumentDB or FerretDB is useful only as a migration/compatibility surface.
- Native PostgreSQL JSONB is sufficient for many workloads and the MongoDB layer adds little value.
- A hybrid architecture is more compelling than a replacement architecture.
- The operational simplicity of one PostgreSQL platform outweighs moderate performance disadvantages.
- The complexity introduced by MongoDB compatibility is not justified.
The research program exists to distinguish among these outcomes.
2. Research Principles
2.1 Correctness Before Performance
A system that returns different results more quickly is not faster at the same task.
Compatibility and semantic behavior must therefore be understood before comparative performance results are treated as meaningful.
2.2 Reproducibility Before Optimization
Every important result must be reproducible from:
- a repository commit
- a machine-readable environment definition
- pinned software versions
- a workload definition
- a dataset definition
- a random seed where relevant
- benchmark configuration
- raw result references
- analysis code
2.3 Compare Architectures, Not Brands
The primary objects of comparison are deployment and access profiles.
Stable profile identifiers from ArchitectureBlueprint.md are used:
M0 Native MongoDB reference
P0 PostgreSQL + native JSONB
P1 PostgreSQL + DocumentDB, native PostgreSQL access
P2 PostgreSQL + DocumentDB + DocumentDB Gateway
P3 PostgreSQL + DocumentDB + FerretDB
Tenant placement identifiers:
T0 POOLED
T1 NAMESPACE
T2 DATABASE
T3 DEDICATED
Workload identifiers:
W0 Basic CRUD
W1 Flexible Schema
W2 Nested Documents
W3 Indexing
W4 Aggregation
W5 Transactions
W6 Change/Event Workloads
W7 Hybrid Relational/Document
W8 Multitenancy
W9 Operations
Additional profiles may be added, but existing identifiers should not be silently redefined.
2.4 One Question Per Experiment Where Possible
An experiment should isolate one major uncertainty.
For example:
P1 vs P2
is useful for estimating gateway/protocol overhead.
By contrast:
different hardware
+ different storage
+ different gateway
+ different indexes
+ different dataset
does not provide a useful causal comparison.
2.5 Negative Results Are Results
Unsupported operations, poor performance, operational complexity, failed upgrades, and undesirable tenant behavior must be documented rather than optimized out of the published evidence.
2.6 Benchmark Realistic Shapes
The benchmark program should include both:
- controlled synthetic workloads for causal isolation
- application-shaped workloads representing plausible production usage
Neither is sufficient alone.
2.7 Measure Tails, Not Only Means
At minimum, performance work should consider:
- median
- p95
- p99
- error rate
- saturation behavior
Multitenant systems should additionally consider per-tenant tail latency.
2.8 Separate Cold and Warm Behavior
Relevant experiments should distinguish:
- cold start
- cold cache
- warm cache
- steady state
- post-failover recovery
2.9 Cost Is an Architectural Metric
CPU, RAM, storage, IOPS, network, operational effort, and idle capacity are part of the outcome.
A system that performs similarly but consumes twice the infrastructure is not equivalent.
2.10 Upstream Changes Are Expected
CloudNativePG, PostgreSQL, DocumentDB, FerretDB, MongoDB, Kubernetes, drivers, and related tooling will evolve.
The research program must therefore support periodic regression runs against newer versions while retaining historical result comparability.
3. Decision Questions
The research program is organized around the following decision questions.
DQ-01 — Functional Viability
Can a CloudNativePG PostgreSQL cluster run the selected document engine and expose a stable MongoDB-compatible interface?
DQ-02 — Semantic Compatibility
Which MongoDB behaviors are identical, equivalent, different, unsupported, or incorrect?
DQ-03 — Application Compatibility
Can representative MongoDB applications and official drivers operate without modification?
DQ-04 — Document Performance
How does PostgreSQL-backed document access compare with native MongoDB under common document workloads?
DQ-05 — Abstraction Cost
What performance and resource costs are introduced by:
- DocumentDB compared with JSONB?
- Mongo-compatible gateways compared with native DocumentDB access?
- DocumentDB Gateway compared with FerretDB?
DQ-06 — PostgreSQL Advantage
Which workloads benefit materially from retaining SQL, relational modeling, PostgreSQL transactions, RLS, and extension compatibility?
DQ-07 — Multitenancy
How do T0–T3 compare in:
- density
- isolation
- noisy-neighbour behavior
- security
- recoverability
- operational complexity
- cost?
DQ-08 — Tenant Mobility
Can a tenant move between placement models without forcing application-level architectural changes?
DQ-09 — Operational Resilience
How does the PostgreSQL-backed system behave under:
- failover
- restart
- backup
- restore
- upgrade
- storage pressure
- node loss?
DQ-10 — Scaling Boundary
At what point does MongoDB sharding or another distributed architecture become materially superior to DocStorePG placement-based scaling?
DQ-11 — Economic Boundary
For which tenant and workload profiles does a shared PostgreSQL data plane offer lower total cost than MongoDB or dedicated placements?
DQ-12 — Product Boundary
What should DocStorePG ultimately provide itself, and what should remain upstream configuration?
4. Research Program Structure
The research program is divided into ten phases.
R0 Research Foundation
R1 Functional Baseline
R2 Compatibility & Conformance
R3 Performance Anatomy
R4 PostgreSQL Differentiation
R5 Multitenancy
R6 Tenant Mobility
R7 Operations & Resilience
R8 Scaling & Economics
R9 Long-Term Regression Program
The phases are ordered.
Not every experiment inside a phase must complete before later exploratory work begins, but each phase has a research gate that should be satisfied before major effort is committed to the next phase.
5. R0 — Research Foundation
Goal
Create a trustworthy experimental environment.
Primary Question
Can two independently executed benchmark runs be meaningfully compared?
Hypotheses
H-R0-01
A benchmark environment can be fully described from version-controlled configuration plus a small amount of machine-discovered metadata.
H-R0-02
Repeated runs of the same workload on the same environment will exhibit bounded variance sufficient for architectural comparison.
H-R0-03
Results can be traced to exact software versions, infrastructure profiles, workload definitions, and repository commits.
Required Outputs
BenchmarkSpecification.md
WorkloadCatalog.md
CompatibilityModel.md
results schema
environment manifest schema
benchmark runner skeleton
MongoDB reference deployment
P0 baseline deployment
P1/P2 deployment skeleton
Experiments
E-R0-001 — Environment Capture
Capture:
- Kubernetes version
- node type
- CPU
- memory
- kernel
- storage class
- volume type
- filesystem
- network configuration
- CNPG version
- PostgreSQL version
- MongoDB version
- DocumentDB version
- gateway version
- driver version
- repository commit
Pass condition: a result can be reproduced without undocumented manual configuration.
E-R0-002 — Run Repeatability
Execute a simple fixed CRUD workload at least 10 times under unchanged conditions.
Measure:
- throughput
- p50
- p95
- p99
- CPU
- memory
Determine variance.
Pass condition: observed variance is characterized and acceptable experimental confidence rules can be defined.
E-R0-003 — Resource Contention Baseline
Run the same benchmark:
- on an idle cluster
- with unrelated CPU load
- with unrelated storage load
Purpose:
Identify how much environmental noise can distort results.
E-R0-004 — Result Provenance
Given any benchmark summary, verify that a reviewer can identify:
- exact workload
- exact dataset
- exact implementation profile
- exact tenant profile
- exact software versions
- raw result location
- analysis code
Research Gate G0
Proceed when:
- environments are reproducible
- results are machine-readable
- variance is understood
- version pinning exists
- M0 and at least P0 can execute the same basic workload
6. R1 — Functional Baseline
Goal
Demonstrate that the proposed architecture works end to end before deep compatibility analysis.
Primary Question
Can a MongoDB client successfully use a CloudNativePG-backed document store?
Hypotheses
H-R1-01
DocumentDB can run reliably within the selected CloudNativePG deployment model.
H-R1-02
DocumentDB Gateway can expose basic MongoDB CRUD through the CloudNativePG backend.
H-R1-03
FerretDB can operate against the same DocumentDB-backed PostgreSQL substrate as an alternative gateway.
H-R1-04
PostgreSQL-native access remains usable alongside the Mongo-compatible surface.
Experiments
E-R1-001 — P0 Native JSONB Baseline
Deploy:
CloudNativePG
+ PostgreSQL
+ JSONB workload
Verify:
- create
- insert
- read
- update
- delete
- index
- backup
- replica startup
E-R1-002 — P1 DocumentDB Native
Deploy:
CloudNativePG
+ PostgreSQL
+ DocumentDB extensions
Verify document operations through PostgreSQL-native interfaces.
E-R1-003 — P2 Gateway Path
Deploy DocumentDB Gateway.
Test with a standard MongoDB client:
- connect
- create collection
- insert
- find
- update
- delete
- index
- basic aggregation
E-R1-004 — P3 FerretDB Path
Repeat E-R1-003 through FerretDB.
E-R1-005 — Concurrent SQL and Mongo Access
Write through Mongo-compatible API and inspect/query through PostgreSQL.
Then reverse where safely supported.
Purpose:
Confirm the document layer remains part of a visible PostgreSQL platform rather than an opaque silo.
E-R1-006 — Restart Survival
Restart:
- gateway
- PostgreSQL pod
- replica
- client
Verify basic recovery.
Required Evidence
For each profile:
deployment manifest
successful smoke test
known failures
startup time
idle CPU
idle memory
storage footprint
Research Gate G1
Proceed when:
- P1 and at least one Mongo-compatible gateway work end to end
- M0 and candidate profiles can execute a shared basic workload
- no unexplained data-corruption or fundamental consistency issue exists
- PostgreSQL-native access has been verified
7. R2 — Compatibility & Conformance
Goal
Determine how compatible the PostgreSQL-backed document surface actually is.
Primary Question
What does "MongoDB compatible" mean for DocStorePG in practice?
Research Model
Every test case should produce one of:
IDENTICAL
EQUIVALENT
DOCUMENTED_DIFFERENCE
UNSUPPORTED
INCORRECT
UNKNOWN
The reference behavior is M0 for MongoDB-specific semantics.
Research Clusters
C-R2-A — BSON and Type Semantics
Test:
- numbers
- strings
- booleans
- null
- missing fields
- arrays
- nested documents
- dates
- ObjectId
- binary
- decimal
- regex
- timestamps
- min/max key where relevant
C-R2-B — CRUD
Test:
- insertOne
- insertMany
- find
- findOne
- updateOne
- updateMany
- replaceOne
- deleteOne
- deleteMany
- findAndModify family
- bulk operations
C-R2-C — Query Operators
Cover representative families:
- equality
- comparison
- logical
- array
- element
- evaluation
- text
- expression
C-R2-D — Aggregation
Cover:
$match$project$group$sort$limit$skip$unwind$lookup$facet- expressions
- nested pipelines
C-R2-E — Index Semantics
Test:
- unique
- compound
- sparse behavior
- partial behavior
- nested fields
- arrays
- sort use
- index errors
- index lifecycle
C-R2-F — Transactions and Sessions
Test:
- single-document atomicity
- multi-document transactions
- abort
- commit
- retry
- concurrent writers
- session behavior
C-R2-G — Errors
Compare:
- duplicate key
- malformed query
- invalid update
- nonexistent namespace
- authorization failure
- transaction conflict
- unsupported operation
Error compatibility should be evaluated independently from successful-response compatibility.
C-R2-H — Drivers
At minimum investigate current official MongoDB drivers for:
- Python
- Node.js
- Go
- Java
Additional languages may be added based on real application demand.
C-R2-I — Tooling
Where practical:
mongosh- migration tools
- common ODM/ORM libraries
- backup/export tools
- schema tools
C-R2-J — Change Streams
Determine:
- availability
- semantics
- resume behavior
- ordering
- failover behavior
- unsupported edge cases
Experiments
E-R2-001 — Differential Primitive Suite
Run identical request sets against:
M0
P2
P3
Compare structured outcomes.
E-R2-002 — Driver Matrix
Run the same functional suite through each target driver.
E-R2-003 — Error Matrix
Generate controlled invalid operations and compare behavior.
E-R2-004 — Aggregation Corpus
Create a growing corpus of aggregation pipelines ranging from trivial to complex.
E-R2-005 — Real Application Smoke Tests
Select at least three open-source MongoDB-backed applications with distinct characteristics.
Prefer applications that exercise:
- normal CRUD
- nested data
- indexes
- aggregation
- transactions or change streams where possible
Run with no code changes first.
Classify required changes.
Outputs
CompatibilityModel.md
compatibility matrix
unsupported-operation registry
behavioral-difference registry
driver compatibility matrix
application compatibility reports
Research Gate G2
Proceed when:
- basic CRUD/query semantics are characterized
- known incompatibilities are explicit
- at least one realistic application has been exercised
- performance benchmarks can avoid comparing semantically different operations unknowingly
No specific compatibility percentage is required to proceed.
The purpose of G2 is understanding, not declaring victory.
8. R3 — Performance Anatomy
Goal
Understand where performance and resource costs originate.
Primary Question
What is the cost of each architectural layer?
Core Comparisons
P0 vs P1
document-engine cost/value
P1 vs P2
DocumentDB Gateway cost
P1 vs P3
FerretDB gateway cost
P2 vs P3
gateway implementation difference
M0 vs P2
MongoDB vs DocumentDB gateway architecture
M0 vs P3
MongoDB vs FerretDB architecture
Hypotheses
H-R3-01
Gateway translation overhead is measurable but not dominant for typical networked application workloads.
H-R3-02
DocumentDB's storage/index model will outperform generic JSONB for some Mongo-shaped queries but may underperform native JSONB for other PostgreSQL-native access patterns.
H-R3-03
The relative performance ranking changes with document shape, index count, concurrency, and query complexity.
H-R3-04
Resource efficiency may differ substantially even when latency is similar.
Benchmark Dimensions
Dataset Size
Example tiers:
S1 100k documents
S2 1m documents
S3 10m documents
S4 larger, where infrastructure permits
Document Size
Example distributions:
small ~1 KB
medium ~10 KB
large ~100 KB
mixed
Concurrency
Example:
1
8
32
128
saturation search
Index Count
0
1
5
20
Cache State
cold
warm
steady-state
Experiments
E-R3-001 — CRUD Latency Curve
For W0:
- insert
- point read
- indexed read
- update
- delete
Across concurrency levels.
E-R3-002 — Document Size Curve
Measure throughput and tail latency as document size increases.
E-R3-003 — Nested Query Cost
Use W2 with increasing nesting and array sizes.
E-R3-004 — Index Write Amplification
Measure:
- insert throughput
- update throughput
- index size
- WAL/write volume
- CPU
as index count rises.
E-R3-005 — Aggregation Complexity
Create graded pipelines:
A1 simple match/project
A2 match/group
A3 unwind/group
A4 lookup
A5 facet
A6 complex mixed pipeline
E-R3-006 — Saturation Point
Increase concurrency until:
- latency rises nonlinearly
- throughput plateaus
- errors increase
Record saturation behavior rather than only peak throughput.
E-R3-007 — Gateway Resource Isolation
Measure gateway CPU/memory separately from PostgreSQL.
E-R3-008 — Storage Efficiency
Compare:
- raw dataset size
- table/collection size
- index size
- WAL/oplog-related write volume where comparable
- backup size
E-R3-009 — Connection Scaling
Measure:
- idle connection cost
- active connection cost
- pool behavior
- reconnect storms
Analysis Rule
Do not publish one global "DocStorePG vs MongoDB" performance score.
Results must remain workload-specific.
Research Gate G3
Proceed when:
- performance anatomy is understood for basic workload classes
- layer overhead can be attributed reasonably
- at least one strength and one weakness of the PostgreSQL-backed approach are documented
- test variance remains controlled
9. R4 — PostgreSQL Differentiation
Goal
Test whether retaining PostgreSQL provides capabilities that justify the architecture beyond MongoDB emulation.
Primary Question
What can DocStorePG do better because it is PostgreSQL?
Research Tracks
C-R4-A — Hybrid Relational/Document Queries
Example model:
customers
contracts
billing
permissions
|
+---- document collections
Test queries that combine relational predicates with document predicates.
C-R4-B — Cross-Model Transactions
Test transactions touching:
- relational row
- document
- related audit record
C-R4-C — Row-Level Security
Evaluate whether pooled document access can benefit from PostgreSQL-enforced tenant policy.
C-R4-D — SQL Analytics
Compare operational document access with downstream SQL analysis.
C-R4-E — Extension Composition
Potential future experiments may include compatibility with relevant extensions such as:
- PostGIS
- pgvector
- full-text/search extensions
- observability extensions
Only test extensions for which there is a plausible workload.
Hypotheses
H-R4-01
Hybrid relational/document queries can reduce duplication or synchronization between separate relational and document systems.
H-R4-02
Cross-model transactions provide a meaningful correctness advantage for some applications.
H-R4-03
Database-enforced tenant policy can provide a stronger pooled-security boundary than application-only tenant filtering.
H-R4-04
The operational value of one PostgreSQL platform may exceed raw document-query performance differences.
Experiments
E-R4-001 — Customer/Document Hybrid Query
Create relational customer/contract data and document-oriented device/application data.
Query:
Find documents for customers satisfying relational commercial conditions and document capability predicates.
Compare implementation complexity and runtime behavior with M0-based alternatives.
E-R4-002 — Cross-Model Atomicity
Update:
contract state
+ document state
+ audit row
in one PostgreSQL transaction.
Document equivalent patterns required under M0.
E-R4-003 — SQL Reporting
Run analytical queries over mixed relational/document data without ETL.
E-R4-004 — RLS Isolation
Attempt cross-tenant access through:
- direct SQL
- pooled application connection
- gateway path
- malicious tenant identifier manipulation
E-R4-005 — Operational Stack Reduction
Compare required operational components for:
PostgreSQL + MongoDB
versus:
DocStorePG on PostgreSQL
Measure:
- operators
- backup systems
- monitoring systems
- secrets
- upgrade procedures
- runbooks
- expertise requirements
This is partly qualitative but should use explicit inventories.
Research Gate G4
Proceed when:
- at least one PostgreSQL-specific benefit has been demonstrated or disproven
- hybrid architecture value can be discussed with evidence rather than assumption
If no meaningful PostgreSQL-specific advantage exists, reconsider whether Mongo compatibility on PostgreSQL is worth pursuing beyond migration use cases.
10. R5 — Multitenancy
Goal
Compare tenant placement strategies and MongoDB multitenancy/sharding approaches.
Primary Question
What is the best continuum between tenant density and tenant sovereignty?
Placement Profiles
T0 POOLED
T1 NAMESPACE
T2 DATABASE
T3 DEDICATED
MongoDB comparison profiles should distinguish at least:
MT0 shared collection with tenant key
MT1 database per tenant
MT2 sharded shared deployment
MT3 dedicated deployment
Exact MongoDB topology definitions belong in BenchmarkSpecification.md.
Research Dimensions
Isolation
- data
- identity
- compute
- memory
- I/O
- failure
- backup
- maintenance
- configuration
Density
- tenants per database
- tenants per cluster
- idle resource cost
- object/catalog overhead
- connection overhead
Operability
- provisioning
- upgrade
- backup
- restore
- tenant deletion
- credential rotation
- migration
Security
- database-enforced boundary
- gateway enforcement
- application enforcement
- privilege escape risk
Performance
- average latency
- p95/p99 per tenant
- noisy-neighbour effects
Experiments
E-R5-001 — Tenant Density Curve
Test:
10 tenants
100 tenants
1,000 tenants
10,000 tenants
where technically practical.
Not every placement mode must reach every tenant count.
Record failure/scaling boundaries.
E-R5-002 — Uniform Tenant Load
All tenants generate similar load.
Purpose:
Establish baseline fairness.
E-R5-003 — Whale Tenant
Canonical distribution:
many small tenants
some medium tenants
one tenant consuming ~50% of workload
Measure:
- small-tenant p99
- medium-tenant p99
- whale latency
- CPU share
- I/O share
- connection share
E-R5-004 — Bursty Tenant
One tenant periodically produces extreme short bursts.
Measure recovery of other tenants after burst completion.
E-R5-005 — Tenant Security Boundary
Attempt:
- missing tenant filter
- manipulated tenant identifier
- reused connection context
- role escalation
- direct database access
- gateway bypass
E-R5-006 — Provisioning Cost
Measure time and resources to provision:
new pooled tenant
new schema tenant
new database tenant
new dedicated tenant
E-R5-007 — Idle Cost
Measure per-tenant marginal cost when inactive.
E-R5-008 — Backup/Restore Granularity
Attempt restoration of a single tenant under T0–T3.
Measure:
- time
- complexity
- collateral impact
E-R5-009 — Failure Blast Radius
Inject failures and observe affected tenants.
SWOT Output
Each placement model should receive a structured SWOT based on measured results.
Example categories:
Strengths
Weaknesses
Opportunities
Threats
Recommended tenant profile
Known scale boundary
Research Gate G5
Proceed when:
- T0, T2, and T3 have measured characteristics
- T1 has either been validated or explicitly rejected
- noisy-neighbour behavior is known
- tenant-security enforcement has been exercised
- a first tenant-placement decision matrix can be written
11. R6 — Tenant Mobility
Goal
Determine whether tenant placement can evolve without redesigning the application.
Primary Question
Can tenant isolation become a runtime/platform decision instead of a permanent application architecture decision?
Target Transitions
Priority order:
T0 -> T2
T2 -> T3
T0 -> T3
Optional later:
T1 -> T2
T3 -> T2
T2 -> T0
Reverse transitions may be operationally useful but have lower initial priority.
Hypotheses
H-R6-01
A tenant can move from pooled storage to an isolated database while retaining the same logical document API.
H-R6-02
A high-volume tenant can move to dedicated infrastructure with bounded interruption.
H-R6-03
Migration complexity depends strongly on change capture and gateway routing capabilities.
Migration Correctness Model
A migration must account for:
snapshot point
delta changes
validation
cutover
rollback
identity
indexes
metadata
permissions
Experiments
E-R6-001 — Offline T0 -> T2
Stop tenant writes.
Export.
Import.
Validate.
Switch routing.
Measure:
- downtime
- data correctness
- operational steps
This establishes the simplest baseline.
E-R6-002 — Online-ish T0 -> T2
Investigate:
- snapshot
- delta capture
- short cutover
E-R6-003 — T2 -> T3
Move isolated database tenant to dedicated CNPG deployment.
E-R6-004 — Whale Evacuation
Start E-R5-003 whale scenario.
Trigger tenant migration when policy threshold is exceeded.
Measure effect on smaller tenants before, during, and after migration.
E-R6-005 — Rollback
Introduce a target validation failure.
Return traffic to source safely.
E-R6-006 — Endpoint Stability
Determine whether application connection details can remain unchanged through a stable routing abstraction.
Research Gate G6
Proceed when:
- at least one tenant migration path is reliable and documented
- migration correctness can be validated automatically
- downtime and operational cost are measured
- routing requirements are understood
If tenant mobility is prohibitively complex, treat placement as an initial provisioning decision rather than a dynamic property.
12. R7 — Operations & Resilience
Goal
Compare operational behavior under realistic failure and lifecycle events.
Primary Question
Does DocStorePG inherit enough operational strength from CloudNativePG to justify the composed architecture?
Research Tracks
C-R7-A — Failover
C-R7-B — Backup and Recovery
C-R7-C — Upgrade
C-R7-D — Extension Lifecycle
C-R7-E — Gateway Lifecycle
C-R7-F — Disaster Scenarios
Experiments
E-R7-001 — Primary Failure
Terminate PostgreSQL primary.
Measure:
- detection time
- promotion time
- Mongo client-visible errors
- SQL client-visible errors
- recovery time
- lost/ambiguous operations
- gateway reconnection behavior
E-R7-002 — Gateway Failure
Terminate gateway instances progressively.
Measure:
- availability
- reconnect behavior
- load redistribution
E-R7-003 — Node Failure
Remove Kubernetes node hosting database or gateway workload.
E-R7-004 — Planned PostgreSQL Restart
Measure application-visible disruption.
E-R7-005 — Backup
Measure:
- duration
- CPU
- I/O
- storage size
- impact on workload latency
E-R7-006 — Point-in-Time Recovery
Create known writes.
Restore to a target point.
Validate.
E-R7-007 — PostgreSQL Minor Upgrade
Evaluate normal lifecycle.
E-R7-008 — PostgreSQL Major Upgrade
Evaluate selected supported CNPG path and extension compatibility.
E-R7-009 — DocumentDB Extension Upgrade
Measure:
- downtime
- compatibility
- rollback options
- replica behavior
E-R7-010 — Gateway Upgrade
Perform rolling upgrade where supported.
E-R7-011 — Storage Pressure
Approach:
- high disk usage
- reduced IOPS
- delayed storage
Measure behavior.
E-R7-012 — Connection Storm
Restart many clients or gateway pods simultaneously.
E-R7-013 — Recovery Runbook Test
Have a clean environment execute recovery using only repository documentation.
Purpose:
Test operational documentation itself.
Comparative Operations Scorecard
Do not collapse into one score unless weights are explicit.
Compare:
availability
RTO
RPO
backup complexity
restore granularity
upgrade complexity
observability
failure transparency
staff/operator burden
Research Gate G7
Proceed when:
- common failures have reproducible outcomes
- backup/restore works
- upgrade path is understood
- client-visible behavior is documented
- operational weaknesses are explicit
13. R8 — Scaling & Economics
Goal
Define where DocStorePG is economically and architecturally preferable, and where MongoDB or another system is better.
Primary Question
Where are the practical boundaries of the architecture?
Research Tracks
C-R8-A — Vertical Scaling
Measure performance across compute sizes.
C-R8-B — Read Scaling
Evaluate replica-based reads where semantics permit.
C-R8-C — Gateway Scaling
Increase gateway replicas.
C-R8-D — Placement Scaling
Increase independent DocStorePG deployments and distribute tenants.
C-R8-E — MongoDB Sharding Comparison
Create MongoDB sharded reference scenarios.
C-R8-F — Cost Modeling
Estimate total infrastructure and operational cost.
Experiments
E-R8-001 — Vertical Efficiency Curve
For each system/profile:
small
medium
large
infrastructure sizes.
Measure throughput per:
- CPU
- GB RAM
- storage cost unit
E-R8-002 — Read Replica Experiment
Test read-heavy workload with replicas.
Validate consistency expectations explicitly.
E-R8-003 — Gateway Horizontal Scale
Increase gateway replicas until database becomes bottleneck.
E-R8-004 — Multi-Placement Tenant Fleet
Distribute tenants across independent CNPG clusters.
Measure:
- fleet overhead
- operational complexity
- routing
- cost
E-R8-005 — MongoDB Shard Scaling
Compare a workload where MongoDB's native sharding should be advantaged.
Purpose:
Identify a genuine architectural boundary rather than avoiding it.
E-R8-006 — Whale Scaling Response
Compare:
MongoDB shard placement
vs
DocStorePG tenant evacuation to T3
Evaluate:
- time to isolation
- client transparency
- operational complexity
- steady-state cost
E-R8-007 — Cost Per Tenant
For each tenancy mode:
monthly infra cost
idle cost
active cost
backup cost
estimated operational cost
across tenant sizes.
E-R8-008 — Cost Per Million Operations
Normalize infrastructure consumption where meaningful.
E-R8-009 — Dual-Database Avoidance Value
Estimate architecture where an organization would otherwise operate:
PostgreSQL
+ MongoDB
versus:
PostgreSQL / DocStorePG only
Include:
- compute
- storage
- backup
- monitoring
- engineering/operator effort
- failure modes
- data synchronization
This is a model, not a benchmark; assumptions must be explicit.
Research Gate G8
Complete when a defensible architecture decision guide can answer:
- use P0
- use P1
- use P2/P3
- use native MongoDB
- use MongoDB sharding
- use dedicated placement
- use another architecture
for representative workload profiles.
14. R9 — Long-Term Regression Program
Goal
Keep conclusions valid as upstream systems evolve.
Trigger Events
A regression cycle should be considered when there is a meaningful release of:
- PostgreSQL
- CloudNativePG
- DocumentDB
- FerretDB
- MongoDB
- Kubernetes
- major MongoDB drivers
Not every patch release requires full reruns.
Regression Tiers
L0 — Smoke
Run:
- deployment
- connect
- W0 basic CRUD
- backup smoke test
L1 — Compatibility
Run:
- conformance core
- driver core
- known-difference regression
L2 — Performance
Run:
- representative CRUD
- aggregation
- index
- hybrid workload
- tenant benchmark
L3 — Full Program
Run major benchmark suite including operations and multitenancy.
Regression Goals
Track:
compatibility gained
compatibility lost
performance gained
performance lost
resource efficiency changes
operational behavior changes
new capabilities
deprecated assumptions
15. Research Workstream Map
Research can proceed through semi-independent workstreams once foundational gates are satisfied.
flowchart TD
R0[R0 Foundation]
R1[R1 Functional Baseline]
R2[R2 Compatibility]
R3[R3 Performance]
R4[R4 PostgreSQL Differentiation]
R5[R5 Multitenancy]
R6[R6 Tenant Mobility]
R7[R7 Operations]
R8[R8 Scaling & Economics]
R9[R9 Regression]
R0 --> R1
R1 --> R2
R2 --> R3
R2 --> R4
R2 --> R5
R5 --> R6
R3 --> R7
R5 --> R7
R3 --> R8
R4 --> R8
R5 --> R8
R6 --> R8
R7 --> R8
R8 --> R9
This means R3, R4, and R5 may overlap after compatibility is sufficiently understood.
16. Priority Order
The default priority is:
P0 Reproducibility
P1 Functional viability
P2 Compatibility
P3 Multitenancy correctness
P4 Performance
P5 Hybrid PostgreSQL advantages
P6 Operations
P7 Tenant mobility
P8 Scaling/economics
P9 long-term optimization
A key reason multitenancy appears early is security:
A pooled architecture should not be performance-optimized before its tenant boundary is understood.
17. Minimal First Research Cycle
The first practical research cycle should deliberately be small.
Cycle 1 — "Can It Work?"
Implement:
M0
P0
P1
P2
Run:
W0 basic CRUD
small W2 nested data
basic W4 aggregation
Outputs:
deployment manifests
environment manifest
first compatibility table
first performance table
known issues
Cycle 2 — "Is It Really Compatible?"
Add:
P3
driver matrix
error behavior
index semantics
aggregation corpus
Cycle 3 — "Does PostgreSQL Buy Us Anything?"
Add:
W7 hybrid workload
RLS experiments
cross-model transactions
Cycle 4 — "Can We Host Tenants Safely?"
Add:
T0
T2
T3
Run:
uniform tenants
whale tenant
security tests
backup/restore granularity
Cycle 5 — "Can Placement Evolve?"
Implement:
T0 -> T2
T2 -> T3
Cycle 6 — "Where Are the Boundaries?"
Add:
HA
upgrade
large datasets
MongoDB sharding
cost models
This gives the repository a practical path without requiring the whole program to exist before useful findings emerge.
18. Experiment Definition Standard
Each experiment should have a Markdown specification.
Suggested format:
# E-R5-003 Whale Tenant
## Question
What effect does one dominant tenant have on other pooled tenants?
## Hypothesis
...
## Profiles
- P2/T0
- P3/T0
- M0/MT0
- M0/MT2
## Environment
...
## Dataset
...
## Workload
...
## Controlled Variables
...
## Independent Variable
...
## Measurements
...
## Procedure
...
## Acceptance / Interpretation Rules
...
## Results
...
## Conclusion
...
## Follow-Up
...
Experiment definitions should be versioned before results are generated when practical.
19. Evidence Hierarchy
Not all evidence has equal strength.
Use this hierarchy:
E0 Assumption
E1 Documentation evidence
E2 Manual observation
E3 Reproducible functional test
E4 Repeated controlled experiment
E5 Cross-environment replication
E6 Longitudinal production evidence
Architecture claims should state the strongest evidence level supporting them.
Example:
Claim:
P2 supports operation X.
Evidence:
E3 reproducible conformance test.
A benchmark finding based on one local run should not be described with production-level certainty.
20. Result Classification
Every significant finding should be classified.
CONFIRMED
Evidence supports the hypothesis within the tested boundary.
REJECTED
Evidence contradicts the hypothesis.
PARTIAL
Hypothesis holds only under defined conditions.
INCONCLUSIVE
Evidence is insufficient or contradictory.
BLOCKED
Experiment cannot currently be run because of missing functionality or infrastructure.
OBSOLETE
Upstream changes invalidate the experiment design or finding.
21. Performance Statistics
At minimum report:
operations
duration
throughput
p50
p95
p99
maximum
error count
error rate
CPU
memory
storage growth
Where relevant:
WAL volume
network volume
index size
cache hit ratio
replication lag
gateway CPU
gateway memory
Confidence
For important comparisons:
- repeat runs
- report dispersion
- avoid interpreting very small differences as meaningful
- establish a minimum practical effect size
The exact statistical method should be defined in BenchmarkSpecification.md.
22. Multitenant Metrics
Per-tenant metrics are required for W8.
At minimum:
tenant throughput
tenant p50
tenant p95
tenant p99
tenant errors
tenant storage
System-level averages must not hide tenant starvation.
Useful derived metrics may include:
latency fairness
throughput fairness
whale impact factor
isolation recovery time
marginal tenant cost
These should be formally defined before use.
23. Compatibility Metrics
A compatibility summary may report counts such as:
IDENTICAL
EQUIVALENT
DOCUMENTED_DIFFERENCE
UNSUPPORTED
INCORRECT
UNKNOWN
A percentage may be calculated for navigation, but must never replace the detailed matrix.
Weighted compatibility scores are allowed only when:
- weighting is documented
- use case is specified
- unsupported high-impact features remain visible
24. Operational Metrics
Operational experiments should record:
event start
failure detection
service interruption start
recovery start
service usable
full steady state
Derived:
detection time
failover time
client interruption
recovery time
RTO
data loss / RPO
ambiguous operation count
25. Cost Model
Cost analysis should distinguish:
Infrastructure
- compute
- memory
- block storage
- object storage
- network
- backup retention
Platform Overhead
- Kubernetes
- gateways
- replicas
- idle capacity
Operational Overhead
Estimate separately:
- deployment
- upgrades
- monitoring
- backup
- recovery
- incident response
- expertise
Operational estimates must be labeled as modeled assumptions unless measured from actual practice.
26. Research Environments
At least three environment classes are desirable over time.
DEV
Purpose:
- functional work
- rapid iteration
- local/small Kubernetes
Not authoritative for performance.
LAB
Purpose:
- controlled benchmarks
- known hardware
- low environmental noise
- repeatability
Primary environment for published benchmark results.
FIELD
Purpose:
- realistic long-running workload
- operational evidence
- future production-derived findings
Not required initially.
27. Dataset Strategy
Use multiple dataset families.
Synthetic Canonical
Purpose:
- controlled variation
- deterministic generation
Examples:
- flat documents
- nested documents
- large arrays
- heterogeneous schemas
Application-Shaped
Examples:
- product catalog
- device/configuration model
- event metadata
- user/profile model
- content objects
Adversarial
Purpose:
Expose boundaries.
Examples:
- extreme nesting
- huge arrays
- large documents
- high-cardinality indexes
- low-cardinality indexes
- heavy update churn
Synthetic datasets should be generated from seedable code rather than committed as large files.
28. Workload Strategy
Each workload should describe:
read/write ratio
operation distribution
document distribution
query distribution
index assumptions
concurrency
tenant distribution
duration
warmup
think time
retry behavior
Avoid hidden client behavior that changes results.
MongoDB driver retries and PostgreSQL client retries should be understood and reported.
29. Benchmark Anti-Patterns
The research program should reject the following.
AP-01 — Peak TPS Marketing
Publishing only maximum throughput.
AP-02 — Single Run Results
Treating one benchmark run as definitive.
AP-03 — Different Hardware
Comparing systems on materially different infrastructure without normalization.
AP-04 — Semantic Mismatch
Comparing operations that produce different correctness guarantees.
AP-05 — Hidden Index Differences
One system having more or different indexes than the other.
AP-06 — Cache Ambiguity
Mixing cold- and warm-cache results.
AP-07 — Client Bottleneck
Benchmark runner saturates before database.
AP-08 — Average-Only Multitenancy
Hiding small-tenant degradation behind aggregate throughput.
AP-09 — Ignoring Errors
Counting failed operations as successful throughput or omitting error rate.
AP-10 — Unpinned Versions
Results that cannot be recreated.
AP-11 — Hand-Tuned Winner
Aggressively tuning only one system.
Tuning stages should be explicit:
default
reasonable production
expert optimized
AP-12 — Benchmark-Only Architecture
Optimizing implementation specifically for benchmark patterns that are not representative.
30. Architecture Decision Integration
Research outcomes should update architecture through ADRs.
Suggested relationship:
Experiment
|
Finding
|
Architecture Decision Record
|
ArchitectureBlueprint.md update
Example:
E-R5-005
|
RLS through gateway is unsafe under current connection model
|
ADR-012 Reject T0 gateway RLS strategy B
|
MultitenancyModel.md revised
This keeps the blueprint evidence-driven.
31. Research Artifact Map
The program should gradually produce:
INTENT.md
ArchitectureBlueprint.md
ResearchProgram.md
BenchmarkSpecification.md
CompatibilityModel.md
MultitenancyModel.md
WorkloadCatalog.md
docs/
research/
findings/
decisions/
references/
experiments/
R0/
R1/
R2/
R3/
R4/
R5/
R6/
R7/
R8/
results/
manifests/
summaries/
reports/
Possible future files:
TenantPlacementDecisionGuide.md
MongoCompatibilityReport.md
OperationsComparison.md
PerformanceReport.md
CostModel.md
ArchitectureDecisionGuide.md
32. Proposed Experiment Naming
Use:
E-R<phase>-<number>
Examples:
E-R2-001
E-R5-003
E-R7-001
Sub-experiments may use:
E-R5-003-A
E-R5-003-B
Names must remain stable once results are published.
33. Research Issue Labels
Suggested repository labels:
research
experiment
benchmark
compatibility
performance
multitenancy
security
migration
operations
resilience
cost
architecture
upstream
blocked
finding
regression
Experiment lifecycle labels:
experiment:proposed
experiment:designed
experiment:ready
experiment:running
experiment:analyzed
experiment:closed
34. Initial Backlog
Foundation
- Create
BenchmarkSpecification.md - Create
WorkloadCatalog.md - Create
CompatibilityModel.md - Define result JSON/YAML schema
- Define environment manifest
- Implement benchmark runner skeleton
Deployment
- Deploy M0 MongoDB reference
- Deploy P0 PostgreSQL JSONB
- Deploy P1 DocumentDB native
- Deploy P2 DocumentDB Gateway
- Deploy P3 FerretDB
- Pin versions
- Capture deployment metadata
R0
- E-R0-001 Environment Capture
- E-R0-002 Run Repeatability
- E-R0-003 Resource Contention Baseline
- E-R0-004 Result Provenance
R1
- E-R1-001 P0 baseline
- E-R1-002 P1 DocumentDB native
- E-R1-003 P2 gateway
- E-R1-004 P3 FerretDB
- E-R1-005 SQL + Mongo coexistence
- E-R1-006 restart survival
R2
- BSON/type corpus
- CRUD corpus
- query corpus
- aggregation corpus
- index corpus
- transaction corpus
- error corpus
- driver matrix
- first real application test
R3
- CRUD latency curve
- document size curve
- nested query benchmark
- index amplification benchmark
- aggregation complexity benchmark
- saturation benchmark
- gateway isolation benchmark
- storage efficiency benchmark
R4
- hybrid customer/document workload
- cross-model transaction test
- SQL reporting test
- RLS tenant-security test
- operational stack inventory
R5
- pooled T0 deployment
- database T2 deployment
- dedicated T3 deployment
- assess T1 viability
- tenant density curve
- whale benchmark
- bursty tenant benchmark
- tenant isolation attack tests
- provisioning benchmark
- idle-cost benchmark
- tenant restore benchmark
R6
- T0 -> T2 offline migration
- T0 -> T2 delta migration
- T2 -> T3 migration
- whale evacuation test
- rollback test
- stable endpoint experiment
R7
- primary failure
- gateway failure
- node failure
- restart
- backup
- PITR
- minor upgrade
- major upgrade
- DocumentDB extension upgrade
- gateway upgrade
- storage pressure
- connection storm
- runbook validation
R8
- vertical scale curve
- read-replica study
- gateway horizontal scale
- multi-placement fleet
- MongoDB sharding reference
- whale scale response comparison
- cost-per-tenant model
- cost-per-million-operations model
- dual-database avoidance model
35. First Milestone Definition
Milestone M1 — Document Store Baseline
M1 is achieved when:
- M0, P0, P1, and P2 deploy reproducibly
- basic W0 workload runs on all applicable profiles
- first differential compatibility results exist
- benchmark results are machine-readable
- environment provenance is complete
- repeated-run variance has been characterized
M1 should not require performance optimization.
36. Second Milestone Definition
Milestone M2 — Compatibility Characterized
M2 is achieved when:
- CRUD compatibility is well understood
- major query operators are classified
- index behavior is classified
- aggregation corpus exists
- at least three official drivers have been tested
- at least one real application has been tested
- P3 has been evaluated
- unsupported behaviors are explicit
37. Third Milestone Definition
Milestone M3 — Architectural Value Demonstrated
M3 is achieved when:
- performance anatomy exists
- at least one hybrid relational/document workload has been evaluated
- T0, T2, and T3 exist
- first whale-tenant benchmark exists
- PostgreSQL-specific strengths and weaknesses have evidence
At M3 the project should explicitly decide whether deeper investment is justified.
38. Fourth Milestone Definition
Milestone M4 — Operationally Credible
M4 is achieved when:
- HA failover is tested
- backups are tested
- recovery is tested
- upgrades are tested
- tenant security has been exercised
- one tenant migration path works
- operational runbooks exist
39. Fifth Milestone Definition
Milestone M5 — Architecture Decision Guide
M5 is achieved when the evidence supports a practical guide answering:
For this workload and tenant profile, which architecture should be used?
The guide should cover at least:
P0 JSONB
P1 DocumentDB native
P2/P3 Mongo-compatible PostgreSQL
M0 native MongoDB
MongoDB sharded
DocStorePG dedicated placement
M5 marks transition from exploratory research into long-term engineering/productization decisions.
40. Stop / Pivot Criteria
The research program should deliberately support stopping or changing direction.
Pivot P1 — Compatibility Layer Only
Trigger:
MongoDB compatibility is useful for migration but too incomplete for general runtime compatibility.
Outcome:
DocStorePG may become primarily:
migration
interop
document import/export
PostgreSQL convergence
Pivot P2 — JSONB Wins
Trigger:
P0 provides equivalent application value with materially lower complexity than P1–P3 for target workloads.
Outcome:
Focus on PostgreSQL document patterns rather than MongoDB emulation.
Pivot P3 — MongoDB Remains Superior
Trigger:
MongoDB demonstrates material advantages in the workloads that actually matter and PostgreSQL-specific benefits do not offset them.
Outcome:
DocStorePG becomes a decision/benchmark framework rather than production alternative.
Pivot P4 — Multitenancy Becomes Core
Trigger:
Tenant placement and isolation prove substantially more valuable than MongoDB compatibility itself.
Outcome:
DocStorePG evolves toward a PostgreSQL tenant-placement/control-plane project.
Pivot P5 — Upstream Absorbs the Product
Trigger:
DocumentDB Operator or another upstream project provides nearly all runtime functionality needed.
Outcome:
DocStorePG concentrates on:
profiles
benchmarks
multitenancy research
conformance
decision support
rather than duplicating runtime orchestration.
41. Long-Term Success Criteria
doc-store-pg has long-term research value if it can maintain:
- A reproducible reference implementation.
- A versioned compatibility corpus.
- A versioned workload catalog.
- Comparative benchmark history.
- Multitenancy evidence.
- Operational resilience evidence.
- Architecture decision records tied to experiments.
- A clear account of MongoDB's remaining strengths.
- A clear account of PostgreSQL-specific strengths.
- Regression tests that detect upstream improvements or regressions.
The repository should become more useful as the ecosystem evolves, not obsolete after one benchmark campaign.
42. Immediate Next Actions
The recommended execution order from this document is:
1. BenchmarkSpecification.md
2. WorkloadCatalog.md
3. CompatibilityModel.md
4. M0/P0/P1/P2 reproducible deployments
5. R0 reproducibility experiments
6. R1 smoke tests
7. first R2 compatibility corpus
8. first T0 tenant-security experiment
9. only then broader performance benchmarking
This order intentionally puts experimental validity and correctness ahead of speed.
43. Guiding Research Principle
Do not ask whether PostgreSQL can replace MongoDB in the abstract.
Ask instead:
For a defined workload, tenant model, operational environment, and compatibility requirement, which architecture produces the best combination of correctness, performance, isolation, flexibility, and cost?
That is the question doc-store-pg should become increasingly good at answering.