2026-07-30 01:29:13 +02:00
|
|
|
|
# Development Effort Calculator — Concept
|
|
|
|
|
|
|
|
|
|
|
|
Status: Concept v0.1
|
|
|
|
|
|
Date: 2026-07-30
|
2026-07-30 11:02:45 +02:00
|
|
|
|
Workplan: `workplans/TREV-WP-0010-development-effort-calculator.md`
|
2026-07-30 01:29:13 +02:00
|
|
|
|
Primary artifacts: `specs/TargetRevenueFrameworkCore.md` §1.4 (Target Multiple, Initial Target formula), `specs/PhaseManifestSpecification.md` (`target_basis` fields), state hub `get_token_summary`/`get_recent_progress`/`list_tasks` tools
|
|
|
|
|
|
|
|
|
|
|
|
**Scope note:** this is a concept for a **working, explicitly refinable v0
|
|
|
|
|
|
estimator** — not a final, authoritative valuation methodology. Every
|
|
|
|
|
|
formula below is a starting point the maintainer expects to iterate on
|
|
|
|
|
|
once real repos produce real data, consistent with how
|
|
|
|
|
|
`specs/TargetDegenerationPolicyResearch.md` treated its own formula
|
|
|
|
|
|
proposal.
|
|
|
|
|
|
|
|
|
|
|
|
---
|
|
|
|
|
|
|
|
|
|
|
|
## 1. What this calculator produces
|
|
|
|
|
|
|
|
|
|
|
|
It does **not** invent a new target-setting mechanism. It produces
|
|
|
|
|
|
concrete values for the fields `specs/PhaseManifestSpecification.md`
|
|
|
|
|
|
already defines under `phase.target_basis`:
|
|
|
|
|
|
|
|
|
|
|
|
- `estimated_effort_days`
|
|
|
|
|
|
- `daily_rate`
|
|
|
|
|
|
- `approved_direct_costs`
|
|
|
|
|
|
- `target_multiple`
|
|
|
|
|
|
|
|
|
|
|
|
and lets `specs/TargetRevenueFrameworkCore.md` §1.4's existing formula do
|
|
|
|
|
|
the rest:
|
|
|
|
|
|
|
|
|
|
|
|
```
|
|
|
|
|
|
Estimated Development Cost = (estimated_effort_days × daily_rate) + approved_direct_costs
|
|
|
|
|
|
Initial Target = Estimated Development Cost × Target Multiple
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
This calculator's job is narrowly: turn measurable repo activity into
|
|
|
|
|
|
defensible values for the left-hand side, and turn a repo's Target
|
|
|
|
|
|
Multiple *classification* (10x/100x/1000x — Incremental/Product-defining/
|
|
|
|
|
|
Platform-defining) into the right-hand multiplier, per the same guidance
|
|
|
|
|
|
table already in `specs/TargetRevenueFrameworkCore.md` §1.4. It is a
|
|
|
|
|
|
value-input tool feeding an existing formula, not a competing formula.
|
|
|
|
|
|
|
|
|
|
|
|
## 2. Input metrics (v0 candidate set)
|
|
|
|
|
|
|
|
|
|
|
|
Four metric families, per the maintainer's request:
|
|
|
|
|
|
|
|
|
|
|
|
### 2a. Human interaction time
|
|
|
|
|
|
|
|
|
|
|
|
**What:** hours a human spent interacting with the repo in any way
|
|
|
|
|
|
(coding sessions, review, planning conversations).
|
|
|
|
|
|
|
|
|
|
|
|
**Candidate sources, weakest to strongest signal:**
|
|
|
|
|
|
- Git commit timestamp clustering: group commits into sessions (a gap
|
|
|
|
|
|
threshold, e.g. commits within 2 hours of each other belong to the same
|
|
|
|
|
|
session), sum session durations. Cheap, works on any repo's existing
|
|
|
|
|
|
history, but systematically undercounts (thinking/reading time between
|
|
|
|
|
|
commits, non-commit work).
|
|
|
|
|
|
- State hub `get_recent_progress`/session records for repos already
|
|
|
|
|
|
tracked there — likely more accurate for repos actively worked on
|
|
|
|
|
|
through this agent/hub workflow, but not available for repos with no
|
|
|
|
|
|
hub-tracked history.
|
|
|
|
|
|
- Explicit human-logged time (a repo's own time-tracking, if it has one) —
|
|
|
|
|
|
most accurate where available, not universal.
|
|
|
|
|
|
|
|
|
|
|
|
**v0 default:** commit-timestamp clustering, since it works uniformly
|
|
|
|
|
|
across every repo without depending on hub coverage — explicitly a
|
|
|
|
|
|
*floor* estimate (real effort is almost certainly higher), not a
|
|
|
|
|
|
best-effort true count.
|
|
|
|
|
|
|
|
|
|
|
|
### 2b. Workplan/task volume — complexity index component
|
|
|
|
|
|
|
|
|
|
|
|
**What:** number and status distribution of workplans and tasks
|
|
|
|
|
|
associated with the repo (state hub `list_tasks`/`list_workplans` scoped
|
|
|
|
|
|
to the repo, or direct inspection of a `workplans/` directory as this
|
|
|
|
|
|
repo's own convention already does).
|
|
|
|
|
|
|
|
|
|
|
|
**v0 formula sketch:** `complexity_units_workplans = finished_workplans × 3 + finished_tasks × 1`
|
|
|
|
|
|
(a workplan is worth more than a bare task, since it implies scoping and
|
|
|
|
|
|
sequencing work beyond the sum of its tasks) — **weights are placeholders,
|
|
|
|
|
|
not derived from any data yet**, flagged explicitly for refinement once
|
|
|
|
|
|
multiple repos' actual effort-vs-workplan-count relationship is known.
|
|
|
|
|
|
|
|
|
|
|
|
### 2c. Repo size — complexity index component
|
|
|
|
|
|
|
|
|
|
|
|
**What:** file count and line count in the repo, as a raw
|
|
|
|
|
|
complexity/scope proxy.
|
|
|
|
|
|
|
|
|
|
|
|
**v0 formula sketch:** `complexity_units_size = file_count × 0.1 + line_count × 0.001`
|
|
|
|
|
|
— deliberately crude (a config repo with 10,000 generated lines is not
|
|
|
|
|
|
"more effort" than a 500-line algorithm-dense repo), which is exactly why
|
|
|
|
|
|
this is one input among several, not a standalone valuation. **Known
|
|
|
|
|
|
distortion to correct for in refinement:** vendored/generated
|
|
|
|
|
|
files/lines, `node_modules`-style dependency trees, and lockfiles
|
|
|
|
|
|
inflate this signal without reflecting real effort — a v1 refinement
|
|
|
|
|
|
should exclude common generated/vendored paths (`.gitignore`-aware or an
|
|
|
|
|
|
explicit exclude list) before counting.
|
|
|
|
|
|
|
|
|
|
|
|
### 2d. AI token cost — estimated and consumed
|
|
|
|
|
|
|
|
|
|
|
|
**What:** the actual $ cost of AI-assisted implementation work on this
|
|
|
|
|
|
repo, both consumed-to-date and, where a task is still in flight,
|
|
|
|
|
|
estimated to completion.
|
|
|
|
|
|
|
|
|
|
|
|
**Source:** the state hub already tracks this — `get_token_summary(scope="repo", id=<repo_id>)`
|
|
|
|
|
|
aggregates recorded token consumption per repo, populated via
|
|
|
|
|
|
`record_token_event`. This is the **strongest, least speculative** input
|
|
|
|
|
|
in this whole set, since it is a real, metered cost already being
|
|
|
|
|
|
recorded for other purposes (token budget policy), not a proxy.
|
|
|
|
|
|
|
|
|
|
|
|
**v0 formula sketch:** `token_cost_usd = tokens_in × input_price_per_token + tokens_out × output_price_per_token`,
|
|
|
|
|
|
using the actual per-model pricing for whatever model(s) `get_token_summary`
|
|
|
|
|
|
reports were used — this is not an estimate at all where data exists, it
|
|
|
|
|
|
is a direct cost figure; it only becomes an *estimate* for effort not yet
|
|
|
|
|
|
performed (extrapolating from a partial task to its expected completion).
|
|
|
|
|
|
|
|
|
|
|
|
## 3. Combining the metrics — v0 formula (explicitly provisional)
|
|
|
|
|
|
|
|
|
|
|
|
Two candidate combination strategies, presented as alternatives for the
|
|
|
|
|
|
maintainer to weigh rather than a single forced answer (matching
|
|
|
|
|
|
`specs/TargetDegenerationPolicyResearch.md`'s own pattern of presenting,
|
|
|
|
|
|
not silently picking, a design fork):
|
|
|
|
|
|
|
|
|
|
|
|
### Candidate A — labor-cost-anchored (recommended v0 default)
|
|
|
|
|
|
|
|
|
|
|
|
Treat human time as the primary cost driver (`estimated_effort_days`,
|
|
|
|
|
|
`daily_rate` directly), and route token cost into `approved_direct_costs`
|
|
|
|
|
|
(a real, metered direct cost, exactly what that field is for per
|
|
|
|
|
|
`specs/OpenQuestions-WorkingDefaults.md` Q5). The workplan/task-volume and
|
|
|
|
|
|
repo-size complexity indices (§2b, §2c) are used as a **sanity check /
|
|
|
|
|
|
adjustment factor** on the human-time estimate — e.g., if commit-timestamp
|
|
|
|
|
|
clustering suggests very little time but the repo has a large, mature
|
|
|
|
|
|
workplan/task history, that's a signal the time-tracking input is
|
|
|
|
|
|
under-counting and should be revised upward, not a separate dollar figure
|
|
|
|
|
|
of its own.
|
|
|
|
|
|
|
|
|
|
|
|
```
|
|
|
|
|
|
estimated_effort_days = (commit-clustered hours) / 8 [sanity-adjusted per §2b/§2c signal]
|
|
|
|
|
|
daily_rate = a declared, repo-independent rate (not computed from metrics)
|
|
|
|
|
|
approved_direct_costs = token_cost_usd (from §2d)
|
|
|
|
|
|
target_multiple = per the Target Multiple classification (§4)
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
### Candidate B — composite-index-anchored
|
|
|
|
|
|
|
|
|
|
|
|
Convert every metric family into a common "complexity unit" score
|
|
|
|
|
|
(weighted sum across §2a–§2d, all four as scored inputs, not three
|
|
|
|
|
|
checking one), then map the total score to a dollar figure via a
|
|
|
|
|
|
declared $-per-unit rate. More uniform across very different repos
|
|
|
|
|
|
(a repo with almost no logged human time but a huge, real workplan/task
|
|
|
|
|
|
history and heavy AI-token spend wouldn't be under-valued the way
|
|
|
|
|
|
Candidate A might undervalue it), but introduces a second calibration
|
|
|
|
|
|
constant ($-per-complexity-unit) with even less grounding than
|
|
|
|
|
|
`daily_rate` has.
|
|
|
|
|
|
|
2026-07-30 12:58:53 +02:00
|
|
|
|
**Adopted 2026-07-30** (`workplans/TREV-WP-0010-development-effort-calculator.md`
|
|
|
|
|
|
T01, maintainer-accepted): **Candidate A is the confirmed v0 formula.**
|
|
|
|
|
|
Candidate B remains recorded here as the considered alternative, not
|
|
|
|
|
|
adopted — revisit only if Candidate A's labor-cost anchor proves
|
|
|
|
|
|
systematically wrong once real repos produce real data, the same
|
|
|
|
|
|
provisional-but-decided treatment `specs/TargetDegenerationPolicyResearch.md`
|
|
|
|
|
|
gave the degeneration formula fork.
|
2026-07-30 01:29:13 +02:00
|
|
|
|
|
|
|
|
|
|
## 4. Target Multiple classification (the Nx factor)
|
|
|
|
|
|
|
|
|
|
|
|
Reuses `specs/TargetRevenueFrameworkCore.md` §1.4's existing guidance
|
|
|
|
|
|
table unchanged — this calculator does not redefine what 10x/100x/1000x
|
|
|
|
|
|
mean, it only needs a repeatable way to *classify* a given repo/Phase
|
|
|
|
|
|
against that existing table:
|
|
|
|
|
|
|
|
|
|
|
|
| Class | Multiple | Indicative interpretation (unchanged from framework core) |
|
|
|
|
|
|
|---|---:|---|
|
|
|
|
|
|
| Commons | 0x | Immediate permissive release or no development monetization target |
|
|
|
|
|
|
| Recovery | 1x | Direct development cost recovery |
|
|
|
|
|
|
| Incremental | 10x | Material enhancement of an existing use case |
|
|
|
|
|
|
| Product-defining | 100x | Significant commercial differentiator or new product capability |
|
|
|
|
|
|
| Platform-defining | 1000x | New platform, market, ecosystem, or foundational capability |
|
|
|
|
|
|
|
|
|
|
|
|
**v0 approach: human classification, calculator-assisted, not
|
|
|
|
|
|
calculator-decided.** The calculator can surface signals relevant to the
|
|
|
|
|
|
choice (e.g., "this repo's capability is consumed by N other repos" as a
|
|
|
|
|
|
platform-defining signal, drawing on state hub repo-dependency data where
|
|
|
|
|
|
available) but per `specs/TargetRevenueFrameworkCore.md` §1.4's own
|
|
|
|
|
|
framing ("represents a product and commercial hypothesis, not an
|
|
|
|
|
|
objectively measurable claim"), the actual classification choice remains
|
|
|
|
|
|
a human judgment call, not an automated output. This mirrors
|
|
|
|
|
|
`specs/PilotPhaseCandidateSurvey.md`'s own indicative classifications,
|
|
|
|
|
|
which were reasoned, not computed.
|
|
|
|
|
|
|
|
|
|
|
|
## 5. Worked illustration (not a real valuation)
|
|
|
|
|
|
|
|
|
|
|
|
Using `info-tech-canon`'s dry-run draft manifest
|
|
|
|
|
|
(`examples/pilot-candidates/info-tech-canon-service-surface/manifest.json`)
|
|
|
|
|
|
purely as a shape check, not a claim that these numbers are its real
|
|
|
|
|
|
value: 25 estimated effort days × 1000/day = 25,000, + 0 direct costs
|
|
|
|
|
|
(no real token-cost data pulled for this illustration) = 25,000 Estimated
|
|
|
|
|
|
Development Cost × 100 (Product-defining) = 2,500,000 — matching the
|
|
|
|
|
|
draft manifest's existing illustrative figures, which were hand-picked,
|
|
|
|
|
|
not calculator-derived. A real application of this calculator to
|
|
|
|
|
|
`info-tech-canon` (or any repo) is future work once §3's formula choice
|
|
|
|
|
|
is confirmed (workplan T02) and implemented (T03).
|
|
|
|
|
|
|
|
|
|
|
|
## 6. Non-goals
|
|
|
|
|
|
|
|
|
|
|
|
- Choosing Candidate A vs. B — a human decision, workplan T02.
|
|
|
|
|
|
- Implementing the actual calculator (data collection + formula code) —
|
|
|
|
|
|
workplan T03.
|
|
|
|
|
|
- Producing a real Initial Target for any specific repo — that requires
|
|
|
|
|
|
both this calculator's implementation and a human Target Multiple
|
|
|
|
|
|
classification decision per repo, neither done here.
|
|
|
|
|
|
- Replacing or reinterpreting `specs/TargetRevenueFrameworkCore.md` §1.4's
|
|
|
|
|
|
existing formula — this calculator feeds it, it does not compete with it.
|