Adds since/until date-range scoping to cluster_commit_hours() and workplan_task_counts() (threaded through calculate_target_basis()), needed whenever a candidate is one bounded workplan within a repo whose overall history spans much more (net-kingdom, railiance-apps) rather than the whole repo being the candidate (vergabe-teilnahme, info-tech-canon). Fixes a real bug found along the way: workplan_task_counts() only scanned the top level of workplans/, missing net-kingdom's workplans/archived/ convention entirely - silently reported zero finished workplans for NK-WP-0002, which lives there. Fixed to scan recursively; added a regression test. Updates all three draft pilot-candidate manifests with calculator- derived target_basis/initial_target values, replacing the hand-picked placeholders: net-kingdom-local-identity: 200,000 -> 10,000 EUR (floor + sanity warnings) railiance-vergabe-teilnahme: 3,500,000 -> 648,800 EUR (no warnings) info-tech-canon-service-surface: 2,500,000 -> 141,800 EUR (sanity warning) history/260730-EffortCalculator-CandidateApplication.md records full derivation, warnings, and the judgment calls made explicit rather than silently picked (date-scoping windows; measuring vergabe-teilnahme's own repo rather than railiance-apps' deployment-only wiring, with both figures shown). Still draft/non-binding - WP-0008-T05 unaffected. 5 new tests (20 -> now covering since/until scoping and the archived-subdirectory fix). Full suite: 84 passing offline.
9.5 KiB
| id | type | title | domain | repo | status | owner | topic_slug | created | updated | state_hub_workstream_id |
|---|---|---|---|---|---|---|---|---|---|---|
| TREV-WP-0010 | workplan | Development Effort Calculator | infotech | target-revenue | finished | claude | infotech | 2026-07-30 | 2026-07-30 | 2cabf156-dbdf-4e46-a213-fb76800702df |
Development Effort Calculator
Turns measurable repo activity (human interaction time, workplan/task
volume, file/line complexity, AI token cost) into target_basis values
(specs/PhaseManifestSpecification.md), which combined with a human
Target Multiple classification (10x/100x/1000x) produces a proposed
Initial Target via the framework's own existing formula
(specs/TargetRevenueFrameworkCore.md §1.4) — not a new formula. This
calculator is a value-input tool feeding that existing formula, not a
competing valuation methodology.
Split from the former combined TREV-WP-0009-control-plane-and-effort-calculator.md
(2026-07-30) — this deliverable has its own formula decision and
implementation arc, independent of the Target Revenue Control Plane
(workplans/TREV-WP-0009-target-revenue-control-plane.md), even though
the Control Plane's Phase-registration flow is expected to consume this
calculator's output once both exist.
Does not include: declaring any real Phase for any repo, or choosing
a repo's Target Multiple classification unilaterally — per
specs/DevelopmentEffortCalculatorConcept.md §4, that classification
remains a human judgment call the calculator can inform but not make.
Concept: specs/DevelopmentEffortCalculatorConcept.md.
Formula decision (human gate)
id: TREV-WP-0010-T01
status: done
priority: high
human_accept_required: true
human_accepted_by: Bernd
human_accepted_at: "2026-07-30"
state_hub_task_id: "a3e1acbf-8659-4745-b595-ba8c2728f3b5"
specs/DevelopmentEffortCalculatorConcept.md §3 presents two combination
strategies without choosing between them:
- Candidate A (labor-cost-anchored, recommended v0 default) — human
time drives
estimated_effort_days/daily_ratedirectly; AI token cost becomesapproved_direct_costs(a real, metered direct cost); workplan/task-volume and repo-size metrics are a sanity check/adjustment on the human-time estimate, not a separate dollar figure. - Candidate B (composite-index-anchored) — all four metric families scored into one weighted complexity index, converted to dollars via a declared $-per-unit rate.
Confirm one as the adopted v0 approach (or propose a refinement), per the
same human-gate pattern already used for the degeneration policy formula
(workplans/TREV-WP-0007-degeneration-policy-and-canonical-profiles.md T02).
Agents may prepare a recommendation and leave this todo.
Accepted 2026-07-30 by the maintainer (Bernd): Candidate A
(labor-cost-anchored) confirmed as the adopted v0 formula.
estimated_effort_days and daily_rate are driven directly by
commit-clustered human interaction time; AI token cost
(get_token_summary) becomes approved_direct_costs; workplan/task-
volume and file/line-size metrics serve only as a sanity check/adjustment
on the human-time estimate, never their own dollar figure; the Target
Multiple remains a human classification, not computed. T02 implements
this formula, not Candidate B or a hybrid.
Implement the calculator
id: TREV-WP-0010-T02
status: done
priority: high
state_hub_task_id: "b5fb46ea-b737-495f-82b8-5e10f0032cb5"
Using T01's confirmed formula, implement data collection for each metric
family and the combination formula itself, producing target_basis
field values ready to drop into a Phase Manifest:
- Human interaction time — v0 default: commit-timestamp clustering (session-gap grouping over git history), explicitly a floor estimate, not a true count (concept §2a).
- Workplan/task volume — counts via state hub
list_tasks/list_workplansscoped to the repo, or directworkplans/directory inspection where hub coverage is thin (concept §2b). - File/line complexity — raw counts with generated/vendored-path exclusion (concept §2c's flagged distortion: dependency trees, lockfiles, and generated files must not inflate this signal).
- AI token cost —
get_token_summary(scope="repo", id=...), the strongest, least speculative input since it is a real metered cost, not a proxy (concept §2d).
Output must show its work (which inputs produced which numbers) — a
black-box dollar figure with no visible derivation would undermine the
"transparent, non-gameable" goal (specs/TargetRevenueLicenseConcept.md
§4.6) this whole framework is built around.
Result: src/target_revenue/effort_calculator.py implemented, pure/
offline (no network or state-hub dependency — list_tasks/list_workplans
was superseded by direct workplans/ directory parsing so the module
works uniformly on any repo using this repo's own workplan convention,
without requiring hub connectivity). cluster_commit_hours() does
session-gap clustering over git log --all --format=%at, with a 15-minute
floor for single-commit sessions. workplan_task_counts() parses
frontmatter status: fields and this repo's triple-backtick task-block
convention. repo_size_metrics()
excludes generated/vendored path components
(node_modules, .venv, __pycache__, etc.). token_cost_usd() takes
caller-supplied token counts (e.g. from get_token_summary) rather than
fetching them itself. estimate_target_basis() combines these per
Candidate A and returns a TargetBasisEstimate with a derivation dict
(showing every input value) and a warnings list.
1-day manual-work floor, as requested: any raw commit-clustered
estimate below MANUAL_EFFORT_FLOOR_DAYS = 1.0 is floored to 1.0 day
(never reported smaller) and flagged with an explicit warning explaining
this is very likely a measurement gap — commit-clustering is a floor
estimate by design — and should usually be compensated for by a manual
override rather than trusted at face value. A second, independent
sanity-check warning fires when a repo's finished-workplan/task volume is
substantial but the raw time estimate is still low, catching the case
where the floor itself wasn't triggered but the estimate still looks
implausible (demonstrated live: running the calculator against
target-revenue's own history — 7 finished workplans, 57 finished
tasks — correctly flagged its 2.38-day raw estimate as under-counted,
above the 1-day floor but still clearly too low for that much finished
work).
scripts/effort_calculator_cli.py — a CLI wrapper printing the estimate
as JSON, following the same offline-first, no-Phase-declaration pattern
as scripts/trf_onboard.py. tests/test_effort_calculator.py (15 tests,
deterministic — builds throwaway git repos and directory fixtures under
tmp_path rather than depending on any real repo's changing state)
covers commit clustering (empty/single-commit/multi-session), workplan/
task parsing, size-metric exclusion, token-cost pricing, the floor-and-
warning behavior, the independent sanity-check warning, and an
end-to-end smoke test. Full suite: 79 passing offline (64 + 15 new), no
new hard dependency (stdlib + existing pathlib/subprocess/re only).
Apply calculator to real candidate repos
id: TREV-WP-0010-T03
status: done
priority: medium
state_hub_task_id: "74a3db97-1b66-4514-895a-bb9c385008aa"
Once T02 is implemented, run it against the real candidates already
identified (specs/PilotPhaseCandidateSurvey.md: net-kingdom's
NK-WP-0002, railiance-apps's vergabe-teilnahme, and
info-tech-canon's service surface) to produce proposed,
calculator-derived target_basis values, replacing the hand-picked
illustrative figures currently in examples/pilot-candidates/*/manifest.json.
This remains draft/non-binding — updating a draft manifest's numbers is
not a Phase declaration, and does not touch
workplans/TREV-WP-0008-governance-and-pilot-rollout.md T05's own gate.
Result: All three draft manifests updated with calculator-derived
target_basis/initial_target values; full derivation and warnings in
history/260730-EffortCalculator-CandidateApplication.md.
net-kingdom-local-identity: date-scoped to NK-WP-0002's own window
(net-kingdom has since grown far beyond that one workplan) — hit the
1-day floor and the sanity-check warning; Initial Target 200,000 → 10,000
EUR. railiance-vergabe-teilnahme: measured against the vergabe-teilnahme
application repo itself (not railiance-apps, which only holds
deployment wiring — a judgment call recorded explicitly, with the
alternative repo's figure shown too) — no warnings; Initial Target
3,500,000 → 648,800 EUR. info-tech-canon-service-surface: whole-repo
history — sanity-check warning fired (12 finished workplans/51 tasks
against 1.42 raw days) even though above the 1-day floor; Initial Target
2,500,000 → 141,800 EUR. Every calculator-derived figure came out
materially lower than the earlier hand-picked placeholders, as expected
since those were never reasoned from a real measurement.
Bug found and fixed along the way: workplan_task_counts() only
scanned the top level of workplans/, missing net-kingdom's
workplans/archived/ convention entirely (silently reporting zero
finished workplans for NK-WP-0002, which lives there). Fixed to scan
recursively; added a regression test
(test_workplan_task_counts_scans_archived_subdirectory). Also added
since/until date-range scoping to both cluster_commit_hours() and
workplan_task_counts() (threaded through calculate_target_basis()),
needed for this task and reusable for any future candidate that's one
workplan within a larger repo.