Assistant: codex Assistant-Model: gpt-5.6-sol Assistant-Session: 01a02b6f-7db1-7222-918b-e813a6bda38d
128 lines
9.8 KiB
Markdown
128 lines
9.8 KiB
Markdown
# Intent–scope gap assessment — 2026-08-23
|
||
|
||
## Purpose and method
|
||
|
||
This assessment compares the commitments and prohibitions in `INTENT.md`
|
||
against the repository's source, tests, deployment files, completed workplans,
|
||
and the current railiance01 worker. It distinguishes implemented behavior from
|
||
historical proof and from the target architecture. It does not treat completed
|
||
workplan status as proof that every statement in `INTENT.md` is already true.
|
||
|
||
At the time of review there was no actionable repo work: the Custodian brief
|
||
listed no active workstream, every file-backed workplan was done or finished,
|
||
the repo had no matching human-needed task, and the two unread messages were a
|
||
Glas acknowledgement plus a superseded request already delivered in
|
||
`c633291`.
|
||
|
||
## Overall assessment
|
||
|
||
`rein-aharness` is a useful deployed queue consumer and bounded execution
|
||
adapter, but it only partially realizes the intended single shared,
|
||
multi-tenant agent runtime. The strongest implemented boundary is scheduled
|
||
work intake: Activity Core owns scheduling, while this repo claims and closes
|
||
execution rows. The strongest current governed-execution behavior is
|
||
fail-closed profile routing through Glas.
|
||
|
||
The largest gap is that the mature, currently usable paths are a mixture of a
|
||
host-local Claude runner and tenant-specific FI/Binky functions. The generic
|
||
profiled path cannot yet complete a current production session because its
|
||
sandbox execution/runtime/egress contract is unresolved upstream. The repo
|
||
therefore should not claim complete multi-tenant isolation, universal
|
||
credential custody, or comprehensive post-run policy verification.
|
||
|
||
## Intent mapping
|
||
|
||
| Intent commitment or boundary | Current evidence | Assessment |
|
||
| --- | --- | --- |
|
||
| Consume work; do not schedule it | `claim_loop.py` consumes Activity Core `ops_run` rows and contains no scheduler. | Met. |
|
||
| Keep blueprint authoring in kaizen-agentic | `persona.py` calls `kaizen-agentic schedule prepare`; this repo does not author blueprints. | Boundary met, but a missing/failed prepare silently falls back to no persona. |
|
||
| Let consuming repos declare instances | `manifest.py` validates `.kaizen/schedule.yml` runtime extensions and resolves policy for local agent sessions. | Partial: manifests affect local session policy, but profile-driven runs take their execution constellation from Activity Core/Glas and legacy routing is hard-coded here. |
|
||
| Provide one shared multi-tenant runtime | One railiance01 user service can map several repository slugs and processes rows sequentially. | Partial: shared process exists, but there is no tenant registry/API, per-tenant process boundary, repository lease, concurrency control, or safe multi-worker coordination. |
|
||
| Consume Activity Core tasks | Claim, heartbeat, complete, fail/reopen, label filtering, and configured backoff are implemented and tested. | Met for the REST queue contract. Issue-core remains legacy compatibility. |
|
||
| Run bounded sessions through llm-connect | The direct Claude adapter implements llm-connect interfaces with timeout/tool profile/budget hooks; structured executors call llm-connect HTTP. Profiled work delegates the rein/model choice through Glas. | Partial: Claude usage metadata is not demonstrated as accurate token accounting, and the current profiled production route stops before dispatch. |
|
||
| Enforce named tool profiles and grants | Unknown profiles refuse; Claude receives a named `--allowedTools` set and a cwd-pinned prompt. Glas profiles are authoritative for profiled rows. | Partial: lane is advisory, direct runs are host-local, and there is no post-run proof of network/push absence, changed-file scope, or clean repository state. |
|
||
| Be the only credential holder and policy enforcement point | Mail scan acquires IMAP credentials through an AppRole lane; worker/provider tokens arrive through host environment; Glas and sand-boxer own profile and isolation enforcement. | Not literally met. Credential and enforcement ownership is distributed across the worker, llm-connect, Glas, sand-boxer, OpenBao, and deployment environment. `INTENT.md` should eventually describe the governed chain rather than a single holder if this architecture is retained. |
|
||
| Verify a local commit | `runner.py` compares `HEAD` before and after; structured paths stage named output files and commit them. | Partial: a `HEAD` delta is not full verification. There is no repo lock, clean-tree gate, parent/branch check, changed-file allow-list, author/provenance check, or remote/push audit. Metrics are written after the accepted `HEAD` snapshot. |
|
||
| Report to State Hub and kaizen metrics | Progress, tool events, token events, task close, JSONL execution metrics, and summaries exist. | Implemented best-effort. Reporting failures are swallowed, so delivery is not part of success acceptance and there is no durable outbox here. |
|
||
| Remain generic, not a per-project logic container | The profiled queue adapter and local `TaskSpec` runner are generic. | Diverged in the coexistence path: FI research, Binky briefs, mailbox configuration, event names, label matching, and default FI push behavior are tenant-specific code. |
|
||
| Not become an LLM abstraction | Provider selection and HTTP abstraction remain in llm-connect; versioned constellation selection remains in Glas. | Mostly met. The Claude-Code specialization is appropriate rein behavior, though it directly constructs CLI invocation policy. |
|
||
| Not become a state store | Durable run/task history is external; only repo-owned `.kaizen/metrics` is written locally. | Met. Local sand-boxer state belongs to the sibling package, not this repo. |
|
||
| Be workstation-independent | The claim worker is active on railiance01 and accesses cluster services without a long-lived port-forward. | Met for queue consumption and legacy structured work. Not yet met for a current isolated profiled Claude run; the Kubernetes Deployment is also only a sleeper placeholder. |
|
||
|
||
## Gaps requiring future work records
|
||
|
||
No new rein-aharness workplan was opened by this documentation audit. The
|
||
following are gaps, not implicitly authorized implementation tasks. They should
|
||
be triaged into explicit work records before code changes.
|
||
|
||
### 1. Complete the generic profiled execution chain
|
||
|
||
Current production evidence resolves the requested profile and creates and
|
||
destroys the sandbox, but fails at Glas `session_start`. Existing residual
|
||
`GLAS-IN-0002` owns executable bwrap reachability, runtime mounting, and model
|
||
egress/credential delivery. After that owner contract lands, this repo needs a
|
||
fresh bounded Activity Core proof showing dispatch, an accepted commit, full
|
||
normalized evidence, and teardown under the current contracts.
|
||
|
||
### 2. Replace weak commit acceptance with repository-boundary verification
|
||
|
||
Before a direct run, record a clean/dirty baseline and acquire a per-repository
|
||
lease. After execution, verify the new commit's parent, changed paths, author
|
||
and task correlation, working-tree state, and remote/push state. Decide how
|
||
`.kaizen/metrics` joins the commit or moves to a separate evidence write so a
|
||
successful run does not create unexplained dirtiness after verification.
|
||
|
||
### 3. Retire tenant-specific runtime code
|
||
|
||
Move FI/Binky prompt formats, paths, approach matching, completion event names,
|
||
mailbox configuration, and push policy into tenant-owned blueprints/manifests
|
||
or explicit external capabilities. Until migration, label these functions as a
|
||
legacy coexistence surface and do not add another tenant by adding another
|
||
branch to `approaches.py`.
|
||
|
||
### 4. Reconcile credential and enforcement ownership language
|
||
|
||
The implementation is a governed chain, not a single credential holder. Define
|
||
which component obtains, injects, observes, and revokes each credential class;
|
||
which component enforces tool, filesystem, network, model, and repository
|
||
policy; and which evidence proves those controls. Then update `INTENT.md` only
|
||
through an explicit architecture decision if “only holder/only enforcement
|
||
point” is no longer the desired topology.
|
||
|
||
### 5. Make deployment match the runtime claim
|
||
|
||
The active worker is the railiance01 user service. The Kubernetes Deployment
|
||
only sleeps and the image does not package the optional Glas/sand-boxer
|
||
constellation. Either make Kubernetes a real supported claim worker with
|
||
workspace, credentials, runtime, and shutdown semantics, or clearly retire it
|
||
to an image-smoke artifact. Add a repeatable host prerequisite check for the
|
||
bubblewrap/AppArmor policy if the local profile remains supported.
|
||
|
||
### 6. Strengthen durability, concurrency, and conformance
|
||
|
||
Add repository locking or owner leases, defined recovery after lease loss or
|
||
close failure, durable/reporting reconciliation, package-level contract tests
|
||
that do not silently skip optional runtimes, and CI/release gates for the
|
||
manifest/profile/rein contracts. A single sequential process is safe only
|
||
while deployment topology remains singular and target repos do not overlap.
|
||
|
||
## Recommended sequencing
|
||
|
||
1. Let `GLAS-IN-0002` deliver an executable sandbox contract and re-prove one
|
||
current profiled run end to end.
|
||
2. Harden repository acceptance and metrics atomicity before increasing worker
|
||
concurrency or onboarding more tenants.
|
||
3. Extract FI/Binky compatibility logic into owned declarative capabilities.
|
||
4. Reconcile credential/policy ownership in an ADR, then align deployment and
|
||
`INTENT.md` terminology with that decision.
|
||
5. Add CI and conformance gates before declaring the rein generally
|
||
multi-tenant.
|
||
|
||
## Conclusion
|
||
|
||
The implemented product is narrower than the former `SCOPE.md` claimed: a
|
||
deployed sequential Activity Core consumer with a generic queue-to-Glas
|
||
adapter, a host-local Claude task runner, and several proven legacy structured
|
||
executors. That is a coherent and useful v0.1 capability. It is not yet the
|
||
fully isolated, credential-centralizing, tenant-neutral runtime described by
|
||
`INTENT.md`; the revised `SCOPE.md` now makes that distinction explicit.
|