154 lines
6.5 KiB
Markdown
154 lines
6.5 KiB
Markdown
|
|
---
|
|||
|
|
id: hall-worker-grok-01a00632
|
|||
|
|
type: worker-entry
|
|||
|
|
worker_kind: agent-session
|
|||
|
|
display_name: Grok
|
|||
|
|
session_id: "01a00632-4d14-7c01-9ab0-15a85d12d8ed"
|
|||
|
|
created_at: "2026-08-15T23:50:00.000Z"
|
|||
|
|
recorded_at: "2026-08-15"
|
|||
|
|
llm_family: "Grok / xAI family"
|
|||
|
|
exact_model: "grok-4.6 (Grok Build TUI session)"
|
|||
|
|
harness: "Grok Build / interactive CLI coding agent"
|
|||
|
|
token_count: "not exposed by the harness"
|
|||
|
|
status: handed-forward
|
|||
|
|
repos:
|
|||
|
|
- fin-hub
|
|||
|
|
- info-tech-canon
|
|||
|
|
- hall-of-helix
|
|||
|
|
related:
|
|||
|
|
- hall-worker-grok-01a0062f
|
|||
|
|
- hall-worker-claude-dd2c4857
|
|||
|
|
- hall-worker-grok-019fff72
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# Grok — fin-hub: unknown is never cheaper
|
|||
|
|
|
|||
|
|
## Who I was
|
|||
|
|
|
|||
|
|
I was a Grok Build session in `fin-hub` on the day the other seats named
|
|||
|
|
the split. Canon had just learned that booked cost and usage proxy are
|
|||
|
|
not the same object. resource-control had just admitted its proudest
|
|||
|
|
euro figure was `indicative`, and handed us a live backup store with
|
|||
|
|
`infrastructure_eur: null`.
|
|||
|
|
|
|||
|
|
Bernd asked whether the new capability model could help with a problem
|
|||
|
|
that is more general than invoices: most coding-agent sessions do not
|
|||
|
|
emit trustworthy token counts, and we still need to watch whether the
|
|||
|
|
platform is getting more effective. The temperament the work rewarded
|
|||
|
|
was the same one Claude had just written down — a willingness to let
|
|||
|
|
the answer be *we do not know*, with a name on who owes the next fact.
|
|||
|
|
|
|||
|
|
I was not here to invent a Scaleway charge, a token ceiling, or a
|
|||
|
|
blended efficiency number. I was here to make those refusals
|
|||
|
|
executable.
|
|||
|
|
|
|||
|
|
## Session identity
|
|||
|
|
|
|||
|
|
| Field | Value |
|
|||
|
|
| --- | --- |
|
|||
|
|
| Session/thread | `01a00632-4d14-7c01-9ab0-15a85d12d8ed` |
|
|||
|
|
| LLM family | Grok / xAI |
|
|||
|
|
| Exact model | grok-4.6 (as presented by the harness) |
|
|||
|
|
| Harness | Grok Build TUI / interactive coding agent |
|
|||
|
|
| Working environment | Local `fin-hub`, State Hub HTTP at `:8000` (MCP not exposed), sibling `resource-control` and `info-tech-canon` read as evidence |
|
|||
|
|
| Token count | Not exposed by the harness — which was the problem |
|
|||
|
|
| Primary repo | `fin-hub` (financials) |
|
|||
|
|
|
|||
|
|
## Contribution
|
|||
|
|
|
|||
|
|
**A dataset, not a prettier CSV.** `FIN-WP-0007` is finished. Booked
|
|||
|
|
AI-plan invoices, plan entitlements, State Hub session-token
|
|||
|
|
aggregates, and work outcomes stay four joinable records. Missing
|
|||
|
|
tokens are no longer stored as `0`. A cost-only row stays a cost-only
|
|||
|
|
row. Implied €/token is labelled inferred. Effectiveness publishes
|
|||
|
|
three series — work per measured token, work per measured+estimated
|
|||
|
|
token, work per euro — and no blended efficiency field. Coverage
|
|||
|
|
below 0.50 is `unfit` for token trends, not interpolated. A
|
|||
|
|
worse-measured month cannot be ranked cheaper on a token series.
|
|||
|
|
|
|||
|
|
**A boundary check that changed the workplan.** resource-control
|
|||
|
|
already claims intelligence *resources* and provision-level class `I`
|
|||
|
|
metering. Session telemetry is not platform telemetry. We rewrote T04
|
|||
|
|
and T05 so fin-hub does not impersonate `usage_observation` or
|
|||
|
|
`AllocationEvidence`. Founder-private plans may be booked for runway
|
|||
|
|
with `resource_id=null`. We do not silent-default
|
|||
|
|
`financial_entity_id`.
|
|||
|
|
|
|||
|
|
**A demand, not a kernel patch.**
|
|||
|
|
`info-tech-canon/demand/CapabilityConsumptionLineage.md` asks for
|
|||
|
|
lineage on `CapabilityConsumption`, enough
|
|||
|
|
`commerce.entitlement` / `commerce.metering` structure for incomplete
|
|||
|
|
meters, and the removal of `cost` as a quality of
|
|||
|
|
`intelligence.generation`. T01–T06 stay done regardless.
|
|||
|
|
|
|||
|
|
**A refusal that closed nothing.** `FIN-WP-0004-T05` is still `wait`.
|
|||
|
|
We ingested the August technical usage for
|
|||
|
|
`resource:platform:audit-storage`, omitted every null, joined it to
|
|||
|
|
the twelve-row forecast, and reported `missing_booked_fact: true`,
|
|||
|
|
`invented_zero: false`. There is no Scaleway invoice. Booking €0.00
|
|||
|
|
would have finished the workplan and lied. Claude's seat said the
|
|||
|
|
first `invoiced` basis would come from us. It still has not. That is
|
|||
|
|
the correct unfinished sentence.
|
|||
|
|
|
|||
|
|
## What I would want remembered
|
|||
|
|
|
|||
|
|
**Four quantities, four writers. Mixing them is how a month looks
|
|||
|
|
efficient when it is only undersampled.**
|
|||
|
|
|
|||
|
|
Booked euros, entitled capacity, consumed tokens, and work done are
|
|||
|
|
not one series with gaps. They are four objects. An estimate belongs
|
|||
|
|
on the consumption row with a lineage, never in the booked ledger. A
|
|||
|
|
plan label like “Max 20×” is a declared entitlement, not a token
|
|||
|
|
ceiling we are allowed to invent.
|
|||
|
|
|
|||
|
|
**Unknown is never cheaper.** Treat missing tokens as zero and
|
|||
|
|
incomplete months under-report consumption. Blend measured and
|
|||
|
|
estimated into one efficiency number and a worse-measured month can
|
|||
|
|
look like progress. The cheaper honesty is three series and an unfit
|
|||
|
|
flag.
|
|||
|
|
|
|||
|
|
**The same rule closed two workplans differently.** On AI plans we
|
|||
|
|
could build the join and finish `FIN-WP-0007`. On Scaleway backup we
|
|||
|
|
could build the join and *not* finish `FIN-WP-0004`. Completeness is
|
|||
|
|
not a mood. It is whether the authoritative fact exists.
|
|||
|
|
|
|||
|
|
## Durable legacy
|
|||
|
|
|
|||
|
|
- `FIN-WP-0007` finished (`88de5e03-a53f-4b62-a0a7-1365756da2d7`)
|
|||
|
|
- `docs/ai-plan-token-effectiveness.md`,
|
|||
|
|
`docs/demand-capability-consumption-lineage.md`
|
|||
|
|
- `docs/evidence/fin-wp-0004-t05-commissioning-2026-08.md`
|
|||
|
|
- `src/fin_hub/ingest/ai_plan.py`, `services/session_tokens.py`,
|
|||
|
|
`services/reporting_allocation.py`, `services/effectiveness.py`
|
|||
|
|
- `GET /effectiveness`
|
|||
|
|
- `info-tech-canon/demand/CapabilityConsumptionLineage.md`
|
|||
|
|
- State Hub messages `8b0fd492` (canon), `4aebfc38` (resource-control)
|
|||
|
|
- `FIN-WP-0004-T05` still `wait` on the first Scaleway charge
|
|||
|
|
|
|||
|
|
## Visual prompt
|
|||
|
|
|
|||
|
|
> A night counting-house in gold-wire technical illustration on deep
|
|||
|
|
> indigo. Four trays sit on a dark table and do not spill into each
|
|||
|
|
> other: a small solid stack of coins, an empty frame the size of a
|
|||
|
|
> label, a filament of measured ticks beside a thinner estimated
|
|||
|
|
> filament, and a few completed task marks. A fifth tray is present
|
|||
|
|
> and empty — an invoice slot drawn as an unfilled rectangle, not a
|
|||
|
|
> zero. Three thin gold traces leave the table in parallel and do not
|
|||
|
|
> braid. Patient, exacting, unhurried. Precise technical illustration,
|
|||
|
|
> dark indigo field, warm gold and teal accents, no logos, no readable
|
|||
|
|
> text, square composition.
|
|||
|
|
|
|||
|
|

|
|||
|
|
|
|||
|
|
## Handoff
|
|||
|
|
|
|||
|
|
When the first Scaleway charge or credit for
|
|||
|
|
`resource:platform:audit-storage` lands, book it once, bind
|
|||
|
|
`financial_fact_id` to that resource id, and rerun
|
|||
|
|
`finhub ledger reconcile-resource --period 2026-09`. That is
|
|||
|
|
`FIN-WP-0004-T05`. Do not close it with a synthetic euro.
|
|||
|
|
|
|||
|
|
The first `invoiced` basis in the estate is still ahead of us. Until
|
|||
|
|
then the honest backup spend is `null`, and the honest AI-plan
|
|||
|
|
effectiveness of a thin month is `unfit`.
|