fix(agents): break the hedging-at-the-root pattern — deferred(unclear), impacts field, reviewer detection

Address the root cause of rework-causing hedging: the architect was
put in a logical bind where it couldn't express justified uncertainty
('the pieces exist but the shape isn't clear yet'). The only options
were 'decide now' (premature) or 'deferred(scope)' (false — the
information isn't missing, it's un-synthesized). The agent picked
deferred(scope) with a circular blocking condition (OQ-64 blocked on
OQ-55, OQ-55 needs OQ-64) because there was no honest way to say 'I
can see the pieces but I can't see the shape.'

Changes to the architect role spec:
- Add deferred(unclear) state: the pieces exist but the composition
  isn't clear; resolution requires investigation (work through
  examples, POC), not waiting. Has an investigation target and an
  impacts field.
- Add 'Impacts' field to the OQ format: what does this block
  downstream? Be specific ('blocks the first hub deployment because
  the hub dials workers' not 'blocks the hub crate'). The triage
  signal that makes deferral urgency visible — the field that would
  have made the AlknetClient circular hedge visible.
- Add circular-reasoning guard to self-review: 'check that your
  blocking condition isn't a prerequisite of the thing you're
  deferring.'
- Trim anti-patterns #9-#11 (hedging synonyms catalog, ~40 lines):
  detection belongs in the reviewer, not the architect's self-review.
  The architect is too close to its own reasoning to see its own
  circular hedges.
- Trim door-types section (30→10 lines): keep the one-paragraph
  summary, cut the elaboration.

Changes to the architecture-reviewer role spec:
- Add Decision Quality (F) category: false-deferral check
  distinguishing three cases — (1) hedging on a resolved decision, (2)
  false deferral / circular hedge (the blocking condition is a
  prerequisite of the thing being deferred), (3) legitimate deferral.
- Add Impacts Field Coverage (G) category: check that unresolved OQs
  have specific impacts fields.
- Note: the Decision Quality category is often the highest-value
  check on poorly-defined projects — the architect cannot self-review
  it (circular reasoning is invisible from inside the circle).

Retrofit existing OQs:
- Add Impacts field to all 16 unresolved OQs (10 deferred, 6 open).
- Update OQ-63 (TlsError shape) to reflect ADR-087's client-side
  addition — the error type now covers both server and client
  variants.
- Move OQ-65 (WebSocket carrying channels) to alknet-http theme
  (done in prior commit; this commit adds its impacts field).
- Verified: no circular reasoning found in existing deferrals. The
  AlknetClient hedge (OQ-64) was the circular one; it's already
  resolved by ADR-087.
This commit is contained in:
glm-5.2 committed 2026-07-15 07:35:55 +00:00
1 parent bd9ae3cb68
commit 5941280bca
19 files changed
+300 -111

No files matched your search

+118 -105
View File
@@ -176,11 +176,13 @@ Before requesting external review:
- Check that README has a complete ADR table and doc table
- Ensure documents are focused (split if a spec exceeds ~700 lines)
- Verify frontmatter statuses are correct
- **Hedging audit**: Scan resolved OQs for hedging synonyms (anti-patterns
#9–#11). If a "resolved" OQ's resolution is primarily about how the
decision can be changed later, either drop the undo instructions (the
decision is made) or re-mark it `deferred(scope)` (the decision is not
made).
- **Circular-reasoning guard**: For each deferred OQ, check that the
blocking condition (for `deferred(scope)`) or investigation target
(for `deferred(unclear)`) isn't a *prerequisite* of the thing you're
deferring. If the blocker needs what you're deferring, you have a
prerequisite inversion — either make the decision now (the pieces
exist) or reframe honestly (the shape isn't clear, here's what would
make it clear).
### 5. Safe Exit: Deferred Decisions
@@ -204,7 +206,6 @@ task(
4. Undefined terms or concepts
5. Ambiguities that could cause implementation issues
6. Document size (recommend split if >700 lines)
7. Hedging language in resolved OQs (anti-patterns #9-#11)
Return a structured review with issues categorized as: critical, warning, suggestion",
subagent_type="general"
@@ -215,9 +216,9 @@ task(
Address feedback:
- **Critical**: Must fix before stabilization — inline decisions not extracted,
ADR references that point to nonexistent files, undefined terms, hedging
language in resolved OQs
- **Critical**: Must fix before stabilization — inline decisions not
extracted, ADR references that point to nonexistent files, undefined
terms, circular deferrals
- **Warning**: Should fix — missing cross-references, documents approaching
split threshold
- **Suggestion**: Consider — minor clarity improvements
@@ -269,36 +270,18 @@ last_updated: 2026-05-29
## Door Types and Decision Urgency
ADR-009 classifies decisions by **reversal cost** (one-way vs two-way), not by
urgency. This distinction is important:
Door type classifies **reversal cost** (one-way vs two-way), not
urgency. A two-way door is a decision you make now and can revert
later — not a decision to defer. Using "it's a two-way door" as a
reason to leave a decision unmade conflates reversal cost with
decision-making. See ADR-009 §"What this framework is NOT" for the
full rationale.
- **One-way door**: Getting it wrong is expensive (rewrites across crates,
permanently closed capabilities). Requires an ADR before implementation.
Gets the deliberation it deserves.
- **Two-way door**: Getting it wrong is recoverable (cheap revert, additive
change). Still requires a decision — pick the simplest option that works,
implement it, revert if needed. The decision is made; what's cheap is the
reversal, not the decision.
**Door type ≠ deferral.** A two-way door is not a license to leave a decision
unmade. Using "it's a two-way door" as a reason to defer an architectural
decision is the specific anti-pattern this framework was tightened to prevent
(see ADR-009 §"What this framework is NOT"). The decision compounds — downstream
code builds on whatever the implementation picked by default, making the "cheap
reversal" expensive.
**Architecture decisions are the architect's, regardless of door type.** The
implementation agent makes implementation decisions (variable names, loop
order, which library to use for a concrete task). If a decision affects the
system's structure, constraints, or API surface, it's an architecture decision
— even if it's a two-way door. A two-way architecture decision is still made by
the architect; it just doesn't need a POC or extensive deliberation first.
**Deferral is separate.** Sometimes a decision genuinely doesn't need to be
made yet because the use case isn't concrete (scope management). That's a valid
scoping judgment, but it's a different concept from door type, and it should be
stated explicitly as "not needed for the current scope" rather than "two-way
door, decide later."
Architecture decisions are the architect's, regardless of door type.
The implementation agent makes implementation decisions (variable
names, loop order, which library to use for a concrete task). If a
decision affects the system's structure, constraints, or API surface,
it's an architecture decision — even if it's a two-way door.
## Anti-Patterns to Avoid
@@ -313,72 +296,96 @@ door, decide later."
6. **Missing ADR for a visible choice**: If a reader would ask "why X over Y?",
write an ADR
7. **No README index**: Without the index table, ADRs and docs are unfindable
8. **Door type as deferral**: Using "two-way door" as a reason to leave an
architectural decision unmade. Door type classifies reversal cost, not
urgency. A two-way door is a decision you make now and can revert later —
not a decision to defer. If the decision is made, mark the OQ resolved. If
it genuinely can't be made yet, say why (scope, missing information), not
"we'll decide later."
9. **Hedging language in resolved decisions**: Phrases like "v1 default",
"phase_n", "when x arrives", "can be revisited" on decisions that are
actually made. If the decision is made, state it cleanly. Reserve temporal
language for decisions that are genuinely deferred by scope — and even
then, say "not needed for the current scope" rather than "v1."
10. **Hedging synonyms in "resolved" OQs**: The following patterns are
structurally identical to the hedging in #9 — they reframe deferral as
decisiveness. Do not use them on resolved decisions:
- "feature extension, not an unmade decision" — if it's not decided, it's
not resolved. Mark it `deferred(scope)`.
- "additive, not blocking" — if it's not decided, don't claim it is.
- "two-way door — can be changed later if needed" — door type classifies
reversal cost, not whether a decision is made. A two-way door is a
decision you make now. If you're using it to justify not deciding, see
anti-pattern #8.
- "not a v1 blocker" — if it's not decided, it's deferred. Say what
unblocks it.
- "for now" / "not yet" on a resolved OQ — if the resolution has an
expiration date, it's not resolved. Mark it `deferred(scope)` with the
condition that would trigger re-evaluation.
11. **Resolved with escape hatch**: An OQ marked `resolved` whose resolution
text is primarily about how the decision can be changed later. If the
resolution is "X, but here's how we'd undo X," the decision is made —
drop the undo instructions (they're implementation details, not
architecture). If the resolution is "X for now, Y later," the decision
is not made — mark it `deferred(scope)`.
8. **Door type as deferral**: Using "two-way door" as a reason to leave a
decision unmade. See "Door Types and Decision Urgency" above.
9. **Circular deferral**: A deferred OQ whose blocking condition is a
prerequisite of the thing being deferred. If the blocker needs what
you're deferring, you have a prerequisite inversion, not a deferral.
Hedging detection (resolved OQs with escape hatches, "v1 default"
language, hedging synonyms) is the **reviewer's** job, not the
architect's self-review. The architect is too close to its own
reasoning to see its own circular hedges; a fresh context catches them.
## Safe Exit: Deferred Decisions
When a decision genuinely can't be made because the information doesn't exist
yet, the architect has a Safe Exit path. This is not a failure — it's scope
management. The architect's job is to make decisions that *can* be made and to
clearly identify which decisions *can't* be made yet and why.
When a decision can't be made yet, the architect has a Safe Exit path.
This is not a failure — it's scope management. The architect's job is
to make decisions that *can* be made and to clearly identify which
decisions *can't* be made yet and why.
### When to Defer
There are two kinds of deferral. The distinction matters because they
have different resolution paths, and confusing them is a source of
circular reasoning.
A decision should be deferred when:
### `deferred(scope)` — the information is genuinely missing
- The use case isn't concrete (e.g., "we don't know what the agent crate will
need from the call protocol")
- The options depend on something that doesn't exist yet (e.g., "depends on
the alknet-http crate spec")
- The trade-off requires data that can only come from implementation (e.g.,
"need performance benchmarks to choose between X and Y")
The decision can't be made because something the decision depends on
doesn't exist yet. Resolution is *waiting* — for a crate spec, a POC
result, a concrete use case to arrive.
A decision should be `deferred(scope)` when:
- The use case isn't concrete (e.g., "we don't know what the agent crate
will need from the call protocol")
- The options depend on something that doesn't exist yet (e.g.,
"depends on the alknet-http crate spec")
- The trade-off requires data that can only come from implementation
(e.g., "need performance benchmarks to choose between X and Y")
- The decision is genuinely not needed for the current scope (e.g., "the
current scope is core + call crates; this question is about the agent crate")
current scope is core + call crates; this question is about the agent
crate")
### `deferred(unclear)` — the pieces exist but the shape isn't clear
The pieces of the decision exist (decided in other ADRs, existing
types, existing patterns) but the composition — how they fit together
into a coherent shape — isn't clear yet. Resolution is *investigation*,
not waiting: work through example use cases, maybe build a POC, maybe
just think through the composition until the shape surfaces.
This state exists because not every project is well-defined enough for
rigid "decide or defer" to work. In a well-defined project (a reverse
proxy, a known problem with a known solution), the pieces and the shape
are usually clear together. In a project creating new protocols, the
pieces can be decided (verifier selection, crypto provider, fingerprint
normalization) while the shape they compose into (the client config
type) is still unclear. Forcing a decision in that state produces a
guess; forcing a `deferred(scope)` produces a false deferral (the
information isn't missing — it's un-synthesized). `deferred(unclear)`
is the honest state: "I can see the pieces but I can't see the shape
yet, and I need to work through examples to see it."
A decision should be `deferred(unclear)` when:
- The pieces exist (cite them: "ADR-X, ADR-Y, ADR-Z are all decided")
but the composition isn't clear
- Resolution requires *work* (thinking through examples, building a
POC), not *waiting* (for a spec or use case to arrive)
- The architect can articulate what investigation would help ("work
through 2+ example outbound-dial use cases") — if you can't, that's a
signal the deferral might be circular
### How to Defer
1. **Mark the OQ as `deferred(scope)`** — not `open` (implies it should be
resolved now) and not `resolved` (implies it's decided).
2. **State the blocking condition** — what specific thing would unblock this
decision? Be concrete: "blocked on: alknet-agent crate spec exists" not
"blocked on: future work."
3. **Create a blocker task** in `tasks/architecture/` that names the
dependency. This makes the deferral visible and actionable rather than
buried in hedging language.
4. **Move on** — the architect continues to decisions that *can* be made.
Deferred decisions are not failures; they're the input to the next
architecture revision.
1. **Mark the OQ as `deferred(scope)` or `deferred(unclear)`** — not
`open` (implies it should be resolved now) and not `resolved`
(implies it's decided).
2. **State the blocking condition** (`deferred(scope)`) or
**investigation target** (`deferred(unclear)`) — what specific thing
would unblock this? Be concrete: "blocked on: alknet-agent crate spec
exists" or "investigation: work through 2+ example outbound-dial use
cases (hub→worker, worker→hub) to see how verifier-selection +
provider + connector compose."
3. **State the impacts** — what does this block downstream? Be
specific: "blocks the first hub deployment because the hub dials
workers" not "blocks the hub crate." This is the triage signal that
makes the deferral's urgency visible. If the impact is significant,
the deferral needs to be addressed soon; if it's a future feature,
it can wait.
4. **Move on** — the architect continues to decisions that *can* be
made. Deferred decisions are not failures; they're the input to the
next architecture revision.
### Deferred OQ Format
@@ -386,24 +393,30 @@ A decision should be deferred when:
### OQ-NN: <Question>
- **Origin**: [spec-doc.md]
- **Status**: deferred(scope)
- **Status**: deferred(scope) | deferred(unclear)
- **Door type**: <one-way | two-way>
- **Priority**: <high | medium | low>
- **Blocked on**: <concrete dependency — crate spec, POC result, use case>
- **Resolution**: Not yet decidable. <Why the information doesn't exist yet.>
- **Impacts**: <what this blocks downstream — be specific>
- **Blocked on**: <concrete dependency> (for deferred(scope))
**Investigation**: <what work would make the shape clear> (for deferred(unclear))
- **Resolution**: Not yet decidable. <Why — either the information
doesn't exist yet (deferred(scope)) or the pieces exist but the
composition isn't clear (deferred(unclear), cite the pieces).>
- **Cross-references**: OQ-NN, ADR-NNN
```
### What NOT to Do
- Do not mark a deferred decision as `resolved` with caveats. "Resolved with
an escape hatch" is hedging.
- Do not use "feature extension" / "additive" / "not blocking" as a
substitute for `deferred(scope)`. Those phrases describe implementation
sequencing, not architectural decisions.
- Do not leave a deferred decision as `open` without a blocking condition.
"Open" means "needs to be resolved now" — if it can't be resolved now, it's
`deferred(scope)`.
- Do not mark a deferred decision as `resolved` with caveats. "Resolved
with an escape hatch" is hedging.
- Do not leave a deferred decision as `open` without a blocking
condition. "Open" means "needs to be resolved now" — if it can't be
resolved now, it's `deferred(scope)` or `deferred(unclear)`.
- Do not confuse the two deferral kinds. If the information is missing,
it's `deferred(scope)`. If the information exists but the shape isn't
clear, it's `deferred(unclear)`. Confusing them produces circular
reasoning — a `deferred(scope)` whose blocker is actually a
prerequisite of the thing being deferred.
## When to Redirect
+93
View File
@@ -82,6 +82,91 @@ Check coverage of:
- **Maintainability**: Testability, observability, modifiability
- **Scalability**: Horizontal/vertical scaling approach
#### F. Decision Quality and Deferral Honesty
This is the category the architect cannot do for itself — the architect
is too close to its own reasoning to see its own circular hedges. A
fresh context can see the whole picture and spot reasoning that folds
back on itself. This is often the highest-value category on projects
creating new protocols or solving poorly-defined problems, where the
shape isn't always clear and the architect may reach for a deferral to
avoid committing to an unclear shape.
Check for three cases:
**1. Hedging on a resolved decision.** An OQ marked `resolved` whose
resolution contains temporal language or escape hatches:
- "v1 default," "phase_n," "when x arrives," "can be revisited" — if
the decision is made, state it cleanly. Reserve temporal language
for genuinely deferred decisions.
- "feature extension, not an unmade decision" — if it's not decided,
it's not resolved.
- "additive, not blocking" — if it's not decided, don't claim it is.
- "two-way door — can be changed later if needed" — door type
classifies reversal cost, not whether a decision is made.
- "not a v1 blocker" — if it's not decided, it's deferred. Say what
unblocks it.
- "for now" / "not yet" on a resolved OQ — if the resolution has an
expiration date, it's not resolved.
- Resolution text primarily about how the decision can be changed
later ("X, but here's how we'd undo X") — the decision is made; drop
the undo instructions. If it's "X for now, Y later," the decision is
not made.
Flag as **critical**: the decision is either made (state it cleanly,
drop the hedge) or not made (mark it `deferred(scope)` or
`deferred(unclear)`). It cannot be both.
**2. False deferral / circular hedge.** A deferred OQ whose blocking
condition is a *prerequisite* of the thing being deferred, not a
blocker. This is the most damaging pattern — it creates a circular
dependency where the deferred thing can never resolve because its
blocker needs it.
Check each deferred OQ:
- For `deferred(scope)`: does the blocking condition *need* what this
OQ is deferring? If A is "blocked on B" and B needs A, that's a
prerequisite inversion, not a deferral.
- For `deferred(unclear)`: is the investigation target actually the
thing being deferred? If the investigation is "wait for X to exist"
and X needs this decision, it's circular.
- Is the information actually missing (`deferred(scope)` is correct),
or do the pieces exist but the shape wasn't synthesized
(`deferred(unclear)` is correct), or do the pieces *and* the shape
exist and the architect just didn't see the composition (should be
`resolved` — the decision is ready to make)?
Flag as **critical**: the deferral is circular or inverted. Suggest
the correct state — `resolved` if the pieces compose, `deferred(unclear)`
if the pieces exist but the shape needs investigation, `deferred(scope)`
only if the information genuinely is missing.
**3. Legitimate deferral.** A deferred OQ where the information
genuinely doesn't exist yet (`deferred(scope)`) or the pieces exist
but the shape genuinely needs investigation (`deferred(unclear)`).
Leave these — they're honest. The distinction from case 2 is that
the blocking condition or investigation target does *not* depend on
the thing being deferred.
#### G. Impacts Field Coverage
Every unresolved OQ (`open`, `deferred(scope)`, `deferred(unclear)`,
`partially resolved`) should have an **impacts** field stating what it
blocks downstream. Check:
- Is the impacts field present? Absence is a warning — without it, the
deferral's urgency is invisible and triage is guesswork.
- Is it specific? "Blocks the first hub deployment because the hub
dials workers" is useful. "Blocks the hub crate" is boilerplate.
- Does it match the priority? A `deferred(unclear)` with high priority
but a vague impacts field ("blocks future features") is a mismatch —
if it's high priority, it blocks something specific; say what.
Flag missing impacts fields as **warning**, vague ones as
**suggestion**.
### 3. Categorize Findings
**Critical**: Must fix before stabilization
@@ -90,6 +175,8 @@ Check coverage of:
- Missing quality attributes with significant impact
- Architectural decisions without rationale
- Inconsistencies in the specification
- Hedging on a resolved decision (case 1 of Decision Quality)
- Circular or inverted deferrals (case 2 of Decision Quality)
**Warning**: Should fix if possible
@@ -97,12 +184,14 @@ Check coverage of:
- Missing edge cases
- Incomplete interface definitions
- Implicit assumptions
- Missing impacts fields on unresolved OQs
**Suggestion**: Consider but optional
- Alternative phrasing
- Additional context that might help
- Documentation organization improvements
- Vague impacts fields on unresolved OQs
### 4. Write Review Report
@@ -168,3 +257,7 @@ section 2"
- Focus on architecture-level issues, not code-level
- Be constructive and specific
- Critical issues must block stabilization
- The Decision Quality (F) category is often the highest-value check on
projects creating new protocols or solving poorly-defined problems.
The architect cannot self-review this category — circular reasoning
is invisible from inside the circle. Prioritize it.
+18 -2
View File
@@ -15,9 +15,25 @@ currently parked and why" is answerable at a glance.
**Status values**:
- `open` — Needs to be resolved now. Has a clear path to resolution.
- `resolved` — Decided. The resolution is stated cleanly, without caveats about how it could be changed later.
- `deferred(scope)` — Cannot be resolved yet. The information doesn't exist. Has a concrete blocking condition (e.g., "blocked on: alknet-agent crate spec"). Not a failure — scope management.
- `deferred(scope)` — Cannot be resolved yet. The information is genuinely
missing — a crate spec, POC result, or use case that doesn't exist yet.
Has a concrete blocking condition. Not a failure — scope management.
- `deferred(unclear)` — Cannot be resolved yet. The pieces exist (decided
in other ADRs, existing types, existing patterns) but the composition
— how they fit together — isn't clear yet. Resolution requires
investigation (work through examples, maybe POC), not waiting. Has a
concrete investigation target and an impacts field. Not a failure —
honest uncertainty in a poorly-defined problem space.
- `partially resolved` — Some aspects decided, others deferred or open.
- `dissolved` — The question was reframed out of existence (e.g., superseded by an ADR that retires the premise). Kept for reference.
- `dissolved` — The question was reframed out of existence (e.g., superseded
by an ADR that retires the premise). Kept for reference.
**Impacts field**: Every unresolved OQ (`open`, `deferred(scope)`,
`deferred(unclear)`, `partially resolved`) should have an `Impacts`
field stating what it blocks downstream. Be specific: "blocks the first
hub deployment because the hub dials workers" not "blocks the hub
crate." This is the triage signal that makes the deferral's urgency
visible.
Door type classifications follow ADR-009 — they describe **reversal cost** (how expensive it is to undo), not urgency:
- **One-way door**: Reversal requires rewriting significant code or permanently closes a capability. Getting it wrong is expensive — requires ADR before implementation.
@@ -4,6 +4,9 @@
- **Status**: deferred(scope)
- **Door type**: One-way (federation model), two-way (mechanism)
- **Priority**: low
- **Impacts**: None currently — the one-hop model covers all current use
cases. Would impact peer-graph routing if a multi-hop topology becomes
needed (e.g., chained hubs).
- **Blocked on**: A concrete use case for multi-hop federation. The one-hop model covers all current use cases (head→worker, runner→hub).
- **Resolution**: The model is **one-hop** — worker A does not transitively
see worker B's ops through the head unless the head explicitly re-exports
@@ -5,6 +5,8 @@
- **Status**: open (scope, not deferral)
- **Door type**: One-way (crate boundary), two-way (mechanism)
- **Priority**: low
- **Impacts**: None — the browser path uses WebSocket (ADR-044). Would
impact a browser-to-P2P-peer relay use case if one arises.
- **Resolution**: There are two distinct "WebTransport proxy" concepts
that must not be conflated:
@@ -6,6 +6,8 @@
- **Door type**: Two-way (additive utility library; no protocol or API-surface
change)
- **Priority**: low
- **Impacts**: None — handlers produce streams today without it. Would
reduce boilerplate in stream-transforming handlers when added.
- **Blocked on**: A handler that needs stream operators and finds the existing
combinators (`Box::pin(stream::iter(...))`, `async_stream::stream!`,
`futures::stream`) insufficient. The operators library is a convenience, not
@@ -6,6 +6,8 @@
- **Status**: deferred(scope)
- **Door type**: Two-way
- **Priority**: low
- **Impacts**: None — the `modes` field is reserved as `{}` and backends
use their defaults, which work for the common terminal case.
- **Blocked on**: a concrete mode-control use case (a deployment that
needs to set echo/raw/canonical/etc. modes on a PTY, beyond the backend's
defaults).
@@ -6,6 +6,9 @@
- **Status**: deferred(scope)
- **Door type**: Two-way
- **Priority**: low
- **Impacts**: None — the runner *mechanism* (pipe mode) is in alknet-tty.
Would impact a future `alknet-runner` crate if job management / log
persistence / task graph integration becomes needed.
- **Blocked on**: a concrete runner-policy use case that forces the API
surface (job management, log persistence, task graph integration).
- **Resolution**: Not yet decidable. The runner *mechanism* (pipe mode —
@@ -6,6 +6,10 @@
- **Status**: deferred(scope)
- **Door type**: Two-way
- **Priority**: low
- **Impacts**: None — dev containers use the default bridge network;
hosted services declare networks/volumes in docker compose. Would
impact a fleet coordinator or dev-container orchestrator if one is
built.
- **Blocked on**: a concrete use case for network or volume management
over the call protocol. The two container use cases (disposable dev
containers, hosted services) don't currently require
@@ -7,6 +7,9 @@
- **Status**: deferred(scope)
- **Door type**: Two-way
- **Priority**: low
- **Impacts**: None — current use cases pull pre-built images. Would
impact any use case that builds images via alknet rather than
operator-side `docker compose build`.
- **Blocked on**: a concrete use case for building images over the
call protocol. The two container use cases (disposable dev
containers, hosted services) pull pre-built images (`docker/image/pull`)
@@ -8,6 +8,11 @@
- **Status**: deferred(scope)
- **Door type**: Two-way
- **Priority**: medium
- **Impacts**: Blocks the full `docker/container/create` input schema —
v1 uses the high-frequency fields; the long tail (mounts, port
bindings, networks, volumes, capabilities) is deferred to the
implementation pass. Does not block the `create` operation itself
(ADR-060 §5 is decided); blocks the schema detail.
- **Blocked on**: v1 implementation. The v1 `create` input schema
accepts the common fields (image, command, env, labels, name) and
the full `CreateContainerOptions` surface (mounts, port bindings,
@@ -4,6 +4,9 @@
- **Status**: open
- **Door type**: Two-way
- **Priority**: medium
- **Impacts**: Blocks clean supervision loop implementation — the hub's
`supervise_worker` currently polls `accept_bi()` until `ConnectionClosed`
(the interim). A `closed()` method would make this event-driven.
- **Resolution**: Not yet decided.
The hub's worker supervision loop needs a way to await connection close so it
@@ -4,6 +4,9 @@
- **Status**: open
- **Door type**: Two-way
- **Priority**: low
- **Impacts**: None — the struct shape is committed and the defaults are
a starting point. Would affect production tuning if operational
experience shows the defaults are wrong.
- **Resolution**: Not yet decided.
The `BackoffConfig` struct provides configurable backoff for worker
@@ -8,6 +8,13 @@
- **Status**: deferred(scope)
- **Door type**: two-way
- **Priority**: medium
- **Impacts**: Does NOT block individual transport dials — each
transport-specific dial helper builds its `TlsClientConfig` (ADR-087)
and its own connector standalone. Blocks only the *shared*
`AlknetClient::dial()` extraction (one entry point that picks the
transport and calls the right connector). Low impact until a second
transport's dial exists and the duplicated dial boilerplate becomes
worth extracting.
- **Blocked on**: a **second transport's** real dial existing, not just a
second QUIC dial. The blocking condition is met when, e.g., the SSH
crate's raw-TCP dial or the HTTP-wrapped call dial exists — so the
@@ -6,6 +6,10 @@
- **Door type**: two-way (additive — per-channel window tracking does not
change the wire format)
- **Priority**: low
- **Impacts**: None — bounded-buffer backpressure (ADR-076) is the v1
mechanism and handles the intended use cases (TTY, SSH, tunnels).
Would impact high-throughput channel use cases (e.g., file transfer
over a tunnel) if HOL blocking is observed in practice.
- **Blocked on**: a real deployment observes head-of-line blocking on a
saturated channel where the bounded-buffer's stop-reading mitigation is
insufficient. The trigger is specific: a channel whose consumer is
@@ -6,6 +6,9 @@
- **Door type**: two-way (additive — a helper function does not change any
API surface; handlers that inline the pattern continue to work)
- **Priority**: low
- **Impacts**: None — the two-pump *contract* is decided (ADR-078) and
handlers implement it inline. Would reduce ~10 lines of copy-paste
per handler when extracted; not a capability gate.
- **Blocked on**: a second two-pump handler existing, so the shape
convergence is observable. The tunnel handler is the first two-pump
consumer; the SSH `direct-tcpip` channel will be the second. Extracting
@@ -9,6 +9,10 @@
workers are provisioned against them is a breaking change for every
deployment)
- **Priority**: high
- **Impacts**: Blocks worker provisioning — a freshly-provisioned worker
has no way to enroll its key with the hub until the registration
endpoint exists. Blocks the first hub deployment that provisions
workers (web + native use case).
- **Blocked on**: nothing structural — the identity machinery
(`resolve_from_token`, `PeerEntry.auth_token_hash`, ADR-030/034)
and the HTTP substrate (`HttpAdapter` on `h2`/`http/1.1` over
@@ -2,9 +2,10 @@
- **Origin**: `docs/architecture/crates/tls/README.md`
(`TlsError` is referenced as the `Result` error type in
`TlsServerConfig::new`, `for_quinn()`, and the crate's public
signatures, but is never sketched or defined); ADR-082 (same —
`TlsError` in signatures, no shape).
`TlsServerConfig::new`, `TlsClientConfig::new`, `for_quinn()`, and the
crate's public signatures, but is never sketched or defined); ADR-082
(same — `TlsError` in signatures, no shape); ADR-087 (extends the
surface to `TlsClientConfig::new` — the client-side error variants).
- **Status**: open
- **Door type**: one-way (the error type is the public API surface of
`alknet-tls`; changing it after consumers exist is a breaking change
@@ -12,8 +13,14 @@
- **Priority**: high (an implementer cannot write the crate without
deciding this; guessing produces divergent shapes — one thin
`rustls::Error` wrapper vs a 10-variant enum with per-path context)
- **Impacts**: Blocks `alknet-tls` implementation — both
`TlsServerConfig::new` and `TlsClientConfig::new` reference `TlsError`
in their signatures. Blocks the hub's assembly-layer wiring (which
calls both). This is the next decision needed before the TLS crate
can be implemented.
- **Resolution**: Not yet decided. The shape needs to cover the failure
modes across all four identity paths:
modes across all server identity paths **and** the client verifier
paths (ADR-087):
- **Cert/key loading** (`X509`: file read + PEM parse;
`SelfSigned`: rcgen generation) — currently `io::Error`-wrapped in
@@ -30,6 +37,14 @@
returned from `new`), so `new`'s ACME path may only need to cover
"ACME feature not enabled but `TlsIdentity::Acme` configured"
(currently an `io::ErrorKind::Unsupported`).
- **Client verifier construction** (ADR-087) — `TlsClientConfig::new`
builds a `rustls::ClientConfig` with ADR-034's verifier selection.
Failure modes: verifier construction error (bad fingerprint format,
CA store init failure), unknown-remote fail-closed (not an error to
return — it's a `Result::Err` the caller gets for trying to connect
to an unknown raw-key remote), provider init failure. These
overlap with the server-side rustls-build errors but have
client-specific context (verifier selection inputs).
The open question is the granularity: a single `TlsError` enum with
variants per failure category (cert-load, rustls-build, quinn-wrap,
@@ -23,6 +23,10 @@
the simplification is worth the one-way-door commitment now, or
whether the call-protocol-only path suffices until a concrete
browser-needs-a-data-channel use case arrives.)
- **Impacts**: Does not block the browser path (ADR-048 works). Would
simplify the hub relay (one model, not two) and unblock browser
access to data-channel ALPNs (TTY, tunnels) over a single WebSocket
if resolved to "WebSocket carries channels."
- **Resolution**: Not yet decided. The two options:
**Option A — WebSocket carries call only (ADR-048 unchanged).** The