docs(arch): TCP+TLS as first-class owned transport — resolves OQ-60, dissolves OQ-61
ADR-083 revised: TCP+TLS moves from an external sibling loop calling public dispatch to a first-class owned transport via with_tcp_tls(listener, acceptor), running inside run() alongside the quinn and iroh accept loops. The endpoint owns all its accept loops; shutdown() stops them all. The multi-owner shutdown problem (OQ-61) does not arise — dissolved. The reason TCP+TLS was structurally excluded (ADR-010 Am. 1: the endpoint built transports internally, TCP+TLS couldn't fit) is gone after ADR-083 — the endpoint no longer builds transports; it runs accept loops on whatever it's given. TCP+TLS is a listener transport, same shape as quinn and iroh. ADR-010 Amendment 2 supersedes Am. 1's struct-level exclusion. dispatch stays public — but for genuinely external shapes (SSH channels, future WebTransport streams), which are connection-internal multiplexing, not listener transports. The listener-vs-multiplexing distinction is now explicit. OQ-60 resolved: the TCP+TLS loop lives in alknet-core behind a tcp feature (owned by the endpoint); builder functions are inlined by the assembly layer. A alknet-transport crate was rejected — it would contain only trivial builders; the real component (the loop) is in core. Hub- specific composition lives in the hub crate; transport runtimes that any node might need live in core. Updated: ADR-010 (Amendment 2), ADR-082 (TCP+TLS loop location), ADR-083 (revised), core/endpoint.md (struct + dispatch + shutdown), hub/README.md (transport table + assembly example + stale sibling references), tls/README.md (endpoint section + TCP+TLS loop location + references), open-questions.md (OQ-60 resolved, OQ-61 dissolved). Review: zero critical issues, five warnings fixed (stale hub README prose, stale core endpoint.md struct/dispatch listings, stale ADR-082 TCP+TLS loop location, stale TLS README reference entry, hub front-matter date).
This commit is contained in:
1 parent
81bde6f28f
commit
2abe8f1872
10 files changed
+524
-308
No files matched your search
@@ -194,7 +194,7 @@ adapter location map is now consistent: all HTTP-backed adapters
|
||||
| [007](decisions/007-bistream-type-definition.md) | BiStream Type Definition | Accepted |
|
||||
| [008](decisions/008-secret-service-integration.md) | Vault Integration Point | Accepted |
|
||||
| [009](decisions/009-one-way-door-decision-framework.md) | One-Way Door Decision Framework | Accepted |
|
||||
| [010](decisions/010-alpn-router-and-endpoint.md) | ALPN Router and Endpoint | Accepted (Amendment 1 superseded by ADR-083 — TCP+TLS dispatch is first-class via public `dispatch`) |
|
||||
| [010](decisions/010-alpn-router-and-endpoint.md) | ALPN Router and Endpoint | Accepted (Amendment 1 superseded by ADR-083 — TCP+TLS is a first-class owned transport via `with_tcp_tls`) |
|
||||
| [011](decisions/011-authcontext-structure.md) | AuthContext Structure and Resolution Flow | Accepted |
|
||||
| [012](decisions/012-call-protocol-stream-model.md) | Call Protocol Stream Model | Accepted |
|
||||
| [013](decisions/013-rust-canonical-implementation.md) | Rust as Canonical Implementation Language | Accepted |
|
||||
@@ -267,7 +267,7 @@ adapter location map is now consistent: all HTTP-backed adapters
|
||||
| [080](decisions/080-channelclient.md) | ChannelClient — the Client Side of a Channels Connection | Accepted |
|
||||
| [081](decisions/081-channels-subcrate-decomposition.md) | channels Sub-Crate Decomposition | Accepted |
|
||||
| [082](decisions/082-alknet-tls-extraction.md) | alknet-tls Crate Extraction | Proposed (amended — endpoint signature superseded by ADR-083) |
|
||||
| [083](decisions/083-endpoint-as-accept-loop-runner.md) | Endpoint as Pure Accept-Loop Runner with Public Dispatch | Proposed |
|
||||
| [083](decisions/083-endpoint-as-accept-loop-runner.md) | Endpoint as Multi-Transport Accept-Loop Runner with Public Dispatch | Proposed (revised — TCP+TLS is an owned transport, not external) |
|
||||
|
||||
## Open Questions
|
||||
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-07-09
|
||||
last_updated: 2026-07-14
|
||||
---
|
||||
|
||||
# Endpoint
|
||||
@@ -15,17 +15,26 @@ The central runtime type. Manages one or more QUIC connection sources, each feed
|
||||
|
||||
```rust
|
||||
pub struct AlknetEndpoint {
|
||||
// QUIC connection sources — both optional, both can be active simultaneously
|
||||
// One or more connection sources — all optional, all can be active simultaneously
|
||||
quinn: Option<quinn::Endpoint>, // Public QUIC+TLS
|
||||
iroh: Option<iroh::Endpoint>, // P2P relay-assisted
|
||||
#[cfg(feature = "tcp")]
|
||||
tcp_tls: Option<TcpTlsListener>, // TCP+TLS (TcpListener + TlsAcceptor)
|
||||
|
||||
handlers: Arc<HandlerRegistry>,
|
||||
dynamic: Arc<ArcSwap<DynamicConfig>>,
|
||||
identity_provider: Arc<dyn IdentityProvider>,
|
||||
shutdown: watch::Receiver<bool>,
|
||||
shutdown_tx: watch::Sender<bool>,
|
||||
shutdown_rx: watch::Receiver<bool>,
|
||||
drain_timeout: Duration,
|
||||
}
|
||||
```
|
||||
|
||||
See [ADR-083](../../decisions/083-endpoint-as-accept-loop-runner.md) for
|
||||
the full design (the endpoint takes no `StaticConfig` or TLS config;
|
||||
transports are built by the assembly layer and handed via
|
||||
`with_quinn` / `with_iroh` / `with_tcp_tls`).
|
||||
|
||||
### Why multiple connection sources?
|
||||
|
||||
A node can be reachable through different paths depending on its network context:
|
||||
@@ -37,29 +46,28 @@ A node can be reachable through different paths depending on its network context
|
||||
|
||||
These are not interchangeable transports — they are **complementary connectivity modes**. A node behind NAT that also has a public IP can use both simultaneously. Both produce QUIC connections that dispatch through the same `HandlerRegistry` by ALPN string.
|
||||
|
||||
### TCP is NOT an endpoint struct concern (but CAN dispatch through the registry)
|
||||
### TCP+TLS is a first-class owned transport
|
||||
|
||||
Bare TCP (SSH over port 22) does not use QUIC or ALPN. TCP access is not
|
||||
owned by the `AlknetEndpoint` struct — there is no `tcp:
|
||||
Option<TcpListener>` field. The endpoint manages QUIC connection sources
|
||||
(quinn + iroh) only.
|
||||
TCP+TLS is a listener transport, same shape as quinn and iroh. The
|
||||
endpoint owns it via `with_tcp_tls(listener, acceptor)` (behind a `tcp`
|
||||
feature) and runs its accept loop inside `run()` — `tcp.accept()` →
|
||||
`tls.accept()` → extract ALPN + fingerprint → `Connection::from_bidi` →
|
||||
`dispatch`. No external sibling loop, no duplicated dispatch logic.
|
||||
|
||||
This does **not** mean TCP+TLS can't participate in ALPN dispatch. Since
|
||||
[ADR-065](../../decisions/065-connection-from-stream-generic-single-stream.md),
|
||||
`Connection::from_bidi(tls_stream, alpn, remote_addr)` constructs a
|
||||
`Connection` from any `TlsStream<TcpStream>` (or any `AsyncRead +
|
||||
AsyncWrite` pair). A TCP+TLS accept loop built *outside* the endpoint (by
|
||||
the assembly layer or a handler) can wrap each TLS stream as a
|
||||
`Connection` and dispatch through the **same `HandlerRegistry`** the
|
||||
endpoint uses, by the ALPN negotiated in the TLS handshake. This is not a
|
||||
parallel listener bypassing the core — it's the same ALPN dispatch, over
|
||||
a non-QUIC transport. `HttpAdapter`, `TtyAdapter`, and the call handler
|
||||
all work over the single stream unchanged (ADR-065's yield-once
|
||||
`accept_bi` contract). The TCP+TLS accept loop itself is a follow-up
|
||||
commit, not part of `AlknetEndpoint`; the primitive it needs
|
||||
(`from_bidi`) is in place.
|
||||
This reverses ADR-010's original "TCP is not an endpoint struct concern."
|
||||
The reason TCP was excluded — the endpoint built transports internally,
|
||||
and TCP+TLS couldn't fit that shape — is gone (ADR-083: the endpoint no
|
||||
longer builds transports; it runs accept loops on whatever it's given).
|
||||
TCP+TLS fits the same listener shape as quinn and iroh.
|
||||
|
||||
The reference implementation's TCP transport (`alknet-main/crates/alknet-core/src/transport/tcp.rs`) is SSH-specific. It doesn't generalize to the ALPN model.
|
||||
The `dispatch` method is public for transports the endpoint **can't
|
||||
own** — SSH channels (one connection, many channels with different
|
||||
ALPNs — a multiplexing shape, not a listener) and future WebTransport
|
||||
streams (one QUIC connection, many WT streams). These are
|
||||
connection-internal multiplexing, not listener transports.
|
||||
|
||||
See [ADR-083](../../decisions/083-endpoint-as-accept-loop-runner.md)
|
||||
for the full design.
|
||||
|
||||
## HandlerRegistry
|
||||
|
||||
@@ -136,24 +144,39 @@ See iroh's `protocol.rs` (`/workspace/iroh/iroh/src/protocol.rs`) for the refere
|
||||
|
||||
### Dispatch function (shared)
|
||||
|
||||
The public `dispatch` method is the shared dispatch path for every
|
||||
transport — the endpoint's own accept loops (quinn, iroh, TCP+TLS) call
|
||||
it after transport-specific extraction, and external dispatch callers
|
||||
(SSH channels, future WebTransport streams) call it after their own
|
||||
extraction.
|
||||
|
||||
```
|
||||
fn dispatch(connection) {
|
||||
let alpn = connection.alpn();
|
||||
match handlers.get(alpn) {
|
||||
pub fn dispatch(&self, connection: Connection, alpn: Vec<u8>,
|
||||
fingerprint: Option<String>, remote_addr: Option<SocketAddr>) {
|
||||
// ACME guard (transport-agnostic — ADR-083)
|
||||
if alpn == b"acme-tls/1" {
|
||||
connection.close(0, "acme done");
|
||||
return;
|
||||
}
|
||||
match handlers.get(&alpn) {
|
||||
Some(handler) => {
|
||||
let auth = AuthContext::from_connection(&connection);
|
||||
let conn = Connection::from_quinn(connection); // or from_iroh
|
||||
let auth = build_auth_context(&alpn, remote_addr, fingerprint, &identity_provider);
|
||||
tokio::spawn(async move {
|
||||
if let Err(e) = handler.handle(conn, &auth).await {
|
||||
if let Err(e) = handler.handle(connection, &auth).await {
|
||||
// log error, connection closes
|
||||
}
|
||||
});
|
||||
}
|
||||
None => connection.close(0u32, "no handler"),
|
||||
None => { connection.close(0, "no handler"); /* log warning */ }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Synchronous (non-async): spawns the handler on its own task and returns
|
||||
immediately. The caller's accept loop is not blocked. See
|
||||
[ADR-083](../../decisions/083-endpoint-as-accept-loop-runner.md) for the
|
||||
full dispatch contract.
|
||||
|
||||
### What the accept loops do NOT do
|
||||
|
||||
- **No byte-peeking**: ALPN negotiation handles protocol detection. The old `stealth` module's `detect_protocol()` is unnecessary.
|
||||
@@ -260,11 +283,14 @@ impl AlknetEndpoint {
|
||||
}
|
||||
```
|
||||
|
||||
- `shutdown_sender()` returns a clone of the shutdown channel sender. Call `send(true)` to signal shutdown.
|
||||
- `shutdown()` signals all accept loops to stop, waits for in-flight connections with a drain timeout (default: 2 seconds), then forcefully closes remaining connections.
|
||||
- `shutdown_sender()` returns a clone of the shutdown channel sender. Call `send(true)` to signal shutdown. The assembly layer uses this for any external dispatch callers (SSH, future WT); the endpoint's own loops are signaled internally.
|
||||
- `shutdown()` signals all owned accept loops (quinn, iroh, TCP+TLS) to stop, waits for in-flight dispatched handlers with a drain timeout, then forcefully closes remaining connections. One owner, one shutdown — no external loop coordination (ADR-083).
|
||||
- SIGTERM/SIGINT are wired to the shutdown channel by the CLI binary.
|
||||
|
||||
The drain timeout is configurable via `StaticConfig::drain_timeout`.
|
||||
The drain timeout is passed to `AlknetEndpoint::new()` directly (as
|
||||
`drain_timeout: Duration`), not via `StaticConfig` — the endpoint no
|
||||
longer takes `StaticConfig` (ADR-083). The assembly layer reads
|
||||
`StaticConfig::drain_timeout` and passes it in.
|
||||
|
||||
## Error Handling
|
||||
|
||||
@@ -308,9 +334,11 @@ Non-fatal errors within a handler. See [core-types.md](core-types.md) for detail
|
||||
|
||||
| Decision | ADR | Summary |
|
||||
|----------|-----|---------|
|
||||
| Multi-connectivity endpoint (quinn + iroh) | [ADR-010](../../decisions/010-alpn-router-and-endpoint.md) | Both optional, both feed same ALPN router |
|
||||
| Multi-connectivity endpoint (quinn + iroh + TCP+TLS) | [ADR-010](../../decisions/010-alpn-router-and-endpoint.md), [ADR-083](../../decisions/083-endpoint-as-accept-loop-runner.md) | All three optional, all feed same dispatch; endpoint owns all accept loops |
|
||||
| Endpoint takes no TLS config; assembly layer builds transports | [ADR-083](../../decisions/083-endpoint-as-accept-loop-runner.md) | `new()` takes `drain_timeout` + builder methods, no `StaticConfig` or `Arc<TlsServerConfig>` |
|
||||
| TCP+TLS is a first-class owned transport | [ADR-083](../../decisions/083-endpoint-as-accept-loop-runner.md) | `with_tcp_tls(listener, acceptor)` — reverses ADR-010's "TCP is not an endpoint struct concern" |
|
||||
| Public `dispatch` for SSH/WT (multiplexing shapes) | [ADR-083](../../decisions/083-endpoint-as-accept-loop-runner.md) | `dispatch` is public for connection-internal multiplexing, not for listener transports |
|
||||
| Static handler registration | [ADR-010](../../decisions/010-alpn-router-and-endpoint.md) | Two-way door, start static, add ArcSwap later |
|
||||
| TCP is not an endpoint struct concern (but dispatches via `from_stream`) | [ADR-010](../../decisions/010-alpn-router-and-endpoint.md), [ADR-065](../../decisions/065-connection-from-stream-generic-single-stream.md) | `AlknetEndpoint` is QUIC-only (no `tcp` field); a TCP+TLS loop outside the endpoint wraps streams via `from_bidi` and shares the registry |
|
||||
| No byte-peeking, ALPN dispatch only | [ADR-001](../../decisions/001-alpn-protocol-dispatch.md) | TLS layer handles protocol detection |
|
||||
| Stealth mode = HTTP handler on standard ALPNs | [ADR-010](../../decisions/010-alpn-router-and-endpoint.md) | Decoy via ALPN routing, not byte-peek |
|
||||
| Network identity ≠ auth identity | [ADR-010](../../decisions/010-alpn-router-and-endpoint.md) | TLS cert/NodeId = network, SSH key/token = auth |
|
||||
@@ -322,4 +350,6 @@ See [open-questions.md](../../open-questions.md) for full details.
|
||||
|
||||
- **OQ-04**: Resolved — HandlerRegistry is static at startup.
|
||||
- **OQ-05**: Resolved — multi-connectivity endpoint with quinn + iroh, both feature-gated.
|
||||
- **OQ-12**: Resolved — two distinct TLS identity use cases: RFC 7250 raw keys (default, P2P) and X.509 certs (domain-hosted, browsers). ACME auto-provisioning designed in [ADR-027](../../decisions/027-tls-identity-redesign-acme-rawkey-decoupling.md); RawKey decoupled from the `iroh` feature (available in quinn-only builds).
|
||||
- **OQ-12**: Resolved — two distinct TLS identity use cases: RFC 7250 raw keys (default, P2P) and X.509 certs (domain-hosted, browsers). ACME auto-provisioning designed in [ADR-027](../../decisions/027-tls-identity-redesign-acme-rawkey-decoupling.md); RawKey decoupled from the `iroh` feature (available in quinn-only builds).
|
||||
- **OQ-60**: Resolved — transport construction is inlined by the assembly layer; the TCP+TLS loop lives in `alknet-core` behind a `tcp` feature as an owned transport. See [ADR-083](../../decisions/083-endpoint-as-accept-loop-runner.md).
|
||||
- **OQ-61**: Dissolved — the multi-owner shutdown problem does not arise; the endpoint owns all its accept loops. See [ADR-083](../../decisions/083-endpoint-as-accept-loop-runner.md).
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-07-13
|
||||
last_updated: 2026-07-14
|
||||
---
|
||||
|
||||
# alknet-hub
|
||||
@@ -42,11 +42,11 @@ welded to a dial. See "Transport" below.
|
||||
### What the hub provides
|
||||
|
||||
1. **Multi-transport endpoint** — accepts channels connections over
|
||||
QUIC (via the quinn `AlknetEndpoint`) and over TCP+TLS (via an
|
||||
accept loop that wraps each `TlsStream<TcpStream>` as a `Connection`
|
||||
with `from_bidi`, per ADR-065/010). Both feed the same
|
||||
`HandlerRegistry`. Also serves HTTP (`h2`/`http/1.1`) over TCP+TLS
|
||||
for registration and browser access. All three coexist.
|
||||
QUIC and over TCP+TLS (both owned transports on `AlknetEndpoint` via
|
||||
`with_quinn` / `with_tcp_tls`, per ADR-083). Both feed the same
|
||||
dispatch path. Also serves HTTP (`h2`/`http/1.1`) over TCP+TLS for
|
||||
registration and browser access. All three coexist; the endpoint owns
|
||||
all accept loops and `shutdown()` stops them all.
|
||||
|
||||
2. **Peer lifecycle** — accept, dial, disconnect, reconnect with
|
||||
backoff. Identity resolution via `IdentityProvider` (fingerprint or
|
||||
@@ -81,7 +81,10 @@ The hub pattern requires wiring that `alknet-channels` and
|
||||
|
||||
- The hub accepts channels connections over **multiple transports**
|
||||
simultaneously. The channels crate is transport-agnostic (ADR-071,
|
||||
ADR-065); the hub is where the multi-transport accept loop lives.
|
||||
ADR-065); the hub composes the multi-transport endpoint
|
||||
(`AlknetEndpoint` with `with_quinn` / `with_tcp_tls`, ADR-083). The
|
||||
accept loops themselves live in `alknet-core`; the hub provides the
|
||||
handlers and wiring.
|
||||
- The hub relays channels between legs (ADR-079) — terminating
|
||||
channel 0 on each leg, translating `channel/open`, byte-forwarding
|
||||
data channels. The channels crate is ALPN-blind and does not know it
|
||||
@@ -179,30 +182,34 @@ The `Hub` exposes builder methods for optional hooks
|
||||
|
||||
The hub accepts and dials channels connections over any transport
|
||||
the channels protocol supports (ADR-071). In practice, a hub runs
|
||||
multiple accept loops simultaneously:
|
||||
multiple accept loops simultaneously — all owned by `AlknetEndpoint`
|
||||
via builder methods (ADR-083):
|
||||
|
||||
| Endpoint | Transport | What it carries |
|
||||
|----------|-----------|-----------------|
|
||||
| Quinn `AlknetEndpoint` | QUIC (quinn/iroh) | Channels connections from workers with QUIC reachability |
|
||||
| TCP+TLS accept loop | TCP + TLS (`h2`/`http/1.1`/`alknet/channels`) | HTTP registration endpoint, browser access, channels-over-TCP from workers |
|
||||
| Builder method | Transport | What it carries |
|
||||
|----------------|-----------|-----------------|
|
||||
| `with_quinn` | QUIC (quinn) | Channels connections from workers with QUIC reachability |
|
||||
| `with_iroh` | QUIC (iroh, relay-assisted) | Channels connections from workers behind NAT |
|
||||
| `with_tcp_tls` | TCP + TLS (`h2`/`http/1.1`/`alknet/channels`) | HTTP registration endpoint, browser access, channels-over-TCP from workers |
|
||||
| (Future) WebTransport | WebTransport | Browser bidirectional path (deferred per ADR-044; WebSocket is the v1 browser path) |
|
||||
|
||||
All accept loops feed into the same `HandlerRegistry`. The hub's
|
||||
`HttpAdapter` serves `h2`/`http/1.1` over the TCP+TLS path for
|
||||
registration and browser access; the `ChannelsAdapter` (registered on
|
||||
the hub's `HandlerRegistry`) serves
|
||||
`alknet/channels` over the QUIC path (and over TCP+TLS when a worker
|
||||
dials channels-over-TCP). Both ALPNs are registered on the same
|
||||
registry; the TLS handshake on each connection negotiates the ALPN
|
||||
and dispatches to the right adapter.
|
||||
All accept loops run inside `endpoint.run()` and feed the same dispatch
|
||||
path. The hub's `HttpAdapter` serves `h2`/`http/1.1` over the TCP+TLS
|
||||
path for registration and browser access; the `ChannelsAdapter`
|
||||
(registered on the hub's `HandlerRegistry`) serves `alknet/channels`
|
||||
over the QUIC path (and over TCP+TLS when a worker dials
|
||||
channels-over-TCP). Both ALPNs are registered on the same registry;
|
||||
the TLS handshake on each connection negotiates the ALPN and dispatches
|
||||
to the right adapter.
|
||||
|
||||
The TCP+TLS accept loop constructs a `Connection` per accepted
|
||||
`TlsStream<TcpStream>` via `Connection::from_bidi` (ADR-065) and
|
||||
hands it to the same `HandlerRegistry::dispatch` the quinn endpoint
|
||||
uses. This is the accept-loop-outside-the-endpoint pattern from
|
||||
ADR-010's Amendment 1: the `AlknetEndpoint` struct stays QUIC-only
|
||||
(no `tcp` field), and the TCP+TLS loop is a sibling accept source
|
||||
that shares the registry. The hub is where that sibling loop lives.
|
||||
`TlsStream<TcpStream>` via `Connection::from_bidi` (ADR-065) and hands
|
||||
it to the same dispatch path the quinn endpoint uses. After ADR-083,
|
||||
TCP+TLS is a first-class owned transport on `AlknetEndpoint` — the hub
|
||||
hands a `TcpListener` + `TlsAcceptor` to the endpoint via
|
||||
`with_tcp_tls(listener, acceptor)`, and the endpoint runs the accept
|
||||
loop inside `run()` alongside the quinn/iroh loops. No external sibling
|
||||
loop; the endpoint owns all its accept loops and `shutdown()` stops them
|
||||
all.
|
||||
|
||||
#### Dial (outbound workers) — transport-agnostic
|
||||
|
||||
@@ -330,8 +337,8 @@ let callback = WorkerConnectedCallback::new(Arc::clone(&hub), FromCallConfig::ne
|
||||
let channels_adapter = ChannelsAdapter::new(Arc::clone(®istry), /* ... */)
|
||||
.with_worker_connected_callback(callback);
|
||||
// Register channels_adapter on alknet/channels in the HandlerRegistry.
|
||||
// Both the quinn endpoint and the TCP+TLS accept loop dispatch
|
||||
// alknet/channels connections to it.
|
||||
// The endpoint dispatches alknet/channels connections to it — whether
|
||||
// they arrived over quinn, iroh, or TCP+TLS (all owned by the endpoint).
|
||||
```
|
||||
|
||||
### Identity over transports
|
||||
@@ -596,37 +603,38 @@ let hub = Arc::new(Hub::new(
|
||||
Arc::clone(&identity_provider),
|
||||
).with_ownership_provider(ownership_provider));
|
||||
|
||||
// 3. Register the ChannelsAdapter on alknet/channels. Both the quinn
|
||||
// endpoint and the TCP+TLS accept loop dispatch alknet/channels
|
||||
// connections to this adapter.
|
||||
// 3. Register the ChannelsAdapter on alknet/channels. The endpoint
|
||||
// dispatches alknet/channels connections to this adapter — whether
|
||||
// they arrived over quinn, iroh, or TCP+TLS.
|
||||
let callback = WorkerConnectedCallback::new(Arc::clone(&hub), FromCallConfig::new());
|
||||
let channels_adapter = ChannelsAdapter::new(/* ... */)
|
||||
.with_worker_connected_callback(callback);
|
||||
registry.register(b"alknet/channels", Arc::new(channels_adapter));
|
||||
|
||||
// 4. Register the HttpAdapter on h2/http1.1. The TCP+TLS accept loop
|
||||
// dispatches h2/http1.1 connections to this adapter (registration
|
||||
// endpoint, browser access, stealth decoy).
|
||||
// 4. Register the HttpAdapter on h2/http1.1. The endpoint dispatches
|
||||
// h2/http1.1 connections (arriving over TCP+TLS) to this adapter
|
||||
// (registration endpoint, browser access, stealth decoy).
|
||||
let http_adapter = HttpAdapter::new(/* ... */);
|
||||
registry.register(b"h2", Arc::new(http_adapter.clone()));
|
||||
registry.register(b"http/1.1", Arc::new(http_adapter));
|
||||
|
||||
// 5. Start the QUIC accept loop (workers with QUIC reachability)
|
||||
let quinn_endpoint = AlknetEndpoint::new(/* ... */, Arc::clone(®istry));
|
||||
quinn_endpoint.run().await;
|
||||
// 5. Build the quinn endpoint (raw-key config for native clients)
|
||||
let quinn_endpoint = raw_key_tls.for_quinn()?.into_endpoint(listen_addr)?;
|
||||
|
||||
// 6. Start the TCP+TLS accept loop (registration, browser access,
|
||||
// channels-over-TCP from workers)
|
||||
// 6. Build the TCP+TLS listener (X.509/ACME config for HTTPS)
|
||||
let tcp_listener = TcpListener::bind(registration_addr).await?;
|
||||
tokio::spawn(async move {
|
||||
loop {
|
||||
let (stream, _) = tcp_listener.accept().await?;
|
||||
let tls_stream = tls_acceptor.accept(stream).await?;
|
||||
let alpn = tls_stream.alpn()?;
|
||||
let conn = Connection::from_bidi(tls_stream, alpn, /* remote_addr */);
|
||||
registry.dispatch(conn).await; // routes by ALPN to HttpAdapter or ChannelsAdapter
|
||||
}
|
||||
});
|
||||
let tls_acceptor = x509_tls.for_tcp_tls();
|
||||
|
||||
// 7. Construct the endpoint with all owned transports, then run.
|
||||
// TCP+TLS is a first-class owned transport (ADR-083) — the endpoint
|
||||
// runs its accept loop inside run() alongside quinn/iroh. No external
|
||||
// sibling loop; shutdown() stops them all.
|
||||
let endpoint = Arc::new(
|
||||
AlknetEndpoint::new(registry, dynamic, identity_provider, drain_timeout)
|
||||
.with_quinn(quinn_endpoint)
|
||||
.with_tcp_tls(tcp_listener, tls_acceptor),
|
||||
);
|
||||
endpoint.clone().run().await;
|
||||
|
||||
// 7. Dial outbound workers (hub dials workers). The closure produces a
|
||||
// channels Connection — the hub's supervise_worker calls
|
||||
@@ -677,7 +685,7 @@ into `CallAdapter::with_aggregated_env`.
|
||||
| Three peer roles | [ADR-034](../../decisions/034-outgoing-only-x509-and-three-peer-roles.md) | Hub = role-3 `PeerEntry` (mixed fingerprints); browsers not peers; bearer-token identity over TCP/WebTransport |
|
||||
| ChannelClient — transport-agnostic | [ADR-080](../../decisions/080-channelclient.md) | `from_connection` primary, `connect_quic` convenience; the dial path the hub uses |
|
||||
| Channels transport-agnostic | [ADR-071](../../decisions/071-channels-wire-format.md) | Substrate modes; `Connection::from_stream`/`from_bidi` (ADR-065) — the substrate the hub relays |
|
||||
| TCP+TLS dispatch via from_stream | [ADR-010](../../decisions/010-alpn-router-and-endpoint.md) Am. 1 | `AlknetEndpoint` stays QUIC-only; TCP+TLS accept loop shares the registry |
|
||||
| TCP+TLS as first-class owned transport | [ADR-083](../../decisions/083-endpoint-as-accept-loop-runner.md) | `with_tcp_tls(listener, acceptor)` — TCP+TLS is owned by the endpoint, not a sibling loop; supersedes ADR-010 Am. 1 |
|
||||
| Channel 0 pre-negotiated | [ADR-072](../../decisions/072-channel-0-pre-negotiated-call.md) | Channel 0 = `alknet/call`; the `CallAdapter` runs here |
|
||||
| Channel lifecycle operations | [ADR-073](../../decisions/073-channel-lifecycle-operations.md) | `channel/open`/`close`/`control`/`resources/subscribe` — what the hub translates |
|
||||
|
||||
@@ -731,6 +739,8 @@ See [open-questions.md](../../open-questions.md) for full details.
|
||||
- ADR-069: from_call Is a Manual Free Function
|
||||
- ADR-079: Hub Relay — Translate, Not Transparently Forward
|
||||
- ADR-080: ChannelClient (transport-agnostic `from_connection`)
|
||||
- ADR-082: alknet-tls extraction (`TlsServerConfig` — shared across quinn + TCP+TLS)
|
||||
- ADR-083: Endpoint as multi-transport accept-loop runner (`with_tcp_tls` — TCP+TLS owned by the endpoint; the hub composes transports and handlers)
|
||||
- alkapi [hub.md](/workspace/@alkdev/alkapi/docs/architecture/hub.md) —
|
||||
the first hub consumer, the concrete use case that informed this
|
||||
crate
|
||||
|
||||
@@ -268,8 +268,9 @@ production and `rustls::sign` in the test helper `build_ed25519_spki_der`
|
||||
|
||||
`AlknetEndpoint::new()` currently builds `TlsSetup` internally. After
|
||||
the refactor (see [ADR-083](../../decisions/083-endpoint-as-accept-loop-runner.md)),
|
||||
the endpoint takes **no TLS config at all** — it is a pure accept-loop
|
||||
runner with a public `dispatch` method:
|
||||
the endpoint takes **no TLS config at all** — it is a multi-transport
|
||||
accept-loop runner. TCP+TLS is an owned transport (via `with_tcp_tls`),
|
||||
not an external loop:
|
||||
|
||||
```rust
|
||||
impl AlknetEndpoint {
|
||||
@@ -283,6 +284,17 @@ impl AlknetEndpoint {
|
||||
pub fn with_quinn(mut self, endpoint: quinn::Endpoint) -> Self;
|
||||
pub fn with_iroh(mut self, endpoint: iroh::Endpoint) -> Self;
|
||||
|
||||
/// TCP+TLS is a first-class owned transport — same `run()` loop,
|
||||
/// same `shutdown()` as quinn/iroh. Feature-gated on `tcp`.
|
||||
#[cfg(feature = "tcp")]
|
||||
pub fn with_tcp_tls(
|
||||
mut self,
|
||||
listener: tokio::net::TcpListener,
|
||||
acceptor: tokio_rustls::TlsAcceptor,
|
||||
) -> Self;
|
||||
|
||||
/// Public for SSH channels / future WT (connection-internal
|
||||
/// multiplexing, not listener transports).
|
||||
pub fn dispatch(
|
||||
&self,
|
||||
connection: Connection,
|
||||
@@ -292,32 +304,35 @@ impl AlknetEndpoint {
|
||||
);
|
||||
|
||||
pub async fn run(self: Arc<Self>);
|
||||
pub async fn shutdown(&self) -> Result<(), EndpointError>;
|
||||
}
|
||||
```
|
||||
|
||||
The assembly layer builds the `TlsServerConfig`(s), builds the
|
||||
transports (`for_quinn()` → `quinn::Endpoint::server()`,
|
||||
`for_tcp_tls()` → `TlsAcceptor`, the `Ed25519SecretKey` → iroh), and
|
||||
hands the pre-built quinn/iroh endpoints to `AlknetEndpoint` via
|
||||
`for_tcp_tls()` → `TlsAcceptor` paired with a `TcpListener`,
|
||||
`Ed25519SecretKey` → iroh), and hands them to `AlknetEndpoint` via
|
||||
builder methods. A hub serving native clients and browsers holds two
|
||||
`TlsServerConfig`s (raw key + X.509/ACME); the endpoint takes neither —
|
||||
it takes the already-built transport endpoints. The TCP+TLS accept
|
||||
loops (one per config) call `endpoint.dispatch(...)` — the same
|
||||
dispatch path quinn and iroh use. The ACME handle lives on the
|
||||
`TlsServerConfig`, not the endpoint.
|
||||
it takes the already-built transport endpoints. The TCP+TLS listener
|
||||
is owned by the endpoint via `with_tcp_tls`; the endpoint runs its
|
||||
accept loop inside `run()` and stops it on `shutdown()`. The ACME
|
||||
handle lives on the `TlsServerConfig`, not the endpoint.
|
||||
|
||||
This resolves the single-`Arc<TlsServerConfig>` problem: the endpoint
|
||||
has no "the TLS config" to take because a hub has two.
|
||||
has no "the TLS config" to take because a hub has two. It also means
|
||||
shutdown is single-owner — the endpoint owns all its accept loops
|
||||
(quinn, iroh, TCP+TLS); one `shutdown()` stops them all.
|
||||
|
||||
### The TCP+TLS accept loop (out of scope for this crate)
|
||||
|
||||
`alknet-tls` provides `for_tcp_tls() -> TlsAcceptor`. The actual TCP
|
||||
accept loop (`TcpListener::accept` → `TlsAcceptor::accept` →
|
||||
`Connection::from_bidi` → `HandlerRegistry::dispatch`) lives elsewhere —
|
||||
in `alknet-hub` (the hub is the primary multi-transport consumer) or in
|
||||
a future `alknet-core` `TcpTlsAcceptor` module. `alknet-tls` is the cert
|
||||
provider, not the accept loop. This keeps `alknet-tls` focused on TLS
|
||||
setup and cert sharing, not transport accept logic.
|
||||
`Connection::from_bidi` → `endpoint.dispatch()`) lives in `alknet-core`
|
||||
behind a `tcp` feature, as an owned transport on `AlknetEndpoint` (via
|
||||
`with_tcp_tls(listener, acceptor)` — see ADR-083). `alknet-tls` is the
|
||||
cert provider, not the accept loop. This keeps `alknet-tls` focused on
|
||||
TLS setup and cert sharing, not transport accept logic.
|
||||
|
||||
## Crate dependencies (in the dep graph)
|
||||
|
||||
@@ -331,23 +346,21 @@ alknet-core (loses TLS setup code)
|
||||
alknet-call (client-side verifier — unchanged)
|
||||
├── alknet-core (fingerprint.rs)
|
||||
|
||||
alknet-http (future: TCP+TLS accept loop)
|
||||
├── alknet-tls (TlsServerConfig::for_tcp_tls())
|
||||
├── alknet-core (Connection::from_bidi, HandlerRegistry)
|
||||
|
||||
alknet-hub (multi-transport endpoint)
|
||||
├── alknet-tls (TlsServerConfig — shared across quinn + TCP)
|
||||
├── alknet-channels-call (ChannelClient)
|
||||
├── alknet-call (CallAdapter, Dispatcher)
|
||||
├── alknet-http (HttpAdapter)
|
||||
├── alknet-core (AlknetEndpoint, HandlerRegistry, Connection)
|
||||
├── alknet-core (AlknetEndpoint with quinn + iroh + tcp features, HandlerRegistry, Connection)
|
||||
```
|
||||
|
||||
`alknet-tls` depends on `alknet-core` only. No handler crate depends on
|
||||
`alknet-tls` — they depend on `alknet-core` for types and on
|
||||
`alknet-tls` only indirectly through the assembly layer. The assembly
|
||||
layer (the deployment binary) builds the `TlsServerConfig` and passes it
|
||||
to the endpoint and the TCP+TLS accept loop.
|
||||
layer (the deployment binary) builds the `TlsServerConfig`(s), builds
|
||||
the transport endpoints (quinn/iroh/TCP+TLS), and hands them to
|
||||
`AlknetEndpoint` via `with_quinn` / `with_iroh` / `with_tcp_tls`
|
||||
(ADR-083).
|
||||
|
||||
## Design Decisions
|
||||
|
||||
@@ -357,7 +370,7 @@ All design decisions are documented as ADRs in
|
||||
| ADR | Decision | Summary |
|
||||
|-----|----------|---------|
|
||||
| [082](../../decisions/082-alknet-tls-extraction.md) | alknet-tls crate extraction | Extract TLS setup from alknet-core/endpoint.rs; `TlsServerConfig` shareable across quinn + TCP+TLS + iroh; one ACME state machine |
|
||||
| [083](../../decisions/083-endpoint-as-accept-loop-runner.md) | Endpoint as accept-loop runner | `AlknetEndpoint` takes no TLS config; assembly layer builds transports from `TlsServerConfig`s; public `dispatch` method; `acme-tls/1` guard moves to shared `dispatch` |
|
||||
| [083](../../decisions/083-endpoint-as-accept-loop-runner.md) | Endpoint as multi-transport accept-loop runner | `AlknetEndpoint` takes no TLS config; TCP+TLS is an owned transport (`with_tcp_tls`); `dispatch` public for SSH/WT; `acme-tls/1` guard moves to shared `dispatch` |
|
||||
|
||||
## Open Questions
|
||||
|
||||
@@ -370,17 +383,13 @@ See [open-questions.md](../../open-questions.md) for full details.
|
||||
on `alknet-tls` — a new dep edge. If it stays, core keeps a narrow
|
||||
`rustls` dep. Decision-ready — the answer depends on whether we want
|
||||
core to be `rustls`-free.
|
||||
- **OQ-60** (open): Where does transport construction live? The
|
||||
endpoint-as-accept-loop-runner boundary (ADR-083) commits to
|
||||
construction being outside the endpoint, but the destination —
|
||||
assembly layer, `alknet-tls` convenience helpers, or a transport
|
||||
module/crate — is open. `alknet-tls`'s stated job is "TLS setup, not
|
||||
transport accept logic," which is an argument against the helper
|
||||
option.
|
||||
- **OQ-61** (open): Multi-owner shutdown coordination. The endpoint owns
|
||||
dispatched handlers (spawned in `dispatch`); the assembly layer owns
|
||||
spawned accept loops (TCP+TLS). The coordination mechanism (shared
|
||||
`shutdown_sender`, drain semantics) is unspecified.
|
||||
- **OQ-60** (resolved): Where does transport construction live? The
|
||||
TCP+TLS accept loop lives in `alknet-core` behind a `tcp` feature as
|
||||
an owned endpoint transport (`with_tcp_tls`). Builder functions are
|
||||
inlined by the assembly layer. See ADR-083.
|
||||
- **OQ-61** (dissolved): Multi-owner shutdown coordination. The
|
||||
problem does not arise — the endpoint owns all its accept loops
|
||||
(quinn, iroh, TCP+TLS); `shutdown()` stops them all. See ADR-083.
|
||||
|
||||
## References
|
||||
|
||||
@@ -392,8 +401,9 @@ See [open-questions.md](../../open-questions.md) for full details.
|
||||
— client-side verifier selection (CA vs fingerprint pin)
|
||||
- `docs/architecture/decisions/065-connection-from-stream-generic-single-stream.md`
|
||||
— `Connection::from_stream`/`from_bidi` (TCP+TLS path)
|
||||
- `docs/architecture/decisions/010-alpn-router-and-endpoint.md` Amendment 1
|
||||
— TCP+TLS dispatch via `from_stream` (accept loop outside the endpoint struct)
|
||||
- `docs/architecture/decisions/010-alpn-router-and-endpoint.md`
|
||||
Amendment 2 — TCP+TLS is a first-class owned transport
|
||||
(`with_tcp_tls`); supersedes Amendment 1's sibling-loop framing
|
||||
- `docs/architecture/crates/core/endpoint.md` — current endpoint design
|
||||
(TLS section will be amended to point to `alknet-tls`)
|
||||
- `docs/architecture/crates/core/config.md` — `TlsIdentity`, `StaticConfig`
|
||||
|
||||
@@ -254,4 +254,33 @@ related cleanup: it bumps the iroh dep to 1.0, unblocking `alknet-blobs`
|
||||
6 API surface edits in `endpoint.rs` / `types.rs` (the `Endpoint::builder`
|
||||
preset, `SecretKey::from_bytes`/`generate` signatures,
|
||||
`Connection::remote_id`/`alpn` return types). No ADR needed; the endpoint
|
||||
design is unchanged.
|
||||
design is unchanged.
|
||||
|
||||
### Amendment 2 (2026-07-14): TCP+TLS is a first-class owned transport (supersedes Amendment 1's struct-level exclusion)
|
||||
|
||||
Amendment 1 preserved the "not an endpoint struct concern" framing at
|
||||
the struct level — no `tcp: Option<TcpListener>` field on
|
||||
`AlknetEndpoint`. The rationale was that the endpoint built transports
|
||||
internally (quinn, iroh), and TCP+TLS couldn't fit that construction
|
||||
shape, so it was a sibling loop outside the struct.
|
||||
|
||||
[ADR-083](083-endpoint-as-accept-loop-runner.md) removes that rationale:
|
||||
the endpoint no longer builds transports at all — it runs accept loops on
|
||||
whatever it's given via builder methods. TCP+TLS is a listener transport,
|
||||
same shape as quinn and iroh (accept → extract ALPN + fingerprint →
|
||||
`Connection::from_bidi` → `dispatch`). The endpoint now owns it via
|
||||
`with_tcp_tls(listener, acceptor)` (behind a `tcp` feature), runs its
|
||||
accept loop inside `run()`, and stops it on `shutdown()`. The struct gains
|
||||
a `tcp_tls: Option<TcpTlsListener>` field.
|
||||
|
||||
Amendment 1's *dispatch* contribution survives — the public `dispatch`
|
||||
method and `Connection::from_bidi` are what make TCP+TLS dispatch work.
|
||||
Amendment 1's *struct-level exclusion* (no `tcp` field, sibling loop
|
||||
outside) is **superseded**: TCP+TLS is now a first-class owned transport.
|
||||
The `dispatch` method stays public, but for genuinely external shapes
|
||||
(SSH channels, future WebTransport streams) — connection-internal
|
||||
multiplexing, not listener transports.
|
||||
|
||||
This also means shutdown is single-owner: the endpoint owns all its
|
||||
accept loops (quinn, iroh, TCP+TLS); one `shutdown()` stops them all.
|
||||
The multi-owner shutdown coordination problem (OQ-61) does not arise.
|
||||
@@ -174,11 +174,11 @@ transports.
|
||||
|
||||
`alknet-tls` provides `for_tcp_tls() -> TlsAcceptor`. The actual TCP
|
||||
accept loop (`TcpListener::accept` → `TlsAcceptor::accept` →
|
||||
`Connection::from_bidi` → `HandlerRegistry::dispatch`) lives elsewhere —
|
||||
in `alknet-hub` (the primary multi-transport consumer) or in a future
|
||||
`alknet-core` module. `alknet-tls` is the cert provider, not the
|
||||
accept loop. This keeps `alknet-tls` focused on TLS setup and cert
|
||||
sharing, not transport accept logic.
|
||||
`Connection::from_bidi` → `endpoint.dispatch()`) lives in `alknet-core`
|
||||
behind a `tcp` feature, as an owned transport on `AlknetEndpoint` (via
|
||||
`with_tcp_tls(listener, acceptor)` — see ADR-083). `alknet-tls` is the
|
||||
cert provider, not the accept loop. This keeps `alknet-tls` focused on
|
||||
TLS setup and cert sharing, not transport accept logic.
|
||||
|
||||
### One ACME state machine, shared
|
||||
|
||||
|
||||
@@ -1,8 +1,12 @@
|
||||
# ADR-083: Endpoint as Pure Accept-Loop Runner with Public Dispatch
|
||||
# ADR-083: Endpoint as Multi-Transport Accept-Loop Runner with Public Dispatch
|
||||
|
||||
## Status
|
||||
|
||||
Proposed
|
||||
Proposed (revised 2026-07-14: TCP+TLS moved inside the endpoint as a
|
||||
third owned transport, not an external loop. This dissolves the
|
||||
multi-owner shutdown problem — the endpoint owns all its accept loops.
|
||||
`dispatch` stays public for genuinely external shapes: SSH channels,
|
||||
WebTransport streams. See §"TCP+TLS is a first-class owned transport".)
|
||||
|
||||
## Context
|
||||
|
||||
@@ -50,29 +54,45 @@ transport-specific parts are narrow bookends:
|
||||
|
||||
Everything between those bookends — the ACME ALPN guard, the handler
|
||||
lookup, the `build_auth_context` call, the `tokio::spawn` — is identical.
|
||||
Exposing a public `dispatch` method that takes the already-extracted
|
||||
ALPN, fingerprint, and `Connection` lets every transport's accept loop
|
||||
call the same dispatch path. No duplicated `build_auth_context` or
|
||||
handler-lookup logic at the assembly layer.
|
||||
|
||||
### TCP+TLS is a first-class transport, not a sibling afterthought
|
||||
### TCP+TLS is a first-class owned transport
|
||||
|
||||
TCP+TLS is the most common transport after raw-key QUIC, not an edge
|
||||
case: HTTPS for browsers, worker registration over HTTP, raw-key
|
||||
fallback for native clients when UDP is blocked. A hub — and any
|
||||
hub-worker — needs it.
|
||||
|
||||
ADR-010 Amendment 1 made TCP+TLS a "sibling accept loop" outside the
|
||||
endpoint because the endpoint's dispatch logic was crate-private and
|
||||
couldn't be shared. The assembly layer had to duplicate
|
||||
`build_auth_context` and handler lookup, or call `HandlerRegistry::get`
|
||||
directly. The "sibling" framing was a workaround for the endpoint being
|
||||
welded to quinn — not a deliberate design.
|
||||
couldn't be shared. The reason TCP+TLS was *structurally* excluded —
|
||||
the endpoint built transports internally, and TCP+TLS couldn't fit that
|
||||
shape — is gone after this ADR: the endpoint no longer builds
|
||||
transports. It runs accept loops on whatever it's given.
|
||||
|
||||
With a public `dispatch` method, the TCP+TLS accept loop calls into the
|
||||
endpoint — the same path quinn and iroh use. Amendment 1's *ownership*
|
||||
model (the TCP+TLS listener owns its own `TcpListener` and `TlsAcceptor`,
|
||||
living outside the endpoint struct) survives; its *dispatch* workaround
|
||||
(duplicated `build_auth_context` and handler-lookup logic) is retired.
|
||||
TCP+TLS is a listener transport, same shape as quinn and iroh:
|
||||
|
||||
TCP+TLS is the most common transport after raw-key QUIC, not an edge
|
||||
case: HTTPS for browsers, worker registration over HTTP, raw-key fallback
|
||||
for native clients when UDP is blocked.
|
||||
| Transport | What `with_*` takes | Accept loop |
|
||||
|-----------|---------------------|-------------|
|
||||
| quinn | `quinn::Endpoint` | `quinn.accept()` → handshake → `Connection::from_quinn_with_alpn` |
|
||||
| iroh | `iroh::Endpoint` | `iroh.accept()` → alpn+handshake → `Connection::from_iroh` |
|
||||
| TCP+TLS | `TcpListener` + `TlsAcceptor` | `tcp.accept()` → `tls.accept()` → `Connection::from_bidi` |
|
||||
|
||||
All three are: accept → extract ALPN + fingerprint → construct
|
||||
`Connection` → dispatch. The endpoint already runs the first two; the
|
||||
third is the same pattern with a different accept call. Making TCP+TLS
|
||||
an owned transport (via `with_tcp_tls(listener, acceptor)`) instead of
|
||||
an external loop gives the endpoint a single, uniform ownership model:
|
||||
it owns all its accept loops, `shutdown()` stops them all. No
|
||||
coordination between the endpoint and external loops.
|
||||
|
||||
This dissolves the multi-owner shutdown problem (the endpoint owns all
|
||||
loops; one owner, one shutdown). The `dispatch` method stays public —
|
||||
but for genuinely external shapes that the endpoint can't own: SSH
|
||||
channels (one connection, many channels with different ALPNs — a
|
||||
multiplexing shape, not a listener shape) and future WebTransport
|
||||
streams (one QUIC connection, many WT streams). These are not listener
|
||||
transports; they're connection-internal multiplexing. The endpoint
|
||||
can't own them the way it owns a `TcpListener`.
|
||||
|
||||
### The `acme-tls/1` guard
|
||||
|
||||
@@ -109,20 +129,25 @@ transports. `StaticConfig` becomes the canonical deployment config the
|
||||
in `alknet-core` (it's a config type) and stays extensible (future
|
||||
transport config fields land here as the assembly layer needs them).
|
||||
|
||||
This is a conceptual shift from ADR-010's framing ("the CLI binary
|
||||
constructs a `HandlerRegistry` and passes it, with `StaticConfig`, to
|
||||
`AlknetEndpoint::new()`"). The endpoint no longer takes `StaticConfig`
|
||||
at all. The endpoint takes `drain_timeout` and pre-built transport
|
||||
endpoints; the assembly layer is the `StaticConfig` consumer.
|
||||
The "assembly layer" is, in practice, the deployment binary — today,
|
||||
that's primarily the hub. A pure worker (no inbound endpoints) has a
|
||||
trivial assembly layer; a hub-worker combines both. Putting
|
||||
hub-specific composition in the hub crate is appropriate; putting
|
||||
transport loops that any node might need in the hub crate is not. The
|
||||
TCP+TLS loop belongs in `alknet-core` (behind a `tcp` feature), where
|
||||
any node that wants it can enable the feature and call `with_tcp_tls` —
|
||||
no hub dependency.
|
||||
|
||||
## Decision
|
||||
|
||||
### The endpoint is a pure accept-loop runner + a public dispatch method
|
||||
### The endpoint is a multi-transport accept-loop runner + a public dispatch method
|
||||
|
||||
```rust
|
||||
pub struct AlknetEndpoint {
|
||||
quinn: Option<quinn::Endpoint>,
|
||||
iroh: Option<iroh::Endpoint>,
|
||||
#[cfg(feature = "tcp")]
|
||||
tcp_tls: Option<TcpTlsListener>, // (TcpListener, TlsAcceptor)
|
||||
handlers: Arc<HandlerRegistry>, // ADR-010
|
||||
dynamic: Arc<ArcSwap<DynamicConfig>>, // ADR-010, unchanged
|
||||
identity_provider: Arc<dyn IdentityProvider>, // ADR-010, unchanged
|
||||
@@ -142,15 +167,29 @@ impl AlknetEndpoint {
|
||||
pub fn with_quinn(mut self, endpoint: quinn::Endpoint) -> Self;
|
||||
pub fn with_iroh(mut self, endpoint: iroh::Endpoint) -> Self;
|
||||
|
||||
/// Take ownership of a TCP+TLS listener. The endpoint runs the
|
||||
/// accept loop (`tcp.accept()` → `tls.accept()` → extract ALPN +
|
||||
/// fingerprint → `Connection::from_bidi` → `dispatch`). Feature-gated
|
||||
/// on `tcp` (pulls `tokio-rustls`). A hub serving HTTPS and a
|
||||
/// hub-worker serving TCP+TLS both use this; a pure worker doesn't.
|
||||
#[cfg(feature = "tcp")]
|
||||
pub fn with_tcp_tls(
|
||||
mut self,
|
||||
listener: tokio::net::TcpListener,
|
||||
acceptor: tokio_rustls::TlsAcceptor,
|
||||
) -> Self;
|
||||
|
||||
/// Clone of the shutdown watch sender. The assembly layer uses this
|
||||
/// to signal its own accept loops (TCP+TLS) to stop accepting — one
|
||||
/// signal, all loops stop. See OQ-61 for the full coordination model.
|
||||
/// to signal any external dispatch callers (SSH, future WT) to stop.
|
||||
/// The endpoint's own loops (quinn, iroh, TCP+TLS) are signaled
|
||||
/// internally by `shutdown()`.
|
||||
pub fn shutdown_sender(&self) -> watch::Sender<bool>;
|
||||
|
||||
/// Dispatch a pre-established connection by ALPN. Called by the
|
||||
/// endpoint's own accept loops (quinn, iroh) after transport-
|
||||
/// specific extraction, and by external accept loops (TCP+TLS,
|
||||
/// future SSH, future WebTransport) after their own extraction.
|
||||
/// endpoint's own accept loops (quinn, iroh, TCP+TLS) after
|
||||
/// transport-specific extraction, and by external dispatch callers
|
||||
/// (SSH channels, future WebTransport streams) after their own
|
||||
/// extraction.
|
||||
///
|
||||
/// Synchronous (non-async): performs the ACME guard, handler
|
||||
/// lookup, `build_auth_context`, and `tokio::spawn`s the handler.
|
||||
@@ -170,52 +209,83 @@ impl AlknetEndpoint {
|
||||
remote_addr: Option<SocketAddr>,
|
||||
);
|
||||
|
||||
/// Run the endpoint's owned accept loops. Returns when shutdown is
|
||||
/// signaled. The caller drives shutdown via `shutdown_sender()`.
|
||||
/// Run the endpoint's owned accept loops (quinn, iroh, TCP+TLS —
|
||||
/// whichever were configured). Returns when shutdown is signaled.
|
||||
/// The caller drives shutdown via `shutdown()` or `shutdown_sender()`.
|
||||
pub async fn run(self: Arc<Self>);
|
||||
|
||||
/// Signal all owned accept loops to stop, drain in-flight handlers
|
||||
/// for `drain_timeout`, then close. One owner, one shutdown — no
|
||||
/// external loop coordination needed.
|
||||
pub async fn shutdown(&self) -> Result<(), EndpointError>;
|
||||
}
|
||||
```
|
||||
|
||||
The consumer pattern: the assembly layer holds `Arc<AlknetEndpoint>`,
|
||||
clones the `Arc` for `run` (which consumes one clone), and passes
|
||||
`&endpoint` to its TCP+TLS accept loops so they can call `dispatch`.
|
||||
After `run` returns, the assembly layer drops its `Arc`; in-flight
|
||||
dispatched handlers (spawned by `dispatch` via `tokio::spawn`) are
|
||||
owned by the endpoint's runtime and drain per `shutdown` / OQ-61.
|
||||
The consumer pattern: the assembly layer (the deployment binary — today
|
||||
primarily the hub) holds `Arc<AlknetEndpoint>`, clones the `Arc` for
|
||||
`run` (which consumes one clone), and — for external dispatch callers
|
||||
like SSH — passes `&endpoint` to them so they can call `dispatch`. After
|
||||
`run` returns (on shutdown), in-flight dispatched handlers drain per
|
||||
`shutdown()`.
|
||||
|
||||
`AlknetEndpoint::new` takes **no `StaticConfig`** and **no TLS config**.
|
||||
The assembly layer (the deployment binary that composes crates — see
|
||||
ADR-014) reads `StaticConfig`, builds the transports, and hands the
|
||||
quinn/iroh endpoints to `AlknetEndpoint` via builder methods. The
|
||||
`handlers`, `dynamic`, and `identity_provider` fields are unchanged
|
||||
quinn/iroh/TCP+TLS endpoints to `AlknetEndpoint` via builder methods.
|
||||
The `handlers`, `dynamic`, and `identity_provider` fields are unchanged
|
||||
passthroughs from ADR-010 — this ADR does not change auth or dynamic
|
||||
config; it changes who builds transports and where dispatch lives.
|
||||
|
||||
`build_iroh_endpoint` moves out of `alknet-core/endpoint.rs`. Its
|
||||
destination is an open question (see OQ-60) — the assembly layer, an
|
||||
`alknet-tls` convenience helper, or a transport-construction
|
||||
module/crate. This ADR commits to the **boundary** (construction is
|
||||
not in the endpoint); the *where* is tracked separately so it doesn't
|
||||
block the endpoint-shape decision. (`build_quinn_server_config_from_rustls`
|
||||
is different: it's a thin wrapper that converts a `rustls::ServerConfig`
|
||||
into a `quinn::ServerConfig`, which is exactly `TlsServerConfig::for_quinn()`
|
||||
— its destination is `alknet-tls`, decided in ADR-082. Only
|
||||
`build_iroh_endpoint`, which reads `StaticConfig` and builds an
|
||||
`iroh::Endpoint` without a rustls config, is genuinely undecided.)
|
||||
### TCP+TLS is owned, not external
|
||||
|
||||
### `dispatch` is public and transport-agnostic
|
||||
The TCP+TLS accept loop runs inside `run()` alongside the quinn and
|
||||
iroh loops. It is not an external sibling. This means:
|
||||
|
||||
- **Shutdown is single-owner.** `endpoint.shutdown()` stops all owned
|
||||
loops (quinn, iroh, TCP+TLS) and drains all dispatched handlers. No
|
||||
coordination between the endpoint and external accept loops. The
|
||||
multi-owner shutdown problem (formerly OQ-61) does not arise.
|
||||
- **Any node can use TCP+TLS.** A hub, a hub-worker, or any node that
|
||||
accepts inbound TCP+TLS enables the `tcp` feature and calls
|
||||
`with_tcp_tls(listener, acceptor)`. No hub dependency. The loop isn't
|
||||
duplicated per binary.
|
||||
- **The TCP+TLS loop's home is `alknet-core`** (behind a `tcp` feature),
|
||||
not the hub crate or the assembly layer. Core already has `quinn` and
|
||||
`iroh` as feature-gated transport deps; adding `tcp` (pulling
|
||||
`tokio-rustls`) is the same pattern. A deployment that doesn't use
|
||||
TCP+TLS doesn't enable the feature.
|
||||
|
||||
### `dispatch` is public — for genuinely external shapes
|
||||
|
||||
The public `dispatch` method takes the already-extracted ALPN,
|
||||
fingerprint (if available), remote address (if available), and a
|
||||
`Connection`. It performs the ACME guard, the handler lookup, the
|
||||
`build_auth_context` call, and the `tokio::spawn`. `build_auth_context`
|
||||
becomes a private helper called by `dispatch`, not a standalone
|
||||
function — the assembly layer calls `dispatch`, which calls
|
||||
function — external callers call `dispatch`, which calls
|
||||
`build_auth_context` internally.
|
||||
|
||||
`dispatch` is for transports the endpoint **can't own** — shapes that
|
||||
aren't listener-based:
|
||||
|
||||
- **SSH channels** (ADR-065): one SSH connection carries multiple
|
||||
channels with different ALPNs. The SSH handler accepts one connection,
|
||||
then dispatches each channel via `dispatch` with the channel's
|
||||
ALPN. This is connection-internal multiplexing, not a listener loop.
|
||||
- **Future WebTransport streams** (parked per ADR-044): one QUIC
|
||||
connection, many WT streams. Same multiplexing shape.
|
||||
|
||||
TCP+TLS is **not** an external dispatch caller — it's a listener
|
||||
transport the endpoint owns. The distinction: listener transports
|
||||
(quinn, iroh, TCP+TLS) produce connections from an accept loop the
|
||||
endpoint runs; multiplexing transports (SSH, WT) produce connections
|
||||
from within an existing connection, and the endpoint can't own their
|
||||
accept loop.
|
||||
|
||||
Transport-specific extraction (`extract_quinn_alpn`,
|
||||
`extract_quinn_client_fingerprint`, `extract_iroh_client_fingerprint`)
|
||||
stays private in the endpoint — the accept loops call them, then call
|
||||
`extract_quinn_client_fingerprint`, `extract_iroh_client_fingerprint`,
|
||||
`extract_tcp_tls_alpn`, `extract_tcp_tls_client_fingerprint`) stays
|
||||
private in the endpoint — the accept loops call them, then call
|
||||
`dispatch` with the results.
|
||||
|
||||
### The `acme-tls/1` guard moves to `dispatch`
|
||||
@@ -242,6 +312,22 @@ location changes. The `acme-tls/1` ALPN append still happens in
|
||||
transport using the ACME `TlsServerConfig` advertises it, but in
|
||||
practice only the TCP+TLS listener on 443 receives challenges.
|
||||
|
||||
### Feature gates
|
||||
|
||||
```toml
|
||||
[features]
|
||||
quinn = ["dep:quinn"] # with_quinn — quinn accept loop
|
||||
iroh = ["dep:iroh"] # with_iroh — iroh accept loop
|
||||
tcp = ["dep:tokio-rustls"] # with_tcp_tls — TCP+TLS accept loop
|
||||
acme = ["dep:rustls-acme"] # (used by alknet-tls, not core directly)
|
||||
```
|
||||
|
||||
The `tcp` feature pulls `tokio-rustls`. A deployment enables the
|
||||
features for the transports it runs. A pure-QUIC node enables `quinn` +
|
||||
`iroh`; a hub serving HTTPS enables `quinn` + `tcp`; a hub-worker
|
||||
enables all three. This matches the existing pattern (`quinn` and
|
||||
`iroh` are already feature-gated transport deps on `alknet-core`).
|
||||
|
||||
### `StaticConfig` stays in core; the endpoint drops it
|
||||
|
||||
`StaticConfig` remains in `alknet-core/config.rs` as the canonical
|
||||
@@ -252,6 +338,33 @@ Future transport config fields (addresses, relay URLs, identity) land in
|
||||
signature does not change when they're added, because the endpoint
|
||||
doesn't read them.
|
||||
|
||||
### Transport construction: inlined by the assembly layer
|
||||
|
||||
The builder functions are trivial API calls. The assembly layer (the
|
||||
deployment binary — today primarily the hub crate's composition code)
|
||||
inlines them:
|
||||
|
||||
- `build_quinn_endpoint`: `tls_config.for_quinn()` → `quinn::Endpoint::server(addr)` — 2 lines
|
||||
- `build_iroh_endpoint`: iroh builder + key + relay + alpns + bind — 15 lines
|
||||
- `build_tcp_tls`: `TcpListener::bind(addr)` + `tls_config.for_tcp_tls()` — 2 lines
|
||||
|
||||
These are pure configuration, not shared logic. No helper crate or
|
||||
module — 20 lines total across all three, each binary picks which
|
||||
transports it wants. If a future binary duplicates the iroh builder and
|
||||
the pattern drifts, extraction is a two-way door on the function; the
|
||||
*loop* (the runtime with shutdown + dispatch) is in core, so the
|
||||
duplication risk is only on the trivial builder, not the real component.
|
||||
|
||||
Hub-specific composition (wiring `ChannelsAdapter`, `HttpAdapter`,
|
||||
`CallAdapter`, the relay, peer lifecycle, worker registration) lives in
|
||||
the hub crate. That's where multi-transport *composition* happens. The
|
||||
endpoint provides the accept loops; the hub provides the handlers and
|
||||
wiring. A `alknet-transport` crate was considered and rejected — it
|
||||
would contain only trivial builder functions (the real component, the
|
||||
TCP+TLS loop, is in core). A crate for 20 lines of API calls doesn't
|
||||
earn its existence. If hub-specific transport helpers accumulate, they
|
||||
live in the hub crate, not a generic transport crate.
|
||||
|
||||
### What moves out of `alknet-core/endpoint.rs`
|
||||
|
||||
| Code | Destination |
|
||||
@@ -264,19 +377,18 @@ doesn't read them.
|
||||
| `SelfSignedCert` / `generate_self_signed_cert()` | `alknet-tls` (ADR-082) |
|
||||
| `load_cert_chain()` / `load_private_key()` | `alknet-tls` (ADR-082) |
|
||||
| `build_quinn_server_config_from_rustls()` | `alknet-tls` (`for_quinn()`, ADR-082) |
|
||||
| `build_iroh_endpoint()` | Out of core (destination: OQ-60) |
|
||||
| `build_iroh_endpoint()` | Assembly layer (inlined; 15 lines of iroh API calls) |
|
||||
|
||||
### What stays in `alknet-core/endpoint.rs`
|
||||
|
||||
- `AlknetEndpoint` struct (accept-loop runner + public `dispatch`)
|
||||
- `AlknetEndpoint` struct (multi-transport accept-loop runner + public `dispatch`)
|
||||
- `HandlerRegistry`
|
||||
- `dispatch` (public — ACME guard, handler lookup, `build_auth_context`,
|
||||
spawn)
|
||||
- `dispatch_quinn` / `dispatch_iroh` (private — transport-specific
|
||||
extraction, then call `dispatch`)
|
||||
- `run_quinn_accept_loop` / `run_iroh_accept_loop`
|
||||
- `dispatch` (public — ACME guard, handler lookup, `build_auth_context`, spawn)
|
||||
- `dispatch_quinn` / `dispatch_iroh` / `dispatch_tcp_tls` (private — transport-specific extraction, then call `dispatch`)
|
||||
- `run_quinn_accept_loop` / `run_iroh_accept_loop` / `run_tcp_tls_accept_loop`
|
||||
- `extract_quinn_alpn` / `extract_quinn_client_fingerprint` /
|
||||
`extract_iroh_client_fingerprint`
|
||||
`extract_iroh_client_fingerprint` / `extract_tcp_tls_alpn` /
|
||||
`extract_tcp_tls_client_fingerprint`
|
||||
- `build_auth_context` (private helper, called by `dispatch`)
|
||||
|
||||
### ADR-082 amendment
|
||||
@@ -284,19 +396,26 @@ doesn't read them.
|
||||
ADR-082's `AlknetEndpoint::new(..., Arc<TlsServerConfig>, ...)` signature
|
||||
is superseded by this ADR. The endpoint takes no TLS config; the
|
||||
assembly layer builds transports from `TlsServerConfig`s and hands them
|
||||
to the endpoint via `with_quinn` / `with_iroh`. ADR-082 should reference
|
||||
this ADR for the endpoint signature and focus on what `alknet-tls`
|
||||
provides (`TlsServerConfig` and its accessors).
|
||||
to the endpoint via `with_quinn` / `with_iroh` / `with_tcp_tls`. ADR-082
|
||||
should reference this ADR for the endpoint signature and focus on what
|
||||
`alknet-tls` provides (`TlsServerConfig` and its accessors).
|
||||
|
||||
## Consequences
|
||||
|
||||
**Positive:**
|
||||
- The endpoint has one job: dispatch. Transport construction lives
|
||||
outside it, where the multi-`TlsServerConfig` hub case is natural.
|
||||
- TCP+TLS dispatch is first-class — same `dispatch` path as quinn/iroh,
|
||||
no duplicated `build_auth_context` or handler-lookup logic at the
|
||||
assembly layer. ADR-010 Amendment 1's second-class-dispatch workaround
|
||||
is retired.
|
||||
- The endpoint has one job: run accept loops + dispatch. Transport
|
||||
construction lives outside it, where the multi-`TlsServerConfig` hub
|
||||
case is natural.
|
||||
- TCP+TLS is a first-class owned transport — same `run()` loop, same
|
||||
`shutdown()`, same dispatch path as quinn/iroh. No duplicated
|
||||
`build_auth_context` or handler-lookup logic anywhere. ADR-010
|
||||
Amendment 1's sibling-dispatch workaround is retired entirely.
|
||||
- Shutdown is single-owner. The endpoint owns all its loops; one
|
||||
`shutdown()` stops them all and drains. The multi-owner shutdown
|
||||
problem (formerly OQ-61) does not arise.
|
||||
- Any node can use TCP+TLS — enable the `tcp` feature, call
|
||||
`with_tcp_tls`. No hub dependency. A hub-worker serving TCP+TLS is
|
||||
the same code path as a hub serving TCP+TLS.
|
||||
- The `acme-tls/1` guard is transport-agnostic — it works for TCP+TLS
|
||||
(where challenges actually arrive) and any future transport. ADR-027
|
||||
§5's quinn-specific location is corrected.
|
||||
@@ -304,51 +423,61 @@ provides (`TlsServerConfig` and its accessors).
|
||||
the endpoint's. Adding transport config fields doesn't churn the
|
||||
endpoint's `new` signature.
|
||||
- The hub (the first multi-transport consumer) is unblocked: build two
|
||||
`TlsServerConfig`s, build quinn/iroh/TCP+TLS listeners, hand the
|
||||
endpoint the quinn/iroh ones, spawn the TCP+TLS loops calling
|
||||
`endpoint.dispatch`.
|
||||
`TlsServerConfig`s, build quinn/iroh/TCP+TLS listeners, hand all
|
||||
three to the endpoint, `run()`.
|
||||
- `dispatch` stays public for genuinely external shapes (SSH channels,
|
||||
future WT streams) — transports the endpoint can't own because they're
|
||||
connection-internal multiplexing, not listener-based.
|
||||
|
||||
**Negative:**
|
||||
- `AlknetEndpoint::new` signature changes (breaking). Pre-1.0, in-repo
|
||||
consumers only — the assembly layer and tests must update. Expected;
|
||||
this is the point of the refactor.
|
||||
- `build_iroh_endpoint` leaves core. Its destination is an open question
|
||||
(OQ-60), not decided here. The boundary (not in the endpoint) is
|
||||
committed; the *where* is not, so that the endpoint-shape decision
|
||||
isn't blocked on the transport-construction-location decision.
|
||||
(`build_quinn_server_config_from_rustls` is decided — it moves to
|
||||
`alknet-tls` as `for_quinn()` per ADR-082; only `build_iroh_endpoint`
|
||||
is open.)
|
||||
- Multi-owner shutdown: the endpoint owns shutdown of *dispatched
|
||||
handlers* (it spawned them in `dispatch`); the assembly layer owns
|
||||
shutdown of the *accept loops it spawned* (TCP+TLS listeners). The
|
||||
coordination mechanism (shared `shutdown_sender`, drain semantics) is
|
||||
a follow-up design point, not resolved by this ADR — see OQ-61.
|
||||
- `alknet-core` gains a `tcp` feature (pulls `tokio-rustls`). This is
|
||||
the same pattern as the existing `quinn` and `iroh` features — a
|
||||
transport that the endpoint can own. A deployment that doesn't use
|
||||
TCP+TLS doesn't enable it. `alknet-core` already depends on `rustls`
|
||||
(for `fingerprint.rs` types, per OQ-59); `tokio-rustls` is the
|
||||
acceptor wrapper over `rustls::ServerConfig`, not a separate TLS
|
||||
stack.
|
||||
- `build_iroh_endpoint` leaves core (inlined by the assembly layer).
|
||||
This is a one-way dep-graph change: the binary that uses iroh depends
|
||||
on `iroh` directly, not via `alknet-core`. This is correct — the
|
||||
binary *is* the thing that knows which transports it wants. The 15
|
||||
lines are pure iroh API calls; no shared logic is lost.
|
||||
- This ADR revises ADR-010's "TCP is not an endpoint struct concern"
|
||||
more deeply than the original ADR-083 draft. The reason TCP was
|
||||
excluded (the endpoint built transports internally, TCP+TLS couldn't
|
||||
fit) is gone; the endpoint is now a multi-transport accept-loop runner
|
||||
and TCP+TLS is a listener transport that fits the same shape.
|
||||
|
||||
## Door type
|
||||
|
||||
**One-way.** The endpoint's `new` signature, the public `dispatch`
|
||||
contract, and the "endpoint owns dispatch, not construction" boundary
|
||||
are structural. Reversing would mean re-welding transport construction
|
||||
to the endpoint and re-privatizing `dispatch` — breaking every
|
||||
multi-transport consumer (hub, future HTTP, future SSH).
|
||||
contract, the "endpoint owns dispatch, not construction" boundary, and
|
||||
the "endpoint owns TCP+TLS as a first-class transport" decision are
|
||||
structural. Reversing would mean re-welding transport construction to
|
||||
the endpoint, re-privatizing `dispatch`, and pushing TCP+TLS back outside
|
||||
— breaking every multi-transport consumer (hub, hub-worker, future SSH,
|
||||
future WT).
|
||||
|
||||
The `dispatch` signature (`connection, alpn, fingerprint, remote_addr`)
|
||||
is one-way — changing it after consumers exist is a rewrite. The
|
||||
internal implementation (how extraction is factored, how `run` spawns
|
||||
tasks) is two-way.
|
||||
tasks, the TCP+TLS loop's internal structure) is two-way.
|
||||
|
||||
## References
|
||||
|
||||
- ADR-010: ALPN router and endpoint (amended — the endpoint no longer
|
||||
constructs transports; Amendment 1's second-class-dispatch workaround
|
||||
is retired by the public `dispatch` method)
|
||||
constructs transports; "TCP is not an endpoint struct concern" is
|
||||
revised: TCP+TLS is now a first-class owned transport via
|
||||
`with_tcp_tls`; Amendment 1's sibling-dispatch workaround is retired
|
||||
by the public `dispatch` method and the owned TCP+TLS loop)
|
||||
- ADR-014: Secret material flow and capability injection (defines the
|
||||
"assembly layer" term — the deployment binary that composes crates)
|
||||
- ADR-010 Amendment 1: TCP+TLS as sibling (superseded — Amendment 1's
|
||||
*ownership* model (sibling listener outside the endpoint) survives;
|
||||
its *dispatch* workaround (duplicated logic) is retired by this ADR's
|
||||
public `dispatch`)
|
||||
- ADR-010 Amendment 1: TCP+TLS as sibling (superseded — TCP+TLS is now
|
||||
an owned transport, not a sibling; `dispatch` is public for SSH/WT,
|
||||
not for TCP+TLS)
|
||||
- ADR-027 §5: ACME ALPN challenge handling (location amended — guard
|
||||
moves from `dispatch_quinn` to shared `dispatch`)
|
||||
- ADR-065: `Connection::from_stream`/`from_bidi` (the primitive that
|
||||
@@ -362,7 +491,8 @@ tasks) is two-way.
|
||||
- `docs/research/alknet-endpoint-refactor/findings.md` — the analysis
|
||||
that surfaced the conflation and the two-config hub case
|
||||
- `crates/alknet-core/src/endpoint.rs` — the code being refactored
|
||||
- OQ-60: Where does transport construction live? (assembly layer,
|
||||
`alknet-tls` helper, or transport module/crate)
|
||||
- OQ-61: Multi-owner shutdown coordination (endpoint owns dispatched
|
||||
handlers; assembly layer owns spawned accept loops)
|
||||
- OQ-60: resolved — transport construction is inlined by the assembly
|
||||
layer (the deployment binary); the TCP+TLS loop lives in
|
||||
`alknet-core` behind a `tcp` feature as an owned transport
|
||||
- OQ-61: dissolved — the multi-owner shutdown problem does not arise;
|
||||
the endpoint owns all its accept loops (quinn, iroh, TCP+TLS)
|
||||
@@ -77,8 +77,8 @@ Door type is separate from whether a decision is made. A two-way door is a decis
|
||||
| [OQ-14](questions/014-batch-operation-semantics.md) | Batch Operation Semantics | resolved | two | low |
|
||||
| [OQ-55](questions/055-alknetclient-establishment-extraction.md) | AlknetClient / Client Establishment Extraction | deferred(scope) | two | med |
|
||||
| [OQ-59](questions/059-fingerprint-module-location.md) | Should `fingerprint.rs` Stay in Core or Move to `alknet-tls`? | open | two | med |
|
||||
| [OQ-60](questions/060-transport-construction-location.md) | Where Does Transport Construction Live? | open | one | high |
|
||||
| [OQ-61](questions/061-multi-owner-shutdown-coordination.md) | Multi-Owner Shutdown Coordination | open | two | med |
|
||||
| [OQ-60](questions/060-transport-construction-location.md) | Where Does Transport Construction Live? | resolved | one | high |
|
||||
| [OQ-61](questions/061-multi-owner-shutdown-coordination.md) | Multi-Owner Shutdown Coordination | dissolved | two | med |
|
||||
|
||||
### alknet-call
|
||||
|
||||
|
||||
@@ -2,49 +2,69 @@
|
||||
|
||||
- **Origin**: `docs/architecture/decisions/083-endpoint-as-accept-loop-runner.md`
|
||||
(the endpoint refactor commits to the boundary — construction is not
|
||||
in the endpoint — but not the location of `build_iroh_endpoint`).
|
||||
- **Status**: open
|
||||
- **Door type**: one-way (where `build_iroh_endpoint` lives determines
|
||||
who depends on `iroh` for transport construction; moving it later
|
||||
churns the dep graph and every binary's assembly code)
|
||||
in the endpoint — but not the location of transport construction).
|
||||
- **Status**: resolved
|
||||
- **Door type**: one-way (where `build_iroh_endpoint` and the TCP+TLS
|
||||
loop live determines who depends on `iroh` / `tokio-rustls` for
|
||||
transport construction; the dep-graph shape is structural)
|
||||
- **Priority**: high (the hub is the first multi-transport consumer;
|
||||
its assembly code sets the pattern)
|
||||
- **Blocked on**: nothing structural. The three options are clear; the
|
||||
decision is a trade-off, not a missing capability.
|
||||
- **Scope**: This OQ covers **`build_iroh_endpoint`** — the function
|
||||
that reads `StaticConfig` and builds an `iroh::Endpoint`.
|
||||
`build_quinn_server_config_from_rustls` is **decided** (it moves to
|
||||
`alknet-tls` as `TlsServerConfig::for_quinn()` per ADR-082 — it's a
|
||||
thin wrapper over a `rustls::ServerConfig`, which `alknet-tls` owns).
|
||||
Only `build_iroh_endpoint` is genuinely undecided: it reads
|
||||
`StaticConfig` (not a rustls config), builds an `iroh::Endpoint` from
|
||||
an `Ed25519SecretKey` + relay URL, and doesn't fit the `alknet-tls`
|
||||
cert-provider boundary.
|
||||
- **Resolution**: Not yet decided. The options:
|
||||
- **Resolution**: Split answer (resolved by ADR-083 revision, 2026-07-14):
|
||||
|
||||
**Option A: In the assembly layer (the binary).** Each binary reads
|
||||
`StaticConfig` and hand-assembles quinn/iroh/TCP+TLS. Pro: maximal
|
||||
flexibility, core stays lean. Con: every binary duplicates the
|
||||
"build an iroh endpoint from an `Ed25519SecretKey` + relay URL"
|
||||
boilerplate; a 10-step procedure is copy-pasted per binary.
|
||||
**TCP+TLS accept loop → `alknet-core` behind a `tcp` feature (owned
|
||||
by the endpoint).** The endpoint takes a `TcpListener` +
|
||||
`TlsAcceptor` via `with_tcp_tls(listener, acceptor)` and runs the
|
||||
accept loop inside `run()` alongside the quinn and iroh loops. TCP+TLS
|
||||
is a listener transport — same shape as quinn and iroh (accept →
|
||||
extract ALPN + fingerprint → `Connection::from_bidi` → `dispatch`).
|
||||
Making it owned gives the endpoint a single uniform ownership model:
|
||||
it owns all its accept loops, `shutdown()` stops them all. The
|
||||
multi-owner shutdown problem (OQ-61) does not arise. Any node that
|
||||
wants TCP+TLS enables the `tcp` feature and calls `with_tcp_tls` — no
|
||||
hub dependency. This reverses ADR-010's "TCP is not an endpoint struct
|
||||
concern": the reason TCP was excluded (the endpoint built transports
|
||||
internally, TCP+TLS couldn't fit) is gone; the endpoint is now a
|
||||
multi-transport accept-loop runner and TCP+TLS fits the same shape.
|
||||
|
||||
**Option B: In `alknet-tls` as convenience helpers.** `alknet-tls`
|
||||
gains a `build_iroh_endpoint` helper. Pro: one place. Con: `alknet-tls`
|
||||
then depends on `iroh`, bloats a crate whose stated job is "TLS setup,
|
||||
not transport endpoint construction" (ADR-082 scopes `alknet-tls` to
|
||||
cert config and its accessors — `build_iroh_endpoint` doesn't touch
|
||||
certs, it touches iroh's relay/secret-key APIs).
|
||||
**Builder functions (`build_iroh_endpoint`, `build_quinn_endpoint`,
|
||||
`build_tcp_tls`) → inlined by the assembly layer.** These are trivial
|
||||
API calls (2-15 lines each, pure configuration, no shared logic). The
|
||||
assembly layer (the deployment binary — today primarily the hub
|
||||
crate's composition code) inlines them. No helper crate or module —
|
||||
20 lines total across all three. A `alknet-transport` crate was
|
||||
considered and rejected: it would contain only trivial builders (the
|
||||
real component, the TCP+TLS loop, is in core). Hub-specific transport
|
||||
helpers, if they accumulate, live in the hub crate, not a generic
|
||||
transport crate.
|
||||
|
||||
**Option C: A new crate or module — `alknet-transport` / an
|
||||
`alknet-core::transport` module — that owns transport construction.**
|
||||
Pro: clean separation (TLS = certs, transport = endpoints, endpoint =
|
||||
dispatch). Con: another layer in the dep graph.
|
||||
**`dispatch` stays public** — but for genuinely external shapes (SSH
|
||||
channels, future WebTransport streams), not for TCP+TLS. The
|
||||
distinction: listener transports (quinn, iroh, TCP+TLS) produce
|
||||
connections from an accept loop the endpoint owns; multiplexing
|
||||
transports (SSH, WT) produce connections from within an existing
|
||||
connection, and the endpoint can't own their accept loop.
|
||||
|
||||
The ADR-082 boundary ("`alknet-tls` is the cert provider, not the
|
||||
transport constructor") is a strong argument against Option B. Option C
|
||||
is the cleanest separation but adds a layer. Option A is simplest but
|
||||
risks per-binary duplication that drifts.
|
||||
- **Cross-references**: ADR-083 (endpoint refactor), ADR-082
|
||||
(`alknet-tls` — the cert provider boundary; `for_quinn()` is in scope,
|
||||
`build_iroh_endpoint` is not), ADR-010 (original endpoint design,
|
||||
where construction was welded to the endpoint)
|
||||
**Why not the hub crate for the TCP+TLS loop:** "the assembly layer"
|
||||
is, in practice, usually the hub. But a hub-worker serving TCP+TLS
|
||||
shouldn't depend on the hub crate for a transport loop. The loop
|
||||
belongs in core (behind a feature), where any node can use it. The hub
|
||||
crate owns hub-specific *composition* (wiring adapters, relay, peer
|
||||
lifecycle), not transport runtimes.
|
||||
|
||||
**Why not `alknet-tls`:** `alknet-tls`'s stated job is TLS setup —
|
||||
`rustls::ServerConfig`, cert resolvers, ACME (ADR-082). The TCP+TLS
|
||||
accept loop is transport runtime, not TLS setup. It calls
|
||||
`endpoint.dispatch()`, which is core's API. Adding it to `alknet-tls`
|
||||
would make a cert-provider crate depend on the endpoint's dispatch API
|
||||
and own a transport runtime — a category error.
|
||||
|
||||
**Why not `alknet-transport` (new crate):** The real component (the
|
||||
TCP+TLS loop) is in core. What's left for a transport crate is trivial
|
||||
builder functions. A crate for 20 lines of API calls doesn't earn its
|
||||
existence. If future transport runtimes accumulate that don't fit
|
||||
core, a transport crate can be created then — but not speculatively.
|
||||
- **Cross-references**: ADR-083 (endpoint refactor — the revision that
|
||||
resolved this), ADR-082 (`alknet-tls` — the cert provider boundary),
|
||||
ADR-010 (original endpoint design — "TCP is not an endpoint struct
|
||||
concern" is revised), OQ-61 (dissolved — the multi-owner shutdown
|
||||
problem does not arise with TCP+TLS owned by the endpoint)
|
||||
@@ -2,41 +2,28 @@
|
||||
|
||||
- **Origin**: `docs/architecture/decisions/083-endpoint-as-accept-loop-runner.md`
|
||||
(the endpoint owns dispatched handlers; the assembly layer owns
|
||||
spawned accept loops — coordination between them on shutdown is
|
||||
spawned accept loops — coordination between them on shutdown was
|
||||
unspecified).
|
||||
- **Status**: open
|
||||
- **Door type**: two-way (the coordination mechanism is an
|
||||
implementation detail; the ownership boundary — endpoint owns
|
||||
dispatched handlers, assembly layer owns spawned accept loops — is
|
||||
the one-way part, committed in ADR-083)
|
||||
- **Priority**: medium (matters for clean hub shutdown; doesn't block
|
||||
the endpoint-shape decision or the TLS extraction)
|
||||
- **Blocked on**: nothing structural. The question is which
|
||||
coordination primitive (shared `shutdown_sender`, separate channel,
|
||||
drain semantics) — a design choice, not a missing capability.
|
||||
- **Resolution**: Not yet decided. The boundary is clear:
|
||||
- **Status**: dissolved
|
||||
- **Door type**: two-way (the coordination mechanism was an
|
||||
implementation detail)
|
||||
- **Priority**: medium
|
||||
- **Resolution**: Dissolved by ADR-083 revision (2026-07-14). The
|
||||
problem this OQ tracked — coordinating shutdown between the endpoint's
|
||||
owned loops and external TCP+TLS accept loops — does not arise. The
|
||||
TCP+TLS accept loop is now an owned transport (via `with_tcp_tls`),
|
||||
not an external sibling. The endpoint owns all its accept loops
|
||||
(quinn, iroh, TCP+TLS); `shutdown()` stops them all and drains all
|
||||
dispatched handlers. One owner, one shutdown. The `dispatch` method
|
||||
stays public for SSH channels and future WebTransport streams, but
|
||||
those are connection-internal multiplexing callers, not listener
|
||||
loops with independent shutdown needs — they call `dispatch` on
|
||||
connections the endpoint already accepted and dispatched a handler
|
||||
for.
|
||||
|
||||
- The endpoint owns shutdown of **dispatched handlers** (it spawned
|
||||
them via `tokio::spawn` in `dispatch`).
|
||||
- The assembly layer owns shutdown of the **accept loops it spawned**
|
||||
(the TCP+TLS listeners).
|
||||
|
||||
The open question is the **coordination mechanism**:
|
||||
|
||||
- Does the assembly layer use the endpoint's `shutdown_sender()` to
|
||||
signal the TCP+TLS loops, or a separate channel?
|
||||
- Does `endpoint.shutdown()` drain in-flight dispatched handlers
|
||||
(including those from TCP+TLS loops), or only the quinn/iroh ones
|
||||
it spawned via `run()`?
|
||||
- What happens to in-flight `dispatch` calls after the endpoint is
|
||||
shut down — are they rejected, or do they complete?
|
||||
|
||||
A likely shape: the assembly layer uses the endpoint's
|
||||
`shutdown_sender()` for all its accept loops (one signal, all loops
|
||||
stop accepting); the endpoint's `shutdown()` drains all dispatched
|
||||
handlers regardless of which transport spawned them (the endpoint
|
||||
owns them all once `dispatch` is called). But this needs to be
|
||||
written down and verified against the drain semantics.
|
||||
- **Cross-references**: ADR-083 (endpoint refactor — ownership
|
||||
boundary), ADR-010 (original shutdown design — single-owner, the
|
||||
simpler case being generalized)
|
||||
The premise of this OQ (external TCP+TLS loop + endpoint-owned
|
||||
dispatch = multi-owner shutdown) was a consequence of TCP+TLS being
|
||||
external. ADR-083's revision made TCP+TLS internal, retiring the
|
||||
premise.
|
||||
- **Cross-references**: ADR-083 (endpoint refactor — TCP+TLS is now
|
||||
owned), OQ-60 (resolved — the TCP+TLS loop lives in core)
|
||||
Reference in new issue
Block a user