The application kernel for TypeScript —
boot a @btravstack/di module into a running
process, and stop it again without losing work.
di proves an application's wiring before the process exists. start owns
when that already-proven graph is constructed and torn down, and nothing
more: one lifecycle state machine, one unit-of-work registry, one Runtime
contract. It knows nothing about HTTP, AMQP or Temporal.
It owns the things every backend process gets wrong on its own — a graceful
drain that survives Kubernetes' eventually-consistent endpoint removal,
liveness and readiness that answer from the state machine rather than a
transport, and a teardown that runs on every path. Nothing throws: start
returns an unthrownResult, and it
never calls process.exit.
pnpm add @btravstack/start @btravstack/di unthrown@btravstack/di and unthrown are peer dependencies — install all three.
The kernel itself has no runtime dependencies beyond node: builtins.
import{Module,Port,Provider}from"@btravstack/di";import{runMain,start,typeRuntime,typeServing}from"@btravstack/start";import{Ok,OkAsync}from"unthrown";classGreeterextendsPort("Greeter")<{readonlygreet: (name: string)=>string;}>{}constAppModule=Module("App")({provides: [Provider(Greeter)({value: {greet: (name: string)=>`hello, ${name}`}}),],exports: [Greeter],});// A runtime owns the transport; the kernel owns the lifecycle. This one is a// timer, so the sample stays self-contained — no published runtime models a// timer, and `@btravstack/start-http` would pull in a real dependency this// sample doesn't need.constticker: Runtime<typeofGreeter>={name: "ticker",needs: [Greeter],start: (host)=>{consttimer=setInterval(()=>{// Every piece of work goes through `host.run`: that is what makes it// count towards the drain, and what gives it an `AbortSignal`.//// The unit's `Result` is the runtime's to map — the kernel hands it back// and stays out of it. A timer has nowhere to return one, so it observes// it instead; dropping it would hide the work's `Err` *and* a `Defect`.voidhost.run({kind: "tick",id: `${Date.now()}`},(ctx,signal)=>signal.aborted ? Ok("") : Ok(ctx.get(Greeter).greet("world")),).tapFailure((failure)=>{process.stderr.write(`${JSON.stringify({tick: failure.tag})}\n`);});},1_000);constserving: Serving={// Stop accepting new work. In-flight units are the kernel's business.drain: ()=>{clearInterval(timer);returnOkAsync();},stop: ()=>OkAsync(),};returnOkAsync(serving);},};awaitrunMain(start(AppModule,{runtime: ticker}));start returns immediately with a RunningApp; runMain awaits its
exited and turns the outcome into a process exit code. The runtime's declared
needs are checked against the module's exports at compile time — booting
ticker against a module that does not export Greeter is a type error at the
start call, not a boot-time crash.
Every code sample on this page is compiled by
packages/start/src/docs-examples.test-d.ts,
so a sample that stops compiling fails the build.
NestJS's NestFactory.create(AppModule)is the wiring step: Nest reads
decorator metadata at runtime, resolves tokens, builds the injector graph, and
throws at boot when something is missing. The graph is discovered while the
process starts.
di proves the graph before the process exists, so start does not wire.
| NestJS | di + start | |
|---|---|---|
| Wiring declared by | decorators + metadata reflection | explicit Module/Provider values |
| Missing dependency | boot-time exception | compile error |
| Module privacy | enforced at runtime | compile error |
| Cycles | forwardRef() | detected pre-construction, as a defect |
| Request scope | request-scoped providers bubble up the injection chain | forkScope — parent services seeded, not reconstructed |
| Lifecycle hooks | 5 interfaces + enableShutdownHooks() | onStart/onStop per provider; start owns signals |
| Failures | thrown | Result |
The accepted cost is that there is no auto-discovery: the provider and its dependency array are written out, and that array is what buys the compile-time checking.
The kernel knows several kinds of runtime. A process boots exactly one.
An api deployment, a consumer deployment and a worker deployment are three
processes booting the same module with a different runtime. They scale, fail
and deploy independently, and it removes a whole class of design problem: there
is never a question of how two runtimes in one process share a drain deadline,
or whose failure takes the process down.
examples/ proves this rather than asserting it: one
clean-architecture application, booted by an oRPC runtime and by a queue-worker
runtime, with the application and persistence layers unchanged between them —
and the same DuplicateOrder arriving as a typed CONFLICT on one and as a
dead-letter on the other.
typeRuntime<NeedsextendsAnyPort,Info=never>={readonlyname: string;readonlyneeds: readonlyNeeds[];readonlystart: (host: RuntimeHost<Needs>,)=>AsyncResult<Serving<Info>,RuntimeStartFailed>;};typeRuntimeHost<NeedsextendsAnyPort>={readonlyctx: Context<InstanceType<Needs>>;readonlyrun: RunUnit<Needs>;};typeServing<Info=never>={readonlydrain: (signal: AbortSignal)=>AsyncResult<void,never>;readonlystop: ()=>AsyncResult<void,never>;readonlyinfo?: Info;};A runtime receives a host, not a bare Context: it needs both the
application services and the kernel's run, and handing it a Context alone
would leave it inventing its own unit tracking — the thing the kernel exists to
own.
Serving.drain returns void, not a report. Only the kernel can see the unit
registry, so the kernel — not the runtime — owns the accounting. drain means
"stop accepting"; the AbortSignal it receives fires when the kernel's own
deadline passes, so a runtime never does arithmetic on time.
Runtime, RuntimeHost and RunUnit are parameterised by port classes
(Needs extends AnyPort) but hand out Context<InstanceType<Needs>>, because
di parameterises Context<in R> by port instance types.
A runtime that binds port: 0 knows which port it got, and nothing else does.
Serving.info is the channel for that, and RunningApp.runtimeInfo() is where
the caller reads it — so no runtime has to invent an onListening hook of its
own.
typeHttpInfo={readonlyport: number};consthttpish: Runtime<typeofGreeter,HttpInfo>={name: "httpish",needs: [Greeter],start: ()=>OkAsync({drain: ()=>OkAsync(),stop: ()=>OkAsync(),// Whatever the runtime actually bound. A queue consumer has no port and// would publish `{ queue, prefetch }` instead — the shape is its own.info: {port: 8080},}),};constapp=start(AppModule,{runtime: httpish});constinfo=awaitapp.runtimeInfo();// Result<HttpInfo | undefined, never>Info defaults to never, so publishing is optional: a runtime with
nothing to say omits info and its type is unchanged. runtimeInfo() is
probePort() one layer up — the same deferred, settled when the runtime starts
serving and undefined on every route that never gets there, so it can never
hang.
Every runtime does the same thing in a loop: take one piece of work, run it, produce an outcome. The kernel names that a unit and owns it, so three runtimes do not each invent it.
constsubmitOne=(run: RunUnit<typeofGreeter>,meta: UnitMeta,): AsyncResult<string,never>=>run(meta,(ctx,signal)=>signal.aborted ? Ok("") : Ok(ctx.get(Greeter).greet("world")),);run is transparent to the work's own channels: whatever Result the handler
produces is what the runtime receives back. The kernel observes only that the
unit settled — and in-flight tracking falls out of this. The kernel counts
open units, so drain means only "stop accepting" and DrainReport.abandoned
is accurate without any cooperation from the runtime.
The kernel never maps an outcome to a transport.Result → HTTP status
belongs to the handler an application hands the HTTP runtime (oRPC, Hono, a
bare function — @btravstack/start-http itself declines that mapping),
Result → ack/nack/DLQ to the AMQP runtime, Result → activity failure to
the Temporal runtime. The kernel hands back the Result and stays out of it.
Per-unit ports are not wired yet: run currently hands the work the
applicationContext. RunUnit is typed so a Module.forkScope call can
land there without a signature change.
Neither is checkable by the kernel, and building the first real runtime on this contract hit both.
Flush the response inside the unit. A unit is closed the instant its
Result settles; an idle registry is what the drain waits for, and going idle
is the kernel's permission to call Serving.stop(). So a runtime that resolves
the unit and then writes its response is racing stop() tearing the transport
down. With a small body the write usually wins; with a large one it does not
(measured with an 8 MB body: UND_ERR_SOCKET: other side closed). A unit is not
"compute the answer" — it is "compute the answer and get it out of the
process".
constserveOne=(host: RuntimeHost<typeofGreeter>,meta: UnitMeta,send: (body: string)=>Promise<void>,): AsyncResult<string,never>=>// Flushed inside the work callback. Sending after `await host.run(...)`// returns is the race: the unit is already closed by then.host.run(meta,async(ctx,signal)=>{constbody=signal.aborted ? "" : ctx.get(Greeter).greet("world");awaitsend(body);returnOk(body);});UnitMeta.id must be unique per unit, unless you pass a traceId — that is
what traceId defaults to. An HTTP runtime that submits the route template
("POST /orders") as the id gives every request the same trace id, and the
ambient record's whole purpose, telling one unit apart from another in a log
line, is silently defeated. A route template is a kind, not an id; a broker
message id or a queue job id is already unique and needs nothing more.
The kernel cannot check this — it would have to remember every id it had ever
seen. What it does guarantee is UnitRecord.unitId, minted per unit and always
unique, so a reader that only needs to tell two units apart already has one.
traceId is the correlation id, which is why it is the one a runtime may
supply: it carries an id from outside the process (a traceparent header, a
message property) so a line logged here joins a trace that started elsewhere.
Not only the fallible ones. AsyncResult<T, never> is how this package spells
"async, and cannot fail" — exactly what fromSafePromise produces — so
app.probePort(), clock.sleep(ms), clock.advance(ms),
registry.awaitIdle(), runtime.untilStarted() and a probe server's close()
all await into a Result. A caller never has to remember which async surfaces
returned a Result and which returned a bare value.
Three surfaces are deliberately outside it:
runMainreturnsPromise<void>. Its job is to leave the Result world and become a process exit code — it is the boundary.UnitWork'sPromise<Result<T, E>>arm, which exists to accept your ownasynchandler.withAppand itsusecallback.useis the test body: a thrown assertion failure must reach the test runner, and anAsyncResultnever rejects — so wrapping it would turn a failingexpectinto aDefectyou can forget to unwrap, which is a green test that asserted nothing.
building ──▶ starting ──▶ serving ──▶ draining ──▶ stopping ──▶ exited
│ │ │ ▲
└────────────┴────────────┴────────────────────────┘
(any failure short-circuits to stopping)
The tracker is monotonic: a phase can only ever move forward, and re-entering one is a no-op.
- Readiness flips
false, and the unit counts are sampled — synchronously, before anything else. - The kernel waits
preDrainDelayMs(default5_000) before telling the runtime to stop accepting. - In-flight work gets
drainTimeoutMs(default20_000) to finish; whatever is still open at that deadline is aborted and reported asabandoned.
Beat 2 looks like a pointless sleep and is not. Kubernetes endpoint removal is eventually consistent, so a pod that stops accepting the instant SIGTERM lands rejects traffic the ingress is still routing to it. That window is what the delay closes.
drainTimeoutMs sits deliberately under the Kubernetes
terminationGracePeriodSeconds default of 30s, leaving headroom for stopping
before SIGKILL. Raise one and you must raise the other.
Only a signal drains. A plain stop() and an uncaught exception both go
straight to stopping, leaving ExitReport.drainundefined. A second
signal skips the drain — double Ctrl-C in development, an operator's escape
hatch in production.
Draining produces a value, not a log line:
typeDrainReport={readonlyinFlightAtStart: number;// units in flight when the drain beganreadonlycompleted: number;// units that closed during the drainreadonlyabandoned: number;// units still open at the deadline};completed may exceed inFlightAtStart if in-flight work spawned more units
during the drain. That is honest reporting, not a bug — it is counted from a
monotonic total precisely so it can never go negative.
Liveness and readiness are process-level concerns, not transport-level ones, so
the kernel runs its own node:http probe server on a separate port (default
9000, probes: false to disable, { port: 0 } to let the OS choose and read
it back from app.probePort(), an AsyncResult<number | undefined, never>):
| Route | 200 | 503 |
|---|---|---|
GET /livez | ok — any phase before exited | unavailable |
GET /readyz | ready — serving, and not forced unready | unavailable |
Anything else answers 404. The server binds 127.0.0.1 only and is
unref'd, so it never keeps the event loop alive. A bind failure is a startup
failure: it stops the graph being built at all, and surfaces as
RuntimeStartFailed({ runtime: "probes" }).
This is how a Temporal worker pod with no HTTP runtime still gets probes, and
why an HTTP runtime never has to expose /healthz on the public port. There is
deliberately no separate startup probe: /livez answers from building onward,
so a slow-building graph is covered by /readyz alone.
Readiness is a one-way latch — once forced false, by a drain or an uncaught
exception, it never returns to true. app.ready() reads the same predicate
synchronously, which is what an embedder wires into a health endpoint of its own
when probes: false.
The kernel opens one AsyncLocalStorage store per unit holding a small, fixed
record — { unitId, traceId, tenantId, deadline } — and nothing else. Services
never go in it.
constlog=(message: string): void=>{constunit=currentUnit();process.stderr.write(`${JSON.stringify({ message,traceId: unit?.traceId})}\n`,);};The line holds because what di exists to prevent is hidden dependencies:
code that secretly needs a collaborator it never declared and cannot be tested
without it. A trace id is not a collaborator — there is no substitutability
question and no test double. A repository pulled from an ambient store is the
untestable coupling; a tenant id read by the Postgres adapter is not.
Legitimate readers are infrastructure adapters only — logger, OTel exporter, database adapter. Application code reading the store is meant to be a lint error; that rule is not written yet (it needs a convention for identifying an adapter, which this stack has not established), so for now it is a convention, not an enforcement.
runMain is the single sanctioned place this package decides a process's fate.
It sets process.exitCode and never calls process.exit(), so pending
output is flushed and an embedding host keeps control of its own lifetime.
| Code | Meaning |
|---|---|
0 | exited cleanly, with nothing abandoned and teardown clean |
1 | startup failure (a modeled Err) |
2 | drained with work abandoned, or exited with teardown errors |
70 | stopped by an uncaught exception or unhandled rejection |
70 | a defect |
The two 70s are the same statement — sysexits(3)'s EX_SOFTWARE, an internal
software error — reached through the two channels a bug can take. A crash takes
precedence over abandoned work.
2 is the one code an operator reads as "we stopped, but not cleanly", and two
facts earn it: work the drain ran out of time for, and a finaliser that failed
on the way out. The second matters as much as the first — a connection pool that
could not flush is exactly the shutdown an orchestrator must not be told
succeeded — which is why a non-empty ExitReport.teardownErrors is never a 0.
start installs uncaughtException and unhandledRejection handlers, and
installing either suppresses Node's own default exit code of 1. So an
embedder that uses startwithoutrunMain and sets no exit code of its own
gets a silent exit 0 after a crash.
Use runMain, or decide the code yourself:
constembed=async(): Promise<void>=>{constapp=start(AppModule,{runtime: ticker,signals: true});constreport=awaitapp.exited;process.exitCode=report.match({ok: (exit)=>(exit.reason==="uncaught" ? 70 : 0),errCases: (matcher)=>matcher.with(P.tag("RuntimeStartFailed"),()=>1),defect: ()=>70,});};(signals: false turns off the uncaught handlers and the signal handlers
together, which is the other way out — at the cost of no signal-driven drain.)
Failures are classified by phase, and each phase has one honest channel.
- Startup — a modeled
Err.startreturnsAsyncResult<ExitReport, E | RuntimeStartFailed>, whereEis the application module's own error type, passed through unwrapped and still typed. The kernel adds onlyRuntimeStartFailed, which is genuinely its own (a port in use, a broker unreachable, a probe port taken). - Wiring — a
Defect, untouched. A cycle or a duplicate provider arrives fromdias a defect and stays one. A wiring bug is not something a caller branches on. - Teardown — visible, never masking.
diguarantees a failing finaliser cannot overwrite the real failure; the kernel collects them intoExitReport.teardownErrors, and a failing release never rewrites the reason the application stopped. - Unit failures never reach the kernel. A handler's
Erris the runtime's to map. uncaughtException/unhandledRejection— readiness false, then straight tostopping, skipping the drain. Deliberately harsher than the signal path: after an uncaught throw the process state may be corrupt, so draining in-flight work risks completing it wrongly. Half-finished correct work beats confidently-wrong finished work. Only the first one is reported.
The kernel emits structured events and takes no logger dependency:
| Event | Payload |
|---|---|
building | — |
serving | runtime |
draining | inFlight |
drained | report |
stopping | — |
exited | — |
teardownError | port, cause |
uncaught | cause |
The default sink writes one JSON line per event to stderr. A throwing sink is swallowed: a broken reporter must not take the process down mid-shutdown.
@btravstack/start/testing ships the deterministic half of the lifecycle.
constdrainTest=async(): Promise<void>=>{constclock=createFakeClock();construntime=testRuntime();constreport=awaitwithApp(AppModule,{ runtime, clock },async(app)=>{awaitruntime.untilStarted();constunit=runtime.submit<string>();app.requestDrain();awaitclock.advance(5_000);// the pre-drain delayunit.settle(Ok("done"));// The unit's own outcome, asserted rather than awaited for its timing// alone — a bare `await unit.result;` would drop it.expect(awaitunit.result).toBeOkWith("done");returnawaitapp.exited;});// `{ inFlightAtStart: 1, completed: 1, abandoned: 0 }`.expectTypeOf(report.getOrThrow().drain).toEqualTypeOf<DrainReport|undefined>();};createFakeClock()— time moves only when the test says so, so a drain test is instant rather than twenty-five seconds long.advancebrackets itself with a real macrotask at each end, so the code under test has reacted by the time it resolves.testRuntime()— an in-memoryRuntimethat lets units be held open, so drain behaviour can be proved with no transport.withApp(module, options, use)— starts, hands the app touse, and stops it again whateverusedoes.signalsandprobesare forced off whatever the caller passes: process-wide handlers would fight across a test file, and a probe port would collide between tests. It rethrows aDefectonexited, so a shutdown that blew up fails the test even whenusenever looked atexited; a modeledErrpasses through untouched, being an outcome a test may legitimately be asserting. A test that wants to assert the defect itself callsstartdirectly.
There is deliberately no overrideProvider. Swapping an adapter is composing a
different module, which di already documents and the type checker already
verifies.
The Runtime contract is the whole of what this package owes the transports.
Three have shipped. @btravstack/start-http: bind, one
unit per request, a drain that retires busy keep-alive connections, stop —
routing, middleware and Result → HTTP status are deliberately not included,
see its README's "What it does not do".
@btravstack/start-temporal: a Temporal worker,
one unit per activity attempt, and a drain that releases the kernel at the
kernel's deadline rather than Temporal's shutdownForceTime.
@btravstack/start-amqp: an amqp-contract worker,
one unit per delivery, and a drain with exactly one deadline — the library is
told to wait forever, so there is no second timeout to keep in sync at all.
The rest are planned, not published:
| Planned package | Would own |
|---|---|
| an observability package | logger and OpenTelemetry, binding to KernelEvent |
The ticker runtime above is a complete one, in roughly forty lines. The
three shipped packages are not: they exist because the lifecycle underneath a
real transport — a listener, a worker, a broker connection — is not forty
lines done well.
See packages/start for the package README,
examples/ for an eleven-package clean-architecture application
booted under four different runtimes, and CLAUDE.md for the
authoritative spec: the theses, the public surface and the conventions. The
load-bearing invariants with the test that guards each, and the internal design
notes, live in
packages/start/CLAUDE.md.
MIT © Benoit TRAVERS