All work

Developer Tooling

Agent Session Observability

Watching a dozen AI coding sessions without tabbing through windows

2026Working tool

At a glance

Credit

Built by
Denandro YusufDesign · Build contract · Implementation · Directing AI coding agents

Run the demo synthetic data

Problem
Several AI coding sessions at once, and no way to see which one was blocked.
What I did
I built a zero-dependency dashboard that reads each session's state from files already on disk.
Outcome
Status, progress, a rough ETA and equivalent cost for every session, at a glance.

Explore

The component graph, from the source repository. Follow the arrows: left to right, and down where two steps share a column. The burgundy bars are where decisions are made; the numbers on the lines are the key flows, written out under the graph.

Agent Session Observability: architectureData flows through 13 components: Session files on disk (Source: Registry, JSONL transcripts, IDE locks); Process table (Source: Liveness + child processes); Sidecar agent transcripts (Source: Background work after the main turn); Incremental digest (Ingest: Streams from stored byte offset); Liveness + activity (Ingest: Guards against recycled pids); Agent scanner (Ingest: mtime walk, then bounded head/tail); State · ETA · cost (Decision logic: 13-rule machine, confidence band); fs.watch + safety poll (Ingest: 150ms debounce, self-rearming); Digest cache (Store: Atomic write + rename; mtime revalidated); Assembly + hash (Service: Broadcasts only on real change); HTTP + SSE on loopback (Service: 6 JSON routes, never routable); Vanilla JS dashboard (Interface: Inline-SVG charts, no build step); Tray · badge · notifications (Output: Consumes its own loopback SSE). Connections: Session files on disk to Incremental digest (appended bytes); Session files on disk to Liveness + activity; Process table to Liveness + activity (verify pid); Sidecar agent transcripts to Agent scanner; Incremental digest to Digest cache (persist / revalidate); Incremental digest to Assembly + hash (digest); Liveness + activity to Assembly + hash; Agent scanner to Assembly + hash (agent activity); Assembly + hash to State · ETA · cost; State · ETA · cost to Assembly + hash (state, ETA, cost); fs.watch + safety poll to Assembly + hash (debounced refresh); Assembly + hash to HTTP + SSE on loopback (change event); HTTP + SSE on loopback to Vanilla JS dashboard (REST + SSE); HTTP + SSE on loopback to Tray · badge · notifications (loopback SSE).Session files on disk — Registry, JSONL transcripts, IDE locksSession files ondiskRegistry, JSONLtranscripts, IDE locksProcess table — Liveness + child processesProcess tableLiveness + childprocessesSidecar agent transcripts — Background work after the main turnSidecar agenttranscriptsBackground work afterthe main turnIncremental digest — Streams from stored byte offsetIncremental digestStreams from storedbyte offsetLiveness + activity — Guards against recycled pidsLiveness + activityGuards againstrecycled pidsAgent scanner — mtime walk, then bounded head/tailAgent scannermtime walk, thenbounded head/tailState · ETA · cost — 13-rule machine, confidence bandState · ETA · cost13-rule machine,confidence bandfs.watch + safety poll — 150ms debounce, self-rearmingfs.watch + safetypoll150ms debounce,self-rearmingDigest cache — Atomic write + rename; mtime revalidatedDigest cacheAtomic write + rename;mtime revalidatedAssembly + hash — Broadcasts only on real changeAssembly + hashBroadcasts only onreal changeHTTP + SSE on loopback — 6 JSON routes, never routableHTTP + SSE onloopback6 JSON routes, neverroutableVanilla JS dashboard — Inline-SVG charts, no build stepVanilla JSdashboardInline-SVG charts, nobuild stepTray · badge · notifications — Consumes its own loopback SSETray · badge ·notificationsConsumes its ownloopback SSE12345678910

Key flows

  1. Session files on disk to Incremental digest: appended bytes
  2. Process table to Liveness + activity: verify pid
  3. Incremental digest to Digest cache: persist / revalidate
  4. Incremental digest to Assembly + hash: digest
  5. Agent scanner to Assembly + hash: agent activity
  6. State · ETA · cost to Assembly + hash: state, ETA, cost
  7. fs.watch + safety poll to Assembly + hash: debounced refresh
  8. Assembly + hash to HTTP + SSE on loopback: change event
  9. HTTP + SSE on loopback to Vanilla JS dashboard: REST + SSE
  10. HTTP + SSE on loopback to Tray · badge · notifications: loopback SSE

Go deeper

The full account

Problem
Running several AI coding sessions at once leaves no way to know which one is working, which is silently blocked on a permission prompt, which errored and which finished — short of tabbing through every terminal and editor window. There is no status API. The only signal is the session registry, the process table, and multi-megabyte JSONL transcripts on disk.
What was built
A zero-dependency Electron dashboard that derives full session state purely by observing files — an incremental transcript digester, a 13-rule state machine, per-model cost arithmetic, and a probabilistic ETA with a confidence band.
Role
Architect and the only human engineer. Wrote the build contract that pinned the module shapes, directed AI coding agents to implement the modules against it in parallel, then reviewed and integrated them.
What changed
Answers at a glance what previously required opening every window: what is each session doing, how far along, roughly how much longer, and what the token usage would have cost.

Context

This is the personal project, and it is the one that best explains how I work: the problem was mine, the observation surface was whatever the tool already wrote to disk, and the constraint was that nothing could be instrumented, injected or sent anywhere.

Transcripts reach tens of megabytes, so any naive read-the-file-and-report approach is too slow to poll. That single performance constraint shaped the entire architecture.

Architecture and the system

Pure observation. Three read-only sources feed incremental ingesters, a derivation layer computes state, ETA and cost, and a change-hashed store pushes over SSE to a framework-free UI and an Electron shell that consumes its own loopback stream.

Read only what changed

The transcript digester streams only the bytes appended since the last tick, from a stored byte offset, and persists digests to a local cache keyed on size and mtime. A cold scan of a large transcript is a one-time cost; every subsequent tick is nearly free.

Digests are written atomically — write then rename — so a crash mid-write cannot leave a corrupt cache behind.

Thirteen rules, in a fixed order

State — working, waiting for input, blocked on a permission prompt, stalled, errored, idle, done — is resolved by a numbered rule sequence, with the order itself part of the spec. Where the implementation deviates from the contract, the module header says so and argues the case.

A separate scanner watches each session's sidecar agent directories, so a session whose main turn has ended but whose background agents are still running is correctly reported as busy rather than finished.

Cost arithmetic that respects the pricing model

Cache reads, five-minute cache writes and one-hour cache writes carry different multipliers, so a blended rate is simply wrong. The price table resolves suffixed model ids by longest prefix, applies date-bounded intro pricing per message timestamp, attributes per-model within mixed-model transcripts, and excludes synthetic placeholder messages.

Degrade, not die

Timeouts fence every per-session scan, refresh is re-entrancy guarded, and every I/O path has a degraded handler. The server binds 127.0.0.1 only and is never routable, which is why it needs no auth — a decision documented in the module header rather than assumed.

The store broadcasts only when a clock-independent hash of the payload actually changes, so an idle machine produces no traffic and no repaints.

What was hard

Polling a file that grows to tens of megabytes

Re-reading is not an option at poll frequency. Streaming from a stored byte offset, skipping pathologically large single lines, and caching digests against size and mtime is what makes the whole dashboard viable — every other feature depends on that one decision.

A session that looks finished but is not

When the main turn ends but background agents keep running, the obvious signal says done. Watching the sidecar agent directories — with bounded head and tail reads rather than full parses — is what makes the status honest.

A build contract instead of a codebase

The spec pins the exact serialized shape, the numbered rule order, the ETA algorithm, performance budgets, the full route table, and the design tokens down to hex values — enough that modules could be implemented independently against it. That is what let the work parallelise, and it is the same technique used on the larger platform projects.

AI and automation

AI

The app makes no model calls and holds no API key. It is an observability layer over AI agent sessions, not an AI product — its AI-specific work is domain modelling of the pricing and the session lifecycle. It is also the clearest artifact of spec-driven, AI-assisted development: a written build contract, parallel implementation, then integration.

Automation

Native notifications fire on transitions into waiting-for-input, blocked, errored and done, with per-session debounce and a startup quiet window so pre-existing sessions are seeded rather than firing retroactively. The packaging script builds a real double-clickable macOS app — bundle, Info.plist, icns, ad-hoc signing — without any packaging dependency.

Direction and delivery

The spec is written as a binding build contract: exact serialized shapes, numbered rule order, performance budgets (a cold scan of ~116 MB under six seconds, never blocking the event loop more than ~50ms), the full HTTP route table, and design tokens. Nearly every module opens with a header explaining why — including the decisions that look wrong until explained: why the shell is CommonJS, why it reads its own loopback stream instead of importing the store, why the watches are non-recursive.

Size signals

Lines of source
~15,700
Runtime dependencies
0
State machine rules
13
HTTP endpoints
6 + SSE

Counted from the source repository. No impact metric is claimed that the source does not prove.

What I would tell the next person

  • A performance budget written into the spec changes the architecture. 'Under six seconds cold' is what forced incremental digestion.
  • Document the decisions that look wrong. A reader who cannot see the constraint will 'fix' them.
  • Not shipping auth can be correct — if you can state precisely why the surface is unreachable, and you write that down.

Technologies and access

  • Node.js 22 (ESM)
  • Electron 34
  • Server-Sent Events
  • node:fs streams
  • JSONL stream parsing
  • Vanilla JavaScript
  • Inline SVG
  • Bash

Integrations

  • Local session registry and JSONL transcripts
  • macOS process table
  • macOS tray, dock and notification APIs

Local-only tool. Binds 127.0.0.1 and is never routable.