Architecture
The graph needs a clock.
A snapshot can only say what is true right now. We rebuilt Rhei's graph so every answer also says since when, on what evidence, and which work changed it.
Rhei Team7 min read
Most coding agents principally operate on source. Agent-memory products principally operate on continuity. Observability products principally operate on causality.
Rhei is trying to bind all three while preserving their independent authority.
What a snapshot can’t answer
Every repository index answers one question well: what is true right now. Ask it what changed since Tuesday, and it rebuilds. Ask whether an edge is a fact or a guess, and it has no opinion. Ask which of yesterday’s conclusions survived this morning’s refactor, and there is nothing to ask. Ask who made a change and why, and it points at git blame and hopes for the best. Freshness, in most tools, means re-scanning everything and hoping nobody noticed the gap.
An agent returning after 50 hours of work should not need the full transcript. It should see which repository state it stands on, what moved while it was gone, and which of its old beliefs still hold. We ran into this with our own agents. A graph that can’t carry time can’t carry any of that, so we put the time dimension inside the graph itself, next to the symbols and edges. This post is what that looks like, including the parts that don’t work yet.
The approach, end to end
Rhei now has two connected layers. Rhei Intelligence answers what is true about the repository now. Rhei Delta Intelligence adds time and work: WorkSessions, GoalRuns, agent activity, intent, chapters, tool events, edits, and validation evidence, compiled into a typed, privacy-bounded work trace. Authority flows one way, down the stack. Source and Git identity feed a Rust kernel that publishes immutable generations. An evidence layer sits on top of it and decides which cross-service edges get to exist. Above that, a work graph folds typed events about what humans and agents did onto the exact generations those events touched. Clients read sealed answers off the bottom, each pinned to one generation, so two tools never reason from different states.
The Delta story binds the work graph to the flow diff between the exact generations the work touched. Readers never see a mixed state. The CLI reads this projection today; public MCP and API delivery follow the current gates.
The bet: three orthogonal axes
It also means Rhei is quietly becoming something different from the original “code intelligence” concept. The working model has three orthogonal axes, and each keeps its own authority:
| Axis | Question | Mechanism |
|---|---|---|
| Source | What exists? | FlowDiff / Git / code graph |
| Causality | Why did it happen? | WorkTrace / WorkGraph / Chapters |
| Continuity | What must the next actor know? | Cursor / D0-D4 / DeltaStory |
SOURCE
▲
/ \
/ \
/ \
CONTINUITY ───── CAUSALITYEach axis keeps its own authority inside one boundary: generations own source identity, the verifier decides which edges are facts, the work graph owns causal attribution, and cursors own what the next actor sees. That is the bet we continue to make.
Generations, not snapshots
The kernel never mutates a published graph. A commit SHA plus the digests of the files that landed produce a new immutable generation, and readers keep reading the old one until the new one seals. Derived capabilities don’t have to land together: search and relations publish first, dead code and reuse follow in explicit epochs, and every read names which capabilities were actually ready. A caller can tell “not supported yet” apart from “checked, found nothing”, which sounds small until you’ve watched an agent act on the wrong one.
The space between two generations is data, not damage. A flow diff exposes the added, removed, moved and unchanged nodes between sealed states as checkable facts. A client holding last week’s generation asks what moved instead of re-scanning the world, and an edit re-verifies only the conclusions it actually invalidated.
Warm updates keep this cheap. When a lineage proof shows the underlying store unchanged, a new generation publishes as an overlay on top of it rather than a revalidated copy of the entire database. On our own repository that path measures 94 ms p95, against 890 ms p95 for the full-validation route. Publication is only one part of a CRUD operation, but it’s the part we used to pay on every commit.
Edges that carry their proof
Above the kernel sits a harder question: which relationships deserve to exist at all. Take service connectivity, the question of which repo talks to which service over which port. We used to answer it the way most tooling answers it, with weighted similarity. Path resemblance scored 0.70, symbol names 0.50, shared graph relations 0.30, date overlap 0.20. That machinery spent every two-second cycle collision-correcting eighty-plus connections across 500 mappers and five builders, and the final checkpoint still held 75 stale ABORTED verdicts. It guessed fluently. It proved nothing.
We deleted it rather than leave it behind a flag: the detector, the separately generated candidate artifact, the TypeScript policy interpreting its output, and a truth-candidate coupling that existed only for this one edge type, all cut in a single commit. What replaced it is a verifier with one acceptance boundary. An edge is admitted only from syntax, resolved symbols, manifests, callsites, literals, control and data flow, pinned to an exact repository identity, and it carries its minimal proof subgraph along with counterevidence and invalidation rules. The verdict is always one of four states: verified, candidate_set, abstained, not_applicable. No edge is ever “mostly true”. Fuzzy signals may still propose candidates. They can never verify one.
Because the evidence graph is temporal, a connection keeps its identity across provider migrations and renames instead of dying and coming back as a fresh guess. We’re not claiming detection is finished; coverage and calibration are open work, and we’re still measuring both. The boundary itself is fixed, though. Belief costs proof.
The work graph
The layer above the evidence graph is where the work lives. Agents leave transcripts, and a transcript is the worst format awareness can take: unverifiable, unordered, padded with reasoning that never touched a file. So the runtime records typed events instead. A session opened, a tool ran, a file was edited, a commit landed, one lane forked from another, a human redirected the goal. Each event is producer-observed and carries digests of what it touched.
Rust alone decides what gets in. An event that fails generation, sequence, privacy, ownership or digest validation isn’t stored and flagged for later review; it never exists. Surviving events pack into content-addressed journal segments, retained at the same boundary where GoalRuns already keep their artifacts, and ordered by causal links rather than wall-clock time. A dropped tool result and a gap between sessions both render as unknown. No interpolation, no plausible story.
Fold the events and you get a WorkGraph: intent, agents, attribution, the who and the why. Bind it against the flow diff between the generations the work touched and you get the what. That binding is the Delta story. Attribution inside it comes from receipts only. The current compiler renders source changes as unattributed because the native producer has not persisted an exact matching receipt yet. It will name a contributor only after that owner supplies the binding.
Depth, not transcripts
The flow interface exposes only the depth an agent needs: current state, continuation, causality, exact proof, or paged history. Orientation should stay cheap. The next lifecycle slice will attach the bounded view at resume, compaction, and handoff; agents already request deeper levels through the explicit flow read. Delivered is not yet understood, though. A planned understanding cursor will separate context an agent received from context it acknowledged, and until that lands we won’t claim the second.
Where the numbers stand
Some of this is frozen in the current beta and some is still characterization evidence on one repository, so here is the honest ledger. Every row comes from a measured run in our own tree.
| Measure | Before | Now | Reading |
|---|---|---|---|
| Native graph p95 (671k edges) | 0.179 ms | 0.168 ms | kernel held; we cut work around it |
| Prefix search p95 | 188.69 ms | 9.32 ms | 20.2× faster |
| Natural search p95 | 220.72 ms | 9.51 ms | 23.2× faster |
| Immutable publication p95 | 890 ms | 94 ms | lineage-proved overlay |
| Continuity scale run | n/a | 782 segments | 100k events, 500 simulated hours |
| Cold-fold replay | target 35.8 s | 94.7 s | missed, named task |
OpenClaw characterization on Apple arm64. The scale run measured bounded correctness, memory, retention and zero spill, which all passed and left 96 events after compaction. The cold-fold row failed its target, and it stays in the table next to the wins because a failed measurement is still a measurement.
What an agent gets back
An agent should inherit the work, not the transcript. Reconnecting, it reads the generation it stands on, the diff since the generation it knew, the intent and chapters of the work in flight, which evidence still applies, and which questions the graph refuses to answer. That last line matters most to us. A system that never says “I don’t know” trains everything sitting on top of it to never ask.
What isn’t proven yet
The weak spots are named. Exact search and Define regressed slightly in the latest run and need guardrails. Impact and reuse still spend more time hydrating evidence than walking the graph. The cold-fold gap is open, cross-tool benchmark claims stay unpublished until we have an identity-matched packaged build, and the 240-repository continuity benchmark is still ahead of us. Read-only federation and the other successor slices each wait behind their own authorization. We’ll publish those runs the same way as this table: receipts, including the ugly ones.
Where this goes
Three legs, in order. First, finished foundations: close the cold-fold gap in the open, keep the 94 ms overlay path, and let the Delta continuity benchmark go green before we publish anything from it. Second, agent-trustable reads: flow becomes a public tool, the gates hold across 240 repositories, and understanding cursors separate delivered context from acknowledged. Four independent review rounds stand between any number and its publication. Third, product integration: Spotlight becomes a real client of the same projection, the harness drives agents through the same automatic orientation and explicit depths, and every surface we add consumes the same authority.
The direction stays simple. Things change and entropy increases, so we built Rhei to follow, notice, and adapt: agents act on facts instead of transcripts, and the humans who own the outcome decide and redirect. The graph keeps the clock.
