A clue found at one stop can make a completely different route useful three stops later. That is the behavior I wanted to test in a graph retrieval system: let the evidence collected along a path change the decision at the next junction.

I built a small open-source prototype with Jev to explore it. The result is a bounded, branching graph walk with inspectable decisions, recorded model responses, and deliberately awkward test cases. There is also a browser replay so you can follow the experiment without an API key.

An investigator consults a notebook at a branching network of bridges; a remembered clue illuminates possible routes ahead.
Evidence gathered on the way changes which path looks useful next. AI-generated conceptual illustration.

A graph walk that remembers why it arrived

Imagine following a case through a collection of connected records. An early handoff contains a reference code, T91. Later, a registry offers several routes. The original question does not contain that code. The useful choice becomes visible because an earlier record supplied it.

If every junction is scored against only the original question and the current node, the system can lose the connection. If it carries the relevant evidence forward, the later decision has a different basis.

That gives us a simple loop: read a node, retain its source facts, inspect the outgoing edges, score their usefulness with the accumulated context, and continue along a limited number of promising paths. Each newly visited node changes the context for the next decision.

The prototype implements the retrieval stage of a possible GraphRAG pipeline. It returns evidence and traversal traces. It does not generate a final prose answer. This boundary makes it easier to see whether the required sources were actually found before judging the fluency of an answer built from them.

Jev supplies judgments; code controls the walk

TypeSafe’s typed primitives make a useful interface for this experiment. The application asks Jev for constrained judgments and then applies ordinary code to those results.

In this implementation, each candidate edge receives a usefulness score. Several edges can be useful at the same time: an incident question might need an operational record, an ownership record, and a policy record. A single winning edge can discard necessary evidence even when its local ranking looks reasonable.

The walker therefore supports branching and a bounded frontier. Code enforces the maximum depth, node count, candidate count, model requests, and context size. A separate sufficiency check asks whether the collected evidence supports each requested requirement. Its confidence threshold is a stopping policy, not proof that an answer is correct.

Two kinds of state matter. Each active path carries the source facts and relationships that explain how it got there. A shared evidence ledger records what the traversal has collected. Keeping those roles distinct lets the system combine evidence across branches without pretending every branch followed the same route.

Cycles are excluded within a path. When two paths reach the same node, their contexts remain distinct. That is necessary if arrival history can affect the next decision. It also creates more work, so the frontier still needs explicit limits. The design notes describe these choices and their constraints.

What the recorded experiment shows

The published snapshot contains 55 runs across 11 cases and control variants, using five traversal strategies. It records 74 unique real Jev requests with model version jev-1.13.0; identical requests are reused rather than billed again. The repository includes the fixtures, recordings, results, and replay checks.

For the five answerable core fixtures, the results were:

StrategyCases with all required source recordsTotal nodes visited across the five cases
Breadth-first search5/543
Lexical routing4/525
Jev, greedy path3/518
Jev, without prior path context4/523
Jev, cascading path context5/525

“Complete” means all designated source records were retrieved. It does not mean a generated answer was correct; there is no answer generator in this experiment. These are small, authored mechanism tests, developed iteratively, rather than a held-out benchmark. The full results include the intentionally failing controls.

Three observations are useful enough to investigate further.

First, branching mattered in the incident fixture. The cascading walker retrieved all three required records, while the greedy walker retrieved one. A question that requires evidence from several parts of a graph is an awkward fit for a policy that repeatedly commits to one route.

Second, carrying context mattered in the relay fixture. The cascading walker retained the earlier reference and found the required evidence; the variant without prior path context did not. Lexical routing also succeeded on the ordinary relay case. The experiment isolates a failure caused by forgetting a clue; it does not establish that model-based routing is always necessary to follow one.

Third, context retention can become the failure point. In a deliberately constrained variant, a 320-character fact budget removed the route clue before the registry decision. The cascading walker then failed too. The relevant implementation question is which evidence survives the budget, and whether the system can recover evidence it discarded.

The failures are part of the design

A scoring model cannot select an edge that the application never exposes. A disconnected fixture and a candidate cap that hides the useful route both defeat all five strategies. A shallow hop budget prevents them from reaching the answer as well. Graph connectivity, candidate generation, and traversal limits define the model’s opportunity to help.

Sufficiency introduces another uncertainty. In the negation fixture, the walker collected the required sources, but the sufficiency judgment stayed around 0.83, below the configured 0.9 threshold. It eventually stopped because no useful edges remained. Retrieving the evidence and confidently recognizing that it is enough are separate tasks.

More context did not uniformly improve efficiency either. Adding traversed relationships to the context increased the cascading walk on the hub fixture from four nodes to five. The reconvergence case checks that distinct path contexts survive a merge, but the variant without prior context also succeeds there. That fixture verifies a state-management property; it does not demonstrate an accuracy advantage from preserving those contexts.

These details make the trace valuable. A failed walk should expose whether it exhausted its budget, lost a clue, excluded a candidate, or made an unhelpful judgment. Otherwise, a single success rate conceals several different engineering problems.

Fewer graph visits is only one part of the cost

The cascading variant matched breadth-first retrieval coverage on these five core cases while visiting fewer nodes. That is a reason to investigate the approach on larger or more expensive graphs.

It is not evidence of a latency or cost win. Breadth-first search over these tiny local JSON graphs is cheap and makes no model calls. The recorded experiment used 81,549 input tokens and 3,194 output tokens across the 74 unique requests. Model scoring adds work that a local traversal avoids.

The possible benefit would have to come from avoiding more expensive operations: remote graph queries, document fetches, or downstream processing. Whether that trade pays off depends on the application. It needs measurements that include graph access, model requests, caching, and the cost of missed evidence.

The next useful evaluation would use unseen questions and larger graphs, vary the budgets, repeat live runs, and measure both evidence coverage and end-to-end resource use. Context selection deserves its own comparison: recency, relevance, explicit unresolved clues, and the ability to retrieve a discarded source again.

Where this fits

Jev-guided branching already has relevant precedent. TypeSafe’s hierarchical classification cookbook explores greedy and beam-style traversal through a taxonomy. William Lyon has also published work on using Jev for knowledge graph extraction.

This prototype explores a narrower composition: traversal over an existing graph, with accumulated source evidence influencing later routing, several active paths, and explicit controls for losing context or hiding a route. It is a starting point for experiments rather than a claim to have invented graph reasoning.

The MIT-licensed repository includes the implementation and reproduction instructions. The interactive demo replays recorded decisions entirely in the browser. It makes no live Jev calls and contains no API key. Running new live experiments requires your own key.

The architectural idea is worth testing: evidence can shape the retrieval process while it is happening. Once the walk can branch, remember, forget, and stop, each of those behaviors becomes a policy we can inspect and improve.