Personal project

claude-memory-graph

The problem. Claude has no memory between sessions. Months of my own engineering conversations sit in local transcript files, and nothing can query them.

What you are looking at. Every dot is an entity pulled out of my own Claude Code history — a tool, a file, a concept, a problem, a decision. Every line is a relationship between two of them. Roughly four months of work, as a graph.

  1. Ingest — parse local transcripts into a normalised store. One chunk = one exchange. 25 sessions → 411 chunks.
  2. Knowledge graph — an LLM extracts (entity, relation, entity) triples against a closed vocabulary of 7 entity kinds and 11 predicates; anything outside it is dropped rather than stored. Further edges come straight from tool calls, so they cannot be hallucinated. 2,579 entities, 2,895 relations.
  3. Vector index — embed every chunk; cosine search in pure Python, no vector database.
  4. MCP server — three tools, so Claude can query all of it mid-conversation.

The result was not the one I set out to prove. The plan was to show that hybrid retrieval — fusing graph and vector search — beats plain vector search. I built two evaluation sets with deliberately opposite biases so neither could flatter the architecture, and measured it.

Hybrid lost to both of its own inputs.

modesingle-hopcross-sessionoverall
vector only0.5760.0740.353
graph only0.1910.4200.293
hybrid (fixed fusion)0.3490.2340.298
adaptive routing0.5380.313 0.438

Mean reciprocal rank over 45 questions. Higher is better.

The two retrievers turned out to be complementary rather than additive: vector search is ~8× better at “what did this conversation say”, the graph ~5× better at “where else did this come up”. Blending them at a fixed ratio is worse than picking one.

So I replaced fusion with routing: let vector search's own confidence decide how much graph to mix in. When nothing in the corpus closely matches the query, the answer is more likely reachable than stateable — so lean on the graph. That is +24% overall against the best fixed strategy, and it matches the graph's cross-session recall while keeping the vector's.

Reported honestly: the routing thresholds are tuned on the same 45 questions they are scored on, and 45 questions from one person's history is a small sample. The README documents this and the other caveats rather than burying them.

Built with the Python standard library — SQLite for both the graph and the vectors. No ORM, no vector database, no graph database. The MCP SDK is the only third-party dependency.

Source and full write-up on GitHub →

Entity names here are limited to well-known public tools; everything drawn from private conversations is masked.