Expérimental95

Hipocampo

I measured what actually reaches Claude Code's context window over 61 days of my own transcripts

r/ClaudeCodeu/JhouHate30 septembre 2026

Capture du projet

Résumé

Le projet Hipocampo mesure ce qui atteint la fenêtre de contexte de Claude Code sur 61 jours de transcripts personnels. Il utilise des fonctionnalités natives telles que les événements hook et les paramètres pour décider ce qui est inclus dans la fenêtre de contexte de chaque session. Les résultats montrent que la compaction est agressive et que seuls 11% des données injectées atteignent le modèle.

Pourquoi c’est intéressant

Ce projet est intéressant car il fournit des informations précieuses sur la façon dont Claude Code traite les données de contexte et comment la compaction affecte les performances du modèle. Les résultats pourraient être utiles pour améliorer la conception de systèmes utilisant Claude Code.

Comment Claude est utilisé

Le projet utilise Claude Code pour mesurer ce qui atteint la fenêtre de contexte de Claude Code sur 61 jours de transcripts personnels. Il utilise également des fonctionnalités natives telles que les événements hook, les paramètres, l'index http://MEMORY.md et les compétences.

Idées dérivées

  1. 01

    Mesure de la compaction dans les conversations multi-agents

    Un projet similaire pourrait être réalisé pour mesurer l'efficacité de la compaction dans d'autres contextes, tels que les conversations avec plusieurs agents.

  2. 02

    Outil de visualisation des données de contexte

    Un autre projet pourrait consister à développer un outil pour visualiser et analyser les données de contexte de Claude Code, afin de mieux comprendre comment les informations sont traitées et utilisées.

  3. 03

    Comparaison des performances des modèles de langage

    Un projet pourrait également être réalisé pour comparer les performances de différents modèles de langage, tels que Claude Code et d'autres modèles, dans des tâches spécifiques.

Afficher le post original
I run several projects alone with Claude Code, often with two or more sessions open at once. Each session starts with nothing from the previous one, and compaction replaces the history with a summary the product writes. Over the last two months I built a set of pieces around that, which I call Hipocampo, to decide what goes into the window of each session, at what moment, and how to check that it got in. What it is Everything uses native features: hook events, settings, the http://MEMORY.md index, http://CLAUDE.md, skills, subagents and compaction. Nothing installed, no MCP. The commit guards run in the git pre-commit hook. The organizing idea is that each thing lives in one of three regimes: • Resident: always arrives, before the first decision (a map loaded by hook, the memory index, a state file). • Paged: only arrives if something opens it (CLAUDE.md in subfolders, docs, the memory files themselves). • Interrupt: arrives when the command or the file touched matches a registered source, on that event and only on it. The three regimes: resident, paged, interrupt. (https://preview.redd.it/wcypckoi6nsh1.png?width=2000&format=png&auto=webp&s=d211693422a85656dc79c95e86e97988f10e2f33) What I measured 61 days of transcripts from this installation (190 main sessions, 918 subagent files), with a positive and a negative control and the population declared for every number. 0 injected recalls in 190 conversations. A memory only got in when the agent opened its file. 209 of 432 memories were opened at least once (a lower bound). Compaction is aggressive. In one event in August, 96.6% of the context was removed (607,378 → 20,766 tokens), and 8 of my 36 messages from before the summary left no trace in it. The starting numbers. (https://preview.redd.it/lf98uvhk6nsh1.png?width=2000&format=png&auto=webp&s=4d0013053e3a537501bd27e729dea627b5233c19) Three green instruments, 11% delivered. On Aug 13, the hook that loads my project map emitted 18,057 bytes and about 11% reached the model. The runtime said hook success, the script log said complete, the script's selftest passed 9 of 9. Hook output above 10,000 units per command goes to a file, and only a preview of about 2 KB reaches the model. The only check that caught it was comparing, byte by byte, what the script emitted with what the transcript recorded. The map now goes in 5 slices, each under the ceiling. The hook reported success three ways; 11% of the output reached the model. (https://preview.redd.it/976elh3m6nsh1.png?width=2000&format=png&auto=webp&s=07f14393cfa5abcbd2f9b7eb4b9777ab4d717465) The compaction ceiling setting changed the shape of sessions. Until Sep 18, conversations went up to almost 1M tokens before the summary. Since Sep 21, 64 of 65 compactions happened at or below 365K. Tokens before each of the 139 compactions. (https://preview.redd.it/7m6rktyn6nsh1.png?width=2000&format=png&auto=webp&s=11df3a0fad887d6b3d5e99c5ea500a4d51563f84) A guard only counts after it has rejected a defect planted on purpose. On Sep 25, 52 of 74 guards had that proof. The rest are declared as debt. Guards in CI vs guards proven failing, Aug 18 to Sep 25. (https://preview.redd.it/kkp2nxpv6nsh1.png?width=2000&format=png&auto=webp&s=6f635635ce663e97c955b9f9c9c270b6951712a5) How to build it The article ends with the order I would follow, in 7 steps, each with the native feature it relies on and the red that proves it works. Every step ends by breaking the piece on purpose and waiting for the failure. The order, in seven steps. (https://preview.redd.it/ay6xrqvx6nsh1.png?width=2000&format=png&auto=webp&s=b9a3baa3020a720c21254209e32876fcd615d352) Limits One operator, one installation. The study does not say whether the memory that arrives is correct or whether it was useful. Links • Full article (architecture piece by piece, the measurement bench, 33 figures, limitations): https://ksmit.com.br/en/blog/hipocampo • Code (hooks as a settings example, the memory guards with their failing fixtures, the measuring scripts, memory and map templates; MIT): https://github.com/JhouCode/hipocampo Written and tested on Linux with bash and python3. If you run the measuring scripts on your own transcripts, I'd like to know what number you get for injected recalls.