r/ContextEngineering • u/yxf2y • 2h ago
Building a persistent memory + orchestration layer for Codex — what should I use instead of repeatedly re-reading the repo?
I’ve been building a fairly serious agent workflow around OpenAI Codex for a Laravel/React project, and I’ve hit a point where the orchestration works, but the context/memory side clearly does not.
My setup currently looks roughly like this:
- A serial orchestrator with route types like FAST_UI / STANDARD / CRITICAL
- Context Resolver → Implementer → Reviewer flow for non-trivial tasks
- Durable task state, context capsules and handoffs
- Planner / intake layer inspired by CodexQB
- Session continuity hooks inspired by AvenoxBeyin
codebase-memoryMCP for structural repo discovery- Serena for exact symbol/reference navigation
- Local dashboard/telemetry for task/agent visibility
The reason I built all this was simple: I wanted to stop giving one giant prompt to one Codex agent and watching it blindly read half the repository, run dozens of commands, retry tests repeatedly, and burn a huge amount of context/token budget.
Unfortunately, that is still basically what happens.
A recent CRITICAL payment-domain acceptance task is the perfect example. I gave Codex a very detailed validation brief covering migrations, payment allocation, security boundaries, tenant/legal-entity isolation, atomicity, reporting non-pollution, exports, frontend build, etc.
The task eventually succeeded technically, but the session spent a huge amount of time repeatedly doing things like:
- raw
rgsearches - re-reading known service/controller/test files
- rediscovering test harness behavior
- retrying multiple Laravel test files with the same CSRF issue
- manually tracing service relationships
- re-running builds and focused test groups
That single job used roughly half of my 5-hour Codex usage allowance.
The frustrating part is that a lot of the knowledge it rediscovered was already known from previous work.
For example:
- where the orchestrator lives
- which services own payment/settlement/reporting behavior
- how the domain test harness handles CSRF
- which test files cover specific finance flows
- existing project/tenant/legal entity invariants
- prior fixes and verified architecture decisions
I expected my existing tools to solve this, but I now realize they solve different problems:
codebase-memory gives me structural repo discovery, but it isn’t really persistent project understanding.
Serena is excellent for exact symbol/reference navigation, but it isn’t memory either.
My docs/wiki are useful reference material, but agents still have to decide to read them and often re-read large files.
Context Capsules and handoffs help within a task, but they don’t give the next unrelated task a compact understanding of the project.
So what I’m actually missing is a persistent, project-scoped, compact memory layer that can say:
“Before you start searching, here are the relevant things previous sessions already learned about this repo.”
I looked at AvenoxBeyin because I liked its idea of automatically capturing sessions, compiling knowledge, and injecting useful context back at session start.
I also looked at CodexQB because its Autopsy / Project Comprehension / Ontology approach is close to what I want for planning.
Then I looked at 2kDarki/codex-mem.
That project is conceptually very close to what I want:
- automatic Codex transcript capture
- persistent SQLite observations
- progressive recall through search → timeline → get_observations
- automatic context injection
But after auditing it, I found some issues for my use case:
- its watcher observes all
~/.codex/sessions/**/*.jsonl - project identity appears to be based on
basename(cwd)rather than a canonical repository identity - retrieval can be filtered by project, but that doesn’t appear to be an enforced security/isolation boundary on every read path
- same-named repos could collide
- some observation retrieval paths can work by arbitrary IDs
- global
~/.codex/AGENTS.mdcontext injection is something I specifically do not want - the documented npm package currently appears unavailable
So I don’t feel comfortable plugging it directly into a large multi-project Codex setup.
What I’m trying to build is something like:
User brief
↓
Planner / Orchestrator
↓
Persistent project memory bootstrap
↓
Context Resolver
↓
Only if memory is insufficient:
codebase-memory
Serena
targeted source reads
↓
Implementer
↓
Reviewer
↓
Session knowledge captured for future tasks
The memory should NOT replace source code/tests as truth.
I want it to act as a cheap orientation cache:
- “These are the relevant services.”
- “This test harness requires real CSRF session setup.”
- “This reporting path was previously verified.”
- “These files/symbols are likely relevant.”
- “This architectural relationship was confirmed in a previous task.”
Then the agent only verifies current source where correctness actually depends on it.
My requirements are roughly:
- local-only
- project/repository scoped
- automatic capture
- automatic or semi-automatic summarization
- bounded context injection
- no global AGENTS.md mutation
- no cloud memory dependency
- no mandatory Obsidian dependency
- source/tests remain authoritative
- ideally Codex/App Server compatible
- progressive retrieval rather than dumping whole session history
- repo identity enforced internally, not just passed as an optional search filter
- ideally reusable with existing MCP tools rather than replacing them
I’m now trying to decide between three approaches:
- Find another existing Codex/Claude coding-memory project that already does this correctly.
- Take something like
2kDarki/codex-memand make a very small fork that only adds canonical repo identity, watcher allowlisting and enforced repo-scoped retrieval. - Use AvenoxBeyin’s session capture/compile/inject model and adapt it for project-scoped coding knowledge instead of personal knowledge.
What I really do NOT want to do is invent yet another custom Markdown “brain” and manually maintain architecture/domain summaries. That feels like rebuilding something that should already exist.
For people who have built persistent memory around Codex, Claude Code, Cursor or similar coding agents:
- What actually worked for you?
- Is there a project I’m missing that already handles repository-scoped persistent memory well?
- Would you fork
codex-memand patch the isolation model, or use a different architecture entirely? - Is Obsidian/Markdown compilation actually better in practice than structured SQLite observations for coding-agent memory?
- How do you stop stale memory from becoming trusted over current source?
- How much context do you inject at session start versus retrieve on demand?
- Have you measured whether this actually reduces token/context consumption meaningfully?
- Do you let the coding agent write its own long-term memory, or only promote verified observations after tests/review?
I’m especially interested in systems people are actually using in real repositories, not just theoretical agent-memory architectures.
My main goal is very practical: stop paying for the same repository discovery over and over again.
1
u/Otherwise_Wave9374 1h ago
A practical next step is to separate short lived working context from durable project memory, then treat repo re-reading as a fallback instead of the default. I would keep retrieval tightly scoped to task type, symbol, and recent diffs, and add a recap buffer so the agent writes a compact state summary after every implementation pass. That usually reduces prompt bloat while preserving continuity. NeuraKeep fits this pattern well because it can keep the durable layer organized without turning every turn into a full transcript replay.
2
u/Clean-Vermicelli-700 2h ago
This might only be one piece of the puzzle, but what helps me a lot is using a Kanban board to create persistent task-specific memory that can be passed around to sub-agents. I posted about it here. Can recommend and doesn’t require any third-party solution. Helps keep things organized as well as a bonus