How it works, and what it refuses to do.
Mnemon stores what you and your agents say, verbatim, and hands the relevant parts back to any tool that speaks MCP. This page explains the mechanism rather than the pitch: how results are actually scored, why there is no schema for relationships, what the benchmark does and does not prove, and what happens to private content you would rather not have leave the building.
Mnemon is a Laravel application you run yourself. It keeps two things: every piece of content you or your agents decide is worth keeping, stored word-for-word and never overwritten; and a set of compiled pages that synthesise that raw material into something readable. Agents reach both through 14 tools over the Model Context Protocol.
The point is that the memory is yours and shared. One instance serves Claude Code, Claude Desktop, Cursor, ChatGPT and anything that can make an authenticated HTTP request, so context captured in one tool is available in the next. There is no hosted version, no account, and nothing phones home.
If you are evaluating this, the useful question is not whether it stores text — everything does. It is whether retrieval hands back the right text, and what it gives up to do that. The rest of this page is about those two things.
which one you are reading.
The palace is append-only storage, organised as wings → rooms → drawers. A drawer holds content exactly as it was written. Nothing is summarised at ingest and nothing is rewritten later, which means the palace can always answer “what was actually said”.
The wiki is the opposite: pages compiled from drawers, typed as
person:, project:, concept:,
decision: or synthesis:, and rewritten whenever enough new
material accumulates. A wiki page is an interpretation, and it tracks how stale it is and
which drawers it came from.
Keeping them separate is the whole design. A system that only summarises loses the evidence; one that only stores raw text makes the reader do all the work every time. The split means a compiled claim can always be traced back to the verbatim record that produced it — and when the two disagree, the drawer wins.
This instance currently holds 4,420 drawers across 33 wings, compiled into 33 wiki pages.
A query fans out to three independent legs. Each pulls three times the requested number of candidates, and the results are merged under weights you can change in configuration.
Semantic (0.6) — cosine distance between the query's embedding and each drawer's, in pgvector. This is what finds a drawer about deployment when you asked about shipping. Full-text (0.3) — PostgreSQL's own index, applied one term at a time with stop-words removed, so an exact identifier or error string still wins. Recency (0.1) — a thumb on the scale for anything from the last seven days, never a filter.
“How are you searching data with complex relationships without defining a full-blown ontology?”
Short answer: it doesn't. There are no joins, no multi-hop traversal in the retrieval path, and no schema constraining what a relationship may be. That is a position, not an oversight.
The bet is that relationships live in compiled prose, and the reader is a language
model. You don't need (alice)-[:MANAGES]->(bob) in a graph store
if a retrievable page says “Alice manages Bob” and retrieval reliably surfaces that page.
The relational reasoning happens at read time, in the model, over text — rather than at
write time, in a schema.
An ontology-first design asks you to know your entity types before you have data. For a system ingesting arbitrary working transcripts, you never do. Every relationship type you failed to anticipate becomes a migration, and every fact that doesn't fit the schema gets dropped at the door.
What the graph actually contains
There is a typed relationship store, and being straight about it matters more than defending it. The schema defines 5 edge types and the traversal code will walk five hops. Here is what is in it right now:
| Quantity | Count |
|---|---|
| Drawers (verbatim) | 4,420 |
| Compiled wiki pages | 33 |
| Typed edges | 86 |
Generic references edges | 86 |
| Edge types in use | 1 of 5 |
Nearly every edge is the most generic term in the vocabulary. uses,
depends-on, caused and contradicts are defined and
effectively unwritten. So what ships is a shallow, largely untyped link graph beside a
retrieval engine that doesn't consult it — and the typed vocabulary is, honestly,
aspirational.
What the choice buys
- It works on arbitrary transcripts from the first ingest — no modelling phase.
- A new kind of relationship needs no migration. Prose absorbs it.
- Nothing is silently dropped for not fitting a schema.
- Only the compile step must be clever, and it is a prompt rather than a data model.
What it costs
- No multi-hop queries. “Which projects depend on a library Alice owns” is unanswerable.
- No contradiction detection, though the edge type exists to express it.
- Fidelity equals whatever the compiler noticed. Nothing forces the question, so an unobserved relationship isn't there.
- Edges are navigational, not semantic.
The route is not “define an ontology” — it is to make the compile step emit the typed edges it already has a schema for. The table, traversal, depth limits, vocabulary and MCP tool all exist and sit unused. That makes it an extraction-and-prompting problem rather than a data-modelling one, and it keeps the original bet intact: edges become an index over prose that still carries the meaning, rather than a cage the prose is flattened into.
Most memory systems assert that semantic retrieval helps. This one measures it. The full 500-question LongMemEval-S set was run twice over identical content — once with embeddings off, once on — so the only variable is the embedding.
Embeddings improve every metric at every depth. Total spend was $36.76 against a $37 estimate extrapolated from a two-question smoke test — within one percent at 250× the sample size.
An earlier 25-question subset had suggested embeddings only re-ranked results rather than finding more evidence. The full run overturned that. The subset was not merely imprecise, it was misleading, and the correction is documented rather than quietly replaced.
The benchmark measures whether the right evidence is retrieved from the palace layer. It is not a test of relational reasoning, of the wiki layer, or of long-horizon agent behaviour, and it should not be cited as one. The answer-accuracy figure is also judged by the same model family that produced the answers. Full methodology and three further caveats ship in the repository.
would rather not share.
The instance is yours, so the baseline is simple: content goes to your database on your host, and nothing is transmitted anywhere unless you configure an embedding provider. With the default settings nothing leaves the machine at all.
Secrets are stripped before storage
Captured content passes a sanitiser on the way in. It redacts OpenAI-style keys,
GitHub tokens, bearer tokens, database passwords embedded in connection URLs, and
PASSWORD= / SECRET= / API_KEY= assignments. This
runs at write time, so an accidentally pasted credential is redacted in the stored copy
rather than sitting in your memory forever.
Agents are restricted per device
Each device gets its own OAuth credential, and each credential is scoped to the wings it
may read. An agent restricted to work cannot read a personal
drawer — the check runs before the tool does, and a denial is written to the audit log.
Every tool call is recorded with which credential made it and when.
One caveat worth knowing before you compile
Wiki pages take their wing from the page name, not from the drawers they were compiled from. A page named after the wing it synthesises is scoped correctly. A page named generically but compiled from restricted material would be scoped by its name instead of its content. Name pages after the wing they belong to and the isolation holds.
This list is maintained in the repository and kept current. If something here is a dealbreaker, better to find out now than after an afternoon of setup.
- 01Wiki wing scope follows the page name, not the sources it was compiled from — as described above.
- 02Single-tenant. Any registered user of the admin panel is an administrator. Agent isolation is per-credential; human isolation does not exist.
- 03Semantic search requires PostgreSQL with pgvector. SQLite works and falls back to full-text plus recency, with no semantic ranking.
- 04No server-push streaming. The MCP transport is request/response; a GET on the endpoint returns 405, which the specification permits.
- 05No reverse-proxy TLS support.
X-Forwarded-*headers are not processed, so run the bundled HTTPS path rather than terminating TLS upstream. - 06Drawers cannot be hard-deleted through the API. Removal is an admin-panel action; the agent-facing layer is read-and-append.
- 07Access tokens last one hour, refresh tokens ninety days. Enrolled devices renew themselves; a static token does not.
- 08Word counts are ASCII-only, so multi-byte content under-counts. Cosmetic, and documented.
The shipped default is MNEMON_EMBEDDING_DRIVER=none: no account,
no API key, no spend. Retrieval falls back to full-text plus recency, which works —
it is simply the weaker leg the benchmark above measures against.
Turning on semantic search means choosing a provider. OpenAI
(text-embedding-3-small) is the path the benchmark used. Ollama
(nomic-embed-text) runs locally against a model on your own hardware, so it
costs nothing and sends nothing out; it needs an Ollama server reachable from the app.
Embedding cost is incurred twice: once per drawer when it is stored or re-embedded, and once per query. Backfilling an existing corpus is the expensive moment — a measured run embedded roughly two drawers per second, so a few thousand drawers is a job to start and walk away from, not something to wait on.
Beyond that it is a PHP application and a Postgres database. It runs comfortably on the smallest instance any host offers.
The stack binds to 127.0.0.1 by default, so a trial is never
exposed to the network.
# clone, configure, start
git clone https://github.com/coopers98/mnemon.git
cd mnemon
cp .env.docker.example .env
docker compose up -dThen open http://localhost:8080. The admin password is generated on first
boot and written inside the container — there is no reset flow, so save it. Connecting an
agent, serving a real hostname with automatic certificates, and the native install are all
covered in the user guide.
The source, the benchmark harness, and the limitations list above all live at github.com/coopers98/mnemon under the MIT licence.