agentic-rag
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@agentic-ragfind the decision we made about error handling in the API"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
agentic-rag
Provider-neutral long-term memory and compaction continuity — in a real database.
Coding sessions end and long contexts compact. agentic-rag preserves both: a canonical, searchable knowledge base in local PostgreSQL + pgvector, plus bounded checkpoints that let Claude Code and Codex resume after compaction. Hybrid vector + full-text search, lifecycle hooks, and a provider CLI you control — without a hosted RAG service.
New in v0.4.0: Claude compaction continuity — six Claude hooks, the managed 1M/500K
autoCompactWindowpolicy, thecompact_summaryhandoff, andrag install --check/--restore. Read What’s New in 0.4.0.0.3.0: Codex compaction continuity, native-memory policy, recoverable global installation, and provider-neutral mining. Read What’s New in 0.3.0.
Most "RAG memory" tools are a cloud retrieval layer you feed documents to: you push, you query, you pay per call. agentic-rag flips both halves. It stores knowledge in local Postgres + pgvector — real HNSW approximate-nearest-neighbour search blended with bilingual full-text — and it populates itself from supported coding sessions. It also stores compact, audited continuation checkpoints at Claude Code and Codex compaction boundaries — on Claude including Claude's own compact summary as a bounded handoff.
Every content write funnels through one gateway: it strips secret-shaped tokens, chunks and embeds the text with a local model, resolves the document's links into a typed knowledge graph, and logs the change — all in one transaction. When a session ends, a single-writer worker uses the configured Codex or Claude CLI to turn the bounded transcript digest into durable, findable memories. The core data model and provider seam are provider-neutral; integrations adapt each coding agent's lifecycle and output contracts.
It runs on your machine and uses your configured CLI account for LLM-assisted mining, curation, and bounded checkpoint enrichment: Codex with ChatGPT login or Claude with its supported OAuth or API-key authentication. Mining prompts can also include all matching pin bodies; mining secret-strips the provider-bound copies without mutating stored pin text. Embeddings are always local (Ollama), so retrieval does not call either provider. It's RAM-lean by design: no always-on daemon beyond Postgres and Ollama, and an idle footprint near zero between sessions. And it's built data-safety-first — it archives rather than deletes, writes through a least-privilege role matrix, audits every change, and periodically restore-tests its own backups.
Your data stays under your control, with explicit provider calls. This repository is code only — it ships no content. The canonical store lives in your PostgreSQL database, but the configured CLI intentionally sends these provider inputs: mining sends a bounded, secret-stripped transcript digest, live domain names, and secret-stripped copies of all matching pin bodies without mutating stored pin text; curation sends selected stored documents and contradiction evidence; checkpoint enrichment sends a secret-stripped transcript delta and validates the returned checkpoint content before persistence. Optional synced backups copy data only to a directory you configure. agentic-rag has no separate hosted RAG backend.
Why · Quick start · What's different · How it works · Comparison · Configuration · 📖 Handbook · Status · Acknowledgments · License
Why agentic-rag
🔎 Hybrid search that actually ranks. Vector ANN over pgvector (HNSW, cosine) — multilingual by way of bge-m3 embeddings — blended with GIN keyword full-text into one ranked query. Search in any language; not a file scan, not lexical-only.
🌱 It turns sessions into durable knowledge and continuation state. Mining
uses your configured Codex or Claude CLI. On Claude Code and Codex,
PreCompact also captures a fast checkpoint so SessionStart(source="compact")
can restore the goal, blockers, next action, repository state, and evidence
references; on Claude the checkpoint also carries Claude's compact summary.
♻️ It curates itself. A near-duplicate gate stops the store from bloating; rag review surfaces duplicates, dangling links, and stale pins; refuting a fact archives it (with a reason and evidence), never hard-deletes it.
🔒 Local-first, on your own account. Canonical knowledge and checkpoints live in your Postgres. LLM-assisted work runs through the local Codex or Claude CLI you configured. Embeddings are always local (Ollama), so search and retrieval do not call either provider.
⚡ RAM-lean. A single-writer worker (flock singleton), no long-lived daemon of its own. Between sessions the footprint is essentially Postgres + Ollama idling — nothing else.
Related MCP server: rawthink
Quick start
agentic-rag is a rag command-line tool with provider integrations. The
no-option install wires two MCP servers, six lifecycle hooks, and the managed
compaction window into Claude Code; the explicit Codex target installs
continuity configuration and hooks.
Prerequisites:
PostgreSQL 17 with the
pgvectorextension (the schema useshalfvec, pgvector ≥ 0.7).Ollama with the embedding model pulled —
ollama pull bge-m3(1024-dim, fixed to the schema).An authenticated LLM CLI: Codex (
codex login) or Claude (claude -p). Claude/Haiku remains the package default for compatibility; select the provider in[llm].uvand Python ≥ 3.13.
Install the common foundation, then choose an integration:
uv sync
uv run rag init-db # creates the DB + schema + roles, seeds the 'general' domain
uv run rag domain add programming --description "Software engineering notes"
uv run rag install --check # preview the Claude settings merge; writes nothing
uv run rag install # Claude MCP/hooks + macOS backup schedule; omit for Codex-onlyrag init-dbcreates the database if needed, applies the migrations insql/, creates the three least-privilege roles, and seeds the built-ingeneraldomain. Run it first —rag installdoes not create the database.rag domain add <name>adds any domains you want to organize documents under (generalalways exists; add more anytime).rag install --checkpreviews the Claude merge (managed: autoCompactWindow=500000, the would-change path, policy warnings) and writes nothing.The no-option
rag installis the Claude target: it registers theagentic-rag(read-write) andagentic-rag-ro(read-only) MCP servers, merges six hooks (SessionStart,UserPromptSubmit,Stop,PreCompact,PostCompact,SessionEnd) plusautoCompactWindow = 500000into~/.claude/settings.json, backs the file up to a uniquesettings.json.bak.<id>, and prints arag install --restore <record>rollback command. On macOS it also schedules nightly backup; omit this command for a Codex-only setup.
If you ran the Claude target, hooks reload live; start a new Claude Code
session so it picks up the MCP servers, then review the handlers with /hooks
and confirm /autocompact reports 500000 tokens from settings. For humans the
same store is available through the CLI:
rag save --title "Postgres VACUUM tuning" --domain programming \
--dtype lesson --body "autovacuum_vacuum_scale_factor tradeoffs..."
rag search "vacuum tuning" --domain programming
rag get <slug-or-id> # body + incoming/outgoing graph edges
rag status # counts, queue health, last backup/curationFor Codex continuity, preview before touching your user configuration, install, then inspect and trust the changed commands in Codex:
uv run rag install --codex --check
uv run rag install --codex
# Start Codex, run /hooks, inspect all six agentic-rag commands, then trust them.
uv run rag statusThe Codex transaction manages only ~/.codex/config.toml,
~/.codex/hooks.json, and ~/.codex/compact_prompt.md. It prints every
changed path, backup, validation result, and a ready-to-run rollback command:
uv run rag install --codex --restore /absolute/path/to/codex-rollback-<id>.jsonUse the exact rollback-record pathname printed by the successful install.
Check mode writes nothing. Repository support does not mean a particular
machine has completed hook trust or live compaction verification; review
/hooks and confirm rag status after every installation.
The Codex target never installs a scheduler. On a Codex-only macOS setup, use
uv run rag backup --install-launchd for scheduled database backups; Linux
uses the handbook's cron/systemd recipes.
What makes it different
Four things that, together, set it apart from both file-based knowledge wikis and hosted RAG stacks:
1. Real hybrid search over a curated graph
Documents are chunked and embedded into halfvec(1024) columns indexed with HNSW, and each chunk is embedded with the multilingual bge-m3 model — so semantic recall works in any language — and also carries generated tsvectors for English/German keyword full-text. A single query runs vector ANN and full-text together and returns one ranked list. Documents aren't an undifferentiated pile: they carry a type (concept, lesson, signal, synthesis, reference, …) and connect through a typed edge graph (references, extends, depends_on, supersedes, contradicts, …), so rag get shows you not just a document but its neighbourhood.
2. Automatic session-mining — the star feature
This is what makes agentic-rag feel like it grows rather than sits there. When a supported coding session ends, a lifecycle hook enqueues the transcript. A single-writer worker drains the queue and calls the configured Codex or Claude CLI to pull out durable memories, lessons, and signals, each saved through the write gateway behind a near-duplicate gate. A fix you discovered today becomes something a future session can recall, with no "remember to write this down" step. It reads only your local session transcripts.
3. Knowledge domains you grow and curate
Domains are just data — a label for where to look (general is seeded at init). Add them with rag domain add, scope any search with --domain, and let the importer derive them from an existing store's topics. Curation is first-class: rag review reports near-duplicates, dangling links, and stale pins; refuting a fact archives it with a required reason + evidence; rag purge removes only already-refuted documents, and only as rag_admin.
4. Built-in maintenance, backup, and restore-testing
rag backup runs pg_dump -Fc locally (plus an optional copy to a synced directory you configure), with rotation. rag maintenance is a tiny, single-flight, always-exit-0 job that ticks the worker, rotates logs, and — weekly — runs a report-only restore-test: it restores your newest dump into an isolated scratch database, compares row counts, and drops it. A backup you've never restored isn't a backup; this one checks itself.
How it works
Supported coding sessions rag save · migrate import · MCP write tools
(queued & mined on session end) (you, or an agent)
│ │
└───────────────────┬───────────────────────────┘
│
one audited write gateway
strips secret-shaped tokens · chunks + embeds (local Ollama) · resolves edges · logs
│
┌───────────────┴────────────────┐
PostgreSQL + pgvector typed knowledge graph
documents · chunks halfvec(1024) edges: references · extends · …
HNSW ANN + EN/DE full-text
│
one hybrid ranked search ── vector ⊕ full-text
│
session-start context · prompt-time recall · read-only MCP behind a privilege boundaryClaude Code and Codex continuity use a separate operational path. The Claude flow:
PreCompact ──► bounded deterministic snapshot ──► audited checkpoint
│ └──► priority enrichment job ──► provider CLI
└──► stdout: versioned compact instructions (+ checkpoint id)
Claude compacts (instructions appended)
│
PostCompact ──► mark boundary + store compact_summary as bounded handoff
│
SessionStart(source="compact") ──► checkpoint + handoff, ≤ 10,000 chars ──► next requestThe Codex flow (PreCompact stays silent; PostCompact stores no handoff):
PreCompact ──► bounded deterministic snapshot ──► audited checkpoint
│ └──► priority enrichment job ──► provider CLI
▼
Codex compacts
│
PostCompact ──► mark boundary only (cannot inject context)
│
SessionStart(source="compact") ──► bounded checkpoint context ──► next requestNative Codex memories are complementary, not the canonical record. With the
installed policy they remain enabled and can be inspected with /memories;
agentic-rag is canonical for durable searchable knowledge, audit history, and
explicit continuation checkpoints.
Your chosen CLI provider. Every LLM call goes through the single
agentic_rag.llmseam and the configured local Codex or Claude command. Transcript digests and checkpoint deltas are character-bounded; each curation call covers one selected document/evidence set. Mining also sends secret-stripped copies of all matching pin bodies without changing the stored pins. Embeddings never leave the box (local Ollama), so retrieval is independent of provider authentication.One audited write path. Every change — a manual
save, a mined memory, an import — funnels through a single gateway that strips secret-shaped tokens, regenerates chunks + embeddings in one transaction, resolves dangling edges, and writes an audit row. Embeddings fail open (queued for retry if Ollama is down); nothing else does.Least privilege, by role. Three login roles enforce a destruction-protection matrix:
rag_reader(SELECT only, used by search and the read-only MCP),rag_writer(INSERT/UPDATE but no DELETE/TRUNCATE/DROP), andrag_admin(migrate, purge, restore).It steps aside, not in front. If Ollama is down, search degrades to full-text-only and returns a warning rather than failing; the maintenance job always exits 0.
The full story is in the handbook — the mental model, everyday use, configuration, importing an existing wiki, and the architecture and design rationale.
Comparison
agentic-rag sits between two worlds: the file-based LLM-Wiki family (human-readable Markdown with a lint/graph layer) and typical RAG stacks (hosted or API-driven retrieval you feed documents to). Every cell below is marked honestly — including the rows where each of them beats us.
Legend: ✅ shipped in code and operationally established · 🧪 shipped and installed, live verification pending · ⚠️ partial / caveated · ❌ absent
vs LLM-Wiki systems (file-based knowledge wikis)
Capability | agentic-rag | File-based LLM-Wiki |
Hybrid vector + full-text ranked search (ANN at scale) | ✅ | ⚠️ lexical/graph, file-scan |
Bilingual full-text (EN + DE) + semantic recall | ✅ | ⚠️ |
Transactional, audited writes through one gateway | ✅ | ⚠️ |
Scales to a large corpus (HNSW index) | ✅ | ⚠️ file-scan slows |
Human-readable, git-diffable plain-text store | ⚠️ import/export MD; store is Postgres | ✅ |
Zero-infrastructure (no DB/service to run) | ❌ needs Postgres + Ollama | ✅ |
Imports an existing llm-wiki store | ✅ | ✅ it is one |
Bottom line: if you want a git-tracked pile of Markdown, a file-wiki wins on its home turf. If you want fast hybrid recall over a growing corpus with transactional safety, agentic-rag wins — and it can import your existing llm-wiki to get you there.
vs typical RAG systems (retrieval frameworks / hosted memory)
Capability | agentic-rag | Typical RAG stack |
Local-first store, provider CLI under your control, no hosted RAG service ¹ | ✅ | ⚠️ usually a hosted service |
Auto-populates from Claude Code sessions (mining) | ✅ | ❌ you feed it |
Codex session mining and continuity | 🧪 | ❌ you feed it |
Claude Code compaction continuity (checkpoint + handoff) | 🧪 | ❌ |
Self-curation (dedup, near-dup gate, refute/archive) | ✅ | ⚠️ |
Typed knowledge graph (edges) alongside vector search | ✅ | ⚠️ |
One audited write gateway with secret stripping | ✅ | ❌ |
Read/write privilege boundary for subagents (RO MCP) | ✅ | ⚠️ |
Turnkey managed hosting / large ecosystem ² | ⚠️ self-host, young | ✅ |
Bottom line: a hosted RAG stack wins on turnkey scale and ecosystem. agentic-rag wins on being local-first, using a provider CLI under your control with no hosted RAG service in the loop, self-populating-from-your-own-work, and self-curating — a memory that fills and tidies itself instead of one you have to keep feeding.
Configuration
Config lives in one TOML file at ~/.agentic-rag/config.toml. Every key is optional — omit a section to keep its defaults.
Setting | Default | What it does |
|
| Database name. |
|
| Empty = local unix socket; set it for a networked/remote server. |
|
| Ollama embedding model tag. |
|
| Fixed to the schema ( |
|
| Local Ollama endpoint. |
|
| Where |
| — (unset) | Opt-in copy to a synced/cloud directory; unset means backups remain local. Provider-bound LLM inputs are disclosed above. |
| auto-resolved | Only needed if |
Roles are created passwordless by default, relying on local peer/trust auth (Postgres and agentic-rag on the same machine). For a networked or shared instance, set role passwords with ALTER ROLE … and let libpq authenticate via ~/.pgpass or PGHOST/PGPORT/PGPASSWORD — see the handbook's privacy chapter.
The Claude target separately manages autoCompactWindow = 500000 in
~/.claude/settings.json — a 1M context compacting at 500K with a [1m]
model; model is reported, never rewritten, and long-context requests above
200K input tokens cost more on API billing. The Codex target separately
manages a 600000 context window and a 500000 total-token compaction threshold, leaving a 100K reserve, plus native memories
and the compact prompt. Official GPT-5.6 capacity is 1.05M, but inputs above
272K are subject to higher provider pricing and may add latency; see
Configuration and
Privacy, cost & control.
📖 Documentation / Handbook
The full story lives in the agentic-rag Handbook — a single, progressively-ordered read from the mental model through everyday use, configuration, importing, and the engine's architecture and design rationale. A few key chapters:
Understand — What is agentic-rag? · The mental model
Use — Quick start · Working with your memory · Session mining & curation
Configure — Privacy, cost & control
Develop — Architecture · Reference — CLI & MCP
Start at the handbook index for the one-line "what you'll learn" map of every chapter.
Status
agentic-rag is young but solid — a real engine, openly developed. Repository and rollout state are listed separately below:
✅ Storage & search: PostgreSQL + pgvector schema, HNSW ANN blended with EN/DE full-text into one ranked list, the typed edge graph, the three-role destruction-protection matrix.
✅ The document write gateway: secret stripping on document inputs and generated document writes, one-transaction chunk + embed + edge-resolve + audit, embeddings that fail open with a retry queue.
✅ Session mining: hooks → queue → single-writer worker → configured Codex/Claude CLI → gateway, with a near-duplicate gate and provider-outage circuit breaker.
✅ Curation & safety:
rag review, refute-as-archive, and admin-onlyrag purge(removes only already-refuted documents, asrag_admin).✅ Maintenance & backups:
pg_dumpbackups with rotation, the tiny always-exit-0 maintenance job, and the weekly report-only restore-test. macOS auto-schedules vialaunchd; Linux uses the documented cron/systemd recipes.✅ Claude integration: two user-scope MCP servers (read-write + a read-only server behind a privilege boundary for subagents), idempotent install that preserves foreign hooks.
✅ Claude continuity in code: six Claude hooks, the
PreCompactstdout compact prompt, the boundedcompact_summaryhandoff, the 10,000-character SessionStart cap, the managed 1M/500K policy, andrag install --check/--restore.🔒 Claude continuity live rollout:
rag install,/hooksreview,/autocompact, manual/automatic compaction, and SessionEnd tail capture on the maintainer machine remain open (backlog 0.3).✅ Codex continuity in code: audited checkpoints, bounded capture and restoration, asynchronous enrichment, all six lifecycle handlers, a versioned compact prompt, recoverable installer/check mode, and checkpoint health in
rag status.🔒 Codex continuity live rollout: the pre-install whole-diff/security review and provider-bound pin hardening are complete. The live global install,
/hookstrust, manual/automatic compaction, provider-recovery, and SessionEnd smoke tests remain open. SeeFEATURES.mdand blocker-firstBACKLOG.md.✅ Quality: a content-free repository with a comprehensive local test suite; exact verification counts belong in rollout evidence, not a static badge.
The clearest gap relative to the field is maturity: it's newly public and self-hosted, without the turnkey hosting or large ecosystem of established RAG stacks.
Acknowledgments
agentic-rag builds on other people's ideas and tools:
Andrej Karpathy — the LLM-Wiki idea that shaped the durable-knowledge model this import path speaks to.
The llm-wiki format — topic-partitioned Markdown with an optional memory store;
rag migrateimports it wholesale, so an existing wiki carries straight over.pgvector and Ollama (
bge-m3) — the local vector search and embeddings underneath everything.Anthropic — Claude Code, its hooks, and the MCP integration agentic-rag plugs into.
Contributing
Tests come first (TDD), and docs/ is kept in step with the code. A warn-only doc-reminder hook ships under .githooks/: if a commit touches agentic_rag/ or sql/ without touching docs/, it prints a reminder — it never blocks. Enable it once per clone:
git config core.hooksPath .githooksRun the suite with uv run pytest. See the handbook's Contributing chapter for dev setup, the test database, and code layout.
License
MIT. The repository is code-only and content-free — your documents, embeddings, and config stay in your own PostgreSQL database, on your own machine.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
No tool schema history has been recorded yet.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Persistent memory for Claude Code and Cursor. Stop re-explaining your project every session.
Shared memory for coding agents. Stop re-explaining your codebase every session.
Persistent cross-session memory shared by Codex, Claude Code, ChatGPT, and other AI agents.
Private persistent memory for Claude, ChatGPT & Gemini via MCP - semantic search, zero-code setup.
Related MCP Servers
- AlicenseAqualityBmaintenancePersistent memory for Claude Code. Automatically indexes every conversation and provides production-grade hybrid search (BM25 + vectors + reranker) via MCP tools. 100% local, zero config, zero API keys, zero invoice.16577MIT
- AlicenseAqualityAmaintenancePersistent memory for Claude Code — hybrid search, knowledge graph, session lifecycle.17MIT
- AlicenseNot gradedqualityCmaintenanceA local, persistent, semantically-aware knowledge graph for AI coding agents like Claude Code, providing efficient session memory with minimal token cost and zero runtime network calls.MIT
- AlicenseNot gradedqualityDmaintenanceProvides persistent, searchable memory for Claude Code using local SQLite, semantic embeddings, and full-text search, enabling Claude to recall and retrieve context across sessions and projects without external services.194MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/phense/agentic-rag'
If you have feedback or need assistance with the MCP directory API, please join our Discord server