Skip to main content
Glama
phense
by phense

agentic-rag

Provider-neutral long-term memory and compaction continuity — in a real database.

Coding sessions end and long contexts compact. agentic-rag preserves both: a canonical, searchable knowledge base in local PostgreSQL + pgvector, plus bounded checkpoints that let Claude Code and Codex resume after compaction. Hybrid vector + full-text search, lifecycle hooks, and a provider CLI you control — without a hosted RAG service.

License: MIT version Python 3.13 PostgreSQL + pgvector

New in v0.4.0: Claude compaction continuity — six Claude hooks, the managed 1M/500K autoCompactWindow policy, the compact_summary handoff, and rag install --check/--restore. Read What’s New in 0.4.0.

0.3.0: Codex compaction continuity, native-memory policy, recoverable global installation, and provider-neutral mining. Read What’s New in 0.3.0.

Most "RAG memory" tools are a cloud retrieval layer you feed documents to: you push, you query, you pay per call. agentic-rag flips both halves. It stores knowledge in local Postgres + pgvector — real HNSW approximate-nearest-neighbour search blended with bilingual full-text — and it populates itself from supported coding sessions. It also stores compact, audited continuation checkpoints at Claude Code and Codex compaction boundaries — on Claude including Claude's own compact summary as a bounded handoff.

Every content write funnels through one gateway: it strips secret-shaped tokens, chunks and embeds the text with a local model, resolves the document's links into a typed knowledge graph, and logs the change — all in one transaction. When a session ends, a single-writer worker uses the configured Codex or Claude CLI to turn the bounded transcript digest into durable, findable memories. The core data model and provider seam are provider-neutral; integrations adapt each coding agent's lifecycle and output contracts.

It runs on your machine and uses your configured CLI account for LLM-assisted mining, curation, and bounded checkpoint enrichment: Codex with ChatGPT login or Claude with its supported OAuth or API-key authentication. Mining prompts can also include all matching pin bodies; mining secret-strips the provider-bound copies without mutating stored pin text. Embeddings are always local (Ollama), so retrieval does not call either provider. It's RAM-lean by design: no always-on daemon beyond Postgres and Ollama, and an idle footprint near zero between sessions. And it's built data-safety-first — it archives rather than deletes, writes through a least-privilege role matrix, audits every change, and periodically restore-tests its own backups.

Your data stays under your control, with explicit provider calls. This repository is code only — it ships no content. The canonical store lives in your PostgreSQL database, but the configured CLI intentionally sends these provider inputs: mining sends a bounded, secret-stripped transcript digest, live domain names, and secret-stripped copies of all matching pin bodies without mutating stored pin text; curation sends selected stored documents and contradiction evidence; checkpoint enrichment sends a secret-stripped transcript delta and validates the returned checkpoint content before persistence. Optional synced backups copy data only to a directory you configure. agentic-rag has no separate hosted RAG backend.

Why · Quick start · What's different · How it works · Comparison · Configuration · 📖 Handbook · Status · Acknowledgments · License


Why agentic-rag

🔎 Hybrid search that actually ranks. Vector ANN over pgvector (HNSW, cosine) — multilingual by way of bge-m3 embeddings — blended with GIN keyword full-text into one ranked query. Search in any language; not a file scan, not lexical-only.

🌱 It turns sessions into durable knowledge and continuation state. Mining uses your configured Codex or Claude CLI. On Claude Code and Codex, PreCompact also captures a fast checkpoint so SessionStart(source="compact") can restore the goal, blockers, next action, repository state, and evidence references; on Claude the checkpoint also carries Claude's compact summary.

♻️ It curates itself. A near-duplicate gate stops the store from bloating; rag review surfaces duplicates, dangling links, and stale pins; refuting a fact archives it (with a reason and evidence), never hard-deletes it.

🔒 Local-first, on your own account. Canonical knowledge and checkpoints live in your Postgres. LLM-assisted work runs through the local Codex or Claude CLI you configured. Embeddings are always local (Ollama), so search and retrieval do not call either provider.

RAM-lean. A single-writer worker (flock singleton), no long-lived daemon of its own. Between sessions the footprint is essentially Postgres + Ollama idling — nothing else.


Related MCP server: rawthink

Quick start

agentic-rag is a rag command-line tool with provider integrations. The no-option install wires two MCP servers, six lifecycle hooks, and the managed compaction window into Claude Code; the explicit Codex target installs continuity configuration and hooks.

Prerequisites:

  • PostgreSQL 17 with the pgvector extension (the schema uses halfvec, pgvector ≥ 0.7).

  • Ollama with the embedding model pulled — ollama pull bge-m3 (1024-dim, fixed to the schema).

  • An authenticated LLM CLI: Codex (codex login) or Claude (claude -p). Claude/Haiku remains the package default for compatibility; select the provider in [llm].

  • uv and Python ≥ 3.13.

Install the common foundation, then choose an integration:

uv sync
uv run rag init-db          # creates the DB + schema + roles, seeds the 'general' domain
uv run rag domain add programming --description "Software engineering notes"
uv run rag install --check  # preview the Claude settings merge; writes nothing
uv run rag install          # Claude MCP/hooks + macOS backup schedule; omit for Codex-only
  • rag init-db creates the database if needed, applies the migrations in sql/, creates the three least-privilege roles, and seeds the built-in general domain. Run it firstrag install does not create the database.

  • rag domain add <name> adds any domains you want to organize documents under (general always exists; add more anytime).

  • rag install --check previews the Claude merge (managed: autoCompactWindow=500000, the would-change path, policy warnings) and writes nothing.

  • The no-option rag install is the Claude target: it registers the agentic-rag (read-write) and agentic-rag-ro (read-only) MCP servers, merges six hooks (SessionStart, UserPromptSubmit, Stop, PreCompact, PostCompact, SessionEnd) plus autoCompactWindow = 500000 into ~/.claude/settings.json, backs the file up to a unique settings.json.bak.<id>, and prints a rag install --restore <record> rollback command. On macOS it also schedules nightly backup; omit this command for a Codex-only setup.

If you ran the Claude target, hooks reload live; start a new Claude Code session so it picks up the MCP servers, then review the handlers with /hooks and confirm /autocompact reports 500000 tokens from settings. For humans the same store is available through the CLI:

rag save --title "Postgres VACUUM tuning" --domain programming \
    --dtype lesson --body "autovacuum_vacuum_scale_factor tradeoffs..."
rag search "vacuum tuning" --domain programming
rag get <slug-or-id>          # body + incoming/outgoing graph edges
rag status                    # counts, queue health, last backup/curation

For Codex continuity, preview before touching your user configuration, install, then inspect and trust the changed commands in Codex:

uv run rag install --codex --check
uv run rag install --codex
# Start Codex, run /hooks, inspect all six agentic-rag commands, then trust them.
uv run rag status

The Codex transaction manages only ~/.codex/config.toml, ~/.codex/hooks.json, and ~/.codex/compact_prompt.md. It prints every changed path, backup, validation result, and a ready-to-run rollback command:

uv run rag install --codex --restore /absolute/path/to/codex-rollback-<id>.json

Use the exact rollback-record pathname printed by the successful install. Check mode writes nothing. Repository support does not mean a particular machine has completed hook trust or live compaction verification; review /hooks and confirm rag status after every installation.

The Codex target never installs a scheduler. On a Codex-only macOS setup, use uv run rag backup --install-launchd for scheduled database backups; Linux uses the handbook's cron/systemd recipes.


What makes it different

Four things that, together, set it apart from both file-based knowledge wikis and hosted RAG stacks:

1. Real hybrid search over a curated graph

Documents are chunked and embedded into halfvec(1024) columns indexed with HNSW, and each chunk is embedded with the multilingual bge-m3 model — so semantic recall works in any language — and also carries generated tsvectors for English/German keyword full-text. A single query runs vector ANN and full-text together and returns one ranked list. Documents aren't an undifferentiated pile: they carry a type (concept, lesson, signal, synthesis, reference, …) and connect through a typed edge graph (references, extends, depends_on, supersedes, contradicts, …), so rag get shows you not just a document but its neighbourhood.

2. Automatic session-mining — the star feature

This is what makes agentic-rag feel like it grows rather than sits there. When a supported coding session ends, a lifecycle hook enqueues the transcript. A single-writer worker drains the queue and calls the configured Codex or Claude CLI to pull out durable memories, lessons, and signals, each saved through the write gateway behind a near-duplicate gate. A fix you discovered today becomes something a future session can recall, with no "remember to write this down" step. It reads only your local session transcripts.

3. Knowledge domains you grow and curate

Domains are just data — a label for where to look (general is seeded at init). Add them with rag domain add, scope any search with --domain, and let the importer derive them from an existing store's topics. Curation is first-class: rag review reports near-duplicates, dangling links, and stale pins; refuting a fact archives it with a required reason + evidence; rag purge removes only already-refuted documents, and only as rag_admin.

4. Built-in maintenance, backup, and restore-testing

rag backup runs pg_dump -Fc locally (plus an optional copy to a synced directory you configure), with rotation. rag maintenance is a tiny, single-flight, always-exit-0 job that ticks the worker, rotates logs, and — weekly — runs a report-only restore-test: it restores your newest dump into an isolated scratch database, compares row counts, and drops it. A backup you've never restored isn't a backup; this one checks itself.


How it works

   Supported coding sessions                rag save · migrate import · MCP write tools
   (queued & mined on session end)          (you, or an agent)
            │                                               │
            └───────────────────┬───────────────────────────┘
                                │
                    one audited write gateway
       strips secret-shaped tokens · chunks + embeds (local Ollama) · resolves edges · logs
                                │
                ┌───────────────┴────────────────┐
        PostgreSQL + pgvector              typed knowledge graph
        documents · chunks halfvec(1024)   edges: references · extends · …
        HNSW ANN  +  EN/DE full-text
                                │
                one hybrid ranked search  ──  vector ⊕ full-text
                                │
   session-start context · prompt-time recall · read-only MCP behind a privilege boundary

Claude Code and Codex continuity use a separate operational path. The Claude flow:

PreCompact ──► bounded deterministic snapshot ──► audited checkpoint
     │                    └──► priority enrichment job ──► provider CLI
     └──► stdout: versioned compact instructions (+ checkpoint id)
Claude compacts (instructions appended)
     │
PostCompact ──► mark boundary + store compact_summary as bounded handoff
     │
SessionStart(source="compact") ──► checkpoint + handoff, ≤ 10,000 chars ──► next request

The Codex flow (PreCompact stays silent; PostCompact stores no handoff):

PreCompact ──► bounded deterministic snapshot ──► audited checkpoint
     │                    └──► priority enrichment job ──► provider CLI
     ▼
Codex compacts
     │
PostCompact ──► mark boundary only (cannot inject context)
     │
SessionStart(source="compact") ──► bounded checkpoint context ──► next request

Native Codex memories are complementary, not the canonical record. With the installed policy they remain enabled and can be inspected with /memories; agentic-rag is canonical for durable searchable knowledge, audit history, and explicit continuation checkpoints.

  • Your chosen CLI provider. Every LLM call goes through the single agentic_rag.llm seam and the configured local Codex or Claude command. Transcript digests and checkpoint deltas are character-bounded; each curation call covers one selected document/evidence set. Mining also sends secret-stripped copies of all matching pin bodies without changing the stored pins. Embeddings never leave the box (local Ollama), so retrieval is independent of provider authentication.

  • One audited write path. Every change — a manual save, a mined memory, an import — funnels through a single gateway that strips secret-shaped tokens, regenerates chunks + embeddings in one transaction, resolves dangling edges, and writes an audit row. Embeddings fail open (queued for retry if Ollama is down); nothing else does.

  • Least privilege, by role. Three login roles enforce a destruction-protection matrix: rag_reader (SELECT only, used by search and the read-only MCP), rag_writer (INSERT/UPDATE but no DELETE/TRUNCATE/DROP), and rag_admin (migrate, purge, restore).

  • It steps aside, not in front. If Ollama is down, search degrades to full-text-only and returns a warning rather than failing; the maintenance job always exits 0.

The full story is in the handbook — the mental model, everyday use, configuration, importing an existing wiki, and the architecture and design rationale.


Comparison

agentic-rag sits between two worlds: the file-based LLM-Wiki family (human-readable Markdown with a lint/graph layer) and typical RAG stacks (hosted or API-driven retrieval you feed documents to). Every cell below is marked honestly — including the rows where each of them beats us.

Legend: ✅ shipped in code and operationally established · 🧪 shipped and installed, live verification pending · ⚠️ partial / caveated · ❌ absent

vs LLM-Wiki systems (file-based knowledge wikis)

Capability

agentic-rag

File-based LLM-Wiki

Hybrid vector + full-text ranked search (ANN at scale)

⚠️ lexical/graph, file-scan

Bilingual full-text (EN + DE) + semantic recall

⚠️

Transactional, audited writes through one gateway

⚠️

Scales to a large corpus (HNSW index)

⚠️ file-scan slows

Human-readable, git-diffable plain-text store

⚠️ import/export MD; store is Postgres

Zero-infrastructure (no DB/service to run)

❌ needs Postgres + Ollama

Imports an existing llm-wiki store

rag migrate

✅ it is one

Bottom line: if you want a git-tracked pile of Markdown, a file-wiki wins on its home turf. If you want fast hybrid recall over a growing corpus with transactional safety, agentic-rag wins — and it can import your existing llm-wiki to get you there.

vs typical RAG systems (retrieval frameworks / hosted memory)

Capability

agentic-rag

Typical RAG stack

Local-first store, provider CLI under your control, no hosted RAG service ¹

⚠️ usually a hosted service

Auto-populates from Claude Code sessions (mining)

❌ you feed it

Codex session mining and continuity

🧪

❌ you feed it

Claude Code compaction continuity (checkpoint + handoff)

🧪

Self-curation (dedup, near-dup gate, refute/archive)

⚠️

Typed knowledge graph (edges) alongside vector search

⚠️

One audited write gateway with secret stripping

Read/write privilege boundary for subagents (RO MCP)

⚠️

Turnkey managed hosting / large ecosystem ²

⚠️ self-host, young

Bottom line: a hosted RAG stack wins on turnkey scale and ecosystem. agentic-rag wins on being local-first, using a provider CLI under your control with no hosted RAG service in the loop, self-populating-from-your-own-work, and self-curating — a memory that fills and tidies itself instead of one you have to keep feeding.


Configuration

Config lives in one TOML file at ~/.agentic-rag/config.toml. Every key is optional — omit a section to keep its defaults.

Setting

Default

What it does

[db] name

agentic_rag

Database name.

[db] host

"" (local socket)

Empty = local unix socket; set it for a networked/remote server.

[embed] model

bge-m3

Ollama embedding model tag.

[embed] dim

1024

Fixed to the schema (halfvec(1024)); init-db refuses a mismatch.

[ollama] url

http://localhost:11434

Local Ollama endpoint.

[backup] local_dir

~/.agentic-rag/backups

Where pg_dump archives are written.

[backup] cloud_dir

— (unset)

Opt-in copy to a synced/cloud directory; unset means backups remain local. Provider-bound LLM inputs are disclosed above.

[pg] bin_dir

auto-resolved

Only needed if pg_dump/pg_restore/psql aren't on PATH (e.g. Postgres.app).

Roles are created passwordless by default, relying on local peer/trust auth (Postgres and agentic-rag on the same machine). For a networked or shared instance, set role passwords with ALTER ROLE … and let libpq authenticate via ~/.pgpass or PGHOST/PGPORT/PGPASSWORD — see the handbook's privacy chapter.

The Claude target separately manages autoCompactWindow = 500000 in ~/.claude/settings.json — a 1M context compacting at 500K with a [1m] model; model is reported, never rewritten, and long-context requests above 200K input tokens cost more on API billing. The Codex target separately manages a 600000 context window and a 500000 total-token compaction threshold, leaving a 100K reserve, plus native memories and the compact prompt. Official GPT-5.6 capacity is 1.05M, but inputs above 272K are subject to higher provider pricing and may add latency; see Configuration and Privacy, cost & control.


📖 Documentation / Handbook

The full story lives in the agentic-rag Handbook — a single, progressively-ordered read from the mental model through everyday use, configuration, importing, and the engine's architecture and design rationale. A few key chapters:

Start at the handbook index for the one-line "what you'll learn" map of every chapter.


Status

agentic-rag is young but solid — a real engine, openly developed. Repository and rollout state are listed separately below:

  • Storage & search: PostgreSQL + pgvector schema, HNSW ANN blended with EN/DE full-text into one ranked list, the typed edge graph, the three-role destruction-protection matrix.

  • The document write gateway: secret stripping on document inputs and generated document writes, one-transaction chunk + embed + edge-resolve + audit, embeddings that fail open with a retry queue.

  • Session mining: hooks → queue → single-writer worker → configured Codex/Claude CLI → gateway, with a near-duplicate gate and provider-outage circuit breaker.

  • Curation & safety: rag review, refute-as-archive, and admin-only rag purge (removes only already-refuted documents, as rag_admin).

  • Maintenance & backups: pg_dump backups with rotation, the tiny always-exit-0 maintenance job, and the weekly report-only restore-test. macOS auto-schedules via launchd; Linux uses the documented cron/systemd recipes.

  • Claude integration: two user-scope MCP servers (read-write + a read-only server behind a privilege boundary for subagents), idempotent install that preserves foreign hooks.

  • Claude continuity in code: six Claude hooks, the PreCompact stdout compact prompt, the bounded compact_summary handoff, the 10,000-character SessionStart cap, the managed 1M/500K policy, and rag install --check/--restore.

  • 🔒 Claude continuity live rollout: rag install, /hooks review, /autocompact, manual/automatic compaction, and SessionEnd tail capture on the maintainer machine remain open (backlog 0.3).

  • Codex continuity in code: audited checkpoints, bounded capture and restoration, asynchronous enrichment, all six lifecycle handlers, a versioned compact prompt, recoverable installer/check mode, and checkpoint health in rag status.

  • 🔒 Codex continuity live rollout: the pre-install whole-diff/security review and provider-bound pin hardening are complete. The live global install, /hooks trust, manual/automatic compaction, provider-recovery, and SessionEnd smoke tests remain open. See FEATURES.md and blocker-first BACKLOG.md.

  • Quality: a content-free repository with a comprehensive local test suite; exact verification counts belong in rollout evidence, not a static badge.

The clearest gap relative to the field is maturity: it's newly public and self-hosted, without the turnkey hosting or large ecosystem of established RAG stacks.


Acknowledgments

agentic-rag builds on other people's ideas and tools:

  • Andrej Karpathy — the LLM-Wiki idea that shaped the durable-knowledge model this import path speaks to.

  • The llm-wiki format — topic-partitioned Markdown with an optional memory store; rag migrate imports it wholesale, so an existing wiki carries straight over.

  • pgvector and Ollama (bge-m3) — the local vector search and embeddings underneath everything.

  • Anthropic — Claude Code, its hooks, and the MCP integration agentic-rag plugs into.


Contributing

Tests come first (TDD), and docs/ is kept in step with the code. A warn-only doc-reminder hook ships under .githooks/: if a commit touches agentic_rag/ or sql/ without touching docs/, it prints a reminder — it never blocks. Enable it once per clone:

git config core.hooksPath .githooks

Run the suite with uv run pytest. See the handbook's Contributing chapter for dev setup, the test database, and code layout.

License

MIT. The repository is code-only and content-free — your documents, embeddings, and config stay in your own PostgreSQL database, on your own machine.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

No tool schema history has been recorded yet.

Maintenance

ActivityMaintained
ResponsivenessUnresponsive

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Persistent memory for Claude Code. Automatically indexes every conversation and provides production-grade hybrid search (BM25 + vectors + reranker) via MCP tools. 100% local, zero config, zero API keys, zero invoice.
    16
    57
    7
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Persistent memory for Claude Code — hybrid search, knowledge graph, session lifecycle.
    17
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    A local, persistent, semantically-aware knowledge graph for AI coding agents like Claude Code, providing efficient session memory with minimal token cost and zero runtime network calls.
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides persistent, searchable memory for Claude Code using local SQLite, semantic embeddings, and full-text search, enabling Claude to recall and retrieve context across sessions and projects without external services.
    19
    4
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/phense/agentic-rag'

If you have feedback or need assistance with the MCP directory API, please join our Discord server