> ## Documentation Index
> Fetch the complete documentation index at: https://chatcli.edilsonfreitas.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Bootstrap and Persistent Memory

> Customize the agent's identity with bootstrap files and maintain context across sessions with long-term memory and daily notes.

ChatCLI offers two complementary systems for customizing and contextualizing the agent: **bootstrap files** to define personality and rules, and **persistent memory** to maintain context across sessions.

<Note>
  The bootstrap and memory system is **fully connected** to the system prompt flow. Files are automatically loaded and injected into all interactions — in chat mode as well as in `/agent` and `/coder` modes.
</Note>

***

## Bootstrap Files

Bootstrap files are Markdown documents automatically loaded into the agent's system prompt. They define who the assistant is, how it behaves, and which rules it should follow.

### Supported Files

ChatCLI loads **exactly 5 bootstrap files**, in this order. All are optional — if they don't exist, they are simply skipped:

| File | Purpose | When to use |
| :- | :- | :- |
| `AGENTS.md` | Sub-agent definitions and their roles | When you want to instruct the orchestrator about available agents and how to use them |
| `SOUL.md` | Assistant personality, tone and style | To define "who" the assistant is — how it speaks, thinks and behaves |
| `USER.md` | User preferences and project context | To inform the stack, conventions, preferred tools and project context |
| `IDENTITY.md` | Agent identity and capabilities | To define "what" the assistant is — name, capabilities, limitations |
| `RULES.md` | Explicit rules and restrictions | For strict guardrails — what it MUST and MUST NOT do |

<Warning>
  File names are **exact and case-sensitive**. ChatCLI looks for `AGENTS.md`, `SOUL.md`, `USER.md`, `IDENTITY.md` and `RULES.md`. A directory without `AGENTS.md` may carry `CHATCLI.md` or `CLAUDE.md` in its place (that order); any other name (`README.md`, …) is **not loaded** by the bootstrap system.
</Warning>

### Instruction hierarchy and imports

`AGENTS.md` is not a single file: ChatCLI reads one from **every directory between the project root and the working directory**, root first, so a package-level file refines the repository-level one (the same layout Codex and Claude Code use). Nested files are prefixed with a `<!-- path -->` comment so the model knows where each block came from; the assembled text is capped at **32 KiB** (truncated at a line boundary with a visible note). The global `~/.chatcli/AGENTS.md` is **always merged first** (Claude Code semantics; it used to be a fallback only). `AGENTS.local.md` / `CLAUDE.local.md` / `CHATCLI.local.md` append per-machine, gitignored additions after their shared file; `AGENTS.override.md` replaces a directory's `AGENTS.md` outright (the Codex convention); `.claude/rules` is read below `.chatcli/rules` (a same-name rule in `.chatcli/rules` wins). `@import` targets are symlink-resolved and must live under the workspace or the global instruction directory — an out-of-tree import is logged and left as text — and each imported file is capped at 32 KiB on its own before the hierarchy cap, so a cloned repository's instruction file cannot pull arbitrary files from your disk into the prompt. A watched context watches at most 2,000 directories, skips directory names matched by the root's `.gitignore`, and warns when the watch is partial.

Any bootstrap file can pull another in with an import line of its own:

```markdown theme={"system"}
@docs/conventions.md
@import ../shared/security-rules.md
@~/.chatcli/snippets/review-checklist.md
```

Paths resolve relative to the importing file (absolute and `~` paths work too), up to **4 hops** deep and never the same file twice; a missing target leaves the line as written, and lines inside code fences are ignored. Imported files take part in cache invalidation like the bootstrap files themselves.

### Loading Priority

Files are searched at two levels, with the workspace taking priority:

<Steps>
  <Step title="Workspace (project root)">
    Project-specific configurations. Takes **priority** over global. The project root is detected automatically (see below).
  </Step>

  <Step title="Global (~/.chatcli/)">
    User default configurations. Serves as a **fallback** when the file does not exist in the workspace.
  </Step>
</Steps>

If the same file exists at both levels, the workspace version prevails. Global files serve as fallback.

#### Automatic Workspace Detection

ChatCLI uses `detectProjectDir()` to find the real project root. Instead of simply using the current directory (CWD), it **walks up the directory tree** looking for project markers:

1. Checks if the current directory contains `.git/` or `.agent/`
2. If not found, moves to the parent directory and repeats
3. Continues until a marker is found or it reaches the filesystem root
4. If no marker is found, falls back to the CWD

This means you can run ChatCLI from **any subdirectory** of your project and bootstrap files at the root will be found normally.

<Info>
  Recognized markers: `.git` (Git repository) and `.agent` (explicit ChatCLI marker). Only one needs to exist to define the workspace root.
</Info>

#### Example Scenarios

| CWD at startup | Marker found | Detected workspace | Files loaded |
| :- | :- | :- | :- |
| `~/project/` | `~/project/.git` | `~/project/` | `~/project/SOUL.md`, etc. |
| `~/project/src/pkg/` | `~/project/.git` | `~/project/` | `~/project/SOUL.md` (walks up 2 levels) |
| `~/project/src/pkg/` | none | `~/project/src/pkg/` | Global only (`~/.chatcli/`) |
| `~/monorepo/services/api/` | `~/monorepo/.git` | `~/monorepo/` | `~/monorepo/SOUL.md`, etc. |
| `~/tmp/` | none | `~/tmp/` | Global only (`~/.chatcli/`) |

### Detailed Examples

<Tabs>
  <Tab title="SOUL.md">
    Defines **personality and tone**. Place at `~/.chatcli/SOUL.md` (global) or `./SOUL.md` (project):

    ```markdown theme={"system"}
    # Personality

    You are a technical assistant specialized in software engineering.
    Be concise and direct. Prefer practical examples over theoretical explanations.
    When suggesting code, use best practices and tests.

    # Tone

    - Professional but approachable
    - Prefer short and objective responses
    - Use bullet points for lists
    - Default language: English
    ```
  </Tab>

  <Tab title="USER.md">
    Defines **user and project context**. Ideal for `./USER.md` in the project directory:

    ```markdown theme={"system"}
    # Project Context

    - Stack: Go 1.25, gRPC, Kubernetes
    - Database: PostgreSQL 16
    - CI/CD: GitHub Actions
    - Style: conventional commits, trunk-based development

    # Preferences

    - Always use tables for comparisons
    - Prefer simple solutions without over-engineering
    - Tests with idiomatic Go table-driven tests
    ```
  </Tab>

  <Tab title="IDENTITY.md">
    Defines **what the assistant is**. Usually global at `~/.chatcli/IDENTITY.md`:

    ```markdown theme={"system"}
    # Identity

    You are ChatCLI, an intelligent terminal assistant.

    ## Capabilities

    - Code reading and editing via @coder plugin
    - Shell command execution with user approval
    - Log analysis and error diagnosis
    - Git operations (status, diff, log, commit)

    ## Limitations

    - You do NOT have internet access
    - You CANNOT install packages without approval
    - Your patches may fail if context has changed
    ```
  </Tab>

  <Tab title="RULES.md">
    Defines **strict rules and guardrails**. Can be global or per-project:

    ```markdown theme={"system"}
    # Mandatory Rules

    1. NEVER execute `rm -rf` without explicit confirmation
    2. NEVER commit directly to the main branch
    3. ALWAYS run tests after modifying code
    4. ALWAYS use conventional commits (feat:, fix:, chore:)

    # Security Restrictions

    - Do not expose secrets, tokens or API keys in logs
    - Do not modify files outside the project directory
    - Do not execute commands with sudo
    ```
  </Tab>

  <Tab title="AGENTS.md">
    Defines **sub-agents and their roles** for the orchestrator:

    ```markdown theme={"system"}
    # Custom Agents

    ## @devops
    Infrastructure specialist for Docker, Kubernetes and CI/CD.
    Use for deployment, monitoring and pipeline configuration tasks.

    ## @dba
    PostgreSQL database specialist.
    Use for queries, optimization, migrations and performance analysis.

    ## @security
    Security auditor focused on OWASP Top 10.
    Use for code review with focus on vulnerabilities.
    ```
  </Tab>
</Tabs>

### Where to Place Files

"Project root" means the directory containing `.git/` or `.agent/` -- automatically detected by `detectProjectDir()`, not necessarily the CWD.

```text theme={"system"}
# GLOBAL configuration (applies to all projects)
~/.chatcli/SOUL.md
~/.chatcli/IDENTITY.md
~/.chatcli/RULES.md
~/.chatcli/USER.md
~/.chatcli/AGENTS.md

# Per-PROJECT configuration (overrides global)
# "Project root" = directory with .git/ or .agent/
<project-root>/SOUL.md          # In project root
<project-root>/USER.md          # In project root
<project-root>/RULES.md         # In project root (project rules)
<project-root>/IDENTITY.md      # In project root (rare)
<project-root>/AGENTS.md        # In project root (project agents)
<project-root>/services/AGENTS.md   # Nested: read when the cwd is under services/
<project-root>/services/CLAUDE.md   # Fallback name for the same role (CHATCLI.md first)
```

<Warning>
  `CHATCLI_BOOTSTRAP_DIR` only overrides the **global** directory (`~/.chatcli/`), not the workspace detection. Project-level files (detected via `.git` or `.agent`) still take priority over global ones, regardless of `CHATCLI_BOOTSTRAP_DIR`.
</Warning>

<Tip>
  Recommended strategy: `SOUL.md` and `IDENTITY.md` global (they're about the assistant), `USER.md` and `RULES.md` per-project (they're about the work context).
</Tip>

### Smart Cache

Bootstrap files use mtime-based (modification time) caching:

* On the first read, the content is cached in memory
* Subsequent reads check if the mtime has changed
* If the file was modified, the cache is automatically invalidated
* `IsStale()` checks if any file has changed since the last load

***

## Persistent Memory

The memory system maintains context across ChatCLI sessions using **structured storage** with multiple components that learn about you and your work over time.

### System Architecture

```text theme={"system"}
Conversation -> memoryWorker (3min) -> LLM extraction -> ProcessExtraction()
                                                              |
                    +───────────+───────────+─────────────────+────────────+
                    v           v           v                 v            v
              FactIndex    Profile    TopicTracker    ProjectTracker  DailyNote
              (scored)     (JSON)     (JSON)          (JSON)          (.md)
                    |
                    v
              Compactor (6h check, 24h cycle)
              |-- Level 1: Score-based pruning + archive
              +-- Level 2: LLM consolidation
                    |
                    v
              MEMORY.md (regenerated, never source of truth)
```

### Resilient extraction — nothing is lost silently

Extraction depends on a background LLM call — and a provider outage **must not cost the conversation its memory**. Three layered defenses:

1. **Fallback chain**: extraction tries the session's active client and, on failure, walks `CHATCLI_MEMORY_FALLBACK_PROVIDERS` (or `CHATCLI_FALLBACK_PROVIDERS`), with a per-attempt timeout.
2. **Durable on-disk queue**: a segment that fails on **every** provider is written to `~/.chatcli/memory/pending/` (atomic writes) and retried on later runs, oldest first — it **survives restarts**. The worker drains the backlog **at boot and on every extraction tick** (a `-p`-only user no longer fills the queue until the oldest segments are dropped), a segment that fails repeatedly is dropped after two attempts, the queue is capped (100 segments) and corrupt files are dropped without wedging the rest. On exit the unextracted tail of the session is queued for the next one, and shutdown waits up to 5 s for an in-flight extraction to finish its write (a hung provider call never holds the exit).
3. **A visible notice**: two consecutive failures print a one-liner in the terminal (`memory: extraction failing…`) — days of silent fact loss can no longer happen.

<Note>
  The gateway consults this memory too: the daemon's persona calls `@memory recall` before answering "I don't know" to personal questions. See [Chat Gateway](/gateway/chat-gateway).
</Note>

### Storage Structure

All memory lives in `~/.chatcli/memory/`:

```text theme={"system"}
~/.chatcli/memory/
|-- MEMORY.md              # Human-readable summary (regenerated from FactIndex)
|-- memory_index.json      # Facts with relevance scores
|-- user_profile.json      # User profile (name, role, expertise)
|-- topics.json            # Recurring topics with frequency
|-- projects.json          # Active projects with context
|-- episodes.json          # Episodic timeline (dated work entries)
|-- graph.json             # Persisted knowledge-graph cache (see below)
|-- usage_stats.json       # Usage patterns and statistics
|-- memory_archive.json    # Archived facts (low score)
|-- weekly/                # Weekly digests (consolidated from daily notes)
|   +-- 2026-W27.md
|-- monthly/               # Monthly digests (kept forever)
|   +-- 2026-06.md
|-- 202603/                # Daily notes for March 2026
|   |-- 20260301.md
|   +-- 20260306.md
+-- 202602/
    +-- 20260228.md
```

### Components

<AccordionGroup>
  <Accordion title="FactIndex -- Long-Term Memory" icon="database">
    Replaces the old append-only MEMORY.md. Each fact has:

    * **Unique ID** via SHA-256 content hash (automatic deduplication)
    * **Category**: architecture, pattern, preference, gotcha, project, personal
    * **Temporal score**: `(1 + log(accessCount)) * exp(-days * ln2 / halfLife)`
    * **Tags** for keyword search

    Frequently accessed and recent facts get higher scores. Old, never-accessed facts naturally decay.
  </Accordion>

  <Accordion title="UserProfile -- User Profile" icon="user">
    Automatically detected by the AI during extraction and editable by you/the model in any mode:

    * Name, role, expertise level, company, location
    * Preferred language and communication style
    * Lists with lifecycle: **certifications, skills, goals, interests and directives** (restating an item **supersedes** the old one instead of duplicating; `_replace`/`_done`/`_remove` suffixes rewrite)
    * **Directives** with severity (hard rules vs preferences) and **per-project scope** (`[scope:<project>] rule` is only injected when that workspace is active)
    * **Stances** — technical positions recorded **with the why** (`"position :: reason"`)
    * **Milestones** — dated timeline of what happened
    * Structured **Environment** (`env_os`, `env_shell`, ...) — auto-migrated from legacy preferences
    * **Per-field provenance and freshness**: each attribute knows whether it came from you or from extraction and when it was last confirmed; fields unconfirmed for 120+ days are flagged as possibly stale in the prompt
    * **Privacy tier**: finance/health/family/document keys are auto-tagged `[sensitive]` — they personalize answers but never enter code, examples or generated artifacts (`sensitive_mark`/`sensitive_unmark` for manual control)
    * Most used commands (top 10) and general preferences

    View your profile with `/memory profile`. Old profiles polluted by append-only versions **self-heal on the next load** (fragments, progress duplicates and completed goals leave on their own).
  </Accordion>

  <Accordion title="TopicTracker -- Recurring Topics" icon="tags">
    Tracks technical topics discussed:

    * Mention frequency
    * Recency (recent topics weigh more)
    * Links to related facts

    View with `/memory topics`.
  </Accordion>

  <Accordion title="ProjectTracker -- Projects" icon="folder-open">
    Tracks projects you work on:

    * Name, path, description
    * Technologies used
    * Status (active, paused, completed)
    * Last activity

    View with `/memory projects`.
  </Accordion>

  <Accordion title="PatternDetector -- Usage Patterns" icon="chart-line">
    Analyzes how you use ChatCLI:

    * Total sessions and average duration
    * Peak activity hours
    * Preferred features (chat, agent, coder)
    * Common errors and resolutions

    View with `/memory stats`.
  </Accordion>
</AccordionGroup>

### Smart Retrieval

Instead of dumping all memory into the system prompt, ChatCLI uses **intelligent retrieval**:

1. Extracts keywords from the last few conversation messages
2. Searches relevant facts in FactIndex by keyword match + temporal score
3. Respects a configurable budget (default: 4000 characters)
4. Prioritizes: Profile > Projects > Topics > Relevant facts > Recent notes > Episodes > **Related (graph)** > Trajectory (digests) > Usage patterns

The **`## Related (graph)`** section is the knowledge graph participating in recall: the top-ranked facts are expanded one hop through the graph, surfacing structurally adjacent memories (a neighboring fact, the episode from the day it was learned, a connected topic or session) that keyword and vector ranking missed. It is purely additive — it never displaces a ranked fact, only spends leftover budget — and it appears both in `full`-mode injection and in every `@memory recall` pull.

<Info>
  Facts accessed by the retriever automatically get their scores bumped, creating a virtuous cycle: the more useful a fact is, the more it appears.
</Info>

### Injection mode: push vs pull

How memory reaches the model — in **every mode**: chat, `/agent` and `/coder` — is controlled by `CHATCLI_MEMORY_MODE`:

| Mode | Behavior |
| - | - |
| `index` (**default**) | Injects only a **compact, stable index** (profile summary + top topic/project names + fact tally by category + episode count) and lets the model pull detail on demand via the memory recall tool. Bounded, cacheable per-turn cost. |
| `full` | Injects the full Smart Retrieval (above) every turn — the classic "push" behavior. |
| `off` | Injects no memory; bootstrap files still apply. |

**Persistent-memory bootstrap card** — every surface's system prompt opens with a `[PERSISTENT MEMORY — REAL, CROSS-SESSION]` card: a session-start snapshot of live counts (long-term facts, dated episodes, saved conversations with the latest one named, in-flight work-board cards) plus an explicit anti-amnesia directive — *never claim you have no memory of past sessions; pull first* (`@memory recall`, `@session search`, `@board list`; in chat, the card points the user at `/session attach` instead of tools chat does not have). The card is frozen per session so it lives in the **cached** prefix. Proactive recall also fires from the **very first turn** of a fresh session: the current message feeds the ranking hints, purely referential questions ("what did we talk about earlier?") list the most **recent** sessions instead of lexically strong old ones, and the ambient recall window covers the 60 most recently saved sessions.

**Proactive auto-recall** closes the pull model's gap: models sometimes skip the recall call even when a stored gotcha is exactly what they need. In `index` mode, each turn also ranks the fact index against keywords from the recent messages and rides the **top 3 matches** into the prompt as a tiny `[MEMORY AUTO-RECALL]` block (700-byte cap; labeled as data, not instructions, each fact on one line — an external provider's answer rides under its own `[EXTERNAL MEMORY]` fence), followed by a one-line **`Related (graph):`** pointer naming what sits next to those facts in the [knowledge graph](#knowledge-graph-memory-neighbors--map) — so the model knows there is more to pull before it claims ignorance. The block lives in the *uncached* trailing region next to the wall-clock context, so the stable digest stays byte-identical and prompt caching is unaffected. Gated by `CHATCLI_MEMORY_AUTORECALL` (**on** by default).

Two quality rules govern what the block may carry. **Relevance floor:** a fact found by keywords alone must match roughly a third of the turn's hints (`computeRelevance ≥ 0.34`), and a fact found by the vector index must clear the cosine floor (`min_cosine_score`, 0.25) — one incidental token can no longer inject an unrelated fact into every turn. **Vectors when available:** with an embedding provider configured (`CHATCLI_EMBED_PROVIDER`), the block ranks with the same blended semantic + lexical + temporal ranker the recall tool uses, embedding the turn's message once (skipped for messages under 8 characters), so a paraphrase with zero keyword overlap can surface; keyless setups keep the lexical ranker. **No reinforcement on push:** facts injected proactively are no longer marked accessed — reinforcing a fact merely for being pushed was self-entrenching (a spuriously surfaced fact climbed its own ranking). Access is reinforced when the model actually pulls detail through the memory tool.

**Reinforcement waits for evidence.** The facts a turn injected are remembered, and after the reply only those the answer evidently drew on — enough of the fact's own significant terms appear in the reply (two, or a third of them for long facts) — get their access counter bumped. A fact the model ignored keeps its score; nothing is demoted. **Episodes ride along:** when the turn's message is long enough, up to two timeline episodes ranked by BM25 over the episodic store (summary, outcome, project, refs) are appended as dated bullets, so "what did we do about the invoice migration" surfaces the work unit, not just the facts around it. The same BM25 ranking now powers `@memory timeline` queries inside a date window.

The pull mode (`index`) shrinks the per-turn memory block by \~88% on a 500-fact store without losing access to detail — see [Token Efficiency › Pull-first memory](/context/token-efficiency) for the full measurement. **Chat honors `index` too**: the sanctioned chat memory tool exception (`CHATCLI_CHAT_MEMORY`, on by default) gives chat the same pull path, so it receives the compact digest plus a recall hint instead of the full retrieval. Only when that exception is disabled does chat fall back to `full` — chat without a pull path must never be starved of memory. Chat also receives the proactive `[MEMORY AUTO-RECALL]` and `[SESSION RECALL]` blocks and the knowledge-graph card, previously agent/coder-only (the session-recall wording in chat suggests `/session attach`, since chat has no session tool). The recall tool uses HyDE + vector search, so pulled detail matches push quality. Check the active mode with `/config memory`.

### Memory Configuration

The memory system has tunable parameters via environment variables:

| Variable | Default | Description |
| - | - | - |
| `CHATCLI_MEMORY_MODE` | `index` | Injection mode in every mode (chat, agent, coder): `index` (pull), `full` (push) or `off` |
| `CHATCLI_MEMORY_AUTORECALL` | `true` | Proactive top-3 fact injection + `Related (graph):` line in `index` mode (see above) |
| `CHATCLI_MEMORY_GRAPH` | `true` | Persisted knowledge-graph cache (`graph.json`) + recall neighborhood expansion; `off` restores derive-per-call with no expansion |
| `CHATCLI_GRAPH_INDEX` | `true` | Per-turn injection of the graph map-of-content card (`@memory map`/`neighbors` pull is always available regardless) |
| `CHATCLI_CHAT_MEMORY` | `true` | Chat memory tool exception — the pull path that lets chat honor `index` mode |
| `CHATCLI_MEMORY_MAX_SIZE` | `32768` (32KB) | Max size of rendered MEMORY.md |
| `CHATCLI_MEMORY_RETENTION_DAYS` | `30` | Daily note retention before cleanup |
| `CHATCLI_MEMORY_MAX_FACTS` | `500` | Max facts in the FactIndex |
| `CHATCLI_MEMORY_RETRIEVAL_BUDGET` | `4000` | Max chars of memory injected in system prompt (`full` mode) |

In addition to environment variables, the internal `Config` struct defines additional defaults:

| Parameter | Default Value | Description |
| - | - | - |
| `CompactionInterval` | 24 hours | Minimum interval between full compactions |
| `DecayHalfLifeDays` | 30.0 | Temporal decay half-life for fact scores |
| Check interval | 6 hours | How often the system checks if compaction is needed |

### How Memories Are Created

The background worker now extracts **6 types of information** (previously only 2):

1. **DAILY** -- What was done (files, commands, errors, tasks)
2. **LONGTERM** -- New facts to remember permanently
3. **PROFILE\_UPDATE** -- Information about the user (name, role, expertise)
4. **TOPICS** -- Technical topics discussed
5. **PROJECTS** -- Projects worked on
6. **EPISODES** -- Dated, durable units of completed work (see [Episodic memory](#episodic-memory-the-work-timeline))

The worker fires after 4+ new messages with a 2-minute cooldown, and also every 3 minutes during long sessions.

**What a call costs, and what it is worth.** The extraction request is shaped for the prompt cache: the instructions are one cached system block, the workspace line and the existing long-term memory another, and only the new conversation segment travels as the user message. On a provider with a prefix cache the three-minute cadence reads the first two back from a warm entry and pays the write for the segment alone; nothing is sent twice. The worker runs on the session model unless `CHATCLI_COMPACT_MODEL` names a cheaper one, which serves the worker as well as compaction — a fast model there is the single largest lever on background spend. `/cost` shows the balance under the worker's line: facts and episodes written, cost per item, and how many proactive recalls put stored facts back in front of the model. A worker that writes memory the session never reads is what that line exposes.

#### Extraction Process

The Memory Worker follows this internal flow:

1. **EnhancedExtractionPrompt**: Sends the recent conversation history to the LLM with a structured prompt requesting information extraction
2. **Expected output**: The LLM returns text with well-defined section headers:
   * `## DAILY` -- Summary of what was done in the session
   * `## LONGTERM` -- New facts for long-term memory
   * `## PROFILE_UPDATE` -- User profile updates
   * `## TOPICS` -- Technical topics identified
   * `## PROJECTS` -- Projects mentioned or worked on
3. **ParseEnhancedResponse()**: Parses the response and extracts each section individually
4. **Deduplication**: Each fact receives a unique ID via SHA-256 content hash. Facts with a hash matching an existing one are automatically discarded
5. **Profile merging with lifecycle**: `PROFILE_UPDATE` changes are merged with the existing profile. List fields **upsert** (restating an item supersedes the old one instead of duplicating), and extraction may emit rewrite operations: `goals_done=`/`goals_remove=` remove finished goals (moving them to `milestone=`/`certifications=`), `goals_replace=` replaces the whole list — the same suffixes work for certifications, skills, interests and directives. Your instructions about the profile ("remove X from my goals") are **applied**, never recorded as if they were facts

### Automatic Compaction

The system runs periodic compaction to prevent uncontrolled growth:

<Steps>
  <Step title="Check (every 6 hours)">
    Checks if the fact count exceeds 80% of the limit or if 24h have passed since the last compaction.
  </Step>

  <Step title="LLM Compaction (preferred)">
    Sends all facts to the AI with instructions to: merge duplicates, remove obsolete, consolidate related. Preserves original metadata.
  </Step>

  <Step title="Score-based Fallback">
    If the LLM call fails, archives facts with scores below 0.1 to `memory_archive.json`.
  </Step>

  <Step title="Daily Note Cleanup">
    Removes notes older than the retention period (default: 30 days). Empty directories are cleaned up.
  </Step>

  <Step title="MEMORY.md Regeneration">
    Rewrites MEMORY.md from the FactIndex -- always up-to-date, never the source of truth.
  </Step>
</Steps>

### Weekly and Monthly Digests (Trajectory)

Daily notes expire after \~30 days — **rollups** preserve the long-range narrative before that:

* When an ISO week ends, that week's daily notes consolidate into `weekly/2026-W27.md` (LLM summary with a deterministic fallback — never depends on the provider being up).
* When a month ends, its weekly digests consolidate into `monthly/2026-06.md`.
* Retention: weeklies keep \~26 weeks; **monthlies are kept forever** (minimal cost).
* A bounded **Trajectory** section enters the memory context with the latest monthly + recent weeklies — that is how "what you have been doing these past months" survives daily-note cleanup.

The process is idempotent and runs on its own in the memory worker (startup check plus every 12h).

### Automatic Migration

When starting for the first time with the new system, ChatCLI detects if a legacy `MEMORY.md` exists (without `memory_index.json`) and migrates automatically:

1. Each line/bullet is converted to an individual fact
2. Categories are detected from markdown headers
3. Tags are extracted by technical keywords
4. The original file is saved as `MEMORY.md.bak`

### `/memory` Command

| Subcommand | Description |
| - | - |
| `/memory` or `/memory today` | Show today's notes |
| `/memory yesterday` | Show yesterday's notes |
| `/memory 2026-03-04` | Show notes from a specific date |
| `/memory week` | Show notes from the last 7 days |
| `/memory longterm` | Show MEMORY.md content |
| `/memory list` | List all memory files (includes structured JSONs) |
| `/memory load <date>` | Load a day's notes into conversation context |
| `/memory profile` | Show detected user profile |
| `/memory profile set <field>=<value>` | Set/update a profile field manually |
| `/memory remember <fact>` | Explicitly add a long-term fact (accepts a `[category]` prefix) |
| `/memory forget <substring>` | Remove long-term facts containing the substring |
| `/memory topics` | Show tracked topics with frequency |
| `/memory projects` | Show tracked projects with status |
| `/memory stats` | Full statistics (sessions, peak hours, errors, features) |
| `/memory facts [category]` | List facts with scores (filter by category) |
| `/memory timeline [q]` | Dated episode history; `q` filters by time ("3 meses atrás", "abril", "2026-04") and/or content |
| `/memory compact` | Force immediate compaction (LLM + note cleanup) |
| `/memory export [path]` | Write facts, episodes, topics, projects and profile as JSONL (default `~/.chatcli/exports/memory-<ts>.jsonl`, mode 0600; sealed line by line when `CHATCLI_ENCRYPTION_KEY` is set) |
| `/memory import <path>` | Merge an export into this machine's memory: every fact is sanitized (one line, bounded, secrets masked) and its id re-derived from the content (known content keeps the stronger metadata; a chosen id cannot smuggle a duplicate), confidence is clamped, tags capped and categories validated; episodes dedup, topics/projects/profile merge field by field; a file whose header declares a newer schema is refused; nothing is deleted |
| `/memory recall <query>` | Dry-run of auto-recall: the facts the query would inject, each with its score and **why** (keyword/semantic/recency signals, use count, provenance, source project) |
| `/memory why` | The last turn's auto-recall with the same reasons |

Every store also reconciles with its file before rewriting it — facts and episodes already did; the **profile, topics and projects** now merge too (fresher field wins, lists union), so a REPL, a gateway daemon and an MCP server writing the same memory never erase each other's learning. Daily notes are appended through an atomic rename, so a crash mid-write cannot tear a note. A fact learned in another project is labeled `(from: ~/path)` in the recalled block, a git root and its packages counting as one project.

#### Manual editing and extended profile

Automatic detection doesn't always catch everything, so you can **edit memory explicitly**. Beyond name/role/expertise/company/location, the profile covers **certifications, skills, goals, interests, directives, stances, milestones and environment (env\_\*)**.

List fields have a **lifecycle**, they are not append-only: new items enter, a restated item (same text apart from status/parenthetical) **supersedes** the old one in place, and operation suffixes rewrite — `_replace` overwrites the whole list (empty value clears it), `_done`/`_remove` delete matching items. Commas inside parentheses are safe (`"Quiz X (Provider, 60 questões)"` stays whole).

```text theme={"system"}
> /memory profile set company=ACME Corp
> /memory profile set certifications=CKA, AWS SAA        # upsert, deduplicated
> /memory profile set goals_done=earn the CKA            # finished goal leaves the list
> /memory profile set milestone=Earned the CKA           # ...and becomes a dated milestone
> /memory profile set goals_replace=launch product Y     # replaces ALL goals
> /memory profile set stance=prefer keyless backends :: less setup friction
> /memory profile set directives=[scope:my-repo] always run the linter before pushing
> /memory profile set env_shell=zsh
> /memory profile set sensitive_mark=monthly_income      # marks as private
> /memory remember [preference] Prefers Go over Python for CLIs
> /memory forget Python                                  # removes FACTS containing "Python"
```

<Info>
  `forget` only acts on **facts**; to remove profile items use the suffixes (`goals_remove=...`). Directives with `[scope:<project>]` are only injected when that workspace is active; without the tag they apply globally.
</Info>

#### `@memory` tool (agent, coder and chat)

Inside `/agent` and `/coder`, the model can persist and explore memory on its own via the **`@memory`** tool (cmds `remember`, `profile`, `forget`, `recall`, `timeline`, `neighbors`, `map`):

```text theme={"system"}
<tool_call name="@memory" args='{"cmd":"remember","args":{"content":"User earned the AWS Solutions Architect certification","category":"personal"}}' />
```

So when you tell the agent something new (e.g. a fresh certification), it records it into your profile/long-term facts without you running `/memory` manually. The `profile` subcommand accepts every lifecycle operation (`goals_done`, `goals_replace`, `stance`, `milestone`, `env_*`, `sensitive_mark`, ...).

**In chat too**: memory/profile updates are the **fourth sanctioned exception** of tool-less chat (alongside `ask_user`, read-only knowledge and `@graphview`). When you reveal or correct a durable fact about yourself, the model persists it **in the same turn** — no more "I'll keep that in mind" without writing anything. The exception also carries the **read side**: it is the pull path that lets chat honor `index` memory mode (compact digest + on-demand recall) and receive the proactive auto-recall, session-recall and graph-card blocks. Gated by `CHATCLI_CHAT_MEMORY` (**on by default**), flippable with `/config chat memory on|off` — disabling it reverts chat to the full-push memory model.

### Episodic memory: the work timeline

Facts capture what is **true**; episodes capture what **happened**. An episode is a dated, durable unit of completed work — summary, outcome and refs (files, PRs, commands) — extracted automatically by the background worker (`## EPISODES` section) and stored forever in `~/.chatcli/memory/episodes.json`. Unlike daily notes (30-day retention) and their digests, episodes never expire, so "what did we do three months ago?" finally has a structured answer.

* **Storage** follows the fact-index disciplines: atomic writes, quarantine on corruption, and reconciling multi-process rewrites under a cross-process file lock (`<store>.lock` next to each store, taken for the merge-and-write critical section, so the REPL, the gateway daemon and the MCP server never erase each other — the same lock guards the context store and the vector index). The same work item re-extracted on the same day **dedups and enriches** outcome/refs instead of duplicating. Capped at 2000 episodes (oldest dropped first).
* **`@memory timeline`** is the model-facing chronological view: `{project?, from?, to?, query?, limit?}`. `from`/`to` accept ISO dates (`2026-04`, `2026-04-12`) **or natural expressions in English and Portuguese** ("3 months ago", "há 3 meses", "abril", "last week") — a time expression inside `query` works too.
* **Temporal recall**: a `@memory recall` query carrying a time expression ("o que fizemos em abril?") automatically prepends the matching timeline slice to the result — relevance ranking alone cannot answer "when". When the window came from natural language and the leftover words match nothing, the window wins and returns unfiltered (the date was the real signal).
* **`/memory timeline [q]`** is the human view — same temporal parsing, rendered in the terminal.
* **Rollups are grounded in episodes**: weekly/monthly digests now lead with the period's episodes and use daily notes as color, so the long-range narrative stops decaying — and a week whose notes already expired still gets its digest from episodes alone.

```text theme={"system"}
<tool_call name="@memory" args='{"cmd":"timeline","args":{"query":"3 months ago"}}' />
<tool_call name="@memory" args='{"cmd":"timeline","args":{"project":"chatcli","from":"2026-04","to":"2026-06"}}' />
```

#### Knowledge graph (`@memory neighbors` / `map`)

What ChatCLI knows about you is not a flat list: facts, topics, projects, skills, tags — and, since the persisted graph landed, **episodes, saved sessions and knowledge bases** — form a graph (Obsidian-style, in the core) derived from the relationships the stores already record (topic↔fact, fact→project, episode↔project, episode↔same-day facts, session↔project, session↔episode, KB↔tags, skill triggers) plus `[[wikilinks]]` in note text. Access is through `@memory` itself:

* `recall` → search by **content** ("which facts match these words?") — the output now ends with a `## Related (graph)` section of one-hop neighbors.
* `neighbors <subject>` → **local graph**: backlinks + related notes for a subject ("what connects to this?").
* `map` → overview (counts by kind + hubs).

**The graph is a persisted cache**, not a per-call derivation: it lives at `~/.chatcli/memory/graph.json` and rebuilds only when a source actually changes (a fact is learned or forgotten, a session is saved, a skill is installed, a context is created — whether by you or by the model through its tools). On the next boot the cache is adopted instantly, so `@memory map` and the per-turn card cost nothing. The file is self-healing: a corrupt or stale copy is simply discarded and re-derived. Gate: `CHATCLI_MEMORY_GRAPH` (**on** by default; `off` restores the old derive-per-call behavior with no recall expansion). Node/edge counts are visible in `/config memory`.

Context discipline: only a tiny, deterministic index card rides each turn (cache-friendly); depth is pulled on demand. To visualize, use `/graph` (renders the graph to an image via embedded go-graphviz).

#### Fact quality (confidence, provenance, reconciliation)

Each fact carries **confidence** and **provenance**: what you state directly outranks a background-extraction guess, and a re-observed fact climbs — confidence scales the score (ranking and survival against decay/pruning). On write, ChatCLI **reconciles**: a rephrasing reinforces the existing fact instead of duplicating, and a same-subject update of equal-or-higher confidence **supersedes** the stale one (conservatively — a weak guess never wipes a strong fact). Topics gain a **rolling summary** of what was discussed, becoming real knowledge nodes. Legacy memory indexes are enriched once on startup, without losing anything.

<AccordionGroup>
  <Accordion title="What goes in each component?" icon="layer-group">
    * **FactIndex**: Stable, long-lasting facts -- decisions, patterns, gotchas, preferences
    * **UserProfile**: Who you are -- name, role, expertise, language
    * **TopicTracker**: What you talk about -- Go, Docker, K8s, etc.
    * **ProjectTracker**: What you work on -- chatcli, my-app, etc.
    * **PatternDetector**: How you work -- schedules, features, common errors
    * **Daily notes**: What happened today -- temporal and specific
  </Accordion>

  <Accordion title="Can I edit memories manually?" icon="file-pen">
    Yes! All files are plain JSON or Markdown:

    ```bash theme={"system"}
    # View profile
    cat ~/.chatcli/memory/user_profile.json | jq .

    # View facts with scores
    cat ~/.chatcli/memory/memory_index.json | jq '.[0:5]'

    # Edit today's note
    vim ~/.chatcli/memory/$(date +%Y%m)/$(date +%Y%m%d).md
    ```

    JSON changes are loaded on next startup. MEMORY.md is regenerated and should not be edited directly.
  </Accordion>

  <Accordion title="How does fact scoring work?" icon="chart-simple">
    Each fact has a score calculated by:

    ```
    score = (1 + log(1 + accessCount)) * exp(-daysSinceAccess * ln(2) / halfLifeDays)
    ```

    * **accessCount**: How many times the fact was used by the retriever
    * **daysSinceAccess**: Days since last access
    * **halfLifeDays**: Decay half-life (default: 30 days)

    Frequently and recently accessed facts get high scores. Never-accessed facts decay to \~0 after 3-4 half-lives.
  </Accordion>
</AccordionGroup>

### What Gets Injected into the Prompt

The ContextBuilder assembles the following block and injects it as a system prompt prefix:

```text theme={"system"}
## AGENTS.md

[content of AGENTS.md]

---

## SOUL.md

[content of SOUL.md]

---

## USER.md

[content of USER.md]

---

## IDENTITY.md

[content of IDENTITY.md]

---

## RULES.md

[content of RULES.md]

---

# Memory

## Long-term Memory

[content of MEMORY.md]

## Recent Daily Notes

### 2026-03-04

[content of March 4th note]

### 2026-03-05

[content of March 5th note]

### 2026-03-06

[content of today's note]
```

Empty sections (missing files) are **automatically omitted** — only what exists is injected.

***

## Dynamic Context (CWD + Disambiguation)

ChatCLI automatically injects into the system prompt:

* The current **date and time**
* The current **working directory** (the process CWD)
* A **disambiguation instruction**: the model is told to ALWAYS resolve "here", "this project" and relative paths against the current CWD — never against paths found in long-term memory

<Info>This solves the common problem where long-term memory holds paths of earlier projects and the model mistakes "the current project" for one that was in memory. Facts from other projects are annotated with `(from: /other/project)` so their origin is clear.</Info>

***

## Path-Specific Rules

Besides the global `RULES.md`, you can create **conditional rules per path** in `.chatcli/rules/`:

```text theme={"system"}
.chatcli/rules/
├── go-style.md          # Rules for Go files
├── api-conventions.md   # Rules for APIs
└── testing.md           # Rules for tests
```

Each file can carry a `paths:` frontmatter that sets which files the rule applies to:

```markdown theme={"system"}
---
paths: ["*.go", "src/**"]
---

Always wrap errors with fmt.Errorf("%w", err) in Go.
Never ignore returned errors.
```

Rules **without** a `paths:` frontmatter apply globally. Workspace rules (`.chatcli/rules/`) take **priority** over global rules (`~/.chatcli/rules/`) with the same name.

<Tip>See [Path-Specific Rules](/context/path-rules) for the full documentation.</Tip>

***

## Configuration

<Tabs>
  <Tab title="Environment Variables">
    ```bash theme={"system"}
    CHATCLI_BOOTSTRAP_ENABLED=true
    CHATCLI_BOOTSTRAP_DIR=/path/to/bootstrap/files
    CHATCLI_MEMORY_ENABLED=true
    ```

    | Variable | Default | Description |
    | - | - | - |
    | `CHATCLI_BOOTSTRAP_ENABLED` | `true` | Enable/disable bootstrap file loading |
    | `CHATCLI_BOOTSTRAP_DIR` | `~/.chatcli/` | Alternative directory for **global** bootstrap files. Use this when you want to keep your files (SOUL.md, RULES.md, etc.) in another location, such as a versioned repository or a directory shared across machines |
    | `CHATCLI_MEMORY_ENABLED` | `true` | Enable/disable the persistent memory system |

    <Note>
      `CHATCLI_BOOTSTRAP_DIR` only overrides the **global** directory (`~/.chatcli/`). Project-level files (detected via `.git` or `.agent` markers) still take priority over global ones.
    </Note>
  </Tab>

  <Tab title="Via Helm Chart">
    ```yaml theme={"system"}
    # values.yaml
    bootstrap:
      enabled: true
      definitions:
        SOUL.md: |
          You are a DevOps assistant...
        USER.md: |
          The user prefers Go...

    memory:
      enabled: true
      subPath: memory   # default: memory's own directory of the sessions PVC
    ```

    The Helm chart creates ConfigMaps for bootstrap files and mounts them at `/home/chatcli/.chatcli/bootstrap/`; a change to `bootstrap.definitions` rolls the pods on `helm upgrade`, while an edited `bootstrap.existingConfigMap` needs `kubectl rollout restart`. `bootstrap.enabled` with neither `definitions` nor `existingConfigMap` mounts nothing.

    With persistence on, memory lives in its own directory of the sessions PVC (`memory.subPath`, default `memory`, mounted at `~/.chatcli/memory`) while the sessions stay at the PVC root; without persistence it is a 200Mi emptyDir. Earlier charts (up to 1.211.x) kept memory at the PVC root, among the session files: on the first start after the upgrade the server copies memory's own files from there into the memory directory, once and without moving or overwriting anything (the chart sets `CHATCLI_MEMORY_LEGACY_DIR` for this). `memory.subPath: ""` keeps the old shared layout. Details in [Kubernetes deployment](/start/docker-deployment#memory-on-the-sessions-pvc).
  </Tab>
</Tabs>

***

## Context Injection Optimization (Prompt Caching)

ChatCLI optimizes token costs when contexts are attached using three complementary strategies:

### Unified System Prompt with Cache Hints

Contexts attached via `/context attach` are injected as **system prompt**, not as user messages. This enables provider-level prompt caching:

| Provider | Mechanism | Discount |
| - | - | - |
| Anthropic | `cache_control: ephemeral` | \~90% |
| OpenAI | Automatic prompt caching | \~50% |
| Google | Context caching API | Variable |

The system prompt block contains:

1. Bootstrap (SOUL.md, USER.md, etc.)
2. Memory (MEMORY.md + daily notes)
3. **Attached Contexts** (new — previously injected as user messages)
4. K8s Watcher (if active)

Since the system prompt is identical across turns, the provider caches it and charges tokens at a discount.

### Smart Compaction

Injected context messages (`/memory load`, summarized contexts) are **automatically truncated** during compaction (Level 1 — trimming). This prevents old reference context from consuming valuable token budget.

### Token Visibility

The `/context attached` command now shows:

* Estimated tokens per context
* Total tokens per turn
* Cache hints per provider
* Warnings for oversized contexts

When running `/context attach`, the feedback includes the estimated cost per turn.

***

## Best Practices

<CardGroup cols={2}>
  <Card title="Global SOUL.md, per-project USER.md" icon="layer-group">
    Keep your preferred personality globally and technical context per project.
  </Card>

  <Card title="Keep MEMORY.md concise" icon="compress">
    Keep only stable and confirmed facts — not session-specific ones.
  </Card>

  <Card title="Daily notes for journaling" icon="calendar-day">
    Use them to record decisions, solved problems, and temporal context.
  </Card>

  <Card title="Don't duplicate CLAUDE.md" icon="copy">
    If you already use CLAUDE.md or project instructions, avoid duplicating them in bootstrap.
  </Card>
</CardGroup>

<Tip>Periodically review your memories and remove outdated ones to keep the context relevant.</Tip>

***

## Next Steps

<CardGroup cols={2}>
  <Card title="Conversation Control" icon="clock-rotate-left" href="/usage/conversation-control">
    Use /compact and /rewind to manage conversation size and state.
  </Card>

  <Card title="Sessions" icon="floppy-disk" href="/context/session-management">
    Save and reuse conversations across projects.
  </Card>
</CardGroup>

## External memory provider (MCP)

`CHATCLI_MEMORY_PROVIDER=mcp-only:<server>` uses the external provider alone (no embedded recall block). The embedded memory stays the default. An organization that already runs a memory service plugs it in through MCP with no ChatCLI code: `CHATCLI_MEMORY_PROVIDER=mcp:<server>` where `<server>` is a configured MCP server exposing two tools:

| Tool | Arguments | Role |
| - | - | - |
| `memory_recall` | `query`, `hints[]`, `budget_chars` | returns text appended to the auto-recall block of every turn (chat and agent/coder), bounded to 900 chars, 5 s |
| `memory_store` | `messages[{role, content}]`, `session` | receives every turn's new messages (system messages excluded), forwarded asynchronously and best-effort, 10 s |

The embedded facts, episodes and graph keep working alongside; the provider's answer rides after them. Any failure or timeout adds nothing and never blocks a turn. Shown in `/config memory`.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.