> ## Documentation Index
> Fetch the complete documentation index at: https://chatcli.edilsonfreitas.com/llms.txt
> Use this file to discover all available pages before exploring further.

# #4 RAG + HyDE — Semantic Retrieval

> Two retrieval layers over memory.Fact: Phase 3a generates an LLM hypothesis and extracts keywords; Phase 3b adds vector embeddings (Voyage / OpenAI / Bedrock Titan-Cohere) + pure-Go cosine search with blended ranking (semantic+lexical+temporal), a cosine floor, auto-migration and an evaluation harness (recall@k/MRR/nDCG). Reusable vindex primitive, JSON-persisted, zero CGO.

**HyDE** (Hypothetical Document Embeddings) is the classic technique of generating a **hypothetical answer** to the user's question and using that answer as an additional retrieval signal. In ChatCLI, HyDE operates in two complementary phases: **3a** expands keywords via LLM hypothesis, **3b** adds vector cosine search.

<Info>HyDE is **opt-in** (`CHATCLI_QUALITY_HYDE_ENABLED=true`) to keep the steady-state with no additional cost. Phase 3a costs +1 cheap LLM call; Phase 3b requires configuring an embedding provider.</Info>

***

## The problem HyDE solves

The pre-pipeline `memory.Fact` retrieval was **keyword-only**: the scorer matches tokens extracted from recent messages against tags and content of stored facts. Works well when vocabulary matches exactly — fails when the user uses synonyms or asks abstract questions.

Gap example:

<Tabs>
  <Tab title="Without HyDE">
    User: `how to do X in Go?`<br />
    Extracted keywords: `[do, go]`<br />
    Stored fact: `"use goroutines for concurrency in X pipelines"`<br />
    **Match: ❌** — "do" and "go" don't literally appear in the fact.
  </Tab>

  <Tab title="With HyDE 3a">
    User: `how to do X in Go?`<br />
    LLM hypothesis: `"In Go you would use goroutines and channels for X, leveraging the concurrency primitives..."`<br />
    Expanded keywords: `[do, go, goroutines, channels, concurrency, primitives]`<br />
    **Match: ✅**
  </Tab>

  <Tab title="With HyDE 3b">
    Hypothesis → embed (1024 dim via Voyage) → cosine search in vector store → top-3 semantically close facts.<br />
    **Match: ✅** — even with no shared keyword.
  </Tab>
</Tabs>

***

## Phase 3a — Hypothesis-based keyword expansion

<Steps>
  <Step title="User types query">
    The query enters `cli_llm.go` or `agent_mode.go`.
  </Step>

  <Step title="HyDEAugmenter.Augment">
    ```go theme={"system"}
    augmenter := memory.NewHyDEAugmenter(cfg, llmCallback, logger)
    expanded := augmenter.Augment(ctx, query, originalHints)
    ```
  </Step>

  <Step title="LLM generates short hypothesis">
    Prompt: "Write a 2-4 sentence plausible answer that uses the technical nouns that would appear in any matching note. Bilingual if the query mixes languages."
  </Step>

  <Step title="ExtractKeywords from the hypothesis">
    The same extractor already used in chat mode (en+pt stop words, min 3 chars).
  </Step>

  <Step title="Merge unique + lower-case">
    Original keywords + top-N from hypothesis, cap configurable via `CHATCLI_QUALITY_HYDE_NUM_KEYWORDS` (default 5).
  </Step>

  <Step title="FactIndex.Search uses the expanded set">
    Existing keyword-based scorer operates over richer hints → much higher recall.
  </Step>
</Steps>

<Tip>Phase 3a **works without configuring an embedding provider**. It's the recommended default if the cost of +1 light LLM call is acceptable.</Tip>

***

## Phase 3b — Vector embeddings + blended ranking

Adds **cosine similarity** search over fact embeddings — and, since the ranking refactor, **fuses the cosine score** with the lexical and temporal signals into a single ranking.

<Info>**What changed (PR #1027):** previously the vector hits were dissolved back into keywords and the cosine score was **discarded** — you paid for an embedding call and threw away the very semantic signal. Now the cosine flows straight into the `SearchBlended` ranker, so a fact matched only by paraphrase (zero lexical overlap) can still rank.</Info>

### Architecture

```text theme={"system"}
┌──────────────────┐
│ User query       │
└────────┬─────────┘
         │  Phase 3a: LLM hypothesis → expanded keywords
         ▼
┌──────────────────────────┐
│ EmbeddingProvider.Embed  │  (Voyage / OpenAI / Bedrock / Null)
└────────┬─────────────────┘
         │  vector float32
         ▼
┌──────────────────────────────────────┐
│ vindex.Index.Search(q, k, floor)     │  cosine top-K, O(n log k), pure-Go
└────────┬─────────────────────────────┘
         │  []Hit{ID, Score}   ← the cosine score is PRESERVED
         ▼
┌────────────────────────────────────────────────┐
│ FactIndex.SearchBlended(keywords, semantic, w)  │
│   final = wSem·cosine + wLex·keyword            │
│         + wTemp·recency    (min-max normalized)  │
└────────┬───────────────────────────────────────┘
         │
         ▼
    Facts ranked by FUSED relevance
```

### Blended ranking (`SearchBlended`)

Three independent, complementary signals, each **min-max normalized** across the candidate set before the weighted sum:

| Signal | Source | Captures |
| :- | :- | :- |
| **semantic** | cosine from the vector store | synonymy, paraphrase |
| **lexical** | keyword/tag overlap | exact terms, identifiers, file names |
| **temporal** | recency decay × access frequency | what the user actually uses |

Default weights (`DefaultRankWeights`): **semantic 0.55 · lexical 0.30 · temporal 0.15** — semantic-first, because the embedding call was already paid for. Fusion is **additive** (not multiplicative) on purpose: a fact with high cosine and zero keyword overlap **still ranks** — a product would zero it out. Min-max normalization is what makes the weights **provider-agnostic**: Voyage 1024-d and OpenAI 1536-d cosine both land in \[0,1] after normalizing.

<Tip>**Cosine floor** (`MinCosineScore`, default `0.25`): embeddings over normalized text are almost always weakly positive, so the old `> 0` cutoff admitted near-orthogonal noise. The floor keeps only genuinely related facts in the top-K.</Tip>

### Supported providers

<Tabs>
  <Tab title="Voyage (recommended)">
    ```bash theme={"system"}
    export CHATCLI_QUALITY_HYDE_ENABLED=true
    export CHATCLI_QUALITY_HYDE_USE_VECTORS=true
    export CHATCLI_EMBED_PROVIDER=voyage
    export VOYAGE_API_KEY=pa-...
    # Optional: pick model (default voyage-4, 1024-dim)
    export CHATCLI_EMBED_MODEL=voyage-4
    ```

    **Why Voyage:** it's the provider **recommended by Anthropic** for Claude workflows, offers high quality/\$ for retrieval, and the default model (`voyage-4`, 1024-dim) is the general-purpose tier of the voyage-4 family — same price as voyage-3, strictly better. `voyage-4-large` (quality flagship), `voyage-4-lite` (cheapest) and `voyage-code-4` (code retrieval) share the same embedding space.
  </Tab>

  <Tab title="OpenAI">
    ```bash theme={"system"}
    export CHATCLI_QUALITY_HYDE_ENABLED=true
    export CHATCLI_QUALITY_HYDE_USE_VECTORS=true
    export CHATCLI_EMBED_PROVIDER=openai
    export OPENAI_API_KEY=sk-...
    # Optional: pick model and custom dimension
    export CHATCLI_EMBED_MODEL=text-embedding-3-small
    export CHATCLI_EMBED_DIMENSIONS=512   # can reduce to save storage
    ```

    **Models:** `text-embedding-3-small` (default, 1536-dim, \$0.02/1M tokens) or `text-embedding-3-large` (3072-dim, higher quality). The native dimension is auto-detected from the model — only set `CHATCLI_EMBED_DIMENSIONS` if you want to **truncate** via Matryoshka (e.g. 512 to save storage).
  </Tab>

  <Tab title="Google (Gemini)">
    ```bash theme={"system"}
    export CHATCLI_QUALITY_HYDE_ENABLED=true
    export CHATCLI_QUALITY_HYDE_USE_VECTORS=true
    export CHATCLI_EMBED_PROVIDER=google
    export GEMINI_API_KEY=AIza...
    # Optional: pick model and truncated dimension
    export CHATCLI_EMBED_MODEL=gemini-embedding-2
    export CHATCLI_EMBED_DIMENSIONS=1536  # 128-3072 via Matryoshka
    ```

    **Models:** `gemini-embedding-2` (default, 3072-dim native, GA, multimodal-ready, free tier available) or the older `gemini-embedding-001` (superseded; shuts down 2028-05-14). Batches ride `:batchEmbedContents`; truncated dims are L2-normalized locally, so any `CHATCLI_EMBED_DIMENSIONS` between 128 and 3072 is safe for cosine similarity.
  </Tab>

  <Tab title="Bedrock (Titan + Cohere)">
    ```bash theme={"system"}
    export CHATCLI_QUALITY_HYDE_ENABLED=true
    export CHATCLI_QUALITY_HYDE_USE_VECTORS=true
    export CHATCLI_EMBED_PROVIDER=bedrock

    # Reuses the chat client's credentials chain — no extra API key.
    # Set region/profile if not already in the environment.
    export BEDROCK_REGION=us-east-1
    export AWS_PROFILE=my-profile        # optional

    # Optional: pick model
    export CHATCLI_EMBED_MODEL=amazon.titan-embed-text-v2:0   # default
    export CHATCLI_EMBED_DIMENSIONS=1024                      # Titan v2: 256/512/1024
    ```

    **Why Bedrock:** same AWS credentials chain as the chat — single billing, AWS compliance (CloudTrail, guardrails), VPC endpoints. Ideal for companies already running chat through Bedrock and wanting retrieval inside the same perimeter.

    **Supported families:**

    | Model ID | Family | Supported dims | Native batching |
    | :- | :- | :- | :- |
    | `amazon.titan-embed-text-v2:0` (default) | Titan v2 | 256, 512, 1024 | No — parallelized with 8-worker pool |
    | `amazon.titan-embed-text-v1` | Titan v1 | 1536 (fixed) | No — parallelized |
    | `cohere.embed-english-v3` | Cohere v3 (en) | 1024 (fixed) | Yes — one call per batch |
    | `cohere.embed-multilingual-v3` | Cohere v3 (multi) | 1024 (fixed) | Yes — one call per batch |
    | `cohere.embed-v4:0` | Cohere v4 (multi, 128K ctx) | 1536 (default) | Yes — one call per batch |
    | `amazon.nova-2-multimodal-embeddings-v1:0` | Nova MME | 256, 384, 1024, 3072 (default 3072) | No — parallelized with 8-worker pool |

    Dispatch between Titan and Cohere is automatic by the model id prefix. `CHATCLI_EMBED_DIMENSIONS` only applies to Titan v2 (256/512/1024); invalid values are rejected in `NewBedrock` before any AWS call.

    **No Converse API for embeddings:** embeddings only live in `InvokeModel`, so each family keeps its dedicated body schema — unlike chat, where Converse covers nearly everything.

    <Tip>
      **Hot-swap via `/reload`:** changing `CHATCLI_EMBED_PROVIDER` / `CHATCLI_EMBED_MODEL` / `CHATCLI_EMBED_DIMENSIONS` in `.env` and running `/reload` rebuilds the provider on the spot — rewiring `/context --rag` semantic retrieval and the memory vector index without restarting ChatCLI. The index auto-migrates when provider/dimension change.
    </Tip>

    See [AWS Bedrock](/providers/bedrock) for the full credentials chain config (SSO, IAM role, corporate proxy, VPC endpoint).
  </Tab>

  <Tab title="Null (default)">
    If `CHATCLI_EMBED_PROVIDER` isn't set, `NullProvider` is returned. `Embed()` returns an error — the code checks `embedding.IsNull(p)` before calling, so it gracefully falls back to keyword-only. **No surprises, no silent failures.**
  </Tab>
</Tabs>

### Pure-Go vector store — generic `vindex` primitive

<Info>**No CGO, no SQLite-vec, no external deps.** Just `float32[]` + cosine + JSON persistence in `~/.chatcli/memory/vector_index.json`.</Info>

The cosine index lives in a **generic, reusable** package, `llm/embedding/vindex`, extracted once a second consumer (the [semantic `/context` retrieval](/context/persistent-context)) appeared. `memory` is just a **thin adapter** over it — no vector machinery duplicated per package:

```go theme={"system"}
// llm/embedding/vindex — the single, agnostic primitive
type Index struct { /* provider, dim, entries map[string][]float32, path */ }
func (x *Index) Upsert(ctx, items map[string]string) error
func (x *Index) Search(query []float32, k int, minScore float64) []Hit  // top-K via min-heap
func (x *Index) Prune(keep map[string]struct{})                         // never serve stale vectors

// cli/workspace/memory/vector_store.go — fact-domain adapter
type VectorIndex struct { idx *vindex.Index; provider embedding.Provider }
type ScoredFact struct { ID string; Score float64 }
```

Top-K selection uses a **min-heap O(n log k)** (not a full O(n log n) sort), so cost scales with `k`, not corpus size — the scale ceiling is no longer "hundreds of facts". For the typical chatcli case linear search completes in microseconds; no HNSW or IVFFlat.

### Provider/dimension auto-migration

Switching provider or dimension (Voyage 1024 → OpenAI 1536, or voyage→cohere at the same 1024) **no longer needs a manual `rm`**. On load, the index detects the mismatch — of **dimension** (cosine between different arities is undefined) **or of provider** (two embedding spaces aren't comparable, even at the same dimension) — and **auto-clears** the cache, removing the file so backfill repopulates:

```text theme={"system"}
WARN vindex provider changed — auto-clearing for re-embed  on_disk_provider=voyage provider=cohere
WARN vindex dimension mismatch — auto-clearing for re-embed  on_disk_dim=1024 provider_dim=1536
```

<Warning>The provider→provider case at the same dim (e.g. voyage 1024 → cohere 1024) was previously **served silently as garbage** — cosine between different spaces is meaningless. It is now detected and re-embedded.</Warning>

### Lazy backfill

When retrieving a fact, if it has no vector (fact predates embeddings activation), the index spawns a **detached** goroutine to embed the top-500 visible facts:

```go theme={"system"}
// cli/workspace/memory/store.go:120
go func(items map[string]string) { //#nosec G118 -- detached on purpose
    if err := m.vectors.BackfillFacts(context.Background(), items); err != nil {
        m.logger.Warn("vector backfill failed", zap.Error(err))
    }
}(items)
```

<Tip>Backfill is **bounded** and configurable: at most `Config.BackfillBatchMax` facts per retrieve (default 500). With Voyage (\~$0.05/M tokens) that costs roughly US$ 0.001 for a fully cold cache. In a normal session, most of the index is embedded on the first interaction.</Tip>

***

## The @embed agent tool

The embedding backend is also exposed directly to the model as the **`@embed`** tool (agent/coder modes), the same way `@image` exposes image generation:

```text theme={"system"}
<tool_call name="@embed" args='{"cmd":"similarity","args":{"a":"reset a password","b":"recover account access"}}' />
```

| Subcommand | Purpose |
| - | - |
| `similarity {a, b}` | cosine similarity between two texts, in \[-1, 1] |
| `rank {query, candidates, k?}` | order candidate texts by semantic relevance to a query |
| `text {text\|texts, out?}` | embed texts; with `out`, save `{index, text, embedding}` JSONL for external pipelines |
| `status` | show provider, model and dimension |

It uses the same `CHATCLI_EMBED_PROVIDER` configuration described above. For **persistent RAG over documents**, the model is steered to `@context` (build/attach a knowledge base) + `@knowledge` (query it) instead — `@embed` covers the ad-hoc semantic operations that don't need an index.

## Evaluation — proving retrieval (not assuming it)

There used to be no way to **measure** whether retrieval was any good. Now there is a dependency-free evaluation harness in `cli/workspace/memory/eval` — standard macro-averaged IR metrics:

| Metric | What it measures |
| :- | :- |
| **recall\@k** | fraction of relevant facts retrieved in the top-k |
| **precision\@k** | fraction of the top-k that was relevant |
| **MRR** | mean reciprocal rank of the first hit |
| **nDCG\@k** | normalized discounted cumulative gain |

The versioned A/B (in `ranking_test.go`, with a deterministic embedding provider) compares keyword-only vs. blended ranking on queries phrased with **synonyms** absent from the fact text:

```text theme={"system"}
baseline (keyword): recall@1=0.2857  MRR=0.2857  nDCG@1=0.2857
candidate (blended): recall@1=1.0000  MRR=1.0000  nDCG@1=1.0000
delta              : recall@1=+0.7143
```

<Info>The harness runs the same way in CI (deterministic provider, reproducible) or against a real backend — it never imports a provider, so it stays neutral across all 14 supported ones. It is what turns "looks good" into "is good, measured", and guards against regressions.</Info>

### Ranking tunables (no new env vars)

The blended-ranking parameters live as `Config` fields with strong defaults — **no** new env vars were introduced (and the pre-existing `CHATCLI_MEMORY_*` vars, previously ignored on the structured path, are now applied via `ConfigFromEnv` + clamping):

| `Config` field | Default | Controls |
| :- | :- | :- |
| `RankWeights` | `{0.55, 0.30, 0.15}` | semantic / lexical / temporal weights |
| `MinCosineScore` | `0.25` | cosine floor in the top-K |
| `VectorTopK` | `12` | vector candidates per query |
| `BackfillBatchMax` | `500` | cap of facts embedded per retrieve |

***

## Full configuration

| Env var | Default | Effect |
| - | - | - |
| `CHATCLI_QUALITY_HYDE_ENABLED` | `false` | Master switch (phase 3a) |
| `CHATCLI_QUALITY_HYDE_USE_VECTORS` | `false` | Enable phase 3b (requires provider) |
| `CHATCLI_QUALITY_HYDE_NUM_KEYWORDS` | `5` | Hypothesis keyword cap in phase 3a |
| `CHATCLI_EMBED_PROVIDER` | — | `voyage` / `openai` / `bedrock` / `null` — **single source of truth** for provider selection |
| `CHATCLI_EMBED_MODEL` | provider default | Voyage: `voyage-4`. OpenAI: `text-embedding-3-small` / `-large`. Bedrock: `amazon.titan-embed-text-v2:0` (default), `amazon.titan-embed-text-v1`, `cohere.embed-english-v3`, `cohere.embed-multilingual-v3`, `cohere.embed-v4:0`, `amazon.nova-2-multimodal-embeddings-v1:0`. |
| `CHATCLI_EMBED_DIMENSIONS` | model native | OpenAI: truncate via Matryoshka. Bedrock Titan v2: 256 / 512 / 1024 (rejects others). Bedrock Titan v1 / Cohere v3: fixed dimension, ignored. |
| `BEDROCK_REGION` / `AWS_REGION` | `us-east-1` | AWS region — only used when `CHATCLI_EMBED_PROVIDER=bedrock`. |
| `AWS_PROFILE` | — | AWS profile — only used when `CHATCLI_EMBED_PROVIDER=bedrock`. |

***

## `/config quality` surfaces state

```text theme={"system"}
── RAG + HyDE (#4)
  CHATCLI_QUALITY_HYDE_ENABLED    : enabled
  CHATCLI_QUALITY_HYDE_USE_VECTORS: enabled
  CHATCLI_EMBED_PROVIDER          : bedrock
  CHATCLI_EMBED_MODEL             : amazon.titan-embed-text-v2:0
  CHATCLI_EMBED_DIMENSIONS        : 1024
  CHATCLI_QUALITY_HYDE_NUM_KEYWORDS: 5
  Vector provider                : bedrock:amazon.titan-embed-text-v2:0
  Vector entries                 : 127
```

***

## Integration with Reflexion

HyDE amplifies Reflexion's value: lessons persisted by #3 are retrieved with much higher recall when the next task doesn't use the exact same keywords. Workflow:

<Steps>
  <Step title="Turn 1: auth.go refactor fails (timeout)">
    Reflexion persists lesson: `"use Edit tool for large files"`, tags `[go, refactor, edit-tool]`.
  </Step>

  <Step title="Turn 5 (days later): 'help me split pkg/engine'">
    Query doesn't contain `refactor` or `edit`. Keyword-only would **miss** the lesson.
  </Step>

  <Step title="HyDE 3a generates hypothesis">
    `"To split a Go package, identify logical groupings and use refactor patterns with Edit tool for surgical changes..."`<br />
    Extracted keywords: `[split, package, refactor, edit, patterns, …]`
  </Step>

  <Step title="Match!">
    Lesson appears in system prompt. Coder picks `Edit` over `write` from the start.
  </Step>
</Steps>

***

## Caveats and tuning

<Warning>**Token cost** of phase 3a: \~200 tokens per retrieval turn. In workflows with many read turns, the cost compounds. Use `CHATCLI_QUALITY_HYDE_NUM_KEYWORDS=3` for tighter budget.</Warning>

<Warning>**Privacy**: the user query is sent to the embedding provider. For sensitive workloads in corporate environments, **Bedrock is the preferred path** — it stays inside the AWS perimeter (CloudTrail, VPC endpoint, IAM). For local workloads with no network, consider self-hosting (roadmap: Ollama-embedding provider).</Warning>

<Info>**Graceful fallback**: if the LLM fails or the embedding provider returns an error, retrieval falls back to keyword-only silently. No turn is aborted by HyDE failure.</Info>

***

## See also

<CardGroup cols={2}>
  <Card title="#3 Reflexion" icon="lightbulb" href="/agents/harness/reflexion">
    The lessons that HyDE retrieves with higher recall.
  </Card>

  <Card title="Bootstrap Memory" icon="brain" href="/context/bootstrap-memory">
    The layer underneath: how memory.Fact is populated and maintained.
  </Card>

  <Card title="Persistent Context" icon="database" href="/context/persistent-context">
    `/context attach` for explicit file contexts.
  </Card>

  <Card title="Full configuration" icon="sliders" href="/agents/harness/configuration">
    All envs and slashes.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.