--- theme: seriph title: "Lesson 11 — Why Does Semantic Search Miss Exact Answers?" info: "English video course for AI Agents in Depth" author: Bojie Li transition: slide-left mdc: true lineNumbers: false monaco: false aspectRatio: 16/9 canvasWidth: 980 layout: cover class: cover ---
Build · Chapter 3 · Memory and Knowledge
# Why Does Semantic Search Miss Exact Answers?

Chunking, dense retrieval, sparse retrieval, and evaluation

Lesson 11 of 42 · 19 minutes · RAG Basics; Document Chunking; Dense Embeddings; Sparse Embeddings
--- layout: center class: text-center ---
The central question
Why can a vector search understand a topic yet miss the exact identifier the user needs?
--- # Why this problem matters

Chunking

Defines the atomic units that can be found.

Dense retrieval

Matches meaning and paraphrase.

Sparse retrieval

Matches exact words, numbers, and identifiers.

--- # Three ideas to keep in view

Recall@k

Did the relevant item enter the candidate set?

ANN index

Trade exact search for speed and memory

BM25

Weight exact terms with saturation and length normalization

--- # The book's visual model BM25 scoring mechanism for exact lexical retrieval
BM25 scoring mechanism for exact lexical retrieval
--- # Dense vs. Sparse

Dense

Sparse

The failure modes are complementary.
--- # Measure retrieval before generation ~~~python candidates = index.search(query, k=10) recall = any(doc.id in relevant_ids for doc in candidates) for rank, doc in enumerate(candidates, 1): print(rank, doc.score, doc.id) ~~~ --- # Test the claim
3-42 min

Compare ANN index behavior

Observe: Latency, recall, memory, and incremental-update trade-offs

3-52 min

Explain one BM25 score

Observe: Per-term TF, IDF, saturation, and length effects

Demo budget: 4 minutes · one contiguous terminal block
--- class: course-terminal ---
Live demo
# Switching to the terminal ~~~bash $ uv run python chapter3/dense-embedding/cli.py --compare-ann -k 10 $ uv run python chapter3/sparse-embedding/cli.py -q "model distillation" --explain ~~~
Run the command(s), narrate decisions, and point to the observation—not just the output.
--- # What the evidence supports

Finding 1

A retrieval failure can begin at chunk boundaries rather than the model.

Finding 2

ANN algorithms differ in update behavior as well as speed.

Finding 3

Exact and semantic search solve different parts of the problem.

--- layout: center ---
Where the claim stops
# Boundary condition
A higher retrieval score does not prove that the retrieved passage answers the question.
--- layout: center ---
Engineering takeaway
# Design rule
Evaluate the candidate set independently before asking whether generation is good.
--- # Continue the experiment
HNSW structure book-en/images/fig3-7.svg BM25 implementation chapter3/sparse-embedding/ Dense model comparison chapter3/dense-embedding/
--- layout: center class: text-center ---
Pause and apply
# Your turn
Which queries in your domain are dominated by identifiers rather than semantics?
--- layout: center class: text-center ---
Next · Lesson 12
Fuse complementary retrievers, then organize knowledge beyond flat chunks.