---
theme: seriph
title: "Lesson 11 — Why Does Semantic Search Miss Exact Answers?"
info: "English video course for AI Agents in Depth"
author: Bojie Li
transition: slide-left
mdc: true
lineNumbers: false
monaco: false
aspectRatio: 16/9
canvasWidth: 980
layout: cover
class: cover
---
Build · Chapter 3 · Memory and Knowledge
# Why Does Semantic Search Miss Exact Answers?
Chunking, dense retrieval, sparse retrieval, and evaluation
Lesson 11 of 42 · 19 minutes · RAG Basics; Document Chunking; Dense Embeddings; Sparse Embeddings
---
layout: center
class: text-center
---
The central question
Why can a vector search understand a topic yet miss the exact identifier the user needs?
---
# Why this problem matters
Chunking
Defines the atomic units that can be found.
Dense retrieval
Matches meaning and paraphrase.
Sparse retrieval
Matches exact words, numbers, and identifiers.
---
# Three ideas to keep in view
Recall@k
Did the relevant item enter the candidate set?
ANN index
Trade exact search for speed and memory
BM25
Weight exact terms with saturation and length normalization
---
# The book's visual model
BM25 scoring mechanism for exact lexical retrieval
---
# Dense vs. Sparse
Dense
- Semantic similarity
- Handles paraphrases
- May miss rare identifiers
Sparse
- Exact lexical match
- Transparent term scores
- Misses synonyms
The failure modes are complementary.
---
# Measure retrieval before generation
~~~python
candidates = index.search(query, k=10)
recall = any(doc.id in relevant_ids for doc in candidates)
for rank, doc in enumerate(candidates, 1):
print(rank, doc.score, doc.id)
~~~
---
# Test the claim
3-42 min
Compare ANN index behavior
Observe: Latency, recall, memory, and incremental-update trade-offs
3-52 min
Explain one BM25 score
Observe: Per-term TF, IDF, saturation, and length effects
Demo budget: 4 minutes · one contiguous terminal block
---
class: course-terminal
---
Live demo
# Switching to the terminal
~~~bash
$ uv run python chapter3/dense-embedding/cli.py --compare-ann -k 10
$ uv run python chapter3/sparse-embedding/cli.py -q "model distillation" --explain
~~~
Run the command(s), narrate decisions, and point to the observation—not just the output.
---
# What the evidence supports
Finding 1
A retrieval failure can begin at chunk boundaries rather than the model.
Finding 2
ANN algorithms differ in update behavior as well as speed.
Finding 3
Exact and semantic search solve different parts of the problem.
---
layout: center
---
Where the claim stops
# Boundary condition
A higher retrieval score does not prove that the retrieved passage answers the question.
---
layout: center
---
Engineering takeaway
# Design rule
Evaluate the candidate set independently before asking whether generation is good.
---
# Continue the experiment
---
layout: center
class: text-center
---
Pause and apply
# Your turn
Which queries in your domain are dominated by identifiers rather than semantics?
---
layout: center
class: text-center
---
Next · Lesson 12
Fuse complementary retrievers, then organize knowledge beyond flat chunks.
→