Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
7.2 KiB
7.2 KiB
theme, title, info, author, transition, mdc, lineNumbers, monaco, aspectRatio, canvasWidth, layout, class
| theme | title | info | author | transition | mdc | lineNumbers | monaco | aspectRatio | canvasWidth | layout | class |
|---|---|---|---|---|---|---|---|---|---|---|---|
| seriph | Lesson 11 — Why Does Semantic Search Miss Exact Answers? | English video course for AI Agents in Depth | Bojie Li | slide-left | true | false | false | 16/9 | 980 | cover | cover |
Build · Chapter 3 · Memory and Knowledge
Why Does Semantic Search Miss Exact Answers?
Chunking, dense retrieval, sparse retrieval, and evaluation
Lesson 11 of 42 · 19 minutes · RAG Basics; Document Chunking; Dense Embeddings; Sparse Embeddings
layout: center class: text-center
The central question
Why can a vector search understand a topic yet miss the exact identifier the user needs?
Why this problem matters
Chunking
Defines the atomic units that can be found.
Dense retrieval
Matches meaning and paraphrase.
Sparse retrieval
Matches exact words, numbers, and identifiers.
Three ideas to keep in view
Recall@k
Did the relevant item enter the candidate set?
ANN index
Trade exact search for speed and memory
BM25
Weight exact terms with saturation and length normalization
The book's visual model
BM25 scoring mechanism for exact lexical retrieval
Dense vs. Sparse
Dense
- Semantic similarity
- Handles paraphrases
- May miss rare identifiers
Sparse
- Exact lexical match
- Transparent term scores
- Misses synonyms
The failure modes are complementary.
Measure retrieval before generation
candidates = index.search(query, k=10)
recall = any(doc.id in relevant_ids for doc in candidates)
for rank, doc in enumerate(candidates, 1):
print(rank, doc.score, doc.id)
Test the claim
3-42 min
Compare ANN index behavior
Observe: Latency, recall, memory, and incremental-update trade-offs
3-52 min
Explain one BM25 score
Observe: Per-term TF, IDF, saturation, and length effects
Demo budget: 4 minutes · one contiguous terminal block
class: course-terminal
Live demo
Switching to the terminal
$ uv run python chapter3/dense-embedding/cli.py --compare-ann -k 10
$ uv run python chapter3/sparse-embedding/cli.py -q "model distillation" --explain
Run the command(s), narrate decisions, and point to the observation—not just the output.
What the evidence supports
Finding 1
A retrieval failure can begin at chunk boundaries rather than the model.
Finding 2
ANN algorithms differ in update behavior as well as speed.
Finding 3
Exact and semantic search solve different parts of the problem.
layout: center
Where the claim stops
Boundary condition
A higher retrieval score does not prove that the retrieved passage answers the question.
layout: center
Engineering takeaway
Design rule
Evaluate the candidate set independently before asking whether generation is good.
Continue the experiment
HNSW structure
book-en/images/fig3-7.svg
BM25 implementation
chapter3/sparse-embedding/
Dense model comparison
chapter3/dense-embedding/
layout: center class: text-center
Pause and apply
Your turn
Which queries in your domain are dominated by identifiers rather than semantics?
layout: center class: text-center
Next · Lesson 12
Fuse complementary retrievers, then organize knowledge beyond flat chunks.
→