Files
ai-agent-book/chapter10/book-translation/sample_book/chapter1.md
T
liqiang b119135836
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
2026-08-20 13:12:50 +00:00

1.3 KiB

Chapter 1: Foundations of LLM Inference

A large language model turns text into numbers before it can reason about anything. Each chunk of text is first split into a token, the smallest unit the model consumes. Every token is then mapped to an embedding, a dense vector that captures its meaning in a high-dimensional space.

When a user sends a request, the text they write is called a prompt. The process of running the model over that prompt to produce an answer is called inference. The time between sending the prompt and receiving the first response is the latency that users feel directly.

A minimal inference call looks like this:

def generate(prompt: str, model) -> str:
    tokens = model.tokenize(prompt)      # split prompt into tokens
    embeddings = model.embed(tokens)     # map each token to an embedding
    output = model.forward(embeddings)   # run inference
    return model.detokenize(output)

Two numbers dominate the user experience. First, the number of tokens in the prompt, because a longer prompt costs more compute. Second, the latency of the first token, because a slow first token makes the whole system feel sluggish. Throughout this book we keep returning to these ideas: token, embedding, prompt, inference, and latency. Getting their definitions right now will save confusion later.