LLM & RAG Systems

Begin with one hand-checkable next-token distribution, then earn prompts, embeddings, indexes, retrieval, context, citations, abstention, evaluation, security, operation, and recovery as separate verifiable layers.

Course details and reading size
Tutorial
Reading comfortAdjust lesson text without changing code or interface size.

Language models, tokens, and probabilistic generation

Objective Run one hand-checkable next-token baseline and distinguish token, context, probability, selection, evidence, and unsupported input.

Core explanation

A language model receives a sequence of tokens and produces a probability distribution for the next token. Repeating that operation creates text, but it does not consult a guaranteed internal database before each claim. A token is a model-specific piece rather than necessarily a character or word; this first experiment deliberately uses visible whitespace-separated pieces so every count can be checked by hand. The context is the sequence available to the current prediction, while a context window is the maximum supported length and not permanent memory. The baseline counts which piece followed each observed piece, divides each count by the total, and exposes one certain transition and one two-way distribution. A modern transformer uses learned representations and much longer context instead of this one-step table, but it retains the next-token prediction contract. Sampling chooses from a distribution; it cannot turn unsupported context into truth. When the baseline receives an unseen prefix, it abstains instead of inventing a transition. Later chapters add prompt contracts, retrieval, citations, evaluation, security, operation, and model adapters only after this capability-versus-evidence boundary is observable.

Fluent next-token prediction is a probabilistic capability; authoritative evidence and an honest abstention path remain separate responsibilities.

Follow generation one conditional choice at a time

A language model consumes the token sequence available in the current request and returns scores for possible next tokens. Software converts those scores into probabilities and chooses one token according to the decoding rule, appends it, and repeats. The complete answer is therefore a chain of conditional choices. It is not fetched as one stored paragraph, and the probability of a token is not the probability that the resulting claim is true.

Modern transformers calculate these scores through learned representations and attention over the available context. The teaching model does not require pretending every internal unit has an interpretable fact. Observe the external contract: input tokens, output distribution or selected token, stop condition, latency, and usage. Hold prompt and model constant, repeat requests, and record where wording or conclusions vary. That experiment makes probabilistic behavior concrete without claiming access to hidden reasoning.

Treat tokenization as a versioned model boundary

A tokenizer maps text to model-specific integer tokens. Common words may be one token, unusual identifiers may split into many pieces, whitespace can matter, and scripts or combining characters behave differently across vocabularies. Characters, bytes, words, and tokens are separate units. Count with the tokenizer used by the target model when setting input limits or estimating cost rather than multiplying a word count by a universal constant.

Truncation is a semantic operation. If an API silently drops the beginning or end, it can remove instructions, definitions, exceptions, or citations. Reserve space for trusted instructions, question, retrieved evidence, and output, then enforce each budget before the call. Test long identifiers, code, tables, several writing systems, and boundary-length inputs. Record tokenizer and model versions because an upgrade can change both count and interpretation.

Separate context, training influence, and application memory

The context window contains tokens supplied or generated for one invocation. Training shaped model parameters before the request, but it is not a searchable source with reliable provenance. Application memory is data a system deliberately stores and retrieves across turns. Calling all three “memory” hides different authority, privacy, staleness, deletion, and capacity rules. A larger context accepts more tokens; it does not guarantee the model will use every detail or resolve conflicts correctly.

Design conversations from explicit state. Retain the minimum needed user-approved facts, summarize only under a traceable policy, and retrieve authoritative records when current truth matters. Never infer that a model remembers a prior request merely because one interface sends history automatically. Test a critical condition near the beginning, middle, and end of context, with distractors and contradictions, then measure the actual task outcome.

Use decoding controls for behavior, not truth certification

Greedy decoding selects the highest-scoring next token, while sampling can use temperature, top-p, or other restrictions to vary choices. Fixed seeds may improve reproducibility in some systems but do not promise identical output across infrastructure or model revisions. Lower variability can help extraction and classification, while controlled diversity can help drafting. Neither setting validates claims or citations.

Build an evidence table with prompt version, model, decoding configuration, repeated outputs, structural validity, claim support, and task score. Compare settings on the same cases rather than deciding from one attractive answer. When a task requires an exact identifier, arithmetic result, permission check, or database state, compute it with deterministic code and ask the model only for the bounded language transformation it can safely provide.

CURRICULUM CONTEXTRelated courses and the course concept model