AI Agent Engineering
Begin with one scripted proposal that ordinary code can allow or reject, then earn schemas, tools, loops, retrieval, memory, planning, approval, identity, testing, observability, failure recovery, and live providers as separately verifiable layers.
Agent systems, autonomy, and useful boundaries
Objective Run one bounded proposal-to-tool decision and explain why the application, not the model, owns authority and stopping.
Core explanation
An agent combines a model or other proposer with state, tools, and a control loop that may select later work from earlier observations. A proposal is an untrusted candidate action, a tool is a narrow application capability, and authority is the permission to use that capability on a current resource. Begin with the user outcome and compare a deterministic function, fixed workflow, retrieval interface, and agent before accepting the extra uncertainty, cost, and attack surface of agent control. The first laboratory substitutes two scripted proposals for a live model so the safety boundary is repeatable. Ordinary Python checks the action against an allowlist and a one-call budget before any tool runs. The read proposal completes once, while the write proposal is rejected without a tool call. The application owns validation, policy, execution, evidence, and the terminal state; fluent model text cannot grant itself permission. Later chapters add schemas, adapters, state machines, retrieval, memory, planning, approval, evaluation, operation, and provider interfaces only after this boundary is visible.
Distinguish a model, workflow, and agent system
A model maps an input context to a candidate output. A retrieval system finds material. A tool executes ordinary application code. A workflow follows transitions selected by deterministic rules. An agent system combines these pieces with a control loop in which model-generated proposals can influence later steps. Calling one function or producing JSON does not make an agent, and adding a loop does not make a useful product. Draw the components and ask which one owns identity, authority, state, validation, budget, effect, evidence, and termination. Those responsibilities remain in ordinary code even when a model suggests the next action.
Use the least flexible mechanism that meets the outcome. A fixed function is best for a known calculation; a form and query serve structured search; retrieval answers from a corpus; a deterministic workflow coordinates known stages; an agent helps when paths cannot be enumerated economically and model judgment among bounded alternatives produces real value. Compare these choices with the same acceptance cases, latency, cost, accessibility, privacy, failure modes, and maintenance. Agent autonomy adds uncertainty and attack surface, so “the framework supports tools” is not a justification.
Define autonomy as a bounded grant of capabilities
Autonomy is not one switch. Specify what the system may observe, infer, propose, retrieve, call, change, communicate, and repeat. A course-support assistant might read published course material and propose a study plan while being unable to read private progress, contact anyone, enroll a learner, publish content, or change an account. Bind the grant to an authenticated actor, tenant, resource scope, purpose, time, and release. Separate read, draft, approve, and commit capabilities. A model never gains authority because it states a reason or because untrusted content requests a tool.
Write prohibited actions and escalation rules beside permitted actions. Unknown policy, conflicting evidence, private-data request, unavailable dependency, repeated failure, excessive cost, low-confidence consequential decision, and any action outside scope should reach a named terminal or approval state. Enforce total turns, time, token use, tool calls, bytes, parallelism, and cost before work is admitted. A kill switch must remove risky capabilities or route to a deterministic fallback without depending on the same failing model. Test the boundary by offering tempting but forbidden tools and proving they are absent or denied.
Model outcomes, ownership, and evidence before implementation
Start with one user journey and a state table. For each state, name accepted events, validation, effect, next state, visible meaning, owner, evidence, timeout, cancellation, and cleanup. Outcomes should distinguish completed, rejected, failed, cancelled, exhausted, awaiting approval, and uncertain. “The model answered” is not completion if a tool failed or a citation is unsupported. “The request timed out” is not known failure if an external write may have committed. Stable operation, attempt, resource, actor, policy, tool, model, prompt, corpus, and release identities allow operators to reconstruct what happened without storing every private payload.
Define success at the product boundary. A helpful answer must be supported by authorized evidence; a plan must contain allowed actions and completion criteria; a draft must exist under the expected identity; an effect must be confirmed by its authoritative system. Retain safe structured evidence such as transition, outcome class, latency, budget use, source IDs, command type, policy decision, and resource version. Exclude secrets, raw private documents, and unnecessary conversations. Assign an owner for model behavior, each tool, policy, data, evaluation, operations, and user appeal so failures do not fall into an undefined gap between teams.
Prove that agent flexibility earns its operational cost
Build a deterministic baseline and an agent candidate against the same versioned cases. Include ordinary requests, ambiguous goals, insufficient evidence, forbidden actions, malformed outputs, slow tools, lost acknowledgements, cancellation, hostile content, high load, and account changes. Score task outcome, required and forbidden actions, evidence, policy, latency, cost, accessibility, recovery, and operator burden separately. A higher subjective answer rating cannot compensate for unauthorized calls or unverifiable effects. Preserve hidden and adversarial cases so the system does not optimize only for examples shown during development.
Write an autonomy decision record containing outcome, alternatives considered, why flexibility helps, granted and prohibited capabilities, state and trust diagram, limits, evaluation evidence, approval points, privacy, fallback, rollback, incident owner, and retirement trigger. Review it after real usage. If most successful paths become predictable, move them into deterministic workflows. If a capability is rarely needed but high risk, remove it or place it behind explicit approval. An agent is justified only while measured value exceeds the combined uncertainty, cost, security, governance, and recovery burden.