Agent Orchestration and Tool Use for AI-103

Agents are a major part of the AI-103 exam because Microsoft expects Azure AI engineers to build workflows that can reason, retrieve knowledge, call tools, preserve conversational state, and coordinate multiple agents. The challenge is not simply making the model autonomous. It is creating an operating loop that remains predictable enough to secure, test, and troubleshoot.

The general model in AI agents fundamentals is a useful starting point. AI-103 adds the Microsoft Foundry context: agent roles, conversation tracking, knowledge and search tools, custom functions, multi-agent orchestration, safeguards, monitoring, and error analysis.

Define the agent’s role before adding tools

Give the agent a bounded responsibility and a clear finish condition. A customer support agent may retrieve policy, inspect account status, create a case, and escalate exceptions. An engineering agent may diagnose a service, propose a change, and stop for approval before deployment.

Vague roles make tool selection noisy because many actions can appear relevant. Clear responsibility also makes evaluation easier: you can tell whether the agent completed its job without inventing additional work.

Treat conversation tracking as managed state

Decide what needs to persist across turns: user intent, identifiers, decisions, pending approvals, or previous tool results. Keep durable business facts in systems of record rather than relying on the conversation as the only copy.

Long workflows benefit from compact structured state. This reduces context growth and makes it easier to resume or hand off the task without replaying every prior message.

Design tool schemas around business actions

A tool should expose a specific capability with typed inputs and predictable outputs. Broad tools such as ‘manage account’ give the model too much ambiguity. Narrow actions such as ‘get account status’ or ‘create support case’ make selection and authorization easier.

The tool-use and function-calling principles apply directly: the agent can choose a tool, but the execution boundary validates inputs, permissions, and side effects.

Use knowledge and tools for different kinds of facts

Static or slowly changing explanatory material fits a knowledge or retrieval source. Transactional facts such as current balance, inventory, or ticket status should usually come from a live tool or API. Mixing those roles can create stale answers.

AI-103 scenarios often become clearer when you ask whether the agent needs evidence to explain something or authority to perform an action. Retrieval and execution are related but not interchangeable.

Integrate search where grounding matters

Foundry agents can use Azure AI Search or other knowledge sources to ground responses. Preserve source metadata and access controls so the agent can cite or explain which evidence supported the response.

Evaluate retrieval independently from agent reasoning. If the correct document is missing, changing the orchestration prompt will not fix the root problem.

Use custom functions for narrow deterministic operations

Custom functions are appropriate when the application needs a precise operation that is not already exposed through a suitable managed tool. Define the input contract, output contract, error behavior, and authorization independently of the model.

Keep irreversible changes behind explicit checks. A generative decision can determine that a function is relevant, but the function should still enforce business rules.

Orchestrate multi-agent systems by responsibility

Split agents when responsibilities, toolsets, data access, or expertise are meaningfully different. A coordinator can delegate to specialist agents, but the handoff should define what context is passed and what output is expected.

Avoid creating several agents that all appear able to do the same work. Overlapping roles make routing unpredictable and can cause duplicated actions or contradictory results.

Use safeguards for autonomous and semiautonomous flows

AI-103 includes autonomous and semiautonomous workflows with approval controls. Decide which actions can run unattended, which require review, and which should never be available to the agent.

Use consequence and reversibility as the decision criteria. A read-only lookup can often be automatic; a high-impact write may require a human to approve the exact proposed change.

Handle tool errors as structured workflow states

Different failures require different responses. A permission denial is not a timeout. A not-found result can be a valid business outcome. An invalid argument may need correction rather than retry.

Return structured error information and preserve enough context for the agent to choose the next safe action. Avoid generic messages that force the model to guess.

Use parallel tool calls only when dependencies allow them

Independent lookups can run in parallel to reduce latency. Dependent operations should remain ordered. If one step creates the identifier needed by another, parallel execution can introduce race conditions or meaningless failures.

Measure whether parallelism actually improves the complete workflow. Some tool calls may be cheap individually but create heavier downstream work when executed unnecessarily.

Monitor agent behavior, not only final answers

AI-103 explicitly includes monitoring deployed agents and performing error analysis. Track tool choices, failed calls, step count, latency, safety events, token usage, grounding quality, and whether the agent stopped at the right time.

The AI evaluation framework is useful because an agent should be scored on task success and trajectory quality, not merely on fluent output.

Design escalation as part of the happy path

An agent will sometimes lack evidence, authority, or confidence. Define the handoff: what information a human or another system receives, which steps already completed, and what question remains unresolved.

Escalation should preserve state so the next actor can continue without reconstructing the conversation. A good handoff is evidence of reliable orchestration, not a failure of autonomy.

Keep orchestration logic observable

When an agent selects a tool, delegates to another agent, or changes plan, preserve enough metadata to understand why the path changed. You do not need to expose private chain-of-thought; you do need operational events that identify the action, state transition, and result.

This makes multi-step behavior diagnosable. Without traceable state changes, a long agent run can fail far from the original mistake.

Separate planning from execution permissions

An agent may be allowed to consider many possible actions while holding permission to execute only a subset. This is useful because planning can remain flexible while execution stays bounded. The orchestration layer can propose a deployment change, for example, while a separate approval or service identity controls whether the change actually occurs.

That separation is one of the strongest ways to preserve autonomy without giving the model unnecessary authority.

Use handoff contracts between agents

When one agent delegates to another, define the expected input and output instead of forwarding an unstructured conversation. Include the task, relevant evidence, constraints, and required result shape.

A handoff contract reduces context size and makes errors easier to isolate. It also prevents a specialist agent from inheriting unrelated instructions or state from the coordinator.

Test agents with tool unavailability

Disable a dependency or force a timeout during testing. The agent should not invent the missing result or loop indefinitely. It should choose an allowed fallback, escalate, or preserve state for later continuation.

Failure testing is especially important for orchestration because the workflow may otherwise look reliable only while every external service behaves perfectly.

Set a maximum action budget

Limit the number of model turns, tool calls, or state-changing operations one run can perform. Budgets contain loops and make cost more predictable. They also create a natural point where the agent must summarize progress and ask for help rather than continuing indefinitely.

The budget should reflect the workflow: research can tolerate more read operations than a transaction-oriented agent.

Use one final integration pass

After specialist agents or tools complete their work, a coordinator should reconcile outputs and check that the overall goal is actually satisfied. Parallel sub-results can each be correct while still conflicting at the system level.

A final integration step is where duplicated work, inconsistent assumptions, and missing dependencies can be caught before the response or action is finalized.

Design memory around the task horizon

Short conversational context, task state, durable user preferences, and authoritative business data are different kinds of memory. Do not store them all in one transcript. Decide how long each piece should live and which system owns it.

Agent orchestration becomes more reliable when the coordinator retrieves the state it needs instead of carrying every historical detail into every turn.

Use explicit stop conditions for autonomous runs

Define when the agent is done, when it must ask for help, and when a tool failure ends the run. Stopping rules should be based on task state rather than a vague instruction to continue until satisfied.

Bounded autonomy protects cost and prevents loops that keep producing activity without meaningful progress.

Use end-to-end scenarios to study orchestration

Build a small Foundry agent that retrieves one knowledge source, calls one API, handles a failed call, and routes one case for approval. Then explain the identity and data boundary at every step.

That exercise covers more of the exam than memorizing a list of agent features because it forces you to reason about roles, tools, memory, safeguards, monitoring, and recovery as one system.

  • img