CCA-F: Hands-On Skills to Practice
Claude Certified Architect: Foundations is an applied architecture credential, so hands-on preparation matters more than collecting definitions. The site’s established target is labeled CCA-F, while candidates will encounter CCAR-F in Anthropic’s current Claude Certified Architect: Foundations exam material. The practical skills behind the blueprint matter more than that naming difference: agentic control flow, tool interfaces, MCP integration, Claude Code configuration, prompt design, structured output, context management, and reliability.
A useful lab program should make those behaviors visible. If a framework hides the message loop, tool results, context boundaries, and failure handling, it may help you build an application but teach you less about the decisions the exam tests. Start small enough that you can explain every component, then add complexity only when the exercise requires it.
The exercises below are designed to complement the CCA-F exam blueprint. They are not a substitute for the current exam guide; they are a way to turn its task statements into engineering practice.
Your first lab should expose the basic agentic protocol. Give Claude one or two tools, send a request, inspect the response, execute a requested tool, return the tool result, and continue until the model reaches the correct stopping condition.
Do not begin with a large agent framework. The point of this exercise is to understand what the framework would otherwise do for you. Record the assistant message containing the tool request, the corresponding tool result, and the next model response. Observe how the conversation history changes after each iteration.
Create three cases: one where no tool is needed, one where a tool is called once, and one where the task requires more than one step. Then deliberately return a tool error. Your implementation should make it obvious why the loop continues, why it stops, and how the failure becomes part of the next decision.
AI agents fundamentals—goals, state, tools, and feedback loops—become much clearer when implemented as concrete Claude control flow.
Once one agent works, create a coordinator that delegates two different jobs to specialized subagents. A simple research example works well: one subagent gathers facts, another checks a document set, and the coordinator combines the results.
The important part is not the number of agents. It is context discipline. Subagents should receive the information they actually need, and the coordinator should understand what comes back. Do not assume a subagent automatically inherits every detail from the coordinator’s conversation.
Run the same task with too little context and with too much context. The first version should fail because a required detail is missing. The second may become noisy or inefficient. Then design a concise handoff object containing the task, constraints, source references, and expected output.
Add one failure case in which a subagent cannot complete its task. Decide whether the coordinator should retry, use another path, return a partial result, or escalate. That is closer to the reliability judgment the exam expects than simply demonstrating that parallel agents can run.
Tool selection is difficult to study when every tool has an obvious purpose. Create two intentionally adjacent tools—for example, one that retrieves a customer profile and one that retrieves an order. Give them weak descriptions first and test prompts where the correct choice is not stated explicitly.
Then improve the names, descriptions, parameter definitions, and boundaries. Rerun the same requests. The goal is to see how interface quality affects model behavior.
Next, add permission checks and typed errors. The model should not be able to bypass authorization because a prompt says the action is allowed. A validation error, permission denial, missing record, and transient service failure should be distinguishable.
Reliable tool use and function calling requires schemas, permissions, failures, idempotency, and observability. Your lab should then narrow those principles to the CCA-F decisions around tool boundaries and selection.
For MCP practice, use a small server rather than a huge catalog. Connect it through a Claude workflow and inspect which tools become available, how they are described, and how the agent behaves when the server returns structured data or an error.
Create a test where a tool name is ambiguous and another where authentication or permissions prevent the operation. Your objective is not to learn every MCP transport option. The exam focuses on how MCP capabilities fit into Claude Code and agent workflows, how errors are represented, and how tool availability is scoped.
It is useful to compare direct tool definitions with MCP-exposed tools. Ask what remains the same: descriptions still matter, schemas still matter, tool results still enter context, and the application still needs a sensible response when a capability fails.
Avoid turning this lab into a cloud-hosting project. Provider-specific deployment and MCP server infrastructure are not the center of the foundation exam.
Create or use a repository with several directories and different conventions. Add a project-level CLAUDE.md, then introduce path-specific rules for one portion of the codebase. Use a reusable skill or command for a repeated task.
Now test what happens when instructions live at the wrong scope. A rule that applies only to tests should not crowd every unrelated coding task. A repository-wide standard should not need to be pasted into each prompt. The lab is successful when you can explain why each instruction belongs where you placed it.
Try one task in plan mode and another through direct execution. Choose tasks where the trade-off is meaningful. A broad refactor with architectural consequences may benefit from an explicit plan; a tightly bounded edit may not.
Finally, run a review workflow that would make sense in CI/CD. Give Claude concrete review criteria and inspect false positives. This prepares you for the intersection between Claude Code configuration and prompt design.
Create a review or classification task with a small test set. Write an intentionally vague first prompt, run the cases, and record the failures. Then introduce explicit criteria and compare the result.
Add few-shot examples only after you understand the failure. Pick examples that demonstrate genuinely ambiguous boundaries. Do not use ten easy examples that simply repeat the instruction.
Prompt engineering fundamentals become more useful when the prompt is treated as a component under test. Change one major variable at a time and measure whether the behavior improves.
Include at least one requirement that should not be enforced solely in the prompt. For example, if an action requires authorization, make the tool or application enforce the rule. That reinforces the distinction between model guidance and deterministic controls.
Use several documents containing fields that may be present, absent, or ambiguous. Define a schema and have Claude produce structured results. Validate each result before accepting it.
The most useful cases are imperfect. Include a document where a value is missing. Include one where two possible values appear. Include one with irrelevant text that resembles a field. The system should represent uncertainty or absence rather than fabricate a convenient answer.
Add a validation and retry loop. When a result fails, return precise feedback that allows another attempt. Then compare that behavior with simply asking the model to “please return valid JSON.”
This exercise teaches structured output, schema discipline, validation, retry, and reliability in one workflow. It also demonstrates why downstream applications need stronger guarantees than formatting instructions alone.
Create a workflow that accumulates many tool results or explores a moderately sized codebase. Let the conversation become noisy enough that preserving every detail is no longer desirable. Then decide what must remain visible to the model.
Move stable facts into structured state or a concise working artifact. Keep important constraints explicit. Remove or summarize results that no longer influence the task. Compare the quality of later decisions before and after the cleanup.
The exam is not primarily about memorizing the latest context-window size. It tests whether you understand how to preserve critical information, avoid losing provenance, and keep long-running work reliable.
Also practice resumption. Stop the work, save the state that matters, and continue later. A robust system should not require replaying an entire noisy history simply to remember what has already been established.
Create a support-style scenario with low-risk questions, ambiguous requests, and one action that should require human review. Define the information the agent needs before it acts. When the information is incomplete, the system should ask for clarification or escalate rather than invent a missing fact.
Then deliberately create a conflict between tool results. Decide how the system represents uncertainty and what evidence reaches the human reviewer. This prepares you for the Context and Reliability tasks around escalation, confidence calibration, and provenance.
Human review should not be a vague “someone checks it” step. Define what triggers review, what context the reviewer receives, and how the decision is recorded for the rest of the workflow.
Create a compact evaluation set for the labs you built. A tool-use case should check whether the correct tool and arguments were selected. A structured-output case should verify schema validity and required business rules. A multi-agent task should verify whether the final result preserved evidence from the correct source.
AI evaluation fundamentals such as representative cases and release thresholds turn impressions into evidence. Use that discipline without turning CCA-F preparation into a generic evaluation course.
Keep a short failure log. For each failure, record whether the root cause was prompt wording, tool design, context, workflow enforcement, validation, or escalation. That classification is valuable exam practice because scenario questions often present several technically plausible fixes and ask which one addresses the actual failure.
The Claude ecosystem is broader than the CCA-F blueprint. The current exam guide excludes several topics that may still be useful in real projects, including embedding models and vector databases, vision/image analysis, custom-model training, and provider-specific deployment configuration.
Do not let an interesting RAG project, cloud deployment, or multimodal demo consume the hours you need for agentic loops, tool design, Claude Code, structured output, and context reliability. Those broader topics can support your Claude expertise, but the hands-on program for this exam should follow the actual objectives.
If you can implement the labs above and explain why each architectural choice was made, you are practicing the kind of applied reasoning Claude Certified Architect: Foundations is designed to validate.
