Amazon AWS AI Practitioner AIF-C01 Foundation Model And Application Evaluation Practice Test

 

AIF-C01 skills 3.4 | 26 original questions

This AWS Certified AI Practitioner AIF-C01 practice test focuses on foundation model and application evaluation through original scenario-based questions aligned to AWS Exam Guide version 1.1 published April 30, 2026. Use the full ExamSnap AIF-C01 collection for broader practice across all five current exam domains. For broader exam preparation, review the Amazon AWS Certified AI Practitioner AIF-C01 Exam Dumps page.

Instructions: Select the best answer for each question. Review the rationale after answering. Each distractor includes a brief explanation of why it is not the strongest fit for the stated scenario.

Question 1

  1. Datum Research has completed discovery for a contact-center transformation. Before implementation, the operations manager must decide how to evaluate on real task examples that reflect production inputs and acceptance criteria. Which choice best satisfies that requirement? Assume the required AWS capabilities are available in the selected Region and normal governance controls are in place. The pilot uses 699 representative records from 5 approved data sources.
  2. Benchmark dataset
  3. Application-specific test set
  4. Failure-mode evaluation
  5. Agent trajectory evaluation
  6. RAG retrieval evaluation

Correct answer: B

Why: Generic benchmarks should be supplemented with workload-specific evaluation. It directly addresses the requirement in this scenario.

Option review:

A: Benchmarks enable repeatable comparisons when they reflect the actual application needs. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

B: Generic benchmarks should be supplemented with workload-specific evaluation. It directly addresses the requirement in this scenario.

C: Robust applications must be evaluated on more than ordinary happy-path requests. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

D: Agent quality depends on decisions and actions across the workflow, not only the final text. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

E: RAG systems should evaluate retrieval quality separately from generation quality. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

Learning point: Application-specific test set – Generic benchmarks should be supplemented with workload-specific evaluation.

Question 2

While planning a operations automation program, Coho Winery identifies this requirement: use a suitably controlled model to score generated responses against a rubric or comparison task. Which option should the AI product manager prioritize if the goal is to control recurring cost? The review committee wants a direct mapping from the requirement to the chosen capability. The first release supports 2 departments and is reviewed every 736 days.

  1. Return on investment (ROI)
  2. Cost per user
  3. LLM-as-a-judge
  4. ROUGE
  5. BLEU

Correct answer: C

Why: LLM-as-a-judge can scale qualitative evaluation but should itself be validated for bias and reliability. It directly addresses the requirement in this scenario.

Option review:

A: ROI evaluates whether the economic return justifies investment. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

B: Cost per user is a business/operational metric rather than a model-quality metric. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

C: LLM-as-a-judge can scale qualitative evaluation but should itself be validated for bias and reliability. It directly addresses the requirement in this scenario.

D: ROUGE emphasizes recall-oriented overlap with reference text and is frequently used for summarization evaluation. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

E: BLEU is traditionally used for machine-translation evaluation. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

Learning point: LLM-as-a-judge – LLM-as-a-judge can scale qualitative evaluation but should itself be validated for bias and reliability.

Question 3

A proof of concept at Lucerne Retail exposed a design decision for the operations manager: the solution must measure how often the AI workflow still requires manual intervention when the objective is safe automation. Which option most directly solves that problem? The solution will serve multiple internal teams, so the recommendation should be reusable without changing the core requirement. The service has a 773-millisecond internal response target for the affected workflow.

  1. Resolution time
  2. Task completion rate
  3. Productivity gain
  4. Escalation rate
  5. Task success rate

Correct answer: D

Why: Escalation rate helps determine whether the system is producing practical operational value. It directly addresses the requirement in this scenario.

Option review:

A: Resolution time can reveal operational efficiency gains or bottlenecks. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

B: Task completion directly reflects whether the AI application accomplishes the target job. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

C: Productivity metrics connect model output to actual work outcomes. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

D: Escalation rate helps determine whether the system is producing practical operational value. It directly addresses the requirement in this scenario.

E: Task success is often more meaningful than a generic language-model metric. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

Learning point: Escalation rate – Escalation rate helps determine whether the system is producing practical operational value.

Question 4

Tailspin Toys is documenting the target state for a internal search upgrade. The AI product manager needs a solution that can include adversarial, ambiguous, missing-data, and unsafe-input cases in the test set. Which option is the strongest fit? The decision must follow the workload characteristics rather than a preference for the largest model or newest service. The team is comparing 4 candidate designs after a 810-day proof of concept.

  1. Agent trajectory evaluation
  2. RAG retrieval evaluation
  3. Application-specific test set
  4. Benchmark dataset
  5. Failure-mode evaluation

Correct answer: E

Why: Robust applications must be evaluated on more than ordinary happy-path requests. It directly addresses the requirement in this scenario.

Option review:

A: Agent quality depends on decisions and actions across the workflow, not only the final text. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

B: RAG systems should evaluate retrieval quality separately from generation quality. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

C: Generic benchmarks should be supplemented with workload-specific evaluation. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

D: Benchmarks enable repeatable comparisons when they reflect the actual application needs. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

E: Robust applications must be evaluated on more than ordinary happy-path requests. It directly addresses the requirement in this scenario.

Learning point: Failure-mode evaluation – Robust applications must be evaluated on more than ordinary happy-path requests.

Question 5

City Power and Light is reviewing a document-intelligence project. The operations manager has one primary requirement: measure how long the AI-assisted workflow takes to resolve the target task. Which choice best fits the requirement? The security baseline is already defined; the decision here concerns the specific capability described in the requirement. The control owner requires evidence from 9 test groups before the 847-day release review.

  1. Resolution time
  2. User satisfaction
  3. Task completion rate
  4. Escalation rate
  5. Task success rate

Correct answer: A

Why: Resolution time can reveal operational efficiency gains or bottlenecks. It directly addresses the requirement in this scenario.

Option review:

A: Resolution time can reveal operational efficiency gains or bottlenecks. It directly addresses the requirement in this scenario.

B: User satisfaction captures perceived value that technical metrics may miss. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

C: Task completion directly reflects whether the AI application accomplishes the target job. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

D: Escalation rate helps determine whether the system is producing practical operational value. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

E: Task success is often more meaningful than a generic language-model metric. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

Learning point: Resolution time – Resolution time can reveal operational efficiency gains or bottlenecks.

Question 6

During a design review for Consolidated Messenger, the AI product manager must compare models on a consistent representative test set. The team also wants to use current managed AWS capabilities. What should the team choose? The recommendation must solve the stated requirement without introducing unrelated platform complexity. The project has 6 downstream consumers and a monthly review of approximately 884 sampled interactions.

  1. Amazon Bedrock Model Evaluation
  2. Benchmark dataset
  3. RAG retrieval evaluation
  4. Agent trajectory evaluation
  5. Workflow end-to-end evaluation

Correct answer: B

Why: Benchmarks enable repeatable comparisons when they reflect the actual application needs. It directly addresses the requirement in this scenario.

Option review:

A: Bedrock Model Evaluation helps compare model performance using configured evaluation approaches. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

B: Benchmarks enable repeatable comparisons when they reflect the actual application needs. It directly addresses the requirement in this scenario.

C: RAG systems should evaluate retrieval quality separately from generation quality. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

D: Agent quality depends on decisions and actions across the workflow, not only the final text. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

E: End-to-end tests catch failures that isolated model benchmarks miss. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

Learning point: Benchmark dataset – Benchmarks enable repeatable comparisons when they reflect the actual application needs.

Question 7

Nod Publishers is moving a claims-processing redesign from pilot to production. The key decision is how to compare generated text with reference translations using n-gram precision-oriented measures. Which option is the strongest fit if the team wants to limit exposure of sensitive data? The design must remain supportable after launch, but no additional feature is required beyond the stated need. The rollout spans 3 application teams, each using the same approved requirement set for the next 921 days.

  1. ROUGE
  2. F1 score
  3. BLEU
  4. BERTScore
  5. Precision

Correct answer: C

Why: BLEU is traditionally used for machine-translation evaluation. It directly addresses the requirement in this scenario.

Option review:

A: ROUGE emphasizes recall-oriented overlap with reference text and is frequently used for summarization evaluation. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

B: F1 summarizes precision and recall when both types of error matter. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

C: BLEU is traditionally used for machine-translation evaluation. It directly addresses the requirement in this scenario.

D: BERTScore captures semantic similarity beyond exact token overlap. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

E: Precision is useful when false positives are especially costly. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

Learning point: BLEU – BLEU is traditionally used for machine-translation evaluation.

Question 8

A workshop at Fabrikam Health focuses on a single decision: how to measure whether intended users adopt and continue using the AI experience. Which option should the AI product manager recommend? A short pilot window means the team prefers an approach that can be evaluated with clear success criteria. The evaluation set contains examples from 8 business workflows and 958 recent production cases.

  1. User satisfaction
  2. Productivity gain
  3. Task success rate
  4. User engagement
  5. Resolution time

Correct answer: D

Why: Engagement can indicate utility but should be interpreted alongside quality and safety. It directly addresses the requirement in this scenario.

Option review:

A: User satisfaction captures perceived value that technical metrics may miss. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

B: Productivity metrics connect model output to actual work outcomes. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

C: Task success is often more meaningful than a generic language-model metric. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

D: Engagement can indicate utility but should be interpreted alongside quality and safety. It directly addresses the requirement in this scenario.

E: Resolution time can reveal operational efficiency gains or bottlenecks. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

Learning point: User engagement – Engagement can indicate utility but should be interpreted alongside quality and safety.

Question 9

For the developer-productivity pilot at Wingtip Logistics, stakeholders need to assess whether an agent selects appropriate tools, follows policy, and completes required steps. Which concept, service, or technique most directly addresses this goal? The architecture board will reject a choice that addresses a different problem from the one described. The initial rollout covers 995 internal users across 5 business units.

  1. Workflow end-to-end evaluation
  2. Amazon Bedrock Model Evaluation
  3. Failure-mode evaluation
  4. Application-specific test set
  5. Agent trajectory evaluation

Correct answer: E

Why: Agent quality depends on decisions and actions across the workflow, not only the final text. It directly addresses the requirement in this scenario.

Option review:

A: End-to-end tests catch failures that isolated model benchmarks miss. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

B: Bedrock Model Evaluation helps compare model performance using configured evaluation approaches. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

C: Robust applications must be evaluated on more than ordinary happy-path requests. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

D: Generic benchmarks should be supplemented with workload-specific evaluation. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

E: Agent quality depends on decisions and actions across the workflow, not only the final text. It directly addresses the requirement in this scenario.

Learning point: Agent trajectory evaluation – Agent quality depends on decisions and actions across the workflow, not only the final text.

Question 10

Trey Research is comparing alternatives for its fraud-review pilot. The AI product manager needs to collect ratings or feedback on usefulness, trust, and experience. Which option is most appropriate while trying to keep the design easy to explain? Budget has been approved for the project, but the team still wants to avoid unnecessary recurring consumption. The workload processes about 72 requests during its busiest hour and has a documented fallback path.

  1. User satisfaction
  2. Productivity gain
  3. Task success rate
  4. Cost per interaction
  5. Escalation rate

Correct answer: A

Why: User satisfaction captures perceived value that technical metrics may miss. It directly addresses the requirement in this scenario.

Option review:

A: User satisfaction captures perceived value that technical metrics may miss. It directly addresses the requirement in this scenario.

B: Productivity metrics connect model output to actual work outcomes. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

C: Task success is often more meaningful than a generic language-model metric. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

D: Cost per interaction connects architecture choices to scalable unit economics. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

E: Escalation rate helps determine whether the system is producing practical operational value. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

Learning point: User satisfaction – User satisfaction captures perceived value that technical metrics may miss.

Question 11

An architecture review at Bellows College has narrowed a analytics modernization decision to one requirement: use managed Bedrock capabilities to evaluate supported models with automatic or human workflows. What should the operations manager select? The team will validate the result with representative production examples before rollout. The pilot uses 109 representative records from 7 approved data sources.

  1. Benchmark dataset
  2. Amazon Bedrock Model Evaluation
  3. Human-in-the-loop evaluation
  4. Failure-mode evaluation
  5. RAG retrieval evaluation

Correct answer: B

Why: Bedrock Model Evaluation helps compare model performance using configured evaluation approaches. It directly addresses the requirement in this scenario.

Option review:

A: Benchmarks enable repeatable comparisons when they reflect the actual application needs. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

B: Bedrock Model Evaluation helps compare model performance using configured evaluation approaches. It directly addresses the requirement in this scenario.

C: Human evaluation is valuable for usefulness, safety, nuance, and domain correctness. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

D: Robust applications must be evaluated on more than ordinary happy-path requests. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

E: RAG systems should evaluate retrieval quality separately from generation quality. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

Learning point: Amazon Bedrock Model Evaluation – Bedrock Model Evaluation helps compare model performance using configured evaluation approaches.

Question 12

The AI product manager at Blue Yonder Airlines is preparing a recommendation for a compliance-assistant prototype. The recommendation must compare generated and reference text using contextual embedding similarity. Which choice is the best match? The pilot has representative data, and the team will measure the selected approach against an agreed acceptance threshold. The first release supports 4 departments and is reviewed every 146 days.

  1. Recall
  2. Precision
  3. BERTScore
  4. F1 score
  5. Accuracy

Correct answer: C

Why: BERTScore captures semantic similarity beyond exact token overlap. It directly addresses the requirement in this scenario.

Option review:

A: Recall is useful when false negatives are especially costly. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

B: Precision is useful when false positives are especially costly. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

C: BERTScore captures semantic similarity beyond exact token overlap. It directly addresses the requirement in this scenario.

D: F1 summarizes precision and recall when both types of error matter. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

E: Accuracy is total correct predictions divided by all predictions. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

Learning point: BERTScore – BERTScore captures semantic similarity beyond exact token overlap.

Question 13

Woodgrove Bank has completed discovery for a forecasting initiative. Before implementation, the operations manager must decide how to measure whether the model-assisted workflow actually completes the intended business task correctly. Which choice best satisfies that requirement? The team will document the rationale for auditors and wants the recommendation to be defensible from the scenario facts. The service has a 183-millisecond internal response target for the affected workflow.

  1. User satisfaction
  2. Resolution time
  3. Cost per interaction
  4. Task success rate
  5. Productivity gain

Correct answer: D

Why: Task success is often more meaningful than a generic language-model metric. It directly addresses the requirement in this scenario.

Option review:

A: User satisfaction captures perceived value that technical metrics may miss. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

B: Resolution time can reveal operational efficiency gains or bottlenecks. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

C: Cost per interaction connects architecture choices to scalable unit economics. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

D: Task success is often more meaningful than a generic language-model metric. It directly addresses the requirement in this scenario.

E: Productivity metrics connect model output to actual work outcomes. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

Learning point: Task success rate – Task success is often more meaningful than a generic language-model metric.

Question 14

While planning a customer-support modernization, Wide World Importers identifies this requirement: test the complete application including retrieval, model calls, guardrails, business logic, and external integrations. Which option should the AI product manager prioritize if the goal is to control recurring cost? The team wants the least complex technically correct choice that satisfies the requirement. The team is comparing 6 candidate designs after a 220-day proof of concept.

  1. Agent trajectory evaluation
  2. Application-specific test set
  3. RAG retrieval evaluation
  4. Failure-mode evaluation
  5. Workflow end-to-end evaluation

Correct answer: E

Why: End-to-end tests catch failures that isolated model benchmarks miss. It directly addresses the requirement in this scenario.

Option review:

A: Agent quality depends on decisions and actions across the workflow, not only the final text. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

B: Generic benchmarks should be supplemented with workload-specific evaluation. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

C: RAG systems should evaluate retrieval quality separately from generation quality. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

D: Robust applications must be evaluated on more than ordinary happy-path requests. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

E: End-to-end tests catch failures that isolated model benchmarks miss. It directly addresses the requirement in this scenario.

Learning point: Workflow end-to-end evaluation – End-to-end tests catch failures that isolated model benchmarks miss.

Question 15

A proof of concept at VanArsdel Media exposed a design decision for the operations manager: the solution must measure average model and application cost for each completed user interaction. Which option most directly solves that problem? The workload has passed basic feasibility checks, so the remaining question is which approach best matches the requirement. The control owner requires evidence from 3 test groups before the 257-day release review.

  1. Cost per interaction
  2. Productivity gain
  3. Task completion rate
  4. Escalation rate
  5. Task success rate

Correct answer: A

Why: Cost per interaction connects architecture choices to scalable unit economics. It directly addresses the requirement in this scenario.

Option review:

A: Cost per interaction connects architecture choices to scalable unit economics. It directly addresses the requirement in this scenario.

B: Productivity metrics connect model output to actual work outcomes. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

C: Task completion directly reflects whether the AI application accomplishes the target job. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

D: Escalation rate helps determine whether the system is producing practical operational value. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

E: Task success is often more meaningful than a generic language-model metric. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

Learning point: Cost per interaction – Cost per interaction connects architecture choices to scalable unit economics.

Question 16

Datum Dynamics is documenting the target state for a contact-center transformation. The AI product manager needs a solution that can have qualified reviewers judge qualities that are difficult to capture with a single automatic metric. Which option is the strongest fit? Stakeholders have ruled out a broad redesign and want the choice that most precisely addresses the stated need. The project has 8 downstream consumers and a monthly review of approximately 294 sampled interactions.

  1. Failure-mode evaluation
  2. Human-in-the-loop evaluation
  3. Amazon Bedrock Model Evaluation
  4. Agent trajectory evaluation
  5. Benchmark dataset

Correct answer: B

Why: Human evaluation is valuable for usefulness, safety, nuance, and domain correctness. It directly addresses the requirement in this scenario.

Option review:

A: Robust applications must be evaluated on more than ordinary happy-path requests. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

B: Human evaluation is valuable for usefulness, safety, nuance, and domain correctness. It directly addresses the requirement in this scenario.

C: Bedrock Model Evaluation helps compare model performance using configured evaluation approaches. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

D: Agent quality depends on decisions and actions across the workflow, not only the final text. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

E: Benchmarks enable repeatable comparisons when they reflect the actual application needs. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

Learning point: Human-in-the-loop evaluation – Human evaluation is valuable for usefulness, safety, nuance, and domain correctness.

Question 17

Alpine Ski House is reviewing a operations automation program. The operations manager has one primary requirement: score overlap between generated and reference text, commonly for summarization. Which choice best fits the requirement? Operational ownership is already assigned, so the team is comparing technical fit rather than staffing models. The rollout spans 5 application teams, each using the same approved requirement set for the next 331 days.

  1. Return on investment (ROI)
  2. Cost per user
  3. ROUGE
  4. F1 score
  5. BLEU

Correct answer: C

Why: ROUGE emphasizes recall-oriented overlap with reference text and is frequently used for summarization evaluation. It directly addresses the requirement in this scenario.

Option review:

A: ROI evaluates whether the economic return justifies investment. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

B: Cost per user is a business/operational metric rather than a model-quality metric. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

C: ROUGE emphasizes recall-oriented overlap with reference text and is frequently used for summarization evaluation. It directly addresses the requirement in this scenario.

D: F1 summarizes precision and recall when both types of error matter. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

E: BLEU is traditionally used for machine-translation evaluation. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

Learning point: ROUGE – ROUGE emphasizes recall-oriented overlap with reference text and is frequently used for summarization evaluation.

Question 18

During a design review for Humongous Insurance, the AI product manager must measure whether users complete target work faster or with less effort. The team also wants to use current managed AWS capabilities. What should the team choose? Existing application interfaces can accommodate any of the listed choices, so functional fit is the deciding factor. The evaluation set contains examples from 2 business workflows and 368 recent production cases.

  1. Escalation rate
  2. User satisfaction
  3. Resolution time
  4. Productivity gain
  5. Task success rate

Correct answer: D

Why: Productivity metrics connect model output to actual work outcomes. It directly addresses the requirement in this scenario.

Option review:

A: Escalation rate helps determine whether the system is producing practical operational value. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

B: User satisfaction captures perceived value that technical metrics may miss. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

C: Resolution time can reveal operational efficiency gains or bottlenecks. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

D: Productivity metrics connect model output to actual work outcomes. It directly addresses the requirement in this scenario.

E: Task success is often more meaningful than a generic language-model metric. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

Learning point: Productivity gain – Productivity metrics connect model output to actual work outcomes.

Question 19

Graphic Design Institute is moving a internal search upgrade from pilot to production. The key decision is how to measure whether retrieval returns relevant, grounded source passages before judging the final answer. Which option is the strongest fit if the team wants to limit exposure of sensitive data? Assume the required AWS capabilities are available in the selected Region and normal governance controls are in place. The initial rollout covers 405 internal users across 7 business units.

  1. Application-specific test set
  2. Benchmark dataset
  3. Amazon Bedrock Model Evaluation
  4. Human-in-the-loop evaluation
  5. RAG retrieval evaluation

Correct answer: E

Why: RAG systems should evaluate retrieval quality separately from generation quality. It directly addresses the requirement in this scenario.

Option review:

A: Generic benchmarks should be supplemented with workload-specific evaluation. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

B: Benchmarks enable repeatable comparisons when they reflect the actual application needs. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

C: Bedrock Model Evaluation helps compare model performance using configured evaluation approaches. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

D: Human evaluation is valuable for usefulness, safety, nuance, and domain correctness. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

E: RAG systems should evaluate retrieval quality separately from generation quality. It directly addresses the requirement in this scenario.

Learning point: RAG retrieval evaluation – RAG systems should evaluate retrieval quality separately from generation quality.

Question 20

A workshop at Relecloud focuses on a single decision: how to measure the percentage of intended workflows that finish successfully. Which option should the AI product manager recommend? The review committee wants a direct mapping from the requirement to the chosen capability. The workload processes about 442 requests during its busiest hour and has a documented fallback path.

  1. Task completion rate
  2. Escalation rate
  3. Task success rate
  4. Cost per interaction
  5. User engagement

Correct answer: A

Why: Task completion directly reflects whether the AI application accomplishes the target job. It directly addresses the requirement in this scenario.

Option review:

A: Task completion directly reflects whether the AI application accomplishes the target job. It directly addresses the requirement in this scenario.

B: Escalation rate helps determine whether the system is producing practical operational value. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

C: Task success is often more meaningful than a generic language-model metric. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

D: Cost per interaction connects architecture choices to scalable unit economics. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

E: Engagement can indicate utility but should be interpreted alongside quality and safety. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

Learning point: Task completion rate – Task completion directly reflects whether the AI application accomplishes the target job.

Question 21

For the knowledge-assistant rollout at Adventure Works Manufacturing, stakeholders need to evaluate on real task examples that reflect production inputs and acceptance criteria. Which concept, service, or technique most directly addresses this goal? The solution will serve multiple internal teams, so the recommendation should be reusable without changing the core requirement. The pilot uses 479 representative records from 9 approved data sources.

  1. RAG retrieval evaluation
  2. Application-specific test set
  3. Workflow end-to-end evaluation
  4. Human-in-the-loop evaluation
  5. Benchmark dataset

Correct answer: B

Why: Generic benchmarks should be supplemented with workload-specific evaluation. It directly addresses the requirement in this scenario.

Option review:

A: RAG systems should evaluate retrieval quality separately from generation quality. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

B: Generic benchmarks should be supplemented with workload-specific evaluation. It directly addresses the requirement in this scenario.

C: End-to-end tests catch failures that isolated model benchmarks miss. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

D: Human evaluation is valuable for usefulness, safety, nuance, and domain correctness. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

E: Benchmarks enable repeatable comparisons when they reflect the actual application needs. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

Learning point: Application-specific test set – Generic benchmarks should be supplemented with workload-specific evaluation.

Question 22

Proseware Services is comparing alternatives for its claims-processing redesign. The AI product manager needs to use a suitably controlled model to score generated responses against a rubric or comparison task. Which option is most appropriate while trying to keep the design easy to explain? The decision must follow the workload characteristics rather than a preference for the largest model or newest service. The first release supports 6 departments and is reviewed every 516 days.

  1. Accuracy
  2. Cost per user
  3. LLM-as-a-judge
  4. Precision
  5. BLEU

Correct answer: C

Why: LLM-as-a-judge can scale qualitative evaluation but should itself be validated for bias and reliability. It directly addresses the requirement in this scenario.

Option review:

A: Accuracy is total correct predictions divided by all predictions. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

B: Cost per user is a business/operational metric rather than a model-quality metric. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

C: LLM-as-a-judge can scale qualitative evaluation but should itself be validated for bias and reliability. It directly addresses the requirement in this scenario.

D: Precision is useful when false positives are especially costly. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

E: BLEU is traditionally used for machine-translation evaluation. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

Learning point: LLM-as-a-judge – LLM-as-a-judge can scale qualitative evaluation but should itself be validated for bias and reliability.

Question 23

An architecture review at Lucerne Publishing has narrowed a personalization program decision to one requirement: measure how often the AI workflow still requires manual intervention when the objective is safe automation. What should the operations manager select? The security baseline is already defined; the decision here concerns the specific capability described in the requirement. The service has a 553-millisecond internal response target for the affected workflow.

  1. Cost per interaction
  2. Task completion rate
  3. User engagement
  4. Escalation rate
  5. User satisfaction

Correct answer: D

Why: Escalation rate helps determine whether the system is producing practical operational value. It directly addresses the requirement in this scenario.

Option review:

A: Cost per interaction connects architecture choices to scalable unit economics. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

B: Task completion directly reflects whether the AI application accomplishes the target job. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

C: Engagement can indicate utility but should be interpreted alongside quality and safety. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

D: Escalation rate helps determine whether the system is producing practical operational value. It directly addresses the requirement in this scenario.

E: User satisfaction captures perceived value that technical metrics may miss. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

Learning point: Escalation rate – Escalation rate helps determine whether the system is producing practical operational value.

Question 24

The AI product manager at Lamna Healthcare is preparing a recommendation for a developer-productivity pilot. The recommendation must include adversarial, ambiguous, missing-data, and unsafe-input cases in the test set. Which choice is the best match? The recommendation must solve the stated requirement without introducing unrelated platform complexity. The team is comparing 8 candidate designs after a 590-day proof of concept.

  1. Benchmark dataset
  2. RAG retrieval evaluation
  3. Workflow end-to-end evaluation
  4. Human-in-the-loop evaluation
  5. Failure-mode evaluation

Correct answer: E

Why: Robust applications must be evaluated on more than ordinary happy-path requests. It directly addresses the requirement in this scenario.

Option review:

A: Benchmarks enable repeatable comparisons when they reflect the actual application needs. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

B: RAG systems should evaluate retrieval quality separately from generation quality. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

C: End-to-end tests catch failures that isolated model benchmarks miss. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

D: Human evaluation is valuable for usefulness, safety, nuance, and domain correctness. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

E: Robust applications must be evaluated on more than ordinary happy-path requests. It directly addresses the requirement in this scenario.

Learning point: Failure-mode evaluation – Robust applications must be evaluated on more than ordinary happy-path requests.

Question 25

Contoso Retail has completed discovery for a fraud-review pilot. Before implementation, the operations manager must decide how to measure how long the AI-assisted workflow takes to resolve the target task. Which choice best satisfies that requirement? The design must remain supportable after launch, but no additional feature is required beyond the stated need. The control owner requires evidence from 5 test groups before the 627-day release review.

  1. Resolution time
  2. Task success rate
  3. User engagement
  4. Escalation rate
  5. User satisfaction

Correct answer: A

Why: Resolution time can reveal operational efficiency gains or bottlenecks. It directly addresses the requirement in this scenario.

Option review:

A: Resolution time can reveal operational efficiency gains or bottlenecks. It directly addresses the requirement in this scenario.

B: Task success is often more meaningful than a generic language-model metric. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

C: Engagement can indicate utility but should be interpreted alongside quality and safety. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

D: Escalation rate helps determine whether the system is producing practical operational value. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

E: User satisfaction captures perceived value that technical metrics may miss. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

Learning point: Resolution time – Resolution time can reveal operational efficiency gains or bottlenecks.

Question 26

While planning a analytics modernization, Fourth Coffee identifies this requirement: use managed Bedrock capabilities to evaluate supported models with automatic or human workflows. Which option should the AI product manager prioritize if the goal is to control recurring cost? A short pilot window means the team prefers an approach that can be evaluated with clear success criteria. The project has 2 downstream consumers and a monthly review of approximately 664 sampled interactions.

  1. Workflow end-to-end evaluation
  2. Amazon Bedrock Model Evaluation
  3. Benchmark dataset
  4. Agent trajectory evaluation
  5. Failure-mode evaluation

Correct answer: B

Why: Bedrock Model Evaluation helps compare model performance using configured evaluation approaches. It directly addresses the requirement in this scenario.

Option review:

A: End-to-end tests catch failures that isolated model benchmarks miss. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

B: Bedrock Model Evaluation helps compare model performance using configured evaluation approaches. It directly addresses the requirement in this scenario.

C: Benchmarks enable repeatable comparisons when they reflect the actual application needs. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

D: Agent quality depends on decisions and actions across the workflow, not only the final text. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

E: Robust applications must be evaluated on more than ordinary happy-path requests. This can be appropriate in another scenario, but it does not most directly satisfy the requirement described here.

Learning point: Amazon Bedrock Model Evaluation – Bedrock Model Evaluation helps compare model performance using configured evaluation approaches.

Popular posts

img