ISTQB CT-AI v2.0: Testing Probabilistic Systems With Evidence

ISTQB CT-AI v2.0 is the current AI Testing certification and replaces the earlier v1.0 syllabus. The older version is still in a formal sunset period, with English exams scheduled to end on April 21, 2027 and non-English exams on October 21, 2027. Candidates preparing now should therefore verify the syllabus version with their exam provider instead of assuming that any material labeled “AI Testing” represents the current learning objectives.

The current ISTQB CT-AI v2.0 exam contains 40 questions worth 44 points, requires 29 points to pass, and allows 60 minutes before any approved non-native-language extension. An ISTQB Foundation Level certificate is a prerequisite. The syllabus now explicitly includes generative AI and large language models while retaining deep attention to machine-learning data, model behavior, functional performance metrics, neural networks, and lifecycle-specific testing activities.

AI testing changes the tester’s relationship with expected results. Many conventional systems are designed to produce deterministic outputs for defined inputs. AI-based systems can be probabilistic, data-dependent, difficult to explain, and sensitive to distribution changes. The central preparation challenge is learning how to build credible evidence when “the same input always produces the same exact output” is no longer a safe assumption.

AI quality starts before model execution

Testing an AI-based system begins with understanding the qualities that matter. Accuracy may be important, but it is rarely sufficient. Robustness, explainability, fairness, privacy, safety, security, and suitability for the intended context can all affect whether a system is acceptable. The relative importance depends on the product. A low-stakes recommendation feature and a model supporting a medical decision should not share the same risk assumptions.

Candidates should practice converting broad quality concerns into observable acceptance conditions. “The model should be fair” is too vague until protected groups, relevant outcomes, measurement methods, and tolerances are defined. “The model should be explainable” needs a stakeholder and purpose: an engineer debugging a model may need different information from a customer challenging a decision. AI quality is contextual evidence, not a single score.

AI-specific quality characteristics expand the test oracle. Functional correctness remains relevant, but qualities such as explainability, robustness, fairness, transparency, and autonomy can affect whether a system is acceptable. These qualities are not automatically measurable with one universal score. Teams need operational definitions that fit the use case, affected population, and risk. A useful acceptance criterion states what evidence is expected rather than merely declaring that the system should be “fair” or “explainable.”

Quality characteristics can also conflict. A more interpretable model may sacrifice some predictive performance; stricter safety controls may reduce coverage of legitimate edge cases; a highly personalized model may create privacy concerns. Testing helps make these tradeoffs visible. Candidates should recognize that maximizing one metric without considering the system objective can produce a locally impressive but globally poor decision.

Data is part of the system under test

Machine-learning behavior is inseparable from data. Training, validation, and test datasets can contain missing values, duplicates, leakage, sampling bias, mislabeled examples, outdated patterns, or representation gaps. These are not merely data-engineering inconveniences; they can directly change model behavior. Testing therefore includes examining whether the data is suitable for the claims the organization wants to make about the model.

A candidate should understand why a model can perform well on a familiar benchmark and poorly in production when the real population differs. Input-data testing asks whether incoming data respects expected structure and semantics, while development testing considers the datasets and processes used to build the model. The goal is to catch quality problems before a performance metric disguises them behind an attractive average.

Data quality also changes over time. A dataset that represented users well last year may become less representative after a product expands into new markets or user behavior changes. That makes monitoring and re-evaluation part of the testing story. Candidates should understand the idea of distribution shift even when the exam question does not use a complex mathematical treatment: evidence collected on one population may weaken when the operating population changes.

Metrics need interpretation, not worship

Classification metrics are a good example of why AI testing requires context. Accuracy can look excellent when one class dominates the data. Precision and recall answer different questions about false positives and false negatives. Threshold changes can improve one outcome while weakening another. A test report that lists metrics without relating them to risk can therefore mislead decision makers.

Candidates should work through small confusion-matrix examples until they can explain what each metric means for the product. Then add cost asymmetry: what if a missed fraud case is much more expensive than an unnecessary review, or a false medical alert creates its own harm? The correct metric and threshold depend on the decision the system supports, not on which number is easiest to maximize.

Data leakage is a particularly dangerous source of misleading confidence. If information from the evaluation set influences training, feature construction, tuning, or selection, measured performance can overstate how the model will behave on genuinely unseen data. Candidates should understand the purpose of separating training, validation, and test evidence and why repeated experimentation against the same holdout set can gradually turn that set into part of the development process.

Representative challenge data should include more than average cases. Rare classes, boundary conditions, missing fields, corrupted inputs, distribution shifts, and subgroups with different error costs can expose weaknesses hidden by aggregate performance. The objective is not to create an impossibly exhaustive dataset but to align challenge sets with the risks that matter for the product and the people affected by its decisions.

Non-determinism changes oracle design

Traditional test cases often rely on a precise expected result. For an AI system, the oracle may instead be a range, distribution, invariant, relationship, or quality threshold. A generative model may produce different valid wording on repeated runs. A recommendation system may reorder acceptable items. A vision model may behave differently after a small but legitimate change in lighting. Test design must account for this variability without accepting arbitrary behavior.

This is where disciplined evidence matters. Repeated execution, statistical analysis, metamorphic relations, curated challenge sets, robustness tests, and human evaluation may all contribute. The correct approach depends on the claim. A single successful example cannot demonstrate reliability across a population, while a broad metric can hide severe failures in a critical subgroup.

Oracle design can also combine automated and human judgment. Automated checks may enforce schema, prohibited content, factual grounding against a trusted source, or consistency constraints. Human reviewers may be needed for usefulness, nuance, or domain-sensitive harm. The key is to define where each type of evaluator is credible. A human rating without criteria can be inconsistent, while a rigid automated score may miss whether an answer is actually helpful.

Generative AI introduces new failure modes

The current syllabus includes testing generative AI and large language models, where hallucination, prompt sensitivity, unsafe output, privacy leakage, bias, and inconsistent reasoning become practical concerns. The AI hallucinations illustrates why testing must examine the whole solution, including retrieval, grounding, validation, and product controls rather than only the base model.

Candidates should also distinguish testing AI-based systems from using generative AI to assist testing. ISTQB CT-AI v2.0 is primarily about the former. The related ISTQB CT-GenAI qualification focuses on applying generative AI within the testing process. The two paths overlap in terminology but answer different professional questions.

AI-specific test levels follow the model lifecycle

The syllabus introduces testing perspectives that map to machine-learning development rather than only conventional software layers. Input data, the model itself, and the ML development process each provide different opportunities to find problems. A defect in feature preparation can invalidate later model evaluation. A model can satisfy aggregate metrics while failing important edge cases. A sound model can still be integrated into a product that uses its output incorrectly.

Preparation should therefore trace an issue through the lifecycle. Where could the problem be introduced? What evidence would reveal it earliest? Which artifact should be examined? How would the test differ before and after deployment? This way of reasoning prevents “AI testing” from collapsing into one final benchmark run.

Testing AI systems also requires attention to repeatability of the test setup. Model versions, prompts, system instructions, retrieval indexes, decoding parameters, data snapshots, and safety layers can all change observed behavior. When those elements are not recorded, a failure may be difficult to reproduce and an improvement may be impossible to attribute. Configuration discipline therefore becomes part of AI test credibility, even when the system itself remains probabilistic.

Robustness testing asks how behavior changes under perturbation. Small input changes, noise, unusual formatting, adversarial manipulation, or environmental variation may reveal that a model is relying on fragile patterns rather than the intended signal. For generative systems, prompt injection and untrusted retrieved content create related concerns because instructions can arrive through data channels the application did not intend to treat as authoritative. Testing should examine these boundaries as part of the complete AI-enabled system.

Risk and governance shape acceptable evidence

AI creates organizational risks in addition to model-performance risks. Data provenance, privacy, security, regulatory obligations, change control, monitoring, and human oversight can all determine whether a system is safe to operate. A candidate does not need to become a lawyer or data scientist, but should recognize that test objectives are driven by these constraints and that different stakeholders may require different evidence.

The broader AI and GenAI is useful for connecting model behavior to application architecture, agents, security, and governance. Testing becomes strongest when it follows those real system boundaries instead of evaluating an isolated model as though it were the whole product.

A final useful distinction is between model capability and product behavior. A powerful model can still support a poor product if prompts, retrieval, orchestration, permissions, or fallback behavior are weak. Conversely, product controls can reduce some model risks without changing the model itself. Exam preparation should repeatedly ask what layer is being tested and whether the proposed evidence supports a claim about that layer or about the whole system.

For each experiment, record enough context to reproduce what happened: the data slice, model or service version, important configuration, prompt or input, evaluator, and acceptance threshold. This habit exposes a central lesson of ISTQB CT-AI v2.0: an observation becomes useful test evidence only when its conditions and interpretation are understood. Reproducibility will never make a probabilistic system deterministic, but it lets a team distinguish meaningful behavioral change from noise, compare releases fairly, and investigate failures without guessing which part of the setup produced them.

Another useful drill is to take one quality claim, such as fairness, robustness, or functional performance, and write the evidence needed to support it. Then identify what that evidence would still fail to prove. This forces candidates to distinguish a metric from a conclusion and a test result from a broader assurance claim. That distinction is central to responsible AI testing because aggregate success can coexist with serious subgroup, edge-case, security, or operational failures.

Prepare with small experiments and explanations

A productive study plan alternates concepts with concrete experiments. Build a tiny classification example and calculate several metrics. Create a dataset with an obvious imbalance and predict how the measures will change. Test a generative prompt repeatedly and record the dimensions on which responses vary. Define an acceptance criterion for a quality such as robustness or fairness, then challenge whether the criterion actually supports the business claim.

Use the ISTQB certifications to keep the path clear: Foundation Level supplies core testing language, ISTQB CT-AI develops the specialist skills for testing AI-based systems, and other qualifications cover adjacent roles. The aim is not to memorize AI vocabulary. It is to reason about evidence when behavior is probabilistic, data-dependent, and capable of failing in ways conventional test oracles may not expose.

  • img