ISTQB CT-GenAI: Prompting, Evaluation, and Responsible Test Use
ISTQB CT-GenAI addresses a problem that many testing teams are already facing: generative AI can accelerate analysis, design, automation, and reporting, but useful output is not the same as trustworthy output. The current specialist syllabus is version 1.1 and the examination uses 40 questions, 46 available points, a passing score of 30 points, and 60 minutes of standard testing time. Foundation Level certification is the prerequisite.
The qualification is deliberately different from a general course about artificial intelligence. Its focus is how testers and quality engineers use large language models and related generative systems inside the test process. That includes prompting, evaluating and refining generated material, understanding hallucination and bias risks, protecting sensitive information, selecting architectural patterns for LLM-powered testing solutions, and planning adoption at organizational scale.
Candidates should therefore connect the syllabus to real test work. A prompt that generates twenty test cases is not automatically productive if the cases repeat one another, miss the risky behavior, invent requirements, or expose confidential data. The strongest preparation treats generative AI as a powerful but fallible collaborator whose outputs require objectives, context, evaluation criteria, and human accountability.
Generative AI can reduce the time required to produce first drafts of test ideas, data variations, scripts, summaries, or defect descriptions. That speed is attractive because testing often contains language-heavy tasks that consume expert attention. The more important question is what happens after the draft appears. A fast output that must be extensively corrected may move effort rather than remove it, and an apparently polished answer can make weak reasoning harder to notice.
The generative AI fundamentals help explain this behavior. Large language models generate responses from learned statistical patterns and supplied context; they do not possess a guaranteed, authoritative model of the product being tested. Candidates should understand tokens, context, inference, grounding, and the difference between model capability and the application controls wrapped around the model.
A useful prompt establishes the task, relevant context, constraints, expected format, and sometimes examples. For testing work, that may include a requirement, risk statement, user persona, interface contract, test technique, or list of prohibited assumptions. Prompt design becomes stronger when the tester can explain why each element is present. Adding detail without purpose can increase noise just as easily as it can improve an answer.
Candidates should practice iterative prompting rather than searching for one magical instruction. A first response can reveal missing context or an ambiguity in the task. The next prompt can constrain scope, request alternatives, require traceability to source information, or ask the model to critique its own draft. This refinement process resembles exploratory testing: the tester learns from each interaction and uses the observation to design the next one.
Reusable prompt patterns can improve consistency, but templates should not become unquestioned automation. A prompt that works for API test ideas may be inappropriate for accessibility risks or regulated acceptance criteria. The tester still owns the decision about what evidence is needed and whether the generated output contributes to that evidence.
Structured output can reduce one source of ambiguity. Requiring a schema, fields for assumptions, source references, or explicit uncertainty makes generated material easier to validate and integrate with downstream tools. Structure does not guarantee truth, but it gives automated checks something concrete to inspect and makes missing information visible. Candidates should see this as a quality-control mechanism around the model interaction rather than proof that the model understood the request correctly.
Prompt context should also be treated as controlled test input. Requirements, examples, retrieved documents, tool outputs, and system instructions can conflict or contain untrusted text. A robust workflow establishes precedence and limits what retrieved content is allowed to instruct the system to do. This becomes especially important when a testing assistant can call tools, modify artifacts, or take actions beyond producing text.
Generative output creates an oracle problem because several responses can be acceptable while none is guaranteed to be correct. Evaluation therefore needs dimensions such as factual accuracy, relevance, completeness, diversity, traceability, safety, or compliance with a required structure. The dimensions should reflect the testing task. A generated defect summary may prioritize fidelity to observed evidence, while generated exploratory charters may prioritize useful diversity and risk coverage.
Human evaluation can be valuable when judgment is inherently contextual, but human review needs criteria as well. Without a rubric, two reviewers may reward different qualities and create an unstable benchmark. Automated checks can support the evaluation by verifying schemas, required references, prohibited phrases, duplicate content, or deterministic calculations. A robust evaluation design often combines these methods instead of treating either humans or automation as universally superior.
A model may confidently produce a nonexistent requirement, API parameter, vulnerability, source, or expected result. The practical risk depends on how the organization consumes that output. If a human expert reviews a brainstorming draft, the cost may be small. If generated assertions flow directly into an automated regression suite, false assumptions can become repeatable misinformation and create a misleading sense of coverage.
Techniques for reducing hallucinations include grounding responses in controlled sources, retrieving relevant material, validating claims, constraining output, and designing product workflows that make uncertainty visible. Candidates should understand that mitigation belongs to the complete solution. Prompt wording alone cannot guarantee factual behavior.
Bias, privacy, intellectual-property concerns, and security add other failure modes. Test artifacts may contain personal data, proprietary code, unreleased features, credentials, or customer information. Before sending such material to an external model, teams need explicit data-handling rules. Responsible test use therefore begins with governance rather than with clever prompting.
Organizations can move beyond interactive prompting and build systems that retrieve internal requirements, call tools, generate test assets, execute workflows, and retain state. Each added component creates both capability and a new place where behavior can fail. Retrieval can return irrelevant context, orchestration can call the wrong tool, permissions can expose restricted data, and a generated action can be executed with more authority than intended.
The syllabus treats architectural approaches as part of the testing problem because a quality engineer needs to know where controls belong. Source selection, retrieval quality, prompt construction, model configuration, output validation, human approval, logging, and rollback may all contribute to trust. Evaluating only the final text ignores the mechanisms that produced it.
Evaluation sets need maintenance just like regression suites. If teams tune prompts repeatedly against the same examples, they can optimize for that set while missing new product behavior. A healthier practice preserves representative cases, adds real failures as they are discovered, and periodically introduces unseen challenges. This keeps evaluation connected to the evolving testing problem rather than rewarding memorization of a fixed benchmark.
Human review should have an explicit purpose as well. Reviewers may validate requirements traceability, technical correctness, safety, or usefulness, but asking one person to judge everything informally can create inconsistent approval. Clear review roles, escalation rules, and sampling strategies help organizations decide when human oversight is mandatory and when automated validation is sufficient for a low-risk task.
LLMOps matters because generative behavior can change when the model, system prompt, retrieval corpus, tool set, safety layer, or decoding configuration changes. A test organization should be able to identify which configuration produced a result and compare releases against a representative evaluation set. Without that discipline, teams can observe improvement or regression without knowing what caused it.
Monitoring after deployment is equally important. A prompt library that worked well for one product release may degrade when terminology, requirements, or workflows change. Usage data can reveal where users repeatedly edit outputs, abandon suggestions, or override generated decisions. Those signals help identify where the solution needs better context, a different interaction design, or a narrower scope.
The distinction between this qualification and ISTQB CT-AI is useful for study planning. ISTQB CT-AI focuses on testing AI-based systems: data, models, probabilistic behavior, AI-specific quality characteristics, and the evidence needed to assess those systems. ISTQB CT-GenAI focuses on using generative AI to improve or support testing work. A professional may eventually need both perspectives, but they begin from different objectives.
This difference prevents a common study mistake. Knowing how to test a large language model does not automatically mean a tester knows how to deploy that model responsibly inside a test organization. Conversely, a skilled prompt user may still lack the techniques required to evaluate an AI product. The wider set of ISTQB certifications makes these specializations complementary rather than interchangeable.
Generative AI changes roles, review practices, tooling, data flows, and expectations about speed. A sensible adoption roadmap begins with bounded use cases where value and risk can be measured. Teams can establish baselines, define acceptable output quality, identify prohibited data, assign review responsibility, and decide what must remain human-controlled before expanding to more autonomous workflows.
Change management matters because poor incentives can defeat good technical controls. If a team is rewarded only for producing more tests faster, people may accept generated volume without examining coverage or maintainability. If reviewers do not know which parts were model-generated, accountability becomes blurred. Successful adoption makes the new workflow transparent and measures whether it improves useful outcomes rather than merely increasing artifact count.
Cost and latency also belong in evaluation. A model that produces excellent output may still be a poor fit if every request is expensive, slow, or dependent on context windows that are difficult to maintain. Smaller models, retrieval, caching, constrained workflows, or selective human use may create a better operational result. Candidates should therefore evaluate a generative testing solution as a system with quality attributes and tradeoffs, not merely rank models by the apparent sophistication of one response.
Teams should also define fallback behavior. If the model service is unavailable, returns low-confidence output, violates a guardrail, or cannot access required context, the workflow needs a safe alternative. That can mean stopping an automated action, routing work to a person, using a deterministic rule, or continuing without the optional AI assistance. Resilience matters because an accelerator should not become a single point of failure for the testing process.
Auditability completes that control picture. Teams should be able to tell which model, context, prompt pattern, and reviewer produced an important testing artifact, especially when the output influences release or risk decisions.
A strong study routine uses a generative system as an object of disciplined experimentation. Give it an ambiguous requirement and observe the assumptions it invents. Add controlled context and compare the output. Ask for test cases, then score them against a rubric for risk coverage, duplication, traceability, and feasibility. Change one prompt element at a time so the effect can be understood rather than guessed.
Then connect those observations back to the official learning objectives. Explain why a mitigation works, what residual risk remains, and which organizational control would be needed before operational use. That approach prepares candidates for the examination while building the habit the qualification is designed to encourage: use generative AI where it creates leverage, but make quality claims only from evidence that has actually been evaluated.
