Computer Vision and Image Generation for AI-901

The current AI-901 exam no longer treats computer vision as a completely separate world from generative AI. Microsoft’s April 2026 objectives combine classic vision workload recognition with multimodal models that interpret visual prompts and generative models that create new images. Candidates also need to understand how a lightweight Foundry application can use those capabilities.

For the AI-901 exam, the most important distinction is between understanding existing visual content and generating new visual content. A model that identifies what is in an uploaded image is solving an interpretation problem. A model that creates a new illustration from a text prompt is solving a generation problem. A multimodal application can combine those capabilities with text and conversation.

Computer vision turns visual input into useful information

Computer-vision workloads analyze images or video to extract meaning. Common scenarios include classifying an image, detecting or locating objects, reading visible text, describing a scene, or answering questions about visual content. The exact technique depends on the business requirement.

At fundamentals level, identify the workload from the outcome. If the requirement is to determine whether a photo contains a damaged product, the system must interpret the image. If the requirement is to read a serial number from a label, text extraction or content-understanding capabilities may be appropriate. If the requirement is to produce a new marketing image, that is image generation rather than analysis.

AI-901 questions often become easier once the desired output is stated in plain language: label, location, extracted text, description, answer, or newly generated visual.

Multimodal models can reason across text and images in one interaction

A multimodal model accepts more than one kind of input. A user can provide an image and ask a text question such as “What safety issue is visible here?” or “Summarize the chart and identify the largest change.” The model uses the visual content together with the instruction to produce a response.

This interaction is different from a narrow computer-vision API that returns one predefined type of result. Multimodal models can support more flexible visual question answering, descriptions, comparisons, and conversational follow-up. The broader multimodal AI fundamentals show how text, images, audio, and structured data can participate in one application pattern.

That flexibility does not remove the need for good prompts. The application should state what the user wants from the image, what level of detail is appropriate, and any format requirements for the answer.

Image generation creates new visual output from instructions

Image-generation models take a prompt and create new visual content rather than merely describing an existing image. The prompt can specify subject, composition, style, environment, or other desired characteristics. The model turns those instructions into a generated image.

The engineering question is whether generation is actually the required workload. If a retailer needs to categorize existing product photos, generation is unnecessary. If a design team needs concept art or a user wants a custom illustration, generation is appropriate.

Generated images also raise responsible-AI questions. The application may need content controls, clear disclosure, restrictions on unsafe requests, and policies for sensitive or misleading content. At AI-901 level, recognize that generation changes both capability and risk.

Visual prompting should be specific about the task

When a multimodal model receives an image, the model still needs an instruction. “Look at this” is not a useful application contract. A better prompt tells the model what to identify, how to format the answer, and which evidence matters.

For example, a maintenance application might ask the model to list visible defects, separate confirmed observations from uncertain possibilities, and return a short structured response. A document assistant might ask for specific fields rather than a general description of the page.

Prompt design is especially important when the image contains irrelevant detail. The instruction should focus the model on the business task and reduce the chance that the response drifts into speculation.

A lightweight Foundry vision app has a simple request path

The implementation objective in AI-901 is intentionally modest. A lightweight application needs a deployed model or Foundry capability, an input path for the image and prompt, a client that sends the request, and code that handles the result. The candidate should understand the components and flow rather than design a complex production media pipeline.

For visual interpretation, the client supplies the image together with the user instruction to a deployed multimodal model. For image generation, the application sends a generation prompt to the appropriate model and handles the returned visual asset. The application may then display, store, or pass the result into another approved workflow.

The current Microsoft AI certification direction matters here because Foundry is the common implementation environment across the updated objectives, not an optional side topic.

Content Understanding is different from free-form visual reasoning

The current study guide also includes information extraction from documents, images, audio, and video with Azure Content Understanding. That objective overlaps visually with computer vision, but the intent is different. Content Understanding is about extracting structured information from content, while a multimodal model can answer flexible questions or reason about what it sees.

If the requirement is to pull known fields from forms, receipts, images, or recorded media into structured output, think information extraction. If the requirement is to interpret a scene or answer an open-ended visual question, think multimodal reasoning. If the requirement is to create a new picture, think image generation.

A useful exam distinction is the shape of the expected result. Structured extraction begins with a defined information need: invoice number, date, total, speaker segment, product label, or another field that should be returned predictably. Visual reasoning is more flexible; the application may need a description, comparison, explanation, or answer based on what appears in an image. Generation reverses the direction entirely by taking instructions and producing new visual content. The same source image can therefore participate in different workloads, but the business question determines the correct capability. Candidates who first identify whether the solution must extract, understand, or generate are less likely to choose a service merely because the scenario contains an image.

Being able to separate those workloads is more useful than memorizing a long service list because it maps directly to scenario questions.

Evaluation should match the kind of visual task

A vision system cannot be evaluated with one universal metric. Classification, object detection, visual question answering, and image generation have different notions of quality. At fundamentals level, the main lesson is to test with representative inputs and verify that the output is useful for the intended scenario.

A product-inspection model should be tested on the range of lighting, camera angles, and defect types expected in production. A multimodal assistant should be tested on images with ambiguous or missing information so the team can see whether it admits uncertainty. An image generator should be evaluated for prompt adherence, safety, and suitability for the business purpose.

The more advanced multimodal generation and understanding for AI-103 goes further into production implementation. AI-901 candidates should keep the conceptual foundation: choose the correct visual workload and understand how a lightweight Foundry solution uses it.

Responsible AI is especially visible in image workloads

Images can contain biometric information, personal documents, private environments, children, medical information, or other sensitive content. Privacy and security therefore matter before the model ever produces an answer. The application should collect and retain only what it legitimately needs and should protect access to uploaded content.

Fairness can appear when performance differs across people or environments. Transparency matters when generated media could be mistaken for authentic content. Reliability and safety matter when visual interpretation influences high-impact decisions. Accountability matters when an organization uses AI output in an operational process.

These principles are not extra decoration around the model. They influence which workload is appropriate, how inputs are handled, how results are reviewed, and whether a human must remain in the decision loop.

Prepare by classifying scenarios before thinking about products

Take a set of use cases and label each one: visual interpretation, image generation, structured extraction, or another workload. Then identify whether the application needs a multimodal model, a generation model, or a content-extraction capability. Only after that should you think about the Foundry implementation.

Examples make the distinctions clear. “Describe damage in an uploaded photo” is visual interpretation. “Create a product concept image” is generation. “Extract invoice number and total from a photographed receipt” is information extraction. “Respond to a user who uploads a chart and asks a question” is multimodal interaction.

If you can recognize those workload boundaries, explain the basic request flow through Foundry, and identify the responsible-AI issues that accompany visual data, you have the computer-vision and image-generation objectives at the depth AI-901 expects.

Visual workloads also benefit from separating observation from inference. If an image clearly shows a red indicator, the model can report that observation. Deciding that the equipment is unsafe may require additional context, policy, or sensor data. In higher-risk applications, the prompt and user experience should make that boundary clear so a plausible visual interpretation is not mistaken for a verified operational conclusion.

For study, include edge cases such as blurry images, poor lighting, partially visible objects, screenshots with dense text, and prompts that ask for information not actually present in the image. A well-designed system should handle uncertainty instead of fabricating detail. That behavior connects visual workload knowledge with reliability, safety, and transparency.

  • img