AI-300 Study Guide: Machine Learning Operations Engineer Associate
AI-300 is the exam behind the Microsoft Certified: Machine Learning Operations Engineer Associate certification. It tests your ability to set up infrastructure for machine learning operations (MLOps) and generative AI operations (GenAIOps) on Azure, together referred to by Microsoft as AI operations (AIOps). If DP-100 was your target and you are now redirecting your study time, this guide walks through what AI-300 actually covers and how to prepare for it.
What AI-300 Tests
AI-300 candidates need subject matter expertise in setting up infrastructure for MLOps and GenAIOps solutions on Azure. That means experience training, optimising, deploying, and maintaining traditional machine learning models with Azure Machine Learning, plus experience deploying, evaluating, monitoring, and optimising generative AI applications and agents with Microsoft Foundry.
You'll also want a data science background with Python programming experience, and an entry-level understanding of DevOps practices, including tools like GitHub Actions and command-line interfaces. On top of that, you need knowledge and experience in MLOps using:
- Azure Machine Learning
- Microsoft Foundry
- GitHub Actions
- Infrastructure as code (IaC) practices with Bicep and Azure CLI
The role covers designing and implementing MLOps infrastructure, implementing machine learning model lifecycle and operations, designing and implementing GenAIOps infrastructure, implementing generative AI quality assurance and observability, and optimising generative AI systems and model performance. Day to day, that means working with data scientists, DevOps teams, and stakeholders to deliver scalable AI solutions with automation and monitoring built in.
Who Should Take AI-300
This certification sits at intermediate level under the AI Engineer role. It suits:
- MLOps engineers who own the infrastructure that trains, deploys, and monitors machine learning models in production
- AI engineers building and operating generative AI applications and agents on Microsoft Foundry
- Data scientists moving from model-building into the operations side of the lifecycle
- DevOps engineers extending CI/CD and infrastructure-as-code practice into ML and GenAI workloads
If you were preparing for the retired DP-100, much of your Azure Machine Learning work (workspaces, MLflow, automated machine learning, pipelines, endpoints) carries over to AI-300's first two domains. If you're mapping the whole path rather than just this exam, our Azure Data Scientist career path guide covers where AI-300 sits between AI-901 and AI-103.
Exam Format
The AI-300 exam is proctored and administered through Pearson VUE.
- Duration: 120 minutes
- Passing score: 700 or greater, on a scale of 1 to 1,000
- Certification level: Associate (Intermediate), role: AI Engineer
- Language: English
- Retakes: available 24 hours after a failed first attempt; the wait varies for later retakes
Most questions test features that are generally available, though a few may cover preview features that are already in common use. Treat the bullets under each skill as examples of how it gets assessed, not a complete list of everything that could appear.
AI-300 Skills Measured and Weighting
The exam breaks into five domains:
| Domain | Weight |
|---|---|
| Design and implement an MLOps infrastructure | 15-20% |
| Implement machine learning model lifecycle and operations | 25-30% |
| Design and implement a GenAIOps infrastructure | 20-25% |
| Implement generative AI quality assurance and observability | 10-15% |
| Optimize generative AI systems and model performance | 10-15% |
Model lifecycle and operations carries the heaviest weight, so plan your study time accordingly, but do not neglect the GenAIOps domain: between GenAIOps infrastructure, quality assurance and observability, and optimisation, generative-AI-specific content makes up close to half the exam.
Domain 1: Design and Implement an MLOps Infrastructure (15-20%)
Create and manage resources in a Machine Learning workspace
- Create and manage a workspace
- Create and manage datastores
- Create and manage compute targets
- Configure identity and access management for workspaces
Create and manage assets in a Machine Learning workspace
- Create and manage data assets
- Create and manage environments
- Create and manage components
- Share assets across workspaces using registries
Implement IaC for Machine Learning
- Configure GitHub integration with Machine Learning to enable secure access
- Deploy Machine Learning workspaces and resources using Bicep and Azure CLI
- Automate resource provisioning using GitHub Actions workflows
- Restrict network access to Machine Learning workspaces
- Manage source control for machine learning projects using Git
Domain 2: Implement Machine Learning Model Lifecycle and Operations (25-30%)
This is the largest domain, and it covers the full path from training through production monitoring.
Orchestrate model training
- Configure experiment tracking with MLflow
- Use automated machine learning to explore optimal models
- Use notebooks for experimentation and exploration
- Automate hyperparameter tuning
- Run model training scripts
- Manage distributed training for large and deep learning models
- Implement training pipelines
- Compare model performance across jobs
Implement model registration and versioning
- Package a feature retrieval specification with the model artifact
- Register an MLflow model
- Evaluate a model using responsible AI principles
- Manage model lifecycle, including archiving models
Deploy machine learning models for production environments
- Deploy models as real-time or batch endpoints with managed inference options
- Test and troubleshoot model endpoints
- Implement progressive rollout and safe rollback strategies
Monitor and maintain machine learning models in production
- Detect and analyze data drift
- Monitor performance metrics of models deployed to production
- Configure retraining or alert triggers when thresholds are exceeded
Domain 3: Design and Implement a GenAIOps Infrastructure (20-25%)
Implement Foundry environments and platform configuration
- Create and configure Foundry resources and project environments
- Configure identity and access management with managed identities and role-based access control (RBAC)
- Implement network security and private networking configurations
- Deploy infrastructure using Bicep templates and Azure CLI
Deploy and manage foundation models for production workloads
- Deploy foundation models using serverless API endpoints and managed compute options
- Select appropriate models for specific use cases
- Implement model versioning and production deployment strategies
- Configure provisioned throughput units for high-volume workloads
Implement prompt versioning and management with source control
- Design and develop prompts
- Create prompt variants and compare performance across different prompts
- Implement version control for prompts using Git repositories
Domain 4: Implement Generative AI Quality Assurance and Observability (10-15%)
Configure evaluation and validation for generative AI applications and agents
- Create test datasets and data mapping for comprehensive model evaluation
- Implement AI quality metrics, including groundedness, relevance, coherence, and fluency
- Configure risk and safety evaluations for harmful content detection
- Set up automated evaluation workflows using built-in and custom evaluation metrics
Implement observability for generative AI applications and agents
- Examine continuous monitoring in Foundry
- Monitor performance metrics, including latency, throughput, and response times
- Track and optimize cost metrics, including token consumption and resource usage
- Configure detailed logging, tracing, and debugging capabilities for production troubleshooting
Domain 5: Optimize Generative AI Systems and Model Performance (10-15%)
Optimize retrieval-augmented generation (RAG) performance and accuracy
- Optimize retrieval performance by tuning similarity thresholds, chunk sizes, and retrieval strategies
- Select and fine-tune embedding models for domain-specific use cases and accuracy improvements
- Implement and optimize hybrid search approaches combining semantic and keyword-based retrieval
- Evaluate and improve RAG system performance using relevance metrics and A/B testing frameworks
Implement advanced fine-tuning and model customization
- Design and implement advanced fine-tuning methods
- Create and manage synthetic data for fine-tuning
- Monitor and optimize fine-tuned model performance
- Manage a fine-tuned model from development through production deployment
From DP-100 to AI-300: What Changed
DP-100 had four domains: design and prepare a machine learning solution (20-25%), explore data and run experiments (20-25%), train and deploy models (25-30%), and optimize language models for AI applications (25-30%). Its language-model domain already covered prompt flow, RAG and fine-tuning.
AI-300 keeps the Azure Machine Learning core (workspaces, MLflow, automated machine learning, hyperparameter tuning, pipelines, endpoints) and adds operations work DP-100 did not name: infrastructure as code with Bicep, Azure CLI and GitHub Actions, progressive rollout and rollback, drift-triggered retraining, Foundry environment configuration, provisioned throughput, prompt version control in Git, and generative AI observability. If you prepared for DP-100, plan extra study time for the IaC skills and for domains 3 and 4.
Study Plan for AI-300
A structured plan that follows the exam's own domain weighting:
Weeks 1-2: Workspace and IaC foundations
Set up an Azure Machine Learning workspace and a Microsoft Foundry project. Practise creating datastores, compute targets, and data assets. Deploy a workspace using Bicep and Azure CLI, then automate that provisioning with a GitHub Actions workflow.
Weeks 3-4: Model training and lifecycle
Configure MLflow experiment tracking. Run training jobs from notebooks and from scripts, then compare runs. Register an MLflow model, evaluate it against responsible AI principles, and practise archiving older model versions.
Weeks 5-6: Deployment and production monitoring
Deploy models to real-time and batch endpoints. Practise a progressive rollout with a rollback plan. Set up data drift detection and configure retraining triggers.
Weeks 7-8: Foundry and GenAIOps infrastructure
Configure a Foundry project's identity, RBAC, and private networking. Deploy a foundation model behind a serverless API endpoint and behind managed compute, and configure provisioned throughput for a high-volume scenario.
Weeks 9-10: Prompts, evaluation, and observability
Version prompts in Git and compare variants. Build a test dataset and run an evaluation workflow scoring groundedness, relevance, coherence, and fluency, including a risk and safety pass for harmful content. Wire up monitoring for latency, throughput, and token cost.
Weeks 11-12: RAG optimisation, fine-tuning, and practice
Tune retrieval: chunk sizes, similarity thresholds, hybrid search. Run an advanced fine-tuning job, including synthetic data generation, and monitor the fine-tuned model through to production. Close out with full-length practice tests across all five domains.
Common Pitfalls
Treating this as a DP-100 refresh. Training, tuning and MLflow tracking still appear in domain 2, but they sit alongside infrastructure, deployment strategy and GenAIOps work that DP-100 did not ask for. Split your time the way the weightings do.
Skipping GenAIOps hands-on work. Domains 3 to 5 together carry 40-55% of the exam. A candidate strong on Azure Machine Learning but who has never touched Microsoft Foundry will lose marks across nearly half the exam.
Under-preparing observability. Groundedness, relevance, coherence, fluency, latency, throughput, and token cost are all named explicitly in the skills list. Generic "monitor your app" knowledge will not cover the specific metrics Microsoft names.
Ignoring IaC. Bicep and Azure CLI deployment, plus GitHub Actions automation, appear in both the MLOps and GenAIOps domains. If you have only used the Azure portal UI, budget time to practise the IaC path.
Worked Example: Deciding When to Retrain
Here is the kind of scenario to expect on the model lifecycle domain, the heaviest-weighted part of the exam. This is an example we wrote, not a real exam question.
A fraud-detection model has run as a real-time endpoint in production for six months. Monitoring shows precision has slipped from 92% to 78% over the last four weeks, and the input distribution for transaction amount has shifted noticeably from the training data. What should you configure?
- A) Redeploy the same model version to reset its metrics
- B) Configure a data drift alert with a retraining trigger tied to the feature shift, and monitor precision alongside it
- C) Increase the endpoint's compute allocation
- D) Roll back to the previous model version with no further changes
B is correct. Falling precision alongside a shifted input distribution is the signature of data drift, and the model lifecycle domain specifically calls for configuring retraining or alert triggers when thresholds are exceeded rather than watching the dashboard and waiting. Redeploying the same version (A) or adding compute (C) does nothing about drifted inputs, and rolling back (D) only helps if the earlier version saw inputs matching training data, which this scenario does not establish.
Practice Questions
azureprep.com offers practice questions across current Azure certifications, including AI-300. Work through domain-by-domain practice sets first; save the long mixed tests for later, so you can see which of the five domains needs more time before you move to full-length exams.
Ready to start? Practise for AI-300 on azureprep.com.