Role-Based Programme
RB1506

AI-Powered DevOps & Site Reliability Engineering

Smarter Automation, Reliability & Incident Management

IMAGE REQUIRED
Duration
16 Hours
Level
Intermediate
Delivery
Instructor-Led
Format
Capability Training

Programme Objectives

  • Apply AI across DevOps and Site Reliability Engineering activities including CI/CD, infrastructure operations, observability, incident management, and reliability improvement.
  • Use AI-assisted techniques to analyse deployment data, logs, alerts, configuration information, operational metrics, and recurring service failures more efficiently.
  • Develop structured workflows for deployment reviews, incident triage, root-cause analysis, change coordination, reliability monitoring, and post-incident improvement.
  • Improve service reliability through AI-assisted SLI/SLO analysis, anomaly identification, capacity insights, automation planning, and operational reporting.
  • Apply responsible AI practices covering production safety, credentials, access control, configuration validation, sensitive technical data, and human oversight.

Tools covered

Generative AI AssistantsCI/CD Analysis SupportInfrastructure Automation SupportLog & Monitoring AnalysisIncident Response SupportReliability AnalyticsConfiguration ReviewDocumentation SupportReporting AssistanceWorkflow Automation

Who should attend

  • DevOps Engineers
  • Site Reliability Engineers
  • DevOps Leads
  • SRE Leads
  • Platform Engineers
  • Cloud Engineers
  • Infrastructure Engineers
  • Release Engineers
  • Build & Deployment Engineers
  • Cloud Operations Professionals
  • Production Support Engineers
  • Reliability Analysts
  • Automation Engineers
  • Information Technology Team Leads

Prerequisites & Participant Readiness

  • Working knowledge of DevOps, cloud, infrastructure, application operations, or software delivery
  • Familiarity with CI/CD, monitoring, logs, deployments, incidents, or infrastructure concepts is helpful
  • Basic understanding of scripting, command-line tools, and technical documentation is helpful
  • Basic awareness of Generative AI is helpful
  • No advanced AI or machine-learning knowledge required

TOC Modules

Concepts
  • Understanding Generative AI, analytics, automation, and their role in DevOps and SRE
  • Identifying AI applications across delivery pipelines, operations, observability, incidents, and reliability engineering
  • Understanding AI assistance versus DevOps Engineer and SRE accountability
  • Recognising risks related to production systems, credentials, unsafe commands, and uncontrolled automation
Practical activities
  • Mapping DevOps and SRE workflows to AI-assisted opportunities
  • Identifying repetitive operational and engineering tasks suitable for AI support
  • Comparing traditional and AI-assisted reliability workflows

Scenarios

Deployment Failure to Reliable Recovery

Deployment → AI-Assisted Log & Metric Review → Failure Detection → Incident Triage → Root-Cause Analysis → Rollback / Recovery → Post-Incident Review → Preventive Actions

Participants analyse a simulated production deployment failure, organise technical evidence, identify likely failure causes, coordinate safe recovery, and prepare a structured post-incident improvement plan.

Reliability Data to SRE Improvement Plan

Availability + Latency + Incident Data + Deployment Metrics + Capacity Data → AI Analysis → Reliability Gaps → SLO Risks → Priority Improvements → Automation Opportunities → Management Report

Participants consolidate SRE and DevOps performance information, identify recurring reliability concerns and operational bottlenecks, and prepare a management-ready reliability improvement plan with clear actions, owners, and priorities.

Continue with programmes from the same capability area.

Take the next step

Ready to make this programme work for your team?

Customise modules, duration and business scenarios for your team.

Instructor-ledVirtualHybrid

Designed around your roles, tools and real workflows.