Role-Based Programme
RB1505

AI-Powered DevOps & Site Reliability Engineering

Smarter Automation, Monitoring, Incident Response & Reliability

IMAGE REQUIRED
Duration
8 Hours
Level
Basic
Delivery
Instructor-Led
Format
Workshop

Programme Objectives

  • Understand how AI can support DevOps and Site Reliability Engineering across CI/CD, infrastructure operations, monitoring, incident response, and reliability improvement.
  • Apply AI-assisted techniques to analyse logs, alerts, deployment information, operational incidents, configuration data, and reliability metrics.
  • Use structured prompting for troubleshooting, deployment review, incident analysis, runbook creation, post-incident reviews, and operational reporting.
  • Explore AI-supported approaches for identifying recurring failures, performance bottlenecks, reliability risks, and automation opportunities.
  • Build responsible AI-assisted DevOps and SRE workflows while maintaining production safety, security, change control, observability, and human oversight.

Tools covered

Generative AI AssistantsAI Search & ResearchCI/CD AnalysisInfrastructure & Configuration AssistanceLog & Observability AnalysisIncident Response AIReliability AnalysisRunbook GenerationReporting & Workflow Automation

Who should attend

  • DevOps Engineers
  • Site Reliability Engineers
  • Platform Engineers
  • Cloud Engineers
  • Infrastructure Engineers
  • Build & Release Engineers
  • CI/CD Engineers
  • Cloud Operations Professionals
  • Production Support Engineers
  • Systems Engineers
  • Application Support Engineers
  • Reliability Engineers
  • DevOps Managers
  • DevOps & SRE Team Leads

Prerequisites & Participant Readiness

  • Basic understanding of software delivery, infrastructure, cloud, or IT operations
  • Familiarity with deployments, pipelines, logs, monitoring, incidents, or configuration management is helpful
  • Basic computer and command-line awareness is beneficial
  • No AI or advanced programming knowledge required
  • No previous AI training required

TOC Modules

Concepts
  • Understanding Generative AI and its relevance to DevOps and Site Reliability Engineering
  • Identifying AI applications across CI/CD, infrastructure operations, observability, incidents, and documentation
  • Understanding AI assistance versus engineer judgement and operational accountability
  • Recognising limitations such as incorrect commands, incomplete context, hallucinations, and unsafe recommendations
Practical activities
  • Mapping a typical DevOps and SRE workflow
  • Identifying repetitive and information-intensive activities suitable for AI assistance
  • Comparing a traditional operations task with an AI-assisted approach

Scenarios

Failed Deployment to Safe Recovery

Deployment Failure → AI-Assisted Log Review → Failure Summary → Dependency Check → Recovery Options → Engineer Validation → Rollback / Fix → Verification

Participants use AI to organise sample deployment evidence, identify investigation areas, and prepare a safe recovery workflow for engineering validation.

Reliability Data to SRE Improvement Plan

Alerts + Incidents + Availability + Latency + Recovery Data → AI Analysis → Recurring Patterns → Reliability Gaps → Automation Opportunities → Improvement Plan

Participants use AI to analyse sample operational and reliability data, identify recurring service issues, and prepare a management-ready DevOps & SRE improvement plan.

Continue with programmes from the same capability area.

Take the next step

Ready to make this programme work for your team?

Customise modules, duration and business scenarios for your team.

Instructor-ledVirtualHybrid

Designed around your roles, tools and real workflows.