Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps
- Defining AIOps and its strategic importance
- Contrasting traditional monitoring with AIOps-driven observability
- Examining AIOps architecture and its essential components
Collecting and Normalising Operational Data
- Types of observability data: metrics, logs, and traces
- Ingesting data from diverse sources such as servers, containers, and cloud environments
- Utilising agents and exporters (e.g., Prometheus, Beats, Fluentd)
Data Correlation and Anomaly Detection
- Applying time series correlation and statistical methods
- Deploying ML models for anomaly detection
- Identifying incidents within distributed systems
Alerting and Noise Reduction
- Crafting intelligent alert rules and thresholds
- Implementing suppression, deduplication, and alert grouping
- Integration with Alertmanager, Slack, PagerDuty, or Opsgenie
Root Cause Analysis and Visualisation
- Using dashboards to visualise metrics and uncover trends
- Analysing events and timelines for RCA (Root Cause Analysis)
- Tracing issues across system layers using distributed tracing tools
Automation and Remediation
- Triggering automated scripts or workflows in response to incidents
- Integrating with ITSM systems (e.g., ServiceNow, Jira)
- Practical use cases: self-healing, auto-scaling, and traffic rerouting
Open Source and Commercial AIOps Platforms
- Overview of key tools: Prometheus, Grafana, ELK, Moogsoft, Dynatrace
- Establishing evaluation criteria for selecting an AIOps platform
- Demonstrations and hands-on exercises with a chosen stack
Summary and Recommended Next Steps
Requirements
- A solid grasp of IT operations and system monitoring concepts
- Practical experience with monitoring tools or dashboards
- Familiarity with standard log and metric formats
Target Audience
- Operations teams managing infrastructure and applications
- Site Reliability Engineers (SREs)
- IT teams focused on monitoring and observability