AI-Driven Observability: From Logs to LLM-Powered Insights Training Course
Traditional observability relies on dashboards, threshold alerts, and manual log diving. AI-driven observability transforms this with natural language querying of telemetry data, LLM-powered root cause analysis, anomaly detection using foundation models, and automated incident summaries that understand context.
This instructor-led, live training (online or onsite) is aimed at observability and SRE engineers who want to integrate LLMs and AI into their monitoring, alerting, and incident analysis workflows.
By the end of this training, participants will be able to:
- Build natural language interfaces for querying Prometheus, Elasticsearch, and SQL-based observability stores.
- Implement LLM-powered log analysis and anomaly detection pipelines.
- Generate automated incident summaries and postmortem drafts from raw telemetry.
- Design AI-assisted root cause analysis workflows with evidence chaining.
- Integrate foundation models for time-series anomaly detection and forecasting.
- Deploy an AI-augmented on-call experience with smart alert enrichment.
Format of the Course
- Interactive lecture and discussion.
- Lots of exercises and practice.
- Hands-on implementation in a live-lab environment.
Course Customization Options
- To request a customized training, please contact us to arrange.
Course Outline
The AI Observability Landscape
- From dashboards to conversations: the shift toward AI-augmented observability
- LLM capabilities relevant to observability: summarization, reasoning, pattern matching
- Architecture patterns: embedding AI into existing observability stacks
Natural Language Telemetry Querying
- Text-to-PromQL: translating natural language into monitoring queries
- NL querying for Elasticsearch, OpenSearch, and Loki log stores
- SQL generation from natural language for structured telemetry
- Building a query assistant agent with tool use and context awareness
LLM-Powered Log Analysis
- Automated log parsing and structuring with LLMs
- Anomaly detection in log streams using embedding similarity
- Log clustering and pattern discovery at scale
- Generating human-readable explanations from raw log sequences
Intelligent Alerting and Incident Enrichment
- Alert correlation and deduplication with semantic understanding
- Automated incident context gathering from runbooks, past incidents, and docs
- Smart alert routing based on content understanding and team expertise
- Reducing alert fatigue with AI-driven noise reduction
AI-Assisted Root Cause Analysis
- Hypothesis generation from multi-source telemetry correlation
- Evidence chaining: connecting symptoms across metrics, logs, and traces
- Guided troubleshooting with interactive AI diagnosis sessions
- Building a root cause analysis agent with progressive investigation
Automated Incident Response and Communication
- Generating incident summaries and status updates from telemetry
- Automated postmortem drafting with timeline reconstruction
- Stakeholder communication tailored to technical and executive audiences
- Runbook suggestion and automated remediation recommendations
ML for Observability
- Time-series forecasting for capacity planning and anomaly prediction
- Foundation models for zero-shot anomaly detection on metrics
- Embedding-based service dependency mapping and topology discovery
- Training and deploying lightweight ML models alongside observability pipelines
Production Deployment and Ethics
- Latency and cost considerations for real-time AI observability
- Data privacy: ensuring LLMs do not leak sensitive telemetry
- Human oversight: when AI diagnosis needs operator validation
- Measuring impact: MTTD, MTTR, and on-call experience metrics
Requirements
- Experience with observability tools such as Prometheus, Grafana, Datadog, or OpenTelemetry.
- Familiarity with log management and metrics concepts.
- Basic Python scripting for data processing.
Audience
- SRE and observability engineers adopting AI-enhanced tooling.
- Platform engineers building next-generation monitoring pipelines.
- DevOps leads evaluating LLM integration into incident workflows.
Open Training Courses require 5+ participants.
AI-Driven Observability: From Logs to LLM-Powered Insights Training Course - Booking
AI-Driven Observability: From Logs to LLM-Powered Insights Training Course - Enquiry
AI-Driven Observability: From Logs to LLM-Powered Insights - Consultancy Enquiry
Provisional Upcoming Courses (Require 5+ participants)
Related Courses
Agentic Development with Gemini 3 and Google Antigravity
21 HoursGoogle Antigravity serves as a specialized agentic development environment, empowering users to create autonomous agents that leverage Gemini 3's multimodal strengths for planning, reasoning, coding, and execution.
Delivered as an instructor-led, live training session (either online or onsite), this program is tailored for advanced technical professionals looking to design, build, and deploy autonomous agents utilizing Gemini 3 and the Antigravity ecosystem.
By the end of this training, participants will be equipped to:
- Construct autonomous workflows that harness Gemini 3 for reasoning, strategic planning, and task execution.
- Create agents within Antigravity capable of analysing tasks, generating code, and interacting with various tools.
- Integrate Gemini-driven agents seamlessly with enterprise systems and APIs.
- Refine agent behaviour to ensure safety and reliability within complex operational environments.
Course Format
- A blend of expert-led demonstrations and interactive discussions.
- Hands-on experimentation focused on autonomous agent development.
- Practical implementation exercises using Antigravity, Gemini 3, and associated cloud tools.
Course Customization Options
- Should your team require specific domain-focused agent behaviours or bespoke integrations, please reach out to us to tailor the program to your needs.
Advanced Antigravity: Feedback Loops, Learning & Long-Term Agent Memory
14 HoursGoogle Antigravity serves as a sophisticated framework for experimenting with long-lived agents and emergent interactive behaviours.
This live, instructor-led training—available online or onsite—is designed for advanced professionals seeking to design, analyse, and optimise agents that can retain memories, adapt through feedback, and evolve over extended operational periods.
Upon completing this course, participants will develop the capability to:
- Design long-term memory structures to ensure agent persistence.
- Implement effective feedback loops to influence and shape agent behaviour.
- Evaluate learning trajectories and monitor model drift.
- Integrate memory mechanisms into complex multi-agent ecosystems.
Course Format
- Expert-led discussions accompanied by technical demonstrations.
- Hands-on exploration via structured design challenges.
- Application of concepts within simulated agent environments.
Customisation Options
- If your organisation requires tailored content or case-specific examples, please contact us to customise this training.
Advanced Mastra Integrations: APIs, Tools, Enterprise Data & External Systems
21 HoursMastra is a framework facilitating deep integration between AI agents, APIs, enterprise applications, and external data systems.
This instructor-led training, available online or onsite, is designed for intermediate-level engineers looking to build reliable, secure, and scalable integrations between Mastra agents and the wider enterprise ecosystem.
Upon completion of this training, participants will be equipped to:
- Implement API-driven integrations between Mastra agents and external services.
- Link enterprise data systems and tools to automated agent workflows.
- Apply secure data exchange and authentication best practices.
- Design integration layers that are scalable, maintainable, and ready for production.
Course Format
- Interactive lectures and discussions.
- Hands-on integration engineering and API exercises.
- Live-lab implementations using real-world enterprise scenarios.
Course Customisation Options
- Custom API scenarios, enterprise system mappings, or data-integration workshops are available upon request.
Interactive AI Agents: AgentCore Memory, Code Interpreter & Browser Tool in Action
14 HoursAgentCore equips AI agents with memory persistence, a secure code interpreter, and a browser tool, enabling the delivery of highly interactive, dynamic, and context-aware user experiences.
Delivered as a live, instructor-led session (either online or onsite), this training is tailored for intermediate to advanced technical professionals looking to design and deploy AI agents that retain long-term context, perform on-the-fly calculations, and directly engage with web user interfaces.
Upon completion of this training, participants will be equipped to:
- Implement AgentCore memory to establish stateful, context-aware workflows.
- Utilise the secure code interpreter to drive dynamic calculations and data transformations.
- Integrate the browser tool to facilitate real-time data retrieval and direct UI interaction.
- Design interactive agents specifically for analytics, customer support, and research applications.
Course Format
- Interactive lectures and facilitated discussions.
- Practical lab exercises focusing on AgentCore memory and tools.
- In-depth case studies covering analytics, automation, and customer support scenarios.
Customisation Options
- Please contact us to arrange a customised training experience tailored to your specific requirements.
Accelerating AI Agent Deployment with AgentCore Runtime & Gateway
14 HoursThe AgentCore Runtime & Gateway is a pairing of AWS services designed for packaging, deploying, and securely exposing AI agents, featuring streamlined integrations with external systems.
\rThis instructor-led live training, available online or onsite, targets intermediate-level engineering teams looking to transition agent prototypes into production. Participants will master the AgentCore Runtime for deployment and the Gateway for secure connectivity and API integration.
Upon completing this training, participants will be able to:
- Deploy AgentCore Runtime environments and package agents for release.
- Expose agents via the Gateway using authenticated, rate-limited endpoints.
- Integrate external tools and APIs into agent workflows using stable contracts.
- Set up observability, logging, and usage monitoring for production operations.
Course Format
- Interactive lectures and discussions.
- Hands-on labs covering Runtime deployments and Gateway integrations.
- Practical exercises focused on reliability, security, and rollout strategies.
Customisation Options
- To request tailored training for this course, please contact us to make arrangements.
Antigravity for Developers: Building Agent-First Applications
21 HoursAntigravity serves as a purpose-built development platform for crafting AI-driven, agent-first applications.
This instructor-led, live training session—available either online or onsite—is tailored for intermediate-level developers seeking to construct real-world solutions using autonomous AI agents within the Antigravity ecosystem.
Upon completion, participants will be fully equipped to:
- Build applications that leverage autonomous and coordinated AI agents.
- Leverage the Antigravity IDE, editor, terminal, and browser for comprehensive end-to-end development.
- Orchestrate multi-agent workflows via the Agent Manager.
- Integrate agent capabilities seamlessly into production-grade software systems.
Course Delivery Format
- A blend of presentations and in-depth technical demonstrations.
- Extensive hands-on practice complemented by guided exercises.
- Practical implementation work conducted directly within the live Antigravity environment.
Customisation Options
- To align content specifically with your development stack, please get in touch to arrange a tailored version of this training.
Getting Started with Antigravity: An Introduction to Agent-First IDEs
14 HoursGoogle Antigravity represents a new paradigm in development environments, specifically engineered as an agent-first platform to optimise engineering workflows through intelligent automation.
This live, instructor-led training session (available online or onsite) is tailored for entry-level practitioners eager to investigate the core principles of Antigravity and discover how agent-driven coding environments can significantly boost productivity.
By the end of this course, participants will be equipped to:
- Install and configure Google Antigravity.
- Navigate and comprehend both the Editor View and Manager View interfaces.
- Collaborate effectively with agents to automate straightforward development tasks.
- Leverage Antigravity to generate, refine, and manage project files.
Course Format
- Instructor-led explanations complemented by live, real-time demonstrations.
- Guided, hands-on exercises designed to deepen practical agent usage.
- Practical exploration of key Antigravity features within a controlled lab setting.
Customisation Options
- Should you require a bespoke version of this training, please contact us to arrange a tailored program.
Antigravity for Web Automation & Browser-Based Tasks
21 HoursGoogle Antigravity serves as a robust platform for developing agents that can engage with web applications, browser environments, and complex multi-surface workflows.
This instructor-led, live training—available both online and on-site—is specifically designed for intermediate-level professionals looking to build, automate, and test browser-based workflows using Google Antigravity.
By the end of the training, participants will be equipped to:
- Develop agents that interact effectively with web applications within the browser surface.
- Automate end-to-end workflows across various browser contexts.
- Validate and troubleshoot agent behaviour in UI-driven environments.
- Implement cross-surface automation strategies leveraging Antigravity.
Course Format
- Guided instruction complemented by live demonstrations.
- Practical, hands-on activities alongside scenario-based exercises.
- Implementation of agent workflows within an interactive lab environment.
Course Customisation Options
- Please contact us to tailor the course to your specific objectives if you require customised training.
Building Fully Managed AI Agents with AgentCore: From Concept to Production
14 HoursAgentCore streamlines the creation, optimisation, and oversight of fully managed AI agents by offering an integrated suite of services designed for scalable deployment.
This live, instructor-led training session—available online or in person—is designed for practitioners ranging from beginner to intermediate levels who are eager to acquire practical experience in developing production-grade AI agents using AgentCore.
Upon completion of this training, participants will be equipped to:
- Comprehend the fundamental capabilities of AgentCore for developing AI agents.
- Architect and configure straightforward AI agents leveraging managed services.
- Incorporate workflows to augment agent functionality.
- Deploy and monitor AI agents within production environments.
Course Delivery Model
- Interactive lectures and group discussions.
- Practical laboratory work focusing on AgentCore services.
- Guided exercises covering the journey from agent concept to deployment.
Tailoring the Course
- For bespoke training arrangements, please get in touch with us to discuss your specific requirements.
AI Agent Development with Mastra
14 HoursThis instructor-led live training, delivered online or on-site, is designed for intermediate-level software developers and engineering teams seeking to build scalable, observable AI systems using Mastra.
Upon completion of this training, participants will be equipped to:
- Comprehend Mastra’s architecture and its integration capabilities with LLMs and external APIs.
- Architect and implement AI agents and workflows using TypeScript.
- Leverage Mastra’s observability and memory tools to monitor and enhance agent performance.
- Deploy production-ready AI applications by capitalising on Mastra’s framework features.
Mastra Debugging, Evaluation & Quality Assurance for AI Agents
21 HoursMastra is a framework offering structured tools to evaluate, debug, and assure the reliability of AI agents operating across complex workflows.
This instructor-led, live training (available online or onsite) is designed for intermediate-level practitioners who want to rigorously test agent behaviour, enhance reliability, and implement measurable evaluation processes.
Upon completion of this training, participants will confidently:
- Apply debugging techniques to identify and correct issues with agent behaviour.
- Evaluate agents using structured metrics, benchmarks, and quality scores.
- Implement tooling and workflows to track reliability, drift, and hallucinations.
- Design QA strategies that ensure consistent and predictable agent performance.
Course Format
- Interactive lectures and discussions.
- Hands-on debugging and evaluation exercises.
- Live-lab analysis of agent behaviours using observability tools.
Course Customisation Options
- Customised reliability testing scenarios and industry-specific QA methods can be arranged upon request.
Mastra Ops & Production Engineering: Deploying and Scaling AI Agents
21 HoursMastra serves as an operational framework aimed at streamlining the deployment, scaling, and lifecycle management of AI agents within production environments.
This instructor-led live training, available online or onsite, is tailored for intermediate to advanced technical professionals seeking to reliably and efficiently operationalise AI agents across their production systems.
Upon completing this training, participants will be equipped to:
- Deploy Mastra-based AI agents into controlled, production-grade environments.
- Scale agents horizontally and vertically using platform-native primitives.
- Implement observability pipelines to monitor agent behaviour and performance.
- Optimise runtime configurations to reduce latency, costs, and operational risks.
Format of the Course
- Interactive lecture and discussion.
- Hands-on exercises focused on real deployment scenarios.
- Live-lab implementation using containerized and orchestrated environments.
Course Customisation Options
- Customisation of topics, hands-on labs, or industry-specific scenarios is available upon request.
Mastra Workflow Automation & Multi-Agent Orchestration
21 HoursMastra is a framework designed to enable sophisticated workflow automation and coordination across multiple AI agents operating within distributed systems.
This instructor-led, live training (available online or onsite) is tailored for intermediate-level practitioners seeking to design, orchestrate, and manage multi-agent workflows at scale.
Upon completing this training, participants will acquire the skills to:
- Design complex workflows leveraging Mastra’s orchestration capabilities.
- Coordinate multiple agents executing parallel or dependent tasks.
- Implement monitoring and debugging tools for workflow execution.
- Optimise orchestration logic to enhance reliability, throughput, and automation efficiency.
Course Format
- Interactive lectures and discussions.
- Hands-on exercises in workflow design and automation.
- Practical implementation within a containerised live-lab environment.
Course Customisation Options
- Customised automation scenarios, enterprise integrations, or workflow patterns can be provided upon request.
Managing Agent Workflows in Google Antigravity: Orchestration, Planning and Artifacts
14 HoursGoogle Antigravity serves as an agent-centric development platform designed to coordinate, supervise, and orchestrate AI-driven coding and automation processes.
Delivered by expert instructors in a live format—either online or on-site—this programme targets intermediate-level professionals seeking to design, manage, and refine multi-agent workflows within the Google Antigravity environment.
By the end of this training, participants will have developed the competency to:
- Define agent responsibilities and construct orchestration pipelines via the Manager interface.
- Create and analyse Antigravity artifacts, such as task lists, strategic plans, system logs, and browser session recordings.
- Establish verification protocols to ensure that agent actions are both transparent and subject to audit.
- Enhance multi-agent collaboration to address complex development and operational requirements.
Delivery Methodology
- Guided presentations complemented by practical live demonstrations.
- Scenario-driven exercises addressing real-world workflow challenges.
- Active, hands-on experimentation within a live Antigravity workspace.
Customisation Options
- Should you require a bespoke version of this curriculum, please reach out to discuss tailored customisation options.
Testing & Verifying Agent-Driven Code: Quality Assurance in Antigravity
14 HoursAntigravity is a framework that represents advanced agent-driven development workflows.
This instructor-led, live training (online or onsite) is aimed at intermediate to advanced professionals who wish to verify, validate, and secure the output produced by AI agents working within Antigravity-driven environments.
Upon completing this training, participants will be able to:
- Assess the accuracy and safety of agent-generated code artifacts.
- Use structured techniques to verify agent-executed tasks.
- Analyze browser recordings and trace agent activity effectively.
- Apply QA and security principles to ensure the reliability of agent workflows.
Format of the Course
- Instructor-guided technical briefings and discussions.
- Practical exercises focused on verifying real agent workflows.
- Hands-on testing and validation within a controlled lab environment.
Course Customization Options
- Adaptation of scenarios, workflows, and testing examples is available upon request.