At re:Invent 2025, AWS announced a series of significant improvements for Cloud Operations. This year’s focus revolves around applying Generative AI for observability and troubleshooting, while simplifying the management of large-scale operational data through centralization features.
Below are detailed notes on the top 10 announcements that help optimize operational processes on AWS.
AWS launched comprehensive observability capabilities for Generative AI applications. This feature provides detailed insights into latency, token usage, and errors occurring in the AI stack. Notably, it integrates seamlessly with Amazon Bedrock AgentCore and open-source frameworks like LangChain or CrewAI without requiring manual instrumentation code.

[Figure 1] Amazon CloudWatch Generative AI dashboard
The Agent Management View also helps track the agent’s workflow from end-to-end:

[Figure 2] Agent Management View
Previously, creating an Application Map often required complex instrumentation setup. Now, CloudWatch Application Signals can automatically detect and visualize application topology, displaying service dependencies instantly. This feature helps operations teams gain a system overview more quickly.

[Figure 3] CloudWatch Application Map
This is a major advancement in AIOps. CloudWatch Investigations leverages Generative AI to automate the root cause analysis process. Instead of manually aggregating data, the system generates an interactive incident report.

[Figure 4] Amazon CloudWatch Investigations Incident Report
Most notably, the “5 Whys” analysis process is built-in, simulating Amazon’s internal Correction of Errors (COE) methodology, helping to systematically identify the underlying causes of issues.

[Figure 5] 5 Whys Analysis in the CloudWatch investigations Incident Report
CloudWatch and Application Signals now support Model Context Protocol (MCP) servers. This acts as a bridge, allowing AI assistants to naturally interact with observability data (metrics, logs, traces). This enables us to build autonomous operational workflows and integrate CloudWatch data into AI-powered development tools.
To better support developers, CloudWatch Application Signals has been integrated directly into GitHub Actions. This feature provides observability insights right within Pull Requests and CI/CD pipelines.
It helps identify performance issues or system errors without leaving the GitHub development environment:

[Figure 6] Automated Root Cause Analysis in the GitHub issues
The system can even suggest automated bug fixes through Pull Requests:

[Figure 7] Automated GitHub Pull Request to Fix the Issue
Amazon OpenSearch Service introduces significant improvements for Piped Processing Language (PPL). Log analysis becomes faster and more intuitive. Enhanced query capabilities help process complex analytical queries efficiently, while seamlessly integrating with CloudWatch Logs for unified log analysis.
Amazon CloudWatch RUM extends real user experience monitoring to mobile platforms. We can now track performance, user journeys, and client-side errors on iOS and Android applications, ensuring consistent experiences across all devices and geographic locations.
To address the massive volume of logs from API activity, AWS CloudTrail adds the Event Aggregation feature. Instead of logging each individual entry, the system summarizes high-frequency activities into aggregated reports every 5 minutes.
This helps:
This is a highly anticipated feature for large organizations. CloudWatch Logs Centralization allows collecting logs from multiple Accounts and Regions into a single destination account.
@aws.account and @aws.region context to logs for easy source tracking.Similar to logs, CloudWatch Database Insights also supports centralized monitoring. We can track the performance of Amazon RDS, Amazon Aurora, and Amazon DynamoDB across the entire AWS organization from a single monitoring account. This makes it easier to correlate database performance with application health.
The 2025 announcements show that AWS is focused on solving the “operational data overload” problem by using AI to filter noise and automate analysis. From monitoring GenAI, to using GenAI to fix errors, and centralization capabilities, these tools help operators shift from a reactive state to proactively controlling systems.

Nereida is a WW Specialist Solutions Architect in Cloud Operations focusing on Centralized Operations Management and Application operations on AWS. When she isn't working, she enjoys traveling to attend music concerts.

Calvin Weng is a Product Marketing Manager for AWS Cloud Operations, focusing on observability and monitoring services. Outside of work, Calvin travels, practices pottery, plays ping pong competitively, and explores the Pacific Northwest with his dog Kai.

Raviteja Sunkavalli is a Senior Worldwide Specialist Solutions Architect at Amazon Web Services, specializing in AIOps and GenAI observability. He helps global customers implement observability and incident management solutions across complex and distributed cloud environments. Outside of work, he enjoys playing cricket and exploring new cooking recipes.