AI's Rapid Enterprise Adoption Reveals Operational Gaps for Reliability Teams
A new study indicates that as artificial intelligence integrates more deeply into enterprise operations, Site Reliability Engineering and platform teams face escalating challenges.
The accelerated integration of artificial intelligence into enterprise operations is highlighting critical breaking points, placing significant new demands on site reliability engineering and platform teams who are now central to ensuring AI's trustworthiness and scalability.
The Shifting Landscape of Enterprise AI Adoption
Artificial intelligence, once a niche technology, is rapidly becoming a foundational component across large enterprises. Its pervasive integration is not only reshaping business operations but also fundamentally altering the responsibilities of key technical teams. New research, focusing on the state of Site Reliability Engineering (SRE) and platform engineering, reveals that these teams are increasingly burdened with the complex task of ensuring AI systems are reliable, scalable, and trustworthy. The study, conducted by Dynatrace, surveyed nearly a thousand IT leaders and underscores a critical shift: the operational complexities introduced by AI workloads are forcing organizations to re-evaluate their approaches to scale, automation, and control. This evolution means that SRE and platform engineering professionals are at the forefront, tasked with developing new benchmarks, tools, and capabilities specifically tailored for AI, integrating them seamlessly into existing reliability and development frameworks.
Indeed, the importance of these roles is growing exponentially. Industry analysts project a substantial increase in SRE practice adoption, with an anticipated 80% of enterprises embracing these methodologies by 2028, a significant jump from just 30% in 2024. With strong executive backing and a shared sense of ownership, SRE and platform engineering teams now bear considerable accountability for the success or failure of AI initiatives. Their mandate extends to continually evolving platforms, refining tooling, and establishing new standards to accommodate the demands of advanced AI.
Bridging the AI Development and Operations Divide
One of the most pressing challenges emerging from the widespread adoption of AI is the gap between its development and its operational deployment. The study highlights that monitoring AI models has become the primary concern for 67% of SREs. Furthermore, a substantial 58% of SREs already utilize AI-powered capabilities specifically for evaluating model performance and accuracy. This surge in demand for robust AI evaluation mechanisms is outpacing the current suite of tools available to manage such sophisticated requirements. Despite the promise of AI to enhance efficiency, its implementation is currently falling short in key areas such as cost reduction and shortening the mean time to resolution for incidents.
A significant barrier reported by over a third of platform engineers is the difficulty of integrating various tools. This fragmentation creates inefficiencies and complicates the seamless flow of data and insights between AI development and operational environments. Recognizing this critical need, Dynatrace recently announced its intention to acquire Arize. This strategic move aims to integrate AI-native evaluation directly into observability platforms, creating a unified data source for both AI model builders and operational teams. This integration seeks to eliminate the need for disparate systems, fostering a more cohesive and effective approach to managing AI in production.
The Ascendancy of Scale as an Enterprise Hurdle
The research unequivocally demonstrates the entrenched status of SRE and platform engineering across large enterprises. A striking 92% of organizations report that their SRE initiatives enjoy robust support from executive leadership. Similarly, among organizations employing platform engineering, 89% have successfully implemented an internal developer platform, with 60% reporting its widespread adoption across various departments. These figures indicate a significant enterprise-wide investment in enhancing reliability, streamlining automation, and boosting developer productivity. The current phase demands that SRE and platform engineering teams leverage this strong foundation to navigate the next wave of digital transformation, particularly as AI workloads increasingly become part of core production infrastructure.
However, this integration is not without its difficulties. The increasing complexity of AI systems, the emergence of new forms of telemetry data, and novel ways in which AI systems can fail are placing unprecedented demands on observability frameworks. Simply put, AI is raising the bar for what constitutes robust reliability and effective oversight within enterprise operations.
Elevating Reliability and Oversight for AI
The introduction of agentic AI, which involves autonomous AI systems, is driving a new set of priorities and challenges for technical teams. For SREs, 89% are now employing service-level objectives (SLOs) across at least some of their teams or systems. As mentioned, monitoring AI models for performance and drift is now their top use case. For platform engineers, a significant 55% prioritize equipping developers with AI-powered tools, such as coding copilots and intelligent chatbots. While AI technologies are largely meeting expectations for improving overall reliability and developer productivity, their impact on cost reduction and reducing mean time to resolution has been less pronounced than anticipated.
This discrepancy highlights a critical need for more sophisticated system-level intelligence and enhanced workflow orchestration, moving beyond merely incorporating AI tools into existing environments. Nearly half of SRE respondents indicated that an excessive number of data sources and metrics impedes their ability to define and manage effective SLOs. Significantly, teams are deliberately prioritizing enhanced visibility and human oversight before expanding automation. This cautious approach is reflected in the fact that monitoring AI systems for model performance and accuracy remains the most common AI-powered capability utilized by SREs.
Observability: The Core of AI-Driven Operations
As enterprises increasingly move towards greater automation and AI-assisted operations, observability is solidifying its role as a fundamental pillar for AI governance, reliability, and optimization within SRE and platform engineering. However, significant obstacles persist, particularly concerning integration, complexity, and fragmented data. More than a third, specifically 37%, of platform engineers identify the integration of new tools with existing systems as their primary challenge. Furthermore, only 40% of platform engineers report that observability is fully embedded across all deployment stages, indicating a significant gap in comprehensive monitoring.
In a notable development, half of all SREs are now utilizing AI-powered capabilities for automated incident response. This trend signals a clear shift towards agentic operations, where observability must serve as the central control plane, dictating when and how autonomous actions are initiated. As Steve Tack, Chief Product Officer at Dynatrace, commented, SRE and platform engineering initially established the groundwork for modern digital reliability. However, AI is fundamentally changing the rules of engagement. Enterprises must transition from merely managing systems to expertly orchestrating them, seamlessly connecting observability, automation, and agentic AI. This integrated approach is essential to operate at the rapid pace demanded by current initiatives, effectively transforming insights into scalable actions. He also noted that this research further validates the company's decision to acquire Arize, emphasizing that the traditional separation between AI engineering teams' evaluation tools and operations teams' monitoring systems is no longer sustainable as AI becomes deeply embedded in enterprise production, according to the Financial Times.
Why it matters
For those in infrastructure, telecommunications, and data center operations, the escalating demands on SRE and platform engineering teams highlighted by this research are directly relevant. As AI workloads become integral to enterprise applications, the underlying infrastructure must adapt to support their unique characteristics, such as unpredictable resource consumption, complex data flows, and novel failure modes. Ensuring the reliability and performance of AI systems necessitates advanced monitoring and observability across physical and virtual infrastructure. This impacts network bandwidth, compute capacity, storage solutions, and energy management in data centers. Telecommunications providers, in particular, face pressure to deliver low-latency, high-throughput networks capable of supporting distributed AI inference and training. The ability to effectively orchestrate, automate, and observe these intricate systems directly affects service uptime, operational efficiency, and the overall success of enterprise AI deployments, fundamentally redefining the skill sets and tools required by in-field technicians and operational staff to maintain critical digital infrastructure. The integration of AI into operational pipelines promises efficiency but demands a robust and adaptable foundational infrastructure to succeed, placing the spotlight on the teams responsible for its continuous reliability and performance. This shift necessitates a re-evaluation of current practices, emphasizing proactive monitoring, automated remediation, and intelligent resource allocation to prevent breaking points as AI scales across the enterprise landscape.
More from Trends
RSS
AI Deciphers Crucial DNA 'On Switch,' Advancing Genetic Understanding
Researchers harnessed AI to identify the precise DNA sequence that acts as an 'on switch' for roughly 60% of human genes, a breakthrough with implications for disease prediction and synthetic biology.

Access Denied: Source Article Behind Paywall
The requested source article, 'Bill Gates calls for 'human reserved' jobs to protect labour force from AI' from the Financial Times, is currently behind a paywall and inaccessible.

Inaccessible Source Content
The provided source content is behind a paywall, preventing access to the article for rewriting.