infrastructure•5 min read

Keysight Unveils Advanced AI Infrastructure Validation at OCP Global Summit

Keysight Technologies is showcasing comprehensive validation solutions for cutting-edge AI infrastructure at the OCP Global Summit, addressing the complex demands of next-gen AI systems.

Abstract representation of data flowing through a complex network, symbolizing AI infrastructure validation.

Keysight Technologies is demonstrating new validation tools and methodologies crucial for building robust, scalable, and high-performing AI infrastructure, particularly focusing on network resilience and workload emulation at the OCP Global Summit.

The Imperative of Validating AI at Scale

The rapidly accelerating evolution of artificial intelligence, from foundational model training to widespread inferencing applications, places unprecedented demands on underlying computational and networking infrastructure. Building and deploying these sophisticated AI systems, often referred to as 'AI factories,' requires meticulous validation to ensure performance, reliability, and interoperability across a diverse ecosystem of hardware and software components. Keysight Technologies, a prominent player in test and measurement solutions, recently presented its latest advancements in AI infrastructure validation at the 2026 OCP Global Summit in San Jose. The company's demonstrations, conducted in collaboration with industry leaders such as Alibaba Cloud, Cisco, HPE, and NVIDIA, underscored the critical need for rigorous testing to simulate real-world AI workloads and network conditions before deployment.

Ram Periakaruppan, Vice President and General Manager of Network Test & Security Solutions at Keysight, emphasized the shift from validating individual components to verifying highly integrated systems. He highlighted that AI infrastructure now spans 'Scale-Up, Scale-Out, and Scale-Across' architectures, each presenting unique validation challenges. The goal of Keysight's solutions is to enable engineering teams to create realistic AI environments, identify potential bottlenecks early in the development cycle, and confidently transition designs into operational deployment.

Next-Generation Network Interconnects

A significant focus of Keysight's showcase was on validating emerging network technologies essential for future AI deployments. As the industry moves towards higher speeds and greater data density, standards like 224G SerDes and 1.6T Ethernet interconnects are becoming vital. Keysight unveiled two new 1.6T Ethernet test platforms: the SERT 1600GE data center switch and edge router test solution, and the AresONE Plus 1600GE platform. The AresONE Plus is notable for being the first in the industry to offer eight 1.6T Ethernet ports within a compact 2U appliance, capable of generating stateful, realistic AI workloads including a mix of collective operations.

Beyond raw speed, interoperability is a cornerstone of scalable AI infrastructure. Keysight demonstrated multi-vendor compatibility for emerging Ethernet Scale-Up technologies. This included a live public presentation of Link Layer Retry (LLR) functioning over 200G SerDes-based 1.6T speeds, a crucial step toward ensuring robust and error-free high-speed communication within AI clusters. Such validation ensures that different vendors' equipment can seamlessly integrate and perform optimally, preventing costly failures and performance degradation in large-scale AI environments.

Validating Across Distributed Architectures

Modern AI workloads are not confined to single data centers. They often span multiple facilities, creating 'Scale-Across' architectures that demand sophisticated validation. Keysight showcased its capabilities in assessing the resiliency and performance of RoCEv2 (RDMA over Converged Ethernet) AI workloads across these distributed networks. Engineers can use these tools to evaluate how factors like latency, congestion, buffering, and packet loss affect workload performance when AI infrastructure extends across data centers. This is particularly important for distributed training and inferencing, where consistent and low-latency communication is paramount.

Will Eatherton, Senior Vice President of Engineering for Data Center and Internet Infrastructure at Cisco, commented on the challenges of AI clusters expanding beyond single data center limits. He explained that distributed AI workloads necessitate immense bandwidth, high resiliency, and the capacity to handle traffic bursts without compromising application performance. Keysight's collaboration with Cisco allows for validation of these architectures under realistic RoCEv2 workloads, congestion scenarios, and diverse network conditions, ensuring that solutions like Cisco Silicon One P200 can deliver predictable performance.

Realistic Workload Emulation and Performance Monitoring

To truly understand how AI systems will behave in production, it is essential to emulate realistic workloads at scale. Keysight demonstrated its ability to simulate 64x800GE NIC and GPU endpoints driving a mix of AI collective workloads across a 2x2 CLOS fabric. This large-scale emulation allows developers to analyze fabric utilization, congestion patterns, load distribution, and overall workload performance under conditions that closely mirror real-world AI traffic. Such insights are invaluable for optimizing network designs and preventing performance bottlenecks before deploying extensive GPU infrastructure.

The company also highlighted its AI Inference Builder, a tool designed to generate realistic inference workloads against live inference stacks. This system correlates workload performance with critical metrics from compute, GPU, memory, and network telemetry. This holistic view provides a comprehensive understanding of an AI system's behavior, enabling engineers to pinpoint performance bottlenecks and ensure that inference models can be served efficiently at scale.

Another advanced technique showcased was Scale-Out validation using MRC (Multi-Path Routing and Congestion Management) traffic with SRv6 uSID-based source routing. This demonstrated resilient path and plane failover capabilities, congestion mitigation through packet trimming, and the maintenance of sustained application performance, all vital for robust and uninterrupted AI operations.

Keysight experts also participated in several speaking sessions at the OCP Summit, covering topics ranging from AI infrastructure validation at scale and digital twins for AI factory design to vendor-neutral benchmarking of switching ASICs and validation methodologies for the ESUN ecosystem, as noted by the Financial Times. These sessions further illustrated Keysight's commitment to advancing the understanding and implementation of high-performance AI infrastructure.

Why it matters

The sophisticated validation solutions offered by Keysight Technologies are fundamental for the successful deployment and operation of advanced AI and augmented reality systems. In the context of in-field AI, technicians often rely on robust and predictable network performance for real-time data processing and decision-making. For telco and data centre operations, the ability to emulate massive AI workloads, validate multi-vendor interoperability, and ensure network resilience across distributed architectures is critical for maintaining service quality, managing resources efficiently, and scaling infrastructure cost-effectively. These validation tools help preempt costly failures, optimize performance, and accelerate the development cycle for the next generation of AI-driven services and applications.

#ai infrastructure#network validation#data center#ocp#keysight#ethernet

More from Trends

RSS