Machine Learning Authors: Pat Romanski, William Schmarzo, Yeshim Deniz, Stackify Blog, Jason Bloomberg

Related Topics: Java IoT, Microservices Expo, Microsoft Cloud, Machine Learning , Agile Computing, @DXWorldExpo

Java IoT: Article

Integrated Load Test Analysis

Using Compuware APM Web Load Test and PureStack Technology

Andreas Grabner described how he used the Compuware APM PureStack technology to identify the server-side performance issues during a recent load test run against the Compuware APM Community Portal, a production application used by our customers. He was able to quickly identify the CPU bottleneck that caused the performance degradation in the server environment, leading to an almost immediate resolution of the issue.

Bridging the Gap between Ops and Apps Data by adding Context: One picture that shows the Hotspots of "Horizontal" Transaction as well as the "Vertical" Stack.

But what about the external performance recorded during this load test? What would a customer have experienced if they had tried to access the site during this time? Well, at the peak of the test, I used WebPageTest to capture a video of the APM Community Homepage loading (Note: The video has been advanced to 50 seconds already).

The external performance degraded badly at the peak of the test - this isn't a surprise given what Andreas already pointed out. But how can the person running the external load - in this case, me, using the Compuware APM Web Load Testing service - make use of the data captured from outside the firewall and the rich data set covering system/infrastructure health and its effect on user experience and application performance available from the Compuware APM PureStack Technology? This post will show how I used a subset of the PureStack data to build charts that helped correlate key events on the server side to performance events in the Web Load Test (WLT) data.

I always like to start with the "Why?" of a load test. The goal of this load test was to determine if the APM Community Portal could handle a substantial increase in traffic as it had just been designated as the central hub for product documentation and customer discussions. To be absolutely sure, the APM Community Portal team wanted to determine if the application could support up to 200 concurrent visitors, an increase of nearly 10X from its current peak traffic.

For WLT to achieve this load volume is easy. Doing it in a controlled way meant that we needed to come up with a plan that effectively tested the application, but provided critical information at all stages of the test. The Portal team wanted a load test that ramped up to a maximum of 200 virtual users (VUs) over the course of 2 hours, with load distributed around the globe. This slow ramping of the load would help diagnose critical performance issues in a controlled fashion, as performance events can be directly tied to the amount of load and the activities occurring on the server at that time.

Test Ramping to 200 VUs used in the April 14 2013 APM Community Portal Load Test

In addition to ramping the load, the global distribution of load generation and traffic types had to be determined. Not all of the virtual users (VUs) would be executing the same test script - four test scripts were created, with each testing a core part of the infrastructure. The Portal team decided on the load and test script distribution, and required only some very small adjustments before the configuration was finalized.

Compuware APM Web Load Test Global Traffic and Script Distribution for APM Community Test Execution - April 14 2013

The test was run on a Sunday morning when traffic and customer impact would be low, which turned out to be a good thing. As load began to increase, performance began to degrade dramatically before the halfway point of the test. Transaction response times began to skyrocket.

Response Times Increasing as Load Increases until the site becomes so slow that it appears unresponsive to visitors

Transaction response times degraded right up to 09:49 EDT when the system began reporting a nearly 100% error rate. Most of the performance analysis here is focused on the time between 08:10 and 09:49 EDT.

Using the PureStack Technology, Andreas has detailed the server-side diagnostic process he went through to diagnose the server-side performance effects. The data captured inside the firewall aligns perfectly with the external data. By comparing the external response time of the transactions to the time required for the server PurePaths, a very clear and direct correlation between the amount of load on the system, the effect on server processing times, and the degradation in performance experienced by customers can be drawn.

A comparative chart that shows Web Load Test Average Transaction Response Time v. VUs v. Average Server PurePath Time

One item not discussed in the previous assessment was the performance event that was detected between 08:50 and 08:55 EDT. During that period, both external transaction and server PurePath response times increased noticeably. As it stands alone in the load test, it was clear that the causes of this event were different than those that eventually caused the overall failure of the system.

By aligning the total transactions per minute being executed by the load test system to the percentage of CPU being consumed at the web server layer, the cause of the 08:50-08:55 EDT spike becomes clear: something at the web server was suddenly consuming 100% of the total available CPU. This had the effect of decreasing the number of transactions that were processed, and caused the WLT response times and PurePath times to increase.

A comparative graph showing Transactions per Minute v. VUs v. CPU Percentage for Web Server - April 14 2013 Load Test

While the eventual failure of the test can also be related to CPU exhaustion, this anomalous event seems completely unrelated to the volume of traffic occurring at that time. The timing indicated that a scheduled job that occurs either daily or hourly was the cause of the spike. Finding these scheduled jobs that may be either undetected or forgotten by system administrators is not unusual during load tests. Digging deeper into the system found that the Atlassian/Confluence application layer, the software that controls much of the core functionality of APM Community, spiked almost exactly in the middle of the recorded issue, indicating that the job was related to something in this layer.

Atlassian Execution CPU Time during the April 14 2013 Load Test

What makes the integrated approach to load testing critical to those of us who have only had access to the external, Web Load Test data in the past is that we can immediately draw correlations between events inside the datacenter and the performance effects we are capturing outside the firewall. By integrating a few key Web Load Test metrics (Average Response Time, Transactions per Minute, and Total VUs) with select PureStack metrics (Number of Confluence Requests in the last 10 seconds and CPU percentages), the team was quickly able to have in-depth information available to them throughout the load test. Finding this high load job was a bonus of the load test, which clearly pointed out that the system was undersized for the load that the Portal team was expecting. But this conclusion could only be found by correlating multiple layers of data into a coherent whole that provided the team with the information they needed to identify critical issues.

The chart below shows how this would appear to someone monitoring the load test.

Comparative Web Load Test and PureStack Metrics - April 14 2013

In one chart, multiple critical metrics are available to identify potential problem hotspots. For example, while the ultimate application bottleneck is a critical issue to resolve, without the correlating data the event between 08:50 and 08:55 EDT may have been overlooked, leaving the Portal team with a potential user experience problem that could surface at a later date.

With all of this data available to teams running load tests, it is recommended that care be taken not to drown them in a flood of data. Here, we took six key metrics and were easily able to show that the issue was a bottleneck at the web server CPU as traffic increased. These metrics were:

  1. WLT Response Time
  2. WLT Transactions per Minute
  3. Server side PurePath time
  4. CPU percentage on the web server
  5. Number of requests to the Confluence application layer
  6. The number of VUs deployed at each minute

Choosing five to six key metrics is the most critical element in this process. These metrics should be able to directly indicate problem areas or point the load test team in the right direction to begin to resolve the issue. For example, the sudden decrease in requests to Confluence during the 08:50-08:55 EDT period may not give you the root cause, but immediately posed the question "Why is this component suddenly showing signs of degradation?"

Another perspective would be to add in database statistics, as the database layer is often the cause of performance issues under heavy load. What is interesting in this case is that an amalgamated view of the load test data shows completely the opposite - when response times and CPU % begin to spike, the number of database queries and the total time spent at the database layer decreases dramatically.

Another integrated view that includes database metrics, showing that the database is likely not an issue in this test.

This last chart provides the team with a key metric: At 09:05 EDT and 90 VUs, the application layer became so congested that it effectively stopped passing requests through to the database. At the same time, WLT response times crossed 20 seconds and the CPU percentage crossed 90%. With this integrated view, the Portal team now has a very clear picture of the end-to-end application and its effect on customers.

We showed two potential methods for PureStack and Web Load Test metrics to produce a complete picture of a load test. Your key metrics may not be the same as ours, and may include number of bytes in and out, Disk I/O, memory usage, total Web requests, third-party performance, or other metrics that are meaningful to your application. But with the PureStack Technology, integrating any of these datapoints directly with the Compuware APM Web Load Testing service becomes easy. PureStack allows you to link the external performance of the application under load to the server-side effects on key components, building a complete end-to-end model of performance for your application during load testing events.

More Stories By Stephen Pierzchala

With more than a decade in the web performance industry, Stephen Pierzchala has advised many organizations, from Fortune 500 to startups, in how to improve the performance of their web applications by helping them develop and evolve the unique speed, conversion, and customer experience metrics necessary to effectively measure, manage, and evolve online web and mobile applications that improve performance and increase revenue. Working on projects for top companies in the online retail, financial services, content delivery, ad-delivery, and enterprise software industries, he has developed new approaches to web performance data analysis. Stephen has led web performance methodology, CDN Assessment, SaaS load testing, technical troubleshooting, and performance assessments, demonstrating the value of the web performance. He noted for his technical analyses and knowledge of Web performance from the outside-in.

Comments (0)

Share your thoughts on this story.

Add your comment
You must be signed in to add a comment. Sign-in | Register

In accordance with our Comment Policy, we encourage comments that are on topic, relevant and to-the-point. We will remove comments that include profanity, personal attacks, racial slurs, threats of violence, or other inappropriate material that violates our Terms and Conditions, and will block users who make repeated violations. We ask all readers to expect diversity of opinion and to treat one another with dignity and respect.

@CloudExpo Stories
Business professionals no longer wonder if they'll migrate to the cloud; it's now a matter of when. The cloud environment has proved to be a major force in transitioning to an agile business model that enables quick decisions and fast implementation that solidify customer relationships. And when the cloud is combined with the power of cognitive computing, it drives innovation and transformation that achieves astounding competitive advantage.
Poor data quality and analytics drive down business value. In fact, Gartner estimated that the average financial impact of poor data quality on organizations is $9.7 million per year. But bad data is much more than a cost center. By eroding trust in information, analytics and the business decisions based on these, it is a serious impediment to digital transformation.
Digital Transformation: Preparing Cloud & IoT Security for the Age of Artificial Intelligence. As automation and artificial intelligence (AI) power solution development and delivery, many businesses need to build backend cloud capabilities. Well-poised organizations, marketing smart devices with AI and BlockChain capabilities prepare to refine compliance and regulatory capabilities in 2018. Volumes of health, financial, technical and privacy data, along with tightening compliance requirements by...
Andrew Keys is Co-Founder of ConsenSys Enterprise. He comes to ConsenSys Enterprise with capital markets, technology and entrepreneurial experience. Previously, he worked for UBS investment bank in equities analysis. Later, he was responsible for the creation and distribution of life settlement products to hedge funds and investment banks. After, he co-founded a revenue cycle management company where he learned about Bitcoin and eventually Ethereal. Andrew's role at ConsenSys Enterprise is a mul...
"NetApp is known as a data management leader but we do a lot more than just data management on-prem with the data centers of our customers. We're also big in the hybrid cloud," explained Wes Talbert, Principal Architect at NetApp, in this SYS-CON.tv interview at 21st Cloud Expo, held Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA.
"Since we launched LinuxONE we learned a lot from our customers. More than anything what they responded to were some very unique security capabilities that we have," explained Mark Figley, Director of LinuxONE Offerings at IBM, in this SYS-CON.tv interview at 21st Cloud Expo, held Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA.
DXWordEXPO New York 2018, colocated with CloudEXPO New York 2018 will be held November 11-13, 2018, in New York City and will bring together Cloud Computing, FinTech and Blockchain, Digital Transformation, Big Data, Internet of Things, DevOps, AI, Machine Learning and WebRTC to one location.
DXWorldEXPO LLC announced today that "Miami Blockchain Event by FinTechEXPO" has announced that its Call for Papers is now open. The two-day event will present 20 top Blockchain experts. All speaking inquiries which covers the following information can be submitted by email to [email protected] Financial enterprises in New York City, London, Singapore, and other world financial capitals are embracing a new generation of smart, automated FinTech that eliminates many cumbersome, slow, and expe...
Evan Kirstel is an internationally recognized thought leader and social media influencer in IoT (#1 in 2017), Cloud, Data Security (2016), Health Tech (#9 in 2017), Digital Health (#6 in 2016), B2B Marketing (#5 in 2015), AI, Smart Home, Digital (2017), IIoT (#1 in 2017) and Telecom/Wireless/5G. His connections are a "Who's Who" in these technologies, He is in the top 10 most mentioned/re-tweeted by CMOs and CIOs (2016) and have been recently named 5th most influential B2B marketeer in the US. H...
DXWorldEXPO | CloudEXPO are the world's most influential, independent events where Cloud Computing was coined and where technology buyers and vendors meet to experience and discuss the big picture of Digital Transformation and all of the strategies, tactics, and tools they need to realize their goals. Sponsors of DXWorldEXPO | CloudEXPO benefit from unmatched branding, profile building and lead generation opportunities.
The best way to leverage your Cloud Expo presence as a sponsor and exhibitor is to plan your news announcements around our events. The press covering Cloud Expo and @ThingsExpo will have access to these releases and will amplify your news announcements. More than two dozen Cloud companies either set deals at our shows or have announced their mergers and acquisitions at Cloud Expo. Product announcements during our show provide your company with the most reach through our targeted audiences.
DevOpsSummit New York 2018, colocated with CloudEXPO | DXWorldEXPO New York 2018 will be held November 11-13, 2018, in New York City. Digital Transformation (DX) is a major focus with the introduction of DXWorldEXPO within the program. Successful transformation requires a laser focus on being data-driven and on using all the tools available that enable transformation if they plan to survive over the long term. A total of 88% of Fortune 500 companies from a generation ago are now out of bus...
With 10 simultaneous tracks, keynotes, general sessions and targeted breakout classes, @CloudEXPO and DXWorldEXPO are two of the most important technology events of the year. Since its launch over eight years ago, @CloudEXPO and DXWorldEXPO have presented a rock star faculty as well as showcased hundreds of sponsors and exhibitors! In this blog post, we provide 7 tips on how, as part of our world-class faculty, you can deliver one of the most popular sessions at our events. But before reading...
Modern software design has fundamentally changed how we manage applications, causing many to turn to containers as the new virtual machine for resource management. As container adoption grows beyond stateless applications to stateful workloads, the need for persistent storage is foundational - something customers routinely cite as a top pain point. In his session at @DevOpsSummit at 21st Cloud Expo, Bill Borsari, Head of Systems Engineering at Datera, explored how organizations can reap the bene...
Cloud Expo | DXWorld Expo have announced the conference tracks for Cloud Expo 2018. Cloud Expo will be held June 5-7, 2018, at the Javits Center in New York City, and November 6-8, 2018, at the Santa Clara Convention Center, Santa Clara, CA. Digital Transformation (DX) is a major focus with the introduction of DX Expo within the program. Successful transformation requires a laser focus on being data-driven and on using all the tools available that enable transformation if they plan to survive ov...
As you move to the cloud, your network should be efficient, secure, and easy to manage. An enterprise adopting a hybrid or public cloud needs systems and tools that provide: Agility: ability to deliver applications and services faster, even in complex hybrid environments Easier manageability: enable reliable connectivity with complete oversight as the data center network evolves Greater efficiency: eliminate wasted effort while reducing errors and optimize asset utilization Security: implemen...
@DevOpsSummit New York 2018, colocated with CloudEXPO | DXWorldEXPO New York 2018 will be held November 11-13, 2018, in New York City. From showcase success stories from early adopters and web-scale businesses, DevOps is expanding to organizations of all sizes, including the world's largest enterprises - and delivering real results.
With tough new regulations coming to Europe on data privacy in May 2018, Calligo will explain why in reality the effect is global and transforms how you consider critical data. EU GDPR fundamentally rewrites the rules for cloud, Big Data and IoT. In his session at 21st Cloud Expo, Adam Ryan, Vice President and General Manager EMEA at Calligo, examined the regulations and provided insight on how it affects technology, challenges the established rules and will usher in new levels of diligence arou...
The dynamic nature of the cloud means that change is a constant when it comes to modern cloud-based infrastructure. Delivering modern applications to end users, therefore, is a constantly shifting challenge. Delivery automation helps IT Ops teams ensure that apps are providing an optimal end user experience over hybrid-cloud and multi-cloud environments, no matter what the current state of the infrastructure is. To employ a delivery automation strategy that reflects your business rules, making r...
"We started a Master of Science in business analytics - that's the hot topic. We serve the business community around San Francisco so we educate the working professionals and this is where they all want to be," explained Judy Lee, Associate Professor and Department Chair at Golden Gate University, in this SYS-CON.tv interview at 21st Cloud Expo, held Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA.