
The Three Root Causes Behind Most Performance Failures - Rao Dhaligadoo
transcript
show notes
Slow systems don't just annoy users, they cost companies real money, and most teams only find out when the phone starts ringing. With Rao Dhaligadoo I talk about what it actually takes to catch performance problems before customers do. We get into how a 70 percent CPU threshold can trigger an automated chain of monitoring, ticket creation, and root cause analysis, and why that matters more than it sounds. Rao also shares what it feels like to spend 45 hours in a war room at a bank, watching a CIO walk in at 1 a.m. while the whole team sits in silence, and how that experience shaped the way he thinks about proactive testing.
"In a perfect world, they never see the loading spinner or the 404." - Rao Dhaligadoo
Vasudev Rao Dhaligadoo is a Quality Engineering Lead with over 10 years of experience driving test management, automation strategy, and QA capability across enterprise systems in banking, payroll, and large-scale platforms. He has led end-to-end testing initiatives, built scalable automation frameworks, and integrated quality practices into CI/CD pipelines to improve release confidence and delivery speed.
He is also a QA and Test Automation trainer, regularly mentoring engineers and delivering hands-on training on modern testing tools and practices. Passionate about advancing quality engineering, he focuses on combining test management, observability, automation, and AI-assisted insights to diagnose complex system issues and strengthen software reliability.
Highlights:
- A 70% CPU or resource threshold triggers automated alerts before systems degrade to failure, keeping end users away from slowdowns and 404 errors entirely.
- Slow systems cost real money even without full outages: a combined YouTube and Azure slowdown incident cost more than 70 million US dollars.
- AI-powered monitoring tools like Datadog's Bits AI pinpoint the exact database query or API endpoint causing a performance issue, cutting analysis time from hours to minutes.
- Database query optimization and load balancer configuration under high traffic are the most common root causes of production performance problems, based on Rao Dhaligadoo's field experience.
- Reactive incident response without proactive monitoring forces teams into war-room situations, with one real case stretching to 45 hours before production was restored.
📌 Testing is a people business, humanity as a superpower. That is my talk on October 7 at HUSTEF 2026 in Budapest: See programme and tickets





