Site Reliability Engineer

Location: Charlotte (NC) | Employment Type: Full Time - 40 hours per week | Job Level: T3 | Work Preference: Onsite at work location address | Job Code: 245605

Job Description:

Senior Site Reliability Engineer (SRE) – PagerDuty / Moogsoft Migration

Position Overview

Lead a critical observability and incident-management transformation by migrating from the legacy Moogsoft AIOps platform to PagerDuty. This role focuses on redesigning incident management workflows, implementing event orchestration, integrating monitoring platforms, and enabling a seamless transition to PagerDuty for live production monitoring.

What You'll Do / Key Responsibilities

  • Lead the migration from Moogsoft and implement PagerDuty as the primary incident intelligence and alerting platform.

  • Connect existing monitoring tools, including Datadog, New Relic, Splunk, and AWS CloudWatch, directly into PagerDuty.

  • Configure PagerDuty Event Orchestration, deduplication rules, and alert suppression to minimize alert fatigue.

  • Build automated incident response workflows, on-call schedules, escalation policies, and bidirectional ChatOps integrations with Slack and Teams.

  • Create runbooks for the new system.

  • Train engineering teams on PagerDuty best practices.

  • Support a seamless transition with zero downtime to live production monitoring.

Required Qualifications

  • Deep, hands-on experience with PagerDuty, including architectural design, event routing, and advanced configurations.

  • Familiarity with Moogsoft, including clustering, Situation Room, and ingestion logic.

  • Strong understanding of how monitoring tools feed into incident management platforms.

  • Experience working with observability/monitoring tools such as Datadog, New Relic, Splunk, and AWS CloudWatch.

  • Proficiency in Python, Bash, or Go for utilizing PagerDuty APIs and developing custom integrations.

Preferred Qualifications

  • Experience managing PagerDuty configurations using Terraform / Infrastructure as Code (IaC).

What Makes HTC A Great Place To Build Your Future

HTC Global Services wants you to join our team. Come build new things with us and advance your career. At HTC Global, you’ll collaborate with experts, work alongside clients, and be part of high-performing teams driving success together. You’ll have long-term opportunities to grow your career and develop skills in the latest emerging technologies.

At HTC Global Services, our employees have access to a comprehensive benefits package. Benefits can include Group Health (Medical, Dental, and Vision), Paid Time Off, Paid Holidays, 401(k) matching, Group Life and Disability insurance, Professional Development opportunities, Wellness programs, and a variety of other perks.

Our success as a company is built on inclusion and diversity. HTC Global Services is committed to providing a workplace free from discrimination and harassment, where every employee is treated with dignity and respect. We celebrate differences and believe that diverse cultures, perspectives, and skills drive innovation and success. HTC is an Equal Opportunity Employer and a proud National Minority Supplier. We seek to empower each individual, fostering an environment where everyone feels valued, included, and respected.

#SRE #SiteReliabilityEngineer #DevOps #DevOpsEngineer #PagerDuty #Moogsoft #AIOps 

At HTC Global Services, our culture is an embodiment of who we are – a value-led organization committed to success of our people and customers.

  • Hybrid and Workplace flexibility
  • Work-Life-Balance
  • Well-defined career development plan
  • Rewards & Recognition program
  • L&D focuses on upskilling
  • Hands-on experience on Emerging Technologies and Digital Transformation
  • Career Mobility programs

Join Our Talent Community

Tell us about yourself, and we will keep you informed about opportunities that match your interests.

Register