
Site Reliability Engineer
Charlotte, NCJobSeen todaySeen in employer's feed today
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near Charlotte, NC, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Job overview
The Senior Site Reliability Engineer will lead a transformation from Moogsoft to PagerDuty for incident management and live production monitoring. The role redesigns incident workflows, implements event orchestration, integrates monitoring platforms, and supports a seamless transition. The engineer will also build incident-response workflows, create runbooks, and train engineering teams on PagerDuty best practices.
Skills & qualifications
Skills
Benefits
Full job description
Senior Site Reliability Engineer (SRE) – PagerDuty / Moogsoft Migration Position Overview
Lead a critical observability and incident-management transformation by migrating from the legacy Moogsoft AIOps platform to PagerDuty. This role focuses on redesigning incident management workflows, implementing event orchestration, integrating monitoring platforms, and enabling a seamless transition to PagerDuty for live production monitoring.
What You'll Do / Key Responsibilities
-
Lead the migration from Moogsoft and implement PagerDuty as the primary incident intelligence and alerting platform.
-
Connect existing monitoring tools, including Datadog, New Relic, Splunk, and AWS CloudWatch, directly into PagerDuty.
-
Configure PagerDuty Event Orchestration, deduplication rules, and alert suppression to minimize alert fatigue.
-
Build automated incident response workflows, on-call schedules, escalation policies, and bidirectional ChatOps integrations with Slack and Teams.
-
Create runbooks for the new system.
-
Train engineering teams on PagerDuty best practices.
-
Support a seamless transition with zero downtime to live production monitoring.
Required Qualifications
-
Deep, hands-on experience with PagerDuty, including architectural design, event routing, and advanced configurations.
-
Familiarity with Moogsoft, including clustering, Situation Room, and ingestion logic.
-
Strong understanding of how monitoring tools feed into incident management platforms.
-
Experience working with observability/monitoring tools such as Datadog, New Relic, Splunk, and AWS CloudWatch.
-
Proficiency in Python, Bash, or Go for utilizing PagerDuty APIs and developing custom integrations.
Preferred Qualifications
- Experience managing PagerDuty configurations using Terraform / Infrastructure as Code (IaC).
What Makes HTC A Great Place To Build Your Future
HTC Global Services wants you to join our team. Come build new things with us and advance your career. At HTC Global, you’ll collaborate with experts, work alongside clients, and be part of high-performing teams driving success together. You’ll have long-term opportunities to grow your career and develop skills in the latest emerging technologies.
At HTC Global Services, our employees have access to a comprehensive benefits package. Benefits can include Group Health (Medical, Dental, and Vision), Paid Time Off, Paid Holidays, 401(k) matching, Group Life and Disability insurance, Professional Development opportunities, Wellness programs, and a variety of other perks.
Our success as a company is built on inclusion and diversity. HTC Global Services is committed to providing a workplace free from discrimination and harassment, where every employee is treated with dignity and respect. We celebrate differences and believe that diverse cultures, perspectives, and skills drive innovation and success. HTC is an Equal Opportunity Employer and a proud National Minority Supplier. We seek to empower each individual, fostering an environment where everyone feels valued, included, and respected.
#SRE #SiteReliabilityEngineer #DevOps #DevOpsEngineer #PagerDuty #Moogsoft #AIOps
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Heavy Check EngineerPSA Airlines · Charlotte, NCPosted 2w agoPosted 2w ago
Principal Site Reliability Engineer (SRE)Ally · Charlotte, NC (Hybrid) · $110–180K/yr
CTIO - Site Reliability Engineer- Senior AssociatePwC · Charlotte, NC · $55–151K/yr
Principal Engineer, Platform, Cloud Engineering and DevOpsWells Fargo · CHARLOTTE, NC (Hybrid)
Principal - Cloud EngineerAlly · Remote · US · $110–180K/yr
You've read the whole posting — now see how you match it.