Insight Global logo

SRE - PERM REMOTE

Insight Global

Dunwoody, GARemoteJobNo compensation foundPosted 1mo agoSeen in employer's feed 5 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
No compensation found
Location
Dunwoody, GARemote
Work Authorization
US work authorization required

Requirements

Credentials this posting asks for.

Bachelor's degree

Job overview

Insight Global is hiring a SRE - PERM REMOTE. This Site Reliability Engineer (SRE) / DevOps Engineer will support and scale a Microsoft Azure environment. The role involves blending infrastructure, automation, deployment support, and reliability engineering to ensure highly available, scalable systems. This individual will improve observability, monitor system health, diagnose issues, and proactively prevent outages, evolving into a balanced SRE/DevOps role with increased ownership of reliability initiatives, capacity planning, root cause analysis, and system performance optimization.

Key focus areas include Support and maintain cloud infrastructure within Microsoft Azure, Build and manage CI/CD pipelines and deployment automation, and Monitor application and system health using Azure observability tools.

Successful candidates bring 2-5 Years DevOps Engineer Experience, Microsoft Azure Environments Experience, and .NET Ecosystem Support Experience. Important skills include Azure, CI/CD Pipelines, Deployment Automation, Azure Observability Tools, Log Analysis, and Incident Diagnosis. Preferred (not required): Implementing Observability Solutions.

Skills & qualifications

RequiredNice to have

Skills

AzureCI/CD PipelinesDeployment AutomationAzure Observability ToolsLog AnalysisIncident DiagnosisTroubleshooting Production IssuesSystem ObservabilityApplication ReliabilityApplication PerformanceInfrastructure as CodePreventative MaintenanceSystem OptimizationCapacity PlanningProactive Reliability EngineeringOperational ExcellenceProblem-SolvingSystems-Thinking.NET EcosystemAzure DevOpsAzure MonitorApplication InsightsLog AnalyticsAzure CLIAzure Container AppsYAML-Based Pipeline DeploymentsBicepGit RepositoriesSource Control Best PracticesSystem Health MonitoringObservability MindsetKusto Query Language (KQL)ARM TemplatesAdditional Azure Platform ExpertiseImplementing Observability SolutionsManaging Observability SolutionsRoot Cause AnalysisPost-Incident ReviewSRE InitiativesReliability Programs

Qualifications

2-5 Years DevOps Engineer ExperienceBachelor's Degree in Computer ScienceBachelor's Degree in EngineeringBachelor's Degree in Information SystemsBachelor's Degree in Related Field

Full job description

Job Description

We are seeking a Site Reliability Engineer (SRE) / DevOps Engineer to join a growing team responsible for supporting and scaling a Microsoft Azure environment. This role is ideal for an engineer who enjoys blending infrastructure, automation, deployment support, and reliability engineering to ensure highly available, scalable systems.

As the organization continues to grow its customer base, this individual will play a key role in improving observability, monitoring system health, diagnosing issues, and proactively preventing outages before they occur. While the position currently leans more heavily toward DevOps and infrastructure support, it will evolve into a balanced SRE/DevOps role with increased ownership of reliability initiatives, capacity planning, root cause analysis, and system performance optimization.

Responsibilities

Support and maintain cloud infrastructure within Microsoft Azure

Build and manage CI/CD pipelines and deployment automation

Monitor application and system health using Azure observability tools

Analyze logs, diagnose incidents, and troubleshoot production issues

Improve monitoring, alerting, and overall system observability

Partner with engineering teams to improve application reliability and performance

Implement Infrastructure-as-Code (IaC) solutions

Participate in preventative maintenance, system optimization, and capacity planning efforts

Contribute to a culture of proactive reliability engineering and operational excellence

We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to [email protected] learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.

Skills and Requirements

Experience

2–5 years of experience in a DevOps Engineer, Site Reliability Engineer (SRE), Cloud Engineer, or related role

Experience working within Microsoft Azure environments

Strong troubleshooting, problem-solving, and systems-thinking abilities

Experience supporting applications in a .NET ecosystem

Azure & Cloud Technologies

Azure DevOps

Azure Monitor

Application Insights

Log Analytics

Azure CLI

Azure Container Apps

Infrastructure & Automation

YAML-based pipeline deployments

Bicep

Git repositories and source control best practices

Infrastructure-as-Code (IaC) experience

Reliability & Operations

Log analysis and troubleshooting

Incident diagnosis and resolution

System health monitoring

Strong observability mindset Kusto Query Language (KQL)

ARM Templates

Additional Azure platform expertise

Experience implementing or managing observability solutions

Capacity planning experience

Root cause analysis and post-incident review experience

Previous ownership of SRE initiatives or reliability programs

Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field

You've read the whole posting — now see how you match it.