Insight Global logo

Overnight SQL Production Support DBA (Event Management / Incident Management / Production Support)

Insight Global

Sandy Springs, GA · FlexibleJobSeen 1w agoSeen in employer's feed 3 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
Sandy Springs, GAFlexible
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

The overnight SQL Production DBA monitors batch jobs, blocking, and performance, while the after-hours Azure Infrastructure Monitor watches the company’s e-commerce website and Azure data warehouse. The role triages alerts, performs authorized first-line actions, and escalates incidents to on-call personnel. Coverage is especially important during the Thanksgiving–Christmas holiday season, when after-hours traffic and batch/ETL load are highest.

Skills & qualifications

RequiredNice to have

Skills

Microsoft AzureAzure MonitorLog AnalyticsApplication InsightsAzure AlertsWeb Application InfrastructureLoad BalancersNetworking BasicsAzure SQL DatabaseAzure Synapse AnalyticsPerformance Metrics AnalysisWritten CommunicationIndependent JudgmentVigilanceAttention to DetailCalm Under PressureMethodical ApproachReliabilitySelf-DirectionRetail Platform SupportE-Commerce Platform SupportServiceNowOn-Call Paging PlatformsETL MonitoringAzure Data FactoryPowerShellAzure CLI

Qualifications

2+ Years IT Operations ExperienceMicrosoft Azure Certification or EquivalentOvernight Weekend and Holiday Availability

Full job description

Job Description

The overnight role is a SQL Production DBA – someone monitoring batch jobs for failure including watching for blocking and some minor performance tuning.

The After-Hours Azure Infrastructure Monitor provides real-time, hands-on-keyboard monitoring of The Company’s Azure-hosted environment outside standard business hours, covering both the public e-commerce website and the Azure data warehouse. This role watches dashboards and alerts as they happen, recognizes early signs of performance degradation or outages, performs first-line triage, and escalates to the appropriate on-call engineer, DBA, or manager with enough detail to act immediately. Coverage is especially critical during The Company’s seasonal order peaks (the Thanksgiving–Christmas holiday season), when after-hours traffic and batch/ETL load are highest. Key Responsibilities

  • Actively monitor Azure Monitor, Log Analytics, Application Insights, and Azure Service Health dashboards for the company website and the Azure data warehouse throughout the shift.
  • Triage incoming alerts in real time, distinguishing routine fluctuations from genuine incidents (elevated error rates, latency spikes, failed health checks, throttling, capacity thresholds).
  • Identify performance bottlenecks such as high CPU/memory utilization, database DTU/vCore or storage pressure, slow-running queries, connection pool exhaustion, and App Service or VM health issues.
  • Monitor data warehouse job health, including failed or delayed ETL/pipeline runs, long-running queries, and missed load windows that could affect morning reporting.
  • Follow documented runbooks to perform authorized first-line actions (e.g., restarting a service, scaling an App Service plan, clearing a stuck queue) before escalating.
  • Escalate incidents to the correct on-call IT personnel (application engineering, database, network/infrastructure, or security) per the escalation matrix, with clear timestamps, affected systems, and symptoms.
  • Maintain a detailed incident log and prepare a shift-handoff summary for the day team.
  • Recognize and escalate security-relevant anomalies (unusual traffic patterns, repeated failed authentication attempts, unexpected access) to the security/on-call team.
  • Track incidents through to acknowledgment and, where applicable, confirm resolution before shift end.
  • Flag recurring or unresolved issues to the IT Operations Manager for follow-up during business hours.

Skills and Requirements

  • 2+ years of experience in IT operations, a network/security operations center (NOC/SOC), systems administration, or a similar monitoring-focused role.
  • Hands-on experience with Microsoft Azure, including Azure Monitor, Log Analytics, Application Insights, and Azure Alerts.
  • Working knowledge of web application infrastructure (App Service, load balancers/CDN, networking basics) and relational or data-warehouse platforms (Azure SQL Database, Azure Synapse Analytics, or comparable).
  • Ability to read performance metrics and dashboards (CPU, memory, latency, error rate, queue depth, throughput) and judge what warrants escalation versus routine variation.
  • Strong written communication skills for clear, actionable incident documentation and escalation.
  • Ability to work independently and exercise sound judgment while unsupervised during overnight hours.
  • Availability to work overnight, weekend, and holiday-season shifts as business needs require.

Key Competencies

  • Vigilant and detail-oriented — catches meaningful signals in a stream of routine metrics and alerts.

  • Calm and methodical under pressure during active incidents.

  • Clear, concise communicator, especially in writing and when escalating to engineers who were asleep five minutes ago.

  • Reliable and self-directed during unsupervised overnight shifts. Work Environment This role may be performed remotely or on-site depending on company policy and requires extended periods at a computer monitoring live dashboards, with the ability to respond to alerts promptly throughout the shift. Reliable internet connectivity (for remote coverage) and a distraction-free environment during working hours are required.

  • Microsoft Azure certification (e.g., AZ-104 Azure Administrator, AZ-500 Security Engineer, or equivalent).

  • Experience supporting retail or e-commerce platforms with pronounced seasonal traffic peaks.

  • Familiarity with ITSM/ticketing tools (e.g., ServiceNow) and on-call/paging platforms

  • Exposure to data pipeline/ETL monitoring (e.g., Azure Data Factory, Synapse pipelines).

  • Basic scripting ability (PowerShell or Azure CLI) to support first-line remediation and reporting.

We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal employment opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment without regard to race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or the recruiting process, please send a request to [email protected].

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.