ICONMA, LLC logo

Google Cloud Site Reliability Engineer (GCP SRE)

ICONMA, LLC

Buffalo Grove, ILJobNo compensation foundTracked 1mo ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
No compensation found
Location
Buffalo Grove, IL
Work Authorization
Not specified

Requirements

Credentials this posting asks for.

EDP-SRE (Google Site Reliability Engineer)

Job overview

ICONMA, LLC is hiring a Google Cloud Site Reliability Engineer (GCP SRE). The Google Cloud Site Reliability Engineer (GCP SRE) will be responsible for incident detection, logging, and meeting agreed SLAs. This role involves bridge activation, communication, and postmortem preparation. Key duties include critical monitoring activities, problem management, and Grafana integration. The engineer will also participate in on-call rotations, handle incidents, and drive timely mitigation and recovery.

Key focus areas include Responsible for Incident Detection & Logging and meeting agreed SLA for incident tickets, Responsible for Bridge Activation & Communication (P1–P2), and Postmortem Preparation (Within 24–72 Hours) & Root Cause Analysis.

Successful candidates bring 8+ Years Site Reliability Engineering Experience. Important skills include Incident Detection & Logging, Bridge Activation & Communication, Postmortem Preparation, Root Cause Analysis, Critical Monitoring Activities, and Problem Management.

Skills & qualifications

RequiredNice to have

Skills

Incident Detection & LoggingBridge Activation & CommunicationPostmortem PreparationRoot Cause AnalysisCritical Monitoring ActivitiesProblem ManagementGrafana IntegrationOncall RotationsIncident HandlingTimely Mitigation and RecoveryAutomating Operational WorkOperating Highly Available SystemsOperating Low Latency SystemsOperating Secure SystemsDefining Reliability Through SLIs/SLOsDefining Reliability Through Error BudgetsBuild and Maintain ObservabilityMetricsLoggingTracesDashboardsAlerts for Critical ServicesTune AlertingReduce NoiseRapid Detection of User Impacting IssuesPost Incident ReviewsSupporting Production GCP EnvironmentsEnterprise Incident ManagementCloud OperationsBig QueryCloud StorageDataprocGoogle Kubernetes EngineAirflow/ComposerPub-SubCloud FunctionCloud SQLGitHubVisual Studio CodeMS-CopilotPrometheusSplunkClear Written CommunicationClear Verbal CommunicationCollaboration Across Multiple TeamsInfluence Engineering PracticesStrong CommunicationAnalytical SkillsIncident Management Life Cycle ProcessAgile Model Experience

Qualifications

8+ Years in Site Reliability Engineering8+ Years in DevOps8+ Years in Cloud EngineeringEDP-SRE (Google Site Reliability Engineer)10.00 Years of Experience

Benefits

Medical Insurance

Full job description

Our client, a IT Services and Consulting company, is looking for a Google Cloud Site Reliability Engineer (GCP SRE)”: for their Buffalo Grove, IL/Hybrid location. Responsibilities:

  • Responsible for Incident Detection & Logging and meeting agreed SLA for incident tickets.

  • Responsible for Bridge Activation & Communication (P1–P2).  

  • Postmortem Preparation (Within 24–72 Hours) & Root Cause Analysis.

  • Responsible for critical monitoring activities ,Problem Management & Grafana Integration.

  • Participate in oncall rotations, handle incidents, and drive timely mitigation and recovery.

  • Automating operational work so services can scale without manual toil also operating highly available, low latency & secure systems.

  • Defining and measuring reliability through SLIs/SLOs and error budgets.

  • Build and maintain observability: metrics, logs, traces, dashboards, and alerts for critical services.

  • Tune alerting to reduce noise while ensuring rapid detection of user impacting issues.

  • Lead or contribute to post incident reviews and root cause analysis and ensure follow up actions are implemented to prevent recurrence.

  • Added Advantage if resource is familiar on Tools Tidal, Service Now, Xmatters, Abinitio, Tableau, Opsgenie & Zeke.

Requirements:

  • 8+ years in Site Reliability Engineering, DevOps, or Cloud Engineering

  • Strong experience supporting production GCP environments

  • Experience with enterprise incident management and cloud operations

  • Knowledge/experience in GCP (Big Query, Cloud storage, Dataproc, GKE, Airflow/Composer ,Pub-sub, Cloud function, Cloud SQL etc).

  • Knowledge/experience in Github & Visual Studio code.

  • Knowledge/experience in MS-Copilot.

  • Knowledge/experience in Prometheus, Grafana & Splunk.

  • Knowledge in Python/Pyspark/Machine learning is an added advantage

  • Clear written and verbal communication, particularly under pressure (e.g., during incidents).

  • Ability to collaborate across multiple teams and influence engineering practices through expertise rather than authority.

  • Strong communication, analytical ,Knowledge on entire Incident management life cycle process, Agile model experience, and problem-solving skills

  • EDP-SRE (Google Site Reliability Engineer)

  • Site Reliability Engineers combined software engineering with systems and infrastructure operations to build and run large, reliable, scalable services.

  • Years of Experience: 10.00 Years of Experience

Skills:

  • Category Name Required Importance Experience

  • Data Mgt_Administration GCP Yes 1

  • DevOps SRE Yes 1

  • EPS_NEW GITHUB Yes 1

Why Should You Apply?

  • Health Benefits

  • Referral Program

  • Excellent growth and advancement opportunities

You've read the whole posting — now see how you match it.