Grafana Observability SME / Full Stack Observability Architect
Most applications go out cold — see where you stand first. No sign-up to start.
Don't just apply. Show up ready.
Olive works from this exact posting — no sign-up to start.
At a glance
Job overview
ICONMA, LLC is hiring a Grafana Observability SME / Full Stack Observability Architect. The Grafana Observability SME / Full Stack Observability Architect will be responsible for platform architecture and configuration across all eight in-scope Grafana Cloud modules. This role involves designing tenancy and access, developing application instrumentation strategies, and engineering log pipelines. Key responsibilities also include designing alerting rules, creating a Single Pane of Glass, and partnering on business dashboards and reporting. The architect will also integrate Grafana with ServiceNow ITOM and ensure quality assurance for all technical deliverables.
Key focus areas include Architect and configure platform across all eight in-scope Grafana Cloud modules, Design tenancy and access, including organizations, folders, teams, and RBAC, and Develop application instrumentation strategy by technology stack.
Successful candidates bring 7+ Years Observability Engineering Experience. Important skills include Grafana Cloud, Mimir, Loki, Tempo, Alloy, and Beyla. Preferred (not required): SolarWinds, Uptrends, ServiceNow CSDM, and Service Mapping Governance.
Skills & qualifications
Skills
Qualifications
Benefits
Full job description
Our Client, an IT Services and Consultant company, is looking for a Grafana Observability SME / Full Stack Observability Architect for their Poughkeepsie, NY/Remote location. Responsibilities:
-
Platform architecture and configuration across all eight in-scope Grafana Cloud modules: Grafana 12 (visualization), Mimir (metrics, 13-month retention), Loki (logs), Tempo (distributed tracing via OTLP), Alloy (telemetry collection agent), Beyla (eBPF zero-code auto-instrumentation), Application Observability (OTel-native APM), and Unified Alerting.
-
Tenancy and access design — organizations, folders, teams, role-based access control, dashboard variables, template links, and annotations.
-
Application instrumentation strategy by technology stack: Beyla eBPF as the default zero-code path for Simple and Medium apps; OpenTelemetry SDKs/agents (Java, .NET, Go, Python, Node.js) for Complex apps requiring deeper traces and custom metrics; JMX Exporter, prometheus_client, and runtime-specific exporters where stack-appropriate.
-
Log pipeline engineering via Alloy — structured JSON, Log4j/Logback, Serilog, NLog, Windows Event Log, Winston, Pino, loguru with parsing rules tuned per stack and LogQL-based dashboards and alerts.
-
Alerting design — PromQL/LogQL/TraceQL rules, severity taxonomy, grouping, routing, and notification policies. Build a low-noise, actionable alert feed; tune thresholds iteratively with application owners.
-
Single Pane of Glass — design and deliver a tiered SPoG that surfaces Grafana application telemetry alongside contextual links to SolarWinds and Uptrends.
-
Business Dashboards and Reporting — partner with the Dashboard Lead to define KPI taxonomy and ensure dashboard-as-code patterns and version control.
-
ServiceNow ITOM integration — co-own the design and review of Grafana → ServiceNow Event Management (native inbound integration) flow: event allow-list governance ("deny by default"), enrichment, deduplication, AIOps correlation, automated incident creation with severity mapping and assignment group rules, CMDB CI attachment, and ServiceNow-as-master incident state.
-
Quality assurance authority across all technical deliverables solution architecture document, instrumentation runbooks, dashboard and alert library, integration test results.
-
Phased delivery execution — Mobilise & Discover → Application Foundation (ML1) → Onboarding of 40 Simple apps (ML2) → Medium/Complex apps + ITOM Integration (ML2→3) → SPoG, Dashboards & Reporting (ML3→4) → Stabilisation, KT, and post-deployment support (ML4).
-
Knowledge transfer — produce platform operating procedures and conduct structured handover to the client's run team.
Requirements:
-
7+ years in observability/monitoring engineering with deep, recent hands-on Grafana Cloud experience (not just OSS Grafana).
-
Production expertise across the full Grafana stack: Mimir, Loki, Tempo, Alloy, Beyla, Grafana Application Observability, Unified Alerting.
-
Strong PromQL, LogQL, and TraceQL authoring skills; able to write recording rules and SLO queries from scratch.
-
OpenTelemetry practitioner OTLP, collectors, SDK/agent instrumentation for at least three of Java, .NET, Go, Python, Node.js.
-
eBPF-based auto-instrumentation experience with Beyla (or equivalent Pixie, Cilium Tetragon) in a production context.
-
Experience integrating Grafana alerts into ServiceNow Event Management (native inbound integration, not webhook-only patterns); familiarity with ServiceNow ITOM, AIOps event correlation, and CMDB CI attachment.
-
Multi-environment hosting fluency on-prem, AWS, Azure and Linux/Windows host agent deployment at scale.
-
Dashboard-as-code and GitOps patterns (Grafana provisioning, Terraform provider, or Grizzly).
-
Excellent written communication solution architecture documents, runbooks, and stakeholder-facing status reporting.
-
Nice to Have
-
Grafana Certified Professional or equivalent vendor certification.
-
Prior experience in a regulated utility, energy, or critical-infrastructure environment.
-
Familiarity with SolarWinds and Uptrends (sufficient to design clean boundaries with retained tooling, not to administer them).
-
Experience with ServiceNow CSDM and Service Mapping governance.
-
Exposure to FinOps for observability — cardinality control, log volume management, retention tuning in Mimir/Loki.
Why Should You Apply?
-
Health Benefits
-
Referral Program
-
Excellent growth and advancement opportunities
You've read the whole posting — now see how you match it.