Platform Engineer — Senior

EggAI

Remote · location unlistedFull-timePosted 4mo agoStill listed today

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
Remote · location unlisted
Schedule
Full-time
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

EggAI seeks a Senior Platform Engineer with deep, hands‑on experience designing and operating cloud‑native platforms. The role involves building and running Kubernetes platforms via Terraform, leading architecture discussions, ensuring reliability, and collaborating directly with client engineers on AI‑focused enterprise projects.

Skills & qualifications

RequiredNice to have

Skills

KubernetesTerraformHelmKustomizeAWSGCPAzureArgoCDFluxPrometheusGrafanaOpenTelemetryBashPythonGoCI/CDGitOpsSecurity

Qualifications

5+ Years Infrastructure/Platform/SRE Experience3+ Years Production Kubernetes Experience

Full job description

EggAI · Full-time · Remote (EU only)

The Role

A Senior Platform Engineer with deep, hands-on experience designing and operating cloud-native platforms in production.

You will work in a team of 2 to 6 engineers on a client project — one project at a time, so you go deep rather than spreading yourself across accounts. Clients come to us because we work at the forefront of AI in the enterprise; you build and run the platform their AI systems actually land on, alongside their own infrastructure engineers.

What You'll Do

  • Design, build and operate the Kubernetes platform for the engagement, defined in Terraform — you own it in production, not just at handover

  • Lead architecture discussions for your part of the platform, with the tech lead supporting you

  • Take ownership: pick up the ambiguous problem, chase down the answer, and say so when something is going wrong

  • Build the self-service paths the delivery engineers use, so they aren't blocked on you

  • Own reliability for what you run — SLOs, observability, incident response, cost

  • Work directly with the client's engineers, in a normal team rhythm: daily standups, planning, retrospectives, code and design review

What We're Looking For

  • 5+ years in infrastructure / platform / SRE roles, at least 3 of them running production Kubernetes — cluster design and operation, workload lifecycle, RBAC, secrets, multi-tenancy

  • Infrastructure as Code in production: Terraform (primary), plus Helm or Kustomize — owned over time, not written once

  • A major cloud in depth (AWS, GCP or Azure): networking, compute, managed databases, IAM and least-privilege, and a habit of watching what it costs

  • CI/CD and GitOps: building pipelines and running ArgoCD, Flux or equivalent

  • Observability and reliability: metrics, logs and tracing (Prometheus/Grafana, OpenTelemetry or similar), SLOs, and on-call incident response

  • Scripting and tooling: Bash, plus Python or Go

  • Security as a first-class concern: secrets management, supply-chain and vulnerability scanning

  • Client-facing ability: can hold a technical conversation with a client's architects and engineers

Desired qualities

  • Platform-thinking mindset: builds for internal customers, not just operations tasks

  • Pragmatic about trade-offs between reliability, cost, and engineering velocity

  • Thinks about blast radius before shipping

  • Clear written communication — design docs, trade-offs, post-incident reviews

Nice to have

  • Service mesh, multi-cluster or multi-region experience

  • Consulting or professional-services background

  • AI/ML or agentic workloads in production

  • EU sovereign cloud, data residency or public sector infrastructure experience

About EggAI

Agentic workforces are inevitable. Making them work is our mission.

EggAI is building the engineering methods, operating models, and reusable technology to deploy agentic workforces with the quality and control required for sustained business impact.

Working closely with enterprise clients, we design, implement and operate agentic systems that evolve from augmenting individual tasks to autonomously executing end-to-end processes alongside people—resilient, controlled, and scalable. We begin with the client problem, select the technical approach that best serves it, and remain accountable through deployment, adoption, and operations. Each deployment strengthens the next by turning production learning into reusable capability.

We are an experienced, international team working across Europe. We set high standards, take ownership of outcomes, and value clear thinking, candid feedback, and the ability to turn ideas into tangible results.

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.