Westcott Multimedia logo

DevOps Engineer - KubeOps Team

Westcott Multimedia

Tel Aviv-Yafo, Tel Aviv District, IsraelJobNo compensation foundPosted 1w agoVerified open 4 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
No compensation found
Location
Tel Aviv-Yafo, Tel Aviv District, Israel
Work Authorization
Not specified

Job overview

Westcott Multimedia is hiring a DevOps Engineer - KubeOps Team. Westcott Multimedia seeks a DevOps Engineer for the KubeOps team to manage the lifecycle of large‑scale Kubernetes environments, build and evolve platform components such as architecture, Karpenter configurations, networking and autoscaling, and integrate AI‑driven tools to accelerate investigation, planning and coding. The role requires ownership of ambiguous tasks and delivery of high‑quality outcomes.

Key focus areas include Manage the lifecycle of large‑scale Kubernetes environments without disrupting services., Own and evolve cluster building blocks including architecture, Karpenter configurations, addons, networking, and autoscaling., and Lead integration of AI‑driven tools to accelerate investigation, planning, and coding for the team..

Important skills include Kubernetes, Cloud Computing, Reliability, Prometheus (Software), Grafana, and Linux.

Skills & qualifications

RequiredNice to have

Skills

KubernetesCloud ComputingReliabilityPrometheus (Software)GrafanaLinuxSoftware Project ManagementAmazon EKSAWSTerraformPulumiCrossplaneGoPythonClaude CodeCodexCursorKarpenterCluster AutoscalingIAMObservabilityNetworking

Qualifications

2+ Years Managing Infrastructure2+ Years Production Kubernetes2+ Years Major Cloud Provider (AWS)2+ Years Infrastructure-as-CodeExperience Building Internal Tooling in Go or PythonHands-on Experience With AI-Assisted Development Tools

Full job description

Job Description

  • Master the Cluster: Manage the lifecycle of our large-scale Kubernetes environments, without disrupting the thousands of services running on it.
  • Build the Platform: Own and evolve the cluster building blocks — Kubernetes Architecture, Karpenter configurations, addons, networking, autoscaling. Where it makes sense, platformize them so other infra teams can consume them cleanly
  • Pioneer AI-Driven Ops: Lead the charge in integrating AI into our ecosystem—building LLM-powered tools that accelerate investigation, planning, and coding for the whole team.
  • Full-Stack Ownership: You’ll take ambiguous tasks and transform them into high-quality outcomes, owning every decision along the way.

Qualifications

  • +2 years in managing infrastructure, operating production systems at large scale
  • 2+ years running production Kubernetes — EKS preferred with a strong understanding of core control-plane and cluster components, with hands-on experience debugging, upgrading, and tuning autoscaling and cluster networking
  • 2+ years with a major cloud provider (AWS preferred) — solid grasp of core compute, networking, and IAM.
  • 2+ years with Infrastructure-as-Code, writing and maintaining reusable modules (Terraform, Pulumi, Crossplane, etc.)
  • Familiarity with observability at scale — Prometheus, Grafana, alert design
  • Experience building internal tooling in Go or Python
  • Hands-on experience with AI-assisted development tools and agents (Claude Code, Codex, Cursor) for knowledge gathering, debugging, coding, planning, and design — with the ability to integrate these into concrete workflows that demonstrably accelerate your work.

Advantage:

  • Cluster autoscaling in production
  • Cost / capacity optimization at fleet scale — Karpenter consolidation, bin-packing, spot strategy
  • Linux internals and node-level debugging — when a Kubernetes problem turns into a Linux problem, you can take it from there
  • Track record of leading projects end-to-end: scoping, execution and delivery.

Additional Information

KubeOps operates Wix's production Kubernetes platform at scale: a multi-DC EKS fleet with clusters running up to ~50,000 pods and ~2,000 nodes, hosting thousands of services across the company. We own the platform end-to-end — reliability, scale, upgrades, autoscaling and networking.

You've read the whole posting — now see how you match it.