Felix logo

Staff Platform Engineer (Runtime & Security)

Felix

United StatesJobPosted 1mo agoStill listed 4 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
United States
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

Félix is a hyper‑growth Series C fintech building an AI‑powered financial companion for US Latinos, offering remittances, credit, savings and wallet services. The company seeks a Staff Platform Engineer to lead platform and security foundations for its internal AI teammate, Maestro, focusing on multi‑tenant runtime, identity, isolation, and audit at scale.

Skills & qualifications

RequiredNice to have

Skills

PythonMicrosoft AzureMulti Tenant SystemsCloud Native ArchitectureEarly-StageThreat ModelingInfrastructure as Code (IaC)Google Kubernetes Engine (GKE)Distributed SystemsAnchorIncident ResponseEncryptionHelmProduct SecurityFinancial TechnologyCMovementGoogle Cloud PlatformObservabilityAuditGithubPagerDutyOpenIDKubernetesKey Management ServiceControl PlaneTerraformInfrastructureOAuthTeam LeadershipArtificial IntelligencePerformance TuningAmazon Web ServicesSystem ArchitectureCross-Functional Team LeadershipServiceGogVisorIstioSPIFFEOAuth 2.0OpenID ConnectCloud KMSOpenTelemetryNew RelicBigQueryCloud LoggingCI/CD

Qualifications

8+ Years Experience

Benefits

Medical Insurance
Dental Insurance
Vision Insurance
401(k) Match
Paid Time Off
Parental Leave

Full job description

About Us

At Félix, we are building the indispensable financial companion for Latinos in the US. We combine an AI-powered, conversational-first interface with real-time financial infrastructure to make cross-border money movement as easy as sending a text. Starting with fast, affordable remittances powered by AI and crypto rails, we are expanding into credit, savings, and wallet services to support the complete immigrant financial journey. Our ambition is to deliver a white-glove financial experience with the simplicity of a conversation to a community the traditional financial system has historically overlooked.

We are a hyper-growth Series C company, backed by over $300 million in funding from top-tier global investors, including Andreessen Horowitz, QED, Castle Island, Switch Ventures, HTwenty, Monashees, General Catalyst Customer Value Fund. This isn't just about the numbers; it's a testament to the trust our investors have in our vision and our team. Additionally, Félix was selected as an “Endeavour Entrepreneur” and was a recipient of the CrossTech Fintech Startups Award.

Joining Félix means you will be part of a team building a legacy, a company that will outlive us all. This is a rare opportunity to apply your skills to a deeply meaningful mission—serving a community that has been underserved for too long because we are obsessed with our customers. We get things done with urgency and focus, driven by extreme ownership over our impact. We collaborate without ego, fostering radical transparency and fierce loyalty so we can grow together. Because we aim for insanely great rather than just good enough, we stay insatiably curious—always experimenting, building the future today, and delivering a product that truly makes our users' lives better.

About the Role

We're looking for a highly technical Staff Platform Engineer to lead the platform and security foundations of Maestro — Félix's internal, identity-aware AI teammate. Maestro already runs at meaningful scale (500+ per-user pods on private GKE) and lives where work happens: Slack, incident rooms, and an emerging agentic intranet, acting across our toolchain (GitHub, Google Workspace, ClickUp, Notion, PagerDuty, New Relic).

The interesting problems here are not prompts or models. Once an AI teammate can open a pull request, page an engineer, or query production, the hard questions become identity, credentials, isolation, blast radius, and audit. This is a platform and security role for an agentic system — you'll own the secure, multi-tenant runtime that makes delegated AI work safe at scale.

You'll be the technical anchor for Maestro's infrastructure and security within the AI team: architecting the control plane, hardening the runtime, running the fleet, and shaping what the platform needs next as adoption grows. AI is the domain you'll operate in — deep platform and security engineering is the craft we're hiring for.

Responsibilities

  • Own the Maestro platform architecture. Design, build, and operate the multi-tenant control plane and per-user runtime on private GKE — the Kubernetes operators (CRDs/controller-runtime), Helm charts, gVisor-sandboxed pods, and per-user isolation primitives (KSA/GSA, Workload Identity, NetworkPolicy, per-user workspaces) that reconcile one user into a fully wired, isolated environment.
  • Lead the security model end to end. Treat the LLM and its tools as adversarial. Own identity separation (requester / actor / persona), JIT short-lived scoped tokens, an encrypted OAuth refresh-token vault (CMEK/Cloud KMS), zero-credential egress patterns, and a policy layer that decides whose credentials an agent uses — never the prompt.
  • Harden service-to-service trust. Enforce mesh identity with Istio mTLS + SPIFFE, signed request claims (JWS) to prevent confused-deputy issues, and deny-by-exception networking across the fleet (Istio AuthorizationPolicy + Kubernetes NetworkPolicy).
  • Operate the fleet, not the bot. Build fleet health, scale-to-zero, resource packing, and safe operational tooling for 500+ pods, with the SRE-grade availability, latency, and recovery the platform demands.
  • Own IaC and delivery. Drive Terraform for the dedicated GCP projects (VPC, private GKE, GPU/gVisor node pools, Cloud SQL, Memorystore, GSM, Artifact Registry), plus CI/CD and progressive delivery for control-plane and runtime components.
  • Make audit a product feature. Own the OpenTelemetry pipeline (logs/metrics/traces) fanning out to Cloud Logging, BigQuery, and New Relic, capturing gateway, kernel/gVisor syscall, and real-time SecOps events so "who asked, which persona acted, which credentials were used, did the user confirm?" is always answerable.
  • Enforce human-in-the-loop and guardrails. Build the approval flows for irreversible actions (writes, merges, admin ops) and the read-only, model-immutable guardrail mounts (identity, instructions, curated skills).
  • Integrate the agent layer, safely. Partner on the OpenClaw gateway, controlled tool wrappers, and model routing (Vertex AI for stakes, self-hosted Ollama for volume) — ensuring every agent capability is a governed capability, not a raw CLI or API key.
  • Set the technical direction. Define platform and security best practices, mentor senior and mid-level engineers, and map Maestro's next infrastructure needs as it scales.

Requirements

  • Experience: 8+ years in software/infrastructure engineering, with a proven track record owning large-scale, security-critical distributed systems end to end.
  • Platform & Kubernetes mastery (Staff bar): Deep, hands-on Kubernetes in production — operators/CRDs and controller-runtime, Helm, runtime isolation (gVisor or equivalent), multi-tenancy, and fleet operations at scale. Strong cloud-native architecture on GCP (or AWS/Azure), and IaC with Terraform.
  • Security & identity depth (Staff bar): Strong applied security engineering — workload identity (SPIFFE/SPIRE, Workload Identity Federation), service mesh mTLS (Istio), OAuth 2.0 / OIDC, token exchange, JIT/short-lived scoped credentials, secrets/KMS envelope encryption, least-privilege and zero-trust patterns, and threat modeling for adversarial workloads (confused-deputy, prompt injection, data exfiltration).
  • Systems & code: Excellent Go and/or Python, with deep system-architecture judgment. Comfortable owning services, operators, and tooling in production.
  • Observability & LLMOps: Production-grade monitoring, tracing, and audit design (OpenTelemetry), plus SRE fundamentals — SLOs, incident response, cost/performance engineering.
  • Ownership & leadership: High autonomy in an early-stage squad — independently diagnose bottlenecks, propose architecture, and ship it. Proven ability to grow engineers through architectural guidance, not just code review, and to align technical decisions with stakeholders across Product, Security, and Leadership.
  • Nice to Have — AI Exposure: Prior hands-on experience with agentic systems or LLM infrastructure is a plus, but strong platform and security fundamentals take priority.
  • These are the applicable requisites, although equivalent competencies in any of the above will also be considered.

What We Offer

  • Competitive salary
  • Initial stock options grant
  • Annual performance bonus
  • Health, dental, and vision plans
  • 401(k) with employer match
  • Continuous learning opportunities
  • Unlimited PTO
  • Paid parental leave
  • Empowering opportunities for growth in a dynamic entrepreneurial environment

What We Offer

  • Competitive salary
  • Initial stock options grant
  • Annual performance bonus
  • Health, dental, and vision plans
  • Continuous learning opportunities
  • 401(k) with employer match
  • Unlimited PTO
  • Paid parental leave
  • Empowering opportunities for growth in a dynamic entrepreneurial environment

Equal Opportunity Employer

At Félix, we are committed to providing equal employment opportunities to all qualified employees and applicants without regard to race, religion, nationality, sex, sexual orientation, gender identity, age, or disability. This policy applies to all terms and conditions of employment, including recruitment, hiring, placement, promotion, training, compensation, benefits, and termination.

Want to learn more about our privacy practices? Check out our Privacy Policy.

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.