Fabrion logo

DevOps Engineer (Founding Team)

Fabrion

San Francisco, CAFull-timePosted 9mo agoStill listed 1w ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
San Francisco, CA
Schedule
Full-time
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

Fabrion is hiring a DevOps Engineer (Founding Team). Fabrion is seeking a DevOps Engineer to join its founding team, focusing on building an AI-native, multi-tenant enterprise platform. This role involves owning infrastructure from day one, automating CI/CD, observability, cloud governance, and security. The engineer will work with a highly technical team on real-time AI pipelines and multi-agent systems, ensuring the platform runs fast, secure, reliable, and explainable.

Key focus areas include Build and maintain scalable cloud infrastructure across AWS/GCP/Azure, Own and evolve CI/CD systems with progressive rollout, testing, and rollback flows, and Establish observability tooling across services, agents, and pipelines.

Successful candidates bring 4+ Years DevOps Experience. Important skills include CI/CD, Observability, Cloud Governance, Security, AWS, and GCP. Preferred (not required): LLM Orchestration Frameworks, LangChain, LangGraph, and Dust.

Skills & qualifications

RequiredNice to have

Skills

CI/CDObservabilityCloud GovernanceSecurityAWSGCPAzureGitHub ActionsArgo CDOpenTelemetryPrometheusGrafanaSentryPolicy-as-CodeOPARegoRBACAudit LoggingIAMAmazon VPCCryptographyKey ManagementImage ScanningSecrets RotationTerraformHelmDockerKubernetesAmazon EKSGoogle Kubernetes EnginePulumiGitOps PracticesProgressive DeliverySecure SDLCMonitoringAlertingFailure SimulationReliabilityLatencyUptimeRepeatabilityCompliance-ConsciousProactiveCollaborationLLM Orchestration FrameworksLangChainLangGraphDustReAct AgentsRetrieval-Augmented Generation Pipelines

Qualifications

4-10+ Years in DevOps4-10+ Years in Platform Engineering4-10+ Years in SREExperience Deploying Distributed Cloud-Native SystemsExperience Running LLM Orchestration FrameworksExperience Building RAG PipelinesExperience Deploying RAG PipelinesExperience Serving Fine-Tuned LLMsExperience Serving Open-Source LLMs

Full job description

DevOps Engineer (Founding Team)

Location: San Francisco Bay Area

Type: Full-Time

Compensation: Competitive salary + meaningful equity (founding tier)

Backed by 8VC, we're building a world-class team to tackle one of the industry’s most critical infrastructure problems.

About the Role We're building an AI-native, multi-tenant enterprise platform for complex domains in industrial verticals. In this architecture, DevOps isn't just about shipping features — it's about operationalizing intelligent agents , ensuring traceability across AI systems , and supporting mission-critical ML infrastructure at scale.

We're looking for a DevOps engineer who can own infrastructure from Day 1 — automating everything from CI/CD and observability to cloud governance and security. You’ll work with a highly technical team building real-time AI pipelines and multi-agent systems. If you want to be the person who makes the platform run — fast, secure, reliable, and explainable — this is your role.

Responsibilities

  • Build and maintain scalable cloud infrastructure across AWS/GCP/Azure with a focus on secure, tenant-isolated deployments

  • Own and evolve CI/CD systems (e.g. GitHub Actions, ArgoCD) with progressive rollout, testing, and rollback flows

  • Establish observability tooling across services, agents, and pipelines (OpenTelemetry, Prometheus, Grafana, Sentry)

  • Implement policy-as-code (OPA, Rego) for deployment safety, RBAC, audit logging, and approval workflows

  • Define and enforce SLAs, uptime targets (99.99%+), incident response, and remediation workflows

  • Secure infrastructure: IAM, VPC, encryption, key management, image scanning, secrets rotation

  • Automate deployments, infrastructure provisioning (Terraform, Helm), and environment replication

What We’re Looking For Core Experience:

  • 4–10+ years in DevOps, platform engineering, or SRE in production-grade systems

  • Strong experience with Docker, Kubernetes (EKS/GKE), Terraform or Pulumi

  • Hands-on experience deploying and monitoring distributed cloud-native systems

  • Familiar with GitOps practices, CI/CD design, progressive delivery, and secure SDLC

  • Clear understanding of how to implement monitoring, alerting, and failure simulation in dynamic environments

Engineering Mindset:

  • Obsessed with reliability, latency, uptime, and repeatability

  • Security-aware and compliance-conscious

  • Proactive — you don’t wait for alerts to fix things

  • Comfortable collaborating with backend, AI, and data teams

Bonus: Agent-Native / ML Ops Capabilities

  • We’re building an agentic, AI-native platform from the ground up. Experience here isn’t required, but would be a strong differentiator:

  • Experience running LLM orchestration frameworks (e.g. LangChain, LangGraph, Dust, ReAct agents)

  • Building retrieval-augmented generation (RAG) pipelines — and deploying them safely and repeatably

  • Familiarity with vector DBs (Weaviate, Qdrant, Pinecone) and embedding pipelines

  • Monitoring and governing long-running or multi-agent chains

  • Auditability and replay systems for agent decision-making

  • Serving fine-tuned or open-source LLMs with model versioning and GPU scaling (e.g. vLLM, TGI)

  • Interest in auto-remediation using agents (e.g. observability + alert → insight → response via LLM)

Why This Role Matters DevOps is the nervous system of the platform — every agent, every data fabric component, every pipeline flows through what you build. This is a rare opportunity to design that system early, the right way, and future-proof it for scale, compliance, and trust.

If you're excited by intelligent systems, distributed data, and deeply technical infrastructure problems — and you want your work to have immediate real-world impact — we’d love to hear from you.

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.