StackAI logo

Senior Software Engineer, Engine & Distributed Systems

StackAI

New York, NYFull-timePosted 3mo agoStill listed 5 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
New York, NY
Schedule
Full-time
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

StackAI is hiring a Senior Software Engineer, Engine & Distributed Systems. StackAI (by Asana) is seeking a Senior Software Engineer to own the core execution engine of its no-code AI workflow platform. This role involves deep systems work, focusing on durable runtime, scheduling, and sub-agent parallelization. The engineer will ensure agents run correctly at any scale, building checkpointing, resumption, and recovery mechanisms. They will also shape the execution model and engineer for scale and reliability.

Key focus areas include Own the execution engine, runtime, scheduling, and sub-agent parallelization, Make long-running work durable with checkpointing, resumption, and recovery, and Shape the execution model for scheduling, queuing, and asynchronous movement.

Successful candidates bring 5+ Years Backend Systems Experience. Important skills include Distributed Systems, Durable Execution, Workflow Orchestration, Temporal, Cadence, and Idempotency. Preferred (not required): Event-Driven Architectures, Message Queues, PydanticAI, and LangGraph.

Skills & qualifications

RequiredNice to have

Skills

Distributed SystemsDurable ExecutionWorkflow OrchestrationTemporalCadenceIdempotencyState MachinesFailure RecoveryConcurrencyMessage QueueRetriesFault TolerancePythonFastAPIDatabase FundamentalsPostgreSQLCorrectness ProblemsEvent-Driven ArchitecturesMessage QueuesPydanticAILangGraphAI RuntimesAgent RuntimesTool-CallingSub-Agent OrchestrationStreamingPerformance OptimizationCost OptimizationHigh-Throughput Backends

Qualifications

5+ Years Building Backend Systems in ProductionStartup or Growth-Stage Experience

Full job description

About the Company StackAI (by Asana) is a no-code AI workflow platform that empowers companies to build, deploy, and scale AI-powered workflows easily. With thousands of users and rapidly growing enterprise adoption, our mission is to democratize access to LLMs and bring AI into the hands of every business operator, not just developers. An Asana Company, we're building the foundational platform for the AI-driven future of work

The role Enterprises run real work on AI agents, and at StackAI that work runs on a single engine. Some agents finish in a second. Others run for days, fan out into dozens of sub-agents, pause, resume, and recover from failures without losing a step. We're hiring a Senior Software Engineer, Engine & Distributed Systems to own that engine: the durable runtime at the core of the platform that has to be correct every time, at any scale.

This is deep systems work at the heart of the product. When the engine is solid, agents simply run — and getting it there is one of the more interesting distributed-systems problems in AI today. You'll own it end to end, from the execution model to how it behaves in production.

What you'll do

  • Own the execution engine. The runtime, scheduling, and sub-agent parallelization that run every agent on the platform.

  • Make long-running work durable. Build checkpointing, resumption, and recovery so agents survive failures and restarts and pick up exactly where they left off.

  • Shape the execution model. Decide how work is scheduled, queued, and moved from synchronous to asynchronous, so the platform stays correct and responsive as load grows.

  • Engineer for scale and reliability. Hold the engine to strict health targets for worker freshness, deploy safety, and drain time, and keep latency and throughput strong as volume grows.

  • Keep the engine open to the ecosystem. Make it straightforward to bring new agent harnesses, orchestration frameworks, and model capabilities into the runtime.

What we're looking for

  • 5+ years building backend systems in production, with real depth in distributed systems.

  • Hands-on experience with durable execution or workflow orchestration (Temporal, Cadence, or equivalent), with a way of thinking rooted in idempotency, state machines, and failure recovery.

  • Strong command of concurrency, queueing, retries, and fault tolerance under load.

  • Strong in Python and modern backend frameworks (FastAPI or similar), with sound database fundamentals (Postgres or similar).

  • You're drawn to the correctness problems that everything else quietly depends on.

Distributed systems is broad. If you're strong on most of this and excited to grow into the rest, we'd like to hear from you, even if you don't check every box.

Bonus points

  • Operating Temporal at scale.

  • Event-driven architectures and message queues.

  • Experience with PydanticAI, LangGraph, or similar.

  • AI or agent runtimes: tool-calling, sub-agent orchestration, streaming.

  • Performance and cost optimization of high-throughput backends.

  • Startup or growth-stage experience.

Why StackAI You'll join a lean, high-impact team and own the engine that every customer's agents run on. Your work ships fast and is felt across the whole product.

StackAI is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.