AI ML Engineer
Bangalore, Karnataka, IndiaJobNo compensation foundPosted 7mo agoVerified open 2 days ago
Most applications go out cold — see where you stand first. No sign-up to start.
Don't just apply. Show up ready.
Olive works from this exact posting — no sign-up to start.
At a glance
Compensation
No compensation found
Location
Bangalore, Karnataka, India
Work Authorization
Not specified
Job overview
Oolka is seeking an AI/ML Engineer to design and implement asynchronous multi‑agent orchestration, optimize end‑to‑end latency, and build resilient inference pipelines for real‑time AI workloads. The role includes implementing intelligent request routing, migrating conversation flows to dedicated services, creating observability, and optimizing caching strategies for high‑throughput AI systems.
Skills & qualifications
RequiredNice to have
Skills
Artificial IntelligenceRESTCommunication ProtocolsContent Distribution NetworksTensorFlowDatabase CachingMachine LearningRedisAsync Event‑Driven ArchitecturesHigh Concurrency SystemsML AI Inference ScalingCaching StrategiesIn‑Memory CachingCDNMessage QueuesReal‑Time Communication ProtocolsLLM AI Model IntegrationTensorFlow ServingTritonAI Inference OptimisationBatchingModel QuantisationConversation State ManagementProduction Debugging Under LoadWebSocket StreamingLoad BalancingCircuit BreakersObservability
Qualifications
6+ Years Production Systems Experience
Full job description
Responsibilities:
- Design and implement asynchronous multi-agent orchestration.
- Own end-to-end latency from user message to AI response.
- Build resilient inference pipelines that gracefully degrade under load.
- Implement intelligent request routing and load balancing for AI workloads.
- Migrate critical AI conversation flow from monolith to dedicated services.
- Implement WebSocket/streaming infrastructure for real-time chat.
- Design circuit breakers and fallback strategies for AI model failures.
- Build comprehensive observability for AI system performance.
- Optimise credit data retrieval and caching strategies.
Requirements:
- 6+ years building production systems handling > 10k concurrent users.
- Proven experience with async/event-driven architectures (not just REST APIs).
- Hands-on experience scaling ML/AI inference in production.
- Deep understanding of caching strategies (Redis, in-memory, CDN).
- Experience with message queues and real-time communication protocols.
- Built systems integrating multiple LLM/AI models in production.
- Experience with AI model serving frameworks (TensorFlow Serving, Triton, etc. )
- Understanding of AI inference optimisation (batching, caching, model quantisation).
- Knowledge of conversation state management and context handling.
- Has debugged production issues under high AI inference load.
You've read the whole posting — now see how you match it.