AI Engineer
Bangalore, Karnataka, IndiaJobNo compensation foundPosted 1mo agoVerified open 2 days ago
Most applications go out cold — see where you stand first. No sign-up to start.
Don't just apply. Show up ready.
Olive works from this exact posting — no sign-up to start.
At a glance
Compensation
No compensation found
Location
Bangalore, Karnataka, India
Work Authorization
Not specified
Job overview
The AI Engineer will design and implement asynchronous multi-agent orchestration, own end-to-end latency, and build resilient inference pipelines that degrade gracefully under load. Responsibilities include intelligent request routing, WebSocket streaming, circuit breaker design, observability, and optimizing credit data retrieval and caching strategies for high‑throughput AI workloads.
Skills & qualifications
RequiredNice to have
Skills
Artificial IntelligenceRESTCommunication ProtocolsContent Distribution NetworksTensorFlowDatabase CachingMachine LearningRedisAsync/Event‑Driven ArchitecturesProduction SystemsContent Delivery NetworksMessage QueuesWebSocket/StreamingTensorFlow ServingTritonAI Model Serving FrameworksInference BatchingModel QuantisationConversation State Management
Qualifications
3-5 Years Production Systems Experience
Full job description
Responsibilities:
- Design and implement asynchronous multi-agent orchestration.
- Own end-to-end latency from user message to AI response.
- Build resilient inference pipelines that gracefully degrade under load.
- Implement intelligent request routing and load balancing for AI workloads.
- Migrate critical AI conversation flow from monolith to dedicated services.
- Implement WebSocket/streaming infrastructure for real-time chat.
- Design circuit breakers and fallback strategies for AI model failures.
- Build comprehensive observability for AI system performance.
- Optimise credit data retrieval and caching strategies.
Requirements:
- 3-5 years building production systems handling > 10k concurrent users.
- Proven experience with async/event-driven architectures (not just REST APIs).
- Hands-on experience scaling ML/AI inference in production.
- Deep understanding of caching strategies (Redis, in-memory, CDN).
- Experience with message queues and real-time communication protocols.
- Built systems integrating multiple LLM/AI models in production.
- Experience with AI model serving frameworks (TensorFlow Serving, Triton, etc. ).
- Understanding of AI inference optimisation (batching, caching, model quantisation).
- Knowledge of conversation state management and context handling.
- Has debugged production issues under high AI inference load.
You've read the whole posting — now see how you match it.