Oolka logo

Senior AI ML Engineer

Oolka

Bangalore, Karnataka, IndiaJobNo compensation foundPosted 7mo agoVerified open 2 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
No compensation found
Location
Bangalore, Karnataka, India
Work Authorization
Not specified

Job overview

The Senior AI ML Engineer will design and implement asynchronous multi‑agent orchestration, own end‑to‑end latency for AI responses, and build resilient inference pipelines that degrade gracefully under load while optimizing credit data retrieval and caching.

Skills & qualifications

RequiredNice to have

Skills

TensorFlowDatabase CachingMachine LearningCommunication ProtocolsContent Distribution NetworksRedisRESTArtificial IntelligenceAsync/Event‑Driven ArchitecturesML/AI Inference ScalingCaching StrategiesContent Delivery NetworksMessage QueuesReal‑Time Communication ProtocolsWebSocket StreamingLLM/AI Model IntegrationTensorFlow ServingTriton Inference ServerBatchingModel QuantisationConversation State Management

Qualifications

6+ Years Production Systems Experience

Full job description

Responsibilities:

  • Design and implement asynchronous multi-agent orchestration.
  • Own end-to-end latency from user message to AI response.
  • Build resilient inference pipelines that gracefully degrade under load.
  • Implement intelligent request routing and load balancing for AI workloads.
  • Migrate critical AI conversation flow from monolith to dedicated services.
  • Implement WebSocket/streaming infrastructure for real-time chat.
  • Design circuit breakers and fallback strategies for AI model failures.
  • Build comprehensive observability for AI system performance.
  • Optimise credit data retrieval and caching strategies.

Requirements:

  • 6+ years building production systems handling > 10k concurrent users.
  • Proven experience with async/event-driven architectures (not just REST APIs).
  • Hands-on experience scaling ML/AI inference in production.
  • Deep understanding of caching strategies (Redis, in-memory, CDN).
  • Experience with message queues and real-time communication protocols.
  • Built systems integrating multiple LLM/AI models in production.
  • Experience with AI model serving frameworks (TensorFlow Serving, Triton, etc. )
  • Understanding of AI inference optimisation (batching, caching, model quantisation).
  • Knowledge of conversation state management and context handling.
  • Has debugged production issues under high AI inference load.

You've read the whole posting — now see how you match it.

Senior AI ML Engineer at Oolka | Olive Jobs