Oolka logo

Data Science / ML Engineer

Oolka

Bangalore, Karnataka, IndiaJobNo compensation foundPosted 5mo agoVerified open 3 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
No compensation found
Location
Bangalore, Karnataka, India
Work Authorization
Not specified

Job overview

Oolka is seeking a Data Science / ML Engineer to design and implement asynchronous multi‑agent orchestration, own end‑to‑end latency, and build resilient inference pipelines for AI workloads, while optimizing caching and real‑time communication infrastructure.

Skills & qualifications

RequiredNice to have

Skills

RedisRESTMachine LearningTensorFlowCommunication ProtocolsDatabase CachingArtificial IntelligenceContent Distribution NetworksAsync/Event‑Driven ArchitecturesML/AI Inference ScalingCaching StrategiesMessage QueuesReal‑Time Communication ProtocolsTensorFlow ServingTritonLLM IntegrationModel Quantization

Qualifications

3+ Years Production Systems Experience

Full job description

Responsibilities:

  • Design and implement asynchronous multi-agent orchestration.
  • Own end-to-end latency from user message to AI response.
  • Build resilient inference pipelines that gracefully degrade under load.
  • Implement intelligent request routing and load balancing for AI workloads.
  • Migrate critical AI conversation flow from monolith to dedicated services.
  • Implement WebSocket/streaming infrastructure for real-time chat.
  • Design circuit breakers and fallback strategies for AI model failures.
  • Build comprehensive observability for AI system performance.
  • Optimize credit data retrieval and caching strategies.

Requirements:

  • 3+ years building production systems handling > 10k concurrent users.
  • Proven experience with async/event-driven architectures (not just REST APIs).
  • Hands-on experience scaling ML/AI inference in production.
  • Deep understanding of caching strategies (Redis, in-memory, CDN).
  • Experience with message queues and real-time communication protocols AI-Specific.

Expertise:

  • Built systems integrating multiple LLM/AI models in production.
  • Experience with AI model serving frameworks (TensorFlow Serving, Triton, etc. ).
  • Understanding of AI inference optimization (batching, caching, model quantization).
  • Knowledge of conversation state management and context handling.
  • Has debugged production issues under high AI inference load.

Growth Path:

  • Direct impact on customer subscription retention through performance.
  • Exposure to cutting-edge AI infrastructure challenges.
  • Ownership of technical decisions affecting revenue-generating conversations.
  • Path to leading an AI platform team as you scale.

You've read the whole posting — now see how you match it.