Oolka logo

AI Engineer

Oolka

Bangalore, Karnataka, IndiaJobNo compensation foundPosted 1mo agoVerified open 2 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
No compensation found
Location
Bangalore, Karnataka, India
Work Authorization
Not specified

Job overview

The AI Engineer will design and implement asynchronous multi-agent orchestration, own end-to-end latency, and build resilient inference pipelines that degrade gracefully under load. Responsibilities include intelligent request routing, WebSocket streaming, circuit breaker design, observability, and optimizing credit data retrieval and caching strategies for high‑throughput AI workloads.

Skills & qualifications

RequiredNice to have

Skills

Artificial IntelligenceRESTCommunication ProtocolsContent Distribution NetworksTensorFlowDatabase CachingMachine LearningRedisAsync/Event‑Driven ArchitecturesProduction SystemsContent Delivery NetworksMessage QueuesWebSocket/StreamingTensorFlow ServingTritonAI Model Serving FrameworksInference BatchingModel QuantisationConversation State Management

Qualifications

3-5 Years Production Systems Experience

Full job description

Responsibilities:

  • Design and implement asynchronous multi-agent orchestration.
  • Own end-to-end latency from user message to AI response.
  • Build resilient inference pipelines that gracefully degrade under load.
  • Implement intelligent request routing and load balancing for AI workloads.
  • Migrate critical AI conversation flow from monolith to dedicated services.
  • Implement WebSocket/streaming infrastructure for real-time chat.
  • Design circuit breakers and fallback strategies for AI model failures.
  • Build comprehensive observability for AI system performance.
  • Optimise credit data retrieval and caching strategies.

Requirements:

  • 3-5 years building production systems handling > 10k concurrent users.
  • Proven experience with async/event-driven architectures (not just REST APIs).
  • Hands-on experience scaling ML/AI inference in production.
  • Deep understanding of caching strategies (Redis, in-memory, CDN).
  • Experience with message queues and real-time communication protocols.
  • Built systems integrating multiple LLM/AI models in production.
  • Experience with AI model serving frameworks (TensorFlow Serving, Triton, etc. ).
  • Understanding of AI inference optimisation (batching, caching, model quantisation).
  • Knowledge of conversation state management and context handling.
  • Has debugged production issues under high AI inference load.

You've read the whole posting — now see how you match it.