Choctaw Nation of Oklahoma logo

AI Operations and Monitoring Engineer

Choctaw Nation of Oklahoma

Durant, OK · HybridFull-timeSeen 1w agoSeen in employer's feed 1w ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
Durant, OKHybrid
Schedule
Full-time
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Requirements

Credentials this posting asks for.

Bachelor's degree

Job overview

The AI Operations and Monitoring Engineer ensures reliability, performance, documentation, and compliance of production AI/ML systems, maintaining system uptime and driving effective incident response to minimize outages and protect service quality.

Skills & qualifications

RequiredNice to have

Skills

KubernetesDockerCloud ServicesObservability ToolsCI/CD PipelinesMonitoringDashboardingAlertsAutomationSecurityComplianceGovernanceRunbooksAudit Logs

Qualifications

Bachelor's in Computer Science or Related Field or 4 Years Relevant Experience3 Years Experience in DevOps SRE or MLOps

Full job description

Monday-Friday 8:00AM-4:30PM| Hybrid Position| Weekly Earned Wage Access is an option for this position.

Job Purpose or Goals: The AI Operations and Monitoring Engineer is responsible for ensuring the reliability, performance, documentation, and compliance of production AI/ML systems, keeping them stable and functioning as intended. They play a key role in maintaining system uptime and driving effective incident response to minimize outages and protect overall service quality.

Tasks:

  1. Monitor AI/ML system health, drift, and anomalies by continuously tracking model performance, identifying unexpected behavior, and ensuring systems operate within expected parameters.

  2. Build dashboards, alerts, and observability tooling to give teams real-time insight into system performance and enable rapid detection of issues before they impact operations.

  3. Implement CI/CD pipelines for models to streamline deployments, maintain version control, and ensure that model updates move reliably from development to production.

  4. Maintain training and deployment environments by keeping infrastructure stable, updated, and optimized so models can be trained and deployed efficiently.

  5. Develop automation for scaling, retraining and reducing manual workloads and ensuring models automatically adapt to new data or demand.

  6. Respond to incidents and outages to restore services quickly, conduct rootcause analysis, and prevent future disruptions to AI/ML workloads.

  7. Collaborate on security, compliance, and governance to ensure AI systems follow organizational standards and regulatory requirements.

  8. Write runbooks and maintain audit logs to document procedures, support operational consistency, and provide traceability for all system actions.

  9. Performs other duties as may be assigned.

Job Requirements:

Bachelor's degree in computer science or related field, or 4 years relevant professional experience

3 years' experience in DevOps, SRE, or MLOps Proficiency with tools like Kubernetes, Docker, and cloud services

Experience with observability tools

Familiarity with ML deployment workflows

Job Identification: 31230

Job Category: Information Technology

Posting Date: 09/03/2026, 7:29 PM

Job Schedule: Full time

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.