Senior Machine Learning Engineer
Remote · USFull-time$130–160K/yrPosted 1w agoStill listed 4 days ago
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New remote roles like this one, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Job overview
Candid’s Data Science and AI team seeks a Senior Machine Learning Engineer to own operational aspects of deployed ML services, improve inference performance, and build observability, while collaborating with data scientists and engineering teams across the United States.
Skills & qualifications
Skills
Qualifications
Full job description
In A Nutshell Location Remote Anywhere in United States
Salary $130,000 - $160,000 / year
Job Type Full-time
Experience Level Mid-level
Deadline to apply October 16, 2026
Candid’s Data Science and AI team builds and ships machine learning and AI models across many concurrent projects. The Senior ML Engineer exists to take operational ownership of Candid’s deployed ML and AI services to free up data science capacity.
Responsibilities
- Take operational ownership of Candid’s deployed machine learning and AI services — monitoring for degradation, managing retraining cadences, coordinating handoffs from data scientists, and serving as the accountable point of contact for models in production.
- Improve inference performance for deployed models, including complex graph inference models, applying techniques such as quantization, artifact slimming, batching, and efficient serialization.
- Design and operate experiment tracking, model versioning, and artifact management so data scientists have a consistent, low-friction way to hand off work.
- Build and maintain observability for ML and AI services — centralized logging, metrics, dashboards, and alerting for the services that matter most — and proactively reduce incidents through good deployment hygiene.
- Establish and iteratively improve a repeatable deployment path for new ML services on AWS, including CI/CD integration and infrastructure-as-code patterns.
- Monitor and actively manage AWS spend across ML workloads, including Bedrock token usage, compute sizing, S3 lifecycle, and publish cost visibility to the team.
- Work closely with data scientists to understand model behavior, surface operational insights back to them, and translate research code into production-ready deployments.
- Serve as the primary technical liaison between the Data Science team and Candid’s product and software engineering teams when ML or AI services are integrated into products and systems, defining integration contracts, APIs, latency and reliability expectations, and input/output schemas, and supporting those teams through integration.
- Build and operate Amazon Bedrock-backed services and integrations where they are the right tool for the job.
- Contribute to Candid’s secure-by-default posture for ML services, including IAM scoping, secrets handling, and compliance-aligned tagging.
- Participate in technical planning and roadmap discussions for the Data Science team.
Skillset
- 4+ years of professional software engineering experience, including at least 2 years in a role where production ML systems were a primary responsibility (MLOps, ML engineering, or production-ML-focused data science).
- Strong proficiency in Python, including writing production-quality service code.
- Hands-on, operational experience with experiment tracking and model lifecycle tooling in production, such as MLflow, Weights & Biases, or comparable systems.
- Hands-on experience deploying PyTorch models in production, including familiarity with serving optimization techniques such as quantization, batching, ONNX Runtime, model serialization, container/artifact slimming, or cold-start mitigation on Lambda/Fargate.
- Demonstrated experience deploying and monitoring ML models in production, with an understanding of model degradation, drift signals, retraining triggers, and artifact management.
- Working knowledge of deploying Python-based ML services on AWS including Lambda, ECS/Fargate, S3, IAM, and CloudWatch. Deep infrastructure expertise is not required; comfort with the AWS ML deployment stack is.
- Experience building or operating CI/CD pipelines for ML services.
- A demonstrated track record of improving production reliability through observability and disciplined deployment and maintaining healthy systems over time.
- Ability to work closely with data scientists, understanding their outputs, inheriting their work, and communicating operational decisions back to them clearly.
- Demonstrated ability to work cross-functionally with software or product engineering teams, translating ML service capabilities into integration-ready contracts and supporting those teams through adoption.
- Comfort owning work independently.
- Strong written and verbal communication.
- Willingness to perform other duties and special projects as needed/requested.
- Sensitivity and respect for racial, gender, sexual orientation, and cultural differences.
- Champions and represents Candid’s core values: We’re driven, direct, accessible, curious, and inclusive.
Apply Now Spot any inaccurate information? Have a job to share? Let us know.
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Senior Machine Learning Engineer, Autonomous DefenseHorizon3.ai · Remote · US · $211–249K/yrPosted 3w agoPosted 3w ago
Senior Machine Learning Engineer- Generative AI (REMOTE)Home Depot · Remote · US · $100–180K/yrPosted 2w agoPosted 2w ago
Senior Staff Machine Learning Engineer Zscaler · Remote · US · $158–225K/yrPosted 1w agoPosted 1w ago
Senior Staff Machine Learning Engineer, ML Understandingreddit · Remote · location unlisted · $266–372K/yrPosted 1w agoPosted 1w ago
Senior Staff Machine Learning Engineer, Agentic AI Platform (Ads Ranking) reddit · Remote · location unlisted · $293–410K/yrPosted 3w agoPosted 3w ago
You've read the whole posting — now see how you match it.