Agtonomy logo

Senior/Staff Machine Learning Engineer, Perception

Agtonomy

San Francisco, CAFull-time$200–280K/yrPosted 1mo agoVerified open 6 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
$200–280K/yr
Location
San Francisco, CA
Schedule
Full-time
Work Authorization
Not specified

Requirements

Credentials this posting asks for.

Doctorate

Job overview

Agtonomy is hiring a Senior/Staff Machine Learning Engineer, Perception. Agtonomy is seeking a skilled Machine Learning Engineer to develop perception systems for autonomous heavy machinery in rugged environments. The role involves building computer vision and machine learning systems to interpret camera and LiDAR data for robust 3D scene understanding. This hands-on position requires writing production-grade software, optimizing models for embedded hardware, and validating work on real machines globally. The engineer will be at the forefront of developing learned, dense scene representations and foundation-model-driven data engines.

Key focus areas include Develop real-time perception models for open-world obstacle and terrain understanding, Build multi-modal fusion combining camera and LiDAR into a unified 3D/BEV representation, and Optimize models for low-latency inference on resource-constrained hardware.

Successful candidates bring MS/PhD In Computer Science, AI, Or Related Field, 6+ Years Industry Experience Building Vision-Based Perception Systems, and Publications At Top-Tier Perception/Robotics Venues. Important skills include Computer Vision, Machine Learning, 3D Scene Understanding, Real-Time Perception Models, Multi-Modal Fusion, and Model Optimization. Preferred (not required): Building Auto-Labeling/Data-Engine Flywheels and Vision-Language-Action (VLA) Models.

Skills & qualifications

RequiredNice to have

Skills

Computer VisionMachine Learning3D Scene UnderstandingReal-Time Perception ModelsMulti-Modal FusionModel OptimizationAuto-Labeling PipelinesFoundation ModelsTeacher-Student DistillationData and Evaluation PipelinesAlgorithm IterationModern Perception ModelsDetectionSegmentationMono/Stereo/Metric DepthBEV/OccupancySensor FusionFine-Tuning Large Pre-Trained Vision ModelsMulti-Sensor IntegrationCalibrationSpatiotemporal SyncCross-Modal FusionHandling Large DatasetsPythonPyTorchTensorFlowOpenCVProduction-Ready CodeReal-Time SystemsExperiment DesignMetrics AnalysisMarketing AutomationIoULatency/ThroughputAgilityCollaborationOwnershipArchitecting Multi-Sensor ML SystemsBuilding Auto-Labeling/Data-Engine FlywheelsCompute-Constrained PipelinesTensorRTModel QuantizationPredictive World ModelsVision-Language-Action (VLA) ModelsWorld-Action (WAM) ModelsCustom CUDA Operations

Qualifications

MS/PhD in Computer Science, AI, or Related Field6+ Years Industry Experience Building Vision-Based Perception SystemsPublications at Top-Tier Perception/Robotics Venues

Benefits

Medical Insurance
Dental Insurance
Vision Insurance
401(k) Match

Full job description

About Us At Agtonomy, we’re not just building tech—we’re transforming how vital industries get work done. Our Physical AI and fleet services turn heavy machinery into intelligent, autonomous systems that tackle the toughest challenges in agriculture, turf, and beyond. Partnering with industry-leading equipment manufacturers, we’re creating a future where labor shortages, environmental strain, and inefficiencies are relics of the past. Our team is a tight-knit group of bold thinkers—engineers, innovators, and industry experts—who thrive on turning audacious ideas into reality. If you want to shape the future of industries that matter, this is your shot.

About the Role

We're looking for a skilled ML engineer to build the perception systems that give our autonomous machines human-like awareness in rugged, unstructured environments. You'll develop computer vision and machine learning systems that turns noisy camera and LiDAR data into robust 3D scene understanding — enabling heavy equipment to operate safely through dust, glare, occlusion, and whatever messy conditions a working site throws at it.

The field is moving past bounding-box detection and hand-tuned tracking toward learned, dense scene representations, foundation-model-driven data engines, and uncertainty-aware perception. You'll be at the center of that shift. This role is hands-on: you'll write production-grade software, distill and optimize models for embedded hardware, and validate your work on real machines at operating around the world.

What You'll Do

  • Develop real-time perception models for open-world obstacle and terrain understanding.

  • Build multi-modal fusion that combines camera and LiDAR into a unified 3D/BEV representation, robust to occlusions, sensor degradation, and GNSS outages.

  • Optimize models for low-latency inference on resource-constrained hardware, balancing accuracy and performance.

  • Design auto-labeling pipelines that leverage foundation models and teacher-student distillation to scale labeling and close the loop from real-world field interventions.

  • Design data and evaluation pipelines that curate large multi-sensor datasets and surface failures fast, with strong visualization and debugging tooling.

  • Analyze performance metrics and iterate on algorithms to improve accuracy and efficiency of various perception subsystems.

What You'll Bring

  • A MS/PhD in Computer Science, AI, or a related field, or 6+ years of industry experience building vision-based perception systems.

  • Deep expertise developing and deploying modern perception models: detection, segmentation, mono/stereo/metric depth, BEV/occupancy, sensor fusion, and 3D scene understanding.

  • Fluency adapting, fine-tuning, and distilling large pre-trained vision and vision-language models.

  • Strong grounding in multi-sensor integration (camera, LiDAR, radar): calibration, spatiotemporal sync, and cross-modal fusion.

  • Experience handling large datasets efficiently and organizing them for labeling, training and evaluation.

  • Fluency in Python with PyTorch/TensorFlow/OpenCV and the ability to write efficient, production-ready code for real-time systems.

  • Proven ability to design experiments, analyze metrics (mAP, IoU, latency/throughput, and calibration/ECE), and optimize to meet stringent real-world performance and safety requirements.

  • An eagerness to get your hands dirty and agility in a fast-moving, collaborative, small team environment with lots of ownership.

What Makes You a Strong Fit

  • Experience architecting multi-sensor ML systems from scratch.

  • Experience building auto-labeling / data-engine flywheels at scale.

  • Experience with compute-constrained pipelines including optimizing models to balance the accuracy vs. performance tradeoff, leveraging TensorRT, model quantization, etc.

  • Familiarity with emerging predictive world models for anticipation, anomaly detection, or closed-loop simulation, and adjacent policy paradigms such as Vision-Language-Action (VLA) and World-Action (WAM) models.

  • Experience with compute-constrained deployment: TensorRT, model quantization, and custom CUDA operations.

  • Publications at top-tier perception/robotics venues (CVPR, ICRA, CoRL, RSS, etc.).

  • Passion for how we feed, build, move, and maintain the world.

Benefits  

  • 100% covered medical, dental, and vision for the employee (partner, children, or family is additional)
  • Commuter Benefits
  • Flexible Spending Account (FSA or HSA)
  • Life Insurance
  • Short- and Long-Term Disability
  • 401k Plan
  • Stock Options
  • Collaborative work environment working alongside passionate mission-driven team!

You've read the whole posting — now see how you match it.