AI Evaluators: Assessing A Shopping Assistant
United StatesJob$50/hrPosted 1mo agoStill listed 3 days ago
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Job overview
Terac is hiring AI evaluators to assess the accuracy and helpfulness of a digital shopping assistant, reviewing interaction traces, identifying logical failures, and creating structured rubrics, requiring 20+ hours per week remote work.
Skills & qualifications
Skills
Full job description
What We're Researching We're hiring AI evaluators to assess the accuracy and helpfulness of a new digital shopping assistant. This project focuses on understanding how well the system handles real-world e-commerce queries and where it falls short in its logic. Your analysis will directly feed into improving the underlying model and its response quality. How It Works You will review real interaction traces between users and the shopping assistant within our custom platform. As you analyze these conversations, you will pinpoint specific failures, logical errors, or unhelpful product recommendations. From there, you will create structured rubrics and verifiers to consistently judge future response quality. This is an ongoing remote engagement requiring 20+ hours per week. Who This Is For This opportunity is ideal for quality assurance specialists, AI data evaluators, and e-commerce professionals with a strong eye for detail. We welcome applicants with prior experience in prompt engineering, complex data annotation, or software testing. You should be comfortable analyzing text interactions deeply and building structured evaluation frameworks from scratch. What You'll Do
-
Review real user interaction traces with an AI shopping assistant
-
Identify logical failures, inaccuracies, or poor recommendations in the text
-
Create structured rubrics and verifiers to judge response quality
-
Commit to 20+ hours per week of evaluation work on our internal platform
Who Should Apply
-
Experience in data evaluation, quality assurance, or AI training
-
Strong analytical skills with the ability to spot subtle errors in text
-
Familiarity with e-commerce search and digital shopping experiences
-
Ability to commit to a sustained workload of 20+ hours per week
Compensation $50 per hour
Ready to participate? Start your paid interview now
About Terac Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Learn more at terac.com or on YouTube at @jointerac.
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Director, Shopping Revenue StrategyApartment Therapy Media · New York, NY · $130–150K/yrPosted 1 day agoPosted 1 day ago
Sr. Associate, eCom/Digital ShoppingWPP Media · New York, NY (Hybrid) · $45–90K/yrPosted 1 day agoPosted 1 day ago
Business Development Assistant (Hardlines)Dick's Sporting Goods · Customer Support Center (Hybrid)Posted 3w agoPosted 3w ago
Technical Program Manager II, Shopping PlatformsPinterest · Remote · US · $104–214K/yrPosted 2w agoPosted 2w agoFullfillment packing assistantCVS Health · Redlands, CA · $19/hrPosted 1 day agoPosted 1 day ago
You've read the whole posting — now see how you match it.