Terac logo

Multilingual Data Contributors: PDF Collection for AI Training

Terac

United StatesJobNo compensation foundPosted 1w agoVerified open 3 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
No compensation found
Location
United States
Work Authorization
Not specified

Job overview

Terac is hiring a Multilingual Data Contributors: PDF Collection for AI Training. Terac is conducting a paid study to collect high‑quality, legally usable multilingual PDFs for AI training, seeking fluent readers to source, verify licensing, upload, and annotate documents in designated languages.

Key focus areas include Source legally usable public PDF documents in designated language, Verify each document meets open‑source or public domain licensing requirements, and Upload collected files to research platform.

Important skills include Multilingualism, Telugu, Odia, Gujarati, Malayalam, and Japanese.

Skills & qualifications

RequiredNice to have

Skills

MultilingualismTeluguOdiaGujaratiMalayalamJapaneseKoreanDigital Document SearchingFile UploadingMetadata EntryPublic Domain Licensing KnowledgeDetail Oriented

Qualifications

Fluent Telugu Reading ComprehensionFluent Odia Reading ComprehensionFluent Gujarati Reading ComprehensionFluent Malayalam Reading ComprehensionFluent Japanese Reading ComprehensionFluent Korean Reading ComprehensionAccess to Reliable Computer and Internet Connection

Full job description

What We're Researching We're running a paid study on multilingual document sourcing to improve AI text recognition and generation. High-quality, legally usable PDFs in various languages are essential for training robust machine learning models. Your contributions will directly support the development of better language processing tools.

How It Works You will work asynchronously to find and submit public, legally usable PDF documents in your designated language. During this process, you will verify that each document meets our quality and licensing requirements. You will upload the files through our secure platform and provide basic metadata for each submission. We will review your uploaded documents to ensure they match the project guidelines before approving the task.

Who This Is For We are hiring fluent readers of Telugu, Odia, Gujarati, Malayalam, Japanese, and Korean who know how to source public documents online. Ideal candidates are detail-oriented individuals comfortable navigating digital archives, public records, or open-source repositories. We welcome data annotators, researchers, and general language contributors who understand basic copyright and licensing rules.

What You'll Do

  • Source legally usable, public PDF documents in your designated language

  • Verify that each document meets open-source or public domain licensing requirements

  • Upload the collected files to our research platform

  • Provide basic descriptive information for each submitted document

Who Should Apply

  • Fluent reading comprehension in Telugu, Odia, Gujarati, Malayalam, Japanese, or Korean

  • Comfortable searching for and downloading digital documents online

  • Basic understanding of public domain or open-source licensing

  • Access to a reliable computer and internet connection for uploading files

Compensation $150 per task

Ready to participate? Start your paid interview now

About Terac Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.

Learn more at terac.com or on YouTube at @jointerac .

You've read the whole posting — now see how you match it.