Cognichip Inc. logo

Senior Data Scientist, AI Training Data

Cognichip Inc.

Redwood City, CAJobNo compensation foundPosted todayVerified open today

Most applications go out cold — see where you stand first. No sign-up to start.

At a glance

Compensation
No compensation found
Location
Redwood City, CA
Work Authorization
Not specified

Requirements

Credentials this posting asks for.

Master's degree

Job overview

Cognichip is seeking a senior data scientist to own and curate the specialized data that powers its AI models for semiconductor engineering. The role involves designing data pipelines, synthetic generation workflows, and governance practices while collaborating with domain engineers and AI researchers to improve model performance and product outcomes.

Skills & qualifications

RequiredNice to have

Skills

Data ScienceMachine LearningDistributed ComputingData CurationData GovernanceStatisticsData ProcessingCollaborationExploratory Data AnalysisVersion ControlPythonSQLSparkPyTorchTensorFlowScikit‑LearnData OrchestrationVersioned StorageSynthetic Data GenerationRetrieval SystemsVector DatabasesAnnotation ToolingLLMs

Qualifications

MS or PhD in Computer Science Data Science Statistics or Related Field5–10 Years of Hands‑on Data Science or ML Data Work

Full job description

Job Title Senior Data Scientist, AI Training Data About Cognichip We build AI-native tools for semiconductor engineering, combining large proprietary models, agentic workflows, and domain-specific intelligence to help engineers design, verify, and optimize chips faster. About the Role We're looking for a Senior Data Scientist to own the data our models learn from. Semiconductor engineering data is specialized, scarce, and often tightly licensed — very different from the general text and code used to train most large models. Turning it into training-ready, evaluation-ready datasets is one of the highest-leverage inputs to our model quality. In this role, you'll design the curation, synthetic data generation, and quality-modeling work that turns raw technical material into usable datasets, running at scale on our internal data infrastructure (managed by a dedicated platform team, so you can focus on the data itself). You'll work closely with domain engineers to figure out what the models actually need, and with our AI team to connect dataset improvements to measurable model performance gains. Key Responsibilities - Curate licensed and open-source technical datasets — collection, cleaning, annotation, and integration across the engineering lifecycle- Build automated pipelines for sourcing, license classification, and normalization of public data- Design synthetic and augmented data generation workflows to keep pace with model training demand- Develop quality-modeling approaches: deduplication, contamination/leakage detection, license and PII screening, difficulty/diversity scoring, and dataset-to-eval attribution- Write large-scale distributed data processing jobs, partnering with a platform team on infrastructure needs (throughput, versioning, lineage, reproducibility)- Translate observed model weaknesses and feedback into targeted, well-sourced datasets- Build and maintain retrieval/embedding datasets that support product features- Run exploratory analysis and produce insights that guide modeling, product, and go-to-market decisions- Collaborate across engineering, AI research, product, and business teams to turn ambiguous needs into concrete datasets- Establish data governance practices: license provenance, documentation, retention, and compliance Required Qualifications - MS or PhD in Computer Science, Data Science, Statistics, or related field- 5–10 years of hands-on experience in data science or ML data work, with ownership of production datasets used by other teams- Expert Python and strong SQL, with experience processing large datasets using distributed computing frameworks (e.g., Spark)- Practical experience preparing text or code corpora for LLM training, fine-tuning, or evaluation- Solid applied statistics and ML foundations, with experience in a major ML framework (PyTorch, TensorFlow, or scikit-learn)- Familiarity with modern data orchestration and versioned storage systems- Working knowledge of data governance practices — licensing, provenance tracking, handling of confidential/contractual data- Strong ability to work with domain experts and convert ambiguous requests into delivered datasets Preferred Qualifications The following items are not required but are great bonuses: - Exposure to hardware or engineering domain data (e.g., specialized design/verification formats and workflows). You don't need deep prior expertise — just genuine interest in learning a technical domain deeply- Experience building retrieval systems: chunking strategies for technical documents, embedding models, vector databases, and evaluation of retrieval-augmented systems- Experience with synthetic data generation using LLMs, including agentic pipelines built with modern orchestration frameworks- Familiarity with annotation tooling and workflows for expert-labeled data, including inter-annotator agreement and active learning- Contributions to open-source data, hardware, or AI projects; or published research in data-centric AI, dataset curation, or code models- Prior experience at a startup or early-stage team where you helped define a data function rather than inherited an existing one What It’s Like Here - We’re a fast-moving AI startup with a collaborative, high-trust culture- We value technical excellence, ownership, and the freedom to experiment- Our best work happens when our builders and innovators work closely together to turn ambitious ideas into category defining products- We operate on a hybrid schedule with four days in office, one day remote- If you’re excited to build cutting edge tools that empower semiconductor engineers and reshape how chips are designed, you’ll feel right at home

You've read the whole posting — now see how you match it.

Senior Data Scientist, AI Training Data at Cognichip Inc. | Olive Jobs