Job Description
Create high-quality datasets for training and evaluation; run experiments on new datasets (data ablations) to assess their impact and determine the most effective data Develop and maintain scalable data pipelines for data ingestion, pre-processing, filtering, and annotation Analyse real-world multimodal datasets to assess quality, diversity, relevance, and identify areas for improvement Build tools and workflows for dataset auditing, visualization, and versioning Collaborate with Safety, Ethics, and Governance teams to ensure datasets meet standards for quality, privacy, and responsible AI practices Bachelor's Degree in AI, Computer Science, Data Science, Statistics, Physics, Engineering, or a related technical field. Proficiency in Python. Experience with distributed data-processing frameworks such as Spark, Ray and workflow-orchestration tools such as Airflow. Experience with processing datasets at many petabyte scale and with trillion rows. Experience building datasets for training foundation models, including language or multimodal models. Strong experience in data analysis, data engineering, or both. Ability to communicate technical findings clearly and effectively to research, engineering, and product teams. Experience evaluating dataset quality and measuring the impact of data through controlled model-training experiments. Master's degree in Computer Science or a related technical field, or equivalent experience. Experience working with large-scale, real-world datasets that are unstructured or semi-structured Experience evaluating dataset quality and measuring the impact of data through controlled model-training experiments. Experience with multimodal data, such as image, video, or audio data.