Multilingual AI Quality Specialist
Data Science · Machine Learning · Artificial Intelligence
In short
Spotify is seeking a Multilingual AI Quality Specialist to enhance AI-powered experiences across languages and markets. This role involves defining quality frameworks, leading dataset curation, and analyzing AI performance to drive improvements in AI-generated, translated, and recommended content.
Responsibilities
- Define quality frameworks, evaluation rubrics, thresholds, and methodologies for multilingual AI experiences.
- Design and execute structured evaluations for AI-generated, AI-translated, AI-curated, and recommendation-driven experiences.
- Lead multilingual dataset curation, annotation, enrichment, and ground-truth creation to support AI model development and evaluation.
- Analyze evaluation results, identify quality gaps, and provide actionable recommendations to improve multilingual AI quality.
- Support LLM-as-a-judge workflows, evaluator calibration, and human-AI agreement studies.
- Partner closely with Product, Engineering, Data Science, Research, Localization, vendors, and market experts to improve AI quality signals and inform launch decisions.
- Document best practices and help define quality standards across languages, markets, and AI use cases.
- Contribute to building scalable evaluation capabilities that support the next generation of AI-powered experiences across Spotify.
Requirements
- Experience in multilingual quality evaluation, localization, data curation, annotation, AI evaluation, or related fields, including text-to-text and text-to-speech experiences.
- Understanding of language quality, cultural relevance, content quality, and user experience across multiple languages and markets.
- Experience designing or conducting structured evaluations using quality rubrics, audits, annotation projects, or review methodologies.
- Comfortable using qualitative and quantitative data to identify trends, measure quality, and make recommendations.
- Familiarity with large language models (LLMs), generative AI evaluation, human-in-the-loop workflows, or LLM-as-a-judge methodologies.
- Ability to work through ambiguity and turn complex quality challenges into practical evaluation strategies.
- Effective communication skills and ability to thrive in highly cross-functional environments, collaborating with technical and non-technical partners alike.
Desired Qualifications
- Experience with recommendation systems, personalization, search, ranking, machine translation, generative AI, dataset creation, annotation operations, evaluator calibration, prompt testing, model evaluation, SQL, Python, dashboards, or annotation platforms is a plus.
Benefits
- Extensive learning opportunities, through our dedicated team, GreenHouse.
- Flexible share incentives letting you choose how you share in our success.
- Global parental leave, six months off - for all new parents.
- All The Feels, our employee assistance program and self-care hub.
- Flexible public holidays, swap days off according to your values and beliefs.
#AI#Quality#Multilingual#Localization#LLM