Data Engineering IC5/IC6
Teknik, data och digitalt · Data, AI och analys · Data engineering · Mjukvaruutveckling
I korthet
Microsoft AI is seeking a Data Engineer (IC5/IC6) to build the world's most advanced multimodal dataset for training frontier AI models. The role involves developing large-scale AI data infrastructure, AI-native data pipelines, and AI-optimized storage layers. This is a hybrid role requiring 4 days a week in the office, focusing on data collection, ingestion, cleaning, curation, and governance to support AI model training.
Ansvarsområden
- Build and evolve large-scale AI data infrastructure and data engines for frontier AI labs.
- Own data collection, ingestion, cleaning, curation, generation, governance, metadata management, query, analytic, and hybrid search.
- Develop intelligent multimodal data processing systems for text, images, video, and documents.
- Lead automated understanding and processing of multimodal data, including labeling, taxonomy, feature extraction, and quality modeling.
- Build AI-native data pipelines using LLMs, VLMs, and Agents to automate data operations.
- Design and build AI-native data storage and table layers using Lance, Iceberg, Paimon, and Parquet.
- Discover and build rare, high-value datasets for challenging domains.
- Drive the Data–Model–Evaluation Iteration Loop using evaluation feedback and model failure analysis.
- Partner closely with model and training teams on training-data construction and data feedback loops.
Krav
- Master's Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 3+ years experience in business analytics, data science, software development, data modeling, or data engineering OR Bachelor's Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 4+ years experience in business analytics, data science, software development, data modeling, or data engineering OR equivalent experience.
- Experience with distributed data Platforms such as Spark, Flink or Ray.
- Proficient experience in Python and experience with SQL and Shell.
- Experience with Multimodal Data (Text, Image, Video, or Audio).
Önskade kvalifikationer
- Hands-on experience building datasets for LLM/VLM/multimodal model pre-training or post-training.
- Experience with synthetic data, including text generation, image-text synthesis, rendering/diffusion-based generation, or multimodal trajectory generation for GUI, search, file-operation, or agentic tasks.
- Familiarity with major evaluation benchmarks such as MMLU, MMBench, MM-BrowseComp, with practical experience using evaluation results and failure analysis to drive targeted data improvements.
- Experience with open file and table formats such as Lance, Iceberg, Paimon and Parquet, including schema evolution, versioning, transactions, indexing, and performance optimization for multimodal AI workloads.
- Solid understanding of LLMs, speech/audio models, vision models, and multimodal models.
- Hands-on experience with large-scale AI data construction, cleaning, synthesis, or quality evaluation.
- Experience building ETL systems, data models, data pipelines, or data warehouses is strongly preferred.
- Experience processing large-scale text, image, or video datasets is a plus.
- Familiarity with Agents and modern LLM toolchains, with practical experience—or strong interest—in applying LLMs to data production, analysis, quality control, governance, and pipeline automation.
Förmåner
- Eligible for benefits and other compensation.
- Access to additional benefits and pay information available via provided link.
#AI#Data Engineering#Machine Learning#Multimodal Data#Data Pipelines#LLM#VLM#Superintelligence#Microsoft AI