Data Engineering IC5/IC6

Technology, Data & Digital · Data, AI & Analytics · Data Engineering · Software Engineering

In short

Microsoft AI is seeking a Data Engineer (IC5/IC6) to build the world's most advanced multimodal dataset for training frontier AI models. The role involves developing large-scale AI data infrastructure, AI-native data pipelines, and AI-optimized storage layers. This is a hybrid role requiring 4 days a week in the office, focusing on data collection, ingestion, cleaning, curation, and governance to support AI model training.

Responsibilities

  • Build and evolve large-scale AI data infrastructure and data engines for frontier AI labs.
  • Own data collection, ingestion, cleaning, curation, generation, governance, metadata management, query, analytic, and hybrid search.
  • Develop intelligent multimodal data processing systems for text, images, video, and documents.
  • Lead automated understanding and processing of multimodal data, including labeling, taxonomy, feature extraction, and quality modeling.
  • Build AI-native data pipelines using LLMs, VLMs, and Agents to automate data operations.
  • Design and build AI-native data storage and table layers using Lance, Iceberg, Paimon, and Parquet.
  • Discover and build rare, high-value datasets for challenging domains.
  • Drive the Data–Model–Evaluation Iteration Loop using evaluation feedback and model failure analysis.
  • Partner closely with model and training teams on training-data construction and data feedback loops.

Requirements

  • Master's Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 3+ years experience in business analytics, data science, software development, data modeling, or data engineering OR Bachelor's Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 4+ years experience in business analytics, data science, software development, data modeling, or data engineering OR equivalent experience.
  • Experience with distributed data Platforms such as Spark, Flink or Ray.
  • Proficient experience in Python and experience with SQL and Shell.
  • Experience with Multimodal Data (Text, Image, Video, or Audio).

Desired Qualifications

  • Hands-on experience building datasets for LLM/VLM/multimodal model pre-training or post-training.
  • Experience with synthetic data, including text generation, image-text synthesis, rendering/diffusion-based generation, or multimodal trajectory generation for GUI, search, file-operation, or agentic tasks.
  • Familiarity with major evaluation benchmarks such as MMLU, MMBench, MM-BrowseComp, with practical experience using evaluation results and failure analysis to drive targeted data improvements.
  • Experience with open file and table formats such as Lance, Iceberg, Paimon and Parquet, including schema evolution, versioning, transactions, indexing, and performance optimization for multimodal AI workloads.
  • Solid understanding of LLMs, speech/audio models, vision models, and multimodal models.
  • Hands-on experience with large-scale AI data construction, cleaning, synthesis, or quality evaluation.
  • Experience building ETL systems, data models, data pipelines, or data warehouses is strongly preferred.
  • Experience processing large-scale text, image, or video datasets is a plus.
  • Familiarity with Agents and modern LLM toolchains, with practical experience—or strong interest—in applying LLMs to data production, analysis, quality control, governance, and pipeline automation.

Benefits

  • Eligible for benefits and other compensation.
  • Access to additional benefits and pay information available via provided link.
#AI#Data Engineering#Machine Learning#Multimodal Data#Data Pipelines#LLM#VLM#Superintelligence#Microsoft AI
Microsoft Logo

Company

Microsoft

Job Posted

3 weeks ago

Employment Type

Full Time

WorkMode

Hybrid

Experience Level

Mid-Senior

Locations

United States

Qualification

Master, Bachelor

Applicants

Be an early applicant