Applied Science PhD Intern

Technology, Data & Digital · Data, AI & Analytics · Machine Learning · Artificial Intelligence · Data Science

In short

Clipchamp is seeking a PhD intern with expertise in computer vision and generative AI for a 12-week full-time internship. You will work on real-world projects involving multimodal models and agentic systems for video creation and editing, based in Brisbane, Sydney, or Melbourne.

Responsibilities

  • Research, prototype, and evaluate computer vision and multimodal models for video understanding.
  • Design and build agentic workflows using language and multimodal models to invoke tools and act over an editing timeline.
  • Measure the reliability, latency, and quality of agentic systems against real product scenarios.
  • Translate ambiguous product goals into well-defined machine learning tasks and design experiments, baselines, and metrics for rapid iteration.
  • Fine-tune, adapt, and benchmark state-of-the-art foundation models (e.g., vision-language models, diffusion models, LLM-based agents) on domain data.
  • Prepare and curate datasets for training and evaluation, reviewing data for quality and technical constraints.
  • Implement prototypes of scalable AI components and contribute to code reviews, analysis, and technical documentation.
  • Build understanding of the broader research area and industry trends, and share findings through demos, write-ups, and presentations.

Requirements

  • Currently pursuing a Doctorate degree in Computer Science, Applied Science, Statistics, or a related field, with a research focus in computer vision, machine learning, or multimodal AI.
  • Must have at least 1 semester/term remaining following the completion of the internship (graduation between August 2027 and July 2028).
  • Hands-on experience building and training deep learning models in Python with frameworks such as PyTorch, Hugging Face Transformers, or Diffusers.
  • Demonstrated experience in computer vision or video understanding, evidenced by research publications, open-source contributions, or substantial project work.

Desired Qualifications

  • Experience with agentic AI systems, including tool use and function calling, planning and reasoning, multi-step orchestration, or agent evaluation frameworks.
  • Familiarity with state-of-the-art architectures and techniques, such as transformers, attention mechanisms, vision-language models, diffusion models, transfer learning, and parameter-efficient fine-tuning.
  • Publication record at major AI conferences, including NeurIPS, ICLR, CVPR, ICML, ACL, EMNLP, ECCV/ICCV, etc.
  • Experience designing rigorous evaluation methodology for generative or agentic systems, including human evaluation and automated benchmarks.
  • Experience taking research prototypes towards production.
  • Fluency in one or more of Python, C# or TypeScript.

Benefits

  • Gain hands-on experience working at the intersection of computer vision, generative AI, and agentic models.
  • Work alongside experienced applied scientists and engineers on real-world projects.
  • Explore how AI can transform the way people create, edit, and understand video.
  • Collaborate with engineers, product managers, designers, and researchers.
  • Access to cutting-edge generative AI models, LLMs, AI tooling, and research.
  • Deepen expertise in large-scale multimodal and agentic systems.
  • Strengthen skills in experimental design, evaluation, and collaborative research engineering.
  • Make a meaningful contribution from day one.
  • Be part of Microsoft's global intern community with opportunities to build networks, explore interests, and learn from others.
#AI#computer vision#generative AI#machine learning#video editing#internship
Microsoft Logo

Company

Microsoft

Job Posted

3 weeks ago

Expires

in 5 months

Employment Type

Internship

WorkMode

On Site

Experience Level

Student

Locations

Brisbane, Australia

Sydney, Australia

Melbourne, Australia

Qualification

Doctoral

Applicants

Be an early applicant