Applied Science PhD Intern
Technology, Data & Digital · Data, AI & Analytics · Machine Learning · Artificial Intelligence · Data Science
In short
Clipchamp is seeking a PhD intern with expertise in computer vision and generative AI for a 12-week full-time internship. You will work on real-world projects involving multimodal models and agentic systems for video creation and editing, based in Brisbane, Sydney, or Melbourne.
Responsibilities
- Research, prototype, and evaluate computer vision and multimodal models for video understanding.
- Design and build agentic workflows using language and multimodal models to invoke tools and act over an editing timeline.
- Measure the reliability, latency, and quality of agentic systems against real product scenarios.
- Translate ambiguous product goals into well-defined machine learning tasks and design experiments, baselines, and metrics for rapid iteration.
- Fine-tune, adapt, and benchmark state-of-the-art foundation models (e.g., vision-language models, diffusion models, LLM-based agents) on domain data.
- Prepare and curate datasets for training and evaluation, reviewing data for quality and technical constraints.
- Implement prototypes of scalable AI components and contribute to code reviews, analysis, and technical documentation.
- Build understanding of the broader research area and industry trends, and share findings through demos, write-ups, and presentations.
Requirements
- Currently pursuing a Doctorate degree in Computer Science, Applied Science, Statistics, or a related field, with a research focus in computer vision, machine learning, or multimodal AI.
- Must have at least 1 semester/term remaining following the completion of the internship (graduation between August 2027 and July 2028).
- Hands-on experience building and training deep learning models in Python with frameworks such as PyTorch, Hugging Face Transformers, or Diffusers.
- Demonstrated experience in computer vision or video understanding, evidenced by research publications, open-source contributions, or substantial project work.
Desired Qualifications
- Experience with agentic AI systems, including tool use and function calling, planning and reasoning, multi-step orchestration, or agent evaluation frameworks.
- Familiarity with state-of-the-art architectures and techniques, such as transformers, attention mechanisms, vision-language models, diffusion models, transfer learning, and parameter-efficient fine-tuning.
- Publication record at major AI conferences, including NeurIPS, ICLR, CVPR, ICML, ACL, EMNLP, ECCV/ICCV, etc.
- Experience designing rigorous evaluation methodology for generative or agentic systems, including human evaluation and automated benchmarks.
- Experience taking research prototypes towards production.
- Fluency in one or more of Python, C# or TypeScript.
Benefits
- Gain hands-on experience working at the intersection of computer vision, generative AI, and agentic models.
- Work alongside experienced applied scientists and engineers on real-world projects.
- Explore how AI can transform the way people create, edit, and understand video.
- Collaborate with engineers, product managers, designers, and researchers.
- Access to cutting-edge generative AI models, LLMs, AI tooling, and research.
- Deepen expertise in large-scale multimodal and agentic systems.
- Strengthen skills in experimental design, evaluation, and collaborative research engineering.
- Make a meaningful contribution from day one.
- Be part of Microsoft's global intern community with opportunities to build networks, explore interests, and learn from others.
#AI#computer vision#generative AI#machine learning#video editing#internship