Back to jobs

Internship - Machine Learning Research Engineer

Perplexity
Berlin (Hybrid)
Internship
AI tools:
PyTorch
DeepSpeed
FSDP
You apply on Perplexity's own careers site

You’ll research and improve search and retrieval quality through large-scale model training, representation learning, and RAG pipeline development. The internship focuses on retrieval and ranking and calls for experience with PyTorch and distributed training.

Internship
On-site
Internship

Skills & Expertise

PyTorch
PyTorch Distributed
DeepSpeed
FSDP
Distributed training
Representation learning
Contrastive learning
RAG

Key Responsibilities

Train and optimize large-scale retrieval and ranking models using distributed training and hardware acceleration.

Research representation learning approaches for search, including multilingual, contrastive, and multimodal modeling.

Build and optimize RAG pipelines for grounding and answer generation.

Full Description

Internship Program Berlin

Internship program: 12 - 24 weeks, full-time, in-person in the Berlin office.

Responsibilities

• Relentlessly push search quality forward — through models, data, tools, or any other leverage available.

• Train, and optimize large-scale deep learning models using frameworks like PyTorch, leveraging distributed training (e.g., PyTorch Distributed, DeepSpeed, FSDP) and hardware acceleration, with a focus on retrieval and ranking models.

• Conduct research in representation learning, including contrastive learning, multilingual, evaluation, and multimodal modeling for search and retrieval.

• Build and optimize RAG pipelines for grounding and answer generation.

Qualifications

• Understanding of search and retrieval systems, including quality evaluation principles and metrics.

• Strong proficiency with PyTorch, including experience in distributed training techniques and performance optimization for large models.

• Interested in representation learning, including contrastive learning, dense & sparse vector representations, representation fusion, cross-lingual representation alignment, training data optimization and robust evaluation.

• Publication record in AI/ML conferences or workshops (e.g., NeurIPS, ICML, ICLR, ACL, EMNLP, SIGIR).

Applications are handled on Perplexity's site