Join an R&D team advancing computer vision for robotic endoscopic video technologies, focusing on vision foundation/diffusion models, feature detection, and multimodal video analysis.
Essential Duties
- Explore and experiment with state-of-the-art computer vision models.
- Prototype algorithms and evaluate performance on public/proprietary datasets.
- Conduct literature surveys and summarize key findings in reports and presentations.
Required Skills & Experience
- Hands-on expertise in computer vision, deep learning, video analysis.
- Knowledge in vision-language models, diffusion models, feature detection, or multimodal learning.
- Proficiency in programming with Python or C++.
- Experience with PyTorch, OpenCV, DINO/CLIP, HuggingFace Transformers.
- Strong research and communication abilities.
- Self-driven; ability for rapid prototyping.
Education: currently enrolled in PhD or Master's in Computer Science, Robotics, Engineering, or related field, returning to degree program Spring 2027.
Availability: full-time (40 hours/week) for 10-12 weeks starting August/September 2026.