← all areas
teach machines to see
learn computer vision
from image classification to detection, segmentation, and vision transformers.
curatedbeginner~4 weeks, part-time
computer vision fundamentals
a practical route from your first image classifier to modern detection, segmentation, and transformers — with code you can run today.
4 modules · 12 resources · checkpoint per modulestay current
see the full digest →what's new in computer vision
- VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compressionthis paper proposes visco, a method that uses large language models as intrinsic encoders to compress visual tokens. this is important for practitioners working with vision-language models, as it can significantly reduce the computational cost and memory usage associated with processing large numbers of visual tokens.
- What Images Cannot Say: Language-Guided Olfactory Representation Learningthis research investigates learning olfactory representations guided by language, acknowledging that images alone don't capture all sensory information. it offers a new direction for multimodal ai, allowing practitioners to build systems that understand and reason about non-visual sensory data.
- Straight-Path Flow Matching for Incomplete Multi-View Clusteringthis paper proposes a flow matching method for clustering data where information is incomplete across different views. it helps practitioners analyze complex datasets with missing modalities, improving clustering performance in real-world scenarios where data collection is imperfect.
want something more specific in computer vision? generate a fresh path.
generate a path