One thing this project reinforced is that "computer vision still has enormous possibilities ahead"
It is not just about taking the latest YOLO, SAM, or any other famous model and putting it into an application. The real opportunity lies in understanding what visual perception actually means.
Vision is fundamentally different from language. The visual world is continuous, highly contextual, spatial, temporal, and incredibly rich in information. Humans have evolved to make sense of this complexity almost effortlessly recognizing objects, understanding relationships, interpreting motion, depth, intent, and context, often without even consciously thinking about it.
Projects like this are a small step toward exploring that much larger space: how machines can perceive the world and turn that perception into meaningful interaction.
For me, that's what makes computer vision such an exciting field we are still very far from truly understanding vision.

Log in or sign up for Devpost to join the conversation.