This Space combines Grounding DINO for open-vocabulary object detection with SAM for text-prompted image segmentation. Enter comma-separated labels such as cat, dog, sidewalk, crosswalk.
cat, dog, sidewalk, crosswalk