A new multimodal AI framework fuses RGB images, segmentation masks, depth maps, and text prompts to estimate robotic gripper ...