Image analysis selection: visual question answering vs field extraction vs dense captioning

This sheet compares Visual question answering, Field extraction, Dense captioning, Concise captioning, OCR.

CriterionVisual question answeringField extractionDense captioningConcise captioningOCR
ReturnsA grounded answer to an open question about an imageClassified categories plus generated values per a schemaA caption and region for each distinct objectOne overall sentence for the whole imagePrinted and handwritten text read from the image

The verdict, the full comparison, 3 rules and 2 traps are part of AI-103 access. Unlock AI-103.