Image analysis selection: visual question answering vs field extraction vs dense captioning
This sheet compares Visual question answering, Field extraction, Dense captioning, Concise captioning, OCR.
| Criterion | Visual question answering | Field extraction | Dense captioning | Concise captioning | OCR |
|---|---|---|---|---|---|
| Returns | A grounded answer to an open question about an image | Classified categories plus generated values per a schema | A caption and region for each distinct object | One overall sentence for the whole image | Printed and handwritten text read from the image |
The verdict, the full comparison, 3 rules and 2 traps are part of AI-103 access. Unlock AI-103.