What is computer vision?

Definition

Computer vision covers tasks such as recognizing objects, locating regions and understanding visual relationships. The output depends on the specific task and system rather than a single universal interpretation of an image.

Visual analysis can miss details or assign the wrong meaning to a region. Systems should be evaluated on representative inputs and preserve uncertainty, especially when a mistaken interpretation could affect later processing steps.

ELI5

Computer vision helps software work with what appears in an image. It can identify regions or patterns, but those interpretations can still be mistaken.

For example, a document tool can locate a table in a scanned page before another step extracts its cells. If it misses part of the table, the later result may lose important information, so the complete outcome needs checking.

Does recognizing a region prove its contents were understood?

No. Locating an object or table and interpreting its meaning are different tasks.

Can visual analysis make mistakes on familiar-looking inputs?

Yes. Differences in layout, quality or content can cause errors, so representative evaluation matters.

  1. Adit Abraham in a blue shirt holds open palms beside the blue-and-white headline “DOCUMENTS TO AGENTS” on a black background.
  2. Andrew Dai beside the headline AI STILL CAN'T SEE on a black background.
Definition card for computer vision: What is computer vision?

Computer vision enables software to analyze visual inputs such as images or video and extract information about their contents or structure.