Tool 24 · Session 2 · Images · Shared representations

Image–Caption Match Lab

Predict which caption best matches an image, reveal a transparent similarity ranking, then change the words and test again. Use the human-decision layer to inspect who supplied the categories, captions, and definition of a “match.”

Big idea: an image and a caption can become comparable numerical representations without the system searching for an existing finished picture.
Teaching simulation. The illustrations, cue vocabulary, weights, and scores on this page were authored for instruction. They are not CLIP outputs or measurements from a live model. They make a simplified matching rule inspectable; claims about CLIP require documented CLIP results.
1 · Predict
2 · Reveal
3 · Revise

Candidate captions

Which caption will land closest?

Choose before revealing. The comparison is against this candidate set—not every possible sentence.

Choose a caption to make your prediction.

Change the words

Try changing one noun, attribute, action, or context phrase. Unrecognized words remain visible instead of being silently treated as understood.

Inspect the human-created label layer
Candidate set
Available cues
Definition of matchHigher score means more of the author-selected scene cues were named, minus a small penalty for cues assigned to a different scene.
What gets excluded

WordNet and ImageNet make a related historical layer visible at a much larger scale: people designed category relationships, selected images, supplied annotations, and decided what counted as correct. This lab does not reproduce either dataset.

Investigation note

Make one bounded claim

Describe this authored mechanism first. Then name the evidence a real model claim would require.

Debrief

What moved—and who decided?

  • Which caption did you predict, and which cue did most of the work?
  • What happened when you changed one word or phrase?
  • How would a different candidate-caption set change the result?
  • Who chose the scene cues, labels, examples, and definition of “closest”?
  • What would you need to run before making a claim about CLIP rather than this simulation?