Tool 24 · Session 2 · Images · Shared representations
Image–Caption Match Lab
Predict which caption best matches an image, reveal a transparent similarity ranking, then change the words and test again. Use the human-decision layer to inspect who supplied the categories, captions, and definition of a “match.”
Candidate captions
Which caption will land closest?
Choose before revealing. The comparison is against this candidate set—not every possible sentence.
Choose a caption to make your prediction.
Authored teaching-score ranking
Inspect the human-created label layer
| Candidate set | |
|---|---|
| Available cues | |
| Definition of match | Higher score means more of the author-selected scene cues were named, minus a small penalty for cues assigned to a different scene. |
| What gets excluded |
WordNet and ImageNet make a related historical layer visible at a much larger scale: people designed category relationships, selected images, supplied annotations, and decided what counted as correct. This lab does not reproduce either dataset.
Investigation note
Make one bounded claim
Describe this authored mechanism first. Then name the evidence a real model claim would require.
Debrief
What moved—and who decided?
- Which caption did you predict, and which cue did most of the work?
- What happened when you changed one word or phrase?
- How would a different candidate-caption set change the result?
- Who chose the scene cues, labels, examples, and definition of “closest”?
- What would you need to run before making a claim about CLIP rather than this simulation?