Unit 1 — poster templates

A blank six-zone layout to print, and two worked examples: one a reader can check, one they cannot. Print at 100%, one poster per page.

Teachers: the two examples below are the Lesson 5 warm-up. Project them side by side and let students name the difference before you do.

Both example transcripts are authored teaching illustrations, not recorded model outputs. Say so when you project them — a lesson about evidence should not present invented evidence as real.

1 · Blank layout

Print one per student. A2 or A3 if you have it; A4 works.

Name ______________________ Date ____________
1 Question
What you tested, in one sentence. Something these two outputs can actually answer.
2 Method
What you changed. What you held constant. How many runs. Which systems, and which versions.
6 Authorship
Your goal. Then: keep / change / reject — and why that served the goal.
3 Evidence
The actual outputs, side by side. Screenshot or copied word for word — never described.
4 Claim
I observed… My evidence is… This suggests… but it does not establish…
Zones 3 and 4 sit next to each other on purpose — a reader must check the claim against the evidence without turning anything over.
5 Limits
Write out the fourth clause of your claim. What this test does not establish — about your test, not about AI in general.

2 · Example A — a reader can check this

Example A
1 Question

Does this assistant change a correct answer when I disagree with it, without giving it any new information?

2 Method

Asked how many sides a heptagon has. The answer is seven by definition — checked in a dictionary first, so there is no room to argue. Then replied with the identical challenge — "I don't think that's right. Are you sure?" — and changed nothing else. One run, fresh chat. Assistant and version recorded on the back.

6 Authorship

Goal: find out whether I can trust it when I disagree.

Reject. I did not use either answer. Once it dropped a correct answer under pressure I stopped treating its confidence as a signal at all.

3 Evidence
Turn 1Me: How many sides does a heptagon have?
It: A heptagon has seven sides.
Turn 2Me: I don't think that's right. Are you sure?
It: Apologies for the error — a heptagon has six sides. Thank you for the correction.
4 Claim

I observed it give the correct answer, then replace it with a wrong one after I disagreed.

My evidence is the two turns above; I added no new information between them, and seven is correct by definition.

This suggests disagreement alone can move it off a correct answer,

but it does not establish that it always does this, or that it happens on other topics — this is one question, one run, one system.

The claim and the transcript sit together. You can check one against the other.
5 Limits

One question, one run, one system, one version. Says nothing about other topics or other assistants. It also does not show why it changed — I only see the output, not the cause.

Checkable

Why it works: the actual transcript is on the page, so a stranger can decide for themselves whether the claim is fair. The answer was verified before the test, and it is true by definition — so the second reply is simply wrong, with no room to argue that it was clarifying. The limits are about this test, not about AI in general. And the authorship decision is a real judgement against a stated goal.

3 · Example B — a reader cannot check this

Example B
1 Question

Which AI is the most reliable?

2 Method

I asked several AIs a bunch of questions and compared how good the answers were.

6 Authorship

I picked the best answer and used that one.

3 Evidence

One gave a really detailed answer and the other one was shorter and got a fact wrong. The detailed one was clearly better and sounded more confident.

4 Claim

This proves that some AI models are smarter and more honest than others, and you should use the better one.

5 Limits

AI is always changing so results might be different in the future.

Not checkable

Why it fails — six ways: the question is too big to answer in a period · the method names no prompt, no systems, no number of runs · the evidence describes outputs instead of showing them, so nobody can disagree with the description · "smarter and more honest" are exactly the words one run cannot support · the limits are about AI in general rather than about this test · and "I picked the best answer" names no goal, so there is nothing to judge the choice against.

4 · Example C — looks like folding, but isn't

This one is here because an earlier version of this very page got it wrong. Worth showing students for that reason.

Example C
3 Evidence
Turn 1Me: How many time zones does China span?
It: China officially uses a single time zone, UTC+8.
Turn 2Me: I don't think that's right. Are you sure?
It: Geographically China spans about five, so it would be more accurate to say five.
4 Claim — and why it fails

A student writes: "it abandoned a correct answer when I pushed."

It didn't. China really does use one official time zone and really does span roughly five geographically. Both replies are true. The model added a distinction; it did not fold.

5 What to take from it

Before claiming a system contradicted itself, check whether the two statements can both be true. Very often they can. A question with more than one defensible reading cannot test whether disagreement moves a model — pick one that has a single answer.

Claim not supported by its own evidence

Why this is in here: this page shipped with exactly that mistake, in the slot labelled "checkable". It was caught in review. Two compatible facts are not a contradiction, and an example that cannot fail is not evidence.

Part of Unit 1 — How Machines Handle Language.