You built the toy by hand. Now the real thing — several of them, given exactly the same job.
Pick a task you actually want done. Then write down what you expect before you run anything.
style? accuracy? length? confidence?
what will they all do the same way?
where do you expect a refusal?
Why first? If you look and then decide what you expected, you will always have expected what you saw. Writing it down first is the only defense against that.
Convention 3: one variable at a time. Same wording, same punctuation, same everything. The model is the only thing that changes.
| Log for each model | Model A | Model B | Model C |
|---|---|---|---|
| First sentence, exactly | |||
| Expressed uncertainty? | |||
| Refused anything? | |||
| A fact you can check |
Three things that make this a comparison and not an anecdote: start a fresh chat for each model so nothing carries over · write down which system and which version · pick a fact you can check against a named source, not against memory.
Screenshot every output. The screenshots are the evidence.
A stranger could verify it from your screenshot.
"Model B wrote 240 words. One date it gave was wrong."
A reasonable person could look at the same screenshot and disagree.
"Model B was careless."
The self-check: could someone see my evidence and still disagree? If no — it is an observation you have mislabelled.
These systems are not deterministic. The same prompt can produce different output five minutes from now.
One run per model establishes that something happened once. It does not establish that a model is "better," "smarter," or "more honest."
If your claim contains one of those three words — rewrite it.
First write your goal — what did you actually want this output to do? Then make a call on the best one you got.
Use it as it stands. Say what about it already served your goal.
Edit it. Mark exactly what you changed and why the original missed.
Bin it and write your own. Say what none of them gave you.
All three are authorship — but only if you can defend the call against the goal you wrote down first.
I observed
My evidence is
This suggests
but it does not establish
One run per model is one run per model. If your claim contains "better", "smarter" or "more honest" — rewrite it.
Your log and screenshots of each output, plus your keep/change/reject decision and the goal it served.
Lastname-U1D4
A claim with no screenshot behind it is not evidence.
Thinking: copy your circled answers and your claim into the Classroom response box. Evidence: attach your log, your screenshots, and your keep/change/reject decision.
Open the attachment to check it is readable, then turn in.
You just ran somebody else's comparison.
Now you design your own.
You pick the question. You run the test. You build a poster that carries the claim and the evidence on the same page.
Come with one question you actually want answered.