Home/ Sessions/ Video/ Recap
03recap

Session 3 · Video · Saturday July 25

How machines move

Participants used a human-drawing teaching analogy, creating short animations frame by frame under two reference conditions. They performed one simplified coherence constraint rather than operating or reproducing an image or video model. Dr. Emily Thomforde then traced how the unexplainable core of these systems became something we now treat as magic rather than a problem.

10+ contributors 9–11 am PT Unlisted recording link AI-assisted recap · facilitator-reviewed

Are you named, quoted, or represented here? Use the consent form to choose whether your name, words, or described work may appear in the written recaps and whether words you specifically approve may be used in promotion — with your name, anonymously, or not at all. Session recordings are shared separately by unlisted YouTube link; anyone with the link may be able to view or reshare them.

How this recap was made. AI-assisted tools helped organize and draft this page from facilitator notes, session transcripts and captions, and chat exports. Those source records included participant display names, speaker labels, attributed messages, and descriptions of participant work. The facilitator reviewed public claims, quotations, and attributions against the source records. Consent controls what may appear publicly; it does not undo the earlier use of those records during drafting. Read the full workflow and consent protocol.

Watch the recording

Open the Session 3 recording on YouTube (unlisted)

Start here

5 min

Just the highlights

Read the session overview and the Coherence Animator debrief — the activity the whole session built toward.

15 min

The full arc

Add what the room noticed, Dr. Thomforde's talk on the axiology of mystery, and the participant Q&A.

Deeper path

Read everything

Include the assignment reviews, the teaching notes, and the full links & resources.

Session overview

The camp's one-line argument arrived whole this week: fluent text is not necessarily true, plausible images are not necessarily neutral, and smooth video is not evidence that an event happened. Session 3 focused on the third line and one simplified way to investigate coherence across time.

In the Coherence Animator, participants created short hand-drawn animations under two reference conditions: the previous drawing alone, or the previous drawing plus a fixed opening anchor. The comparison made a question about continuity tangible: how does a stable visual reference change what people preserve or alter from frame to frame? It was a teaching analogy for one coherence problem, not a simulation of a particular video model. Guest Dr. Emily Thomforde then gave the ethical and epistemic backbone in a talk called Artificial Intelligence and the Axiology of Mystery.

Two takeaways from the activity:

  1. In this activity, adding the opening frame as an anchor sometimes steadied background details, but it also pulled some participants toward repeating the opening pose.
  2. Keyframes, timing, and foreground/background order are concepts human animators already use. The parallels give us useful questions to ask about video systems; they do not show that a model represents or solves these problems in the same way.

Who was in the room

A smaller, engaged group this final week. Participants are identified by first name or chosen display name only:

Ryan Shane Yunseo Randy Andy Judy Sharleen Meghan Hope Andrea Wilson Dr. Emily Thomforde guest

Shane, in chat, when both animation runs went sideways: "Just like real life, either way you're doomed lol."

— Chat, 1:30

Assignment reviews (Sessions 1 & 2)

We opened by looking at submitted work — a chance to turn loose intuitions into something concrete to carry into the school year:

  • Ryan (Session 1): drew parallels between narrative game design and simple rule-based dialogue systems like ELIZA, arguing that the versatility of LLMs can overshadow simpler tools that do a lot with a little. His broader point: games make learning memorable because interaction takes more effort than passive consumption. (Judy pushed back in chat that films and plays can be just as transformative — a good tension to keep.)
  • Randy (Session 1): ran an AI summer camp for middle-schoolers using Raspberry Pi's Experience AI curriculum and Machine Learning for Kids. Students built image classifiers (tomatoes vs. apples, dogs vs. cats) and a "which animal do you look like" webcam model. His planned next lesson: using text analysis to help students identify features of fake news.
  • Cierra (Session 2, submitted; absent): compared two Claude-generated frogs — a plain "photo frog" and an "unsettling frog" — showing how one descriptor shifts the model's defaults (darker color, smaller yellow eyes, more warts) while the graphic style and centered placement stay stubborn. In chat, Meghan asked why the frogs weren't more realistic; Ryan suggested they were drawn with SVG graphics rather than generated by diffusion — a reminder to ask which production path actually made an artifact before reading mechanism off it.
  • Andy (Session 2): tried to get Gemini to explain how the text modifies the image in diffusion, and hit the honest wall — explanations from an LLM smooth over exactly the friction where a good teacher would catch a misconception. The key that unlocked it: diffusion is steered by the tokenized/vectorized text space the prompt provides.
  • Ray (Session 1, submitted; absent): a screenshot from the room-guesses-vs-model-prediction activity.

"You can start to go down this misconception rabbit hole … and that's where a good teacher comes in — gets inside your head and says, oh, this is how you're understanding it. I don't think the LLMs are very good at that yet."

— Andy, on explanatory conversations with LLMs

Session 2 feedback, briefly

For most respondents, image mechanics were brand new. What landed: diffusion as denoising — chiseling coherence out of noise — and the idea that a model works with pixel data, not objects or a scene graph. What stayed fuzzy: the concrete route from text tokens to pixels, and where the human label layer (WordNet, ImageNet) sits in the pipeline. One respondent described diffusion as pointillism — building an image dot by dot — which we connected to photography arriving at the same historical moment: artists mimicking systems of seeing. Judy added that generative images need very specific lighting instruction; the model does not author a lighting plot on its own.

The Coherence Animator: performing one constraint, twice

The centerpiece. Participants opened the Coherence Animator and drew a five-frame animation of one shared prompt — a person at work — under two conditions, so the constraint itself would reveal something:

  • Run A — previous frame only. You can see only the frame you just drew.
  • Run B — anchor + previous frame. You can see both the opening frame and the frame you just drew. The additional reference may help preserve some details, but it also adds another constraint to reconcile.

Run A and Run B changed which earlier drawings remained visible. The comparison was designed to make continuity across frames easier to notice.

Then participants compared their own Run A and Run B at different playback speeds and picked one feature to watch. (Most first frames were a stick figure at a desk with a laptop — "nothing profound; it's about what work means to most of us.") The honest caveat, named aloud: this is not a controlled experiment — Run B also comes second, with more practice — so it fails scientific-method muster, but the friction still makes the coherence problem legible.

"My measure of success is whether someone can recognize what action or idea I'm trying to represent."

— Randy, in chat

What the room noticed

The shares did the teaching. They surfaced several hypotheses that we can compare with documented animation and video-generation workflows:

  • Ryan found Run B no better in raw quality, but he changed his drawing order — static background first in Run A, character first in Run B — which shifted line quality and where his attention went. Separating foreground from background is a real human workflow move; whether a model splits work the same way is a hypothesis to check against documented pipelines, not something the activity shows. (His character's arm-to-chin gesture stayed coherent across both runs; he "spent his tokens" on the figure and had little left for the background.)
  • Shane, a professional animator and timing supervisor, saw Run B's background elements come out slightly more coherent — "everything that isn't important is a little more stable when I had something to anchor me." His deeper point: animation timing is semantically driven. Holding a drawing a few extra frames implies thinking, which implies intention, and changes the meaning of the whole sequence.
  • Yunseo found Run A pulled toward vivid, active motion (only the last frame to react to), while Run B pulled toward near-static, deliberate poses — like a "Wimpy Kid" or Powerpuff Girls diary telling a story through still images. Anchoring reduces how much can happen.
  • In a share attributed to Sharleen in the transcript, the participant gave herself a start-and-end-frame constraint and noticed she rushed the middle ("let the fire take over the screen so I have less to draw") — which prompted a facilitator hypothesis, to check against real workflows rather than shown by the activity, that some systems lean on keyframes and interpolation and might need a separate system to avoid over-converging on them.

"Once you get to a certain amount of rules and a certain amount of text, you begin to build software on accident … and once you lose the mental model and just have to trust the agent is doing what you think, I lose everything."

— Shane, on dictating an animation system to an LLM

Shane also described trying to give an LLM the four elements of any move — start pose, end pose, anticipation, overshoot — and watching it overshoot subtle moves because it followed the discrete rule without any sense of the relativity of spacing. The room used human animation concepts as a vocabulary for asking how video pipelines handle timing, spacing, and scene consistency. The activity does not establish that a model represents those problems as a human animator does. The room also circled a desired capability: editing one part of a frame without regenerating the whole. Some node-based workflows (fuser.studio, ComfyUI) connect multiple stages to support targeted generation, masking, or compositing, but the exact production path varies.

Guest spotlight

Dr. Emily Thomforde (she/her · "Em")

Machine learning PhD, University of Edinburgh (School of Informatics) · NLP / computational linguistics · AI educator for teachers, administrators, and policymakers · self-described AI skeptic who does not use AI to generate content

Em's talk, Artificial Intelligence and the Axiology of Mystery, opened with a distinction (from AI for Education, later adopted by the California Department of Education): learning about AI — her field, from CS education and engineering — versus learning with AI, which belongs to educational technology and the learning sciences. She educates the people who have to make the "learning with" decisions.

The explainability problem. "No one really knows how AI works" is false but points at something real. Like the water cycle: we can model it, predict it, simulate it — but ask where did this one drop come from and there is no answer. Neural networks keep no paper trail of which training data led to a given output. That was once disqualifying — you couldn't ship a product or earn a PhD on a system with no accountability — and it is still unsolved. What changed with ChatGPT, she argued, was not technical progress but a cultural and economic shift that now treats mystery as acceptable, even desirable.

The framework. Any discipline has three parts: ontology (what AI is and isn't), epistemology (how it works — including the explainability problem), and axiology (what we value — ethics and aesthetics). AI literacy usually teaches only a thin slice of the ethics; without the ontology and methodology underneath, that ethics is thin and produces poor policy. Her advice for educators: start with bias (where people already have footing), then work up to methodology and ontology. When she quizzed her own university's AI Literacy work group with a research-backed instrument, they bombed it — half thought a chatbot can fully explain its own reasoning.

Against magical thinking. A Journal of Marketing study (published online January 2025) found that people who know less about AI are more likely to adopt it; a later American Marketing Association summary of the work urged educators to "enhance understanding without eroding the excitement that drives adoption" — casting educators as drivers of consumerism. Thomforde argued that there is not yet an evidence base showing AI tools benefit students; in her account, gaps in evidence, technical understanding, and domain expertise leave room for hope or fear to become magical thinking. Her alternative: speculative fiction. Decades of writers have already imagined technology in education; works like Ender's Game and Klara and the Sun give structured imagination for policy.

"Where did this drop of water come from? If you can't tell me, you don't know. That's what people mean by 'no one really knows how AI works' — there's no paper trail."

— Dr. Emily Thomforde, on the explainability problem

Questions from the room

The Q&A turned the guest's framework back toward participants' own decisions about dependence, reasoning, and ownership:

  • Shane described AI as holding an "unnaturally large" place in his mind because its uses remain difficult to bound, then asked whether it was safe to call it a permanent part of life. Thomforde described this as an unusually open-ended exploration stage and argued that the uncertainty is worth retaining: dependence can become monetized, so this is also a moment to build enough expertise to understand and, where possible, own the tools one relies on.
  • Wilson asked whether symbolic AI could play a positive role in a field dominated by neural networks. Thomforde identified neuro-symbolic work as a possible next frontier and argued that symbolic processing could add explicit abstractions, logic, and modularity to sub-symbolic pattern recognition. This was her research judgment, not a settled conclusion of the workshop.
  • Randy asked whether educators and other non-specialists might eventually create or adapt models for their own use cases. Thomforde pointed to open mathematics, open-source communities, and model-sharing platforms such as Hugging Face, while distinguishing freely available knowledge from the compute and technical capacity still required to train and maintain systems.

Teaching notes

Make one coherence constraint tangible

Let learners draw a short sequence under two reference conditions, then compare what persisted and what changed. The activity makes coherence across time available for discussion; it does not run a model or demonstrate how a particular video architecture works.

The two-run comparison is the lesson

Run A (previous frame only) vs Run B (anchor + previous) makes the effect of a changed human reference condition physical. Name the caveat that it isn't a controlled experiment — Run B also comes second — and compare anyway. Use the result to ask what participants tried to preserve and how the visible references shaped those choices, not to evaluate a particular video system.

Add an onion-skin

Randy's suggestion, worth building: let the previous frame's ghost carry the static elements forward so students aren't redrawing the desk every frame. It removes busywork and focuses attention on what actually changes — the motion.

Mine the animators in the room

Shane's timing-supervisor read — a held beat implies intention; subtle moves need proportional spacing an LLM doesn't feel — is the richest bridge to why models struggle. If you have a domain expert (animation, film, music), give them room; their craft language predicts the machine's failure modes.

Temporal Telephone, on paper

The in-person version needs no screens: each person animates a frame and hands it to the next. Separate the "models" so groups can or can't see each other's work, add roles (someone picks colors, someone tracks a feature) — it turns coherence and drift into a game.

Action items

  • Participants: fill out the Session 3 feedback form — including whether you'd present at an August showcase and which Saturday works (the 8th or the 15th)
  • Participants: choose one route from the Session 3 between-session fieldwork. If you're joining the showcase, bring one rough thing to present — a tool, activity, artifact, critique, or set of questions from any of the three sessions
  • UK educators working with 13–16-year-olds: see Andy's shared invitation to register interest in piloting Experience AI's Design Thinking and AI workshop before the end of September 2026 — a one-day workshop that can run in a classroom, summer school, or Code Club, followed by feedback to the Experience AI team
  • Host: schedule the optional August showcase from feedback responses (needs ~3–4 presenters to be worthwhile), and email an end-of-month August office hour to every cohort
  • Host: consider an onion-skin option in the Coherence Animator, and a "fix one part of a frame" concept for a future tool

Session tools (used live)

This summary was drafted with AI assistance and reviewed by a human facilitator. Quotations and attributions were checked against the session transcript and chat export.