The class in brief
First guest of the semester. Geoff Gibbins brought Corrix, his research product on how humans and AI actually collaborate, and the STOP framework for evaluating AI output before trusting it. The class opened by reviewing the previous night's impossible self-portrait homework, then closed with Tony onboarding everyone into Krea, the semester's primary image and video tool, in place of a co-guest from Krea who didn't end up presenting. After this page you can run STOP on any AI output before you read it, and choose approving versus supervising mode by what's actually at stake.
The night at a glance
The exercise, reviewed · 6:29 PM
Dom named the exercise's real point on the spot: it isn't to make a perfect picture, it's to teach you how to use your words, and to see how the machine reacts to them. Getting exactly right is impossible by design.
Homework from Class 2, reviewed live
One classmate explained a technique that Dom praised and named on the spot: starting a brand new chat for each attempt instead of refining a single thread, because a continuous chat kept reverting to earlier versions and cherry-picking old details until the result looked more like a collage than a photo. Dom called it branching, or forking, a cutting taken from the main plant, and flagged it as the exact skill Sydney would build on in a later class about art direction and consistency.
"Long layers" undersells what a model actually understands; "butterfly cut" is the tagged, zeitgeist term the training data was built on. The same swap turned "cool tone" into "soft winter." Vocabulary unlocks surfaced live in the room too: quiff for hair styled up in front, blaze for a dog's white face-stripe, epicanthic fold for East Asian eye shape, terms nobody had reached for until they were named out loud.
Describing yourself in words forces the model toward ethnicity defaults, and it performs best, by the room's own account, for a white man. One student's unprompted output defaulted to a cisgender white male description before any of that was specified. Nobody had a fix for it that night, only the observation, which is exactly the kind of blind spot Wouter Oomen's later class digs into.
The framework · 7:03 PM
The baseline question: working with the model, do you outperform what it would have produced on its own. Across roughly 1,000 people assessed, 82% did.
Are you adding real context, pushing back, challenging the output, or just accepting the first draft. This axis divided people more than any tool choice did.
Do you still understand what the model is doing, or are you losing the underlying skill by delegating it away. Even early adopters with three-plus years of daily use showed their scores quietly regress.

The craft · 7:11 PM
Which archetype are you actually working as: an Explorer, a Partner, a Passenger, or a Briefer?
Corrix's four collaboration archetypes each carry a distinct development path . The data behind them cuts against a few assumptions in the room: people in their fifties scored best, people in their twenties scored worst, the most likely group to passively accept an output without pushing back. Self-assessment had zero correlation with real skill, and the gap ran by gender too: men overconfident, women underconfident, even though women outperformed men on nearly everything, driven mostly by challenging the AI more often.
The judgment · 7:16 PM
Stop. Think. Organize. Proceed. A pause, not a checklist for every output.
Geoff's STOP framework isn't meant to slow every interaction down, it's meant to catch the moments that deserve it: check your own biases before you read the answer, then check the source, who made it, is it fact or opinion . The single practical habit Geoff kept returning to: ask the AI to fact-check its own answer before you even read it. Models are far better at spotting their own errors than at avoiding them in the first place, and it works roughly nine times out of ten.
Pick the mode by the stakes, not by habit. Approving mode means you sign off on every output; supervising mode means you oversee a sample or the underlying framework instead of reviewing each one, the right call for something like a thousand personalized emails, where your judgment belongs on what matters most, not every line.
"It is way, way easier for AI to spot errors in its own work than to avoid making those errors in the first place."
Geoff, on the fact-check-itself habit
Methods and prompts
When a continuous refinement thread starts reverting to old versions or cherry-picking earlier details, stop refining and start a new chat. A wider starting point beats a narrower, more tangled one.
Working prompt
Start a completely new chat. Here is the same brief again: [paste]. Do not reference any previous attempt, generate from scratch. I want to compare this fresh version against my earlier thread to see which starting point actually works better.
You will know it worked whenthe new version reads as a genuinely separate attempt, with no leftover phrasing or ideas carried over from the earlier thread.
Trade literal description for the tagged, zeitgeist vocabulary the training data was actually built on. "Long layers" undersells; "butterfly cut" lands closer to what the model has actually seen.
Working prompt
I'm describing [a look, a style, a feature] and want the vocabulary that image models are actually trained on, not the literal description. Give me the tagged, search-style terms a stylist or the internet would use for this, then I'll pick the ones that match what I mean.
You will know it worked whenthe terms it returns read like search tags or industry shorthand, not literal restatements of the words you already used.
Geoff's single most practical habit. Ask for the self-critique before you look at the answer. It works on research, code, or a creative draft, and works roughly nine times out of ten.
Working prompt
Before I read your answer, fact-check it yourself. Flag anything you're uncertain about, anything that could be wrong, and any place your own bias or a leading assumption in my prompt might have shaped the result. Then show me the answer.
You will know it worked whenthe self-check names a specific uncertain point or a leading assumption in your own prompt before the answer itself appears.
Decide the mode before the task starts, by what's actually at stake, not by default habit. Reviewing every one of a thousand emails is not the same job as reviewing the framework that generated them.
Working prompt
Here's the task: [describe it] and the number of outputs involved: [how many]. I answer first, then you check me: my call on the mode is [approving, review everything, or supervising, review a sample or the framework], because [your reasoning]. Tell me what risk I'm accepting if I'm wrong.
You will know it worked whenit names a specific risk you're taking on if your approving-or-supervising call turns out to be the wrong one for this task.
Feed back your own edited drafts along with what you changed and why, so the model learns the gap between its instinct and what actually felt authentic to you, instead of just extracting one more draft.
Working prompt
Here's a draft you gave me: [paste original] and here's my edited version: [paste edit]. Analyze the gap: what did I change, and what pattern does that reveal about my actual voice versus your default? Apply that pattern to the next draft before I ask.
You will know it worked whenit names a specific pattern in what you changed, not just a list of edits, and the next draft actually reflects that pattern.
The close · 8:45 PM
Tony ran Krea onboarding solo after the planned co-guest from Krea didn't end up presenting. Krea sits as an aggregator: many frontier models and interfaces in one place, so the class isn't juggling a dozen logins, with the left menu holding mood boards, LoRA training, a node editor, and a shared asset library. The demo compared three models side by side on the same still life , and token costs came up as a new consideration: Google Imagen 4 runs about 15 tokens a generation, Nano Banana Pro about 100.
Dom's caution carried straight through from Class 2's sea of sameness: even training on your own brand guidelines only gets you dilution, not distinction, because everyone using the same base model drifts toward the same look. At Coca-Cola they deep-tune a foundation model instead of touching these shared tools, precisely to avoid getting the same images as every other brand on the platform. A LoRA trained on something private, like a classmate's dog, goes into that same shared soup and stays effectively unrecoverable unless it becomes as tagged and famous as a household name.
Try this prompt
Quiz me on the STOP framework and Corrix's three measurement buckets. Then give me a real AI output and make me run STOP on it before you weigh in.
You will know it worked whenit quizzes you on the STOP framework and the three measurement buckets first, then makes you run STOP on a real AI output before it weighs in.
The shelf
22 captures, in order. Click any one to see it full size.





















