A Lesson on Audience Adaptation
What I got wrong teaching a hard idea, and what three years of revision taught me
In the summer of 2022 I arrived in Taipei as the only policy debate coach in the country. Every previous coach had left, COVID had made the program's future uncertain, and the other coaches I was expecting were stuck waiting on visas. My predecessors were well loved by the team. I had two months alone to make a first impression on students and parents who had every reason to wonder whether the program would survive.
So I built a research team. I drafted a quota system, defined the contract and expectations, pitched it to my boss, ran the interviews, and communicated the whole thing to families. My own job on that team was to write the team affirmative — the case every student on the squad would read all semester. The topic that year was security cooperation with NATO, with an artificial-intelligence component. I wanted to write something they would remember.
The idea underneath the idea
On the surface, the case was a conventional national-security proposal: a Cognitive Warfare Early Alert System to detect and deter AI-powered information attacks. Underneath the topical requirements, it was my thesis on AI literacy. I wanted to teach two things:
AI only works if we trust each other.
AI, like humans, needs to
learn to love.
I titled the file 1AC — AI Love. Partly because I thought it would be fun and partly because I want it to be true. But wanting students to learn something is not the same as building something they can learn from, and the gap between those two is the whole subject of this page.
What went wrong
I picked the content myself. The students didn't want a military AI topic, and I hadn't asked. I had confused what I found meaningful with what they would find meaningful, which is the oldest mistake in teaching and I made it in my first semester there.
I miscalibrated difficulty. The file ran to almost 63,000 words and I gave the students almost no scaffolding to climb it. They needed more teaching and less content. I gave them the reverse.
I assumed the feedback loop worked the way it does in the United States — that we would have a squad meeting and address concerns directly. In Taipei, that assumption was wrong.
I had fixed ideas about how debate ought to be learned: independent research, highlight your own evidence, read widely on your own. Those are how I learned. I was teaching my own path rather than theirs.
The evidence, in the files themselves
Both artifacts below are founding curricular materials — the first scripted speech plus supporting scripts for the second. They are roughly the same length. What changed is everything about how they teach. These figures come from the documents themselves:
Highlighted words
493 → 3,182
2022 file → 2025 file
0.8% of the text → 4.8%
Underlined words
12,732 → 7,457
2022 file → 2025 file
20.3% of the text → 11.2%
Median card length
2,168 → 734
2022 file → 2025 file
words per card
Cards in the 1AC
19 → 30
2022 file → 2025 file
the speech as actually read
For scale, Brave New World is 63,766 words, and schools place it at grades 9 to 12. The 2022 file is 62,629.
Word count, ATOS level 7.5 and interest level UG 9–12 from Accelerated Reader BookFinder (Renaissance Learning), quiz no. 8653.
Highlighting shows the words a student actually says out loud; underlining shows the reasoning that supports the claim. In 2022 I underlined a fifth of the file and highlighted less than one percent of it. I had done the analysis thoroughly and the performance work almost not at all — the document was a record of my reading, not an instrument for their speaking. A student who opened it could not tell what to say.
The 2025 file inverts that: six times the highlighting, 40% less underlining, and 56 cards where the 2022 file had 20, each one a third the length. Plus a documentation layer that did not exist before.
The count that matters most is the speech itself. The 2022 first affirmative was 19 cards across two advantages; the 2025 one is 30 cards across two. More cards, each of them shorter, arguing the same amount. There is a third advantage in the 2025 file — 25 cards on NOAA, written and highlighted and ready — that is deliberately not in the script. Having material a student does not have to read is its own kind of calibration.
2022–23 — Cognitive Warfare
- Topic chosen by me
- 19 cards in the 1AC, 20 in the file
- Median card 2,168 words
- No documentation, no notes
- No timing guidance
- Students expected to cut and mark their own
- Abstract thesis: trust and love between humans and machines
2025–26 — Fisheries
- Topic chosen by students from a prototype
- 30 cards in the 1AC, 56 in the file
- Median card 734 words
- Documentation section answering every predictable question
- Every advantage pre-timed to the minute
- Scripts usable on day one
- Concrete thesis: the fish are moving north, so we should go catch them
What I did not change
I did not make the words easier. I made the unit smaller. A student meets CAOFA, cetacean, or ecosystem-based management in a card of about seven hundred words, after a fairy tale has already explained the treaty — rather than on the third page of a card of over two thousand, with nothing marked.
Keeping the vocabulary hard was deliberate, and not only for debate. I work in Taiwan with students learning English as a second language, and vocabulary is one of the highest-value things an ESL classroom can spend time on. Unusual words are the part of the file parents can see: a student who comes home using cetacean or bloc is visibly learning English, not just learning debate. Hard words also generate questions — a student who does not recognize a word asks about it, which is the behavior I most wanted and least reliably got in 2022. And they make good games. A word list is something you can quiz, race, draw, and argue over long before anyone is ready to argue about the Beaufort Sea.
That also solved a problem I have not mentioned yet: the same file had to work for two age groups at once. A squad is not one audience. The word games were for the younger students, who will happily spend twenty minutes racing each other to define cetacean. The older students needed something else — ideas large enough to be worth disagreeing about. Floating cities. Hybrid warfare. Science diplomacy. Those are the phrases that made a sixteen-year-old look up.
A single difficulty setting would have failed one group or the other. Hard words in small units let one document serve both: the younger students climbing it a word at a time, the older ones arguing about whether a floating city counts as a country. Neither group got a different file. They got different footholds on the same one.
What I changed, and why it worked
I prototyped before committing. I built a short arguments file for the Arctic topic and ran it past younger students, letting them pick what interested them. They chose fisheries. The prototype is deliberately plain — taglines like “Melting ice unleashes deadly Arctic diseases, but scientific monitoring can help us quarantine them” instead of anything about affordance or deterrence.
I used an AI agent to source material so I could spend my own hours on the part only I could do: calibrating difficulty and adapting to cultural norms. The research was the cheap part. The teaching was the expensive part.
I wrote the documentation first. The 2025 file opens with a section called Notes, and the first thing in it is a fairy tale. Two ice kingdoms share a frozen sea they cannot divide; the ice melts; an Angry King closes his borders; two fish wizards make a bargain. It is a complete allegory of the policy — the Beaufort Sea boundary dispute, the collapse of scientific cooperation, the role of institutions that answer to evidence rather than to kings. A student who reads the fairy tale understands the case before reading a single piece of evidence.
After the story, the Notes answer the questions students were actually going to ask: What does the plan do? What advantages do I read? Why is the plan text so vague? Is this topical? And then timing, in plain numbers: it takes me four minutes to read the Canada advantage — you're probably faster, with a ranked list of what to cut if you run long.
What I understand now about how people learn
A curriculum does not belong to me. It is not for me, and it is not for any one student. It has to be universal, simple, and consistent enough to pick up conveniently.
That is the biggest thing I took from the three years between these two files. Learning and communication come in many shapes and sizes, and there is plenty of empirical literature on that — but the practical version is architectural. A good curriculum is a foundation, not a finished building. Other teachers have to be able to build on it, students have to be able to add to it, it has to survive a change in topic or staff, and the same idea has to be deliverable through video, story, practice, and debate. The 2022 file was a finished building. Only I had the keys.
That is also why I wanted an empirical curriculum with broad goals rather than a narrow one built around my own thesis. Modeling and prompting are among the few teaching tools validated across nearly every audience — they show up in the evidence-based practice literature for learners with autism, for language learners, and for skill acquisition generally. If the practices that work for the widest range of students are modeling and prompting, then the file itself has to contain the models and the prompts. The topic had to be accessible enough that students would engage with it, and the file had to carry the demonstrations and cues they needed to reach the bigger concepts underneath. That is what the pre-highlighting, the scripts, the fairy tale, and the timing notes actually are: models to imitate and prompts to follow.
On modeling and prompting as evidence-based practices, see the site's EBPs for Academics and EBPs for Social Outcomes pages.
Everything else follows from that:
Motivation is a prerequisite, not a bonus. Explaining that fish are moving north is easier than explaining that we should merge with AI to reduce misinformation risk. Same depth of argument underneath; radically different cost of entry.
Scaffolding is not condescension. I thought pre-highlighting robbed students of the work. It doesn't — it changes which work they do. Freed from deciding what to read, they argued about whether it was true.
Documentation is a teaching act. Every question I answered in the Notes is a question a student would otherwise have had to ask in front of their peers, in a second language. Writing it down removed the social cost of not knowing. It also made the file legible to the next coach, which is the part I had not thought about at all in 2022.
Redundancy across modes is a feature. The fairy tale, the Q&A, the highlighting, and the timing notes all teach the same case. A student who does not get it from one gets it from another, and none of them cost anything to include.
The hard idea can still be in there. The fisheries case is, underneath, an argument about democratic engagement, the authority of scientific institutions, the limits of philosophical skepticism, the proper use of irony, and the place of spirituality in debate. The Great Wizard Council in the fairy tale refuses a king and says its ally is truth and reality itself. I hid nothing; I just built a staircase to it.
I was less confident students would absorb everything underneath. I was much more confident they would stand up and give the speech. That trade is one I would make again.
A note on tools, and two thank-yous
One obstacle I have not mentioned was purely mechanical. Debate software is built for Windows, and my squad was not. Some students worked on a Mac, some on Windows, some on a Chromebook, some on an iPad, and some only on paper. Every one of those setups had slightly different requirements to open the same file, which meant a share of every class went to troubleshooting rather than arguing.
CardMirror fixed all of it, so huge thanks to the team behind it. My classes will include a transition phase for it going forward. Thanks also to the team behind debate-flow, a browser-based flowing tool that solved the same portability problem for note-taking.
The artifacts
Each file contains a sample of the full product.
The document is also only part of the teaching. Each was written to carry a full year, and around it I gave lectures, built presentations, ran private lessons, and judged practice debates across 16 classes a year — eight per semester, spanning three grade levels. A file that has to survive that much repetition, in front of that many different students, cannot be built around what one teacher finds interesting — which is the same lesson from a different direction.
2022–23 — the first draft
The Cognitive Warfare affirmative. A critical eye will notice immediately that it is far too long, lacks documentation, and is largely unhighlighted. I include it because the failure is legible in the file itself.
Team Affirmative 01 · .docx, 1.7 MB Team Affirmative 01 · .cmir, 1.7 MB2025 — the revision
The fisheries affirmative. Documentation, full highlighting, pre-timed advantages, scripts ready to use immediately, and a real path for students to engage with a working government agency in NOAA.
Team Affirmative 02 · .docx, 219 KB Team Affirmative 02 · .cmir, 212 KB Arctic Prototype · .docx, 39 KBIf you have not read a debate file before, How to Read Debate Files explains the anatomy — what the highlighting and underlining mean, and why they are the whole point of this page.
How I calculated the numbers
I used Claude Cowork, pointed it at my docs in the CardMirror and .docx formats, and asked it to count the cards, word counts, underlining, highlighting, and tags. I manually counted the cards for each and compared with Claude's results.
I asked Claude to count using both .docx and .cmir to see if there were any differences. Claude describes them in this sentence:
In Word a card tag, an analytic, and a heading in the notes are all just bold text in the same style, so telling them apart means guessing, and my first pass guessed 63 cards where Gabe counted 56. In .cmir a card is marked as a card and a citation as a citation, so the same count needs no guessing and matches the hand count exactly. Word counts needed a second fix in both formats: each one breaks text wherever the formatting changes, so highlighting half a word stores it as two pieces, and counting those pieces separately inflated every total by about 2%. Reassembling the text before counting gives 62,629 and 66,331 — which is what the .docx and the .cmir now both produce.
I then ran the calculation with GLM 5 Turbo and ChatGPT (Free). Claude was running as Opus 5 on high reasoning throughout. Here is what each attempt produced for the fisheries file, against the count I did by hand:
| Who counted | Cards | Words | H'lighted | U'lined |
|---|---|---|---|---|
| Me, by handreading the document | 56 | — | — | — |
| Claude, first attemptsbefore both corrections | 63–93 | 67,558 | 3,206 | 7,523 |
| Claude, correctedcards with a citation; words reassembled first | 56 | 66,331 | 3,182 | 7,457 |
| ChatGPT, first scriptcounted every card node | 93 | 67,068 | 3,202 | 7,605 |
| ChatGPT, rewrittencards correct; words not reassembled | 56 | 67,558* | 3,206* | 7,523* |
| GLM, first scriptcounted every card node | 93 | 67,558 | 3,206 | 7,523 |
| GLM, rewrittenboth corrections applied | 56 | 66,331 | 3,182 | 7,457 |
| Gap between the two rewritten scripts | 0 | +1.85% | +0.75% | +0.89% |
I asked ChatGPT and GLM to verify Claude's script, and both said the script would yield the correct count. Both then wrote their own scripts from scratch. Both got the card count right, and both reproduced Claude's totals — though ChatGPT reproduced the figures from before the word-splitting fix rather than after it.
GLM then wrote a separate script with improvements to the engineering, and arrived at the same count. It is eight times faster and holds almost nothing in memory, and every figure it reports is identical to the ones published here. It is published below alongside Claude's, unmodified.
Claude published a Python script to analyze the files. My Python is rusty, so if anybody wants to take a look and correct my findings please feel free. It would be great if I could say someone audited the data.
analyze-debate-files.py · Python, no dependencies the same script, commented line by line · if you don't write Python analyze-cmir-glm.py · GLM's version, same numbersI wanted to publish ChatGPT's script here too, for the same reason, but I ran out of tokens before I could get it out. If I get back to it I will add it.
Still to do: a comparison of the vocabulary in each file against English word-frequency bands is written but unpublished, pending a defensible mapping from word frequency to reading difficulty.