Instructor Notes
Teaching Philosophy: The Researcher Stays the Active Reviewer
This lesson helps researchers work with AI coding agents without handing off the thinking. The shift is from writing every line of syntax towards reading, questioning, and validating code the agent produced. Resist framing this as “the AI does the work and you orchestrate”, the episodes deliberately push back on that. The core challenge for learners isn’t syntax; it’s managing cognitive load and staying the active reviewer who can explain and judge the result.
Key Concepts to Emphasize:
- Verification Load: It is often harder to verify code you didn’t write than to write it yourself. Normalise this “friction” as a sign of high-quality research.
- Evidence Mantra: “I do not approve changes; I approve evidence.” This should be the recurring theme of the workshop.
- Sandboxing: Always emphasize the security implications of giving an AI direct access to the filesystem.
The worked example (one project all the way through)
The whole lesson uses one fixture,
learners/files/coastal-water-quality/ (three messy site
CSVs). Learners carry it from spec to clean merge to validation to a
trend plot in the capstone. Two anchors to keep in mind: -
Expected result: a merged
data/master_dataset.csv with 60 rows (3
sites x 20 weekly samples), all dates in 2023. - The planted
trap (the lesson’s climax): site C dates are day-month-year
(05-01-2023 is 5 January). A naive
pd.to_datetime() call with no explicit format actually
raises an error on this file (day values like 19 and 26 can’t
be a month) — that’s a loud, easy failure, and a fine outcome if it
happens. The dangerous case is format="mixed": it silently
swaps day and month for some rows (05-01-2023 becomes May
1) while accidentally parsing others correctly (19-01-2023
still lands in January), so a handful of wrong dates hide among
mostly-right ones. Either way the row count stays 60 and the trend plot
is still wrong. This is the concrete “working code is not trustworthy
code” moment; let it happen and use it.
Episode: Before We Use AI (opening)
- Purpose: Reset expectations before anyone opens a terminal. Some learners arrive expecting a speed demo; this episode says plainly that the workshop measures whether they can explain and validate AI output.
- Set the norm: Make it safe to say “I don’t understand this line.” Treat confusing AI output as a shared teaching artefact, not a personal failure.
- The run / revise / reject checkpoint introduced here is reused in every later episode. Refer back to it by name.
- Sticky-note opener: Collect what learners want from AI and what they fear it will get wrong. Revisit at the end.
Episode 1: Understanding CLI-Based AI
-
Auth Check: Ask learners to run
claude --version. If it returns a version number they are ready. If not, have them launchclaudeonce and complete sign-in before continuing. -
Model Check: Have all learners set the same model
with
/modelbefore starting (pick the current recommended model close to the session date), so outputs are comparable and provenance records are meaningful. -
Starter folder: Everyone works in the provided
coastal-water-qualityfolder (ships atlearners/files/coastal-water-quality/). Confirm learners have it and canls data/to see the three site CSVs before starting. Episode 1 has them inspect a real file and compare the agent’s description to it, the first hands-on win. - The Browser vs. CLI distinction: Use the analogy of a “consultant” (Browser) vs. a “research assistant with keys to the lab” (CLI).
- Discussion: The prompt about “ChatGPT writing code that looks correct but fails” is a great way to bond over shared frustration and set the stage for why we need the CLI (to run and test immediately).
Episode 2: Best Practices for Prompting
- CO-STAR vs. CLEAR: Don’t get bogged down in the acronyms. The goal is intentionality.
- Live Demo Tip: Show a “Bad Prompt” vs. a “Good Prompt” live. Purposely run a vague command and show how it fails or produces messy output before using the refined version.
- Self-Correction: This is the “lightbulb” moment. Demonstrate asking the AI, “Are you sure? Review your code for edge cases.”
Episode 3: Data Cleaning (Live Demo)
- High Intensity: This is the most technically demanding episode.
-
Expected output:
data/master_dataset.csvwith 60 rows, produced byclean_and_merge.py, which learners should not modify. The “Update the script” challenge writes a separateclean_feb_onward.py(12 rows removed, 48 remaining) instead of editing the canonical script in place, so Episodes 4 and 6 can still rely on the full 60-row file. A wrong row count on the filter usually means some site C dates were silently misparsed. -
The “Safety Net”: If a learner’s AI fails to
produce working code after two attempts, have them copy
instructors/files/backup_clean_and_merge.py(it is written for this fixture and parses site C dates correctly). This keeps them on pace for the validate-and-judge steps, which are the point. - Stop and Read: Literally tell the class to “hands off keyboards” for two minutes to read the generated script before they run it.
Episode 4: Validation Best Practices
-
Four-Layer Validation Stack (match the episode
exactly):
- Requirement constraints / No-Go Zones (human-authored ground truth
in
CLAUDE.md) - Executable checks you can run (finish the shipped
validate_data.py) - Metamorphic and invariant checks (60-row invariant; mean unchanged when rows are shuffled)
- Domain plausibility (where the researcher’s expertise is irreplaceable; do not omit this layer)
- Requirement constraints / No-Go Zones (human-authored ground truth
in
-
Build the validator: the core activity. Learners
finish the three TODO checks in
validate_data.py, then deliberately misparse the site C dates and confirm the validator now fails. A validator that cannot fail on a known-bad input is not protecting them.
Episode 5: Limitations and Cautions
- Silent Semantic Drift: This is the most “dangerous” failure. The site C date trap from the worked example is the live version of this; refer back to it.
- Environmental Cost: This is often a new topic for researchers. It grounds the workflow in physical reality.
Episode 6: From AI Output to Research-Ready Code (capstone)
- Mostly doing: budget the time for work, not exposition. Learners assemble the full bundle (spec, plan, code, validator, plot, provenance, approval decision) on the coastal data.
- Expect the date trap to surface for several learners during the plot step; that is the highlight. Collect a few approval decisions and read them aloud, especially the honest “revise” answers.
-
Fallbacks:
backup_clean_and_merge.pyandbackup_plot_trend.pylet a stuck group still reach the validate-and-judge steps.
Episode 7: Resources and Next Steps
- The Toolscape: Acknowledge that tools (Aider, Cursor, Claude Code) change weekly. Focus on the principles (CLI, validation, provenance) rather than specific software.
- Monday workflow card: the takeaway exercise; push learners to name the one check that would catch a silent error in their own data.
- Attribution: Remind learners that while AI can’t be an author, transparency about its use is a core tenet of Open Science.
Suggested first-pilot path (compressed)
For a first pilot, do not teach every section evenly. Protect the practical spine and let the rest be optional. The pilot should answer one question: can learners use Claude Code to clean a small messy dataset, explain what changed, validate it with checks, and decide whether to approve the result?
Teach live: Before We Use AI (briefly); CLI setup,
the early data inspection, and /init; the Living Spec; Data
Cleaning with AI; the validate_data.py approval gate; the
capstone bundle.
Make optional / skim if short on time: CO-STAR and reasoning-model detail; the advanced tool landscape, MCP, and local-model material in Resources; multi-model verification.
If you run short, cut whole objectives (and their assessments), not bits from everywhere. The data-cleaning-to-validation-to-capstone arc is the part that must survive.
Troubleshooting & Common Issues
API Quotas & Limits
If learners hit “Resource Exhausted” errors, it’s likely they’ve exceeded their free tier quota or are prompting too rapidly. Suggest they wait 60 seconds or use a smaller “context” (don’t send every file in the folder).
Maintainer checklist
Run through this before teaching a pilot or merging substantial changes. It encodes the Carpentries guidance on teaching with generative AI.
Keeping the lesson current
The AI tool landscape moves faster than a Carpentries release cycle. The most volatile parts of this lesson are the tool landscape and reputable sources in Episode 7, the failure modes in Episode 5, and any named model or CLI command in Episodes 1-4.
Before each teaching, refresh these by running the
last30days skill (see
~/projects/last30days-skill):
BASH
python3 scripts/last30days.py "vibe coding" --include-web --emit=md --store --save-dir research
python3 scripts/last30days.py "AI coding agents" --include-web --emit=md --save-dir research
Use the output two ways: confirm the tool names and claims in Episodes 5 and 7 are still accurate, and pull one current hype example to debunk live in the Episode 7 “Spotting hype” section. A fresh example beats a canned one.
Tooling status (last checked 2026-06-19)
-
The lesson now uses Claude Code (Anthropic).
Learners run it on one of two backends: a personal Pro/Max plan or API
key for non-sensitive (P1-P3) work, or UCLA Amazon
Bedrock (Anthropic models) for sensitive (P3/P4) research data.
The same commands work on both; only the backend changes. Confirm
Bedrock access and data-tier approval with the learner’s unit before
using real sensitive data. See
learners/setup.md. -
Historical note (why we migrated): Google retired
the free/consumer Gemini CLI on June 18, 2026 (folded into the paid
Antigravity platform), which broke the lesson’s original tooling. Paid
Gemini Code Assist Standard/Enterprise and API-key access continued, but
the free
geminicommand did not. - As of June 2026, the most common research-capable CLI agents are
Claude Code (Anthropic), Codex CLI (OpenAI), Antigravity CLI (Google),
Cursor Agent, and open-source OpenCode/Aider. The lesson’s principles
(Living Spec, plan-first, validation stack, provenance) transfer across
all of them, see the
agentic-research/cli-agent-landscapewiki note.
Before We Use AI: What Are We Practising?
Instructor note: set the tone early
The goal of this episode is to reset expectations before learners open a terminal. Some will arrive expecting a productivity demo. Be explicit that the workshop measures whether they can explain and validate AI output, not how fast they can generate it.
Signs to watch for across the whole workshop:
- Learners who are impressed that the agent can read their files but cannot say what it changed.
- Learners who accept output they cannot explain because “it ran.”
- Learners who turn to the AI before turning to a helper, which hides where they are stuck.
Normalise bringing AI confusion back into the room. Treat confusing AI output as a shared teaching artefact, not a personal failure.
CLI-Based AI
Setup check
Ask all learners to run:
If it returns a version number, they are ready. If the command is not
found, they need to complete the install and sign in (launch
claude once and follow the prompts) before continuing.
Confirm everyone has selected the same model with
/model.
Instructor note: access is not understanding
Learners are often impressed that a CLI agent can read their files and run their code. Impressive access is not the same as a correct result. Watch for learners who can describe what the agent can do but not what it just did.
Before any learner approves a command that changes files, ask them to say out loud: what did the agent read, what is it about to change, and why. If they cannot answer, that is the moment to slow down, not speed up.
Discussion prompt
Ask learners: “Have you ever used ChatGPT to write code that looked correct but failed when you ran it?” This is a good time to introduce the concept of orchestration. The goal is not only to “fix” code, but to ensure the AI’s intent (the spec) is correct.
Best Practices for Prompting
Instructor note: watch for cognitive load
Generated prompts often pull in syntax, libraries, or abstractions the lesson has not introduced. Signs that AI output is adding extraneous load:
- The agent uses an advanced feature (comprehensions, classes, decorators) before it has been taught.
- It imports a library that is not installed locally.
- It writes several files when one short script would do.
- It buries the core logic under heavy comments.
- The answer is correct but the learner cannot explain it.
Interventions: ask the learner to request a simpler version, to remove one abstraction, to trace the code line by line, or to compare it with a minimal reference solution. Slowing down here is the lesson, not a detour.
Teaching tip: Visual aids
Write CLEAR vertically on a whiteboard. As you explain each letter, add the keyword (Concise, Logical, Explicit, Adaptive, Reflective). This helps students remember the framework.
The introspection concept
Emphasize this section. Most learners treat AI output as final. The idea that they can ask the AI to fix its own work is often a new concept. It is like asking a student, “Are you sure you checked your work?”, they often find their own mistakes when asked.
Data Cleaning with AI
Live coding
This episode uses live coding. Learners should follow along by running commands on their own machines.
Backup scripts
If a learner’s AI fails to generate working code on the coastal data,
provide the pre-written versions from instructors/files/: -
backup_inspect_data.py -
backup_clean_and_merge.py
Instructor note: cognitive load in this episode
This is the most technically demanding episode, and generated
cleaning code is where extraneous load shows up most. Watch for scripts
that reach for apply with lambdas, regex date parsing, or
multiple helper files when a short linear script would do. If a learner
cannot explain a block in step E, have them ask the agent for a simpler
version before continuing. The “hands off keyboard” read in step E is a
feedback checkpoint: it is where you find out who is lost.
Challenge tip
This challenge requires modifying existing code. If learners are
stuck, suggest they ask the AI to read clean_and_merge.py
before asking for modifications.
Validation Strategies: The Approval Gate
Teaching tip: approval fatigue
Warn learners about approval fatigue, the tendency to accept AI suggestions without reading them. The four-layer stack is designed to make the AI prove it is correct before you review the code.
Limitations and Cautions
Managing expectations
Current models tend to flag uncertainty more often than older ones, but they still hallucinate. Do not promise learners it won’t happen. * If it refuses: Acknowledge that the model correctly identified its own limitations. * Backup: Have a screenshot of a known hallucination ready to show if the AI performs perfectly during the session.
From AI Output to a Review-Ready Bundle
Instructor note
Budget most of the time for doing, not explaining. Expect the site C
date bug to surface for several learners; treat it as the highlight, not
a snag, it is the lesson’s whole thesis in one concrete failure. If a
group is far behind, have them use
instructors/files/backup_clean_and_merge.py so they still
reach the validate-and-judge steps, which are the point. Collect a few
approval decisions and read them aloud; the honest “revise” answers are
the best teaching.