Part of the Good Data vs Bad Data lesson guide. Teaching a different grade? 🔧 Builder (8–10) · 💻 Hacker (11–14) · ⚡ Architect (15–18)
Gather students close — on the rug, or with devices face-down for a moment so all eyes are on you. This lesson's mascot is Pixel, so read the opening line with warmth and a little wonder, as if Pixel is telling the class a secret:
"What if you only showed A.I. pictures of BIG dogs? 🐕 It would think ALL dogs are big! And it would be WRONG about small dogs! That's why GOOD data matters!"
Pause after "wrong about small dogs" and ask, with genuine curiosity in your voice: "Has anyone here ever seen a REALLY tiny dog? Like, tiny enough to fit in a teacup?" Let a few students describe small dogs they know (chihuahuas, teacup poodles, a neighbor's puppy). Then say: "So if a smart helper only ever saw pictures of big dogs like Great Danes, and never saw a tiny dog like that... what do you think it would say if it saw one?" Take guesses — most students will land on "it wouldn't know it's a dog!" without much prompting. That's the whole lesson in one sentence: a smart helper only knows what it's been shown.
Tell the class: "Today we're going to be data detectives. We're going to figure out what makes examples GOOD for teaching a smart helper, and what makes examples BAD — and why it matters so, SO much."
Before opening the app, do a thirty-second warm-up with an object already in the room. Hold up a shoe and say: "Imagine I only ever taught you the word 'shoe' by showing you MY shoe, this exact one, a hundred times. Then I showed you your friend's sneaker. Would you know that's a shoe too?" Most students will say yes, because people generalize easily — but push a little further: "What if I only showed you high heels, never sneakers or sandals or boots — would 'shoe' in your head only mean high heels?" This gets at the same idea as the big-dog hook from a second angle before the vocabulary gets any harder, so slower processors in the room have two chances to grab the idea rather than one.
Open the lesson together. It has three parts — walk through each one like a game, always asking the class to guess before you reveal the answer.
Part 1 — Data Quality ✨. Show or describe each idea, one at a time, and connect it back to the big-dog example from the hook:
Part 2 — Bias in Data ⚖️. Introduce the word "bias" gently: "Bias is a big word for a simple idea: it means the examples weren't fair to everyone." Walk through the four items, translating each into something concrete:
After these three examples, pause and say plainly: "Notice something — none of these smart helpers were being MEAN on purpose. They just repeated whatever was in the examples they were shown. That's why picking GOOD, FAIR examples matters so much." If a student asks "so is it the smart helper's fault?" — this is a good moment to say clearly: "No. It's not the smart helper's fault, and it's usually not any one person's fault either. It's about making sure the examples include EVERYONE before you ever start teaching it."
Part 3 — Fixing Data Problems 🔧. End on hope, not just problems — this part is about how people fix it:
Give one more concrete example to cement Part 3 before moving on: "Imagine a school is building a smart helper that recommends lunch menus. If the team building it only ever ate rice and adobo growing up, they might forget to include vegetarian options, or halal options, or someone with a peanut allergy. It's not that they don't care — it's that it didn't occur to them, because it wasn't part of THEIR examples. That's exactly why 'Diverse Teams' matters: more kinds of people notice more kinds of missing pieces."
Close the activity by reading the lesson's own summary line together: "A.I. needs fair and complete examples to learn right! If it only sees big dogs, it won't know small dogs. If it only sees boys, it might not work for girls. Good data = good A.I.!" Then let students do the in-app practice round, sorting more examples into fair vs. biased data.
Close with: "You're data detectives now! Every time you see a smart helper — a phone, a game, a video app — you can ask yourself: what did it learn from? Did it get to see EVERYONE and EVERYTHING fairly? That question makes you someone who thinks carefully about smart helpers, not just someone who uses them." Give the class a round of applause for finishing the lesson, and remind them that being a "data detective" is a habit they get to keep using forever, not just something for today's lesson.
Extension activity — "Fair Picture Book": Give small groups a stack of magazine cutouts, printed photos, or drawing paper and ask them to build a tiny "picture book" that teaches a pretend smart helper what a "good pet" looks like — but it has to be FAIR, meaning it needs different kinds of pets (big, small, fluffy, short-haired, common, unusual). Groups then swap books and try to spot anything "biased" or missing in another group's book (a book with only dogs, or only fluffy animals). This turns the abstract idea of representative data into something students physically build and can see is missing pieces — a satisfying way to stretch the 5–15 minute core lesson into a full 30-minute activity.
If time is short, a simpler five-minute version works too: hold up two hand-drawn "example books" you prepared ahead of time — one clearly unfair (ten pictures, all the same kind of animal) and one clearly fair (ten pictures, many different kinds) — and have the whole class vote, thumbs up or thumbs down, on which one would teach a smart helper better. Ask a volunteer to explain their vote out loud before revealing which one you intended as the "trick" example.