← Teaching

Experiments

Playable versions of canonical experiments — designed for solo walkthrough or as the basis for in-class replication.

Zollman (2010)

The Epistemic Benefit of Transient Diversity
Open the tutorial →

An interactive build-up of Zollman's bandit-network model of scientific communities — why more communication can mean less truth — one assumption at a time. Then two dashboards: the original model, and an extension in which each scientist consults a sycophantic or a devil's-advocate LLM, showing that whether an AI helps a community reach the truth turns on the shape of its network. Runs entirely in the browser.

Reuter & Brun (2022)

Empirical Studies on Truth and the Project of Re-engineering Truth
Start experiment →

Is 'true' ambiguous? A three-part replication: read a story where a speaker's answer fits everything they believe but not the facts, then one where it fits the facts but not their beliefs, and say whether each answer was true. A third part runs the paper's checks on whether 'true' really just meant truthful. Reuter & Brun found responses split close to 50/50 — evidence that everyday 'true' has both a correspondence and a coherence sense.

Instructor results dashboard →

Knobe (2003)

Intentional Action and Side Effects in Ordinary Language
Start experiment →

The side-effect effect — the most-cited result in experimental philosophy. Read one vignette, assigned at random, and say whether a chairman who is indifferent to the environment harmed (or helped) it intentionally. The two versions are matched on foresight, indifference and causal structure, and differ only in whether the side effect is bad or good; Knobe found 82% vs 23%. Both of his studies are included, and class results are compiled live against the published figures.

Instructor results dashboard →

Machery (2008)

The Folk Concept of Intentional Action
Start experiment →

A non-moral test of the side-effect effect. Read one randomly assigned smoothie-shop vignette and say whether Joe intentionally paid a dollar more (a cost) or intentionally got a free commemorative cup (a bonus) — then rate whether the outcome was blameworthy, praiseworthy, or neutral. Machery found 95% vs 45% 'intentional' with both outcomes judged morally neutral, evidence that cost-benefit reasoning, not morality, drives the asymmetry. Class results compiled live against the published figures.

Pettit & Knobe (2009)

The Pervasive Impact of Moral Judgment
Start experiment →

Does the moral asymmetry reach beyond 'intentionally'? Read one randomly assigned version of the chairman case and rate, 1–7, whether he 'decided' to help or harm the environment. Pettit & Knobe found stronger agreement that he 'decided' to do it in the harm version (4.6 vs 2.7) — the same asymmetry for an ordinary mental-state word, evidence that moral judgment is a pervasive input to folk psychology. Class results compiled live against the published figures.

Sripada (2010)

The Deep Self and Asymmetries in Intentional Action
Start experiment →

The badness of an outcome, or its fit with who the agent is? Read one randomly assigned, morally neutral rifle-contest vignette — one where the winner has no real stake, one where winning fulfils a lifelong dream — and rate whether the hit was intentional. Sripada found people call the outcome intentional far more when it concords with the agent's settled values, independent of moral valence. Class results compiled live against the published figures.

Uttich & Lombrozo (2010)

Norms Inform Mental State Ascriptions
Start experiment →

Moral badness, or norm-breaking as such? Read one randomly assigned Gizmo-company vignette where a foreseen side effect either conforms to or violates a norm — sometimes a moral norm, sometimes a mere colour convention — and rate, 1–7, how apt it is to call it intentional. Uttich & Lombrozo found norm-violating outcomes rated more intentional even for conventional norms, evidence that norm violation, not moral valence, carries the effect. Class results compiled live against the published figures.

Nadelhoffer (2006)

Bad Acts, Blameworthy Agents, and Intentional Actions
Start experiment →

Does blame bias our judgments of what a person did on purpose? Read one randomly assigned version of a fatal car-swerve case — a fleeing thief killing a pursuing police officer, or an innocent driver killing an armed carjacker — and say whether the death was brought about knowingly, intentionally, and how much blame it deserves. A built-in lesson in experimental design: the two versions differ in more than one way at once. Class results compiled live against the published figures.

Phillips, Luguri & Knobe (2015)

Unifying Morality's Influence: The Relevance of Alternative Possibilities
Start experiment →

Why does morality change judgments that aren't about morality? Read one randomly assigned version of the chairman case, then rate both whether he acted intentionally and whether a bystander's alternative was relevant. The claim: morality shapes which alternative possibilities strike us as relevant, and that is what moves the intentionality verdict — unifying the side-effect effect with parallel effects on causation and freedom. Your class's two measures are compiled live against the published figures.

Lindauer & Southwood (2021)

How to Cancel the Knobe Effect
Start experiment →

Can the side-effect effect be switched off? Read one randomly assigned version of the chairman case and rate your agreement that he did NOT do it intentionally — except one group can also register strong condemnation in the same breath ("…but he knowingly harmed it and should be blamed"). When that option is present, the asymmetry collapses: the survey evidence its authors read as support for the pragmatic account. Compiled live against the published figures.

Frohlich, Oppenheimer & Eavey (1987)

Choices of Principles of Distributive Justice
Start experiment →

A solo playtest of the classic veil-of-ignorance experiment. Pick a principle, see what you'd earn, then deliberate with four simulated co-participants and try to reach unanimous agreement. The original study found 35 of 44 groups chose a principle Rawls explicitly rejected.

Esper (1966)

Social Transmission of an Artificial Language
Join a room →

A live, in-class transmission chain. Each student learns names for eight shape-colour objects, then reproduces them from memory — and their version becomes the language taught to the next student. Across generations a 'totally suppletive' vocabulary drifts toward morphological categories. Run several parallel chains and compare how they diverge. Instructors: open the room from the host console.

Instructor host console →

Inspired by Allen, Quinlan, Andow & Fischer (2021)

What Is It Like to Be Colour-Blind?
Start experiment →

What do you think a red/green colour-blind person sees when they look at something red? Choose between four rival philosophical accounts, name six patches run through a standard simulation of colour blindness, and predict which colours would look new through 'colour-correcting' glasses. Allen et al. interviewed 17 colour-blind participants: all of them named the red patch red, 12 of 17 saw no new colours through the glasses, and nobody saw red or green for the first time — every candidate new colour was a pink or a purple. An original classroom design built on the paper's question, not a replication of its interview method. Class results compiled live against the published findings.

Instructor results dashboard →

Brown & Lenneberg (1954)

Codability and Colour Memory
Start experiment →

A classroom replication of the codability-and-memory study: name 12 colours, then try to recognise a subset after a 30-second arithmetic-filled delay. Anonymous results are compiled live for discussion.

Heider — Focal Colours (1972)

Universals in Colour Naming and Memory
Start experiment →

Rosch Heider's challenge to Brown & Lenneberg. Pick best examples of basic colour names, name 12 chips (focal, internominal, boundary), then recognise a subset from an 80-chip array — testing whether codability drives memory or some colours are simply more distinctive.

Winawer et al. (2007)

Russian Blues and Colour Discrimination
Start experiment →

Does the vocabulary you have change how fast you can tell two colours apart? A speeded matching task on twenty blues spanning the Russian siniy/goluboy border, which English does not mark: pick which of two squares matches the one above, sometimes while rehearsing an eight-digit number, sometimes while holding a grid pattern in mind. Russian speakers were 124 ms faster across the border than within it — and verbal, but not spatial, interference wiped that out. The class runs the paper's English-speaker arm, where the predicted result is a flat line.

Instructor results dashboard →

Roberson, Davies & Davidoff (2000)

Color Categories Are Not Universal: The Triad Task
Start experiment →

The double-dissociation test. See three color chips at a time and click the two that look most alike — one set spans the English green–blue boundary, the other spans the boundary Berinmo (five basic color terms, Papua New Guinea) draws between nol and wor, straight through English green. Roberson et al. found each group's similarity judgments snapped to its own language's boundary and sat at chance on the other's: English speakers 23.0 vs 14.38 of 32, Berinmo 20.75 vs 25.38. The class runs the English arm against the published Berinmo numbers.

Instructor results dashboard →

Berlin & Kay (1969)

Basic Color Terms: Mapping the Munsell Array
Start experiment →

The procedure that launched the universals debate, on the 330-chip World Color Survey Munsell array. List the basic color words of a language you speak, mark every chip each word covers, and pick its single best example. Berlin & Kay found that boundaries wander but foci cluster in the same regions across twenty languages; Roberson, Davies & Davidoff's Berinmo work pushed back. The instructor dashboard lays the class's charts on top of each other, focal points and term regions compared across whatever languages are in the room.

Instructor results dashboard →

Gilbert, Krull & Malone (1990)

Unbelieving the Unbelievable
Start experiment →

A replication of Study 1: learn an invented Hopi vocabulary, get interrupted by an occasional tone, then take the identification test. The diagnostic asymmetry — false propositions misidentified as true under interruption — is what Spinoza predicts and Descartes does not. (From the Speech Attacks readings.)

Hansen & Liao (2026)

Measuring Conceptual Inflation
Start experiment →

A classroom replication of Study 1 on the meaning of 'racist'. Rate the extension and intensity of 'racist', its degree-modified forms, related vocabulary, and a set of thin moral terms — then compare your live audience's pattern with the published representative-sample results.