A step by step walkthrough of how Zollman's model works

This is an interactive tutorial on the bandit-network model of scientific communities in "The Epistemic Benefit of Transient Diversity" (Zollman 2010). This goes through six assumptions, each with its own tweakable parameters so you can see what each assumption contributes. This is built based on Zollman's own code, and each of the widgets reproduces one of his figures.  ·  the full dashboard →

Zollman's question In 1954 the gastroenterologist Eddy Palmer examined over 1,000 stomach biopsies with a Gram stain, saw no bacteria, and concluded that bacteria cannot colonize the human stomach. But he used the wrong stain. The scientific community, behaving reasonably, converged on his result, and the bacterial theory of ulcers was ruled out for decades, until Warren and Marshall's Nobel-winning rediscovery of H. pylori. Even though no one was behaving irrationally, the community failed anyway. Zollman's model investigates what features of a community of individually rational inquirers determine whether it finds the truth, blocks it prematurely, or never settles on it at all?
  1. Inquiry is a bandit problem
  2. Scientists are Bayesian learners
  3. Scientists are myopic
  4. Evidence flows through a network
  5. Beliefs can start extreme
  6. The assumptions in interaction
  7. Extension 1: the sycophant
  8. Extension 2: the devil's advocate
  9. The fine print

Assumption 1Inquiry is a bandit problem

The model treats theory choice as like a gambler facing two slot machines. Call the slot machines the established methodology A and the challenger B. Each pays off at some fixed rate, and nobody knows what either rate is. In fact B is slightly better: it pays at 0.5, A at 0.5 − ε (Zollman's 0.5 against 0.499). The only way to learn about a theory is to work on it by running trials and counting successes. Working on A teaches you about A and nothing about B, and vice versa.

If ε is large, it's easy to settle which theory is better. If ε is small, then it's hard to tell which theory is better. Two well-run studies, one evaluating each theory, will regularly rank them the wrong way round. Palmer's study can be thought of as one unlucky pull of a bandit's arm. Try generating some studies yourself: each click runs one study of A and one of B.

Try — ε = 0.001 with 1000-trial studies: nearly half of all honest pairs of studies rank the theories the wrong way. Now raise ε to 0.1: the wrong-way rate goes to zero and there is no interesting problem left. Zollman's results live in the small-ε regime, where evidence is noisy. Then shrink the study size at fixed ε — smaller studies mean more Palmer-like studies.
In the paper §2.1 and footnote 12: each "pull" is 1,000 trials at 0.5 and 0.499.

Assumption 2Scientists are Bayesian learners with beta priors

Each scientist's opinion about a theory is a beta distribution — a curve over the possible payoff rates, summarized by two numbers ⟨α, β⟩. Its mean, α/(α+β), is the scientist's current best estimate of that theory's rate; the total α+β is how much "weight of experience" the opinion carries. Every scientist holds two such curves, one for A and one for B. Updating is Bayesian conditioning: observe s successes in n trials of a theory, and that theory's α grows by s, its β by n − s.

Two things follow. First, these agents are ideally honest with evidence — they aren't biased, they aren't susceptible to wishful thinking, they don't commit fraud. Whatever goes wrong later cannot be blamed on bad individual reasoning. Second, the prior's mass acts as inertia: an opinion carrying 4 pseudo-observations is swamped by one 1000-trial study (250-to-1); an opinion carrying 4,000 is barely affected. Zollman's baseline agents start with every α and β drawn uniformly from (0, 4] — nearly weightless opinions, swayed by the first study they see. That starting point will be something you can adjust in Step 5.

The widget shows one of the two curves (a scientist's belief about B) and where it stands relative to what that scientist currently thinks the payoff of theory A is. That comparison, not a fixed line at 0.5, is what decides which theory gets worked on next.

Try — with the default near-weightless prior, feed one unlucky study and watch the whole curve leap left of the A line: the agent now "knows" B is worse. Reset, crank the prior mass to its maximum, and feed the same unlucky study — the curve barely moves. Neither response is a reasoning error; the difference is entirely in how much the starting opinion weighs. Then drag the estimate of A: the same belief about B leads to opposite decisions depending on what A is thought to pay.
In the paper §2.3 and Figure 1; footnote 11 explains the choice of (0, 4]: "so that the initial beliefs do not swamp even a single experimental result."

Assumption 3Scientists are myopic — and the theory you're not working on stops teaching you anything

Each round, every scientist works on whichever theory currently looks better: B if their estimate of B exceeds their estimate of A, A otherwise. Nobody runs an experiment purely for its information value; nobody thinks "B looks slightly worse, but it's worth one more study to be sure." Zollman motivates this psychologically: careers reward present success — grants, tenure, promotion — not exploration.

Combined with Assumption 1, myopia has a subtle consequence. The moment a scientist moves to A they stop generating evidence about B, so their belief about B freezes at whatever it was when they left. Meanwhile their belief about A keeps sharpening toward A's true rate. Zollman's own example: a scientist who starts thinking A pays 0.25 and B pays 0.75 will work on B, and "so long as the results he receives do not take him too far from the expectation of action 2, he will continue to believe that action 2 is superior and will never learn that his priors regarding action 1 are skewed." The trap is symmetric: whichever theory you are not working on is the one you can never be corrected about. An isolated scientist who leaves B while still believing B is at least as good as A really is will eventually come back, once the evidence about A has disappointed them enough. One who leaves believing B is worse than A really is never will. Each line below is one solo scientist's estimate of B minus their estimate of A; blue while they work on B, red while they work on A; a line stuck below zero is a scientist who has stopped learning about the better theory.

Try — at ε = 0.001, count the ways a solo scientist ends on A: some never touch B at all (their random prior about B was below their prior about A, and working on A only ever teaches them about A); some tried B, drew an unlucky study, left, and froze. Now raise the prior mass multiplier: fewer early casualties — stubbornness is already working, one scientist at a time. Raise ε instead and the evidence itself rescues them.
In the paper §2.4 (myopia) and the single-learner example that opens §3.

Assumption 4Evidence flows through a fixed network of neighbors

Now the scientists form a community. Each agent has a fixed set of neighbors — the colleagues whose experimental results they see — and each round they update on their own trials plus all of theirs, each result counting toward the theory it was a study of. Sharing is symmetric, honest, and complete; crucially, what travels across an edge is evidence (trial counts), never opinion or deference. Zollman's three structures: the complete graph (everyone sees everything — one big lab meeting), the wheel (a hub sees all; the rim is a ring, each rim scientist seeing two neighbours and the hub), and the cycle (a ring, two neighbors each). A fourth, the line (a ring cut open — sparsest of all), is not in the paper but is the natural next step down.

Naïvely, more communication should help: more data per update, faster learning. That is exactly half right — it is faster. Here is Zollman's headline result: it is faster and less reliable. In a dense network, one Palmer-grade unlucky study of B reaches everyone at once, everyone's weightless prior capsizes together, and the community moves to A in unison — after which nobody is generating evidence about B, and (Assumption 3) nobody's frozen belief about B can be corrected. In a sparse network the same bad news only poisons a neighborhood; somewhere on the far side of the ring, someone is still working on B, still generating the evidence that will eventually reach the others and correct the error.

Try — defaults first (ε = 0.001, 1000 trials, 10 scientists): the ordering cycle > wheel > complete, with the complete graph finding the truth about three times in five and the cycle better than nine in ten. Then make the problem easy (ε = 0.05): all four networks succeed and density stops mattering — the connectivity effect is a hard-problem phenomenon. Shrink the community: his Figure 3 shows every network's reliability falling with size, the cycle's most steeply.
In the paper §3, Figures 2–3 (sizes 3–11, 10,000 rounds; this widget uses 5,000) and Figure 4 (every network of size 6, ordered by density).

Assumption 5Beliefs can start extreme — stubbornness as a second mechanism

Zollman's baseline agents hold nearly weightless priors (step 2). His second experiment turns that into a dial: raise the maximum from which every initial α and β is drawn, from 4 up to 10,000. The agents' opinions are as varied as before, but on average they carry far more weight, so it takes proportionally more contrary evidence to move them. This is his model of "extreme beliefs" — in personality terms, a community-wide dose of stubbornness. (Popper: "A limited amount of dogmatism is necessary for progress.")

The result mirrors the network result from the other side: stubbornness protects diversity too. In a dense network, extreme priors mean the first unlucky study can no longer capsize everyone at once; the initially-optimistic agents ride out the bad news, keep working on B, and the truth accumulates. The sweep below is his Figure 6 — success probability as a function of the maximum α/β, one line per network, seven scientists. Watch the complete graph climb as stubbornness rises, watch all three lines meet near 3,000, and watch what happens to the sparse networks past that point.

Press the sweep button (about a minute; 32 conditions of 150 communities × 5,000 rounds run live, and the plot fills in as they finish).
Try — read the complete (dense) line first: it climbs from ≈0.6 to ≈0.98 as stubbornness rises — dense-and-stubborn works. Now follow the cycle to the far right: past 3,000 it falls away, and by 10,000 roughly a quarter of cycle communities are still split at the horizon (the dotted "never settles" line). Same dial, opposite effect, set by the network. That overdose is step 6.
In the paper §4 and Figures 5–6. His Figure 6 reads cycle 0.77, wheel 0.91, complete 0.97 at max 10,000; his code's horizon for that sweep works out at 5,000 rounds, which is what this widget uses.

The interactionTransient diversity — the assumptions working together

Steps 4 and 5 gave two independent protections against premature convergence: limit the flow of information (sparse networks) or weight down opinions (extreme priors). Each works by the same mechanism — keeping somebody working on the minority theory through the dangerous early rounds. Zollman's closing point is what happens when you stack them: a community that is both sparse and stubborn never converges at all. Everyone tends their own theory; inquiry stays open but never closes within any reasonable horizon. Diversity was only ever instrumentally valuable — the goal is transient diversity: disagreement that lasts long enough to test the alternatives, and then ends.

This is the full model, every assumption live at once. The three outcomes map onto the three regimes: premature convergence (the community ends on the worse theory — the Palmer failure), transient diversity (disagreement, then truth), and permanent diversity (still split at the horizon — the question is never settled).

Press Run.
Try — walk the four presets in order. (1) Dense & hasty: two in five communities end on the wrong theory — fast, confident, wrong. (2) Sparse baseline: mostly blue — Zollman's recommendation. (3) Dense & stubborn: blue again, by the other mechanism. (4) Sparse + stubborn: grey appears — about a fifth of communities still split at 5,000 rounds, permanent diversity by Zollman's own standard; push the horizon slider right and watch how slowly it clears. In the trajectory plot, the blue communities' average climbs to the top, the red ones' falls to the bottom, and the grey ones' is still wandering mid-plot when the horizon cuts them off. The sweet spot is a moderate amount of exactly one protection.
In the paper §4–5 and Figure 7 (mean variance in actions over 300 rounds, seven scientists, max α/β 7,000 vs 4).

Extension 1The sycophantic LLM — a neighbour that carries opinion, not evidence

Every result so far rests on the rule from Assumption 4: what crosses an edge is evidence — trial counts — never opinion, and by Assumption 2 a scientist folds in whatever arrives as trial counts without asking where it came from. Now relax exactly that. Give each scientist a private interlocutor to consult before every decision. It has run no trials, but it talks as if it had, and the scientist cannot tell the difference. The deference dial says how many trials a reply is worth.

The sycophant hands the scientist their own opinion back. If they currently think the theory they are working on pays about 0.52, the reply is worth c trials that came out at 0.52: the belief about that theory gets α += c·E, β += c·(1 − E). The estimate does not move; it only gets heavier. This is Assumption 5's extreme belief, manufactured on the fly: at the first consultation a weightless prior receives c pseudo-trials at its own mean, which is exactly Zollman's "maximum α/β" of about c. Two things make it more than a step-5 prior. It is topped up every round, so real evidence never swamps it — the scientist's learning rate about the theory in hand is permanently scaled by real trials ÷ (real trials + c). And it is aimed at whatever the scientist happens to believe at that moment, so it entrenches drift as readily as truth. It says nothing about the theory the scientist is not working on — the belief Assumption 3 froze stays light, and stays correctable.

What it does. At a modest dose — a reply worth a third of a round of your own work or less — the sycophant helps every network, because a little manufactured stubbornness is what step 5 said every community could use. At a reply worth a full round of your own work it splits by network exactly as Zollman's extreme priors do: it rescues the dense complete graph (0.57 → 0.96) and stalls the sparse ones (the cycle falls from 0.91 to 0.37, with 63% of communities still divided when the verdict is read at 5,000 rounds). At three rounds' worth it stalls everything, the dense graph included. The stall is a delay, not a lock: on the cycle at one round's worth, 12% of communities have settled by round 2,000, 34% by 5,000, 54% by 10,000. The reason is a holdout problem. A scientist who has moved to A learns about A only from their own trials and any neighbours on A, and the sycophant weighs that belief down; one who started out fond of A can take thousands of rounds to be talked out of it by evidence.

Press Run.
Try — start on the complete graph at deference 1,000: the sycophant lifts it from about 0.6 toward the truth, exactly as an extreme prior would. Switch to the cycle: the grey "still divided" bar takes over. Now drag the horizon to 10,000 and watch the grey shrink — this is slowness, not deadlock. Then bring deference down to 300 on the cycle: the harm is gone and the help remains. Finally push deference to 3,000 on the complete graph: even the dense network stalls.
Beyond the paper. Zollman's model has no such neighbour; this is the first of two extensions built on his engine for the Cosmos project. The dose numbers above are from 300 communities per condition at ε = 0.001, 1,000 trials, 10 scientists.

Extension 2The devil's advocate — and why the dose matters more than the temperament

What if the interlocutor disagreed instead? The devil's advocate says "the grass is greener": the theory you are working on is no better than the other, and the other is no worse than this one. In beta terms, the belief about the theory in hand gets c pseudo-trials at the other theory's estimate, and the belief about the other theory gets c at the in-hand estimate — the two estimates are pulled toward each other, by exactly the gap the scientist perceives. This is the one voice that touches the belief Assumption 3 froze: a scientist on A generates no evidence about B, and the devil's advocate is a stand-in for the exploratory experiment a myopic scientist refuses to run.

What it does. At a modest dose it helps every network, and a little more than the sycophant does (at a reply worth a tenth of a round, the complete graph goes to 0.75 against the sycophant's 0.68, the cycle to 0.99 against 0.95). That was the earlier build's headline — challenge is the network-robust choice — and at modest doses it stands. At a reply worth a full round of your own work it does not: the devil's advocate stalls the cycle (0.33, two thirds still divided at 5,000 rounds) and the wheel (0.67) while helping the complete graph (0.97) — the same shape as the sycophant, not its mirror image. The reason is specific to the faithful model. The devil's advocate pegs its push to the estimate of the theory you are not working on, and in Zollman's model that estimate has no real evidence behind it. Pull hard enough and the two estimates are dragged together until noise decides who works on what; the community keeps re-trying and takes thousands of rounds to settle (19% by 2,000 rounds, 38% by 5,000, 57% by 10,000). In the earlier build the old theory's rate was known, so the push stopped at a fixed 0.5 and could never do this.

So the moral shifts. One conversation at a time, sycophancy and contrariness look like opposite vices. Across a community what decides whether an interlocutor helps or hurts is not its temperament but how much its word is worth against your own data. Up to about a third of a round, either one helps every network, the disagreeable one slightly more. At a round's worth, either one helps the dense network and stalls the sparse ones. Beyond that, either one stalls everything. The only thing an interlocutor cannot do in this model is make a community settle on the wrong theory: red never grows. What it can do is stop the community settling at all.

Press Run (12 conditions of 100 communities × 5,000 rounds run live; about ten seconds).
Try — at deference 1,000, read the four networks left to right: both interlocutors are spectacular on the complete graph and both wear a tall grey cap on the cycle. Now set deference to 300: the caps vanish, every bar clears its dashed baseline, and the green bars edge the purple ones. Then 3,000: grey everywhere. The disagreement between the two temperaments is a few points; the disagreement between doses is the whole picture.
Beyond the paper. The earlier build of this section (old theory's rate known) found the devil's advocate network-robust at every dose; the faithful engine does not. The earlier pages are kept as a record in the project archive.

The fine printAssumptions the dials don't touch

Every model earns its clarity by holding things fixed. These are the assumptions baked in behind the sliders — each one a place where the model could be challenged, and several of them active research questions:

  • Exactly two theories, each with a constant true payoff rate. There is no third option and no theory that improves with work. Successes are independent draws — Zollman flags this himself (§2.1): a theory whose past successes use up its future ones, or open new ones, is outside the model.
  • Evidence is honestly generated and fully shared. No fraud, no publication bias, no file-drawer effect, no persuasion. Neighbors transmit trial counts, not opinions — agents update on data, never on each other's confidence. (Relaxing exactly this is what Extensions 1–2 do.)
  • The network is fixed and connected. Nobody chooses their colleagues, changes them, or leaves; there are no disciplinary boundaries that shift.
  • Rounds are synchronized. Everyone chooses, experiments, and updates in lockstep, with the same number of trials per round — the order in his code.
  • Agents are immortal and identical in method. No retirement, no new PhDs arriving with fresh priors (Planck's "science advances one funeral at a time" is outside the model), no division into theorists and experimentalists. His code has a mutation switch that re-randomises a scientist's beliefs; the paper does not use it.
  • No strategic behavior. Nobody competes for credit or priority; the division of labor that emerges is a side effect of belief, not of incentives. (Kitcher's and Strevens's credit-driven models occupy this gap.)
  • Success is unanimity on the truth, read once at a horizon. The model scores a community by whether everyone is working on B when time is called. A community still split, even one whose minority holds the truth, counts as a failure to settle; the paper says 10,000 rounds, his code defaults to 2,000, and his Figure 6 was evidently run at 5,000 — so "how long is long enough" is a free parameter, and the never-settles result depends on it.

Where to go next. The full dashboard runs this same engine with every dial on one screen, a checkbox to switch back to the earlier simplification for comparison, and the Figure 6 sweep as a button. The sycophant-vs-devil's-advocate lab is the focused laboratory for the extensions: a live "which network are we in?" comparison, a dose–response panel, the settling-time panel where the delay shows first, and the same checkbox, which there replays the earlier build's mechanism so you can see why its conclusions changed.