Assumption 1Inquiry is a bandit problem
The model treats theory choice as like a gambler facing two slot machines. Call the slot machines the established methodology A and the challenger B. Each pays off at some fixed rate, and nobody knows what either rate is. In fact B is slightly better: it pays at 0.5, A at 0.5 − ε (Zollman's 0.5 against 0.499). The only way to learn about a theory is to work on it by running trials and counting successes. Working on A teaches you about A and nothing about B, and vice versa.
If ε is large, it's easy to settle which theory is better. If ε is small, then it's hard to tell which theory is better. Two well-run studies, one evaluating each theory, will regularly rank them the wrong way round. Palmer's study can be thought of as one unlucky pull of a bandit's arm. Try generating some studies yourself: each click runs one study of A and one of B.
Assumption 2Scientists are Bayesian learners with beta priors
Each scientist's opinion about a theory is a beta distribution — a curve over the possible payoff rates, summarized by two numbers 〈α, β〉. Its mean, α/(α+β), is the scientist's current best estimate of that theory's rate; the total α+β is how much "weight of experience" the opinion carries. Every scientist holds two such curves, one for A and one for B. Updating is Bayesian conditioning: observe s successes in n trials of a theory, and that theory's α grows by s, its β by n − s.
Two things follow. First, these agents are ideally honest with evidence — they aren't biased, they aren't susceptible to wishful thinking, they don't commit fraud. Whatever goes wrong later cannot be blamed on bad individual reasoning. Second, the prior's mass acts as inertia: an opinion carrying 4 pseudo-observations is swamped by one 1000-trial study (250-to-1); an opinion carrying 4,000 is barely affected. Zollman's baseline agents start with every α and β drawn uniformly from (0, 4] — nearly weightless opinions, swayed by the first study they see. That starting point will be something you can adjust in Step 5.
The widget shows one of the two curves (a scientist's belief about B) and where it stands relative to what that scientist currently thinks the payoff of theory A is. That comparison, not a fixed line at 0.5, is what decides which theory gets worked on next.
Assumption 3Scientists are myopic — and the theory you're not working on stops teaching you anything
Each round, every scientist works on whichever theory currently looks better: B if their estimate of B exceeds their estimate of A, A otherwise. Nobody runs an experiment purely for its information value; nobody thinks "B looks slightly worse, but it's worth one more study to be sure." Zollman motivates this psychologically: careers reward present success — grants, tenure, promotion — not exploration.
Combined with Assumption 1, myopia has a subtle consequence. The moment a scientist moves to A they stop generating evidence about B, so their belief about B freezes at whatever it was when they left. Meanwhile their belief about A keeps sharpening toward A's true rate. Zollman's own example: a scientist who starts thinking A pays 0.25 and B pays 0.75 will work on B, and "so long as the results he receives do not take him too far from the expectation of action 2, he will continue to believe that action 2 is superior and will never learn that his priors regarding action 1 are skewed." The trap is symmetric: whichever theory you are not working on is the one you can never be corrected about. An isolated scientist who leaves B while still believing B is at least as good as A really is will eventually come back, once the evidence about A has disappointed them enough. One who leaves believing B is worse than A really is never will. Each line below is one solo scientist's estimate of B minus their estimate of A; blue while they work on B, red while they work on A; a line stuck below zero is a scientist who has stopped learning about the better theory.
Assumption 4Evidence flows through a fixed network of neighbors
Now the scientists form a community. Each agent has a fixed set of neighbors — the colleagues whose experimental results they see — and each round they update on their own trials plus all of theirs, each result counting toward the theory it was a study of. Sharing is symmetric, honest, and complete; crucially, what travels across an edge is evidence (trial counts), never opinion or deference. Zollman's three structures: the complete graph (everyone sees everything — one big lab meeting), the wheel (a hub sees all; the rim is a ring, each rim scientist seeing two neighbours and the hub), and the cycle (a ring, two neighbors each). A fourth, the line (a ring cut open — sparsest of all), is not in the paper but is the natural next step down.
Naïvely, more communication should help: more data per update, faster learning. That is exactly half right — it is faster. Here is Zollman's headline result: it is faster and less reliable. In a dense network, one Palmer-grade unlucky study of B reaches everyone at once, everyone's weightless prior capsizes together, and the community moves to A in unison — after which nobody is generating evidence about B, and (Assumption 3) nobody's frozen belief about B can be corrected. In a sparse network the same bad news only poisons a neighborhood; somewhere on the far side of the ring, someone is still working on B, still generating the evidence that will eventually reach the others and correct the error.
Assumption 5Beliefs can start extreme — stubbornness as a second mechanism
Zollman's baseline agents hold nearly weightless priors (step 2). His second experiment turns that into a dial: raise the maximum from which every initial α and β is drawn, from 4 up to 10,000. The agents' opinions are as varied as before, but on average they carry far more weight, so it takes proportionally more contrary evidence to move them. This is his model of "extreme beliefs" — in personality terms, a community-wide dose of stubbornness. (Popper: "A limited amount of dogmatism is necessary for progress.")
The result mirrors the network result from the other side: stubbornness protects diversity too. In a dense network, extreme priors mean the first unlucky study can no longer capsize everyone at once; the initially-optimistic agents ride out the bad news, keep working on B, and the truth accumulates. The sweep below is his Figure 6 — success probability as a function of the maximum α/β, one line per network, seven scientists. Watch the complete graph climb as stubbornness rises, watch all three lines meet near 3,000, and watch what happens to the sparse networks past that point.
The interactionTransient diversity — the assumptions working together
Steps 4 and 5 gave two independent protections against premature convergence: limit the flow of information (sparse networks) or weight down opinions (extreme priors). Each works by the same mechanism — keeping somebody working on the minority theory through the dangerous early rounds. Zollman's closing point is what happens when you stack them: a community that is both sparse and stubborn never converges at all. Everyone tends their own theory; inquiry stays open but never closes within any reasonable horizon. Diversity was only ever instrumentally valuable — the goal is transient diversity: disagreement that lasts long enough to test the alternatives, and then ends.
This is the full model, every assumption live at once. The three outcomes map onto the three regimes: premature convergence (the community ends on the worse theory — the Palmer failure), transient diversity (disagreement, then truth), and permanent diversity (still split at the horizon — the question is never settled).
Extension 1The sycophantic LLM — a neighbour that carries opinion, not evidence
Every result so far rests on the rule from Assumption 4: what crosses an edge is evidence — trial counts — never opinion, and by Assumption 2 a scientist folds in whatever arrives as trial counts without asking where it came from. Now relax exactly that. Give each scientist a private interlocutor to consult before every decision. It has run no trials, but it talks as if it had, and the scientist cannot tell the difference. The deference dial says how many trials a reply is worth.
The sycophant hands the scientist their own opinion back. If they currently think the theory they are working on pays about 0.52, the reply is worth c trials that came out at 0.52: the belief about that theory gets α += c·E, β += c·(1 − E). The estimate does not move; it only gets heavier. This is Assumption 5's extreme belief, manufactured on the fly: at the first consultation a weightless prior receives c pseudo-trials at its own mean, which is exactly Zollman's "maximum α/β" of about c. Two things make it more than a step-5 prior. It is topped up every round, so real evidence never swamps it — the scientist's learning rate about the theory in hand is permanently scaled by real trials ÷ (real trials + c). And it is aimed at whatever the scientist happens to believe at that moment, so it entrenches drift as readily as truth. It says nothing about the theory the scientist is not working on — the belief Assumption 3 froze stays light, and stays correctable.
What it does. At a modest dose — a reply worth a third of a round of your own work or less — the sycophant helps every network, because a little manufactured stubbornness is what step 5 said every community could use. At a reply worth a full round of your own work it splits by network exactly as Zollman's extreme priors do: it rescues the dense complete graph (0.57 → 0.96) and stalls the sparse ones (the cycle falls from 0.91 to 0.37, with 63% of communities still divided when the verdict is read at 5,000 rounds). At three rounds' worth it stalls everything, the dense graph included. The stall is a delay, not a lock: on the cycle at one round's worth, 12% of communities have settled by round 2,000, 34% by 5,000, 54% by 10,000. The reason is a holdout problem. A scientist who has moved to A learns about A only from their own trials and any neighbours on A, and the sycophant weighs that belief down; one who started out fond of A can take thousands of rounds to be talked out of it by evidence.
Extension 2The devil's advocate — and why the dose matters more than the temperament
What if the interlocutor disagreed instead? The devil's advocate says "the grass is greener": the theory you are working on is no better than the other, and the other is no worse than this one. In beta terms, the belief about the theory in hand gets c pseudo-trials at the other theory's estimate, and the belief about the other theory gets c at the in-hand estimate — the two estimates are pulled toward each other, by exactly the gap the scientist perceives. This is the one voice that touches the belief Assumption 3 froze: a scientist on A generates no evidence about B, and the devil's advocate is a stand-in for the exploratory experiment a myopic scientist refuses to run.
What it does. At a modest dose it helps every network, and a little more than the sycophant does (at a reply worth a tenth of a round, the complete graph goes to 0.75 against the sycophant's 0.68, the cycle to 0.99 against 0.95). That was the earlier build's headline — challenge is the network-robust choice — and at modest doses it stands. At a reply worth a full round of your own work it does not: the devil's advocate stalls the cycle (0.33, two thirds still divided at 5,000 rounds) and the wheel (0.67) while helping the complete graph (0.97) — the same shape as the sycophant, not its mirror image. The reason is specific to the faithful model. The devil's advocate pegs its push to the estimate of the theory you are not working on, and in Zollman's model that estimate has no real evidence behind it. Pull hard enough and the two estimates are dragged together until noise decides who works on what; the community keeps re-trying and takes thousands of rounds to settle (19% by 2,000 rounds, 38% by 5,000, 57% by 10,000). In the earlier build the old theory's rate was known, so the push stopped at a fixed 0.5 and could never do this.
So the moral shifts. One conversation at a time, sycophancy and contrariness look like opposite vices. Across a community what decides whether an interlocutor helps or hurts is not its temperament but how much its word is worth against your own data. Up to about a third of a round, either one helps every network, the disagreeable one slightly more. At a round's worth, either one helps the dense network and stalls the sparse ones. Beyond that, either one stalls everything. The only thing an interlocutor cannot do in this model is make a community settle on the wrong theory: red never grows. What it can do is stop the community settling at all.
The fine printAssumptions the dials don't touch
Every model earns its clarity by holding things fixed. These are the assumptions baked in behind the sliders — each one a place where the model could be challenged, and several of them active research questions:
- Exactly two theories, each with a constant true payoff rate. There is no third option and no theory that improves with work. Successes are independent draws — Zollman flags this himself (§2.1): a theory whose past successes use up its future ones, or open new ones, is outside the model.
- Evidence is honestly generated and fully shared. No fraud, no publication bias, no file-drawer effect, no persuasion. Neighbors transmit trial counts, not opinions — agents update on data, never on each other's confidence. (Relaxing exactly this is what Extensions 1–2 do.)
- The network is fixed and connected. Nobody chooses their colleagues, changes them, or leaves; there are no disciplinary boundaries that shift.
- Rounds are synchronized. Everyone chooses, experiments, and updates in lockstep, with the same number of trials per round — the order in his code.
- Agents are immortal and identical in method. No retirement, no new PhDs arriving with fresh priors (Planck's "science advances one funeral at a time" is outside the model), no division into theorists and experimentalists. His code has a mutation switch that re-randomises a scientist's beliefs; the paper does not use it.
- No strategic behavior. Nobody competes for credit or priority; the division of labor that emerges is a side effect of belief, not of incentives. (Kitcher's and Strevens's credit-driven models occupy this gap.)
- Success is unanimity on the truth, read once at a horizon. The model scores a community by whether everyone is working on B when time is called. A community still split, even one whose minority holds the truth, counts as a failure to settle; the paper says 10,000 rounds, his code defaults to 2,000, and his Figure 6 was evidently run at 5,000 — so "how long is long enough" is a free parameter, and the never-settles result depends on it.
Where to go next. The full dashboard runs this same engine with every dial on one screen, a checkbox to switch back to the earlier simplification for comparison, and the Figure 6 sweep as a button. The sycophant-vs-devil's-advocate lab is the focused laboratory for the extensions: a live "which network are we in?" comparison, a dose–response panel, the settling-time panel where the delay shows first, and the same checkbox, which there replays the earlier build's mechanism so you can see why its conclusions changed.