Social deduction games are usually described as tests of social reading. Looked at structurally, they produce beliefs that feel better calibrated than they are. The argument here is that the designer's most important lever is the prior probability that a traitor exists at all, more than the traitor's abilities or the voting rule, and that a mixed regime (some rounds with a traitor, some with none) produces strictly more aggregate suspicion than a guaranteed traitor does. The sections below show how to compute the belief dynamics a rule set implies, how to estimate how much information each round leaks, and how to state a fairness constraint that keeps reasoning worth the effort.
The informational structure of hidden-role games
Three properties define the genre. Roles are privately assigned: each player draws a type from a distribution whose form is public and whose outcome is not. Actions are publicly observable to some degree, so the action history becomes a shared record that everyone reasons over. And type can only be communicated through signalling, because speech is unverifiable and therefore free.
Formally, this is a game of incomplete information. The normal-form apparatus of von Neumann and Morgenstern (1944) is where the analysis comes from, but it is not quite the right tool here, since payoffs depend on beliefs about types and not only on strategy profiles. In the MDA framework (Hunicke, LeBlanc & Zubek, 2004), role assignment, voting and reveal rules are mechanics. The suspicion that emerges from them is a dynamic, and the suspense players report is the aesthetic. Designers can only change the mechanics directly.
The genre also turns a classical result on its head. Aumann (1976) showed that agents with a common prior who publicly know each other's posteriors cannot persistently disagree. Hidden roles remove the common prior, and unverifiable speech removes common knowledge of posteriors. A social deduction game breaks both conditions on purpose, so the persistent disagreement at the table is by design and says nothing bad about the players.
Signalling is the only channel left. Following Spence (1973), a signal separates types only when its cost differs by type, and a message that is cheap for everyone conveys nothing at equilibrium. For design, this means innocent players need at least one action a traitor would find expensive, such as a public commitment that constrains later sabotage, a verifiable task completion, or an irreversible allocation of a scarce resource. Without that cost difference, discussion is babble in the game-theoretic sense, however lively it gets.
Bayesian updating as the normative baseline
Write T_i for the event that player i is the traitor and e for an observed action. The odds form of Bayes' rule (Bayes, 1763) is the compact statement of correct updating.
O(T_i | e) = O(T_i) × LR(e), LR(e) = P(e | T_i) / P(e | ¬T_i)
O(T_i) = P(T_i) / (1 − P(T_i)), P(T_i) = p / N
Here N is the number of players and p the existence prior: the probability, before any observation, that a traitor is present at all. Two consequences follow immediately, and the rest of the argument rests on them.
First, the prior enters only through the product p/N. A designer who changes p changes every player's starting odds at once and in the same direction, which no other single mechanic does.
Second, and less obviously, summing P(T_i) over all players gives exactly p. The existence prior is therefore the total amount of suspicion in the system. When p = 1 that total is fixed at one and updating is strictly zero-sum, so every unit of suspicion attached to one player is taken from the others. When p < 1 the total is free, and evidence can inflate it.
How human updating departs from Bayesian inference
Base-rate neglect
Players routinely substitute the likelihood for the posterior. The question they answer at the table is "how traitor-like was that action", which is P(e | T_i), and they treat the answer as if it were P(T_i | e). That substitution throws away p/N entirely. This is the representativeness pattern documented by Tversky and Kahneman (1974), and here it has a recognisable form. "That was a traitor move" is a statement about a likelihood, and it is true far more often than the conclusion people draw from it.
The asymmetry of exculpatory evidence
Once a player has been nominated, later observations get read for consistency with the accusation instead of being scored against both hypotheses. Formally, players apply a likelihood ratio above one to consistent evidence and a ratio near one to inconsistent evidence, when it should be below one, because they attend to P(e | T_i) and neglect P(e | ¬T_i). Belief then only moves in one direction. The effect is severe in games where innocent players produce anomalous actions at a high rate, which means games whose tasks are genuinely difficult. If confused innocents look odd often enough, there is enough material to build a case against any player.
No common knowledge of the base rate
Under a mixed regime players cannot condition on p properly even in principle, because they also have to estimate whether they are in a traitor round, and the same observations bear on both existence and identity. Memory makes this worse. Rounds that had a traitor are memorable and rounds that did not are easily forgotten, so the subjective estimate drifts above the true value. It is the same asymmetric decay that governs how quickly rotated content is forgotten. Even players who update correctly are starting from a base rate that is too high.
The existence prior as the primary design lever
The formal result above can now be read as a design statement. With a guaranteed traitor, aggregate suspicion is conserved. Clearing one player mechanically incriminates everyone else a little, clearing N−1 players identifies the traitor exactly, and the reasoning has an end state that correct inference will reach, as in any deduction problem with a solution.
Under a mixed regime none of that holds. Aggregate suspicion equals p, and any anomalous observation raises the posterior probability that a traitor exists, which lifts the prior for every player at once. So evidence against one player adds suspicion to the system as well as moving it around. There is also no end state, because "several people acted strangely and there was no traitor" stays a live hypothesis until the game reveals otherwise. If the game never reveals the regime, players have no way of learning that this outcome ever happens.
This leads to a result that sounds backwards at first. A group that knows a traitor is present ends up less paranoid than a group that is unsure whether one exists. Knowing makes suspicion zero-sum, and the suspicion then corrects itself as players are cleared. When existence is uncertain, the group can keep producing more suspicion out of ordinary oddities.
We designed a heist prototype around this observation, and later shelved it. A hidden traitor was present in roughly 30% of games and absent in the other 70%, and nobody was told which regime they were in. Its 720 challenges across 7 categories generated the actions players observed, and emergency meetings were where private observations turned into public statements. That makes the meetings a participation structure, deciding whose observations enter the record at all.
Worked example: posterior belief under two traitor regimes
The figures below are a worked example on stated assumptions. Nothing here was measured. Assume N = 6 players, a traitor uniformly distributed among them when present, an observable action class that a traitor produces with probability 0.6 and an innocent with probability 0.2 (so LR = 3), and a single such observation attached to one player. Compare p = 1 with p = 0.30.
| Quantity | Guaranteed traitor (p = 1) | Mixed regime (p = 0.30) |
|---|---|---|
| Prior on a given player | 0.167 | 0.050 |
| Posterior on the accused | 0.375 | 0.136 |
| Posterior on each bystander | 0.125 | 0.045 |
| Aggregate suspicion before | 1.000 | 0.300 |
| Aggregate suspicion after | 1.000 | 0.364 |
| Entropy of the role hypothesis space | 2.58 bits | 1.66 bits |
In the mixed regime the same evidence supports a much weaker case against the accused, 13.6% against 37.5%, yet aggregate suspicion has gone up by about a fifth, from 0.300 to 0.364, while in the guaranteed regime it could not rise at all. The bystanders' posteriors actually fell slightly, from 0.050 to 0.045, so the whole increase sits in the existence term.
Sensitivity to the base rate makes this worse. Suppose a player has quietly set their own existence prior to 0.70 instead of 0.30, which is what the availability pattern described above predicts. From the identical observation they get a posterior of roughly 0.284. Overestimating the base rate slightly more than doubles the strength of the case, and nothing in the evidence they observed points to the error.
Information leakage, entropy and pacing
Treat the state of knowledge as a distribution over hypotheses and measure it in bits. With a guaranteed traitor and N = 6, the hypothesis space has entropy log2(6) = 2.58 bits, all of it about identity. With p = 0.30 the space is {no traitor at 0.70, each player at 0.05} and its entropy is 1.66 bits. The existence bit alone accounts for H(0.30) = 0.88 bits of that, which is more than half.
This has a direct effect on pacing. In a mixed regime most of the uncertainty that can be resolved is about whether, not who, so rounds that shift the existence estimate feel like progress while doing very little to narrow down identity. Pacing that depends on belief rather than on board state is closer to the tension curves and power reversals of social games than to the solution path of a puzzle. If the expected number of rounds to resolution is roughly H divided by the bits leaked per round, the designer sets the numerator through p and N and the denominator through how diagnostic each round's actions are. The two are usually tuned separately, which is a mistake.
How diagnostic a round is depends heavily on the type of task. In a challenge where traitor-optimal and innocent-optimal play look identical, nothing leaks and the round only adds pacing. It may be fun, but players learn nothing from it. Counting how many genuinely distinct observations a finite challenge set can produce is the same accounting problem as replay entropy in a finite card set. So each of the prototype's 7 categories should be graded by how far the two behaviours diverge, and that grade belongs in the decision about when the category gets scheduled.
Measuring leakage from play logs
Leakage can be estimated from ordinary telemetry with three measurements. The first is an accuracy curve: the fraction of games in which the majority vote identifies the traitor, plotted against round index. If the curve is flat, the rounds are leaking nothing. The second is calibration. Bucket players' stated confidence against their actual accuracy and summarise the gap with a Brier score. Some miscalibration is expected in this genre, but its size is a design parameter, and the designer should choose it on purpose instead of inheriting it. The third is diagnosticity per category: for each challenge category, the difference in the distribution of observable outcomes between rounds with a traitor and rounds without one. A category that shows no difference leaks nothing and is there only as decoration.
Epistemic fairness and practical design rules
Call a hidden-role game epistemically fair if a Bayes-rational player can expect a strictly higher payoff than a player who votes at random. When this fails, players eventually notice, because payoffs are the one thing they observe reliably, and they switch to social heuristics: vote the quiet one, vote the newcomer, vote whoever spoke last. When reasoning pays nothing, that is the correct reaction, and once players have learned it, it is hard to unlearn.
Three mechanics commonly break the constraint. Tasks on which innocents produce anomalies at the same rate as traitors have LR = 1 and generate noise that cannot be told apart from signal. Traitor abilities that erase or forge the observable record remove the evidence that inference works from. And reveal rules that never confirm the regime stop players from learning across games at all. Braverman, Etesami and Mossel (2008) show that the balance point of a Mafia-type game can be worked out analytically, and that it does not sit where a naive proportional ratio of factions would put it. What carries over to design is that faction sizes and information channels have to be tuned together.
In practice this gives five rules. Set p first and work out what follows from it. If more than half the entropy sits in the existence bit, the game is about paranoia, and if p = 1 it is about deduction. Both are legitimate designs, but they are different games. Reveal the regime after every game. Hiding whether a traitor existed maximises tension within a game, but it destroys the calibration across games that keeps people playing over the long run. Grade every task by diagnosticity and aim for a deliberate mix instead of the maximum, because rounds that are all diagnostic resolve too fast and rounds with none stall. Give innocents at least one action that genuinely costs a traitor more, or accept that discussion has no separating equilibrium. Finally, keep the return on reasoning positive, and use the accuracy curve to check that it is.
Limitations
No experiment is reported here, so nothing shows that any of these levers changes how people actually play. Every number in the worked example is an assumption chosen to show the arithmetic, and the 30% traitor rate in the prototype was our design choice, not a tuned optimum. The Bayesian baseline treats behaviour as an external signal with fixed likelihoods, which is false. Traitors choose their signals and adapt to whatever inference the group is running, so the likelihood ratios are equilibrium quantities and not constants. The uniform-traitor and independent-observation assumptions are also wrong at real tables, where seating, familiarity and speaking order all correlate with who gets accused. The entropy figures describe information available in principle, not what a human player can hold in mind or use, and we have not modelled memory limits or the cost of deliberation. The biggest gap is that "mixed regimes produce more suspicion" is a statement about aggregate posterior mass under a stated model. It is not a measurement of felt paranoia, and the step from one to the other is asserted here without being demonstrated. Last, epistemic fairness is necessary for reasoning to be worth doing, but it is not enough to make a game worth playing.
References
- Aumann, R. J. (1976). Agreeing to Disagree.
- Bayes, T. (1763). An Essay Towards Solving a Problem in the Doctrine of Chances.
- Braverman, M., Etesami, O., & Mossel, E. (2008). Mafia: A Theoretical Study of Players and Coalitions in a Partial Information Environment.
- Hunicke, R., LeBlanc, M., & Zubek, R. (2004). MDA: A Formal Approach to Game Design and Game Research.
- Spence, M. (1973). Job Market Signaling.
- Tversky, A., & Kahneman, D. (1974). Judgment Under Uncertainty: Heuristics and Biases.
- von Neumann, J., & Morgenstern, O. (1944). Theory of Games and Economic Behavior.