Content-driven games are usually specified by library size, meaning the number of cards, challenges or prompts shipped, and they often disappoint against that specification. Players experience the distribution of situations the items produce, and that distribution is what they judge. We call its entropy replay entropy, and we argue that perceived variety tracks it. Rule-type composition and draw policy drive replay entropy far more than library size does. There is also a timing problem. Designers tend to reason with the coupon-collector timescale, which sits two or three orders of magnitude away from the point where players actually notice repetition. The sections below show how to write a replay-entropy model for your own content, compute both timescales, and decide between adding items and adding factors on the basis of measurement.
Why library size stops predicting perceived variety
Anyone who has shipped a content update knows the pattern. The library grows by thirty percent and the novelty reports do not move. The reason is that players perceive the situation a card creates, and many different cards create the same situation.
Hunicke, LeBlanc and Zubek (2004) separate mechanics from the dynamics they generate and from the aesthetic experience those dynamics produce. Library size is a mechanics-layer quantity and perceived variety is an aesthetics-layer one. Nothing guarantees that a linear increase at the bottom carries through to the top, because the mapping from items to experiences is many-to-one. Salen and Zimmerman treat rules as the locus of meaningful play for the same reason. What a player remembers is the situation the rules created, and rarely the specific token that triggered it.
Formally, let X be the card drawn and let φ be the coarsening a player applies when judging whether a moment feels new. Two cards with different text, art and identifiers, but the same rule type, social configuration and decision structure, collapse to one class under φ. Perceived variety is a property of φ(X), and for any deterministic φ the entropy of the coarsened variable is at most that of the original. So library size buys an upper bound on perceived variety. How much of that bound you actually get depends on how coarse φ is, and that is set by the design more than by some fixed fact about the player.
An entropy formulation of the experience distribution
Here is the model stated precisely. Let E be the set of experience classes induced by φ, and let P be the probability distribution over E that a session actually produces, given the library, the draw policy and the game state. Replay entropy is the Shannon entropy of that distribution (Shannon, 1948):
H(P) = − Σ_{e ∈ E} P(e) · log₂ P(e) bits
V_eff = 2^H(P) effective distinct experiences
Two properties do the work. First, H reaches its maximum of log₂|E| when P is uniform, so a library with a skewed effective distribution behaves like a much smaller one. V_eff, the perplexity of the distribution, is the honest number to put in the headline. Second, H obeys the chain rule, which splits an experience into the parts a designer controls. Writing R for the rule type and C for the card content within that type:
H(R, C) = H(R) + H(C | R)
This is where the model starts to guide decisions. H(R) is bounded by log₂ of the number of rule types, which is in the low tens for most designs. H(C | R) is bounded by log₂ of the average number of cards per type, which runs into the hundreds. On that count content should dominate. But φ discounts the second term heavily and the first barely at all. Swapping which of two hundred prompts appears inside a familiar rule type changes the surface of an experience whose structure the player already knows. Switching rule type changes the structure. If the coarsening keeps only a fraction of a bit per additional card within a type but nearly all of H(R), the term with the smaller theoretical ceiling decides the outcome.
Rule-type composition as the dominant term
There is a second, sharper reason library growth saturates. Effective draw distributions are rarely uniform. Weighting, rarity tiers, unlock gating and players skipping cards all concentrate probability on a minority of items, which gives the heavy-headed, long-tailed shape Zipf described across human selection behaviour. When probabilities fall as a power law with an exponent greater than one, the tail is summable. As the item count grows without bound the distribution converges, and so does its entropy. New items in that tail add arbitrarily little variety, however many there are. With a flatter exponent entropy still grows, but more slowly than the logarithm of library size, and log₂ growth is already a punishing curve on its own: doubling the library buys one bit.
Working memory sets the ceiling on the other side of the decomposition. Miller (1956) established that people can hold and work with only a small number of distinct chunks at once, and rule types are chunks. Each one is a schema the player has to recognise, recall the constraints of and apply under social pressure. Past roughly a dozen types, extra mechanics register as lookup cost instead of variety, and one type too many can lower perceived quality by slowing the table down. The same constraint governs participation architecture in party games, where time spent explaining is time taken from play. Our own party game GUZZL sits in that band on purpose, with 14 distinct rule types across a library of more than 1,080 handcrafted cards. The variety budget goes into recombining a set of mechanics small enough to hold in working memory, and much less into making the card list longer.
Draw policy and the perception of repetition
Once the library and the rule-type mix are fixed, the draw policy determines how often the coarsened experience comes back, and it is the cheapest lever available.
Independent draws with replacement are the worst case, and also the default in most implementations. The number of coincident pairs in a session of n draws from N items grows as n²/(2N), so the probability of at least one exact repeat is approximately 1 − exp(−n²/2N). Here is an illustration with assumed figures, not measured ones: 1,080 items, uniform weighting and a session of 40 draws. Then n²/2N ≈ 0.74 and the repeat probability is close to 0.52. Even with more than a thousand cards, a player has about an even chance of seeing a repeat within one sitting. The problem there is the sampler, and the library can stay as it is.
There are three fixes. Drawing without replacement within a session removes in-session repeats entirely, and it costs one shuffled deck. A cooldown window that forbids the last w draws extends this across sessions and is the stronger control. If w is at or above typical session length, back-to-back recurrence cannot happen. Weighting is the subtler case, because it trades entropy for salience. As an illustration, again assuming independent draws, a rare variant that surfaces 2% of the time contributes about 0.11 bits to H. On the entropy account that is negligible, yet events like that are the moments players talk about afterwards. Entropy measures unpredictability and says nothing about memorability. A composition tuned only to maximise H would strip out exactly the rare, structurally distinct events that give players a story to tell about the session. GUZZL's boss rounds every 10 cards make the same trade from the other side. A deterministic period contributes zero entropy by construction and is there for pacing, which is better explained by tension curves and power reversal than by entropy.
The coupon-collector problem and the wrong timescale
Designers who reach for mathematics usually reach for the coupon-collector problem. It is a standard occupancy result (Feller, 1968) with a known limiting distribution for the completion time (Erdős & Rényi, 1961). The expected number of independent uniform draws needed to see all N distinct items is
E[T_full] = N · H_N ≈ N · (ln N + γ), γ ≈ 0.5772
For N = 1,080 that is roughly 8,170 draws. The number is correct but of little use here, because players do not complain about running out of cards. What they say is "this feels repetitive", and that judgement comes from the first noticeable recurrence. That is the birthday regime, at order √N, far earlier than the collector regime at order N log N.
| Question | Governing regime | Order of magnitude | Illustrative value at N = 1,080 |
|---|---|---|---|
| When is a repeat first noticed | Birthday / coincidence | √N | ≈ 33 draws |
| When is half the library seen | Occupancy | N · ln 2 | ≈ 750 draws |
| When is the library exhausted | Coupon collector | N · (ln N + γ) | ≈ 8,170 draws |
The values assume independent uniform draws over 1,080 items and no cooldown. They are arithmetic from those assumptions, not observed player data. The two relevant timescales differ by more than two orders of magnitude. Optimising against the wrong one, which means buying content to push back an exhaustion date no player will reach, is the most expensive mistake in this area, and the coupon-collector formula is what makes it look reasonable.
Combinatorial interaction as the cheap source of variety
The better approach uses the fact that entropy adds over independent factors. If an experience is made up of several independent components, such as rule type, modifier, target-selection mode and scoring context, then
H(R, M, T, S) = H(R) + H(M) + H(T) + H(S)
Each new orthogonal factor adds its entropy in full, without diluting an existing distribution. A factor with four equally likely levels adds exactly 2 bits. That is the same gain as quadrupling a uniform library, and you get it by writing four rules instead of three thousand cards. A modifier that contradicts a rule type does produce an unplayable state, though, so the real content work is validating the cross-product.
A heist prototype we later shelved was built on this idea. Its 720 challenges spanned 7 categories, but most of the variety came from orthogonal structure layered over them: co-op heist missions, emergency meetings, and a hidden traitor present in roughly 30% of games. The traitor was a single binary factor, worth at most one bit on its own. Yet it changed how players read every other draw in the session, in the way the probability structure of social deduction describes, and that is how interaction terms deliver more than their nominal entropy.
Measuring content saturation from player behaviour
None of this can be settled in a spreadsheet, because φ cannot be observed directly. It has to be estimated from behaviour, and the quantity you can measure is novelty from one session to the next. The library itself tells you little.
For session s, define a novelty rate ν(s): the fraction of drawn experiences the player marks as not seen before, or treats as new in how they play. A single post-session prompt is the direct instrument. Behavioural proxies such as skip rate, time to first action and re-read duration are noisier but continuous, and worth calibrating against self-report on a sample. Fit a decay to ν(s) and report a half-life in sessions. A content roadmap should be held to that half-life, and the card count should come second.
Experimental discipline matters more than the choice of model. Compare cohorts that differ in one composition variable at a time: the same library with and without a cooldown window, the same card count at eleven versus fourteen rule types, the same rule set with and without an added factor. Non-stationarity needs explicit handling. GUZZL's weekly content rotation makes the draw distribution change over time on purpose, and the interaction between content rotation and the forgetting curve is what makes a returning item read as new. Entropy therefore has to be computed over a window long enough to contain a rotation, with cohorts aligned to rotation boundaries. Finally, measure at the coarsening players actually use. If players describe two rule types as the same thing, merge them in the model before computing H.
Decision rules for content composition
Five rules follow from the model. We state them as rules because they go against how content is usually planned.
Report V_eff instead of the item count in every internal specification. A library of 1,080 items whose effective distribution has an entropy of 6 bits has an effective variety of 64, and the gap between 1,080 and 64 is the real design problem.
Fix the sampler before buying content. A cooldown window at or above session length takes a few hours to build and removes a whole class of complaint that a thousand extra cards would only partly hide. Do not make content decisions while draws are still independent and with replacement.
Keep the rule-type count inside the working-memory band and put growth into orthogonal factors. Once H(R) is near its useful ceiling, the next bit costs less as a new modifier axis than as a fifteenth mechanic or another hundred cards.
Add items only where they raise conditional entropy at the coarsening the player uses. Content inside a rule type that is already dense adds variety cheaply only if φ can tell it apart. Content that sets up a new interaction is worth many times its count.
Gate releases on the measured novelty half-life, and not on how much was shipped. A release that moves the half-life has worked. If it does not move it, library size was never the constraint, and the next release will fail in the same way at the same cost.
Limitations
We do not test whether replay entropy predicts retention, satisfaction or revenue. A low-entropy design with strong social framing could well beat a high-entropy one. Entropy measures unpredictability only, and the rare-variant case above shows it parting ways with memorability, so a full account of perceived variety would need a salience term this model lacks. The model also treats φ as fixed and the same for every player. In reality it varies across players and sharpens with experience, so H(φ(X)) drifts downward over a player's lifetime for reasons no content release can fix. Factor additivity assumes independence for convenience. Real modifiers correlate with rule types, and correlated factors deliver less than their summed entropy. The numeric timescales follow from assumptions about uniform independent draws over our own library size and are not observations of player behaviour. The first-party figures (GUZZL's 1,080-plus cards, 14 rule types and 10-card boss period, and the shelved heist prototype's 720 challenges across 7 categories) describe two designs in one genre, which is not a sample you can estimate general parameters from. Last, the framework says nothing about content quality. Entropy counts distinguishable experiences without ranking them, and a library can raise H by adding experiences nobody wants.
References
- Erdős, P., & Rényi, A. (1961). On a Classical Problem of Probability Theory.
- Feller, W. (1968). An Introduction to Probability Theory and Its Applications.
- Hunicke, R., LeBlanc, M., & Zubek, R. (2004). MDA: A Formal Approach to Game Design and Game Research.
- Miller, G. A. (1956). The Magical Number Seven, Plus or Minus Two.
- Salen, K., & Zimmerman, E. (2004). Rules of Play: Game Design Fundamentals.
- Shannon, C. E. (1948). A Mathematical Theory of Communication.
- Zipf, G. K. (1949). Human Behavior and the Principle of Least Effort.