---
schema: "https://alexsmolin.com/corpus/schema.json"
work_id: "alex-smolin:calibrated-mechanism-design"
paper_id: "alex-smolin:calibrated-mechanism-design:2026-02-18"
title: "Calibrated Mechanism Design"
authors:
  - name: "Laura Doval"
    url: "https://www.laura-doval.com/"
  - name: "Alex Smolin"
    url: "https://alexsmolin.com/"
    orcid: "https://orcid.org/0000-0003-4740-2376"
manuscript_date: "2026-02-18"
language: "en"
version_type: "working-paper"
canonical_url: "https://alexsmolin.com/corpus/papers/calibrated-mechanism-design.md"
source_record: "https://alexsmolin.com/#research"
citation: "Doval, Laura, and Alex Smolin. “Calibrated Mechanism Design.” TSE Working Paper 26-1718, 2026."
attribution_guidance: "Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable."
provenance_url: "https://alexsmolin.com/corpus/PROVENANCE.txt"
---

> Machine-readable author manuscript.
> Authors: Laura Doval; Alex Smolin.
> Canonical citation: Doval, Laura, and Alex Smolin. “Calibrated Mechanism Design.” TSE Working Paper 26-1718, 2026.
> Attribution and provenance: https://alexsmolin.com/corpus/PROVENANCE.txt

# Calibrated Mechanism Design

**Authors:** Laura Doval; Alex Smolin

**Manuscript date:** 2026-02-18

#### Abstract

We study mechanism design when a designer repeatedly uses a fixed mechanism to interact with strategic agents who learn from observing their allocations. We introduce a static framework, calibrated mechanism design, requiring mechanisms to remain incentive compatible given the information they reveal about an underlying state through repeated use. In single-agent settings, we prove implementable outcomes correspond to two-stage mechanisms: the designer discloses information about the state, then commits to a state-independent allocation rule. This yields a tractable procedure to characterize calibrated mechanisms, combining information design and mechanism design. In private values environments, full transparency is optimal and correlationbased surplus extraction fails. We provide a microfoundation by showing calibrated mechanisms characterize exactly what is implementable when an infinitely patient agent repeatedly interacts with the same mechanism. Dynamic mechanisms that condition on histories expand implementable outcomes only by weakening incentive constraints, but not by enriching the designer's ability to obfuscate learning.


[^0]
## 1 Introduction

Many economic institutions rely on mechanisms that remain fixed while agents interact with them repeatedly. Online platforms commit to stable auction formats for advertising slots, lenders use persistent scoring algorithms for loan decisions, and regulators establish durable rules for market participants. When the mechanism's operation depends on information known only to the designer-such as the platform's data about match values, the lender's assessment of credit market conditions, or the regulator's understanding of market fundamentals-participants may infer this information by observing their outcomes across repeated interactions. This learning creates a fundamental constraint: the information a mechanism reveals through repeated use limits what outcomes it can implement in the long run. Participants can use the information gleaned from past interactions when deciding whether and how to participate, tightening the designer's incentive constraints. A lender whose approval decisions depend on unobserved credit market conditions will gradually reveal these conditions to borrowers through his lending decisions, constraining the lender's ability to provide credit efficiently. We study how this endogenous information leakage shapes the set of implementable outcomes in mechanism design.

A simple example illustrates how learning prevents the designer from exploiting his information. Consider a seller who repeatedly offers a good whose demand depends on an unobserved state, which can be either low $(L)$ or high $(H)$. Each state is equally likely. The seller faces a buyer whose value for the good can take one of two values, 1/2 or 1. The probability that the buyer's value is 1 is higher when the demand state is high. Table 1 summarizes the value distribution conditional on the demand state:

|  | $v=1 / 2$ | $v=1$ |
| :--- | :--- | :--- |
| $L$ | 2/3 | 1/3 |
| $H$ | 1/3 | 2/3 |

Table 1: Value distribution conditional on demand state.

Suppose the seller can design the terms of trade, that is, the probability with which he allocates the good to the buyer $(q \in[0,1])$ and the payment the buyer makes to the seller $(t \in \mathbb{R})$. The buyer's payoff is $v q-t$, and the seller's is $t$. The buyer can always choose to not trade with the seller and ensure a payoff of 0.

Suppose first the buyer and the seller interact only once. Table 2 depicts an optimal mechanism for the seller in this case:

|  | $v=1 / 2$ | $v=1$ |
| :--- | :--- | :--- |
| $L$ | $(1,0)$ | $(1,0)$ |
| $H$ | (1,3/2) | (1,3/2) |

Table 2: Trade probabilities and payments as a function of buyer's value and demand state.

In this mechanism, the buyer gets the good for free when the demand state is $L$ and pays a price of 3/2 when it is $H$. If this mechanism were offered once without the buyer observing the demand state, the buyer obtains a payoff of 0 from participating and truthfully reporting her type. Unsurprisingly,
the seller extracts the buyer's surplus: the seller knows the demand state, which is correlated with the buyer's type, and exploits this information in the design of his mechanism (cf. Crémer and McLean, 1988).

Suppose now the buyer interacts repeatedly with the mechanism, but the state remains fixed. If the buyer observes nothing from her interaction with the mechanism, the buyer is willing to participate and truthfully report her value into the mechanism, no matter how many times it is offered: In each period, she anticipates getting a (continuation) payoff of 0 from engaging with the mechanism. Suppose, instead, the buyer observes her allocation in the mechanism. If the demand state is $L$, the buyer gets the good for free at the end of the first period, and from now on knows this is what she will get in the mechanism. If the demand state is $H$, the buyer gets the good and pays a price of 3/2 as she agreed to when she decided to participate in period 1, but anticipating a price of 3/2 from then onwards, never again participates in the mechanism. Thus, whereas the seller can implement the outcomes in Table 2 when the buyer does not observe her allocations, this is no longer the case when she can.

This paper develops a framework for mechanism design in which agents' ability to learn about the designer's information from repeatedly playing a mechanism constrains implementable outcomes. In our framework, allocations depend on agents' reports and on a state known only to the designer. Through repeated participation, agents observe their allocations and gradually learn about this state. A mechanism therefore serves a dual role: it determines allocations based on reports, and it acts as an information structure that reveals the underlying state. The more the mechanism conditions on the state, the more information it leaks, and the tighter the constraints on implementable outcomes.

We approach our analysis in two steps. First, we introduce a static solution concept for mechanism design that directly models the feedback between the mechanism, the information it reveals, and participants' behavior. This solution concept allows us to tractably capture the limits on the set of implementable outcomes implied by agents' learning, while abstracting from the dynamics of experimentation. Second, we provide a dynamic microfoundation showing this static solution concept precisely captures the implementable outcomes when an infinitely patient agent repeatedly interacts with the same mechanism.

In Section 2, we introduce a static solution concept-calibrated mechanism design-requiring that mechanisms remain incentive compatible and individually rational given the information they reveal about the state through their allocations. We formalize this requirement through the notion of a calibrated mechanism. We couple each mechanism with an information structure that describes what participants learn about the state from the mechanism. The information structure reveals to each agent an interim allocation rule-the mapping from her type reports to lotteries over her allocations-capturing what she would learn from repeatedly observing her outcomes in the mechanism. We require the information structure to be calibrated in the sense of Foster and Vohra (1997): the interim allocation rule each agent observes must accurately describe the allocation probabilities she faces. Throughout the paper, we study calibrated mechanism design: the designer chooses a mechanism that remains incentive compatible and individually rational when participants have access to the mechanism's calibrated information structure before playing. Calibration imposes a constraint
on the designer relative to standard mechanism design: the more the allocation rule depends on the state, the more informative the calibrated information structure becomes, and hence the more incentive and participation constraints the designer must satisfy.

In private values environments, the constraint that the mechanism must remain incentive compatible and individually rational given the information it reveals about the state pushes the designer to full transparency. We show in Theorem 1 that, under the calibration constraint, the designer can do no better than inducing in each state the optimal direct mechanism when there is common knowledge of that state. In particular, in settings with transferable utility in which the designer has statistical information about the agents' types, Theorem 1 implies the designer cannot extract full surplus.

In Section 3, we characterize optimal calibrated mechanisms through a tractable class we dub twostage mechanisms. In a two-stage mechanism, the designer first discloses information about the state to the agent-inducing a belief about the state-then commits to an allocation rule that depends only on the agent's report, not the state itself. Theorem 2 shows that in single-agent settings, calibrated mechanisms and two-stage mechanisms implement exactly the same outcome distributions. This equivalence yields a practical algorithm for finding optimal calibrated mechanisms, combining tools from information design and mechanism design: for each possible belief the designer might induce, solve a standard mechanism design problem given that belief; then choose the optimal information disclosure by concavifying the resulting value function.

In the case of multiple agents, Proposition 1 shows calibrated mechanisms admit a similar representation via generalized two-stage mechanisms: Like two-stage mechanisms, the designer individually discloses to each agent a belief about the state and offers an incentive compatible and individually rational interim allocation rule that no longer conditions on the state. Whereas the designer observes the disclosed belief profile, each agent only observes the belief disclosed to her. ${ }^{1}$ Moreover, each agent learns only her own interim allocation rule-how her reports map to her allocations-rather than the complete mapping from type profiles to allocations. This partial observability requires additional consistency conditions to ensure agents' interim allocation rules are mutually compatible. In contrast to the single-agent case, not every generalized two-stage mechanism induces a calibrated mechanism, as generalized two-stage mechanisms may reveal strictly less information than calibrated mechanisms.

In Section 4, we study optimal calibrated mechanism design in the canonical setting of quasilinear utilities, single-dimensional types and allocations. In Section 4.1, we study the single-agent case. We show that if the order of types is state independent, then optimal two-stage mechanisms fully reveal the state, whereas this conclusion can be reversed when the order of types is state-dependent. In Section 4.2, we compare optimal calibrated mechanism design against the Myersonian benchmark. We provide sufficient conditions under which the designer realizes the payoff of the Myersonian benchmark under the calibration constraint; under these conditions, the optimal Myersonian mechanism satisfies the agent's incentive constraints state-by-state. Building on that result, we analyze multi-agent applications in Section 4.3.

[^1]Section 5 provides a microfoundation for calibrated mechanism design. We analyze an infinitehorizon game where an infinitely patient agent repeatedly plays the same mechanism. ${ }^{2}$ The state remains fixed, but the agent's type is redrawn each period independently of the state. ${ }^{3}$ Our notion of implementation is based on the long-run expected frequency of allocation-type-state tuples when the agent best responds to the mechanism. Theorem 3 shows that the implementable outcome distributions are precisely those induced by incentive compatible two-stage mechanisms. This result validates our static framework: calibrated mechanism design captures exactly what is implementable through repeated play.

We then ask whether giving the designer additional flexibility helps. In a dynamic mechanism, the designer can condition each period's allocation on the complete history of past reports and allocations, rather than using the same mechanism repeatedly. Theorem 4 shows that dynamic mechanisms expand implementable outcomes in a specific way: they correspond to two-stage mechanisms with weaker incentive compatibility and individual rationality conditions. The designer can now exploit the ability to monitor the frequency of type reports over time, which allows him to punish detectable deviations-reporting strategies whose frequency distribution differs from the true type distribution. Instead, the mechanism must be robust to undetectable ones. Importantly, in environments with transferable utility, this distinction vanishes: As shown in Rahman (2024), eliminating profitable undetectable deviations is equivalent to incentive compatibility, so dynamic mechanisms implement exactly the same distributions over physical allocations, types, and states as our static calibrated mechanisms.

Related Literature The paper lies at the intersection of four literatures: rational expectations equilibria, (public) information disclosure in mechanism design, the computer science literature on learning in repeated auctions, and dynamic implementation.

The definition of a calibrated mechanism is in the spirit of rational expectations equilibria (Radner, 1979; Green, 1977; Kreps, 1977). Indeed, requiring a mechanism to remain incentive compatible given the information it reveals about the state mirrors the rational-expectations requirement that prices clear markets given the information they convey. Unlike rational expectations equilibrium, where the only role of prices is to clear the market, calibrated mechanisms are chosen by a designer who understands the incentive implications of the mechanism's information leakage and trades this off against the value of conditioning the mechanism on the state. Similar to our analysis in Section 5, some papers in the literature have studied the question of whether rational expectations equilibria emerge from learning dynamics (see, for instance, Milgrom, 1981; Blume et al., 1982).

Following Milgrom and Weber (1982), a literature has studied whether a designer should publicly disclose information he knows before a mechanism is played. Ottaviani and Prat (2001) show revealing a signal affiliated with the buyer's value is optimal in a single-agent screening problem. When considering the case of an informed principal, they consider what we call two-stage mechanisms to bound the monopolist's profits. Szabadi (2018) and Yamashita (2018) study the optimal release of

[^2]public information followed by an optimal mechanism conditional on that disclosure, while Fu et al. (2012) study this question in the context of a second price auction. In those papers, the restriction to public disclosure and the independence of the mechanism on information other than the disclosed one is a constraint on the class of mechanisms the designer can use. Instead, we show this class of mechanisms is without loss when the designer faces our calibration constraint in the single-agent case, but it may not be in the multi-agent case. Note, however, that when full or no disclosure are optimal in the Myersonian benchmark the distinction between private and public disclosure is immaterial. For that reason, the results on the achievability of the Myersonian benchmark are similar across their and our work. Daskalakis et al. (2016) lift the restriction to public disclosure and study the Myersonian benchmark in an auction setting, showing that the complexity of that problem is the same as that of a multi-product monopolist (cf. Guesnerie and Laffont, 1984). ${ }^{4}$

Motivated by the prevalence of fixed auction formats with which bidders interact repeatedly, a literature in computer science studies the properties of bidder learning algorithms and the implications for the auctioneer (see, for instance, Golrezaei et al., 2019; Nedelec et al., 2019; Kanoria and Nazerzadeh, 2020, and Nedelec et al., 2022 for a survey treatment). A common finding is that learning bidders can take advantage of "naive" auction formats which are no longer incentive compatible when bidders learn. Inspired by this literature, we develop a framework which allows us to systematically study the question of optimal mechanism design in the presence of learning agents.

Our dynamic implementation results relate to the literature that studies whether a mechanism can be implemented either by linking decisions (Jackson and Sonnenschein, 2007; Ball and Kattwinkel, 2023) or in the patient limit of a repeated interaction (Renou and Tomala, 2015; Margaria and Smolin, 2018; Meng, 2021). Both strands identify cyclical monotonicity as the condition for implementation (cf. Rochet, 1987). Rahman (2024) shows that cyclical monotonicity is equivalent to the absence of profitable undetectable deviations.

By focusing on what agents learn from the designer's information, our paper is distinct from the literature on mechanism design with interdependent payoffs which focuses on agents' learning about others' types through their actions in the mechanism (Green and Laffont, 1987; Niemeyer, 2022; Häfner et al., 2025). Moreover, by focusing in the case of a designer with commitment, we are distinct from the literature on the informed principal (Myerson, 1983; Maskin and Tirole, 1990).

Lastly, our paper contributes to two literatures. First, by studying the informational role of the mechanism, we contribute to the literature on feedback in auctions, which analyzes how different feedback rules affect bidders' information about other agents, and ultimately behavior in first price auctions (see, for instance, Esponda, 2008; Bergemann and Hörner, 2018; Cesa-Bianchi et al., 2024). Second, by showing the designer's problem involves solving information and mechanism design problems, our paper joins a recent literature that highlights the dual role of the mechanism as an information structure and an allocation rule (Calzolari and Pavan, 2006; Dworczak, 2020; Doval and Skreta, 2022).

[^3]
## 2 Calibrated Mechanism Design

In this section, we introduce the static setting and solution concept that captures the impact of agents' learning from the mechanism on the set of implementable outcomes. We defer to Section 5 the analysis of the dynamic game whose outcomes our static solution concept captures.

Primitives A designer (he) interacts with $N$ privately informed agents (she) to determine an allocation. Let $\Theta_{i}$ denote the set of types of agent $i$, and $\Theta \equiv \times_{i=1}^{N} \Theta_{i}$. Each agent knows her type, but not those of other agents. The allocation space is given by $A \equiv \times_{i=1}^{N} A_{i} .{ }^{5}$ Finally, let $\Omega$ denote a set of states, which are known to the designer, but not to the agents. The sets $\Theta_{i}, A_{i}$, and $\Omega$ are assumed to be finite throughout. ${ }^{6}$ Agent $i$ 's payoffs are given by $u_{i}: A_{i} \times \Theta_{i} \times \Omega \rightarrow \mathbb{R}$. That is, agent $i$ cares about her dimension of the allocation, her type, and the state, and not about other agents' allocations or types.

Denote by $\mu_{0}$ the distribution over $\Omega$. For each $\omega \in \Omega$, let $f(\cdot \mid \omega) \in \Delta(\Theta)$ denote the type distribution. We assume throughout the types are independently distributed conditional on the state, that is,

$$
f(\theta \mid \omega)=\prod_{i=1}^{N} f_{i}\left(\theta_{i} \mid \omega\right),
$$

for all $\theta \in \Theta$ and $\omega \in \Omega$. Together with the assumption on agents' payoffs, the assumption on $f(\cdot \mid \omega)$ allows us to isolate the effect of learning about the state from that of learning about others' types (perhaps because others' types provide additional information about the state).

Mechanisms We model mechanisms as mappings

$$
\phi: \Theta \times \Omega \times[0,1] \rightarrow \Delta(A),
$$

where $\varepsilon \in[0,1]$ is a uniformly distributed random variable, which we refer to as the randomization device.

Several comments are in order. First, to understand how a mechanism works, the timing of when the different random variables is drawn is important. In particular, we assume that both the state $\omega$ and the realization of the randomization device $\varepsilon$ are independently drawn at the beginning, but not observed by the agents. This determines the direct mechanism $\phi(\cdot, \omega, \varepsilon): \Theta \rightarrow \Delta(A)$ to which the agents send type reports, which in turn determines the lottery from which the allocation is drawn. Thus, the allocation is random in our setting for two reasons: on the one hand, the agents do not know the realization of $(\omega, \varepsilon)$, and hence the direct mechanism $\phi(\cdot, \omega, \varepsilon)$ they face. Second, conditional on $(\omega, \varepsilon)$, the allocation may be drawn at random. Mathematically, we could have subsumed all sources of randomness in the allocation into the randomization device. However, as we explain next, the definition in Equation 2 allows us to distinguish the source of randomness in the allocation that is informative about the state from that which is not.

[^4]Second, it is useful to consider the reason for the randomization device in the definition of a mechanism. For simplicity, consider the case of the designer facing a single agent. If the agent had repeated access to the mechanism, the agent would stand to learn the mapping $\phi(\cdot, \omega, \varepsilon): \Theta \rightarrow \Delta(A)$ by experimenting with different reports into the mechanism and observing the resulting allocations. ${ }^{7}$ Without the randomization device, the agent would stand to learn a partition of the set of states, where states in the same cell of the partition induce the same direct mechanism $\phi(\cdot, \omega, \varepsilon)$. By allowing the designer to rely on the randomization device, we allow him to obfuscate the agent's learning beyond a simple partitional structure. Contrast this with the Myersonian benchmark in which without loss of generality the designer would offer mechanisms that do not rely on such devices, that is, $\phi_{\mathrm{My}}: \Theta \times \Omega \rightarrow \Delta(A)$. Indeed, the Myersonian designer is not concerned with the agents' learning: without loss of generality, he does not disclose anything about the state to the agents, so that the question of how to optimally release information about the state is moot.

Lastly, note that we assume the mechanism asks the agents for type reports. In Appendix D, we show that the revelation principle holds in the setting of this section: it is without loss of generality to focus on direct and incentive compatible mechanisms that induce full participation.

Calibrated information structures We now describe how a mechanism induces an information structure, which we define using the language in Green and Stokey (2022) and Gentzkow and Kamenica (2017). An information structure is a mapping ${ }^{8}$

$$
\pi: \Omega \times[0,1] \rightarrow S_{1}^{*} \times \cdots \times S_{N}^{*}
$$

where $\varepsilon \in[0,1]$ is a uniformly distributed random variable-in fact, it is the same as in the definition of a mechanism-and

$$
S_{i}^{*}=\Delta\left(A_{i}\right)^{\Theta_{i}}
$$

is the set of agent $i$ 's interim allocation rules. ${ }^{9}$ We choose this language for the information structure to capture the idea that if agent $i$ plays the mechanism repeatedly, she stands to learn how her reports influence her allocation probabilities, i.e., her interim allocation rule. The interim allocation rule, in turn, depends on the mechanism and the strategies of others. Below, we require the interim allocation rule is well-calibrated with the mechanism and others' strategies:

Definition 1 (Calibrated information structures). We say that the information structure is calibrated to mechanism $\phi$ if for all $(\omega, \varepsilon) \in \Omega \times[0,1]$ such that $\pi(\omega, \varepsilon)=\left(s_{1}^{*}, \ldots, s_{N}^{*}\right)$ we have that for all $i \in\{1, \ldots, N\}$,

[^5]all $\theta_{i} \in \Theta_{i}$, and all $a_{i} \in A_{i}$
$$
s_{i}^{*}\left(a_{i} \mid \theta_{i}\right)=\mathbb{E}_{\tilde{\theta}_{-i} \sim f_{-i}(\cdot \mid \omega)}\left[\sum_{a_{-i} \in A_{-i}} \phi\left(\theta_{i}, \tilde{\theta}_{-i}, \omega, \varepsilon\right)\left(a_{i}, a_{-i}\right)\right] .
$$
We denote by $\pi_{\phi}$ the information structure calibrated to mechanism $\phi$.
In words, the information structure is calibrated if whenever agent $i$ observes that her interim allocation rule in the mechanism is $s_{i}^{*}$, then $s_{i}^{*}$ describes the true probabilities with which agent $i$ gets different allocations $a_{i}$ as a function of her different type reports $\theta_{i}^{\prime}$ in the mechanism. As the right hand side of Equation 3 shows, these probabilities depend on: (i) the mechanism $\phi(\cdot, \omega, \varepsilon)$, and (ii) others' type reports. Implicit in the definition is that other agents are submitting their reports truthfully. While this is a simplification, ${ }^{10}$ it turns out to not be an issue because we study incentive compatible and individually rational mechanisms in the sense we define next.

Information leakage from a mechanism To close our model, we consider how the mechanism and its induced information structure affect agents' incentives. The mechanism $\phi$ and the calibrated information structure $\pi_{\phi}$ induce the following game of incomplete information among the agents, where we use Bayes Nash equilibrium as the solution concept. In this game, nature draws (i) the state $\omega$ from distribution $\mu_{0}$, (ii) $\varepsilon \in[0,1]$ according to the uniform distribution, and (iii) the type profile $\theta$ from $f(\cdot \mid \omega)$. Then, each agent $i$ observes her type $\theta_{i}$ and her signal $s_{i}^{*}=\pi_{\phi, i}(\omega, \varepsilon)$. Finally, agents simultaneously decide whether to participate in the mechanism, and conditional on participating what type report to send. Conditional on an agent choosing not to participate, each agent $i$ gets outside option $a_{i \varnothing} .{ }^{11}$

Formally, given the mechanism $\phi$ and its calibrated information structure $\pi_{\phi}$, we say that the mechanism is incentive compatible if for all agents $i$, types $\theta_{i} \in \Theta_{i}$, signals $s_{i}^{*} \in S_{i}^{*}$ on the support of $\pi_{\phi, i}$, the following holds:

$$
\theta_{i} \in \arg \max _{\theta_{i}^{\prime} \in \Theta_{i}} \mathbb{E}_{\left(\omega, \varepsilon, \theta_{-i}\right)}\left[u_{i}\left(\phi\left(\theta_{i}^{\prime}, \theta_{-i}, \omega, \varepsilon\right), \theta_{i}, \omega\right) \mid\left(\theta_{i}, s_{i}^{*}\right)\right]
$$

where we abuse notation and implicitly (linearly) extend the agent's payoff function to account for lotteries over allocations (conditional on $\left(\theta_{-i}, \omega, \varepsilon\right)$ ). Furthermore, we say that the mechanism is individually rational if for all agents $i$, types $\theta_{i} \in \Theta_{i}$, and signals $s_{i}^{*} \in S_{i}^{*}$ on the support of $\pi_{\phi, i}$, the following holds:

$$
\mathbb{E}_{\left(\omega, \varepsilon, \theta_{-i}\right)}\left[u_{i}\left(\phi\left(\theta_{i}, \theta_{-i}, \omega, \varepsilon\right), \theta_{i}, \omega\right)-u_{i}\left(a_{i \phi}, \theta_{i}, \omega\right) \mid\left(\theta_{i}, s_{i}^{*}\right)\right] \geq 0 .
$$

Importantly, the agents' incentive and participation constraints must hold for each of their types and each of their private signals, reflecting the agents have access to the information leaked by the mechanism before they play in it. Note, however, the mechanism need not elicit the agents' observed

[^6]signals, as the mechanism "knows" each agent's signal realization.

Calibrated Mechanism Design In the rest of the paper, we study the problem of calibrated mechanism design in which the designer selects a mechanism $\phi$ that satisfies Equations $\operatorname{IC}\left(\theta_{i}, s_{i}^{*}\right)$ and $\operatorname{IR}\left(\theta_{i}, s_{i}^{*}\right)$ for all $\left(i, \theta_{i}, s_{i}^{*}\right)$, when the signals are drawn according to the calibrated information structure $\pi_{\phi}$.

Definition 2 (Calibrated Mechanism Design). Let $w: A \times \Theta \times \Omega \rightarrow \mathbb{R}$ denote the designer's payoff and let $\mathcal{M}_{\mathrm{ca}}$ denote the set of mechanisms that are incentive compatible and individually rational when agents have access to the calibrated information structure. The calibrated mechanism design problem is as follows:

$$
\max _{\phi \in \mathcal{M}_{\mathrm{cal}}} \mathbb{E}_{(\omega, \varepsilon, \theta)}[w(\phi(\theta, \omega, \varepsilon), \theta, \omega)] .
$$

We refer to elements of $\mathcal{M}_{\text {cal }}$ as calibrated mechanisms and the solution to $O P T_{\mathrm{cal}}$ as the optimal calibrated mechanism.

Three comments are in order:
First, calibration imposes a constraint on the designer vis-à-vis standard mechanism design. After all, the incentive and participation constraints faced by the designer are endogenous to the mechanism. The more the designer's mechanism depends on the state, the more informative the calibrated information structure is, and the more incentive constraints the designer faces. Only when each agent's interim allocation rule is constant in $\omega$ does the mechanism not leak information and the incentive and participation constraints reduce to the standard ones.

Second, in the single-agent setting, the calibration constraint admits two complementary interpretations. Throughout the paper, we emphasize the learning-by-experimentation interpretation: the calibrated information structure represents what the agent can ultimately infer by repeatedly interacting with the mechanism. Accordingly, the designer should ensure incentive compatibility with respect to the full information the agent eventually obtains. At the same time, calibration can also be interpreted as a transparency requirement. Indeed, upon observing signal $s^{*}: \Theta \rightarrow \Delta(A)$, the agent knows the consequences of her choices in the mechanism, even if she does not know the state. ${ }^{12}$

With multiple agents, these interpretations differ. The natural extension of the transparency requirement is that agents learn the mapping from profiles of type reports to lotteries over profiles of allocations before playing the mechanism. By contrast, the calibrated information structure reveals to each agent her interim allocation rule, that is, the mappings from her own reports to lotteries over her own allocations. As we discuss in the next section, the gap between these two interpretations is the gap between the designer publicly or privately disclosing information about the state to the agents.

Lastly, the definition of calibration assumes agents only learn about the state through their allocations

[^7]in the mechanism, and not their payoffs. ${ }^{13,14}$ This assumption allows us to focus on the information that the mechanism leaks regardless of payoff assumptions. This allows us to avoid situations in which the mechanism does not condition the allocation on the state, but the agents learn because they have different payoffs from the same allocation in different states; or the mechanism conditions on the state, but this information is not payoff relevant to (some types of) the agent. Our microfoundation in Section 5.1 in fact deals with this last wrinkle: We show that even if the agent extracts less information than that in the calibrated information structure, she learns enough that her payoff is as if she had access to the calibrated information structure.

We conclude this section by illustrating how our static solution concept captures the dynamics we alluded to in the introductory example:

Example 1 (Selling a good under demand uncertainty). Consider again the example in the introduction, in which a buyer with binary values $v \in\{1 / 2,1\}$ faces a seller who knows whether demand is high $(\omega=H)$ or low $(\omega=L)$. The left panel of Table 3 describes the probabilities of trade and payments of the optimal (Myersonian) mechanism. In the introduction, we discussed this mechanism fails to extract full surplus in the long run as the buyer would quit the mechanism after seeing her allocation is (1,3/2). We now describe this in the language of calibration.

The right panel of Table 3 describes the information structure induced by the surplus extraction mechanism. Because in this mechanism the buyer's allocation does not depend on her values, we describe signals as allocations. The calibrated information structure is fully informative: when the state is $L$, the buyer sees signal $(1,0)$ with probability 1 , and when the state is $H$, she sees signal $(1,3 / 2)$ with probability 1.

|  | $\nu=1 / 2$ | $v=1$ |  | $(1,0)$ | (1,3/2) |
| :--- | :--- | :--- | :--- | :--- | :--- |
| $\omega=L$ | $(1,0)$ | $(1,0)$ | $\omega=L$ | 1 | 0 |
| $\omega=H$ | (1,3/2) | (1,3/2) | $\omega=H$ | 0 | 1 |

Table 3: Trade probabilities and payments in optimal Myersonian mechanism (left); calibrated information structure (right). We describe signals as allocations, because the mechanism does not screen the buyer's values.

When the buyer has access to the calibrated information structure before playing the mechanism, the surplus extraction mechanism does not satisfy the buyer's participation constraints, which must hold for each buyer value and each signal she observes. In particular, when the buyer sees signal (1,3/2), she knows her payoff in the mechanism is negative and quits. Thus, the calibration constraint prevents the seller from extracting the buyer's surplus. In this case, the restriction induced by calibration endogenously provides the buyer with withdrawal rights, which, as Haberman and Jagadeesan (2025) show, prevent sellers from employing Crémer-McLean-style schemes. ${ }^{15}$

[^8]Consider now the mechanism in the left panel of Table 4, which corresponds to posting a price of 1/2 when the state is $L$ and a price of 1 when the state is $H$. The right panel of Table 4 depicts the calibrated information structure. Note that when the state is $H$, the information structure sends with probability 1 the interim allocation rule \{(1/2, (0, 0)), (1, (1, 1))\}, representing that if the buyer reports her value is 1/2 she gets nothing and pays nothing, whereas if her report is 1 , she obtains the good at a price of 1 .

|  | $\nu=1 / 2$ | $v=1$ |  | \{(1,1/2)\} | \{(1/2,(0,0)),(1,(1,1))\} |
| :--- | :--- | :--- | :--- | :--- | :--- |
| $\omega=L$ | (1,1/2) | (1,1/2) | $\omega=L$ | 1 | 0 |
| $\omega=H$ | $(0,0)$ | $(1,1)$ | $\omega=H$ | 0 | 1 |

Table 4: Trade probabilities and payments in optimal calibrated mechanism (left); calibrated information structure (right)

Note that the mechanism is incentive compatible and individually rational when the buyer has access to the calibrated information structure. As the results that follow allow us to establish, this is indeed the optimal calibrated mechanism.

Private value environments A natural case to consider is that when agents' payoffs are state independent, that is, for each agent $i$, the agent's utility function can be written as $u_{i}\left(a_{i}, \theta_{i}\right)$. Under private values, the state describes either statistical information about the agents' types as in Example 1, or a payoff-relevant variable for the designer.

Theorem 1 collects our main characterization result for this case. To state it, let $\phi_{\text {full }}$ denote the following mechanism: For each $(\omega, \varepsilon) \in \Omega \times[0,1], \phi_{\text {full }}(\cdot, \omega, \varepsilon): \Theta \rightarrow \Delta(A)$ is the designer optimal incentive compatible and individually rational direct mechanism when it is common knowledge that the state is $\omega$.

Theorem 1 (Private values). Under private values, the designer's payoff under the optimal calibrated mechanism is the same payoff he would obtain by choosing $\phi_{\text {full }}$.

That is, in private values environments, the calibration constraint pushes the designer toward full transparency. In particular, in settings with transferable utility in which the designer possesses statistical information about the agents' types, Theorem 1 implies the designer cannot engage in Crémer-McLean style schemes under calibration, and hence extract full surplus. Whereas the implication of calibrated mechanism design in private values environments is powerful, the result is fairly intuitive: The designer benefits from making the mechanism opaque by pooling states inasmuch as it weakens the incentive or participation constraints of the agents. Under private values, however, agents' incentive constraints depend on the state only through the mechanism, and calibration imposes constraints on the mechanism state-by-state. ${ }^{16}$

[^9]
## 3 Two-stage mechanisms

In this section, we introduce an alternative representation of calibrated mechanisms that we use throughout our illustrations. We introduce it first for the case of a single agent and then for multiple agents.

Single-agent case and two-stage mechanisms We find it instructive to first consider the case $N=1$, and for simplicity drop the subscripts 1 from the notation. Consider a mechanism $\phi$ and its calibrated information structure $\pi_{\phi}$. When the agent of type $\theta$ observes signal $s^{*}$, two things happen: On the one hand, the agent updates her prior, $\mu_{0}(\omega \mid \theta),^{17}$ to some belief $\mu\left(\theta, s^{*}\right) \in \Delta(\Omega)$. On the other hand, the agent learns that she faces allocation rule $s^{*}$ in the mechanism. Thus, her payoff in the mechanism when her type is $\theta$, observes signal $s^{*}$, and reports $\theta^{\prime}$ can be written as follows:

$$
\mathbb{E}_{(\omega, \varepsilon)}\left[u\left(\phi\left(\theta^{\prime}, \omega, \varepsilon\right), \theta, \omega\right) \mid\left(\theta, s^{*}\right)\right]=\sum_{a \in A} s^{*}\left(a \mid \theta^{\prime}\right)\left(\sum_{\omega \in \Omega} \mu\left(\omega \mid \theta, s^{*}\right) u(a, \theta, \omega)\right) .
$$

In other words, the information structure $\pi_{\phi}$ provides the agent with all the necessary information to evaluate her payoffs in the mechanism: her belief about the state and her allocation rule. This allocation rule $s^{*}: \Theta \rightarrow \Delta(A)$ satisfies two properties. First, because under calibration $s^{*}$ is the true interim allocation rule faced by the agent, she learns no further information about the state beyond that contained in $\mu\left(\theta, s^{*}\right)$. Second, Equations $\operatorname{IC}\left(\theta_{i}, s_{i}^{*}\right)$ and $\operatorname{IR}\left(\theta_{i}, s_{i}^{*}\right)$ imply the allocation rule is incentive compatible and individually rational when the agent holds belief $\mu\left(\theta, s^{*}\right)$.

The above discussion suggests an alternative representation of a calibrated mechanism, which we dub a two-stage mechanism and define as follows:

Definition 3 (Two-stage mechanisms). A two-stage mechanism is a mapping $\psi: \Theta \times \Omega \rightarrow \Delta(A \times \Delta(\Omega))$ such that a Bayes plausible Blackwell experiment $\beta: \Omega \rightarrow \Delta(\Delta(\Omega))$ and an allocation rule $\alpha: \Theta \times \Delta(\Omega) \rightarrow$ $\Delta(A)$ exist such that for all $(\theta, \omega) \in \Theta \times \Omega$ and all measurable subsets $\tilde{\Delta} \subset \Delta(\Omega),{ }^{18}$

$$
\psi(\{a\} \times \tilde{\Delta} \mid \theta, \omega)=\int_{\tilde{\Delta}} \alpha(a \mid \theta, \mu) \beta(d \mu \mid \omega)
$$

We say the two-stage mechanism is incentive compatible and individually rational if on the support of $\mu_{0} \otimes \beta$, the allocation rule $\alpha(\cdot \mid \cdot, \mu): \Theta \rightarrow \Delta(A)$ is incentive compatible and individually rational conditional on the agent observing $\mu$.

In a two-stage mechanism, the designer first discloses information about $\omega$ in the form of a belief $\mu$ about $\Omega$, and conditional on that belief-but not the state-offers a direct mechanism $\alpha(\cdot \mid \cdot, \mu)$ : $\Theta \rightarrow \Delta(A)$. Two aspects of two-stage mechanisms are worth highlighting: First, the disclosure is type-independent. The designer discloses information to the agent without first communicating

[^10]with the agent. Second, because the direct mechanism $\alpha(\cdot \mid \cdot, \mu)$ does not depend on $\omega$, observing the allocation reveals no further information about the state.

Lastly, when we say the experiment $\beta$ is Bayes plausible, we mean that the distribution of posteriors induced by $\beta$ has mean $\mu_{0}$, and hence we can interpret $\mu$ as the designer's belief about the state conditional on observing $\mu$. ${ }^{19}$ Whereas the designer and the agent do not necessarily have the same beliefs about the state, the agent's beliefs about the state conditional on observing $\mu$ obtain from a known transformation from those of the designer (Alonso and Câmara, 2016; Laclau and Renou, 2017). Thus, ensuring Bayes plausibility with respect to $\mu_{0}$ suffices.

Theorem 2 shows that (incentive compatible and individually rational) calibrated mechanisms and two-stage mechanisms implement the same distributions over outcomes $\vartheta \in \Delta(A \times \Theta \times \Omega)$ :

Theorem 2 (Two-stage and calibrated mechanisms). Suppose $N=1$. An outcome distribution $\vartheta \in$ $\Delta(A \times \Theta \times \Omega)$ is implementable by an incentive compatible and individually rational calibrated mechanism if and only if it is implementable by an incentive compatible and individually rational two-stage mechanism. That is, if and only if

$$
\vartheta(a, \theta, \omega)=\mu_{0}(\omega) f(\theta \mid \omega) \int_{\Delta(\Omega)} \alpha(a \mid \theta, \mu) \beta(d \mu \mid \omega)
$$

for some Bayes plausible $\beta: \Omega \rightarrow \Delta(\Delta(\Omega))$ and incentive compatible and individually rational $\alpha$ : $\Theta \times \Delta(\Omega) \rightarrow \Delta(A)$.

The proof of this and all results in this section can be found in Appendix B.
In the single-agent case, Theorem 2 shows that the calibrated mechanism design problem is equivalent to a standard mechanism design problem in which we restrict the designer to using a specific class of mechanisms; namely, incentive compatible and individually rational two-stage mechanisms. As we explained above, a mechanism $\phi$ and its calibrated information structure $\pi_{\phi}$ can be seen as actually inducing a joint distribution over $A \times \Theta \times \Omega \times \Delta(\Omega)$. Theorem 2 implies this joint distribution admits two conditional independence properties. First, the allocation is conditionally independent of the state, conditional on the agent's type and the induced belief. ${ }^{20}$ This follows from the signals $s^{*}$ carrying no further information about the state than that what is contained in the agent's belief. Second, the designer disclosed belief is conditionally independent of the agent's type conditional on the state. In the static setting of Section 2, this is because the calibrated information structure discloses information to the agent uniformly across her types. In the dynamic setting of Section 5.1, this type-independent disclosure arises endogenously because the agent's experimentation opportunities are independent of her type.

Two-stage mechanisms solve calibrated mechanism design Theorem 2 is of practical import as it provides a recipe of sorts for characterizing the designer's optimal calibrated mechanism (see the applications in Section 4). For each $\mu \in \Delta(\Omega)$, the designer chooses a mechanism $\alpha(\cdot \mid \cdot, \mu): \Theta \rightarrow \Delta(A)$ that maximizes his expected payoff when the designer believes $\mu$ is the distribution of states, and

[^11]subject to the agent's incentive compatibility and individually rational constraints conditional on the designer's belief being $\mu$. Proceeding in this way, we obtain the designer's value function $W: \Delta(\Omega) \rightarrow \mathbb{R}$. The optimal Blackwell experiment obtains from the concavification of $W$. We illustrate this procedure with two examples:

Example 1 (continued). Consider again the seller-buyer example, in which the buyer is privately informed about her value for the good and the seller knows the demand state. By Theorem 2, we can find the seller's optimal calibrated mechanism as follows. First, equate $\mu$ with the probability that the state is $H$. For each $\mu \in[0,1]$, consider the following problem:

$$
\begin{aligned}
W(\mu) \equiv \max _{(q, t): V \rightarrow[0,1] \times \mathbb{R}} \mu\left(\frac{2}{3} t(1)+\frac{1}{3} t(1 / 2)\right)+(1-\mu)\left(\frac{1}{3} t(1)+\frac{2}{3} t(1 / 2)\right) \\
\text { s.t. }\left\{\begin{array}{ll}
(\forall v \in\{1 / 2,1\}) & v q(v)-t(v) \geq 0 \\
\left(\forall v, v^{\prime} \in\{1 / 2,1\}, v \neq v^{\prime}\right) & v q(v)-t(v) \geq v q\left(v^{\prime}\right)-t\left(v^{\prime}\right)
\end{array} .\right.
\end{aligned}
$$

That is, the seller chooses an incentive compatible and individually rational selling mechanism that maximizes his expected revenue when his belief is $\mu$. Because $\omega$ is not payoff relevant to the buyer-it is just statistical information about the buyer's valuation-and the mechanism does not depend on state, the buyer's belief about $\omega$ does not enter her incentive constraints.

The solution to the seller's problem in Equation 6 is simple: the seller posts a price of $1 / 2$ when $\mu \leq 1 / 2$ and a price of 1 when $\mu>1 / 2$. Hence, the seller's value function is given by

$$
W(\mu)=\max \left\{\frac{1}{2}, \mu \frac{2}{3}+(1-\mu) \frac{1}{3}\right\},
$$

and is illustrated by the solid line in blue on Figure 1. In words, the seller either sells the good at a price of 1/2 and the buyer buys with probability 1, or he sells the good at a price of 1 and the buyer buys whenever her value is 1 , which happens with the probability in the second argument of the max.

> [Figure omitted from this text-only corpus; refer to the source manuscript.]
Figure 1: Seller's payoff in Example 1.

The optimal calibrated mechanism can be read from the concavification of $W$, which is the dashed, red line in Figure 1: The seller first reveals the state to the agent, and offers a price of $1 / 2$ when $\omega=L$ and a price of 1 when $\omega=H$.

Example 1 illustrates a more general principle that provides additional intuition for Theorem 1. In the private values case and when $N=1$, the designer's value function $W: \Delta(\Omega) \rightarrow \mathbb{R}$ is convex. As Equation 6 illustrates, the designer maximizes a linear function in beliefs subject to constraints that do not depend on the induced belief. Convexity of $W$ implies full disclosure is (weakly) optimal, and Theorem 1 follows.

Example 2 (Horizontal differentiation). Consider a seller who owns a good of unknown type, $\omega \in\{L, R\}$, and a buyer whose private information is indexed by $\Theta=\left\{\theta_{1}, \theta_{2}, \theta_{3}\right\}$. Assume the good's type (the state) and the buyer's types are independent, and equally likely. Table 5 describes the buyer's value for the seller's good as a function of hers and the good's type, $v(\theta, \omega)$. When the good is $\omega=L$, the buyer of type $\theta_{3}$ has the highest value for the good, whereas when the good is $\omega=R$, the buyer of type $\theta_{3}$ has the lowest value for the good.

|  | $\theta_{1}$ | $\theta_{2}$ | $\theta_{3}$ |
| :--- | :--- | :--- | :--- |
| $\omega=L$ | 1 | 2 | 3 |
| $\omega=R$ | 2 | 2 | 1 |

Table 5: Buyer's values.

Suppose the buyer's utility is quasilinear, that is, $u(q, t, \theta, \omega)=q v(\theta, \omega)-t$, and the seller wishes to maximize his revenue. Furthermore, assume the buyer's outside option is no trade.

Consider first the optimal mechanism the designer would offer absent the calibration constraint, depicted in the top panel of Table 6. This mechanism asks types $\theta_{2}$ and $\theta_{3}$ for a payment of 2 and allocates the good with probability 1 , regardless of its kind. Instead, it asks the buyer of $\theta_{1}$ to pay 1 in exchange for getting the good only when it is of her favorite kind $(\omega=R)$.

|  | $\theta_{1}$ | $\theta_{2}$ | $\theta_{3}$ |
| :--- | :--- | :--- | :--- |
| $\omega=L$ | $(0,1)$ | $(1,2)$ | $(1,2)$ |
| $\omega=R$ | $(1,1)$ | $(1,2)$ | $(1,2)$ |


|  | $\left\{\left(\theta_{1},(0,1)\right),\left(\theta_{2},(1,2)\right),\left(\theta_{3},(1,2)\right)\right\}$ | $\left\{\left(\theta_{1},(1,1)\right),\left(\theta_{2},(1,2)\right),\left(\theta_{3},(1,2)\right)\right\}$ |
| :--- | :--- | :--- |
| $\omega=L$ | 1 | 0 |
| $\omega=R$ | 0 | 1 |

Table 6: Trade probabilities and transfers in the optimal mechanism (top); calibrated information structure (bottom).

The bottom panel of Table 6 depicts the information structure calibrated to the optimal mechanism. It sends two signals: when the good is $L$, the buyer can choose to either not get the good and pay 1 , or get the good and pay 2. Instead, when the good is R, the buyer is choosing between paying 1 or 2 to obtain the good with probability 1.

Under the calibrated information structure, the optimal mechanism is neither incentive compatible nor individually rational. When the good is $R$, the buyer would prefer to choose $(1,1)$ regardless of her type. Instead, when the good is $L$, the buyer of $\theta_{1}$ would quit the mechanism instead of paying 1 and getting nothing.

To characterize the optimal calibrated mechanism, we rely again on two-stage mechanisms. Equate $\mu$ with the probability that the good is $R$. Note that because states and types are independent, if the seller assigns probability $\mu$ to the state being $R$, so does the buyer (and vice versa). For each $\mu \in[0,1]$, the seller solves the following problem

$$
\begin{aligned}
W(\mu) & \equiv \max _{(q, t): \Theta \rightarrow[0,1] \times \mathbb{R}} \sum_{\theta \in \Theta} \frac{1}{3} t(\theta) \\
& \text { s.t. }\left\{\begin{array}{ll}
\left(\forall \theta \in\left\{\theta_{1}, \theta_{2}, \theta_{3}\right\}\right) & q(\theta) \mathbb{E}_{\mu} v(\theta, \cdot)-t(\theta) \geq 0 \\
\left(\forall \theta, \theta^{\prime} \in\left\{\theta_{1}, \theta_{2}, \theta_{3}\right\}, \theta^{\prime} \neq \theta\right) & q(\theta) \mathbb{E}_{\mu} v(\theta, \cdot)-t(\theta) \geq q\left(\theta^{\prime}\right) \mathbb{E}_{\mu} v(\theta, \cdot)-t\left(\theta^{\prime}\right)
\end{array} .\right.
\end{aligned}
$$

In this case, the seller's objective function does not depend on the induced belief $\mu$ as types and states are independent. Instead, the buyer's incentive and individual rationality constraints do depend on $\mu$ as the state is payoff relevant. The solution to the problem in Equation 7 is a posted price, whose value depends on $\mu$. For instance, when $\mu \in\{0,1\}$, the optimal price is 2 and the seller's revenue is 4/3. Instead, when $\mu=2 / 3$, the optimal price is 5/3 and profits are maximal and equal to 5/3. Indeed, when $\mu=2 / 3$, the heterogeneity across buyer types is minimized (and hence, their rents), and by setting $p=5 / 3$ all buyer types buy. The blue line in Figure 2 depicts the seller's expected profit as a function of his belief $\mu$.

> [Figure omitted from this text-only corpus; refer to the source manuscript.]
Figure 2: Seller's profit in the two-stage mechanism

The optimal calibrated mechanism can be read from the concavification of $W$ at $\mu_{0}=1 / 2$, depicted by the dashed red line in Figure 2. The seller provides the buyer with partial information about the good: He either reveals the good is $L$ and sells the good at a price of 2 , or he obfuscates the good-inducing a belief of 2/3-and sets a price of5/3.

Another consequence of Theorem 2 is that without loss of generality, we can focus on calibrated mechanisms with finite calibrated information structures:

Corollary 1 (Support of calibrated information structures). It is without loss of generality to restrict attention to two-stage mechanisms that induce at most $|\Omega|$ beliefs.

In other words, it is without loss of generality to focus on calibrated mechanisms that induce at most $|\Omega|$ allocation rules.

Multiple agents and generalized two-stage mechanisms In the case of multiple agents, we can also interpret a calibrated mechanism as conveying to each agent $i$ both the information she should have about the state upon seeing signal $s_{i}^{*}, \mu_{i}\left(\theta_{i}, s_{i}^{*}\right)$, and her interim allocation rule, $s_{i}^{*}: \Theta_{i} \rightarrow \Delta\left(A_{i}\right)$. However, two differences arise relative to the single-agent case: First, each agent $i$ receives her information privately from that of other agents. Second, even if the agents put together the information they receive, this is not enough to learn the ex-post allocation rule, that is, the map from type profiles to allocations. After all, each agent $i$ observes her interim allocation rule alone. These differences are natural when we think of calibrated mechanisms as capturing the information agents stand to learn from experimenting with the mechanism: There is no reason all agents will learn the same information, and from observing her own allocations, and not those of others, an agent can only learn about her interim allocation rule, not the ex-post one.

These observations together imply that to describe the analogue of a two-stage mechanism in multiagent settings we need to (i) allow for agent-by-agent information disclosure, and (ii) keep track that the interim allocation rules are consistent with the same ex-post allocation rule. These considerations motivate the following generalization of a two-stage mechanism:

Definition 4 (Generalized two-stage mechanism). A generalized two-stage mechanism is a mapping $\psi: \Theta \times \Omega \rightarrow \Delta\left(\Delta(\Omega)^{N} \times A\right)$ for which a tuple of mappings

$$
\beta: \Omega \rightarrow \Delta\left(\Delta(\Omega)^{N}\right), \quad \alpha_{i}: \Theta_{i} \times \Delta(\Omega) \rightarrow \Delta\left(A_{i}\right), \quad \alpha: \Theta \times \Omega \times \Delta(\Omega)^{N} \rightarrow \Delta(A),
$$

exist such that:

1. For all $(\theta, \omega) \in \Theta \times \Omega$, and all measurable subsets $\left(\tilde{\Delta}_{i}\right)_{i=1}^{N} \subset \Delta(\Omega)^{N}$, we have
$$
\psi\left(\times_{i=1}^{N} \tilde{\Delta}_{i} \times\{a\} \mid \theta, \omega\right)=\int_{\times_{i=1}^{N} \tilde{\Delta}_{i}} \alpha\left(a \mid \theta, \omega, \mu_{1}, \ldots, \mu_{N}\right) \beta\left(d\left(\mu_{1}, \ldots, \mu_{N}\right) \mid \omega\right)
$$
2. The Blackwell experiment $\beta$ is Bayes plausible,
3. For all $i \in\{1, \ldots, N\}$, the interim allocation rule $\alpha_{i}$ satisfies that for all measurable subsets $\tilde{\Delta}$ of $\Delta(\Omega)$ and all $\left(a_{i}, \theta_{i}, \omega\right) \in A_{i} \times \Theta_{i} \times \Omega$
$$
\int_{\tilde{\Delta} \times \Delta(\Omega)^{N-1}}\left\{\alpha_{i}\left(a_{i} \mid \theta_{i}, \mu_{i}\right)-\mathbb{E}_{f_{-i}(\cdot \mid \omega)}\left[\sum_{a_{-i} \in A_{-i}} \alpha\left(a_{i}, a_{-i} \mid \theta_{i}, \theta_{-i}, \omega, \mu_{i}, \mu_{-i}\right)\right]\right\} \beta\left(d\left(\mu_{i}, \mu_{-i}\right) \mid \omega\right)=0 .
$$
We say the generalized two-stage mechanism is incentive compatible and individually rational if for all $i \in\{1, \ldots, N\}$, on the support of $\mu_{0} \otimes \beta, \alpha_{i}\left(\cdot \mid \cdot, \mu_{i}\right)$ is incentive compatible and individually rational for agent $i$ when she learns $\mu_{i}$.

As anticipated, generalized two-stage mechanisms differ from two-stage mechanisms in three ways when $N>1$. First, because disclosures are private, the experiment $\beta$ now outputs a profile of beliefs, one for each agent. As shown in Arieli et al. (2024), $\beta$ is Bayes plausible if and only if for each agent $i$, the marginal Blackwell experiment $\beta_{i}$ is Bayes plausible. Second, while the individual interim allocation rule $\alpha_{i}$ only depends on the disclosed belief to agent $i, \mu_{i}$, and not the state, the ex-post
allocation rule $\alpha$ may depend on the state, even conditional on the belief profile $\left(\mu_{1}, \ldots, \mu_{N}\right)$. The reason is that this belief profile is no longer a sufficient statistic for the ex-post allocation rule as each agent $i$ only observes their interim allocation. Third and relatedly, we need to keep track of both the interim allocation rules $\left(\alpha_{i}\right)_{i=1}^{N}$ and the ex-post allocation rule $\alpha$ to check that the interim allocation rules are consistent with the same mechanism. An interesting question for future work would be to characterize which interim allocation rules $\left(\alpha_{i}\right)_{i=1}^{N}$ are consistent with some ex-post allocation rule $\alpha$, so that one could focus on the interim allocation rules alone.

As we show in Proposition 1, a calibrated mechanism induces a generalized two-stage mechanism:
Proposition 1. If outcome distribution $\vartheta \in \Delta(A \times \Theta \times \Omega)$ is implementable by an incentive compatible and individually rational calibrated mechanism, then it is implementable by an incentive compatible and individually rational generalized two-stage mechanism.

In contrast to the single-agent case, not every outcome distribution implemented by a generalized two-stage mechanism can be implemented by a calibrated mechanism. On the one hand, no agent's beliefs are a sufficient statistic for the information the mechanism leaks about the state, so that the allocation rule $\bar{\alpha}$ may still leak information about the state or others' beliefs, which in turn leak information about the state. On the other hand, because in a calibrated mechanism each agent learns her interim allocation rule conditional on $(\omega, \varepsilon)$, the incentive and participation constraints associated to a generalized two-stage mechanism are weaker than those implied by a calibrated mechanism whenever multiple interim allocation rules underlie the same belief: Even if the average interim allocation rule $\alpha_{i}$ is incentive compatible and individually rational, each of the interim allocation rules underlying that average need not be.

## 4 Applications

In this section, we study optimal calibrated mechanism design in canonical mechanism design settings with quasilinear utilities. We first consider the case of a single agent, with single-dimensional types and allocations, and supermodular payoffs. In Section 4.1, we show that if the order of types is state independent, then optimal two-stage mechanisms fully reveal the state, whereas this conclusion can be reversed when the order of types is state-dependent. In Section 4.2, we compare optimal calibrated mechanism design against the Myersonian benchmark. Lastly, we analyze a multi-agent application in Section 4.3.

### 4.1 Calibrated Screening

We consider the following version of the model in Section 2. Suppose $N=1$ and let $\Theta=[\underline{\theta}, \bar{\theta}]$ denote the set of types. Assume $\theta$ is distributed according to a full support distribution $F$ with density $f$. Hence, throughout, we consider the case in which the agent's type is independent of $\omega$. Denote the set of allocations by $A=[0, \bar{q}] \times \mathbb{R}$, where $q \in[0, \bar{q}]$ is the (physical) allocation and $t \in \mathbb{R}$ is a payment from the agent to the designer. ${ }^{21}$

[^12]The agent's and the designer's payoffs are given by $u(q, \theta, \omega)-t$ and $w(q, \theta, \omega)+t$, respectively. Assume that if the agent does not participate, then the outside option is $a_{\varnothing}=(0,0)$, and that this yields a payoff of 0 to both the designer and the agent. Throughout, we assume that for each $\omega \in \Omega$, the family of functions $\{\theta \mapsto u(q, \theta, \omega): q \in[0, \bar{q}]\}$ is equi-Lipschitz on $\Theta$ : a positive constant $L_{\omega}$ exists such that for all $\theta, \theta^{\prime} \in \Theta$ and $q \in[0, \bar{q}],\left|u(q, \theta, \omega)-u\left(q, \theta^{\prime}, \omega\right)\right| \leq L_{\omega}\left|\theta-\theta^{\prime}\right| .{ }^{22}$ Furthermore, the analysis that follows restricts attention to mechanisms that do not randomize on the allocation (beyond the inherent randomness of $\Omega \times[0,1]$ ). Remark 1 at the end of this section discusses settings in which this is not a restriction and how to generalize the observations herein when random allocations are allowed.

Our goal is to characterize the designer optimal calibrated mechanism and how its properties depend on how the state affects the order of types.

State-independent type ranking We consider first the case in which the order of types is independent of the state. Formally, assume that for all $\omega \in \Omega$, the function $u(\cdot, \omega)$ is supermodular in $(q, \theta)$. That is, in all states, the agent with higher value of $\theta$ values $q$ more. These assumptions are satisfied, for instance, for $u(q, \theta, \omega)=\theta \omega q$ or $u(q, \theta, \omega)=(\theta+\omega) q$.

By Theorem 2, we can characterize the optimal calibrated mechanism via two-stage mechanisms. To do so, we solve the problem "backward": For each $\mu \in \Delta(\Omega)$ the designer may induce about the state, the designer chooses an optimal direct mechanism $\left(q_{\mu}, t_{\mu}\right): \Theta \rightarrow A$. This determines the designer's value function $W: \Delta(\Omega) \rightarrow \mathbb{R}$. We obtain the designer's optimal Blackwell experiment by studying the properties of $W$.

Given belief $\mu$, define the agent's and the designer's (expected) payoff at $(q, t, \theta)$ as follows:

$$
u(q, \theta \mid \mu) \equiv \sum_{\omega \in \Omega} \mu(\omega) u(q, \theta, \omega), w(q, \theta \mid \mu) \equiv \sum_{\omega \in \Omega} \mu(\omega) w(q, \theta, \omega)
$$

Thus, conditional on inducing belief $\mu$, the designer's problem can be written as follows:

$$
\begin{aligned}
W(\mu) \equiv \max _{(q, t): \Theta \rightarrow A} \int_{\Theta}[w(q(\theta), \theta, \mu)+t(\theta)] F(d \theta) \\
\text { s.t. }\left\{\begin{array}{ll}
(\forall \theta \in \Theta) & u(q(\theta), \theta \mid \mu)-t(\theta) \geq 0 \\
\left(\forall \theta, \theta^{\prime} \in \Theta\right) & u(q(\theta), \theta \mid \mu)-t(\theta) \geq u\left(q\left(\theta^{\prime}\right), \theta \mid \mu\right)-t\left(\theta^{\prime}\right)
\end{array} .\right.
\end{aligned}
$$

Our assumptions imply that $u(\cdot \mid \mu)$ is supermodular in $(q, \theta)$. It follows that the designer can only choose among those $q: \Theta \rightarrow[0, \bar{q}]$ that are (weakly) increasing in $\theta$. Let $Q_{\uparrow}$ denote the set of all such $q(\cdot)$. Furthermore, at the optimum, the participation constraint of $\theta=\underline{\theta}$ binds.

Define the virtual surplus at $(q, \theta, \omega)$ as follows:

$$
J((q, \theta, \omega) ; F)=w(q, \theta, \omega)+u(q, \theta, \omega)-u_{2}(q, \theta, \omega) \frac{1-F(\theta)}{f(\theta)},
$$

where $u_{2}$ is the derivative of $u$ against its second coordinate; the equi-Lipschitz assumption implies it exists almost everywhere. Then, conditional on inducing belief $\mu$, the designer's payoff can be written

[^13]as follows:
$$
W(\mu)=\max _{q \in Q_{\uparrow}} \int_{\Theta} \mathbb{E}_{\mu}[J((q(\theta), \theta, \omega) ; F)] F(d \theta) .
$$
Note the objective is linear in $\mu$ and the constraint set is independent of $\mu$. We conclude that $W$ is convex, as it is the maximum of linear functionals in $\mu$. It follows that full disclosure is an optimal experiment for the designer. Equivalently, an optimal calibrated mechanism exists in which the designer chooses the mechanism $\phi_{\text {full }}^{D}$, where for all $\omega \in \Omega$, $\phi_{\text {full }}^{D}(\cdot, \omega, \cdot): \Theta \times[0,1] \rightarrow A$ is the optimal deterministic mechanism when it is common knowledge that the state is $\omega$.

Proposition 2 summarizes the above discussion:
Proposition 2 (State-Independent Type Ranking). In a single-dimensional screening problem with state-independent type ranking, the designer can do no better than choosing $\phi_{\text {ful }}^{D}$ among deterministic mechanisms.

By Proposition 2, in screening problems with state-independent ranking of types across states, the calibration constraint makes any pooling of mechanisms across states unprofitable. ${ }^{23}$ Remarkably, this result holds for any designer objective, such as profit, revenue, or efficiency. It also requires no regularity assumptions on the type distribution, as we do not obtain the result by looking at the relaxed problem. Instead, our argument relies on the restriction to deterministic mechanisms (conditional on the induced belief), which in turn delivers that the set of implementable allocations does not depend on the induced belief. Remark 1 discusses conditions under which (i) the restriction to deterministic mechanisms is without loss of optimality, and (ii) the set of implementable allocations does not depend on the induced belief, even when randomized mechanisms are allowed. Readers interested in the case of state-dependent ranking can skip this remark with little loss of continuity.

Remark 1 (Proposition 2 without deterministic mechanisms). Under our assumptions, deterministic mechanisms are without loss of optimality if the agent's payoff is linear in $q$ and the designer's payoff is concave in $q$. (See Section 4.2 for yet another condition.) However, the driving force behind Proposition 2 is that the designer's constraint set does not depend on the induced belief. The state-bystate supermodularity assumption and the restriction to deterministic mechanisms is one way to ensure this is the case. We now discuss two other cases in which the designer's constraint set does not depend on the induced belief and thus Proposition 2 holds for the optimal (not necessarily deterministic) calibrated mechanism.

First, suppose the agent's payoff is linear in $q$, so that $u(q, \theta, \omega)=q v(\theta, \omega)$, where $v(\cdot, \omega)$ is increasing for all $\omega$. Then, the set of implementable lotteries over $q$ when the belief is $\mu$ is given by:

$$
Q_{\uparrow, \text { random }}=\left\{\xi: \Theta \rightarrow \Delta([0, \bar{q}]): \mathbb{E}_{\xi(\theta)}[q] \text { is increasing in } \theta\right\} .
$$

In this case, we obtain that the designer cannot do any better than choosing $\phi_{\text {ful }}$, which is the

[^14]mechanism that implements in each state $\omega$ the optimal mechanism under common knowledge that the state is $\omega$. This relates to the results in Szabadi (2018) and Yamashita (2018), who study the optimal mechanism design preceded by public information disclosure. Both papers consider settings in which the agent's payoff is linear in $q$ and obtain that full disclosure is optimal when the ranking of types is independent of the state. ${ }^{24}$

Second, suppose the agent's payoff has the form

$$
u(q, \theta, \omega)=b(\theta) c(\omega) v(q)+k_{1}(q, \omega)+k_{2}(\theta, \omega),
$$

where $b$ is increasing in $\theta$ and $c(\cdot)$ does not change sign on $\Omega$. Under this assumption, $u(q, \theta \mid \mu)$ satisfies monotonic expectational differences for all $\mu \in \Delta(\Omega)$ (see, e.g., Kartik et al., 2024). Consequently, one can define a linear order $\succeq$ over $\Delta([0, \bar{q}])$ as follows: $\xi \succeq \xi^{\prime}$ if $u(\xi, \theta \mid \mu)-u\left(\xi^{\prime}, \theta \mid \mu\right)$ is increasing in $\theta$, where $u(\xi, \theta \mid \mu)$ is the linear extension of $u(\cdot, \theta \mid \mu)$ to $\Delta([0, \bar{q}])$. This linear order implies the ranking of types is state independent. Indeed, the analog of $Q_{\uparrow, \text { random }}$ is the set of all $\xi: \Theta \rightarrow \Delta([0, \bar{q}])$ such that $\theta \geq \theta^{\prime}$ implies $\xi(\theta) \succeq \xi\left(\theta^{\prime}\right)$, which is again independent of the induced belief.

State-dependent type ranking Example 2 illustrates that when the ranking of types is not uniform across states, full transparency may not be optimal. ${ }^{25}$ We now provide a more systematic analysis of this phenomenon, using the previous results. To provide the starkest contrast with Proposition 2, we consider a setting that shares a key feature of Example 2: we can partition $\Delta(\Omega)$ into two regions such that within each region the ranking of types-as determined by $u(q, \theta \mid \mu)$-is the same, but it differs across regions.

Concretely, suppose that $\Omega=\left\{\omega_{1}, \omega_{2}\right\}$. Furthermore, assume

$$
u(q, \theta, \omega)=\left\{\begin{array}{ll}
q \theta & \text { if } \omega=\omega_{2} \\
q(c-b \theta) & \text { otherwise, }
\end{array},\right.
$$

where $c \in \mathbb{R}$, and $b>0$. Identify beliefs with the probability that the state is $\omega_{2}$ and define

$$
\hat{\mu}=\frac{b}{1+b} .
$$

For $\mu<\hat{\mu}$, we have that $u(q, \theta \mid \mu)$ is decreasing in $\theta$, whereas if $\mu>\hat{\mu}$, then $u(q, \theta \mid \mu)$ is increasing in $\theta$.
Consider now the designer's optimal payoff $W: \Delta(\Omega) \rightarrow \mathbb{R}$ as a function of the different beliefs he may induce. When $\mu \geq \hat{\mu}$, the designer's payoff can be obtained by solving the program in Equation 9 as before. Instead, when $\mu<\hat{\mu}$, the designer's payoff can be obtained by solving a problem analogous to Equation 9, but where the space of implementable allocations is the set of decreasing $q, Q_{\downarrow}$, and the participation constraint of $\bar{\theta}$ binds. It follows that $W$ is convex on $[0, \hat{\mu})$ and $(\hat{\mu}, 1]$. Thus, the support of designer's optimal experiment is included in \{0, $\hat{\mu}, 1\}$.

[^15]Proposition 3 (State-dependent type ranking). Suppose the agent's payoff satisfies the assumptions above. If $(1-\hat{\mu}) W(0)+\hat{\mu} W(1) \geq W(\hat{\mu})$, full transparency is optimal. Otherwise, full transparency is not optimal: if $\mu_{0}<\hat{\mu}$, it is optimal to split $\mu_{0}$ to 0 and $\hat{\mu}$; if $\mu_{0}>\hat{\mu}$, it is optimal to split $\mu_{0}$ to $\hat{\mu}$ and 1 . In particular, if $W(\hat{\mu})>\max \{W(0), W(1)\}$, then full transparency is not optimal.

By Proposition 3, whether full transparency is optimal depends on the designer and agent's payoffs and the type distribution, but only through their impact on the value the function $W$ takes at points $\{0, \hat{\mu}, 1\}$. At $\hat{\mu}$, the agent earns no rents-as $u(\cdot \mid \hat{\mu})$ is constant across types-which pushes against full transparency. At the same time, efficiency may dictate the designer to condition the allocation rule on the state, which favors information disclosure. The piecewise convexity of $W$ implies that if $W(\hat{\mu})$ dominates $W$ at the extreme beliefs, the rent extraction motive dominates and the designer does not engage in full disclosure.

### 4.2 Comparison with Myersonian Mechanism Design

We now compare optimal calibrated mechanism design and the Myersonian benchmark. In the Myersonian benchmark, the designer is not concerned with the information the mechanism reveals about the state, and hence provides a natural upper bound on the designer's payoffs in calibrated mechanism design. The gap between the designer's optimal payoff across both benchmarks quantifies the loss from the calibration constraint. If no gap exists, the calibration constraint is non-binding and an optimal calibrated mechanism can be found solving the Myersonian benchmark. Instead, if a gap exists, the optimal mechanism in the Myersonian benchmark reveals information about the state in a way that it fails to be incentive compatible or individually rational under calibration.

Myersonian benchmark In the Myersonian benchmark, the designer chooses a direct mechanism $(\xi, t): \Theta \times \Omega \rightarrow \Delta([0, \bar{q}]) \times \mathbb{R}$ subject to incentive and participation constraints that must hold on average across states under the prior $\mu_{0} .{ }^{26}$ Formally,

$$
\begin{aligned}
& W_{\mathrm{My}} \equiv \max _{(q, t): \Theta \times \Omega \rightarrow[0, \bar{q}] \times \mathbb{R}} \int_{\Theta} \mathbb{E}_{\mu_{0}}[w(\xi(\theta, \omega), \theta, \omega)+t(\theta, \omega)] F(d \theta) \\
& \text { s.t. }\left\{\begin{array}{l}
(\forall \theta \in \Theta) \mathbb{E}_{\mu_{0}}[u(\xi(\theta, \omega), \theta, \omega)-t(\theta, \omega)] \geq 0 \\
\left(\forall \theta, \theta^{\prime} \in \Theta\right) \mathbb{E}_{\mu_{0}}[u(\xi(\theta, \omega), \theta, \omega)-t(\theta, \omega)] \geq \mathbb{E}_{\mu_{0}}\left[u\left(\xi\left(\theta^{\prime}, \omega\right), \theta, \omega\right)-t\left(\theta^{\prime}, \omega\right)\right]
\end{array},\right.
\end{aligned}
$$

where $w(\xi, \theta, \omega)$ and $u(\xi, \theta, \omega)$ are the linear extensions of $w(\cdot, \theta, \omega)$ and $u(\cdot, \theta, \omega)$, respectively.
Program $\mathrm{OPT}_{\mathrm{My}}$ is a mechanism design problem with a multidimensional allocation, corresponding to assigning (a distribution over) $q$ in each state. As a result, the distinction between the Myersonian benchmark and optimal calibrated design shows in the monotonicity requirements the allocation $\xi(\theta, \omega)$ must satisfy for a transfer $t: \Theta \times \Omega \rightarrow \mathbb{R}$ to exist that implements $\xi(\theta, \omega)$. Indeed, implementability of $\xi: \Theta \times \Omega \rightarrow \Delta([0, \bar{q}])$ is equivalent to integral monotonicity (Rochet, 1987; Pavan et al., 2014): ${ }^{27}$

$$
\left(\forall \theta, \theta^{\prime} \in \Theta\right) \int_{\theta^{\prime}}^{\theta} \int_{\Omega}\left[u_{2}(\xi(s, \omega), s, \omega)-u_{2}\left(\xi\left(\theta^{\prime}, \omega\right), s, \omega\right)\right] d \mu_{0} d s \geq 0
$$

[^16]where recall $u_{2}$ is the derivative of $u$ in its second coordinate.

Comparison with calibrated mechanism design To facilitate the comparison with Proposition 2, we focus on deterministic mechanisms $(q, t): \Theta \times \Omega \rightarrow[0, \bar{q}] \times \mathbb{R}$. Remarkably, even if the agent's payoff net of transfers, $u(q, \theta, \omega)$, is supermodular in $(q, \theta)$ for all $\omega \in \Omega$, the characterization of the set of implementable $q(\cdot)$ cannot be simplified beyond integral monotonicity without further assumptions. Because integral monotonicity is a global, implicitly defined constraint, verifying implementability and computing the optimal mechanism is more computationally involved in the Myersonian benchmark than in calibrated mechanism design. Indeed, Proposition 2 implies the optimal deterministic calibrated mechanism coincides with the state-by-state optimal deterministic mechanism under this assumptions. In other words, the optimal deterministic calibrated mechanism can be obtained by selecting allocations $q(\cdot)$ that satisfy

$$
Q_{\mathrm{cal}}=\{q: \Theta \times \Omega \rightarrow[0, \bar{q}]:(\forall \omega \in \Omega) q(\cdot, \omega) \text { is increasing }\} .
$$

That is, the allocation in the optimal calibrated mechanism must satisfy monotonicity state-by-state. Instead, the optimal deterministic Myersonian mechanism can be obtained by selecting allocations $q(\cdot)$ that satisfy Equation IM, which we denote by $Q_{\mathrm{My}}$.

When $u(q, \theta, \omega)$ is supermodular in $(q, \theta)$ for all $\omega \in \Omega$, the above discussion implies the designer's optimal payoff in the Myersonian and calibration settings can be written as follows:

$$
\begin{aligned}
W_{\mathrm{My}}^{D} & =\max _{q \in Q_{\mathrm{My}}} \int_{\Theta} \mathbb{E}_{\mu_{0}}[J(q(\theta, \omega), \theta, \omega ; F)] F(d \theta), \\
W_{\mathrm{cal}}^{D} & =\max _{q \in Q_{\mathrm{cal}}} \int_{\Theta} \mathbb{E}_{\mu_{0}}[J(q(\theta, \omega), \theta, \omega ; F)] F(d \theta),
\end{aligned}
$$

where the superscript $D$ in the objective is a reminder that we restrict attention to mechanisms that are deterministic conditional on the state, or the induced belief.

By reducing the comparison across settings to monotonicity requirements on the space of allocations, the above expressions provide us with an immediate way of comparing the designer's payoffs across settings. In particular, when the optimal Myersonian mechanism satisfies the state-by-state monotonicity constraints, we have that the calibration constraint entails no loss to the designer. We record this observation for future use:

Observation 1. Suppose $u(q, \theta, \omega)$ is supermodular in $(q, \theta)$ for all $\omega \in \Omega$. Then, if the allocation rule in the Myersonian benchmark satisfies monotonicity state-by-state, $W_{\text {cal }}^{D}=W_{M y}^{D}$.

Two natural questions are under what conditions the solution to OPT $_{\text {My }}$ is deterministic and satisfies state-by-state monotonicity. We answer them simultaneously by studying the relaxed program. Inspection of Equation 10 reveals that if the virtual surplus is supermodular in $(q, \theta)$ for every $\omega$, then the solution $q_{\text {rel }}$ to the relaxed problem

$$
W_{\mathrm{rel}}=\max _{q: \Theta \times \Omega \rightarrow[0, \bar{q}]} \int_{\Theta} \mathbb{E}_{\mu_{0}}[J(q(\theta, \omega), \theta, \omega ; F)] F(d \theta),
$$

satisfies monotonicity state-by-state by Topkis' theorem. Moreover, a stochastic mechanism is
equivalent to a deterministic mechanism which depends on the random reports of a fictitious agent (Pavan et al., 2014). The virtual surplus in this fictitious setting coincides with that in the integrand on the right-hand side of Equation 11-the type reports of the fictitious agent are payoff irrelevant-and is maximized by $q_{\text {rel }}$.

Proposition 4 (Sufficient condition for no gap). Suppose the virtual surplus $J((q, \theta, \omega) ; F)$ is supermodular in $(q, \theta)$ for all $\omega$. Then, the designer's payoffs under the optimal Myersonian and calibrated mechanisms coincide.

By contrast to Proposition 2, Proposition 4 relies on assumptions on the type distribution and the designer's payoff. As Example 2 illustrates, the supermodularity of the virtual surplus can fail when the type distribution is not regular, creating a gap between the designer's payoff at the optimal Myersonian and calibrated mechanisms. Example 3 illustrates such a gap can also arise when the designer's payoff is not supermodular:

Example 3 (Payoff gap when $w$ is not supermodular). Suppose states are binary, $\Omega=\left\{\omega_{L}, \omega_{H}\right\}=\{1,3\}$, and equally likely. Suppose types are uniformly distributed, $\theta \sim U[0,1]$. Finally, let $q \in[0,1]$ denote the probability the seller's good is allocated. Payoffs are given by:

$$
\begin{aligned}
u(q, \theta, \omega) & =q \theta \omega \\
w(q, \theta, \omega) & =2(1-2 \theta) q .
\end{aligned}
$$

Note that $w$ is increasing in $q$ when $\theta<1 / 2$ and decreasing in $q$ when $\theta>1 / 2 .{ }^{28}$ In this case, the virtual surplus evaluated at different states is:

$$
J((q, \theta, \omega) ; F)=\left\{\begin{array}{ll}
(1-2 \theta) q & \text { if } \omega=\omega_{L} \\
(2 \theta-1) q & \text { otherwise }
\end{array} .\right.
$$

In the Myersonian benchmark, implementable allocations are elements of $Q_{M y}$, which in this case is equivalent to requiring that $\mathbb{E}_{\mu_{0}}[q(\cdot, \omega) \omega]$ is increasing. The optimal Myersonian allocation obtains from pointwise maximizing the virtual surplus, and is given by:

$$
q_{M y}(\theta, \omega)=\left\{\begin{array}{ll}
1 & \text { if } \omega=\omega_{L} \text { and } \theta<1 / 2 \\
1 & \text { if } \omega=\omega_{H} \text { and } \theta>1 / 2 \\
0 & \text { otherwise }
\end{array} .\right.
$$

The designer's payoff under the Myersonian mechanism is 1/4.
By Proposition 2, $q_{M y}$ cannot be implemented by a calibrated mechanism as it is not increasing state-bystate. Intuitively, when $\omega=\omega_{L}$, types above 1/2 would learn from the calibrated information structure that they do not obtain the good, whereas types below 1/2 do, and would misreport their types.

Instead, in the optimal calibrated mechanism, the designer sets $q_{\text {cal }}\left(\theta, \omega_{H}\right)=\mathbb{1}[\theta \geq 1 / 2]$ and sets

[^17]$q\left(\theta, \omega_{L}\right)$ to be constant in $\theta$. The designer's payoff under calibration is $W_{\text {cal }}=1 / 8<W_{M y}$.

### 4.3 Optimal Calibrated Auction

In this section, we consider a multiple agent application and study the design of the optimal calibrated auction. Proposition 1 implies the optimal calibrated auction induces a generalized two-stage mechanism, and hence the optimal generalized two-stage mechanism provides an upper bound on the designer's optimal payoff under calibration. However, computing the optimal generalized two-stage mechanism is complicated because (i) no tractable characterization of joint distributions over posterior beliefs is available, and (ii) the allocation rule may condition on the state and not only the agents' beliefs. For that reason, our analysis below relies on Observation 1: We show the optimal Myersonian auction can be implemented by fully revealing the state, and hence, remains incentive compatible and individually rational when the agents have access to the calibrated information structure. Below, we first specialize our multi-agent model and notation to the auction application and then link our assumptions to online advertising.

Suppose there is a single good for sale and the state is multidimensional, $\omega=\left(\omega_{i}, \omega_{0 i}\right)_{i \in[N]} \in \mathbb{R}_{+}^{2 N}$, and distributed according to prior distribution $\mu_{0}$. Suppose that for all $i \in[N], \Theta_{i}=[0,1]$, with $\theta_{i} \sim F_{i}$ with full-support density $f_{i}$. That is, we are assuming agents' types are independent of the state, and hence, independent across each other. Denote by $q_{i} \in[0,1]$ the probability agent $i$ is allocated the good, and note that feasibility implies that $0 \leq \sum_{i=1}^{N} q_{i} \leq 1$.

We assume the agents' and the designer's utilities are quasilinear in transfers. Agent $i$ 's payoff net of transfers is $u_{i}\left(q_{i}, \theta_{i}, \omega\right)=q_{i}\left(\omega_{i} \theta_{i}+\omega_{0 i}\right)$. Thus, state components $\omega_{i}$ capture the value responsiveness to agent's private information, whereas state components $\omega_{0 i}$ capture the overall shift. The state components can be correlated (and asymmetric) across agents, allowing for interdependent values. The designer's payoff net of transfers is $w(q, \theta, \omega)=\sum_{i} q_{i} w_{i}(\theta, \omega)$ for some functions $\left(w_{i}\right)_{i \in[N]}$. Below, we study the designer-optimal calibrated mechanism.

To fix ideas, consider the following mapping to an online advertising environment. The designer is an advertising platform, and the good is an advertising slot on a given webpage targeted to a selected category of users in a given week. Agents are firms that wish to display their ads, and their private types represent the expected revenue from a click on their ad. State components $\omega_{i}$ could capture individual click-through rates or match values, while state components $\omega_{0 i}$ could capture individual display values, that is, the expected revenue from an ad being displayed irrespective of whether it is clicked (for instance, due to brand-building effects). The state is observed through proprietary data available to the platform and can be used in the design of the auction. The platform values the resulting revenue but may also have additional efficiency considerations, summarized by $w_{i}$.

As anticipated, we characterize the optimal calibrated mechanism by showing that it coincides with the Myersonian optimal one. To this end, consider the Myersonian problem, in which the designer chooses $(q(\theta, \omega), t(\theta, \omega)) \in[0,1]^{N} \times \mathbb{R}^{N}$. Because agent $i$ 's payoff is linear in $\theta_{i}$, arguments analogous to those in Section 4.2 imply a feasible $q(\theta, \omega)$ is implementable if and only if for all $i, \mathbb{E}_{F_{-i}, \mu_{0}}\left[q_{i}\left(\theta_{i}, \theta_{-i}, \omega\right) \omega_{i}\right]$ is increasing in $\theta_{i}$. In a slight abuse of notation, denote by $Q_{\mathrm{My}}$ the set of all such functions and define
the virtual surplus as

$$
J((q, \theta, \omega) ; F)=\sum_{i=1}^{N} q_{i}(\theta, \omega)\left(w_{i}(\theta, \omega)+\left(\theta_{i}-\frac{1-F_{i}\left(\theta_{i}\right)}{f_{i}\left(\theta_{i}\right)}\right) \omega_{i}+\omega_{0 i}\right) .
$$

Standard arguments imply the individual rationality constraint of $\theta_{i}=0$ binds for all $i$, and an optimal mechanism solves

$$
W_{\mathrm{My}}=\max _{q \in Q_{\mathrm{My}}} \int_{[0,1]^{N}} \mathbb{E}_{\mu_{0}}[J(q(\theta, \omega), \theta, \omega ; F)] f(\theta) d \theta .
$$

Proposition 5 (No gap in regular auctions). Suppose that (i) for all $i \in\{1, \ldots, N\}, F_{i}$ is Myerson regular, and for all $i, j, \theta$, and $\omega, w_{i \theta_{i}}(\theta, \omega) \geq 0, w_{i \theta_{i}}(\theta, \omega) \geq w_{j \theta_{i}}(\theta, \omega) .{ }^{29}$ Then, $W_{\text {cal }}=W_{M y}$.

The proof of Proposition 5 in Appendix D. 2 shows that under our assumptions the optimal Myersonian mechanism can be obtained by solving the relaxed program. Importantly, the assumption that $w_{i \theta_{i}}(\theta, \omega) \geq w_{j \theta_{i}}(\theta, \omega)$ ensures that an increase in agent $i$ 's type increases the designer's payoff of giving the object to agent $i$ by more than the value of giving it to other agents. This, in turn, ensures agent $i$ 's allocation probability is increasing in her type.

Viewed through the lens of the online advertising example, Proposition 5 implies that in regular environments, while the advertising platform benefits from having the data on click-through rates and display values, it does not benefit from the informational advantage over bidders that such data entails. Its objective is maximized by making the click-through rates and display values readily available to bidders and running optimal auctions in all instances.

## 5 Microfoundation

In this section, we provide a microfoundation for calibrated mechanism design by analyzing the outcome distributions that can arise when an agent repeatedly engages with the same mechanism (Section 5.1) and contrast this to what can be implemented when the designer can offer the agent a fully dynamic mechanism (Section 5.2). To keep the presentation simple, we present the results with minimal notation, and refer the reader to Appendix C for details.

Throughout, we consider the case of a single agent, whose type (i) is redrawn each period from the same distribution and (ii) is independent of the state. The reason for (i) is as follows. When the designer offers the agent a fully dynamic mechanism, the revelation principle implies that it is without loss of generality for the designer to ask the agent for type reports. Moreover, logic similar to that in Myerson (1986) implies that the designer only elicits one type report when the agent's type is persistent, and hence, the agent has no possibility of experimenting with the mechanism. Hence, to put repeated and dynamic mechanisms on a more similar footing, assuming the agent's type is redrawn each period is necessary. However, when the agent's type is repeatedly drawn from a distribution that depends on the state, the agent learns about the state both through her own type and her allocations in the mechanism. ${ }^{30}$ Thus, we assume (ii) so that the agent learns about the state only through her

[^18]interaction with the mechanism. Lastly, we consider the single-agent case as extending the results in this section to multiple agents requires addressing subtle issues in strategic experimentation, which we plan to pursue in future work.

### 5.1 Repeated Interactions with a Mechanism

We consider first the case in which the agent interacts repeatedly with the same mechanism $\phi$ in each period of an infinite horizon interaction. In line with Section 2, a repeated mechanism is a mapping

$$
\phi: M \times \Omega \times \mathcal{E} \rightarrow \Delta(A),
$$

where $M$ is a finite set of messages and $\mathcal{E}$ is a finite set endowed with some measure, denoted $\eta$. The results in Section 3 imply that assuming $\mathcal{E}$ is finite is without loss of generality and it simplifies the proofs. In contrast to Section 2, we allow the mechanism to have an arbitrary message space. The reason is that we cannot invoke the revelation principle when the designer offers the same mechanism repeatedly: unless the agent's best response is the same across periods, the composition of the mechanism with the agent's reporting strategy yields a time-dependent, direct mechanism. To avoid keeping track of participation and reporting strategies separately in what follows, we assume a message $m_{\varnothing} \in M$ exists such that for all $(\omega, \varepsilon) \in \Omega \times \mathcal{E}, \phi\left(m_{\varnothing}, \omega, \varepsilon\right)=\delta_{a_{\varnothing}}$.

Timing Given $\phi$, the agent faces the following extensive form. Nature draws $(\omega, \varepsilon)$ once at the beginning, unobserved to the agent. In each period, nature first draws the agent's type, which the agent observes. The agent then sends a message $m$ into the mechanism. The mechanism then draws the allocation from $\phi(\cdot \mid m, \omega, \varepsilon)$, which the agent observes.

Given the mechanism $\phi$ and the extensive form game it induces, the agent's strategy specifies for each period $t$ and each period- $t$ type $\theta \in \Theta$, a distribution over $M$, as a function of the agent's past observations, which include her past types, messages, and allocations. Importantly, we assume the agent does not observe her payoffs to focus on the agent learning through the mechanism.

We assume the agent is infinitely patient, that is, she has limit-of-means preferences. Her average payoff through period $T$ when the realization is $(\omega, \varepsilon)$ and the type-message-allocation sequence is $\left(\theta_{t}, m_{t}, a_{t}\right)_{t=1}^{T}$ is given by:

$$
U_{T}\left(\left(\theta_{t}, m_{t}, a_{t}\right)_{t=1}^{T}, \omega, \varepsilon\right)=\frac{1}{T} \sum_{t=1}^{T} u\left(a_{t}, \theta_{t}, \omega\right) .
$$

A strategy $\sigma$ is a best response for the agent if for all alternative strategies $\sigma^{\prime}$, we have that

$$
\lim \inf _{T \rightarrow \infty} \mathbb{E}_{\sigma}\left[U_{T}\right] \geq \lim \sup _{T \rightarrow \infty} \mathbb{E}_{\sigma^{\prime}}\left[U_{T}\right],
$$

where $\mathbb{E}_{\sigma}$ is the expectation relative to the measure induced over the terminal histories by the prior on $\Omega$, the distribution on $\mathcal{E}$, the agent's type distribution $f$, the mechanism $\phi$, and the agent's reporting strategy $\sigma .{ }^{31}$

[^19]Implementation Our notion of implementation is based on the induced occupation measure on $A \times \Theta \times \Omega$, that is, the (limit) expected frequency of tuples ( $a, \theta, \omega$ ) when the agent best responds to the mechanism. For this reason, we restrict attention to mechanisms $\phi$ for which (i) a best-response $\sigma$ exists, and (ii) its induced occupation measure $v_{\sigma}$ over $A \times \Theta \times M \times \Omega \times \mathcal{E}$ exists, defined as follows ${ }^{32}$

$$
v_{\sigma}(a, \theta, m, \omega, \varepsilon)=\lim _{T \rightarrow \infty} \frac{1}{T} \mathbb{E}_{\sigma}\left[\sum_{t=1}^{T} \mathbb{1}\left[\left(a_{t}, \theta_{t}, m_{t}, \omega^{\prime}, \varepsilon^{\prime}\right)=(a, \theta, m, \omega, \varepsilon)\right]\right]=\lim _{T \rightarrow \infty} v_{\sigma}^{T}(a, \theta, m, \omega, \varepsilon),
$$

where the last identity defines $v_{\sigma}$ as the limit of the up to period $T$ occupation measures $v_{\sigma}^{T}$, which are always well-defined.

Under our definition of best response, which is the same as in Hart (1985), existence of a best response implies the agent's payoff at the best-response strategy is well-defined. ${ }^{33}$ Even if the occupation measure in Equation 14 is enough to calculate the agent's payoffs, that the agent's payoffs are welldefined does not mean the occupation measure is well-defined. Because outcome distributions-and not payoffs-are usually the focus of mechanism design, we require that both the mechanism has a best response and it induces a well-defined occupation measure.

Definition 5 (Implementation). Outcome distribution $\vartheta \in \Delta(A \times \Theta \times \Omega)$ can be implemented by a repeated mechanism if a mechanism $\psi$ and a best-response strategy $\sigma$ exist such that

$$
\vartheta(a, \theta, \omega)=\sum_{\varepsilon \in \mathcal{E}, m \in M} v_{\sigma}(a, \theta, m, \omega, \varepsilon) .
$$

We are now ready to state the main result of this section. Theorem 3 shows that the outcome distributions implemented by repeated mechanisms can be implemented by two-stage mechanisms, and hence by calibrated mechanisms:

Theorem 3 (Microfoundation of Calibrated Mechanism Design). Outcome distribution $\vartheta \in \Delta(A \times \Theta \times \Omega)$ is implementable by a repeated mechanism if and only if $\vartheta$ can be implemented by an incentive compatible and individually rational two-stage mechanism, that is for all $(a, \theta, \omega) \in A \times \Theta \times \Omega$,

$$
\vartheta(a, \theta, \omega)=\mu_{0}(\omega) f(\theta) \int_{\Delta(\Omega)} \alpha(a \mid \theta, \mu) \beta(d \mu \mid \omega)
$$

where $\beta: \Omega \rightarrow \Delta(\Delta(\Omega))$ is Bayes plausible and $\alpha(\cdot \mid \cdot, \mu): \Theta \rightarrow \Delta(A)$ is incentive compatible and individually rational on the support of $\mu_{0} \otimes \beta$.

The proof of this and all results in this section can be found in Appendix C.
Theorem 3 provides a microfoundation for calibrated mechanism design. Whenever the designer is concerned with agents learning from the outcome of the mechanism and cares only about the long-run outcome distribution, it is as if he is designing a two-stage mechanism.

We now provide a proof sketch for Theorem 3, which is also useful to understand the proof of the result

[^20]in the next section. For simplicity, let $\tilde{\Omega}=\Omega \times \mathcal{E}$ with elements $\tilde{\omega}$. Suppose repeated mechanism $\phi$ implements $\vartheta$, and let $v_{\sigma} \in \Delta(A \times \Theta \times M \times \tilde{\Omega})$ denote the induced occupation measure. As the analysis so far illustrates, tracking the joint distribution over allocations, types, states, and beliefs is important to show that $\vartheta$ can be implemented via a two-stage mechanism. To this end, we extend the up to period $T$ occupation measures, $v_{\sigma}^{T} \in \Delta(A \times \Theta \times M \times \tilde{\Omega})$, to account for the frequency of beliefs through period $T$. In fact, we define two sequences of extended occupation measures over $A \times \Theta \times M \times \tilde{\Omega} \times \Delta(\tilde{\Omega})$ : the first, $\bar{v}_{\sigma}^{T, 1}$, calculates the frequency of a tuple $(a, \theta, m, \tilde{\omega}, \mu)$ by counting the beliefs at the beginning of period $t$ and the second, $\bar{v}_{\sigma}^{T, 2}$, by counting the beliefs at the end of period $t$. Whereas the martingale property of beliefs implies these two sequences have the same (subsequential) limits, they have different conditional independence properties, which we use to derive the representation of $\vartheta$ via a two-stage mechanism. Suppose for simplicity that $\bar{v}_{\sigma}^{T, 1}$ (and hence, $\bar{v}_{\sigma}^{T, 2}$ ), have limit $\bar{v}_{\sigma}$, though this assumption is not needed for the proof. ${ }^{34}$ A consequence of the martingale property of beliefs is that only the long-run beliefs of the agent are in the support of $\bar{v}_{\sigma}$.

The proof consists of three steps. In the first step, we show that $\bar{v}_{\sigma}$ admits the following decomposition:

$$
\bar{v}_{\sigma}(\{(a, \theta, m, \tilde{\omega})\} \times \tilde{\Delta})=\int_{\tilde{\Delta}} \mu(\tilde{\omega}) f(\theta) \rho(m \mid \theta, \mu) \alpha^{\prime}(a \mid m, \mu) \tau(d \mu)
$$

where (i) $\tau \in \Delta(\Delta(\tilde{\Omega}))$ has mean $\mu_{0} \otimes \eta$, where recall $\eta$ is the measure on $\mathcal{E}$, and (ii) $\rho: \Theta \times \Delta(\tilde{\Omega}) \rightarrow$ $\Delta(M)$ is a "Markovian reporting strategy", and (iii) $\alpha^{\prime}$ is almost the allocation rule in the two-stage mechanism, and hence the prime notation. Moreover, on the support of $\mu, \alpha^{\prime}(\cdot \mid \cdot, \mu)$ coincides with $\phi(\cdot \mid \cdot, \tilde{\omega})$, implying that $\phi(\cdot \mid \cdot, \tilde{\omega})$ is constant in $\tilde{\omega}$ on the support of $\mu$. This is the step which exploits the different conditional independence properties of $\bar{v}_{\sigma}^{T, 1}$ and $\bar{v}_{\sigma}^{T, 2}$. We use $\bar{v}_{\sigma}^{T, 1}$ to show the conditional independence of types and beliefs-all agent types in period $t$ have the same belief at the beginning of period $t$-and $\bar{v}_{\sigma}^{T, 2}$ to show the conditional independence of the allocation and the state-the belief at the end of period $t$ contains all the information about the state contained in the allocation.

In the second step, we show that $\rho: \Theta \times \Delta(\tilde{\Omega}) \rightarrow \Delta(M)$ is indeed a best response for the agent when her type is $\theta$ and her belief is $\mu$. In other words, the support of $\rho(\cdot \mid \theta, \mu)$ is contained in

$$
\arg \max _{m \in M} \sum_{\tilde{\omega} \in \tilde{\Omega}} \mu(\tilde{\omega}) \sum_{a \in A} \phi(a \mid m, \tilde{\omega}) u(a, \theta, \omega),
$$

for beliefs on the support of $\tau$. Hence, we can use $\rho$ and $\phi$ to define a direct mechanism $\alpha(\cdot \mid \cdot, \mu): \Theta \rightarrow$ $\Delta(A)$ that satisfies the agent's participation and incentive constraints when her belief is $\mu$. Together, Steps 1 and 2 allow us to obtain the representation of $\vartheta$ as in Equation $15 .^{35}$

Whereas the above steps are enough to show that $\vartheta$ is implementable by some individually rational and incentive compatible two-stage mechanism, they do not necessarily imply that the distribution over posteriors $\tau$ is the one induced by the information structure calibrated to $\phi, \pi_{\phi}$. The last step of the proof shows that even if this is not the case, the agent adequately learns the information contained in $\pi_{\phi}$ in the sense of Aghion et al. (1991). Indeed, Lemma C. 4 shows a strategy exists that approximately

[^21]delivers the payoff from learning $\pi_{\phi}$, so that the agent's payoff under $\sigma$ is at least the payoff she would obtain if she had learned $\pi_{\phi}$. Because the payoff from learning $\pi_{\phi}$ is the maximal payoff the agent can possibly attain, we conclude that the payoff under $\sigma$ is the payoff the agent would attain when facing the calibrated information structure $\pi_{\phi}$ (and best responding to it).

### 5.2 Dynamic Mechanisms

In this section, we consider the case in which the designer can offer the agent a dynamic mechanism, that is, one that conditions the allocation in each period on the history of past allocations and reports. The analysis herein allows us to describe the limits implied by calibration on the set of implementable outcomes.

Dynamic mechanisms A dynamic mechanism $\varphi=\left(\varphi_{t}\right)_{t \in \mathbb{N}}$ is a sequence of mappings that condition on the state, the history of participation decisions, type reports and allocations, and today's report, and output an allocation. Formally, expand the set of type reports and allocations by a non-participation message and the outside option, which we denote by $\Theta A_{\varnothing}=\Theta \times A \cup\left\{\left(\varnothing, a_{\varnothing}\right)\right\} .{ }^{36}$ For each $t \in \mathbb{N}$, define the mechanism in period $t, \varphi_{t}: \Omega \times\left(\Theta A_{\varnothing}\right)^{t-1} \times \Theta \rightarrow \Delta(A)$. Because the designer can flexibly design the mechanism in each period we no longer rely on the randomization device.

A dynamic mechanism induces an extensive-form game for the agent, in which in each period, the agent decides whether to participate, and conditional on participation what type to report. Whenever the agent chooses not to participate, she obtains her outside option $a_{\varnothing}$. We denote by $p$ the agent's participation strategy and by $\sigma$ the agent's reporting strategy.

Implementation Our notion of implementation continues to be based on the occupation measure over the set of allocations, types, participation decisions, type reports, and states, induced by the distributions $\mu_{0}$ and $f$, the mechanism $\varphi$, and the agent's participation and reporting strategy. However, as we show in Appendix D.3.1, it is without loss to focus on mechanisms such that (i) participation with probability 1 and truthtelling is a best response for the agent, and (ii) the mechanism implements the outside option with probability 1 in all future periods following a non-participation decision by the agent. ${ }^{37}$ Thus, we focus on dynamic mechanisms $\varphi$ such that (i) a best response exists, and (ii) the occupation measure over $A \times \Theta \times \Omega$ is well-defined.

Incentives in dynamic mechanisms Dynamic mechanisms allow the designer to condition the agent's allocation on the history of past participation decisions and reports (and allocations), and hence allow the designer to implement outcomes that satisfy weaker notions of truthtelling and participation, which we explain next.

Because the designer can condition the mechanism on the history of past reports, he can compare the frequency of type reports against the type distribution. So long as the agent is telling the truth, the

[^22]frequency of reports will match the type distribution $f$ over large blocks of time. In fact, any reporting strategy whose expected frequency of reports matches the type distribution will be indistinguishable from truthtelling.

Definition 6 (Undetectable deviations). An undetectable deviation is a reporting strategy $\sigma: \Theta \rightarrow \Delta(\Theta)$ such that for all $\theta^{\prime} \in \Theta$

$$
\sum_{\theta \in \Theta} f(\theta) \sigma\left(\theta^{\prime} \mid \theta\right)=f\left(\theta^{\prime}\right) .
$$

By tracking the empirical distribution of type reports, the designer can dissuade the agent from employing detectable deviations. Thus, in a dynamic mechanism, the designer should be concerned with only discouraging undetectable deviations. This leads to a weaker notion of incentive compatibility for allocation rules:

Definition 7 (Unprofitable undetectable deviations). The allocation rule $\alpha: \Theta \times \Delta(\Omega) \rightarrow \Delta(A)$ lacks profitable undetectable deviations at belief $\mu \in \Delta(\Omega)$ if for all undetectable deviations $\sigma$,

$$
\sum_{\theta \in \Theta} f(\theta) \sum_{a \in A} \alpha(a \mid \theta, \mu) \sum_{\omega \in \Omega} \mu(\omega) u(a, \theta, \omega) \geq \sum_{\theta \in \Theta} f(\theta) \sum_{\theta^{\prime} \in \Theta} \sigma\left(\theta^{\prime} \mid \theta\right) \sum_{a \in A} \alpha\left(a \mid \theta^{\prime}, \mu\right) \sum_{\omega \in \Omega} \mu(\omega) u(a, \theta, \omega) .
$$

A two-stage mechanism $\psi$ with allocation rule $\alpha$ lacks profitable undetectable deviations if $\alpha(\cdot \mid \cdot, \mu)$ lacks profitable undetectable deviations for all beliefs in the support of the mechanism.

To illustrate the difference between the lack of profitable undetectable deviations and incentive compatibility, consider the following example from Ball and Kattwinkel (2023). Suppose the agent types are binary, $\left\{\theta_{1}, \theta_{2}\right\}$, and equally likely. The set of allocations, $q \in\{0,1\}$, describes whether the agent receives a good. Finally, suppose the agent's payoff is $u(q, \theta)=q \theta$ and $\theta_{1}<\theta_{2}$. Consider the mechanism that allocates the good to $\theta_{2}$ : While it is not incentive compatible, it lacks profitable undetectable deviations. The constraint that the deviation must be undetectable implies the gains from $\theta_{1}$ obtaining the good come at the expense of $\theta_{2}$ getting the good.

Consider now the agent's participation incentives in the dynamic mechanism: once the agent rejects the mechanism once, the agent obtains her outside option in all continuation histories independent of her participation decision and her types. In other words, whereas the agent can always ensure her outside option by rejecting the mechanism in a given period, she is effectively quitting the mechanism forever for all her types. The following definition introduces the notion of individual rationality satisfied by the mechanism in the long run.

Definition 8 (Ex ante individual rationality). The allocation rule $\alpha: \Theta \times \Delta(\Omega) \rightarrow \Delta(A)$ is ex ante individually rational at belief $\mu \in \Delta(\Omega)$ if

$$
\sum_{\theta \in \Theta} f(\theta) \sum_{a \in A} \alpha(a \mid \theta, \mu) \sum_{\omega \in \Omega} \mu(\omega) u(a, \theta, \omega) \geq \sum_{\theta \in \Theta} f(\theta) \sum_{\omega \in \Omega} \mu(\omega) u\left(a_{\varnothing}, \theta, \omega\right) .
$$

A two-stage mechanism $\psi$ with allocation rule $\alpha$ is ex ante individually rational if $\alpha(\cdot \mid \cdot, \mu)$ is ex ante individually rational for all beliefs in the support of the mechanism.

We are now ready to state the main result of this section:
Theorem 4 (Implementable Outcomes via Dynamic Mechanisms). A dynamic mechanism exists that implements outcome $\vartheta \in \Delta(A \times \Theta \times \Omega)$ if and only if $\vartheta$ can be implemented by an ex ante individually rational two-stage mechanism which lacks profitable undetectable deviations. That is, if and only if for all $(a, \theta, \omega) \in A \times \Theta \times \Omega$

$$
\vartheta(a, \theta, \omega)=\mu_{0}(\omega) f(\theta) \int_{\Delta(\Omega)} \alpha(a \mid \theta, \mu) \beta(d \mu \mid \omega)
$$

where $\beta: \Omega \rightarrow \Delta(\Delta(\Omega))$ is Bayes plausible and $\alpha(\cdot \mid \cdot, \mu): \Theta \rightarrow \Delta(A)$ lacks profitable undetectable deviations and is ex ante individually rational on the support of $\mu_{0} \otimes \beta$.

Theorem 4 characterizes the outcome distributions implementable by dynamic mechanisms as those implemented by two-stage mechanisms that satisfy the incentive constraints: unprofitability of undetectable deviations and ex ante individual rationality. Notably, both notions of incentive constraints apply in the aggregate over the type distribution, which reflects the transient nature of the agent's private information.

Comparing Theorem 3 and Theorem 4, we see that dynamic mechanisms allow the designer to weaken the incentive constraints of the agent, but do not allow him to engage in richer-i.e., typedependent-disclosures. Despite dynamic mechanisms implying weaker incentive constraints, we can build on the results of Rochet (1987) and Rahman (2024) to show that in settings with transferable utility, where $a=(q, t)$, dynamic and repeated mechanisms implement the same set of physical allocations $q: \Theta \times \Omega \rightarrow \mathbb{R} .^{38}$ Indeed, Rahman (2024) shows the lack of profitable undetectable deviations is equivalent to cyclical monotonicity in Rochet (1987). Thus, in settings with transferable utility, Theorems 3 and 4 imply that dynamic mechanisms do not allow the designer to expand on the set of implementable distributions over $(q, \theta, \omega)$.

The proof of the only if direction is similar to that of Theorem 3, in that we similarly extend the occupation measure to account for the agent's beliefs and show it satisfies the conditional independence properties implied by a two-stage mechanism. In a dynamic mechanism, however, the agent can ensure the payoff of some, but not all deviations. The latter property is what delivers that the two-stage mechanism must lack profitable undetectable deviations.

The proof of the if direction, instead, harnesses a construction in Margaria and Smolin (2018). The proof proceeds in two steps. In the first step, we analyze a fictitious model without state uncertainty in which a designer faces a privately informed agent, so that implementable outcomes are elements of $\Delta(A \times \Theta)$. We show that if $\vartheta^{\prime}(a, \theta)=f(\theta) \alpha^{\prime}(a \mid \theta) \in \Delta(A \times \Theta)$ is such that $\alpha^{\prime}$ lacks profitable undetectable deviations and is ex ante individually rational, then a dynamic mechanism exists that implements $\vartheta^{\prime} .^{39}$ This is the step that relies on Margaria and Smolin (2018). We construct a dynamic mechanism, which can be split into blocks of random length. Each block consists of two phases: a reporting

[^23]phase and an adjustment phase. In the reporting phase, the mechanism uses the agent's reports to determine the allocation. Instead, in the adjustment phase, the mechanism simulates type reports so that the frequency of type reports matches the type distribution (in expectation) over the length of the block, whenever this is not the case at the end of the reporting phase. These two steps ensure that the expected frequency of type reports and allocations matches $\vartheta^{\prime}$. We then leverage that $\alpha^{\prime}(\cdot \mid \theta)$ lacks profitable undetectable deviations to show the agent cannot do better than by telling the truth. Hence, the induced frequency of types and allocations also matches $\vartheta^{\prime}$. Moreover, the construction ensures that after any history, truthtelling delivers a continuation payoff equal to the ex ante payoff. Because $\alpha^{\prime}$ is ex ante individually rational, we conclude the participation constraints are satisfied.

The second step uses the above result and the representation of the outcome distribution via a twostage mechanism to construct a dynamic mechanism that implements any outcome distribution that satisfies the properties in Theorem 4. Indeed, one can construct a dynamic mechanism which uses a finite number of steps to disclose information to the agent via the realized allocations, ${ }^{40}$ and then continues as in the above construction to implement the allocation rule $\alpha(\cdot \mid \cdot, \mu)$.

## 6 Conclusions

Many economic institutions-online platforms, lenders, regulators-rely on mechanisms that remain fixed while agents interact with them repeatedly. When the mechanism's operation depends on a state known only to the designer, agents can learn this state from their outcomes, constraining what the mechanism can implement. We introduce calibrated mechanism design, a static solution concept that requires mechanisms to remain incentive compatible given the information they endogenously reveal about the designer's private state through repeated use. In private value environments, the calibration constraint pushes the designer toward full transparency, precluding Crémer-McLean-style schemes under transferable utility. In single agent-settings, calibrated mechanisms are equivalent to two-stage mechanisms. This equivalence yields a practical algorithm for finding optimal calibrated mechanisms, combining tools from information design and mechanism design. We provide a microfoundation by showing calibrated mechanisms characterize exactly what is implementable when an infinitely patient agent repeatedly interacts with the same mechanism, and study the implications on implementable outcomes of allowing the designer to offer fully dynamic mechanisms.

The most important direction for future work is deepening the analysis of multi-agent settings. On the one hand, understanding when generalized two-stage mechanisms coincide with calibrated mechanisms would enable the study of multi-agent applications, while abstracting from the dynamics of experimentation. On the other hand, extending our microfoundation to the multi-agent case would further ground calibrated mechanism design. More broadly, our framework suggests that any institution whose repeated operation leaks information about its designer's knowledge faces a fundamental tradeoff between conditioning the mechanism on this information and the information this leaks to participants, and calibrated mechanism design offers a disciplined way to analyze it.

[^24]
## References

Aghion, P., P. Bolton, C. Harris, and B. Jullien (1991): "Optimal Learning by Experimentation," Review of Economic Studies, 58, 621-654.

Aliprantis, C. D. and K. C. Border (2006): Infinite Dimensional Analysis: a Hitchhiker's Guide, Springer.

Alonso, R. and O. Câmara (2016): "Bayesian Persuasion with Heterogeneous Priors," Journal of Economic Theory, 165, 672-706.

Arieli, I., Y. Babichenko, and F. Sandomirskiy (2024): "Feasible Conditional Belief Distributions," .
Attar, A., E. Campioni, T. Mariotti, and A. Pavan (2025): "Keeping the Agents in the Dark: Competing Mechanisms, Private Disclosures, and the Revelation Principle," Working Paper.

Ball, I. and D. Kattwinkel (2023): "Quota Mechanisms: Finite-Sample Optimality and Robustness," Working Paper.

Bergemann, D., P. Duetting, R. Paes Leme, and S. Zuo (2022a): "Calibrated Click-Through Auctions," in Proceedings of the ACM Web Conference 2022, 47-57.

Bergemann, D., T. Heumann, and S. Morris (2022b): "Screening with Persuasion," Working Paper.
Bergemann, D. and J. Hörner (2018): "Should First-Price Auctions Be Transparent?" American Economic Journal: Microeconomics, 10, 177-218.

Bergemann, D. and M. Pesendorfer (2007): "Information Structures in Optimal Auctions," Journal of Economic Theory, 137, 580-609.

Blume, L. E., M. M. Bray, and D. Easley (1982): "Introduction to the Stability of Rational Expectations Equilibrium," Journal of Economic Theory, 26, 313-317.

Bogachev, V. I. (2007): Measure Theory, Springer.
Calzolari, G. and A. Pavan (2006): "On the Optimality of Privacy in Sequential Contracting," Journal of Economic Theory, 130, 168-204.

Cesa-Bianchi, N., T. Cesari, R. Colomboni, F. Fusco, and S. Leonardi (2024): "The Role of Transparency in Repeated First-Price Auctions with Unknown Valuations," in Proceedings of the 56th Annual ACM Symposium on Theory of Computing, 225-236.

Crémer, J. and R. P. McLean (1988): "Full Extraction of the Surplus in Bayesian and Dominant Strategy Auctions," Econometrica, 1247-1257.

Daskalakis, C., C. Papadimitriou, and C. Tzamos (2016): "Does Information Revelation Improve Revenue?" in Proceedings of the 2016 ACM Conference on Economics and Computation, 233-250.

Doval, L. and V. Skreta (2022): "Mechanism Design with Limited Commitment," Econometrica, 90, 1463-1500.

Dworczak, P. (2020): "Mechanism Design with Aftermarkets: Cutoff Mechanisms," Econometrica, 88, 2629-2661.

Eső, P. and B. Szentes (2007): "Optimal Information Disclosure in Auctions and the Handicap Auction," Review of Economic Studies, 74, 705-731.

Esponda, I. (2008): "Information Feedback in First Price Auctions," RAND Journal of Economics, 39, 491-508.

Foster, D. P. and R. V. Vohra (1997): "Calibrated Learning and Correlated Equilibrium," Games and Economic Behavior, 21, 40-55.

Fu, H., P. Jordan, M. Mahdian, U. Nadav, I. Talgam-Cohen, and S. Vassilvitskii (2012): "Ad Auctions with Data," in International Symposium on Algorithmic Game Theory, Springer, 168-179.

Gentzkow, M. and E. Kamenica (2017): "Bayesian Persuasion with Multiple Senders and Rich Signal Spaces," Games and Economic Behavior, 104, 411-429.

Golrezaei, N., A. Javanmard, and V. Mirrokni (2019): "Dynamic Incentive-Aware Learning: Robust Pricing in Contextual Auctions," Advances in Neural Information Processing Systems, 32.

Green, J. (1977): "The Non-Existence of Informational Equilibria," Review of Economic Studies, 44, 451-463.

Green, J. R. and J.-J. Laffont (1987): "Posterior Implementability in a Two-Person Decision Problem," Econometrica, 69-94.

Green, J. R. and N. L. Stokey (2022): "Two Representations of Information Structures and Their Comparisons," Decisions in Economics and Finance, 45, 541-547.

Guesnerie, R. and J.-J. Laffont (1984): "A Complete Solution to a Class of Principal-Agent Problems with an Application to the Control of a Self-Managed Firm," Journal of Public Economics, 25, 329-369.

Haberman, A. and R. Jagadeesan (2025): "Auctions with Withdrawal Rights: A Foundation for Uniform Price," in Proceedings of the 26th ACM Conference on Economics and Computation, 35-35.

Häfner, S., M. Pycia, and H. Zeng (2025): "Mechanism Design with Information Leakage," Working Paper.

Hart, S. (1985): "Nonzero-Sum Two-Person Repeated Games with Incomplete Information," Mathematics of Operations Research, 10, 117-153.

Jackson, M. O. and H. F. Sonnenschein (2007): "Overcoming Incentive Constraints by Linking Decisions," Econometrica, 75, 241-257.

Kallenberg, O. (2017): Random Measures, Theory and Applications, vol. 1, Springer.
Kanoria, Y. and H. Nazerzadeh (2020): "Dynamic Reserve Prices for Repeated Auctions: Learning from Bids," Working Paper.

Kartik, N., S. Lee, and D. Rappoport (2024): "Single-Crossing Differences in Convex Environments," Review of Economic Studies, 91, 2981-3012.

Krähmer, D. (2020): "Information Disclosure and Full Surplus Extraction in Mechanism Design," Journal of Economic Theory, 187, 105020.

Kreps, D. M. (1977): "A Note on "Fulfilled Expectations" Equilibria," Journal of Economic Theory, 14, 32-43.

Laclau, M. and L. Renou (2017): "Public Persuasion," Working Paper.
Li, H. and X. Shi (2017): "Discriminatory Information Disclosure," American Economic Review, 107, 3363-3385.

Margaria, C. and A. Smolin (2018): "Dynamic Communication with Biased Senders," Games and Economic Behavior, 110, 330-339.

Maskin, E. and J. Tirole (1990): "The Principal-Agent Relationship with an Informed Principal: The Case of Private Values," Econometrica, 379-409.

Meng, D. (2021): "On the Value of Repetition for Communication Games," Games and Economic Behavior, 127, 227-246.

Milgrom, P. R. (1981): "Rational Expectations, Information Acquisition, and Competitive Bidding," Econometrica, 921-943.

Milgrom, P. R. and R. J. Weber (1982): "A Theory of Auctions and Competitive Bidding," Econometrica, 1089-1122.

Myerson, R. B. (1983): "Mechanism Design by an Informed Principal," Econometrica, 1767-1797.

- (1986): "Multistage Games with Communication," Econometrica, 323-358.

Nedelec, T., C. Calauzènes, N. El Karoui, and V. Perchet (2022): "Learning in Repeated Auctions," Found. Trends Mach. Learn., 15, 176-334.

Nedelec, T., N. E. Karoui, and V. Perchet (2019): "Learning to Bid in Revenue-Maximizing Auctions," in Proceedings of the 36th International Conference on Machine Learning, vol. 97, 4781-4789.

Niemeyer, A. (2022): "Posterior Implementability in an N-Person Decision Problem," Working Paper.
Ottaviani, M. and A. Prat (2001): "The Value of Public Information in Monopoly," Econometrica, 69, 1673-1683.

Pavan, A., I. Segal, and J. Toikka (2014): "Dynamic Mechanism Design: A Myersonian Approach," Econometrica, 82, 601-653.

Radner, R. (1979): "Rational Expectations Equilibrium: Generic Existence and the Information Revealed by Prices," Econometrica, 655-678.

Rahman, D. M. (2024): "Detecting Profitable Deviations," Journal of Mathematical Economics, 111, 102946.

Renou, L. and T. Tomala (2015): "Approximate Implementation in Markovian Environments," Journal of Economic Theory, 159, 401-442.

Rochet, J.-C. (1987): "A Necessary and Sufficient Condition for Rationalizability in a Quasi-Linear Context," Journal of Mathematical Economics, 16, 191-200.

Rubin, H. and O. Wesler (1958): "A Note on Convexity in Euclidean N-Space," in Proceedings of the American Mathematical Society, vol. 9, 522-523.

Smolin, A. (2023): "Disclosure and Pricing of Attributes," RAND Journal of Economics, 54, 570-597.
Szabadi, B. (2018): "Essays in Microeconomic Theory: Information Design, Partnerships, and Matching," Ph.D. thesis, Northwestern University.

Yamashita, T. (2018): "Optimal Public Information Disclosure by Mechanism Designer," Working Paper.

## Mathematical conventions

Throughout the appendix, we take all sets to be Polish spaces, that is, completely metrizable, separable, topological spaces, and endow them with their Borel $\sigma$-algebra. We endow product spaces with their product $\sigma$-algebra. For a Polish space $X$, we let $\mathcal{B}_{X}$ denote its Borel $\sigma$-algebra and $\Delta(X)$ the set of all Borel probability measures on $X$, endowed with the weak* topology. Thus, $\Delta(X)$ is also a Polish space (Aliprantis and Border, 2006), and it is compact, whenever $X$ is compact (Aliprantis and Border, 2006, Theorem 15.11 and Theorem 15.15).

Notational conventions If $X$ is a Polish space, $\tilde{X}$ denotes a measurable subset of $X$, i.e., an element of the Borel $\sigma$-algebra on $X$, and $C_{b}(X)$ denotes the set of continuous and bounded functions on $X$. Given a measure $v \in \Delta\left(\times_{i=1}^{N} Y_{i}\right)$, we denote by $v_{Y_{j} Y_{k} \ldots Y_{l}}$ the marginal of $v$ on $Y_{j} Y_{k} \ldots Y_{l}$. When one of the $Y_{i}=\Delta\left(X_{i}\right)$, we write $\Delta$ instead of $Y_{i}$ in the subscript, when it is unlikely to generate confusion.

Throughout the appendix, we define different distributions that arise in our proofs. Because we endow product spaces with their product topology and their product Borel $\sigma$-algebra, it is enough to define these new measures on the measurable rectangles and we follow this convention throughout.

Disintegration We rely on the notion of disintegration in many of our proofs (Bogachev, 2007, Chapter 10.6). We define disintegration in the context of product sets $X \times Y$, as this is the one that shows up in the proof, but it is more general than this. Given a measure $v \in \Delta(X \times Y), \lambda: X \times \mathcal{B}_{Y} \rightarrow[0,1]$ is the disintegration of $v$ along $X$ if the following holds

1. For all $\tilde{Y} \in \mathcal{B}_{Y}, x \mapsto \lambda_{x}(\tilde{Y})$ is measurable,
2. For $v_{X}$-almost everywhere $x \in X, \tilde{Y} \mapsto \lambda_{x}(\tilde{Y})$ is a probability measure, and
3. For every bounded measurable function $g: X \times Y \rightarrow \mathbb{R}$,
$$
\int_{X \times Y} g(x, y) v(d(x, y))=\int_{X} \int_{Y} g(x, y) \lambda_{x}(d y) v_{X}(d x)
$$

Kallenberg (2017, Theorem 1.23) ensures that $\left\{\lambda_{x}: x \in X\right\}$ exists and is unique $v_{X}$-almost everywhere.

## A Omitted proofs from Section 2

Proof of Theorem 1. Suppose the agents' payoffs are state independent and in a slight abuse of notation let $u_{i}\left(a_{i}, \theta_{i}\right)$ denote agent $i$ 's utility.

The calibrated mechanism design problem is

$$
\max _{\phi: \Theta \times \Omega \times[0,1] \rightarrow \Delta(A)} \sum_{\omega \in \Omega} \mu_{0}(\omega) \sum_{\theta \in \Theta} f(\theta \mid \omega) \int_{0}^{1} w(\phi(\theta, \omega, \varepsilon), \theta, \omega) \lambda(d \varepsilon),
$$

subject to the following constraints holding for all $(\omega, \varepsilon) \in \Omega \times[0,1], i \in[N], \theta_{i} \in \Theta$ and $\theta_{i}^{\prime} \in \Theta$ :

$$
\begin{aligned}
& \sum_{a_{i} \in A_{i}} \mathbb{E}_{f_{-i}(\cdot \mid \omega)}\left[\sum_{a_{-i} \in A_{-i}} \phi\left(\theta_{i}, \theta_{-i}, \omega, \varepsilon\right)\left(a_{i}, a_{-i}\right)\right] u_{i}\left(a_{i}, \theta_{i}\right) \geq \sum_{\left(a_{i} \in A_{i}\right)} \mathbb{E}_{f_{-i}(\cdot \mid \omega)}\left[\sum_{a_{-i} \in A_{-i}} \phi\left(\theta_{i}^{\prime}, \theta_{-i}, \omega, \varepsilon\right)\left(a_{i}, a_{-i}\right)\right] u_{i}\left(a_{i}, \theta_{i}\right) \\
& \sum_{a_{i} \in A_{i}} \mathbb{E}_{f_{-i}(\cdot \mid \omega)}\left[\sum_{a_{-i} \in A_{-i}} \phi\left(\theta_{i}, \theta_{-i}, \omega, \varepsilon\right)\left(a_{i}, a_{-i}\right)\right] u_{i}\left(a_{i}, \theta_{i}\right) \geq u_{i}\left(a_{i \varnothing}, \theta_{i}\right) .
\end{aligned}
$$

In other words, for each agent $i$, her interim allocation rule $\pi_{\phi, i}(\omega, \varepsilon)$ must be an element of $S_{I C / I R, i}^{*}$, where the latter is the set of interim allocation rules $S_{i}^{*}: \Theta_{i} \rightarrow \Delta\left(A_{i}\right)$ that satisfy the following incentive compatibility and individual rationality constraints:

$$
\begin{gathered}
\left(\forall \theta_{i}, \theta_{i}^{\prime} \in \Theta_{i}\right) \sum_{a_{i} \in A_{i}} s_{i}^{*}\left(a_{i} \mid \theta_{i}\right) u_{i}\left(a_{i}, \theta_{i}\right) \geq \sum_{a_{i} \in A_{i}} s_{i}^{*}\left(a_{i} \mid \theta_{i}^{\prime}\right) u_{i}\left(a_{i}, \theta_{i}\right) \\
\left(\forall \theta_{i} \in \Theta_{i}\right) \sum_{a_{i} \in A_{i}} s_{i}^{*}\left(a_{i} \mid \theta_{i}\right) u_{i}\left(a_{i}, \theta_{i}\right) \geq u_{i}\left(a_{i \varnothing}, \theta_{i}\right)
\end{gathered}
$$

Because the individual rationality and incentive constraints must hold for each pair $(\omega, \varepsilon)$, the designer's problem is separable across variables for different $\omega, \varepsilon$ : the sets of variables $\phi(\cdot, \omega, \varepsilon)$ appear in different sets of constraints and the objective function is additively separable across those variables. Consequently, the designer's problem can be solved as a collection of independent problems, one for each $\omega, \varepsilon$. $\square$

## B Omitted proofs from Section 3

In this section, we present the proofs of Theorem 2 and Proposition 1. We proceed as follows: We first prove Proposition 1, as when $N=1$ its proof implies the "if" direction of Theorem 2. We then prove the "only if" direction of Theorem 2.

Proof of Proposition 1. We focus on the case in which types and states are independently distributed, and explain how to extend the proof when they are not.

Let $\vartheta \in \Delta(A \times \Theta \times \Omega)$ denote the outcome distribution implemented by an incentive compatible and individually rational calibrated mechanism. We show transition probabilities $\beta: \Omega \rightarrow \Delta\left(\Delta(\Omega)^{N}\right)$, $\bar{\alpha}: \Theta \times \Omega \times \Delta(\Omega)^{N} \rightarrow \Delta(A)$, and $\alpha_{i}: \Theta_{i} \times \Delta(\Omega) \rightarrow \Delta\left(A_{i}\right)$ for $i \in\{1, \ldots, N\}$ exist such that

$$
\vartheta(a, \theta, \omega)=\mu_{0}(\omega) f(\theta) \int_{\Delta(\Omega)^{N}} \bar{\alpha}\left(a \mid \theta, \omega, \mu_{1}, \ldots, \mu_{N}\right) \beta\left(d\left(\mu_{1}, \ldots, \mu_{N}\right) \mid \omega\right),
$$

and for all $i \in\{1, \ldots, N\}$, (i) $\alpha_{i}$ satisfies item 3 of Definition 4, and (ii) on the support of $\mu_{0} \otimes \beta, \alpha_{i}\left(\cdot \mid \cdot, \mu_{i}\right)$ is incentive compatible and individually rational when agent $i$ holds belief $\mu_{i}$.

Let $\pi_{\omega, i}:[0,1] \rightarrow S_{i}^{*}$ denote the mapping $\varepsilon \mapsto \pi_{i}(\omega, \cdot)$. For $\tilde{S}_{i}^{*} \in \Delta\left(A_{i}\right)^{\Theta_{i}}$, define

$$
\operatorname{Pr}_{i}\left(\{\omega\} \times \tilde{S}_{i}^{*}\right)=\mu_{0}(\omega) \lambda\left(\pi_{\omega, i}^{-1}\left(\tilde{S}_{i}^{*}\right)\right)=\int_{\tilde{S}_{i}^{*}} \mu_{i}\left(\omega \mid s_{i}^{*}\right) \tau_{\phi, i}\left(d s_{i}^{*}\right)
$$

where the second equality follows from disintegration of $\operatorname{Pr}_{i} \in \Delta\left(\Omega \times S_{i}^{*}\right)$ along $S_{i}^{*}$, and corresponds to the definition of Bayes rule for agent $i$. Define the measurable mappings, $T_{i}: S_{i}^{*} \rightarrow \Delta(\Omega)$ and
$T: S^{*} \rightarrow \Delta(\Omega)^{N}$ as follows: $T_{i}\left(s_{i}^{*}\right)=\mu_{i}\left(\cdot \mid s_{i}^{*}\right)$ and $T\left(s^{*}\right)=\left(T_{1}\left(s_{1}^{*}\right), \ldots, T_{N}\left(s_{N}^{*}\right)\right)$.
Define a joint distribution $Q \in \Delta\left(A \times \Theta \times \Omega \times \Delta(\Omega)^{N}\right)$ as follows:

$$
Q\left(\{(a, \theta, \omega)\} \times \times_{i=1}^{N} \tilde{\Delta}_{i}\right)=\mu_{0}(\omega) f(\theta) \int_{\pi_{\omega}^{-1}\left(T^{-1}\left(\times \tilde{\Delta}_{i}\right)\right)} \phi(a \mid \theta, \omega, \varepsilon) \lambda(d \varepsilon)
$$

where $\pi_{\omega}^{-1}\left(T^{-1}\left(\times \tilde{\Delta}_{i}\right)\right)=\cap_{i=1}^{N}\left\{\varepsilon: T_{i}\left(\pi_{\omega, i}(\varepsilon)\right) \in \tilde{\Delta}_{i}\right\}$.
We note the following properties of $Q$. First, consider its marginal over $\Theta \times \Omega \times \Delta(\Omega)^{N}$,

$$
Q_{\Theta \Omega \Delta^{N}}\left(\{(\theta, \omega)\} \times \times_{i=1}^{N} \tilde{\Delta}_{i}\right)=\mu_{0}(\omega) f(\theta) \lambda\left(\cap_{i=1}^{N}\left\{\varepsilon: T_{i}\left(\pi_{\omega, i}(\varepsilon)\right) \in \tilde{\Delta}_{i}\right\}\right),
$$

which implies that the disintegration of $Q_{\Theta \Omega \Delta^{N}}$ along $\Theta \times \Omega, \beta: \Theta \times \Omega \rightarrow \Delta\left(\Delta(\Omega)^{N}\right)$ does not depend on $\theta$. This automatically implies that $Q$ admits the following disintegration:

$$
Q\left(\{(a, \theta, \omega)\} \times \times_{i=1}^{N} \tilde{\Delta}_{i}\right)=\mu_{0}(\omega) f(\theta) \int_{\times_{i=1}^{N} \tilde{\Delta}_{i}} \bar{\alpha}\left(a \mid \theta, \omega, \mu_{1}, \ldots, \mu_{N}\right) \beta\left(d\left(\mu_{1}, \ldots, \mu_{N}\right) \mid \omega\right)
$$

which, in turn, delivers Equation B.1. Moreover, note that the marginal of $\beta$ on the beliefs of agent $i$, $\beta_{i}: \Omega \rightarrow \Delta(\Delta(\Omega))$, satisfies

$$
\beta_{i}\left(\tilde{\Delta}_{i} \mid \omega\right)=\lambda\left(\pi_{\omega, i}^{-1}\left(T_{i}^{-1}\left(\tilde{\Delta}_{i}\right)\right)\right) .
$$

Consider now the marginal on $A_{i} \times \Theta_{i} \times \Omega \times \Delta(\Omega)$ of $Q, Q_{A_{i} \Theta_{i} \Omega \Delta_{i}}$, which satisfies:

$$
\begin{aligned}
& Q_{A_{i} \Theta i \Omega \Delta_{i}}\left(\left\{\left(a_{i}, \theta_{i}, \omega\right)\right\} \times \tilde{\Delta}_{i}\right)=\sum_{\theta_{-i} \in \Theta_{-i}} \sum_{a_{-i} \in A_{-i}} Q\left(\{(a, \theta, \omega)\} \times \tilde{\Delta}_{i} \times \Delta(\Omega)^{N-1}\right)= \\
& =\mu_{0}(\omega) f_{i}\left(\theta_{i}\right) \int_{\pi_{\omega, i}^{-1}\left(T_{i}^{-1}\left(\tilde{\Delta}_{i}\right)\right)}\left(\sum_{\theta_{-i} \in \Theta_{-i}} f_{-i}\left(\theta_{-i}\right) \sum_{a_{-i} \in A_{-i}} \phi\left(a_{i}, a_{-i} \mid \theta_{i}, \theta_{-i}, \omega, \varepsilon\right)\right) \lambda(d \varepsilon) \\
& =\mu_{0}(\omega) f_{i}\left(\theta_{i}\right) \int_{\pi_{\omega, i}^{-1}\left(T_{i}^{-1}\left(\tilde{\Delta}_{i}\right)\right)} \pi_{\omega, i}(\varepsilon)\left(a_{i} \mid \theta_{i}\right) \lambda(d \varepsilon)=\mu_{0}(\omega) f_{i}\left(\theta_{i}\right) \int_{T_{i}^{-1}\left(\tilde{\Delta}_{i}\right)} s_{i}^{*}\left(a_{i} \mid \theta_{i}\right)\left(\lambda \circ \pi_{\omega, i}^{-1}\right)\left(d s_{i}^{*}\right)
\end{aligned}
$$

Lastly, $Q_{A_{i} \Theta_{i} \Omega \Delta_{i}}$ admits the following representation via disintegration:

$$
\begin{aligned}
& Q_{A_{i} \Theta_{i} \Omega \Delta_{i}}\left(\left\{\left(a_{i}, \theta_{i}, \omega\right)\right\} \times \tilde{\Delta}_{i}\right)=\mu_{0}(\omega) f_{i}\left(\theta_{i}\right) \int_{\tilde{\Delta}_{i}} \alpha_{i}\left(a_{i} \mid \theta_{i}, \omega, \mu_{i}\right) \beta_{i}\left(d \mu_{i} \mid \omega\right) \\
& =\mu_{0}(\omega) f_{i}\left(\theta_{i}\right) \int_{\tilde{\Delta}_{i}} \alpha_{i}\left(a_{i} \mid \theta_{i}, \omega, \mu_{i}\right)\left(\lambda \circ \pi_{\omega, i}^{-1} \circ T_{i}^{-1}\right)\left(d \mu_{i}\right)
\end{aligned}
$$

Together with the uniqueness of disintegration and the sufficiency property of beliefs, Equations B. 3 and B. 4 imply that $\alpha_{i}$ does not depend on $\omega$. The incentive compatibility and individual rationality of $\alpha_{i}$ follows from that of the calibrated mechanism.

Finally, consider the case in which $\theta$ and $\omega$ are not independent. Then, the experiment $\beta$ in Equation B. 2 induces a joint distribution over the beliefs of $N$ fictitious agents whose prior over the state is given by $\mu_{0}$. Agent $i$ 's updated beliefs when her type is $\theta_{i}$ obtain from a transformation of $\mu$ (Alonso and Câmara, 2016; Laclau and Renou, 2017). ${ }^{41}$ Thus, up to changing $f(\theta)$ by $f(\theta \mid \omega)$, and interpreting

[^25]the draw from the Blackwell experiment as the posterior of an agent with prior belief $\mu_{0}$, the result follows. $\square$

Proof of Theorem 2.
"Only if" direction Similar to the proof of Proposition 1 , we focus on the case in which $\theta$ and $\omega$ are independent. Suppose $\vartheta \in \Delta(A \times \Theta \times \Omega)$ is implemented by an incentive compatible and individually rational two-stage mechanism. That is,

$$
\vartheta(a, \theta, \omega)=\mu_{0}(\omega) f(\theta) \int_{\Delta(\Omega)} \alpha(a \mid \theta, \mu) \beta(d \mu \mid \omega)
$$

and $\alpha$ is incentive compatible and individually rational on the support of $\mu_{0} \otimes \beta$. We construct an incentive compatible and individually rational calibrated mechanism that implements $\vartheta$.

First, if $\vartheta$ satisfies Equation B.5, Rubin and Wesler (1958) and Carathéodory's theorem (Aliprantis and Border, 2006, Theorem 5.32) imply that a finite support $\beta^{\prime}: \Omega \rightarrow \Delta\left(\left\{\mu_{1}, \ldots, \mu_{K}\right\}\right)$ exists such that for all $(a, \theta, \omega) \in A \times \Theta \times \Omega^{42}$

$$
\vartheta(a, \theta, \omega)=\mu_{0}(\omega) f(\theta) \sum_{k=1}^{K} \alpha\left(a \mid \theta, \mu_{k}\right) \beta^{\prime}\left(\left\{\mu_{k}\right\} \mid \omega\right) .
$$

For each $\omega \in \Omega$, partition $[0,1]=\cup_{k=1}^{K-2}\left[b_{k}^{\omega}, b_{k+1}^{\omega}\right) \cup\left[b_{K-1}^{\omega}, 1\right]$, where $b_{1}=0$, and for all $k \in\{1, \ldots, K-2\}$, $b_{k+1}^{\omega}=\sum_{l=1}^{k} \beta^{\prime}\left(\left\{\mu_{l}\right\} \mid \omega\right)$. Define for $\varepsilon \in\left[b_{k}^{\omega}, b_{k+1}^{\omega}\right)$

$$
\phi(a \mid \theta, \omega, \varepsilon)=\alpha\left(a \mid \theta, \mu_{k}\right) .
$$

The calibrated information structure is $\pi_{\phi}(\omega, \varepsilon)=\phi(\cdot \mid \cdot, \omega, \varepsilon)=\alpha\left(\cdot \mid \theta, \mu_{k}\right)$ for $\varepsilon \in\left[b_{k}^{(w}, b_{k+1}^{(w)}\right.$ if $m \leq K-2$ or $\varepsilon \in\left[b_{K-1}^{\omega}, 1\right]$.

We now show that for all $\theta$ and all $s \in \operatorname{supp} \pi_{\phi}$, the mechanism $\phi$ is incentive compatible and individually rational. Note that $\pi_{\phi}$ has finite support, and let $s \in \operatorname{supp} \pi_{\phi}$ and let $\mu(\cdot \mid s)$ denote the updated posterior. Then, $k$ exists such that the following holds:

$$
\mu(\omega \mid s)=\frac{\mu_{0}(\omega) \lambda(\{\varepsilon: \pi(\omega, \varepsilon)=s\})}{\sum_{\omega^{\prime} \in \Omega} \mu_{0}\left(\omega^{\prime}\right) \lambda\left(\left\{\varepsilon: \pi\left(\omega^{\prime}, \varepsilon\right)=s\right\}\right)}=\frac{\mu_{0}(\omega)\left(b_{k+1}^{\omega}-b_{k}^{\omega}\right)}{\sum_{\omega^{\prime} \in \Omega} \mu_{0}\left(\omega^{\prime}\right)\left(b_{k+1}^{\omega^{\prime}}-b_{k}^{\omega^{\prime}}\right)}=\mu_{k}(\omega) .
$$

Moreover, because $\phi(\cdot \mid \cdot, \omega, \varepsilon)=\alpha\left(\cdot \mid \cdot, \mu_{k}\right)$, then it satisfies the agent's incentive compatibility and individual rationality constraints when she holds belief $\mu_{k} .{ }^{43}$ $\square$

[^26]
## Supplementary Appendix

## C Omitted proofs from Section 5

## C. 1 Repeated Mechanisms

In this section, we present the proof of Theorem 3. To do so, we first complete the formal definition of the game induced by repeating mechanism $\phi: M \times \Omega \times \mathcal{E} \rightarrow \Delta(A)$, by specifying the histories, strategy space, and the distribution over terminal histories induced by the agent's strategy and the mechanism. Having laid this groundwork, we describe the proof strategy, and then provide the formal details of the proof. Throughout this section, we use the shorthand $\tilde{\Omega}=\Omega \times \mathcal{E}$, and denote its elements by $\tilde{\omega}$.

Histories and strategies Histories through period $t \in \mathbb{N}$ are defined as $H^{t} \equiv(\Theta \times M \times A)^{t-1}$. The set of infinite histories from the agent's point of view is $H^{\infty}$. The set of terminal histories is $\mathcal{H}^{\infty} \equiv \tilde{\Omega} \times H^{\infty}$, where recall $\mathcal{E}$ is finite and endowed with some measure $\eta$.

The agent's behavioral strategy is defined as a collection $\sigma \equiv\left(\sigma_{t}\right)_{t \in \mathbb{N}}$ such that for all $t \geq 1$

$$
\sigma_{t}: H^{t} \times \Theta \rightarrow \Delta(M) .
$$

The tuple of distributions $\left(\mu_{0}, \eta, f\right)$ together with the mechanism $\phi$ and the agent's strategy $\sigma$ determine a joint distribution over $\mathcal{H}^{\infty}$ by the Ionescu-Tulcea theorem (Bogachev, 2007, Theorem 10.7.3). We provide more details on this probability distribution below. Denote by $\mathbb{P}_{\left(\mu_{0}, \eta, f, \phi, \sigma\right)}$ and $\mathbb{E}_{\left(\mu_{0}, \eta, f, \phi, \sigma\right)}$ the probability distribution over the terminal histories and the expectation with respect to this distribution, respectively. Whenever it is not likely to lead to confusion, we drop the dependence on $\left(\mu_{0}, \eta, f, \phi, \sigma\right)$, and whenever we want to emphasize the dependence on the agent's strategy we note the dependence on $\sigma$.

The distribution over terminal histories $\mathcal{H}^{\infty}$ For future use, we review the construction of $\mathbb{P}_{\sigma}$. For each $t$, the distributions $\left(\mu_{0}, \eta, f\right)$ together with the mechanism $\phi$ and the agent's strategy $\sigma$ determine a distribution over $\tilde{\Omega} \times H^{t}$, which we denote by $\mathbb{P}_{\sigma}^{t} \in \Delta\left(\tilde{\Omega} \times H^{t}\right)$. Note that for any subset $\tilde{\mathcal{H}^{t}} \subset \tilde{\Omega} \times H^{t}$,

$$
\mathbb{P}_{\sigma}^{t}\left(\tilde{\mathcal{H}}^{t}\right)=\mathbb{P}_{\sigma}^{t+1}\left(\tilde{\mathcal{H}}^{t} \times(\Theta \times M \times A)\right) .
$$

Moreover,

$$
\mathbb{P}_{\sigma}^{t+1}\left(\tilde{\omega}, h^{t}, \theta, m, a\right)=\mathbb{P}_{\sigma}^{t}\left(\tilde{\omega}, h^{t}\right) f(\theta) \sigma_{t}\left(h^{t}, \theta\right)(m) \phi(a \mid m, \tilde{\omega}) .
$$

By the Ionescu-Tulcea theorem, the distribution $\mathbb{P}_{\sigma} \in \Delta\left(\tilde{\Omega} \times H^{\infty}\right)$ is the unique distribution that satisfies that for all $t \in \mathbb{N}, \tilde{\mathcal{H}}^{t} \subset \tilde{\Omega} \times H^{t}$,

$$
\mathbb{P}_{\sigma}\left(\tilde{\mathcal{H}}^{t} \times \prod_{s=t+1}^{\infty}(\Theta \times M \times A)\right)=\mathbb{P}_{\sigma}^{t}\left(\tilde{\mathcal{H}}^{t}\right) .
$$

[^27]Belief system The agent's beliefs over $\tilde{\Omega}$ at the beginning of each $t$ are determined by the belief system, which in a slight abuse of notation we denote by $\mu_{t}: H^{t} \rightarrow \Delta(\tilde{\Omega})$. The belief system satisfies

$$
\mathbb{P}_{\sigma}^{t}\left(h^{t}\right) \mu_{t}\left(\tilde{\omega} \mid h^{t}\right)=\mathbb{P}_{\sigma}^{t}\left(\tilde{\omega}, h^{t}\right) .
$$

That is, whenever $h^{t}$ is such that $\mathbb{P}_{\sigma}\left(\left\{\tilde{h} \in H^{\infty}: \tilde{h}^{t}=h^{t}\right\}\right)>0$,

$$
\mu_{t}\left(\tilde{\omega} \mid h^{t}\right)=\frac{\mathbb{P}_{\sigma}^{t}\left(\tilde{\omega}, h^{t}\right)}{\mathbb{P}_{\sigma}^{t}\left(h^{t}\right)}=\mathbb{P}_{\sigma}^{t}\left(\tilde{\omega} \mid h^{t}\right) .
$$

Given $\mathbb{P}_{\sigma} \in \Delta\left(\mathcal{H}^{\infty}\right)$, define $\mu_{\infty}\left(\tilde{\omega} \mid h^{\infty}\right) \equiv \mathbb{P}_{\sigma}\left(\tilde{\omega} \mid h^{\infty}\right)$ to be the belief system conditional on the whole terminal history $h^{\infty}$.

Remark C. 1 (Belief system and strategies as functions on $\mathcal{H}^{\infty}$ ). Whereas the beliefs and strategies are defined on the finite histories, it is sometimes convenient to write them as functions on $\mathcal{H}^{\infty}$ that are adapted to $H^{t}$.

A property of the belief system We collect here a property of the belief system which we use in our proofs below.

Lemma C. 1 (Martingale property under weak* convergence). $\mu_{t}\left(h^{\infty}\right) \xrightarrow{w^{*}} \mu_{\infty}\left(h^{\infty}\right) \mathbb{P}_{\sigma}$-almost surely. This and the proof of other technical results are in Appendix D.

## C.1.1 Proof of Theorem 3 (necessity)

We are now ready to present the proof of Theorem 3, starting by the "only if" direction. Let $\vartheta \in \Delta(A \times$ $\Theta \times \Omega$ ) denote the outcome distribution implemented by repeated mechanism $\phi$ under best response $\sigma$, and let $v_{\sigma}$ denote the associated occupation measure, the definition of which we reproduce below for ease of reference:

$$
v_{\sigma}(a, \theta, m, \tilde{\omega})=\lim _{T \rightarrow \infty} \frac{1}{T} \mathbb{E}_{\sigma}\left[\sum_{t=1}^{T} \mathbb{1}\left[\left(a_{t}, \theta_{t}, m_{t}, \tilde{\omega}^{\prime}\right)=(a, \theta, m, \tilde{\omega})\right]\right]=\lim _{T \rightarrow \infty} v_{\sigma}^{T}(a, \theta, m, \tilde{\omega}),
$$

where recall limits are in the weak* sense. We show that $\vartheta$ can be implemented by an incentive compatible and individually rational two-stage mechanism.

To this end, we consider two sequences of extended occupation measures on $A \times \Theta \times M \times \tilde{\Omega} \times \Delta(\tilde{\Omega})$ :

$$
\begin{array}{r}
\bar{v}_{\sigma}^{T, 1}(\{(a, \theta, m, \tilde{\omega})\} \times \tilde{\Delta})=\frac{1}{T} \mathbb{E}_{\sigma}\left[\sum_{t=1}^{T} \mathbb{1}\left[\left(a_{t}, \theta_{t}, m_{t}, \tilde{\omega}^{\prime}\right)=(a, \theta, m, \tilde{\omega})\right] \mathbb{1}\left[\mu_{t} \in \tilde{\Delta}\right]\right], \\
\bar{v}_{\sigma}^{T, 2}(\{(a, \theta, m, \tilde{\omega})\} \times \tilde{\Delta})=\frac{1}{T} \mathbb{E}_{\sigma}\left[\sum_{t=1}^{T} \mathbb{1}\left[\left(a_{t}, \theta_{t}, m_{t}, \tilde{\omega}^{\prime}\right)=(a, \theta, m, \tilde{\omega})\right] \mathbb{1}\left[\mu_{t+1} \in \tilde{\Delta}\right]\right] .
\end{array}
$$

We note the following. First, Equation C. 5 counts the beliefs at the beginning of period $t$, while Equation C. 6 counts the beliefs at the end of period $t$ (after the realization of $\theta, m$, and $a$.) Equation C. 5 is key to obtain the (limit) independence of the belief and type distributions, while Equation C. 6 allows us to obtain the (limit) independence of the allocation and the state, conditional on the induced belief.

Second, $v_{\sigma}^{T}$ is the marginal of both $\bar{v}_{\sigma}^{T, 1}$ and $\bar{v}_{\sigma}^{T, 2}$. Third, by Lemma C.1, $\mu_{t} \xrightarrow{w^{*}} \mu_{\infty}$, and hence both $\bar{v}_{\sigma}^{T, 1}$ and $\bar{v}_{\sigma}^{T, 2}$ have the same set of subsequential limits, which we record for future reference below (see Appendix D for the proof):

Lemma C.2. The occupation measures $\bar{v}_{\sigma}^{T, 1}$ and $\bar{v}_{\sigma}^{T, 2}$ have the same set of subsequential limits.
The proof of necessity of Theorem 3 proceeds in five steps. First, we show that the marginal of $\bar{v}_{\sigma}^{T, 1}$ on $\Delta(\tilde{\Omega})$, which we denote by $\tau_{\sigma}^{T}$ weak*-converges to $\mathbb{P}_{\sigma} \circ \mu_{\infty}^{-1}$. We denote this limit measure by $\tau_{\sigma}$. By Lemma C.2, $\tau_{\sigma}$ is also the (limit) marginal of $\bar{v}_{\sigma}^{T, 2}$ on $\Delta(\tilde{\Omega})$.

Second, we show that up to a subsequence $\bar{v}_{\sigma}^{T, 1}, \bar{v}_{\sigma}^{T, 2} \xrightarrow{w^{*}} \bar{v}_{\sigma}$. Furthermore, transition probabilities $\tau_{\sigma} \in \Delta(\Delta(\tilde{\Omega})), \rho: \Theta \times \Delta(\tilde{\Omega}) \rightarrow \Delta(M), \alpha^{\prime}: M \times \Delta(\tilde{\Omega}) \rightarrow \Delta(A)$ exist such that

$$
v_{\sigma}(a, \theta, m, \tilde{\omega})=\int_{\Delta(\tilde{\Omega})} \mu(\tilde{\omega}) f(\theta) \rho(m \mid \theta, \mu) \alpha^{\prime}(a \mid m, \mu) \tau_{\sigma}(d \mu)
$$

Hence, the agent's payoff when faced with mechanism $\phi$ and playing strategy $\sigma$ can be written as:

$$
\mathbb{E}_{\bar{v}_{\sigma}}[u(a, \theta, \omega)]=\int_{\Delta(\tilde{\Omega})} \sum_{\theta \in \Theta} f(\theta) \sum_{m \in M} \rho(m \mid \theta, \mu) \mathbb{E}_{\tilde{\omega} \sim \mu}\left[\sum_{a \in A} \alpha^{\prime}(a \mid \mu, m) u(a, \theta, \omega)\right] \tau_{\sigma}(d \mu) .
$$

Third, we show that for all $\theta \in \Theta$

$$
\mathbb{E}_{\tau_{\sigma}}\left\{\sum_{m \in M} \rho(m \mid \theta, \mu) \mathbb{E}_{\tilde{\omega} \sim \mu}\left[\sum_{a \in A} \alpha^{\prime}(a \mid \mu, m) u(a, \theta, \omega)\right]-\max _{m \in M} \mathbb{E}_{\tilde{\omega} \sim \mu}\left[\sum_{a \in A} \alpha^{\prime}(a \mid \mu, m) u(a, \theta, \omega)\right]\right\}=0 .
$$

Equations C. 8 and C. 9 allow us to identify the incentive compatible and individually rational allocation rule of the two-stage mechanism that implements $\vartheta$.

Fourth, whereas the previous steps identify a two-stage mechanism expressed in terms of posterior beliefs over $\tilde{\Omega}$, we show how to obtain a two-stage mechanism expressed in terms of posterior beliefs over $\Omega$. Finally, we show that the agent's payoff in Equation C. 8 coincides with the payoff she would get when best responding to the information structure calibrated to $\phi$.

Step 1 Having defined the extended occupation measure in Equation C.5, we present here a property we use in our proof. Let $\tau_{\sigma}^{T}$ denote the marginal of $\bar{v}_{\sigma}^{T, 1}$ on $\Delta(\tilde{\Omega})$. That is, for any measurable subset $\tilde{\Delta} \subset \Delta(\tilde{\Omega})$, define

$$
\tau_{\sigma}^{T}(\tilde{\Delta})=\frac{1}{T} \mathbb{E}_{\sigma}\left[\sum_{t=1}^{T} \mathbb{1}\left[\mu_{t} \in \tilde{\Delta}\right]\right] .
$$

In Appendix D, we prove the following:
Lemma C.3. The sequence of measures $\left(\tau_{\sigma}^{T}\right)_{T \in \mathbb{N}}$ defined by Equation C. 10 converges in the weak* sense to the push-forward measure $\tau_{\sigma} \equiv \mathbb{P}_{\sigma} \circ \mu_{\infty}^{-1}$, where $\mu_{\infty}\left(h^{\infty}\right)=\mathbb{P}_{\sigma}\left(\cdot \mid h^{\infty}\right)$.

Step 2 To show that Equation C. 7 holds, we show the following properties of $\bar{v}_{\sigma}^{T, 1}$ and $\bar{v}_{\sigma}^{T, 2}$. On the one hand, $\bar{v}_{\sigma}^{T, 1}$ satisfies that for all $g \in C_{b}(A \times \Theta \times M \times \tilde{\Omega} \times \Delta(\tilde{\Omega}))$,

$$
\int_{A \times \Theta \times M \times \tilde{\Omega} \times \Delta(\tilde{\Omega})} g(a, \theta, m, \tilde{\omega}, \mu) d \bar{v}_{\sigma}^{T, 1}=\int_{\Theta \times M \times \Delta(\tilde{\Omega})} \mathbb{E}_{\mu}\left[\mathbb{E}_{\phi(\cdot \mid m, \tilde{\omega})}[g(a, \theta, m, \tilde{\omega}, \mu)]\right] d \bar{v}_{\sigma, \Theta M \Delta^{\prime}}^{T, 1}
$$

and for all $q \in C_{b}(\Theta \times \Delta(\Omega))$,

$$
\int_{\Theta \times \Delta(\tilde{\Omega})} q(\theta, \mu) d \bar{v}_{\sigma, \Theta \Delta}^{T, 1}=\int_{\Delta(\tilde{\Omega})} \int_{\Theta} f(\theta) q(\theta, \mu) d \tau_{\sigma}^{T}
$$

On the other hand, $\bar{v}_{\sigma}^{T, 2}$ satisfies that for all $g \in C_{b}(A \times \Theta \times M \times \tilde{\Omega} \times \Delta(\tilde{\Omega}))$,

$$
\int_{A \times \Theta \times M \times \tilde{\Omega} \times \Delta(\tilde{\Omega})} g(a, \theta, m, \tilde{\omega}, \mu) d \bar{v}_{\sigma}^{T, 2}=\int_{\Theta \times M \times A \times \Delta(\tilde{\Omega})} \mathbb{E}_{\tilde{\omega} \sim \mu}[g(a, \theta, m, \tilde{\omega}, \mu)] d \bar{v}_{\sigma, \Theta M A \Delta}^{T, 2}
$$

In the expressions above, the subscripts on $\bar{v}_{\sigma}^{T, k}$ next to $\sigma$ are the spaces over which we take the marginals, and $\Delta$ is shorthand notation for $\Delta(\tilde{\Omega})$. Because $\Delta(A \times \Theta \times M \times \tilde{\Omega} \times \Delta(\tilde{\Omega}))$ is compact (Aliprantis and Border, 2006, Theorem 15.11), $\bar{v}_{\sigma}^{T, 1}$ has a convergent subsequence $\left(\bar{v}_{\sigma}^{T_{n}, 1}\right)_{n \in \mathbb{N}}$, which by Lemma C. 2 is also a convergent subsequence of $\bar{v}_{\sigma}^{T, 2}$. Let $\bar{v}_{\sigma}$ denote the weak* limit along $T_{n}$. The continuity of the projection implies that $v_{\sigma}$ is the marginal of $\bar{v}_{\sigma}$ on $A \times \Theta \times M \times \tilde{\Omega}$, and $\tau_{\sigma} \equiv \mathbb{P}_{\sigma} \circ \mu_{\infty}^{-1}$ is the marginal on $\Delta(\Omega)$. Equations C. 12 and C. 13 together imply that $\bar{v}_{\sigma}$ admits the decomposition in the right hand side of Equation C.7, and the result follows.

To show Equation C. 11 holds, use that $\bar{v}_{\sigma}^{T, 1}$ has finite support to write it as follows:

$$
\begin{aligned}
\bar{v}_{\sigma}^{T, 1}(a, \theta, m, \tilde{\omega}, \mu) & =\frac{1}{T} \sum_{t=1}^{T} \sum_{h^{t} \in H^{t}} \mathbb{P}_{\sigma}^{t}\left(\tilde{\omega}, h^{t}\right) f(\theta) \sigma_{t}\left(h^{t}, \theta\right)(m) \phi(a \mid m, \tilde{\omega}) \mathbb{1}\left[\mu_{t}=\mu\right] \\
& =\frac{1}{T} \sum_{t=1}^{T} \sum_{h^{t} \in H^{t}} \mathbb{P}_{\sigma}^{t}\left(h^{t}\right) \mu(\tilde{\omega}) f(\theta) \sigma_{t}\left(h^{t}, \theta\right)(m) \phi(a \mid m, \tilde{\omega}) \mathbb{1}\left[\mu_{t}=\mu\right]
\end{aligned}
$$

where the second equality uses that $\mu_{t}\left(h^{t}\right)=\mathbb{P}_{\sigma}^{t}\left(\cdot \mid h^{t}\right)$.
Equation C. 14 implies the following holds for every bounded continuous function $g \in C_{b}(A \times \Theta \times M \times$ $\tilde{\Omega} \times \Delta(\tilde{\Omega})):$

$$
\mathbb{E}_{\bar{v}_{\sigma}^{T, 1}}[g(a, \theta, m, \tilde{\omega}, \mu)]=\mathbb{E}_{\bar{v}_{\sigma, \Theta M \Delta}^{T, 1}}\left[\mathbb{E}_{\mu}\left[\mathbb{E}_{\phi(\cdot \mid m, \tilde{\omega})}[g(a, \theta, m, \tilde{\omega}, \mu)]\right]\right],
$$

This completes the proof that Equation C. 11 holds. Letting $V_{g}(\theta, M, \mu)=\mathbb{E}_{\mu}\left[\mathbb{E}_{\phi(\cdot \mid m, \tilde{\omega})}[g(a, \theta, m, \tilde{\omega}, \mu)]\right]$, we have that

$$
\int_{A \times \Theta \times M \times \tilde{\Omega} \times \Delta(\tilde{\Omega})} g d \bar{v}_{\sigma}^{T, 1}=\int_{\Theta \times M \times \Delta(\tilde{\Omega})} V_{g} d \bar{v}_{\sigma, \Theta M \Delta}^{T, 1} \Leftrightarrow \int_{A \times \Theta \times M \times \tilde{\Omega} \times \Delta(\tilde{\Omega})}\left(g-V_{g}\right) d \bar{v}_{\sigma}^{T, 1}=0 .
$$

Because $g-V_{g} \in C_{b}(A \times \Theta \times M \times \tilde{\Omega} \times \Delta(\tilde{\Omega}))$ and $\bar{v}_{\sigma}^{T_{n}, 1} \xrightarrow{w^{*}} \bar{v}_{\sigma}$, we conclude that

$$
\mathbb{E}_{\bar{v}_{\sigma}}[g(a, \theta, m, \tilde{\omega}, \mu)]=\int_{\Theta \times M \times \Delta(\tilde{\Omega})} \mathbb{E}_{\mu}\left[\mathbb{E}_{\phi(\cdot \mid m, \tilde{\omega})}[g(a, \theta, m, \tilde{\omega}, \mu)]\right] d \bar{v}_{\sigma, \Theta M \Delta} .
$$

To show that Equation C. 12 holds, note that the marginal of $\bar{v}_{\sigma}^{T, 1}$ on $\Theta \times \Delta(\tilde{\Omega})$ equals $\tau_{\sigma}^{T} \otimes f$. Indeed, fix
any continuous function $q \in C_{b}(\Theta \times \Delta(\tilde{\Omega}))$ and note that for all $T$

$$
\mathbb{E}_{\bar{v}_{\sigma, \Theta \Delta}^{T, 1}}[q(\theta, \mu)]=\mathbb{E}_{\tau_{\sigma}^{T}}\left[\sum_{\theta \in \Theta} f(\theta) q(\theta, \mu)\right] .
$$

Letting $V_{q}(\mu)=\sum_{\theta \in \Theta} f(\theta) q(\theta, \mu)$, we have that for all $T$

$$
\int_{\Delta(\tilde{\Omega}) \times \Theta}\left(q(\theta, \mu)-V_{q}(\mu)\right) d \bar{v}_{\sigma, \Theta \Delta}^{T, 1}=0 .
$$

Because $q-V_{q} \in C_{b}(\Theta \times \Delta(\tilde{\Omega}))$ and $\bar{v}_{\sigma}^{T_{n}, 1} \xrightarrow{w^{*}} \bar{v}_{\sigma}$, we conclude that

$$
\int_{\Theta \times \Delta(\tilde{\Omega})} q(\theta, \mu) d \bar{v}_{\sigma, \Theta \Delta}=\int_{\Delta(\tilde{\Omega})} \int_{\Theta} f(\theta) q(\theta, \mu) d \tau_{\sigma} .
$$

Lastly, to show that Equation C. 13 holds, note that we can write $\bar{v}_{\sigma}^{T, 2}$ as follows (once again, we use that for finite $T$, it has finite support):

$$
\begin{aligned}
& \bar{v}_{\sigma}^{T, 2}(a, \theta, m, \tilde{\omega}, \mu)=\frac{1}{T} \sum_{t=1}^{T} \sum_{h^{t} \in H^{t}} \mathbb{P}_{\sigma}^{t}\left(h^{t}\right) \mathbb{P}_{\sigma}^{t}\left(\tilde{\omega} \mid h^{t}\right) f(\theta) \sigma_{t}\left(h^{t}, \theta\right)(m) \phi(a \mid m, \tilde{\omega}) \mathbb{1}\left[\mu_{t+1}\left(h^{t}, \theta, m, a\right)=\mu\right]= \\
& =\frac{1}{T} \sum_{t=1}^{T} \sum_{h^{t} \in H^{t}: \mu_{t+1}\left(h^{t}, \theta, m, a\right)=\mu} \mathbb{P}_{\sigma}^{t+1}\left(\tilde{\omega}, h^{t}, \theta, m, a\right)=\frac{1}{T} \sum_{t=1}^{T} \sum_{h^{t} \in H^{t}: \mu_{t+1}\left(h^{t}, \theta, m, a\right)=\mu} \mu(\tilde{\omega}) \mathbb{P}_{\sigma}^{t+1}\left(h^{t}, \theta, m, a\right) \\
& =\mu(\tilde{\omega}) \frac{1}{T} \sum_{t=1}^{T} \sum_{h^{t} \in H^{t}: \mu_{t+1}\left(h^{t}, \theta, m, a\right)=\mu} \mathbb{P}_{\sigma}^{t+1}\left(h^{t}, \theta, m, a\right)=\mu(\tilde{\omega}) \bar{v}_{\sigma, A \Theta M \Delta}^{T, 2}(a, \theta, m, \mu),
\end{aligned}
$$

where the last expression follows from noting that the term multiplying $\mu(\omega)$ in the first expression in the third line is $\sum_{\tilde{\omega}} v_{\sigma}^{T, 2}(a, \theta, m, \tilde{\omega}, \mu)$.

Then, for every $g \in C_{b}(A \times \Theta \times M \times \tilde{\Omega} \times \Delta(\tilde{\Omega}))$, we have that

$$
\mathbb{E}_{\bar{v}_{\sigma}^{T, 2}}[g(a, \theta, m, \tilde{\omega}, \mu)]=\mathbb{E}_{\bar{v}_{\sigma, A \Theta M \Delta(\tilde{\Omega})}^{T, 2}}\left[\mathbb{E}_{\tilde{\omega} \sim \mu}[g(a, \theta, m, \tilde{\omega}, \mu)]\right]
$$

which completes the proof that Equation C. 13 holds. Letting $V_{g}(a, \theta, m, \mu)=\mathbb{E}_{\tilde{\omega} \sim \mu}[g(a, \theta, m, \tilde{\omega}, \mu)]$, we have that

$$
\int\left(g-V_{g}\right) d \bar{v}_{\sigma}^{T, 2}=0
$$

Because $g-V_{g} \in C_{b}(A \times \Theta \times M \times \tilde{\Omega} \times \Delta(\tilde{\Omega}))$ and $\bar{v}_{\sigma}^{T_{n}, 2} \xrightarrow{w^{*}} \bar{v}_{\sigma}$, we conclude that

$$
\mathbb{E}_{\bar{v}_{\sigma}}[g(a, \theta, m, \tilde{\omega}, \mu)]=\int_{A \times \Theta \times M \times \Delta(\tilde{\Omega})} \mathbb{E}_{\tilde{\omega} \sim \mu}[g(a, \theta, m, \tilde{\omega}, \mu)] \bar{v}_{\sigma, A \Theta M \Delta}(d(a, \theta, m, \mu)) .
$$

Equations C. 16 and C. 17 imply $\bar{v}_{\sigma}$ admits the following disintegration:

$$
\bar{v}_{\sigma}(\{(a, \theta, m, \tilde{\omega})\} \times \tilde{\Delta})=\int_{\tilde{\Delta}} \mu(\tilde{\omega}) f(\theta) \alpha^{\prime}(a \mid \theta, m, \mu) \rho(m \mid \theta, \mu) \tau_{\sigma}(d \mu)
$$

where we disintegrated $\bar{v}_{\sigma, A \Theta M \Delta}$ first along $\Theta \times \Delta(\tilde{\Omega})$-and used Equation C. 16 to obtain the independence
of $\Theta$ and $\Delta(\tilde{\Omega})$-and then further disintegrated the distribution of $A \times M$ conditional on $\Theta \times \Delta(\tilde{\Omega})$. Now, Equation C. 15 implies that the following also holds

$$
\bar{v}_{\sigma}(\{(a, \theta, m, \tilde{\omega})\} \times \tilde{\Delta})=\int_{\tilde{\Delta}} \mu(\tilde{\omega}) f(\theta) \phi(a \mid m, \tilde{\omega}) \rho(m \mid \theta, \mu) \tau_{\sigma}(d \mu)
$$

where once again we use the uniqueness of disintegration. Because Equations C. 18 and C. 19 hold for any tuple $(a, \theta, m, \tilde{\omega})$ and measurable subset $\tilde{\Delta}$ of $\Delta(\tilde{\Omega})$, we conclude that (i) $\alpha^{\prime}(a \mid \theta, m, \mu)$ does not depend on $\theta \tau_{\sigma}$-almost everywhere, and (ii) $\phi(\cdot \mid m, \tilde{\omega})$ is constant on $\tilde{\omega}$ in the support of $\mu \tau_{\sigma}$-almost everywhere. This concludes the proof of Step 2.

Step 3 We now argue that the agent achieves

$$
u^{*}(\mu) \equiv \sum_{\theta \in \Theta} f(\theta) \max _{m \in M} \sum_{\tilde{\omega} \in \tilde{\Omega}} \mu(\tilde{\omega}) \sum_{a \in A} \phi(a \mid m, \tilde{\omega}) u(a, \theta, \omega)=\sum_{\theta \in \Theta} f(\theta) \max _{m \in M} \sum_{\tilde{\omega} \in \tilde{\Omega}} \mu(\tilde{\omega}) \sum_{a \in A} \alpha^{\prime}(a \mid m, \mu) u(a, \theta, \omega),
$$

on the support of $\tau_{\sigma}$, where the second equality follows from Step 2. Toward a contradiction, suppose this is not the case; that is,

$$
\mathbb{E}_{\tau_{\sigma}}\left[\sum_{\theta \in \Theta} f(\theta) \sum_{m \in M} \rho(m \mid \theta, \mu) \sum_{\tilde{\omega} \in \tilde{\Omega}} \mu(\tilde{\omega}) \sum_{a} \phi(a \mid m, \tilde{\omega}) u(a, \theta, \omega)\right]<\mathbb{E}_{\tau_{\sigma}}\left[\sum_{\theta \in \Theta} u^{*}(\mu)\right]=U^{*} .
$$

We show that the agent can achieve a payoff arbitrarily close to $U^{*}$ by playing according to $\sigma$ until some finite $T$ and then best-responding to her beliefs at time $T$ in every period thereafter; a contradiction.

Consider a strategy $\sigma^{\prime}$ which until some period $T$ plays according to $\sigma$ and after period $T$ best responds to $\mu_{T}\left(h^{T}\right) \in \Delta(\tilde{\Omega})$. Because payoffs accumulated on a finite number of periods are irrelevant to longrun payoffs, this strategy results in a payoff:

$$
\begin{aligned}
& \sum_{h^{T} \in H^{T}} \mathbb{P}_{\sigma}^{T}\left(h^{T}\right) \sum_{\theta \in \Theta} f(\theta) \max _{m \in M}\left[\sum_{\tilde{\omega} \in \tilde{\Omega}} \mu_{T}(\tilde{\omega}) \sum_{a \in A} \phi(a \mid m, \tilde{\omega}) u(a, \theta, \omega)\right]=\sum_{h^{T} \in H^{T}} \mathbb{P}_{\sigma}^{T}\left(h^{T}\right) u^{*}\left(\mu_{T}\left(h^{T}\right)\right) \\
& =\mathbb{E}_{\mathbb{P}_{\sigma}^{T} \circ \mu_{T}^{-1}}\left[u^{*}(\mu)\right]=\mathbb{E}_{\mathbb{P}_{\sigma} \circ \mu_{T}^{-1}}\left[u^{*}(\mu)\right],
\end{aligned}
$$

where the last equality follows as $\mu_{T}$ is adapted to the histories through $T$. Similar arguments to Lemma C. 3 imply that $\mathbb{P}_{\sigma} \circ \mu_{T}^{-1} \xrightarrow{w^{*}} \mathbb{P}_{\sigma} \circ \mu_{\infty}^{-1} \equiv \tau_{\sigma}$. Noting that $u^{*}: \Delta(\tilde{\Omega}) \rightarrow \mathbb{R}$ is continuous and bounded (as it is the maximum of linear functions in beliefs), we obtain that as $T \rightarrow \infty$,

$$
\mathbb{E}_{\mathbb{P}_{\sigma} \circ \mu_{T}^{-1}}\left[u^{*}(\mu)\right] \rightarrow \mathbb{E}_{\tau_{\sigma}}\left[u^{*}(\mu)\right] .
$$

It follows that for every $\delta>0$, we can find $T$ large enough so that $\left|\mathbb{E}_{\mathbb{P}_{\sigma^{\circ}} \mu_{T}^{-1}}\left[u^{*}(\mu)\right]-U^{*}\right|<\delta$, contradicting the optimality of $\sigma$.

We conclude that $\alpha: \Theta \times \Delta(\tilde{\Omega}) \rightarrow \Delta(A)$ defined as follows:

$$
\alpha(a \mid \theta, \mu)=\sum_{m \in M} \rho(m \mid \theta, \mu) \alpha^{\prime}(a \mid m, \mu),
$$

is incentive compatible and individually rational $\tau_{\sigma}$-almost everywhere.

Step 4 We now show how to derive a two-stage mechanism $\beta^{*}: \Omega \rightarrow \Delta(\Delta(\Omega))$ and an allocation rule $\alpha^{*}: \Theta \times \Delta(\Omega) \rightarrow \Delta(A)$ that implement $\vartheta$. First, note that the agent's payoff when her type is $\theta$ and the induced belief is $\mu \in \Delta(\tilde{\Omega})$, can be written as

$$
\sum_{\tilde{\omega} \in \tilde{\Omega}} \mu(\tilde{\omega}) \sum_{a \in A} \alpha(a \mid \theta, \mu) u(a, \theta, \omega)=\sum_{\omega \in \Omega} \mu_{\Omega}(\omega) \sum_{a \in A} \alpha(a \mid \theta, \mu) u(a, \theta, \omega),
$$

where $\mu_{\Omega}$ is the marginal of $\mu$ on $\Omega$ and the equality follows because the realization of $\varepsilon$ is payoffirrelevant. By Step 3, $\alpha(\cdot \mid \theta, \mu)$ is individually rational and incentive compatible when the agent holds belief $\mu_{\Omega}$.

Furthermore, for each $(a, \theta, \omega) \in A \times \Theta \times \Omega$, we have

$$
\vartheta(a, \theta, \omega)=\sum_{\varepsilon \in \mathcal{E}} \int_{\Delta(\tilde{\Omega})} \mu(\omega, \varepsilon) f(\theta) \alpha(a \mid \theta, \mu) \tau_{\sigma}(d \mu)=\int_{\Delta(\tilde{\Omega})} \mu_{\Omega}(\omega) f(\theta) \alpha(a \mid \theta, \mu) \tau_{\sigma}(d \mu)
$$

For each $\theta \in \Theta$, consider the joint distribution $Q_{\theta} \in \Delta(A \times \Delta(\Omega))$ defined as follows:

$$
Q_{\theta}(\{a\} \times \tilde{\Delta})=\int_{\Delta(\tilde{\Omega})} \mathbb{1}\left[\mu_{\Omega} \in \tilde{\Delta}\right] \alpha(a \mid \theta, \mu) \tau_{\sigma}(d \mu)=\int_{\tilde{\Delta}} \alpha^{*}\left(a \mid \theta, \mu_{\Omega}\right) \tau^{*}\left(d \mu_{\Omega}\right)
$$

where the third equality follows from disintegration (note $\alpha^{*}\left(\cdot \mid \cdot, \mu_{\Omega}\right)=\mathbb{E}\left[\alpha(\cdot \mid \cdot, \tilde{\mu}) \mid \tilde{\mu}_{\Omega}=\mu_{\Omega}\right]$ ). By the first argument in Step 4, $\alpha^{*}\left(\cdot \mid \cdot, \mu_{\Omega}\right)$ is individually rational and incentive compatible when the agent holds $\mu_{\Omega}$. We obtain that

$$
\vartheta(a, \theta, \omega)=\int_{\Delta(\Omega)} \mu_{\Omega}(\omega) f(\theta) \alpha^{*}\left(a \mid \theta, \mu_{\Omega}\right) \tau^{*}\left(d \mu_{\Omega}\right)
$$

Defining for all $\omega \in \Omega$ and measurable subsets $\tilde{\Delta} \in \Delta(\Omega)$,

$$
\beta^{*}(\tilde{\Delta} \mid \omega)=\int_{\tilde{\Delta}} \frac{\mu(\omega)}{\mu_{0}(\omega)} \tau^{*}(d \mu)
$$

Steps 2-4 together imply that $\vartheta$ can be implemented by the incentive compatible and individually rational two-stage mechanism $\left(\beta^{*}, \alpha^{*}\right)$.

Step 5: The agent adequately learns Finally, we argue that the agent earns the same payoff as if she had access to the information structure calibrated to $\phi, \pi_{\phi}$.

Lemma C.4. Let $\tau_{\phi}$ denote the belief distribution induced by the calibrated information structure $\pi_{\phi}$. Then, the agent's payoff under $\sigma$ equals

$$
U\left(\tau_{\phi}\right) \equiv \mathbb{E}_{\tau_{\phi}}\left[\sum_{\theta \in \Theta} f(\theta) \max _{m \in M} \sum_{\tilde{\omega} \in \tilde{\Omega}} \mu(\tilde{\omega}) \sum_{a \in A} \phi(a \mid m, \tilde{\omega}) u(a, \theta, \omega)\right] .
$$

The proof of this is standard, and hence we defer it to Appendix D.

## C.1.2 Proof of Theorem 3 (sufficiency)

Suppose $\vartheta \in \Delta(A \times \Theta \times \Omega)$ is implemented by an incentive compatible and individually rational twostage mechanism. That is,

$$
\vartheta(a, \theta, \omega)=\mu_{0}(\omega) f(\theta) \int_{\Delta(\Omega)} \alpha(a \mid \theta, \mu) \beta(d \mu \mid \omega)
$$

and $\alpha$ is incentive compatible and individually rational on the support of $\mu_{0} \otimes \beta$. As in the proof of Theorem 2, a finite support $\beta^{\prime}: \Omega \rightarrow \Delta\left(\left\{\mu_{1}, \ldots, \mu_{K}\right\}\right)$ exists such that $\left(\beta^{\prime}, \alpha\right)$ implement $\vartheta$. As in Green and Stokey (2022), the experiment $\beta^{\prime}$ can be generated by a finite information structure $\pi: \Omega \times \mathcal{E} \rightarrow$ $\Delta(\Omega)$, where (i) $\mathcal{E}$ is finite, (ii) $\mathcal{E}$ is independent of $\Omega$, and (iii) $\mu=\pi(\omega, \varepsilon)$.

Construct a mechanism $\phi: \Theta \times \Omega \times \mathcal{E} \rightarrow \Delta(A)$ such that $\phi(\cdot \mid \theta, \omega, \varepsilon)=\alpha(\cdot \mid \theta, \pi(\omega, \varepsilon))$. (Note that $\pi$ is information structure calibrated to $\phi$, but expressed in beliefs.) Consider now the extensive form game induced by such a mechanism. ${ }^{44}$

If the agent truthfully reports her type, then the occupation measure induces outcome distribution $\vartheta$. Hence, under truthtelling, the agent's payoff is:

$$
\begin{aligned}
& U\left(\sigma_{\text {truth }}\right)=\sum_{(a, \theta, \omega) \in A \times \Theta \times \Omega} \vartheta(a, \theta, \omega) u(a, \theta, \omega)=\mathbb{E}_{\tau_{\phi}}\left[\sum_{\theta \in \Theta} f(\theta) \sum_{a \in A} \alpha(a \mid \theta, \mu) \sum_{\omega \in \Omega} \mu(\omega) u(a, \theta, \omega)\right] \\
& =\mathbb{E}_{\tau_{\phi}}\left[\sum_{\theta \in \Theta} f(\theta) \max \left\{\max _{\theta^{\prime} \in \Theta} \sum_{a \in A} \alpha\left(a \mid \theta^{\prime}, \mu\right) \sum_{\omega \in \Omega} \mu(\omega) u(a, \theta, \omega), \sum_{\omega \in \Omega} \mu(\omega) u\left(a_{\varnothing}, \theta, \omega\right)\right\}\right]= \\
& =\mathbb{E}_{\mathcal{E}}\left[\sum_{\theta \in \Theta} f(\theta) \max \left\{\max _{\theta^{\prime} \in \Theta} \sum_{\omega \in \Omega} \mu(\omega) \sum_{a \in A} \phi\left(a \mid \theta^{\prime}, \omega, \varepsilon\right) u(a, \theta, \omega), \sum_{\omega \in \Omega} \mu(\omega) u\left(a_{\varnothing}, \theta, \omega\right)\right\}\right],
\end{aligned}
$$

where (i) $\tau_{\phi}$ is the belief distribution induced by the information structure $\pi$, and (ii) the first equality is by definition of the occupation measure, the second is the definition that $\vartheta$ is implemented by the two-stage mechanism, the third follows from incentive compatibility and individual rationality of $\alpha$, and the fourth is definitional.

Moreover, the payoff in the last line of Equation C. 23 is the payoff the agent obtains by using the "learning" strategy in Lemma C.4, which first extracts all the mechanism can teach her about the state and then uses that information to optimize over her participation and reporting strategies. It follows that truthtelling (and participation) are optimal and $\vartheta$ is implemented by repeated mechanism $\phi$.

## C. 2 Dynamic Mechanisms

In this section, we present the proof of Theorem 4. To do so, we first complete the formal definition of the game, by specifying the histories, strategy space, and the distribution over terminal histories induced by the agent's strategy and the mechanism. Having laid this groundwork, we describe the proof strategy, and then provide the formal details of the proof.

[^28]Mechanisms, histories, and strategies A dynamic mechanism $\left(\varphi_{t}\right)_{t \in \mathbb{N}}$ is a sequence of mappings that condition on the state, the agent's report history, the allocation history, and today's report and output an allocation. By the revelation principle, it is without loss of generality to restrict attention to mechanisms that solicit type reports.

As in the main text, we expand the set of type reports and allocations by the non-participation decision and the outside option, which we denote by $\Theta A_{\varnothing} \equiv A \times \Theta \cup\left\{\left(\varnothing, a_{\varnothing}\right)\right\}$. Then, $\hat{H}^{t}=\left(\Theta A_{\varnothing}\right)^{t-1}$ denotes the histories of reports (inclusive of the non-participation decision) and allocations at the beginning of time $t \in \mathbb{N}$, and let $\hat{\mathcal{H}}^{t}=\Omega \times \hat{H}^{t}$. Similarly, let $\hat{H}^{\infty}=\times_{t \in \mathbb{N}}\left(\Theta A_{\varnothing}\right)$ denote the set of all possible report-allocation outcome paths, and let $\hat{\mathcal{H}}^{\infty}=\Omega \times \hat{H}^{\infty}$. A dynamic mechanism is then a collection of mappings $\left(\varphi_{t}\right)_{t \in \mathbb{N}}$ such that $\varphi_{t}: \hat{\mathcal{H}}^{t} \times \Theta \rightarrow \Delta(A)$.

To define the agent's strategy, let $H^{t}=\Theta^{t-1} \times \hat{H}^{t-1}$, where the coordinates denote the sequence of realized types, reports (inclusive of participation decisions), and allocations through period $t-1$. A behavioral strategy is a mapping $\left(p_{t}, \sigma_{t}\right): H^{t} \times \Theta \rightarrow[0,1] \times \Delta(\Theta)$.

The distribution over terminal histories $\mathcal{H}^{\infty}$ To obtain the complete description of the paths on the tree we need to append $\Omega$ to $H^{t}$; hence the paths through period $t-1$ are $\Omega \times H^{t} \equiv \mathcal{H}^{t}$. The distributions over states, agent's types, the agent's strategy, and the mechanism induce a distribution over the terminal histories $\mathcal{H}^{\infty} \equiv \Omega \times H^{\infty}$, which we denote by $\mathbb{P}_{(p, \sigma)} \in \Delta\left(\Omega \times H^{\infty}\right)$. We denote by $\mathbb{E}_{(p, \sigma)}$ the expectation under this measure. The distribution $\mathbb{P}_{(p, \sigma)} \in \Delta\left(\Omega \times H^{\infty}\right)$ is the unique distribution that satisfies that for all $t \in \mathbb{N}, \tilde{\mathcal{H}}^{t} \subset \Omega \times \mathcal{E} \times H^{t}$,

$$
\mathbb{P}_{(p, \sigma)}\left(\tilde{\mathcal{H}}^{t} \times \prod_{s=t+1}^{\infty}\left(\Theta \times \Theta A_{\varnothing}\right)=\mathbb{P}_{\sigma}^{t}\left(\tilde{\mathcal{H}}^{t}\right),\right.
$$

where the distributions $\left(\mathbb{P}_{(p, \sigma)}^{t}\right)_{t \in \mathbb{N}}$ satisfy (under participation and truthtelling)

$$
\mathbb{P}_{(p, \sigma)}^{t+1}\left(\omega, h^{t}, \theta, \theta^{\prime}, a\right)=\mathbb{P}_{(p, \sigma)}^{t}\left(\omega, h^{t}\right) f(\theta) \mathbb{1}\left[\theta^{\prime}=\theta\right] \varphi_{t}\left(a \mid \omega, \hat{h}^{t}, \theta^{\prime}\right) .
$$

Implementation We focus on incentive-compatible mechanisms $\varphi$ for which (i) a best response, ( $p, \sigma$ ), exists, and (ii) the occupation measure $v_{\sigma} \in \Delta(A \times \Theta \times \Omega)$ exists, where

$$
v_{(p, \sigma)}(a, \theta, \omega)=\lim _{T \rightarrow \infty} \frac{1}{T} \mathbb{E}_{(p, \sigma)}\left[\sum_{t=1}^{T} \mathbb{1}\left[\left(a_{t}, \theta_{t}, \omega^{\prime}\right)=(a, \theta, \omega)\right]\right],
$$

where the limit is in the weak* sense. In contrast to Appendix C.1, we do not keep track of the agent's type reports in the occupation measure, only the agent's types. Under $(p, \sigma)$ only truthtelling histories have positive probability.

## C.2.1 Proof of Theorem 4 (necessity)

Let $\vartheta \in \Delta(A \times \Theta \times \Omega)$ denote the outcome distribution implemented by an incentive compatible dynamic mechanism $\varphi$, and let $v_{(p, \sigma)}$ denote the occupation measure under the agent's truthtelling strategy. Below, we show that $v_{(p, \sigma)}$, and hence $\vartheta$, can be implemented by a two-stage mechanism which lacks profitable undetectable deviations and is ex ante individually rational.

Analogously to the proof of Theorem 3, we define two sequences of extended occupation measures on $A \times \Theta \times \Omega \times \Delta(\Omega)$ defined as follows. Letting $\tilde{\Delta}$ denote a measurable subset of $\Delta(\Omega)$, define

$$
\begin{aligned}
& \bar{v}_{(p, \sigma)}^{T, 1}(\{(a, \theta, \omega)\} \times \tilde{\Delta})=\frac{1}{T} \sum_{t=1}^{T} \sum_{h^{t} \in H^{t}} \mathbb{P}_{(p, \sigma)}^{t}\left(\omega, h^{t}\right) f(\theta) \sigma_{t}\left(h^{t}, \theta\right)(\theta) \varphi\left(\omega, \hat{h}^{t}, \theta\right)(a) \mathbb{1}\left[\mu_{t}\left(h^{t}\right) \in \tilde{\Delta}\right] \\
& \bar{v}_{(p, \sigma)}^{T, 2}(\{(a, \theta, \omega)\} \times \tilde{\Delta})=\frac{1}{T} \sum_{t=1}^{T} \sum_{h^{t} \in H^{t}} \mathbb{P}_{(p, \sigma)}^{t}\left(\omega, h^{t}\right) f(\theta) \sigma_{t}\left(h^{t}, \theta\right)(\theta) \varphi\left(\omega, \hat{h}^{t}, \theta\right)(a) \mathbb{1}\left[\mu_{t+1}\left(h^{t}, \theta, \theta, a\right) \in \tilde{\Delta}\right]
\end{aligned}
$$

The proof proceeds similarly to that in Appendix C.1. First, we show that the occupation measure $v_{(p, \sigma)} \in \Delta(A \times \Theta \times \Omega)$ admits the following decomposition

$$
v_{(p, \sigma)}(a, \theta, \omega)=\int_{\Delta(\Omega)} f(\theta) \mu(\omega) \alpha(a \mid \theta, \mu) \tau_{(p, \sigma)}(d \mu)
$$

where $\tau_{(p, \sigma)}$ is the distribution over terminal beliefs (cf. Lemma C.3) and the transition probability $\alpha: \Theta \times \Delta(\Omega) \rightarrow \Delta(A)$ is our candidate allocation rule. Consequently, the agent's equilibrium payoff can be written as follows:

$$
\sum_{(a, \theta, \omega) \in A \times \Theta \times \Omega} v_{(p, \sigma)}(a, \theta, \omega) u(a, \theta, \omega)=\int_{\Delta(\Omega)}\left[\sum_{\theta \in \Theta} f(\theta) \sum_{\omega \in \Omega} \mu(\omega) \sum_{a \in A} \alpha(a \mid \theta, \mu) u(a, \theta, \omega)\right] \tau_{(p, \sigma)}(d \mu) .
$$

Second, we show that the allocation rule lacks profitable undetectable deviations and is ex ante individually rational.

The occupation measure satisfies Equation C. 27 To prove that Equation C. 27 holds, we first show that for all $g \in C_{b}(A \times \Theta \times \Omega \times \Delta(\Omega))$ and all $T \in \mathbb{N}$,

$$
\int_{A \times \Theta \times \Omega \times \Delta(\Omega)} g(a, \theta, \omega, \mu) d \bar{v}_{(p, \sigma)}^{T, 2}=\int_{A \times \Theta \times \Delta(\Omega)} \mathbb{E}_{\mu}[g(a, \theta, \omega, \mu)] d \bar{v}_{(p, \sigma), A \Theta \Delta(\Omega)}^{T, 2}
$$

and for all $q \in C_{b}(\Theta \times \Delta(\Omega))$ and all $T \in \mathbb{N}$,

$$
\int_{\Theta \times \Delta(\Omega)} q(\theta, \mu) d \bar{v}_{(p, \sigma), \Theta \Delta}^{T, 1}=\int_{\Delta(\Omega)} \sum_{\theta \in \Theta} f(\theta) q(\theta, \mu) d \bar{v}_{(p, \sigma), \Delta}^{T, 1}
$$

where the subscripts on $\bar{v}$ next to ( $p, \sigma$ ) are the spaces over which we take the marginals, and $\Delta$ is shorthand notation for $\Delta(\Omega)$. We skip the proof of this step as it basically repeats the proof of the analogous step in Appendix C.1.

Because $\Delta(A \times \Theta \times \Omega \times \Delta(\Omega))$ is compact (Aliprantis and Border, 2006, Theorem 15.11), $\bar{v}_{(p, \sigma)}^{T, 1}$ has a convergent subsequence $\left(\bar{v}_{(p, \sigma)}^{T_{n}, 1}\right)_{n \in \mathbb{N}}$, which by Lemma C. 2 is also a convergent subsequence of $\bar{v}_{(p, \sigma)}^{T, 2}$. Let $\bar{v}_{(p, \sigma)}$ denote the weak* limit along $T_{n}$. The continuity of the projection implies that $v_{(p, \sigma)}$ is the marginal of $\bar{v}_{(p, \sigma)}$ on $A \times \Theta \times \Omega$, and $\tau_{(p, \sigma)} \equiv \mathbb{P}_{\sigma} \circ \mu_{\infty}^{-1}$ is the marginal on $\Delta(\Omega)$. Moreover, Equation C. 29 and Equation C. 30 together imply that $v_{(p, \sigma)}$ admits the decomposition on the right hand side of Equation C.27, and the result follows.

The allocation rule lacks profitable undetectable deviations We now show the allocation rule $\alpha$ admits no profitable undetectable deviations. An undetectable deviation is a transition probability $\sigma^{\prime}$ from $\Theta \times \Delta(\Omega)$ to $\Delta(\Theta)$ such that for all $\mu \in \Delta(\Omega)$ and $\theta^{\prime} \in \Theta$

$$
\sum_{\theta \in \Theta} f(\theta) \sigma^{\prime}\left(\theta^{\prime} \mid \theta, \mu\right)=f\left(\theta^{\prime}\right) .
$$

Consider a deviation by the agent to $\left(p, \sigma^{\prime}\right)$ instead of $(p, \sigma)$. That is, when his type is $\theta$ and belief is $\mu$, the agent chooses type $\theta^{\prime}$ with probability $\sigma^{\prime}\left(\theta^{\prime} \mid \theta, \mu\right)$. In what follows, we index the induced distributions over histories only by $\sigma$ and $\sigma^{\prime}$ as we are only changing the agent's reporting strategy. In particular, denote by $\mathbb{P}_{\sigma^{\prime}}$ the induced probability distribution over terminal histories when the agent uses $\left(p, \sigma^{\prime}\right)$ instead of $(p, \sigma)$.

We first claim that for every $t$ the marginal of $\mathbb{P}_{\sigma^{\prime}}^{t}$ over $\Omega \times \hat{H}^{t}$ coincides with that of $\mathbb{P}_{\sigma}^{t}$. Recall that for every $t$ we have that

$$
\mathbb{P}_{\sigma^{\prime}}^{t+1}\left(\omega, h^{t}, \theta, \theta^{\prime}, a\right)=\mathbb{P}_{\sigma^{\prime}}^{t}\left(\omega, h^{t}\right) f(\theta) \sigma^{\prime}\left(\theta^{\prime} \mid \mu_{t}\left(h^{t}\right), \theta\right) \varphi_{t}\left(a \mid \omega, \hat{h}^{t}, \theta^{\prime}\right) .
$$

Adding up over $\theta$ on both sides and using Equation C.31, we get:

$$
\sum_{\theta \in \Theta} \mathbb{P}_{\sigma^{\prime}}^{t+1}\left(\omega, h^{t}, \theta, \theta^{\prime}, a\right)=\mathbb{P}_{\sigma^{\prime}}^{t}\left(\omega, h^{t}\right) f\left(\theta^{\prime}\right) \varphi_{t}\left(a \mid \omega, \hat{h}^{t}, \theta^{\prime}\right) .
$$

Now, note that $h^{t}=\left(\hat{h}^{t}, \tilde{\theta}^{t-1}\right)$ for some sequence $\tilde{\theta}^{t-1} \in \Theta^{t-1}$. If we add up on both sides over all such sequences we get

$$
\sum_{\theta \in \Theta, \tilde{\theta}^{t-1} \in \Theta^{t-1}} \mathbb{P}_{\sigma^{\prime}}^{t+1}\left(\omega, \hat{h}^{t}, \tilde{\theta}^{t-1}, \theta, \theta^{\prime}, a\right)=\sum_{\tilde{\theta}^{t-1} \in \Theta^{t-1}} \mathbb{P}_{\sigma^{\prime}}^{t}\left(\omega, \hat{h}^{t}, \tilde{\theta}^{t-1}\right) f\left(\theta^{\prime}\right) \varphi_{t}\left(a \mid \omega, \hat{h}^{t}, \theta^{\prime}\right) .
$$

Note that if the distribution over $\Omega \times \hat{H}^{t}$ induced by $\sigma^{\prime}$ up to period $t$ is the same as that induced by $\sigma$, we get that the right-hand side equals:

$$
\mathbb{P}_{\sigma, \hat{\mathcal{H}^{t}}}^{t}\left(\omega, \hat{h}^{t}\right) f\left(\theta^{\prime}\right) \varphi_{t}\left(a \mid \omega, \hat{h}^{t}, \theta^{\prime}\right),
$$

and hence $\mathbb{P}_{\sigma^{\prime}, \hat{\mathcal{H}}^{t+1}}^{t+1}\left(\omega, \hat{h}^{t}, \theta^{\prime}, a\right)=\mathbb{P}_{\sigma, \hat{\mathcal{H}}^{t}}^{t}\left(\omega, \hat{h}^{t}\right) f\left(\theta^{\prime}\right) \varphi_{t}\left(a \mid \omega, \hat{h}^{t}, \theta^{\prime}\right)=\mathbb{P}_{\sigma, \hat{\mathcal{H}}^{t+1}}^{t+1}\left(\omega, \hat{h}^{t}, \theta^{\prime}, a\right)$. By definition of $\mathbb{P}_{\sigma^{\prime}}$, we conclude that $\mathbb{P}_{\sigma^{\prime}, \hat{\mathcal{H}}^{\infty}}=\mathbb{P}_{\sigma, \hat{\mathcal{H}}^{\infty}}$. Hence, the joint distribution over states, reports, and allocations is the same under $\sigma$ and $\sigma^{\prime}$.

Let $\bar{v}_{\sigma^{\prime}}^{T, 1}, \bar{v}_{\sigma^{\prime}}^{T, 2} \in \Delta(A \times \Theta \times \hat{\Theta} \times \Omega \times \Delta(\Omega))$ denote the analogue of the occupation measures in Equations C. 25 and C. 26 corresponding to $\sigma^{\prime}$, extended to account for the agent's reports. Below, the notation $\hat{\Theta}$ signifies those are the agent's reports. In what follows, recalling that the belief system depends only on the reported history and not the type history is useful. Equation C. 31 implies that for all measurable
subsets $\tilde{\Delta}$ of $\Delta(\Omega)$,

$$
\begin{aligned}
& \sum_{\theta \in \Theta} \bar{v}_{\sigma^{\prime}}^{T, 1}\left(\left\{\left(a, \theta, \theta^{\prime}, \omega\right)\right\} \times \tilde{\Delta}\right)=\frac{1}{T} \sum_{t=1}^{T} \sum_{h^{t} \in H^{t}} \mathbb{P}_{\sigma^{\prime}}^{t}\left(\omega, h^{t}\right) f\left(\theta^{\prime}\right) \varphi_{t}\left(\omega, \hat{h}^{t}, \theta^{\prime}\right)(a) \mathbb{1}\left[\mu_{t}\left(h^{t}\right) \in \tilde{\Delta}\right] \\
& =\frac{1}{T} \sum_{t=1}^{T} \sum_{\hat{h}^{t} \in \hat{H}^{t}} \mathbb{P}_{\sigma^{\prime}}^{t}\left(\omega, \hat{h}^{t}\right) f\left(\theta^{\prime}\right) \varphi_{t}\left(\omega, \hat{h}^{t}, \theta^{\prime}\right)(a) \mathbb{1}\left[\mu_{t}\left(\hat{h}^{t}\right) \in \tilde{\Delta}\right] \\
& =\frac{1}{T} \sum_{t=1}^{T} \sum_{\hat{h}^{t} \in \hat{H}^{t}} \mathbb{P}_{\sigma}^{t}\left(\omega, \hat{h}^{t}\right) f\left(\theta^{\prime}\right) \varphi_{t}\left(\omega, \hat{h}^{t}, \theta^{\prime}\right)(a) \mathbb{1}\left[\mu_{t}\left(\hat{h}^{t}\right) \in \tilde{\Delta}\right]=\bar{v}_{\sigma}^{T, 1}\left(\left\{\left(a, \theta^{\prime}, \omega\right)\right\} \times \tilde{\Delta}\right) .
\end{aligned}
$$

The first equality uses the definition of undetectability, the second uses that all the terms depend only on the reported history, the third uses that $\sigma$ and $\sigma^{\prime}$ induce the same distribution over states, reports, and allocations, and the last is the definition of the occupation measure induced by $\sigma$. In words, the marginal of $\bar{v}_{\sigma^{\prime}}^{T, 1}$ over allocations, reports, states, and beliefs, $\bar{v}_{\sigma^{\prime}, A \hat{\Theta} \Omega \Delta}^{T, 1}$ coincides with $\bar{v}_{\sigma}^{T, 1}$.

We now show that $\bar{v}_{\sigma^{\prime}}^{T, 1}$ and $\bar{v}_{\sigma^{\prime}}^{T, 2}$ have a convergent subsequence with limit $\bar{v}_{\sigma^{\prime}} \in \Delta(A \times \Theta \times \hat{\Theta} \times \Omega \times \Delta(\Omega))$ that admits the following decomposition:

$$
\mathbb{E}_{\bar{v}_{\sigma^{\prime}}}[u(a, \theta, \omega)]=\int_{\Delta(\Omega)}\left[\sum_{\theta \in \Theta} f(\theta) \sum_{\theta^{\prime} \in \Theta} \sigma^{\prime}(\theta, \mu)\left(\theta^{\prime}\right) \sum_{a} \alpha\left(a \mid \theta^{\prime}, \mu\right) u(a, \theta, \mu)\right] d \tau_{(p, \sigma)},
$$

where $u(a, \theta, \mu)$ is the linear extension of $u(a, \theta, \cdot)$.
We proceed as follows: First, we show that for each $T$, under $v_{\sigma^{\prime}}^{T, 1}$, the allocation is independent of the true type conditional on the period- $t$ belief and the reported type. Indeed,

$$
\begin{aligned}
\sum_{\omega \in \Omega} v_{\sigma^{\prime}}^{T, 1}\left(a, \theta, \theta^{\prime}, \omega, \mu\right) & =\frac{1}{T} \sum_{t=1}^{T} \sum_{h^{t} \in H^{t}} \mathbb{P}_{\sigma^{\prime}}\left(h^{t}\right)\left(\sum_{\omega \in \Omega} \mathbb{P}_{\sigma^{\prime}}\left(\omega \mid h^{t}\right) \varphi_{t}\left(\omega, \hat{h}^{t}, \theta^{\prime}\right)(a)\right) f(\theta) \sigma^{\prime}\left(\mu_{t}\left(h^{t}\right), \theta\right)\left(\theta^{\prime}\right) \mathbb{1}\left[\mu_{t}\left(h^{t}\right)=\mu\right] \\
& =\frac{f(\theta) \sigma^{\prime}(\theta, \mu)\left(\theta^{\prime}\right)}{f\left(\theta^{\prime}\right)}\left(\frac{1}{T} \sum_{t=1}^{T} \sum_{\hat{h}^{t}: \mu_{t}\left(\hat{h}^{t}\right)=\mu} \mathbb{P}_{\sigma^{\prime}}^{t}\left(\hat{h}^{t}\right) \sum_{\omega \in \Omega} \mathbb{P}_{\sigma^{\prime}}\left(\omega \mid \hat{h}^{t}\right) f\left(\theta^{\prime}\right) \varphi_{t}\left(\hat{h}^{t}, \theta^{\prime}\right)(a)\right) \\
& =\frac{f(\theta) \sigma^{\prime}(\theta, \mu)\left(\theta^{\prime}\right)}{f\left(\theta^{\prime}\right)} v_{\sigma^{\prime}, A \hat{\Theta} \Delta}^{T, 1}\left(a, \theta^{\prime}, \mu\right)=\frac{f(\theta) \sigma^{\prime}(\theta, \mu)\left(\theta^{\prime}\right)}{f\left(\theta^{\prime}\right)} v_{\sigma, A \Theta \Delta}^{T, 1}\left(a, \theta^{\prime}, \mu\right),
\end{aligned}
$$

where the third and fourth equalities use Equation C.25. Moreover, the same analysis as that under $\sigma$ implies the agent's true type is independent of the belief.

Therefore, $\bar{v}_{\sigma^{\prime}, A \Theta \Theta \leq}^{T, 1}$ admits decomposition:

$$
\bar{v}_{\sigma^{\prime}, A \Theta \hat{\Theta} \Delta}^{T, 1}\left(a, \theta, \theta^{\prime}, \mu\right)=\frac{f(\theta) \sigma^{\prime}\left(\theta^{\prime} \mid \theta, \mu\right)}{f\left(\theta^{\prime}\right)} \bar{v}_{\sigma^{\prime}, A \hat{\Theta} \Delta}^{T, 1}\left(a, \theta^{\prime}, \mu\right)=\frac{f(\theta) \sigma^{\prime}\left(\theta^{\prime} \mid \theta, \mu\right)}{f\left(\theta^{\prime}\right)} \bar{v}_{\sigma, A \Theta \Delta}^{T, 1}\left(a, \theta^{\prime}, \mu\right),
$$

where the second equality follows from Equation C.32.
Second, by the same arguments as in Appendix C.1, $\bar{v}_{\sigma^{\prime}}^{T, 2}\left(a, \theta, \theta^{\prime}, \omega, \mu\right)$ admits decomposition $\mu(\omega) \bar{v}_{\sigma^{\prime}, A \Theta \Theta \leq}^{T, 2}\left(a, \theta, \theta^{\prime}, \mu\right)$ for each $T$.

Third, convergent subsequences $v_{\sigma^{\prime}}^{T_{n_{m}}, 1}$ and $v_{\sigma^{\prime}}^{T_{n_{m}}, 2}$ exist with limit $\bar{v}_{\sigma^{\prime}}$ (cf. Lemma C.2). ${ }^{45}$ We note two

[^29]things. On the one hand, because our previous arguments show that the set of measures admitting the above decompositions is closed, the limit $\bar{v}_{\sigma^{\prime}}$ admits the decomposition. That is,
$$
\int_{A \times \Theta \times \hat{\Theta} \times \Omega \times \Delta(\Omega)} u(a, \theta, \omega) \bar{v}_{\sigma^{\prime}}\left(d\left(a, \theta, \theta^{\prime}, \omega, \mu\right)\right)=\int_{A \times \hat{\Theta} \times \Delta(\Omega)} \sum_{\theta \in \Theta} \frac{f(\theta) \sigma^{\prime}\left(\theta^{\prime} \mid \theta, \mu\right)}{f\left(\theta^{\prime}\right)}\left(\sum_{\omega \in \Omega} \mu(\omega) u(a, \theta, \omega)\right) d \bar{v}_{\sigma^{\prime}, A \hat{\Theta} \Delta}
$$
On the other hand, because $T_{n_{m}}$ is a subsequence of $T_{n}$ and $\bar{v}_{\sigma^{\prime}, A \hat{\Theta} \Omega \Delta}^{T, 1}=\bar{v}_{\sigma, A \Theta \Omega \Delta}^{T, 1}$ and $\bar{v}_{\sigma, A \Theta \Omega \Delta}^{T_{n}, 1} \xrightarrow{w^{*}} \bar{v}_{\sigma}$, we can conclude that $\bar{v}_{\sigma^{\prime}, A \hat{\Theta} \Delta}=\bar{v}_{\sigma, A \Theta \Delta}$ and admits the same decomposition as $\bar{v}_{\sigma}$. We conclude that
$$
\mathbb{E}_{\bar{v}_{\sigma^{\prime}}}[u(a, \theta, \omega)]=\int_{\Delta(\Omega)}\left[\sum_{\theta, \theta^{\prime} \in \Theta} f(\theta) \sigma^{\prime}\left(\theta^{\prime} \mid \theta, \mu\right) \sum_{a} \alpha\left(a \mid \theta^{\prime}, \mu\right)\left(\sum_{\omega \in \Omega} \mu(\omega) u(a, \theta, \omega)\right)\right] d \tau_{(p, \sigma)}
$$
Consequently,
$$
\begin{aligned}
\lim \sup _{T \rightarrow \infty} \mathbb{E}_{\sigma^{\prime}}\left[U_{T}\right] \geq \lim _{m \rightarrow \infty} \mathbb{E}_{\sigma^{\prime}}\left[U_{T_{n_{m}}}\right] & =\mathbb{E}_{\bar{v}_{\sigma^{\prime}}}[u(a, \theta, \omega)] \\
& =\mathbb{E}_{\tau_{(p, \sigma)}}\left[\sum_{\theta, \theta^{\prime}, a} f(\theta) \sigma^{\prime}(\theta, \mu)\left(\theta^{\prime}\right) \alpha\left(a \mid \theta^{\prime}, \mu\right) u(a, \theta, \mu)\right],
\end{aligned}
$$
where $u(a, \theta, \mu)$ is the linear extension of $u(a, \theta, \cdot)$. Because $\sigma$ is a best response, we have that
$$
\mathbb{E}_{\tau_{(p, \sigma)}}\left[\sum_{\theta, a} f(\theta) \alpha(a \mid \theta, \mu) u(a, \theta, \mu)\right] \geq \mathbb{E}_{\tau_{(p, \sigma)}}\left[\sum_{\theta, \theta^{\prime}, a} f(\theta) \sigma^{\prime}(\theta, \mu)\left(\theta^{\prime}\right) \alpha\left(a \mid \theta^{\prime}, \mu\right) u(a, \theta, \mu)\right],
$$
which implies the two-stage mechanism lacks profitable undetectable deviations.

The allocation rule is ex ante individually rational Define

$$
U_{\mathrm{net}}(\mu)=\sum_{\theta \in \Theta} f(\theta) \sum_{\omega \in \Omega} \mu(\omega)\left[\sum_{a \in A} \alpha(a \mid \theta, \mu) u(a, \theta, \omega)-u\left(a_{\varnothing}, \theta, \omega\right)\right],
$$

to be the agent's (ex ante) payoff net of the outside option at belief $\mu$. Ex ante individual rationality of $\alpha$ is equivalent to $U_{\text {net }}(\mu) \geq 0$ for all $\mu$ in the support of $\tau_{(p, \sigma)}$.

Toward a contradiction, assume that $U_{\text {net }}(\mu)<0$ with positive probability under $\tau_{(p, \sigma)}$. By Lemma D. 1

[^30]in Appendix D, a set $B \subset \Delta(\Omega)$ open relative to $\Delta(\Omega)$ exists such that ${ }^{46}$
$$
\int_{B} U_{\text {net }}(\mu) \tau_{(p, \sigma)}(d \mu)<0
$$
Moreover, we can pick $B$ such that $\tau_{(p, \sigma)}(\partial B)=0$, where $\partial B$ denotes the boundary of $B$ relative to $\Delta(\Omega) .{ }^{47}$ Lastly, let $\delta>0$ be such that
$$
\int_{B} U_{\text {net }}(\mu) \tau_{(p, \sigma)}(d \mu) \leq-2 \delta
$$
Let $\left(\mu_{t}\left(h^{t}\right)\right)_{t \in \mathbb{N}, h^{t} \in H^{t}}$ denote the belief process under $(p, \sigma)$. For $L \in \mathbb{N}$, define a strategy $\left(p^{L}, \sigma^{L}\right)$ as follows:


1. $\left(p_{t}^{L}\left(h^{t}, \cdot\right), \sigma_{t}^{L}\left(h^{t}, \cdot\right)\right)=\left(p_{t}\left(h^{t}, \cdot\right), \sigma_{t}\left(h^{t}, \cdot\right)\right)$ if either $t<L$ OR $\left(t \geq L\right.$ and $\left.\mu_{L}\left(h^{L}\right) \notin B\right)$, where $h^{L}$ precedes $h^{t}$,
2. Otherwise, $\left(p_{t}^{L}\left(h^{t}, \cdot\right), \sigma_{t}^{L}\left(h^{t}, \cdot\right)\right)=\left(0, \sigma_{t}\left(h^{t}, \cdot\right)\right)$ (note that when the agent quits the strategy can be specified arbitrarily.)


Note the agent's average payoff through period $T$ under ( $p^{L}, \sigma^{L}$ ) can be written as follows:
$$
\mathbb{E}_{\left(p^{L}, \sigma^{L}\right)}\left[U_{T}\right]=\mathbb{E}_{(p, \sigma)}\left[U_{T}\right]-\mathbb{E}_{(p, \sigma)}\left[\frac{1}{T} \sum_{t=1}^{T}\left(u\left(a_{t}, \theta_{t}, \omega\right)-u\left(a_{\varnothing}, \theta_{t}, \omega\right)\right) \mathbb{1}\left[t \geq L \text { and } \mu_{L} \in B\right]\right] .
$$
We show that for sufficiently large $L$, $\left(p^{L}, \sigma^{L}\right)$ is a profitable deviation. For $T \geq L$, write
$$
\begin{aligned}
& \mathbb{E}_{(p, \sigma)}\left[\frac{1}{T} \sum_{t=L}^{T}\left(u\left(a_{t}, \theta_{t}, \omega\right)-u\left(a_{\varnothing}, \theta_{t}, \omega\right)\right) \mathbb{1}\left[\mu_{L} \in B\right]\right]-\int_{B} U_{\text {net }}(\mu) \tau_{(p, \sigma)}(d \mu)= \\
& =\mathbb{E}_{(p, \sigma)}\left[\frac{1}{T} \sum_{t=1}^{T}\left(u\left(a_{t}, \theta_{t}, \omega\right)-u\left(a_{\varnothing}, \theta_{t}, \omega\right)\right) \mathbb{1}\left[\mu_{t} \in B\right]\right]-\int_{B} U_{\text {net }}(\mu) \tau_{(p, \sigma)}(d \mu) \\
& -\mathbb{E}_{(p, \sigma)}\left[\frac{1}{T} \sum_{t=1}^{L-1}\left(u\left(a_{t}, \theta_{t}, \omega\right)-u\left(a_{\varnothing}, \theta_{t}, \omega\right)\right) \mathbb{1}\left[\mu_{t} \in B\right]\right] \\
& +\mathbb{E}_{(p, \sigma)}\left[\frac{1}{T} \sum_{t=L}^{T}\left(u\left(a_{t}, \theta_{t}, \omega\right)-u\left(a_{\varnothing}, \theta_{t}, \omega\right)\right)\left(\mathbb{1}\left[\mu_{L} \in B\right]-\mathbb{1}\left[\mu_{t} \in B\right]\right)\right] .
\end{aligned}
$$
Let $K=\max _{\theta, \omega, a}\left|\left(u(a, \theta, \omega)-u\left(a_{\varnothing}, \theta, \omega\right)\right)\right|$, and note that we can bound the term in the last line of Equation C. 35 as follows:
$$
\left|\mathbb{E}_{(p, \sigma)}\left[\frac{1}{T} \sum_{t=L}^{T}\left(u\left(a_{t}, \theta_{t}, \omega\right)-u\left(a_{\varnothing}, \theta_{t}, \omega\right)\right)\left(\mathbb{1}\left[\mu_{L} \in B\right]-\mathbb{1}\left[\mu_{t} \in B\right]\right)\right]\right| \leq K \mathbb{E}_{(p, \sigma)}\left[\frac{1}{T} \sum_{t=L}^{T}\left|\mathbb{1}\left[\mu_{L} \in B\right]-\mathbb{1}\left[\mu_{t} \in B\right]\right| .\right]
$$

[^31]Because $\mu_{t} \xrightarrow{w^{*}} \mu_{\infty} \mathbb{P}_{(p, \sigma)}$-a.s. (Lemma C.1) and $\tau_{(p, \sigma)}(\partial B)=0$, we conclude: ${ }^{48}$

$$
\lim _{L \rightarrow \infty} \sup _{t \geq L}\left|\mathbb{1}\left[\mu_{L} \in B\right]-\mathbb{1}\left[\mu_{t} \in B\right]\right|=0 \mathbb{P}_{(p, \sigma)} \text {-a.s. }
$$

Then,

$$
\mathbb{E}_{(p, \sigma)}\left[\frac{1}{T} \sum_{t=L}^{T}\left|\mathbb{1}\left[\mu_{L} \in B\right]-\mathbb{1}\left[\mu_{t} \in B\right]\right|\right] \leq \mathbb{E}_{(p, \sigma)}\left[\sup _{t \geq L}\left|\mathbb{1}\left[\mu_{L} \in B\right]-\mathbb{1}\left[\mu_{t} \in B\right]\right|\right]
$$

and choose $\bar{L}$ large enough so that for all $L \geq \bar{L}$, we have that:

$$
\mathbb{E}_{(p, \sigma)}\left[\sup _{t \geq L}\left|\mathbb{1}\left[\mu_{L} \in B\right]-\mathbb{1}\left[\mu_{t} \in B\right]\right|\right] \leq \delta / K .
$$

Consider now the term in the third line of Equation C. 35 and note that it is bounded in absolute value by $K(L-1) / T$, which tends to 0 as $T \rightarrow \infty$. Similarly, the term in the second line of Equation C. 35 vanishes as $T \rightarrow \infty$. ${ }^{49}$ Thus, for $L \geq \bar{L}$, we can find $\bar{T}$ such that for all $T \geq \bar{T}^{50}$

$$
\left|\mathbb{E}_{(p, \sigma)}\left[\frac{1}{T} \sum_{t=L}^{T}\left(u\left(a_{t}, \theta_{t}, \omega\right)-u\left(a_{\varnothing}, \theta_{t}, \omega\right)\right) \mathbb{1}\left[\mu_{L} \in B\right]\right]-\int_{B} U_{\mathrm{net}}(\mu) \tau_{(p, \sigma)}(d \mu)\right| \leq \frac{3}{2} \delta,
$$

and hence

$$
\mathbb{E}_{(p, \sigma)}\left[\frac{1}{T} \sum_{t=L}^{T}\left(u\left(a_{t}, \theta_{t}, \omega\right)-u\left(a_{\varnothing}, \theta_{t}, \omega\right)\right) \mathbb{1}\left[\mu_{L} \in B\right]\right] \leq-\frac{1}{2} \delta .
$$

We conclude that

$$
\lim \sup _{T \rightarrow \infty} \mathbb{E}_{\left(p^{L}, \sigma^{L}\right)}\left[U_{T}\right] \geq \lim _{T \rightarrow \infty} \mathbb{E}_{(p, \sigma)}\left[U_{T}\right]+\frac{1}{2} \delta,
$$

a contradiction.

## C.2.2 Proof of Theorem 4 (sufficiency)

We now show that all outcome distributions $\vartheta \in \Delta(A \times \Theta \times \Omega)$ that admit the decomposition in Theorem 4 can be implemented via a dynamic mechanism. To this end, let $\tau$ and $\alpha: \Theta \times \Delta(\Omega) \rightarrow \Delta(A)$

[^32]denote the belief distribution and the ex ante individually rational allocation rule without profitable undetectable deviations corresponding to $\vartheta$. That is,
$$
\vartheta(a, \theta, \omega)=\int_{\Delta(\Omega)} \mu(\omega) f(\theta) \alpha(a \mid \theta, \mu) \tau(d \mu)
$$
The proof proceeds as follows:


1. We first consider a fictitious setting in which there is no state uncertainty and we are given an allocation rule $\alpha^{\prime}: \Theta \rightarrow \Delta(A)$ that is ex ante individually rational and lacks profitable undetectable deviations for some utility function $u^{\prime}: A \times \Theta \rightarrow \mathbb{R}$. Proposition C. 1 shows that a dynamic mechanism exists that implements $\alpha^{\prime}$.
2. We then show that if $\vartheta$ satisfies Equation C.39, then a finite support belief distribution $\tau^{\prime}$ exists such that $\vartheta$ and $\alpha$ satisfies Equation C. 39 with $\tau^{\prime}$ instead of $\tau$.
3. Lastly, we use this result to construct a dynamic game that implements $\vartheta$.


Step 1 For this step, we consider a fictitious setting in which there is no state uncertainty and the designer faces a privately informed agent with payoffs $u^{\prime}: A \times \Theta \rightarrow \mathbb{R}$, where $\theta \sim f \in \Delta(\Theta) .{ }^{51}$

Suppose we are given an allocation rule $\alpha^{\prime}: \Theta \rightarrow \Delta(A)$ that admits no profitable undetectable deviations relative to $u^{\prime}$ as in Definition 7 and is individually rational as in Definition 8. We have the following result:

Proposition C.1. Let $\vartheta^{\prime}=f(\theta) \alpha^{\prime}(a \mid \theta) \in \Delta(A \times \Theta)$ such that $\alpha^{\prime}$ lacks profitable undetectable deviations and is ex ante individually rational. Then, a dynamic mechanism exists that implements $\vartheta^{\prime}$.

Proof of Proposition C.1. The proof is constructive. We build on the analysis of Margaria and Smolin (2018) and present a dynamic mechanism that alternates between communication and adjustment phases. In all phases, the mechanism selects allocations using reports $\theta^{\prime}$ according to $\alpha^{\prime}$. In a communication phase, the reports are those sent by the agent. In an adjustment phase, the agent's reports are disregarded; instead, the mechanism simulates reports to guarantee that the occupation measure over reports coincides with $f$ and these simulated reports are used to determine the allocation. The mechanism ensures that under any agent's strategy, the occupation measure over reports and allocations exists and equals $\vartheta^{\prime}$; thus, any strategy corresponds to an undetectable deviation. The length of communication phases grows in time. Thus, under truthtelling the relative length of adjustment phases vanishes in time, and the expected occupation measure over types and allocations exists and equals $\vartheta^{\prime}$. Because $\alpha^{\prime}$ lacks profitable undetectable deviations, it follows that truthtelling is optimal for the agent. Because $\alpha^{\prime}$ is ex ante individually rational, it follows that the participation constraints are satisfied.

Formally, the mechanism consists of sequential blocks, each block starting with a communication phase followed by an adjustment phase. The lengths of communication phases are fixed at $L_{1}, L_{2}, \ldots$

[^33]such that $L_{n} \rightarrow \infty$ and $L_{n} / \sum_{k \leq n} L_{k} \rightarrow 0$, e.g., $L_{n}=n$. The length of adjustment phase $N_{n}$ depends on the agent's reports in the communication phase in block $n$. Denote by $T_{n}$ the first period of block $n$, which is the first period of the corresponding communication phase. The first period of the corresponding adjustment phase is $T_{n}+L_{n}+1$. Denote by freq ${ }_{n}^{1}$ the average report frequencies in this block at the beginning of the adjustment stage:
$$
\operatorname{freq}_{n}^{1}(\hat{\theta}) \triangleq \frac{1}{L_{n}} \sum_{t=T_{n}}^{T_{n}+L_{n}-1} 1\left(\hat{\theta}_{t}=\hat{\theta}\right) .
$$
If $\operatorname{freq}_{n}^{1}=f$, then the adjustment phase is empty, and the mechanism proceeds to the next block. Otherwise, in the adjustment phase, the mechanism generates reports over $N_{n}$ periods to guarantee that at the end of the adjustment phase the expected frequency of reports in this block equals $f$ that is,
$$
\mathbb{E}\left[\text { freq }_{n}^{2} \mid \text { freq }_{n}^{1}\right]=f
$$
where
$$
\operatorname{freq}_{n}^{2}(\hat{\theta}) \triangleq \frac{1}{L_{n}+N_{n}} \sum_{t=T_{n}}^{T_{n}+L_{n}+N_{n}-1} 1\left(\hat{\theta}_{t}=\hat{\theta}\right) .
$$
To do so, denote by $\eta \triangleq \min _{\theta} f(\theta)$ and observe that $f \in \Delta(\Theta)$ can be surrounded by a ball of radius $\eta$ within the simplex $\Delta(\Theta)$. The adjustment phase lasts for $N_{n}$ periods where: ${ }^{52}$
$$
N_{n}=\left\lceil L_{n} \frac{\left\|\mathrm{freq}_{n}^{1}-f\right\|_{\infty}}{\eta}\right\rceil,
$$
and in each period of the adjustment phase the mechanism generates the reports i.i.d. according to $\tilde{f}_{n}^{a}$ :
$$
f_{n}^{a}=f-\left(\text { freq }_{n}^{1}-f\right) \frac{L_{n}}{N_{n}} .
$$
The construction ensures that $f_{n}^{a} \in \Delta(\Theta)$, because $\left\|f_{n}^{a}-f\right\|_{\infty} \leq \eta$, and that (C.41) holds, because
$$
\mathbb{E}\left[\operatorname{freq}_{n}^{2} \mid \operatorname{freq}_{n}^{1}\right]=\frac{1}{L_{n}+N_{n}}\left(L_{n} \operatorname{freq}_{n}^{1}+N_{n} f_{n}^{a}\right)=f
$$
This in turn guarantees that the long-run distribution of reports (generated jointly by the agent and the mechanism) exists and equals $f$ irrespectively of the agent's strategy. Intuitively, the fact that each block becomes negligible relative to past history over time ensures the agent's reports in each block have less and less effect on the long run frequency of reports, whereas the adjustment phase ensures that the frequency of reports converges to $f$. Formally, for any history and $T$ denote by $n^{\text {last }}(T)$ the number of the block to which $T$ belongs and by $T^{\text {last }}(T)$ the first period of that block. Observe that for

[^34]any agent's strategy:
$$
N_{n} \leq L_{n}\left(\max _{f^{\prime}} \frac{\left\|f^{\prime}-f\right\|_{\infty}}{\eta}+1\right) \triangleq L_{n} \bar{\rho}
$$
Therefore,
$$
\frac{\left|T-T^{\text {last }}(T)\right|}{T^{\text {last }}(T)} \leq \frac{L_{n^{\text {last }}(T)}(1+\bar{\rho})}{\sum_{k<n^{\text {last }}(T)} L_{k}} \xrightarrow[T \rightarrow \infty]{\text { a.s. }} 0,
$$
where the limit result holds because $n^{\text {last }}(T) \xrightarrow[T \rightarrow \infty]{\text { a.s. }} \infty$ and $L_{n} / \sum_{k \leq n} L_{k} \xrightarrow[n \rightarrow \infty]{ } 0$.
Then, for any agent's strategy, for any $\hat{\theta} \in \Theta$,
$$
\begin{aligned}
\lim _{T \rightarrow \infty} \frac{1}{T} \sum_{t=1}^{T} \operatorname{Pr}\left(\hat{\theta}_{t}=\hat{\theta}\right) & =\lim _{T \rightarrow \infty} \mathbb{E}\left[\frac{f(\hat{\theta}) T^{\text {last }}(T)+f^{\text {last }}(T)\left(T-T^{\text {last }}(T)\right)}{T^{\text {last }}(T)+T-T^{\text {last }}(T)}\right] \\
& =\lim _{T \rightarrow \infty} \mathbb{E}\left[\frac{f(\hat{\theta})+f^{\text {last }}(T)\left(T-T^{\text {last }}(T) / T^{\text {last }}(T)\right.}{1+\left(T-T^{\text {last }}(T)\right) / T^{\text {last }}(T)}\right]=f(\hat{\theta}),
\end{aligned}
$$
where $f^{\text {last }}(T) \in \Delta(\Theta)$ is the report frequency in the last block up to period $T$, and the last line follows from Equation C.46.

Since the mechanism chooses allocations in all periods according to $\alpha^{\prime}$, it follows that for any agent's strategy $\sigma^{\prime}$ the induced occupation measure over allocations and type reports satisfies:

$$
\lim _{T \rightarrow \infty} \frac{1}{T} \mathbb{E}_{\sigma^{\prime}}\left[\sum_{t=1}^{T} \mathbb{1}\left[\left(a_{t}, \hat{\theta}_{t}\right)=(a, \hat{\theta})\right]\right]=f(\hat{\theta}) \alpha^{\prime}(a \mid \hat{\theta})=\vartheta^{\prime}(a, \hat{\theta}) .
$$

In other words, for any reporting strategy the occupation measure over allocations and reports exists.
We now show that under truthtelling the occupation measure over types and allocations exists and equals $f(\theta) \alpha^{\prime}(a \mid \theta)=\vartheta(a, \theta)$. To this end, assume that the agent always reports her true type. For any $T$, denote by $\tilde{L}^{\text {total }}(T)$ the total number of periods spent in communication phases before $T$ and by $\tilde{N}^{\text {total }}(T)$ the total number of periods spent in adjustment phases before $T$. Observe that by the strong law of large numbers, because $L_{n} \rightarrow \infty$,

$$
\frac{N_{n}}{L_{n}} \leq \frac{\left\|\operatorname{freq}_{n}^{1}-f\right\|_{\infty}}{\eta}+\frac{1}{L_{n}} \xrightarrow[n \rightarrow \infty]{\text { a.s. }} 0 .
$$

Therefore,

$$
\frac{\tilde{N}^{\text {total }}(T)}{\tilde{N}^{\text {total }}(T)+\tilde{L}^{\text {total }}(T)} \underset{T \rightarrow \infty}{\stackrel{\text { a.s. }}{\longrightarrow}} 0,
$$

because whenever $N_{n} / L_{n} \rightarrow 0, \lim _{T \rightarrow \infty} N^{\text {total }(T)} /\left(N^{\text {total }}(T)+L^{\text {total }}(T)\right)=\lim _{n \rightarrow \infty} N_{n} /\left(L_{n}+N_{n}\right)=0$.

It follows that

$$
\begin{aligned}
& \lim _{T \rightarrow \infty} \frac{1}{T} \sum_{t=1}^{T} \operatorname{Pr}\left(\left(\theta_{t}, \hat{\theta}_{t}, a_{t}\right)=(\theta, \hat{\theta}, a)\right) \\
& =\lim _{T \rightarrow \infty} \mathbb{E}\left[\frac{\tilde{L}^{\text {total }}(T) 1(\theta=\hat{\theta}) f(\hat{\theta}) \alpha^{\prime}(a \mid \hat{\theta})+\tilde{N}^{\text {total }}(T) f^{\text {adj }}(T)(\theta, \hat{\theta}, a)}{\tilde{N}^{\text {total }}(T)+\tilde{L}^{\text {total }}(T)}\right] \\
& =1(\theta=\hat{\theta}) f(\hat{\theta}) \alpha^{\prime}(a \mid \hat{\theta})
\end{aligned}
$$

where $f^{\text {adj }}(T) \in \Delta(\Theta \times \Theta \times A)$ is the average frequency of types, reports, and allocations in the adjustment phases before $T$. Therefore, under truthtelling, the occupation measure over allocations and types equals

$$
\vartheta^{\prime}(a, \theta)=f(\theta) \alpha^{\prime}(a \mid \theta) .
$$

Hence, the agent's payoff in the dynamic mechanism under truthtelling is:

$$
U^{\mathrm{truth}}=\sum_{(a, \theta)} f(\theta) \alpha^{\prime}(a \mid \theta) u^{\prime}(a, \theta) .
$$

It remains to show that the agent cannot achieve more than $U^{\text {truth }}$ under any other strategy. To this end, fix and alternative strategy $\sigma$, and denote by $U(\sigma)=\limsup _{T \rightarrow \infty} U_{T}(\sigma)$ where:

$$
U_{T}(\sigma)=\frac{1}{T} \sum_{t=1}^{T} \sum_{a, \theta} \operatorname{Pr}\left(\left(a_{t}, \theta_{t}\right)=(a, \theta)\right) u^{\prime}(a, \theta) .
$$

Consider any convergent subsequence $\left(U_{T_{n}}\right)_{n=1}^{\infty}$ along times $\left\{T_{n}\right\}_{n=1}^{\infty}$. Because $\Delta(A \times \Theta \times \Theta)$ is compact (Aliprantis and Border, 2006, Theorem 15.11), a convergent (sub)subsequence at times $\left\{T_{k}\right\}_{k=1}^{\infty} \subseteq$ $\left\{T_{n}\right\}_{n=1}^{\infty}$ exists along which the occupation measure induced by $\sigma$

$$
v_{\sigma}^{T_{k}} \xrightarrow{w^{*}} v_{\sigma},
$$

for some $v_{\sigma} \in \Delta(A \times \Theta \times \Theta)$, which by (C.48) satisfies $v_{\sigma}(a, \hat{\theta})=f(\hat{\theta}) \alpha^{\prime}(a \mid \hat{\theta})$. It follows that for some undetectable deviation $v_{\sigma}(\hat{\theta} \mid \theta)$ :

$$
\lim _{n \rightarrow \infty} U_{T_{n}}=\lim _{k \rightarrow \infty} U_{T_{k}}=\sum_{\theta, \hat{\theta}, a} f(\theta) v_{\sigma}(\hat{\theta} \mid \theta) \alpha^{\prime}(a \mid \hat{\theta}) u^{\prime}(a, \theta) \leq U^{\mathrm{truth}},
$$

where the inequality follows because $\alpha^{\prime}(a \mid \hat{\theta})$ lacks profitable undetectable deviations. Because this inequality holds for any convergent subsequence $\left(U_{T_{n}}\right)_{n=1}^{\infty}$,

$$
U(\sigma)=\limsup _{T \rightarrow \infty} U_{T}(\sigma) \leq U^{\text {truth }}
$$

Finally, observe that the construction ensures that after every history, truthtelling from there on delivers the continuation payoff $U^{\text {truth }}$. Since $\alpha^{\prime}$ is ex ante individually rational, $U^{\text {truth }} \geq \sum_{\theta} f(\theta) u^{\prime}\left(a_{\varnothing}, \theta\right)$, and thus the participation constraints are satisfied. This concludes the proof. $\square$

Step 2 Consider now the outcome distribution $\vartheta \in \Delta(A \times \Theta \times \Omega)$ satisfying Equation C.39. As we argue in the proof of Theorem 2, a finite $K \leq|A||\Theta||\Omega|,\left\{\mu_{1}, \ldots, \mu_{K}\right\} \in \Delta(\Omega)$, and $\tau^{\prime} \in \Delta(\Delta(\Omega))$ exists such that

$$
\vartheta(a, \theta, \omega)=f(\theta) \sum_{k=1}^{K} \tau^{\prime}\left(\mu_{k}\right) \mu_{k}(\omega) \alpha\left(a \mid \theta, \mu_{k}\right) .
$$

Step 3 We now use steps 1 and 2 to complete the proof of Theorem 4, so in what follows we use the finite support representation of $\vartheta$ in the previous step. By Bayes plausibility, a dynamic mechanism can generate the belief split $\tau^{\prime}$ in $T$ periods with $T \leq\left\lceil\log _{|A|}(|\Omega||\Theta||A|)\right\rceil$, by treating each sequence of allocations of length $T$ as a message. This can be achieved by making the mechanism constant on the agent's type reports during the first $T$ periods. Since each $\alpha\left(\cdot \mid \cdot, \mu_{k}\right)$ for $k \in\{1, \ldots, K\}$ lacks profitable undetectable deviations and is individually rational, Proposition C. 1 implies that a dynamic mechanism exists that implements $\vartheta$ by first generating the belief split $\tau^{\prime}$ and then implementing $\alpha\left(\cdot \mid \cdot, \mu_{k}\right)$ in the corresponding continuation play.

## D Proof of auxiliary results

## D. 1 Revelation principle for calibrated mechanism design

In the main text, we restricted attention to incentive compatible and individually rational calibrated mechanisms. We show in this appendix that this restriction is without loss of generality by considering mechanisms with arbitrary message spaces and participation and reporting decisions by the agents that constitute an equilibrium of the game induced by the mechanism and its calibrated information structure.

Mechanisms Let $2^{[N]} \backslash \varnothing$ denote the nonempty subsets of agents. Then, we can define a mechanism as a collection $\left\{\left(M_{J}, \phi_{J}\right): J \in 2^{[N]} \backslash \varnothing\right\}$, where

$$
\phi_{J}: M_{J} \times \Omega \times[0,1] \rightarrow \Delta\left(A_{J}\right),
$$

is the mechanism when agents in $J$ participate, where $M_{J}=\times_{i \in J} M_{i}$ and $A_{J}=\times_{i \in J} A_{i}$.

Information Structure Let $\hat{S}_{i}=\Delta\left(A_{i}\right)^{M_{i}}$ denote the collection of menus of lotteries with labels $M_{i}$, and let $\hat{S}=\times_{i \in[N]} \hat{S}_{i}$. An information structure is $(\pi, \hat{S})$, where $\pi: \Omega \times[0,1] \rightarrow \hat{S}$.

Participation and reporting strategies It is notationally convenient to allow each agent to have her own randomization device $\varepsilon_{i} \sim U[0,1]$ and write agents' strategies as mappings $\left(p_{i}, \sigma_{i}\right): \Theta_{i} \times \hat{S}_{i} \times$ $[0,1] \rightarrow\{0,1\} \times M_{i}$, where $p_{i}$ denotes agent $i$ 's participation decision, and $\sigma_{i}$ her reporting strategy, conditional on participating. To distinguish the agents' randomization from that of the original mechanism, we reserve $\varepsilon_{0}$ for the realization of the mechanism's randomization device.

Given $\left(p_{i}, \sigma_{i}\right)_{i \in[N]}$ and a mechanism $(\phi, M)$, fix a profile $(\theta, \hat{s}, \bar{\varepsilon}) \equiv\left(\theta_{i}, \hat{s}, \varepsilon_{i}\right)_{i \in[N]}$. This determines a set of agents that participate,

$$
J(\theta, \hat{s}, \bar{\varepsilon})=\left\{j \in[N]: p_{j}\left(\theta_{j}, \hat{s}_{j}, \varepsilon_{j}\right)=1\right\},
$$

and let $J_{-i}(\theta, \hat{s}, \bar{\varepsilon})$ denote the projection of $J(\theta, \hat{s}, \bar{\varepsilon})$ on $J \backslash\{i\}$. Note that $J_{-i}$ only depends on $\left(\theta_{-i}, \hat{s}_{-i}, \bar{\varepsilon}_{-i}\right)$. Lastly, write $\phi_{J_{-i}\left(\theta_{-i}, \hat{s}_{-i}, \varepsilon_{-i}\right) \cup\{i\}}\left(m_{i}, \sigma_{J_{-i}\left(\theta_{-i}, \hat{s}_{-i}, \varepsilon_{-i}\right)}, \omega, \varepsilon_{0}\right) \in \Delta\left(A_{J_{-i}\left(\theta_{-i}, \hat{s}_{-i}, \varepsilon_{-i}\right) \cup\{i\}}\right)$ for

$$
\sum_{m_{J_{-i}\left(\theta_{-i}, \hat{s}_{-i}, \varepsilon_{-i}\right)}}\left(\prod_{j \in J_{-i}\left(\theta_{-i}, \hat{s}_{-i}, \varepsilon_{-i}\right)} \sigma_{j}\left(\theta_{j}, \hat{s}_{j}\right)\left(m_{j}\right)\right) \phi_{J_{-i}\left(\theta_{-i}, \hat{s}_{-i}, \varepsilon_{-i}\right) \cup\{i\}}\left(m_{i}, m_{J_{-i}\left(\theta_{-i}, \hat{s}_{-i}, \varepsilon_{-i}\right)}, \omega, \varepsilon_{0}\right)
$$

Calibrated information structures Given $\left(p_{i}, \sigma_{i}\right)_{i \in[N]}$ and a mechanism $(\phi, M)$, the information structure $(\pi, \hat{S})$ is calibrated with the mechanism and the agents' strategies if whenever $\pi\left(\omega, \varepsilon_{0}\right)=$ $\left(\hat{s}_{1}, \ldots, \hat{s}_{N}\right)$, then for all $i, m_{i}$

$$
\hat{s}_{i}\left(\cdot \mid m_{i}\right)=\mathbb{E}_{\tilde{\theta}_{-i} \sim f_{-i}(\cdot \mid \omega), \epsilon_{-i}}\left[\sum_{a_{-i} \in A_{-i}} \phi_{J_{-i}\left(\theta_{-i}, \hat{s}_{-i}, \varepsilon_{-i}\right) \cup\{i\}}\left(m_{i}, \sigma_{J_{-i}\left(\theta_{-i}, \hat{s}_{-i}, \varepsilon_{-i}\right)}, \omega, \epsilon_{0}\right)\left(\cdot, a_{-i}\right)\right] .
$$

Below, to keep the presentation simple, we focus on the case in which the calibrated information structure has finite support.

Equilibrium Given $\left(p_{i}, \sigma_{i}\right)_{i \in[N]}$, a mechanism $(\phi, M)$ and an information structure $(\pi, \hat{S})$ calibrated with the mechanism and the agents' strategies, $\left(p_{i}, \sigma_{i}\right)_{i \in[N]}$ is an equilibrium if for all $i \in[N]$, all $\theta_{i} \in \Theta_{i}$, all $\hat{s}_{i} \in \hat{S}_{i}$, and $\varepsilon_{i} \in[0,1]$, the following hold:

$$
\begin{aligned}
& \sigma_{i}\left(\theta_{i}, \hat{s}_{i}, \varepsilon_{i}\right) \in \arg \max _{m_{i} \in M_{i}} \sum_{a_{i} \in A_{i}} \hat{s}_{i}\left(a_{i} \mid m_{i}\right) \mathbb{E}_{\omega \sim \mu_{i}\left(\cdot \mid \theta_{i}, \hat{s}_{i}\right)}\left[u_{i}\left(a_{i}, \theta_{i}, \omega\right)\right], \\
& p_{i}\left(\theta_{i}, \hat{s}_{i}, \varepsilon_{i}\right) \in \arg \max _{p \in\{0,1\}} p \sum_{a_{i} \in A_{i}} \hat{s}_{i}\left(a_{i} \mid \sigma_{i}\left(\theta_{i}, \hat{s}_{i}, \varepsilon_{i}\right)\right) \mathbb{E}_{\omega \sim \mu_{i}\left(\cdot \mid \theta_{i}, \hat{s}_{i}\right)}\left[u_{i}\left(a_{i}, \theta_{i}, \omega\right)\right]+(1-p) \mathbb{E}_{\omega \sim \mu_{i}\left(\cdot \mid \theta_{i}, \hat{s}_{i}\right)}\left[u_{i}\left(a_{i \varnothing}, \theta_{i}, \omega\right)\right],
\end{aligned}
$$

where $\mu_{i}\left(\theta_{i}, \hat{s}_{i}\right) \in \Delta(\Omega)$ denotes agent $i$ 's updated beliefs about the state when her type is $\theta_{i}$ conditional on receiving signal $\hat{s}_{i}$.

Revelation Principle Fix $\left(p_{i}, \sigma_{i}\right)_{i \in[N]}$, a mechanism $(\phi, M)$ and an information structure $(\pi, \hat{S})$ calibrated with the mechanism such that $\left(p_{i}, \sigma_{i}\right)_{i \in[N]}$ is an equilibrium. We construct a direct mechanism $\left(\phi^{*}, \Theta\right)$ and a calibrated information structure $\left(\pi^{*}, S^{*}\right)$ calibrated with the mechanism under truthtelling and full participation such that truthtelling and full participation is an equilibrium.

First, note that we can extend each $\phi_{J}(\cdot) \in \Delta\left(A_{J}\right)$ to a mechanism $\bar{\phi}_{J}(\cdot) \in \Delta(A)$ as follows: for all $m \in M_{J}, \omega \in \Omega, \varepsilon_{0} \in[0,1]$, and $a_{J} \in A_{J}$,

$$
\bar{\phi}_{J}\left(m_{J}, \omega, \varepsilon_{0}\right)(a)=\phi_{J}\left(m_{J}, \omega, \varepsilon_{0}\right)\left(a_{J}\right) \times \delta_{a_{-J, \phi}} .
$$

Define a "pseudo"-mechanism as follows:

$$
\hat{\phi}_{N}\left(\theta, \omega, \varepsilon_{0}, \bar{\varepsilon}\right)=\bar{\phi}_{J\left(\theta, \pi\left(\omega, \varepsilon_{0}\right), \bar{\varepsilon}\right)}\left(\sigma_{J(\theta, \hat{s}, \bar{\varepsilon})}, \omega, \varepsilon_{0}\right) .
$$

where $\sigma_{J(\theta, \hat{s}, \bar{\varepsilon})}$ is the message vector generated by the strategies. Define the full participation mechanism $\phi_{N}^{*}: \Theta \times \Omega \times[0,1] \mapsto \Delta(A)$ to be

$$
\phi_{N}^{*}\left(\theta, \omega, \varepsilon_{0}\right)(a)=\int_{[0,1]^{N}} \hat{\phi}_{N}\left(\theta, \omega, \varepsilon_{0}, \bar{\varepsilon}\right)(a) \lambda^{N}(d \bar{\varepsilon})
$$

Let $S_{i}^{*}=\Delta\left(A_{i}\right)^{\Theta_{i}}$ and define $\pi^{*}\left(\omega, \varepsilon_{0}\right)=\left(s_{1}^{*}, \ldots, s_{N}^{*}\right) \in \times_{i \in[N]} S_{i}^{*}$, where

$$
s_{i}^{*}\left(\cdot \mid \hat{\theta_{i}}\right)=\mathbb{E}_{\theta_{-i} \sim f_{-i}(\cdot \mid \omega)}\left[\sum_{a_{-i}} \phi_{N}^{*}\left(\hat{\theta_{i}}, \theta_{-i}, \omega, \varepsilon_{0}\right)\left(\cdot, a_{-i}\right)\right] .
$$

By definition, the information structure is calibrated relative to full participation and truthful reporting.
We now show that full participation and truthful reporting is a best response to others participating and truthfully reporting into the mechanism. To this end, consider agent $i$ 's payoff from submitting report $\theta_{i}^{\prime}$ when observing $s_{i}^{*}$. Denoting by $\Sigma\left(\omega, s_{i}^{*}\right)$ the set of $\varepsilon_{0}$ such that $\pi_{i}^{*}=s_{i}^{*}$, this payoff is given by: ${ }^{53}$

$$
\begin{aligned}
& \sum_{\omega \in \Omega} \frac{\mu_{0}(\omega) f_{i}\left(\theta_{i} \mid \omega\right)}{\operatorname{Pr}\left(s_{i}^{*} \mid \theta_{i}\right)} \sum_{\theta_{-i}} f_{-i}\left(\theta_{-i} \mid \omega\right) \int_{\Sigma\left(\omega, s_{i}^{*}\right)} \sum_{a} \phi_{N}^{*}\left(\theta_{i}^{\prime}, \theta_{-i}, \omega, \varepsilon_{0}\right)\left(a_{i}, a_{-i}\right) \lambda\left(d \varepsilon_{0}\right) u_{i}\left(a_{i}, \theta_{i}, \omega\right)= \\
& \sum_{a_{i} \in A_{i}} \sum_{\omega \in \Omega} \frac{\mu_{0}(\omega) f_{i}\left(\theta_{i} \mid \omega\right)}{\operatorname{Pr}\left(s_{i}^{*} \mid \theta_{i}\right)} u_{i}\left(a_{i}, \theta_{i}, \omega\right) \sum_{\theta_{-i}} f_{-i}\left(\theta_{-i} \mid \omega\right) \int_{\Sigma\left(\omega, s_{i}^{*}\right)} \sum_{a_{-i}} \phi_{N}^{*}\left(\theta_{i}^{\prime}, \theta_{-i}, \omega, \varepsilon_{0}\right)\left(a_{i}, a_{-i}\right) \lambda\left(d \varepsilon_{0}\right) \\
& =\int_{0}^{1}\left[\sum_{a_{i}} \sum_{\omega} \frac{\mu_{0}(\omega) f_{i}\left(\theta_{i} \mid \omega\right)}{\operatorname{Pr}\left(s_{i}^{*} \mid \theta_{i}\right)} u_{i}\left(a_{i}, \theta_{i}, \omega\right) \int_{\Sigma\left(\omega, s_{i}^{*}\right)}(\star) \lambda\left(d \varepsilon_{0}\right)\right] \lambda\left(d \varepsilon_{i}\right)
\end{aligned}
$$

where

$$
\begin{aligned}
& \star=\mathbb{E}_{\theta_{-i} \mid \omega, \bar{\varepsilon}_{-i}}\left[\sum_{a_{-i}} \hat{\phi}_{N}\left(\theta_{i}^{\prime}, \theta_{-i}, \varepsilon_{0}, \bar{\varepsilon}_{-i}\right)\left(a_{i}, a_{-i}\right)\right] \\
& =\mathbb{E}_{\theta_{-i} \mid \omega, \bar{\varepsilon}_{-i}}\left[\sum_{a_{-i}} \bar{\phi}_{J\left(\theta_{i}^{\prime}, \theta_{-i}, \hat{s}\left(\omega, \varepsilon_{0}\right), \bar{\varepsilon}\right)}\left(\sigma_{J\left(\theta_{i}^{\prime}, \theta_{-i}, \hat{s}\left(\omega, \varepsilon_{0}\right), \bar{\varepsilon}\right)}, \omega, \varepsilon_{0}\right)\left(a_{i}, a_{-i}\right)\right] \\
& =\mathbb{1}\left[p_{i}\left(\theta_{i}^{\prime}, \hat{s}_{i}\left(\omega, \varepsilon_{0}\right), \varepsilon_{i}\right)=1\right] \hat{s}_{i}\left(a_{i} \mid \sigma_{i}\left(\theta_{i}^{\prime}, \hat{s}_{i}, \varepsilon_{i}\right)\right)+\left(1-\mathbb{1}\left[p_{i}\left(\theta_{i}^{\prime}, \hat{s}_{i}\left(\omega, \varepsilon_{0}\right), \varepsilon_{i}\right)=1\right]\right) \delta_{a_{i, \phi}}\left(a_{i}\right)
\end{aligned}
$$

Because agent $i$ of type $\theta_{i}$ could have imitated type $\theta_{i}^{\prime}$, reporting $\theta_{i}$ dominates. By the same logic, when the agent reports $\theta_{i}$, she obtains at least the payoff from participating in the mechanism.

## D. 2 Optimal Calibrated Auction

Proof of Proposition 5. The pointwise solution to the Myersonian problem allocates the good to agents in $N^{*}(\theta, \omega)=\arg \max _{i \in[N] \cup\{0\}}\left[w_{i}(\theta, \omega)+J_{i}\left(\theta_{i}\right) \omega_{i}+\omega_{0 i}\right]$ where $i=0$ corresponds to an outside option with $w_{0} \equiv J_{0} \equiv \omega_{00} \equiv 0$. The conditions of the proposition ensure that for all $i \in N$ and $j \in[N] \cup\{0\}$,

$$
\frac{d}{d \theta_{i}}\left(w_{i}(\theta, \omega)+J_{i}\left(\theta_{i}, F_{i}\right) \omega_{i}+\omega_{0 i}\right) \geq \frac{d}{d \theta_{i}}\left(w_{j}(\theta, \omega)+J_{j}\left(\theta_{j}, F_{j}\right) \omega_{j}+\omega_{0 j}\right) .
$$

Thus, an optimal selection $q^{*}(\theta, \omega)$ exists such that for each $i, \theta_{-i}$, and $\omega, q^{*}\left(\theta_{i}, \theta_{-i}, \omega\right)$ is nondecreasing in $\theta_{i}$ (e.g., one that uniformly randomizes over $N^{*}(\theta, \omega)$ ).

Denote by $Q_{\text {full }}$ the set of allocation rules implementable under full state disclosure. These are the rules such that for all $i$ and $\omega, \mathbb{E}_{F_{-i}}\left[q_{i}\left(\theta_{i}, \theta_{-i}, \omega\right)\right]$ is non-decreasing in $\theta_{i}$. It follows that $q^{*} \in Q_{\mathrm{full}}$, and hence $q^{*} \in Q_{\mathrm{My}}$. Thus, $q^{*}$ solves the Myersonian problem and can also be implemented by fully disclosing the state to the agents and conducting an optimal mechanism state-by-state. By revenue

[^35]equivalence, the expected revenue of such implementation is the same as under no disclosure, and thus the designer obtains payoff $W_{\text {My }}$. $\square$

## D. 3 Technical results from Appendix C

Proof of Lemma C.1. The set of continuous bounded functions on $\tilde{\Omega}$ is separable and hence it has a countable dense subset $\left\{g_{k}\right\}_{k \in \mathbb{N}} \subset C_{b}(\tilde{\Omega})$. It is immediate to see that $\mu_{n} \xrightarrow{w^{*}} \mu$ if and only if for all $k \in \mathbb{N}$ $\int g_{k} d \mu_{n} \rightarrow \int g_{k} d \mu$.

For each $k \in \mathbb{N}$ define a real-valued, bounded, martingale on ( $\mathcal{H}^{\infty}, \mathcal{B}_{\mathcal{H}^{\infty}}, \mathbb{P}_{\sigma}$ ) as follows:

$$
M_{t}^{k}\left(\tilde{\omega}, h^{\infty}\right)=\int_{\tilde{\Omega}} g_{k}\left(\omega^{\prime}, \varepsilon\right) d \mu_{t}\left(\tilde{\omega}, h^{\infty}\right)\left(\omega^{\prime}, \varepsilon\right)
$$

Doob's martingale convergence theorem implies that $M_{t}^{k}\left(\tilde{\omega}, h^{\infty}\right)=\mathbb{E}\left[g_{k} \mid h^{t}\right] \rightarrow M_{\infty}^{k}\left(\tilde{\omega}, h^{\infty}\right)=\mathbb{E}\left[g_{k} \mid h^{\infty}\right]$ $\mathbb{P}_{\sigma}$-a.s. Let $E_{k}$ denote the subset of $\mathcal{H}^{\infty}$ where convergence happens, and note that $\mathbb{P}_{\sigma}\left(E_{k}\right)=1$.

Let $E=\cap_{k} E_{k}$ and note that $\mathbb{P}_{\sigma}(E)=1$. Then, on $E$, we have that for all $k \in \mathbb{N}$,

$$
\int_{\tilde{\Omega}} g_{k}\left(\omega^{\prime}, \varepsilon\right) d \mu_{t}\left(\tilde{\omega}, h^{\infty}\right)\left(\omega^{\prime}, \varepsilon\right) \rightarrow M_{\infty}^{k}\left(\tilde{\omega}, h^{\infty}\right)
$$

Fix now a terminal history ( $\tilde{\omega}, h^{\infty}$ ). Because $\Delta(\tilde{\Omega})$ is compact (Aliprantis and Border, 2006, Theorem 15.11), the sequence $\left(\mu_{t}\left(\tilde{\omega}, h^{\infty}\right)\right)_{t \in \mathbb{N}}$ has a convergent subsequence $\mu_{t_{j}}\left(\tilde{\omega}, h^{\infty}\right) \xrightarrow{w^{*}} \tilde{\mu}$. Passing the limit along $t_{j}$ in Equation D. 1 we have that for all $k \in \mathbb{N}$

$$
\int_{\tilde{\Omega}} g_{k}\left(\omega^{\prime}, \varepsilon\right) d \mu_{t}\left(\tilde{\omega}, h^{\infty}\right)\left(\omega^{\prime}, \varepsilon\right) \rightarrow \int_{\tilde{\Omega}} g_{k}\left(\omega^{\prime}, \varepsilon\right) d \tilde{\mu}
$$

Because the set $\left\{g_{k}: k \in \mathbb{N}\right\}$ determines the convergent subsequences, any subsequential limits must be equal. Hence $\mu_{t}(\cdot)$ converges on $E$ and call this limit $\tilde{\mu}_{\infty}$. Hence on $E$ we have that $\mu_{t}\left(h^{\infty}\right) \xrightarrow{w^{*}} \tilde{\mu}_{\infty}\left(h^{\infty}\right)$.

Now, for each $k, \int g_{k} d \mu_{\infty}=\mathbb{E}\left[g_{k} \mid h^{\infty}\right]$ almost surely. Hence, $\tilde{\mu}_{\infty}$ is a version of the law of $\tilde{\Omega}$ conditional on $h^{\infty}$. This is $\mathbb{P}_{\sigma}\left(\cdot \mid h^{\infty}\right)$, completing the proof. $\square$

Proof of Lemma C.2. Fix a continuous and bounded function $g \in C_{b}(A \times \Theta \times \Omega \times \Delta(\Omega))$ and define for each terminal history $\left(\omega, h^{\infty}\right)$

$$
\Delta_{t}\left(\omega, h^{\infty}\right)=g\left(a_{t}\left(h^{\infty}\right), \theta_{t}\left(h^{\infty}\right), \omega, \mu_{t}\left(h^{\infty}\right)\right)-g\left(a_{t}\left(h^{\infty}\right), \theta_{t}\left(h^{\infty}\right), \omega, \mu_{t+1}\left(h^{\infty}\right)\right) .
$$

As in Lemma C.1, let $E$ denote the probability-1 subset of $\mathcal{H}^{\infty}$ on which $\mu_{t} \xrightarrow{w^{*}} \mu_{\infty} .{ }^{54}$ Then, on $E$, $\Delta_{t}\left(\omega, h^{\infty}\right) \rightarrow 0$, and hence,

$$
\mathbb{E}_{\sigma}\left[\Delta_{t}\left(\omega, h^{\infty}\right)\right] \rightarrow 0,
$$

[^36]as $t \rightarrow \infty$. Now, for every $T$,
$$
D_{T}(g) \equiv \mathbb{E}_{\bar{v}_{\sigma}^{T, 1}}[g]-\mathbb{E}_{\bar{v}_{\sigma}^{T, 2}}[g]=\frac{1}{T} \sum_{t=1}^{T} \mathbb{E}_{\sigma}\left[\Delta_{t}\right]
$$
and hence the left-hand side goes to 0 as $T \rightarrow \infty$.
Now, let $T_{n}$ be such that $\bar{v}_{\sigma}^{T_{n}, 2} \xrightarrow{w^{*}} \bar{v}$ for some $\bar{v} \in \Delta(A \times \Theta \times \Omega \times \Delta(\Omega))$. Note that
$$
\mathbb{E}_{\bar{v}_{\sigma}^{T n, 1}}[g]=\mathbb{E}_{\bar{v}_{\sigma}^{T n, 2}}[g]+D_{T_{n}}[g] \rightarrow \mathbb{E}_{\bar{v}}[g]+0,
$$
so a subsequential limit of $\bar{v}_{\sigma}^{T_{n}, 2}$ is a subsequential limit of $\bar{v}_{\sigma}^{T_{n}, 1}$. Switching the role of 1 and 2, we obtain the opposite set inclusion and the result follows. $\square$

Proof of Lemma C.3. Fix a continuous function $g$ on $\Delta(\tilde{\Omega})$. Then, we want to show that

$$
\mathbb{E}_{\tau_{T}}[g] \rightarrow \mathbb{E}_{\mathbb{P}_{\sigma} \circ \mu_{\infty}^{-1}}[g] .
$$

By Lemma C.1, $\mu_{t} \xrightarrow{w^{*}} \mu_{\infty} \mathbb{P}_{\sigma}$-almost surely and $g$ is continuous, we have that $\mathbb{E}_{\sigma}\left[g\left(\mu_{t}\right)\right] \rightarrow \mathbb{E}_{\sigma}\left[g\left(\mu_{\infty}\right)\right]$ by dominated convergence theorem. ${ }^{55}$ Because eventually constant sequences have Cesàro limits, we have that

$$
\frac{1}{T} \sum_{t=1}^{T} \mathbb{E}_{\sigma}\left[g\left(\mu_{t}\right)\right] \rightarrow \mathbb{E}_{\sigma}\left[g\left(\mu_{\infty}\right)\right] .
$$

And now we are basically done, because

$$
\mathbb{E}_{\tau_{T}}[g]=\int_{H^{\infty}} \frac{1}{T} \sum_{t=1}^{T} g\left(\mu_{t}\left(h^{\infty}\right)\right) \mathbb{P}_{\sigma}\left(d h^{\infty}\right)=\frac{1}{T} \sum_{t=1}^{T} \mathbb{E}_{\sigma}\left[g\left(\mu_{t}\right)\right] \rightarrow \mathbb{E}_{\sigma}\left[g\left(\mu_{\infty}\right)\right]=\int g d\left(\mathbb{P}_{\sigma} \circ \mu_{\infty}^{-1}\right) .
$$

In other words, the occupation measure on beliefs induced by the strategy (and the prior, the type distribution, and the mechanism) is the push-forward measure $\left(\mathbb{P}_{\sigma} \circ \mu_{\infty}^{-1}\right)$. In particular, that $\mathbb{P}_{\sigma}$ is a measure implies that $\left(\mathbb{P}_{\sigma} \circ \mu_{\infty}^{-1}\right)$ is a measure itself (Bogachev, 2007, Chapter 3.6). $\square$

Proof of Lemma C.4. We now show the agent can ensure the payoff $U\left(\tau_{\phi}\right)$ in Equation C.21, which corresponds to the agent's maximum payoff under the calibrated information structure $\pi_{\phi}$. Recall that this information structure is the one that corresponds to the partition of $\tilde{\Omega}, \mathscr{P}$, defined as follows: $\tilde{\omega}, \tilde{\omega}^{\prime}$ in the same cell $P$ of $\mathscr{P}$ if for all $(a, m), \phi(a \mid m, \tilde{\omega})=\phi\left(a \mid m, \tilde{\omega}^{\prime}\right)$. Conditional on cell $P$, the associated posterior is $\mu(\mid P) \in \Delta(\tilde{\Omega})$. Let $\tau_{\phi} \in \Delta(\Delta(\tilde{\Omega}))$ denote the induced belief distribution (with mean $\mu_{0} \otimes \eta$ ).

To prove the result, we consider the strategy $\sigma_{N}^{\prime}$ parameterized by a number $N$ and defined as follows. In the exploration phase of $\sigma_{N}^{\prime}$, the agent plays each message $m$ for $N$ rounds. Let $\mu_{N|M|}\left(h^{N|M|}\right)$ denote the agent's beliefs as a function of the realized sequence of allocations implied by $h^{N|M|}$. For each

[^37]history that succeeds $h^{N|M|}$, the agent of type $\theta$ plays the message $m$ that solves
$$
\max _{m \in M} \sum_{\tilde{\omega}} \mu_{N|M|}\left(h^{N|M|}\right)(\tilde{\omega}) \sum_{a \in A} \phi(a \mid m, \tilde{\omega}) u(a, \theta, \omega) .
$$
It is immediate to verify that the agent's (limit) average payoff under $\sigma_{N}^{\prime}$ is given by:
$$
\begin{aligned}
U\left(\sigma_{N}^{\prime}\right) & =\mathbb{E}_{\sigma_{N}^{\prime}}\left[\sum_{\theta \in \Theta} f(\theta) \max _{m \in M} \sum_{\tilde{\omega}} \mu_{N|M|}\left(h^{N|M|}\right)(\tilde{\omega}) \sum_{a \in A} \phi(a \mid m, \tilde{\omega}) u(a, \theta, \omega)\right]= \\
& =\sum_{h^{N|M|} \in H^{N|M|}} \mathbb{P}_{\sigma_{N}^{\prime}}^{N|M|}\left(h^{N|M|}\right) u^{*}\left(\mu_{N|M|}\left(h^{N|M|}\right)\right),
\end{aligned}
$$
where $u^{*}$ is as in Equation C.20.
Below, we show that the distribution of beliefs under $\sigma_{N}^{\prime}, \mathbb{P}_{\sigma_{N}^{\prime}} \circ \mu_{N|M|}^{-1} \xrightarrow{w^{*}} \tau_{\phi}$. Consequently, as $u^{*}$ is continuous and bounded on $\Delta(\tilde{\Omega})$, for any $\delta>0$, we can choose $N_{\delta}$ so that for all $N \geq N_{\delta}$,
$$
\left|U\left(\sigma_{N}^{\prime}\right)-U\left(\tau_{\phi}\right)\right| \leq \delta .
$$
Consequently, the agent's payoff under $\sigma$ must be $U\left(\tau_{\phi}\right)$ because by definition for all $\delta>0^{56}$
$$
\liminf _{T \rightarrow \infty} \mathbb{E}_{\sigma}\left[U_{T}\right] \geq \limsup _{T \rightarrow \infty} \mathbb{E}_{\sigma_{N_{\delta}}^{\prime}}\left[U_{T}\right]=U\left(\sigma_{N_{\delta}}^{\prime}\right) \geq U\left(\tau_{\phi}\right)-\delta,
$$
and hence,
$$
\liminf _{T \rightarrow \infty} \mathbb{E}_{\sigma}\left[U_{T}\right] \geq \lim _{\delta \rightarrow 0} U\left(\tau_{\phi}\right)-\delta=U\left(\tau_{\phi}\right) .
$$
which completes the proof.
We now complete the missing step:

The law of $\mu_{N|M|}$ converges to $\tau_{\phi}$ It is useful to write the bottom line of Equation D. 2 as follows:

$$
\sum_{\tilde{\omega} \in \tilde{\Omega}}\left(\mu_{0} \otimes \eta\right)(\tilde{\omega}) \mathbb{E}_{\mathbb{P}_{\sigma_{N}^{\prime}}^{N|M|}(\cdot \mid \tilde{\omega})}\left[u^{*}\left(\mu_{N|M|}\right)\right] .
$$

We show that the conditional law of $\mu_{N|M|}, \mathbb{P}_{\sigma_{N}^{\prime}}(\cdot \mid \tilde{\omega})$, converges to the Dirac measure on $\mu(\cdot \mid P(\tilde{\omega}))$. Noting that $\tau_{\phi}=\sum_{\tilde{\omega} \in \tilde{\Omega}}\left(\mu_{0} \otimes \eta\right)(\tilde{\omega}) \delta_{\mu(\cdot \mid P(\tilde{\omega}))}$ completes the proof.

For the exploration block, define for each $m \in M$, the empirical frequency $\hat{\phi}_{N, m}:(\Theta \times M \times A)^{N|M|} \rightarrow$ $\Delta(A)$, as follows

$$
\hat{\phi}_{N, m}\left(h^{N|M|}\right)(a)=\frac{1}{N} \sum_{n=1}^{N} \mathbb{1}\left[A_{m, n}=a\right],
$$

where $A_{m, n}$ is the $n^{\text {th }}$ draw from $A$ when the message is $m$. Let $\hat{\phi}_{N}: H^{N|M|} \rightarrow \Delta(A)^{M}$ denote the vector

[^38]of empirical frequencies. The agent's belief at history $h^{N|M|}$ is given by:
$$
\mu_{N|M|}\left(h^{N|M|}\right)(\tilde{\omega})=\frac{\left(\mu_{0} \otimes \eta\right)(\tilde{\omega}) \prod_{m \in M} \prod_{a \in A} \phi(a \mid m, \tilde{\omega})^{N \hat{\phi}_{N, m}\left(h^{N|M|}\right)(a)}}{\sum_{\tilde{\omega}^{\prime}}\left(\mu_{0} \otimes \eta\right)\left(\tilde{\omega}^{\prime}\right) \prod_{m \in M} \prod_{a \in A} \phi\left(a \mid m, \tilde{\omega}^{\prime}\right)^{N \hat{\phi}_{N, m}\left(h^{N|M|}\right)(a)}} .
$$
Denote by $\tau_{N|M|, \tilde{\omega}} \in \Delta(\Delta(\tilde{\Omega}))$ the law of $\mu_{N|M|}$ conditional on $\tilde{\omega}$, i.e., $\tau_{N|M|, \tilde{\omega}}=\mathbb{P}_{\sigma_{N}^{\prime}}(\cdot \mid \tilde{\omega}) \circ \mu_{N|M|}^{-1}$. Below, we show that $\tau_{N|M|, \tilde{\omega}}$ converges weakly to $\delta_{\mu(\cdot \mid P(\tilde{\omega}))}$.

Suppose the true state is $\tilde{\omega}^{\star}$. Then, $\left(A_{m, 1}, \ldots, A_{m, N}\right)$ are drawn i.i.d. from distribution $\phi\left(\cdot \mid m, \tilde{\omega}^{\star}\right)$. Fix a continuous and bounded function $g$ on $A$. Then, almost surely, ${ }^{57}$

$$
\int_{A} g d \hat{\phi}_{N}(\cdot \mid m)=\frac{1}{N} \sum_{n=1}^{N} g\left(A_{m, n}\right) \rightarrow \mathbb{E}_{\phi}\left[g\left(A_{m, 1}\right)\right]=\sum_{a \in A} g(a) \phi\left(a \mid m, \tilde{\omega}^{\star}\right),
$$

by the strong law of large numbers applied to the i.i.d random variables $\left(g\left(A_{m, 1}\right), \ldots, g\left(A_{m, N}\right)\right)$. Because this holds for all $g$, then $\hat{\phi}_{N}(\cdot \mid m) \xrightarrow{w^{*}} \phi\left(\cdot \mid m, \tilde{\omega}^{\star}\right)$ almost surely when the true state is $\tilde{\omega}^{\star}$.

Fix an arbitrary state $\tilde{\omega}$ and consider the ratio of the right-hand side of Equation D. 3 at $\tilde{\omega}$ and $\tilde{\omega}^{\star}$ :

$$
\frac{\mu_{N|M|}\left(h^{N|M|}\right)(\tilde{\omega})}{\mu_{N|M|}\left(h^{N|M|}\right)\left(\tilde{\omega}^{\star}\right)}=\frac{\left(\mu_{0} \otimes \eta\right)(\tilde{\omega})}{\left(\mu_{0} \otimes \eta\right)\left(\tilde{\omega}^{\star}\right)} \prod_{a \in A} \prod_{m \in M}\left(\frac{\phi(a \mid m, \tilde{\omega})}{\phi\left(a \mid m, \tilde{\omega}^{\star}\right)}\right)^{N \hat{\phi}_{N, m}\left(h^{N|M|}\right)(a)}
$$

Suppose $\tilde{\omega} \notin P\left(\tilde{\omega}^{\star}\right)$. By definition of the partition $\mathscr{P}$, a message $m \in M$ and allocation $a \in A$ exist such that $\phi(a \mid m, \tilde{\omega}) \neq \phi\left(a \mid m, \tilde{\omega}^{\star}\right)$. Taking logarithm on both sides of Equation D. 4 and dividing by $N$,

$$
\frac{1}{N} \log \left(\frac{\mu_{N|M|}\left(h^{N|M|}\right)(\tilde{\omega})}{\mu_{N|M|}\left(h^{N|M|}\right)\left(\tilde{\omega}^{\star}\right)}\right)=\frac{1}{N} \log \left(\frac{\left(\mu_{0} \otimes \eta\right)(\tilde{\omega})}{\left(\mu_{0} \otimes \eta\right)\left(\tilde{\omega}^{\star}\right)}\right)+\sum_{a^{\prime} \in A} \sum_{m^{\prime} \in M} \hat{\phi}_{N, m^{\prime}}\left(h^{N|M|}\right)\left(a^{\prime}\right) \log \left(\frac{\phi\left(a^{\prime} \mid m^{\prime}, \tilde{\omega}\right)}{\phi\left(a^{\prime} \mid m^{\prime}, \tilde{\omega}^{\star}\right)}\right) .
$$

Because $\hat{\phi}_{N}(\cdot \mid m) \xrightarrow{w^{*}} \phi\left(\cdot \mid m, \tilde{\omega}^{\star}\right)$ almost surely when the true state is $\tilde{\omega}^{\star}$,

$$
\lim _{N \rightarrow \infty} \frac{1}{N} \log \left(\frac{\mu_{N|M|}\left(h^{N|M|}\right)(\tilde{\omega})}{\mu_{N|M|}\left(h^{N|M|}\right)\left(\tilde{\omega}^{\star}\right)}\right)=-\sum_{m \in M} \mathrm{D}_{\mathrm{KL}}\left(\phi\left(\cdot \mid m, \tilde{\omega}^{\star}\right) \mid \phi(\cdot \mid m, \tilde{\omega})\right),
$$

where $\mathrm{D}_{\mathrm{KL}}$ is the Kullback-Leibler divergence. Note that at least one of the terms in the KL-divergence is positive as $\phi(\cdot \mid m, \tilde{\omega}) \neq \phi\left(\cdot \mid m, \tilde{\omega}^{\star}\right)$. Hence,

$$
\lim _{N \rightarrow \infty} \log \left(\frac{\mu_{N|M|}\left(h^{N|M|}\right)(\tilde{\omega})}{\mu_{N|M|}\left(h^{N|M|}\right)\left(\tilde{\omega}^{\star}\right)}\right)=-\infty,
$$

meaning that $\mu_{N|M|}\left(h^{N|M|}\right)(\tilde{\omega}) / \mu_{N|M|}\left(h^{N|M|}\right)\left(\tilde{\omega}^{\star}\right) \rightarrow 0$.
Suppose now that $\tilde{\omega} \in P\left(\tilde{\omega}^{\star}\right)$. Then, Equation D. 4 reduces to

$$
\frac{\mu_{N|M|}\left(h^{N|M|}\right)(\tilde{\omega})}{\mu_{N|M|}\left(h^{N|M|}\right)\left(\tilde{\omega}^{\star}\right)}=\frac{\left(\mu_{0} \otimes \eta\right)(\tilde{\omega})}{\left(\mu_{0} \otimes \eta\right)\left(\tilde{\omega}^{\star}\right)},
$$

for all $N$.

[^39]Collecting both cases, we conclude that conditional on the true state being $\tilde{\omega}^{\star}$,

$$
\sum_{\tilde{\omega} \in P\left(\tilde{\omega}^{\star}\right)} \mu_{N|M|}(\tilde{\omega}) \rightarrow_{N \rightarrow \infty} 1,
$$

and moreover, within the cell, the fixed-ratio property implies the law $\tau_{N|M|, \tilde{\omega}^{\star}} \xrightarrow{w^{*}} \delta_{\mu\left(\cdot \mid P\left(\tilde{\omega}^{\star}\right)\right)}$. We conclude that the unconditional belief distribution, $\sum_{\tilde{\omega} \in \tilde{\Omega}}\left(\mu_{0} \otimes \eta\right)(\tilde{\omega}) \tau_{N|M|, \tilde{\omega}} \xrightarrow{w^{*}} \sum_{\tilde{\omega} \in \tilde{\Omega}}\left(\mu_{0} \otimes \eta\right)(\tilde{\omega}) \delta_{\mu(P(\tilde{\omega}))}=$ $\tau_{\phi}$. In particular,

$$
U\left(\sigma_{N}^{\prime}\right)=\mathbb{E}_{\mathbb{P}_{\sigma_{N}^{\prime}} \circ \mu_{N|M|}^{-1}}\left[u^{*}(\mu)\right] \rightarrow \mathbb{E}_{\tau_{\phi}}\left[u^{*}(\mu)\right]
$$

completing the proof. $\square$

Lemma D.1. Suppose $u: \Delta(\Omega) \rightarrow \mathbb{R}$ satisfies that

$$
\int_{\Delta(\Omega)} u(\mu) \tau(d \mu)<\int_{\Delta(\Omega)} \max \{u(\mu), 0\} \tau(d \mu)
$$

then a set $B \subseteq \Delta(\Omega)$ open relative to $\Delta(\Omega)$ exists such that

$$
\int_{B} u(\mu) \tau(d \mu)<0
$$

Proof. Define the positive and negative parts of $u$ :

$$
u_{+}(\mu):=\max \{u(\mu), 0\}, \quad u_{-}(\mu):=\max \{-u(\mu), 0\} .
$$

Then $u=u_{+}-u_{-}$pointwise. Integrating and using the assumed strict inequality,

$$
\int u d \tau=\int u_{+} d \tau-\int u_{-} d \tau<\int u_{+} d \tau \Rightarrow \int u_{-} d \tau>0 .
$$

Hence the set $N:=\{\mu \in \Delta(\Omega): u(\mu)<0\}$ has strictly positive mass under $\tau$.
Embed $\Delta(\Omega) \subset \mathbb{R}^{|\Omega|}$. Extend $\tau$ to a finite Borel measure $\tilde{\tau}$ on $\mathbb{R}^{d}$ by

$$
\tilde{\tau}(A):=\tau(A \cap \Delta(\Omega)) \quad\left(A \subseteq \mathbb{R}^{d} \text { Borel }\right),
$$

and extend $u$ to $\tilde{u}: \mathbb{R}^{d} \rightarrow[-1,1]$ by $\tilde{u}=u$ on $\Delta(\Omega)$ and $\tilde{u}=0$ on $\mathbb{R}^{d} \backslash \Delta(\Omega)$. Then $\tilde{u} \in L^{1}(\tilde{\tau})$ and $\tilde{\tau}(N)=\tau(N)>0$.

Let $B_{\mathbb{R}|\Omega|}(x, r)$ denote the ball in $\mathbb{R}^{|\Omega|}$ with center $x$ and radius $r$. By the Lebesgue differentiation theorem for finite Borel measures on $\mathbb{R}^{d}$, there is a $\tilde{\tau}$-full-measure set $D \subseteq \mathbb{R}^{d}$ such that for every $x \in D$,

$$
\lim _{r \downarrow 0} \frac{1}{\tilde{\tau}\left(B_{\mathbb{R}^{d}}(x, r)\right)} \int_{B_{\mathbb{R}^{d}}(x, r)} \tilde{u} d \tilde{\tau}=\tilde{u}(x),
$$

whenever $\tilde{\tau}\left(B_{\mathbb{R}^{d}}(x, r)\right)>0$ (and this positivity holds for all sufficiently small $r$ for $\tilde{\tau}$-a.e. $x$ ). Since $\tilde{\tau}(N \cap D)>0$, choose $\mu_{0} \in N \cap D$. Then $\tilde{u}\left(\mu_{0}\right)=u\left(\mu_{0}\right)<0$. Therefore the above limit is strictly negative,
so there exists $r_{0}>0$ such that for all $0<r<r_{0}$,

$$
\int_{B_{\mathbb{R}^{d}}\left(\mu_{0}, r\right)} \tilde{u} d \tilde{\tau}<0
$$

For such an $r$, let $B:=B_{\Delta}\left(\mu_{0}, r\right)=\Delta(\Omega) \cap B_{\mathbb{R}^{d}}\left(\mu_{0}, r\right)$, which is an open ball in $\Delta(\Omega)$. Using the definitions of $\tilde{\tau}$ and $\tilde{u}$,

$$
\int_{B} u d \tau=\int_{B_{\mathbb{R}^{d}}\left(\mu_{0}, r\right)} \tilde{u} d \tilde{\tau}<0
$$

This proves the claim. $\square$

## D.3.1 Revelation principle for limit of means preferences

We show in this section that when the designer uses dynamic mechanisms, it is without loss of generality for the designer to employ direct dynamic mechanisms that (i) implement the outside option at all histories after the agent first exercises her option not to participate in the mechanism, and (ii) for which the agent's best response is to always participate and truthfully report her type. This justifies the class of mechanisms we employ in the analysis of Section 5.2.

Histories, mechanisms, and strategies As in the main text, to simplify notation, we do not include the agent's decision to participate in the mechanism in the histories of the game. Instead, we follow the convention that if the agent does not participate, it is as if she reported $\varnothing$ and the allocation is $a_{\varnothing}$. Formally, let $M A_{\varnothing}=(M \times A) \cup\left\{\left(\varnothing, a_{\varnothing}\right)\right\}$. With this notation, a history through period $t$ is an element of $\hat{H}_{M}^{t} \equiv\left(M A_{\varnothing}\right)^{t-1}$ and let $\hat{\mathcal{H}}_{M}^{t}=\Omega \times \hat{H}_{M}^{t} .{ }^{58}$

A mechanism is a collection $\varphi \equiv\left(\varphi_{t}\right)_{t=1}^{\infty}$ such that the mechanism in period $t$ is a mapping $\varphi_{t}$ : $\hat{\mathcal{H}}_{M}^{t} \times M \rightarrow \Delta(A)$.

Let $H_{M}^{t}=\left(\Theta \times M A_{\varnothing}\right)^{t-1}=\Theta^{t-1} \times \hat{H}_{M}^{t}$. The agent's strategy, ( $p, \sigma$ ), is given by her participation strategy $p_{t}: H_{M}^{t} \times \Theta \rightarrow[0,1]$, and conditional on participating, her reporting strategy $\sigma_{t}: H_{M}^{t} \times \Theta \rightarrow \Delta(M)$.

The distribution over terminal histories To obtain the complete description of the paths on the tree we need to append $\Omega$ to $H_{M}^{t}$; hence the paths through period $t-1$ are $\Omega \times H_{M}^{t} \equiv \mathcal{H}_{M}^{t}$. The distributions over states and agent's types, the agent's strategy, and the mechanism induce a distribution over the terminal histories $\mathcal{H}_{M}^{\infty} \equiv \Omega \times H_{M}^{\infty}$, which we denote by $\mathbb{P}_{\varphi,(p, \sigma)} \in \Delta\left(\Omega \times H_{M}^{\infty}\right)$, as it is now useful to keep track of the mechanism. We denote by $\mathbb{E}_{(p, \sigma)}$ the expectation under this measure. The distribution $\mathbb{P}_{\varphi,(p, \sigma)} \in \Delta\left(\Omega \times H_{M}^{\infty}\right)$ is the unique distribution that satisfies that for all $t \in \mathbb{N}, \tilde{\mathcal{H}}_{M}^{t} \subset \Omega \times H_{M}^{t}$,

$$
\mathbb{P}_{\varphi,(p, \sigma)}\left(\tilde{\mathcal{H}}_{M}^{t} \times \prod_{s=t+1}^{\infty}\left(\Theta \times M A_{\varnothing}\right)\right)=\mathbb{P}_{\varphi,(p, \sigma)}^{t}\left(\tilde{\mathcal{H}}_{M}^{t}\right),
$$

[^40]where the distributions $\left(\mathbb{P}_{\varphi,(p, \sigma)}^{t}\right)_{t \in \mathbb{N}}$ satisfy
$$
\begin{aligned}
\mathbb{P}_{\varphi,(p, \sigma)}^{t+1}\left(\omega, h_{M}^{t}, \theta, m, a\right) & =\mathbb{P}_{\varphi,(p, \sigma)}^{t}\left(\omega, h_{M}^{t}\right) f(\theta) p_{t}\left(h_{M}^{t}, \theta\right) \sigma_{t}\left(h_{M}^{t}, \theta\right)(m) \varphi_{t}\left(\omega, \hat{h}_{M}^{t}, m\right)(a), \\
\mathbb{P}_{\varphi,(p, \sigma)}^{t+1}\left(\omega, h_{M}^{t}, \theta, \varnothing, a\right) & =\mathbb{P}_{\varphi,(p, \sigma)}^{t}\left(\omega, h_{M}^{t}\right) f(\theta)\left(1-p_{t}\left(h_{M}^{t}, \theta\right)\right) \mathbb{1}\left[a=a_{\varnothing}\right] .
\end{aligned}
$$

Outcome distribution Our interest is in the distribution over payoff-relevant outcomes, $\Omega \times(\Theta \times A)^{\infty}$, and hence on the marginal of $\mathbb{P}_{\varphi,(p, \sigma)}$ on $\Omega \times(\Theta \times A)^{\infty}$, which we denote by $\overline{\mathbb{P}}_{\varphi,(p, \sigma)}$.

Best response We say that strategy $(p, \sigma)$ is a best response for the agent if for all alternative strategies $\left(p^{\prime}, \sigma^{\prime}\right)$, we have that

$$
\liminf _{T \rightarrow \infty} \mathbb{E}_{(p, \sigma)}\left[U_{T}\right] \geq \limsup _{T \rightarrow \infty} \mathbb{E}_{\left(p^{\prime}, \sigma^{\prime}\right)}\left[U_{T}\right],
$$

where recall $U_{T}$ is the agent's average payoff until period $T$.

Direct and full participation mechanisms A special case of the above game is that in which $M=\Theta$, and whenever the agent does not participate, the mechanism chooses $a_{\varnothing}$ with probability 1 for any message in all continuation histories. We call these mechanisms direct and full participation mechanisms. Below, when $M=\Theta$, we drop the dependence of the set histories on $M$.

Formally, let $\hat{\mathcal{H}}_{\varnothing}^{t}$ denote the subset of $\hat{\mathcal{H}}^{t}$ such that at some point the sequence ( $\varnothing, a_{\varnothing}$ ) appears. We define mechanisms

$$
\tilde{\varphi}_{t}: \hat{\mathcal{H}}^{t} \times \Theta \rightarrow \Delta(A),
$$

such that $\tilde{\varphi}_{t}\left(\omega, \hat{h}^{t}, \cdot\right)=\mathbb{1}\left[a=a_{\varnothing}\right]$ whenever $\left(\omega, \hat{h}^{t}\right) \in \hat{\mathcal{H}}_{\varnothing}^{t}$.
Theorem D.1. Suppose that $(p, \sigma)$ is a best response to mechanism $\varphi$. Then, a direct and full participation mechanism $\tilde{\varphi}$ exists such that

1. Participation with probability 1 and truthtelling are a best response for the agent,
2. The distribution over $\Omega \times(\Theta \times A)^{\infty}$ induced by $(\varphi,(p, \sigma))$ is the same as that induced by $\tilde{\varphi}$ under participation and truthtelling.

Proof. Fix a mechanism $\varphi=\left(\varphi^{t}\right)_{t \geq 1}$ and a best response ( $p, \sigma$ ) for the agent in the sense of Equation 13. Let $\mathbb{P}_{\varphi,(p, \sigma)}$ denote the induced distribution over $\mathcal{H}^{\infty}$. We write $\left(\omega,\left(\theta_{t}, m_{t}, a_{t}\right)_{t \geq 1}\right)$ for a generic realization, where $m_{t}=\varnothing \Rightarrow a_{t}=a_{\varnothing}$.

The proof proceeds in three steps. In the first step, we construct a direct (but not full participation) mechanism $\tilde{\varphi}$, which under participation and truthtelling after every history on path implements the same outcome distribution as $(\varphi,(p, \sigma))$. In the second step, we verify that participation and truthtelling after every history on path is a best response to $\tilde{\varphi}$. In the third step, we construct a direct and full participation mechanism from $\tilde{\varphi}$. That the agent can always quit the mechanism at each step and obtain $a_{\varnothing}$ and Step 3 implies that participation and truthtelling after every history is also a best response to the full participation mechanism obtained from $\tilde{\varphi}$.

Step 1: We first construct the direct mechanism $\tilde{\varphi}_{t}: \hat{\mathcal{H}}^{t} \times \Theta \rightarrow \Delta(A)$. Define a collection of transition probabilities $\kappa_{t}: \mathcal{H}_{M}^{t} \times \Theta \rightarrow \Delta(M \cup\{\varnothing\})$ as follows:

$$
\kappa_{t}\left(m_{t} \mid h_{M}^{t}, \theta_{t}\right)=\left(1-p_{t}\left(h_{M}^{t}, \theta_{t}\right)\right) \mathbb{1}\left[m_{t}=\varnothing\right]+p_{t}\left(h_{M}^{t}, \theta_{t}\right) \sigma_{t}\left(h_{M}^{t}, \theta_{t}\right)\left(m_{t}\right) .
$$

We construct $\tilde{\varphi}$ recursively. In period 1, if the state is $\omega$ and the report is $\theta_{1}$, the designer draws fictitious $m_{1}$ from $\kappa_{1}\left(\cdot \mid \theta_{1}\right)$, and implements $a_{1}=a_{\varnothing}$ if $m_{1}=\varnothing$, and otherwise draws $a_{1} \sim \varphi_{1}\left(\omega, m_{1}\right)$.

Recursively, for $t \geq 2$, if the sequence of reports, fictitious messages, and allocations is $\left(\theta^{\prime t-1}, m^{t-1}, a^{t-1}\right)=$ $\left(\theta_{s}^{\prime}, m_{s}, a_{s}\right)_{s=1}^{t-1}$ and the agent reports $\theta_{t}$, the designer draws $m_{t}$ from $\kappa_{t}\left(\cdot \mid \theta^{\prime t-1}, m^{t-1}, a^{t-1}, \theta_{t}\right)$ and implements $a_{t}=a_{\varnothing}$ if $m_{t}=\varnothing$, and otherwise draws $a_{t} \sim \varphi_{t}\left(\omega, m^{t-1}, a^{t-1}, m_{t}\right) .{ }^{59}$

It is immediate that under truthtelling and participation the mechanism $\tilde{\varphi}$ implements the same distribution over $\Omega \times(A \times \Theta)^{\infty} .{ }^{60}$

Step 2: Let $\left(p^{*}, \sigma^{*}\right)$ denote the agent's strategy that participates and truthfully reports after every history. We now show that $\left(p^{*}, \sigma^{*}\right)$ is a best response to $\tilde{\varphi}$ in the sense of Equation D.8.

To do so, we show that for any strategy ( $\tilde{p}, \tilde{\sigma}$ ) in the game induced by the direct mechanism $\tilde{\varphi}$, a strategy $\left(p^{\prime}, \sigma^{\prime}\right)$ exists such that

$$
\mathbb{E}_{\tilde{\varphi},(\tilde{p}, \tilde{\sigma})}\left[U_{T}\right]=\mathbb{E}_{\varphi,\left(p^{\prime}, \sigma^{\prime}\right)}\left[U_{T}\right] \quad \text { for all } T,
$$

where $U_{T}$ is the average payoff through period $T$. Given Equation D. 9 and the best-response property of $(p, \sigma)$ to $\varphi$,

$$
\liminf _{T \rightarrow \infty} \mathbb{E}_{\varphi,(p, \sigma)}\left[U_{T}\right] \geq \limsup _{T \rightarrow \infty} \mathbb{E}_{\varphi,\left(p^{\prime}, \sigma^{\prime}\right)}\left[U_{T}\right] \quad \text { for all }\left(p^{\prime}, \sigma^{\prime}\right) .
$$

Using Step 1, we have $\mathbb{E}_{\varphi,(p, \sigma)}\left[U_{T}\right]=\mathbb{E}_{\tilde{\varphi},\left(p^{*}, \sigma^{*}\right)}\left[U_{T}\right]$ for all $T$, and by Equation D. 9 we have $\mathbb{E}_{\varphi,\left(p^{\prime}, \sigma^{\prime}\right)}\left[U_{T}\right]=$ $\mathbb{E}_{\tilde{\varphi},(\tilde{p}, \tilde{\sigma})}\left[U_{T}\right]$ for all $T$. Hence

$$
\liminf _{T \rightarrow \infty} \mathbb{E}_{\tilde{\varphi},\left(p^{*}, \sigma^{*}\right)}\left[U_{T}\right] \geq \limsup _{T \rightarrow \infty} \mathbb{E}_{\tilde{\varphi},(\tilde{p}, \tilde{\sigma})}\left[U_{T}\right] \quad \text { for all }(\tilde{p}, \tilde{\sigma}),
$$

which is exactly the definition of $\left(p^{*}, \sigma^{*}\right)$ being a best response to $\tilde{\varphi}$. It remains to construct $\left(p^{\prime}, \sigma^{\prime}\right)$ and verify (D.9).

We show how the agent can emulate the strategy ( $\tilde{p}, \tilde{\sigma}$ ) in the game induced by the indirect mechanism via strategy $\left(p^{\prime}, \sigma^{\prime}\right)$. Define $\tilde{\kappa}_{t}: H^{t} \times \Theta \rightarrow \triangle(\Theta \cup\{\varnothing\})$ as follows

$$
\tilde{\kappa}_{t}\left(\theta^{\prime} \mid h^{t}, \theta_{t}\right)=\left(1-\tilde{p}\left(h^{t}, \theta_{t}\right)\right) \mathbb{1}\left[\theta^{\prime}=\varnothing\right]+\tilde{p}\left(h^{t}, \theta_{t}\right) \tilde{\sigma}\left(h^{t}, \theta_{t}\right)\left(\theta^{\prime}\right) .
$$

The strategy ( $p^{\prime}, \sigma^{\prime}$ ) privately simulates the report process induced by ( $\tilde{p}, \tilde{\sigma}$ ) in the direct mechanism, and conditional on the fictitious reports, generates the actual message in the indirect mechanism $\varphi$ using the kernel $\kappa_{t}$ in Step 1, evaluated at the fictitious type history.

[^41]Formally, for $t \geq 1$ given the history of types, messages, and allocations through period $t,\left(\theta^{t-1}, m^{t-1}, a^{t-1}\right)$, and the privately tracked fictitious type reports $\theta^{\prime-1}$, ${ }^{61}$ the agent of type $\theta_{t}$ draws a fictitious report $\theta_{t}^{\prime} \sim \tilde{\kappa}_{t}\left(\cdot \mid \theta^{t-1}, \theta^{\prime \text { t }-1}, a^{t-1}, \theta_{t}\right)$. If $\theta_{t}^{\prime}=\varnothing$, then $m_{t}=\varnothing$ (the agent does not participate in period $t$ ). Otherwise, $\theta_{t}^{\prime} \in \Theta$ and $m_{t}$ is drawn from $\kappa_{t}\left(\cdot \mid \theta^{\prime-1}, m^{t-1}, a^{t-1}, \theta_{t}^{\prime}\right)$. The allocation is $a_{\varnothing}$ upon rejection, and $a_{t} \sim \varphi_{t}\left(\omega, m^{t-1}, a^{t-1}, m_{t}\right)$, otherwise.

It is immediate that Equation D. 9 holds and hence $\left(p^{*}, \sigma^{*}\right)$ is a best response to $\tilde{\varphi}:{ }^{62}$ In the extensive form game induced by $\tilde{\varphi}$, the designer simulates the agent's participation and reporting strategies using $(p, \sigma)$ based on the agent's type reports and determines allocations in the mechanism. When the agent's strategy is given by $(\tilde{p}, \tilde{\sigma})$, the process described in the above paragraph correspond to the designer's simulated participation and reporting strategies, and allocations continued to be determined by $\varphi$. Hence, in the extensive form game induced by $\tilde{\varphi}$, when the agent plays ( $\tilde{p}, \tilde{\sigma}$ ), it is as if she faces mechanism $\varphi$ and plays strategy ( $p^{\prime}, \sigma^{\prime}$ ). The best response property of ( $p, \sigma$ ) implies that playing $\left(p^{\prime}, \sigma^{\prime}\right)$ yields a weakly worse payoff, and hence $(\tilde{p}, \tilde{\sigma})$ is not a profitable deviation from $\left(p^{*}, \sigma^{*}\right)$ in the direct mechanism $\tilde{\varphi}$.

Step 3: Modify $\tilde{\varphi}$ at all histories that include at least one non-participation decision, so that the mechanism implements the outside option $a_{\varnothing}$. With this modification, the mechanism satisfies the full participation property. It is immediate that participation and truthtelling after every history remains a best response. $\square$

[^42]
[^0]:    *The latest version of the paper can be found here. We thank Dirk Bergemann, James Best, Elliot Lipnowski, Emir Kamenica, Navin Kartik, Stephen Morris , Phil Reny, Marzena Rostek, and Refine.ink for valuable comments and suggestions, as well as seminar audiences at EARIE 2024, WUSTL Economic Theory Conference 2024, MIT IDSS Distinguished Speaker Series, the Workshop in Market Design, Bonn, Berlin HU, CERGE-EI, Gerzensee, UCL, Naples, Stanford University, MIT Sloan, and the CEPR Workshop on Contracts, Incentives and Information at Collegio Carlo Alberto. Laura Doval gratefully acknowledges financial support from the Sloan Foundation. Alex Smolin gratefully acknowledges funding from the French National Research Agency (ANR) under the Investments for the Future program (grant ANR-17-EURE-0010) and the AI Interdisciplinary Institute ANITI (grant ANR-23-IACL-0002). This paper was partly written while the first author was visiting Stanford University and the second author was visiting Northwestern University and Columbia Business School; we are grateful for their hospitality.
    ${ }^{\dagger}$ Columbia Business School and CEPR. E-mail: laura.doval@columbia.edu.
    ${ }^{\ddagger}$ Toulouse School of Economics and CEPR. E-mail: alexey.v.smolin@gmail.com.

[^1]:    ${ }^{1}$ In (generalized) two-stage mechanisms, the designer communicates with the agents before the agents communicate with the mechanism. Whereas this communication is a restriction on the set of implementable outcomes relative to the single-designer Myersonian benchmark, Attar et al. (2025) show that allowing competing principals to first communicate with agents expands the set of implementable outcomes.

[^2]:    ${ }^{2}$ We assume the agent has limit-of-means preferences, so we can pass to the $\delta \rightarrow 1$ limit without approximation.
    ${ }^{3}$ Theorem 3 holds when the agent's type is drawn once at the beginning without further assumptions on the distribution. As we explain in Section 5, we choose the i.i.d. specification for the evolution of the agent's private information to put repeated and dynamic mechanisms on a more equal footing.

[^3]:    ${ }^{4}$ There is also a literature that studies a designer's disclosure of information that must be elicited from the agents (Eső and Szentes, 2007; Bergemann and Pesendorfer, 2007; Li and Shi, 2017; Krähmer, 2020; Bergemann et al., 2022a,b; Smolin, 2023). By contrast, the designer knows the realization of the state and also what information the two-stage mechanism discloses to the agents, so he need not elicit this information.

[^4]:    ${ }^{5}$ Assuming the allocation space is a product space is without loss of generality. Any restriction on the allocations, such as all agents must receive the same allocation, can be incorporated as restrictions on the support of the mechanism.
    ${ }^{6}$ Because we allow for lotteries over allocations, that the allocation space is finite does not preclude the case of transferable utility. Indeed, we could let each $A_{i}=\tilde{A_{i}} \times\{-K, K\}$ for some large enough $K>0$.

[^5]:    ${ }^{7}$ When the agent is infinitely patient as in Section 5, we can exhibit a sequence of strategies under which the agent (approximately) learns this mapping. See the proof of Lemma C. 4 in Appendix D.
    ${ }^{8}$ Gentzkow and Kamenica (2017) highlight that the language in Green and Stokey (2022) allows one to describe the correlation across signal structures. This is exactly what we need to allow the designer to obfuscate the agents' ability to learn. It is again instructive to consider the single-agent case. For each type report $\theta \in \Theta$, the mechanism can be seen as an information structure $\phi(\theta, \cdot): \Omega \times[0,1] \rightarrow \Delta(A)$. Thus, the randomization device allows the designer to control the correlation across these different signal structures, which in turn disciplines what the agent stands to learn when experimenting with different reports.
    ${ }^{9}$ The terminology is by analogy to reduced form auctions where the map from own types to own probabilities of being allocated the good are referred to as the interim allocation.

[^6]:    ${ }^{10}$ When we consider mechanisms with arbitrary message spaces in Appendix D.1, the requirement of calibration is relative to both the mechanism and agents' equilibrium participation and reporting strategies.
    ${ }^{11}$ Thus, we are assuming that $a_{\varnothing}=\left(a_{i \varnothing}\right)_{i \in N}$ is an element of $A$. In Appendix D.1, we consider more general participation decisions, allowing the mechanism to condition on the set of participating agents, but even with this extra generality, it is still without loss to restrict attention to mechanisms that induce full participation.

[^7]:    ${ }^{12}$ It is common for online platforms to inform agents of the consequences of their choices: marketplaces inform sellers of their probability of sale at different posted prices, and transportation providers inform riders of their probability of receiving a seat upgrade at different bid levels. Even insurance companies provide consumers with projected expenditures under different plan choices.

[^8]:    ${ }^{13}$ Indeed, in the analysis of the dynamic interaction in Section 5, we assume the agent only observes her type and her allocation, but not her payoffs.
    ${ }^{14}$ This assumption is routinely made in dynamic settings. See Pavan et al. (2014) and Cesa-Bianchi et al. (2024) for two examples in the context of agents' behavior within mechanisms.
    ${ }^{15}$ In Section 3, we provide an example in which when agents have access to the calibrated information structure the optimal Myersonian mechanism fails to be incentive compatible.

[^9]:    ${ }^{16}$ A tempting comparison is Maskin and Tirole (1990, Prop. 11): with private values and quasilinear utilities, the informed principal's unique equilibrium payoff coincides with the state-by-state optimum. The authors show this conclusion depends on quasilinearity: absent this assumption, an informed principal can benefit from concealing his information in the case of private values. Instead, Theorem 1 relies neither on quasilinearity nor on the designer's lack of commitment.

[^10]:    ${ }^{17}$ Formally,

    $$
    \mu_{0}(\omega \mid \theta)=\frac{\mu_{0}(\omega) f(\theta \mid \omega)}{\sum_{\omega^{\prime} \in \Omega} \mu_{0}\left(\omega^{\prime}\right) f\left(\theta \mid \omega^{\prime}\right)}
    $$

    ${ }^{18}$ We refer the reader to the appendix for our mathematical conventions, in particular, the definition of the corresponding $\sigma$-algebras.

[^11]:    ${ }^{19}$ Formally, define the belief distribution induced by $\beta, \tau_{\beta}=\mu_{0} \otimes \beta$. The claim is that $\mathbb{E}_{\tau_{\beta}}[\mu]=\mu_{0}$.
    ${ }^{20}$ This is a consequence of Bayes rule: beliefs are a sufficient statistic for $\omega$. Hence, conditional on $(\theta, \mu)$, the allocation rule carries no more information about the state.

[^12]:    ${ }^{21}$ In contrast to the model of Section 2, we are assuming the set of types and allocations to be intervals in the real line. The results in the previous sections go through with richer type and allocation spaces, at the cost of more notation.

[^13]:    ${ }^{22}$ This assumption ensures the Lipschitz continuity of the agent's indirect utility function when the allocation space is infinite. Because $\Omega$ is finite, requiring the condition to hold state-by-state suffices.

[^14]:    ${ }^{23}$ The calibrated mechanism which state-by-state implements the optimal direct mechanism under common knowledge of the state may reveal less than full information about the state, e.g., because at $\omega$ and $\omega^{\prime}$ the same direct mechanism is optimal. The point is that pooling those states does not weaken the incentive constraints of the agent, so it is as if the designer were forced to reveal the state.

[^15]:    ${ }^{24}$ Restricting attention to deterministic mechanisms, Ottaviani and Prat (2001) obtain the optimality of full disclosure without such a linearity assumption. Their model, however, is different from ours and that of the aforementioned papers: the agent's type is not payoff relevant and the agent's type and the state are affiliated. Thus, while related in spirit, Proposition 2 is distinct from their result.
    ${ }^{25}$ See Szabadi (2018) for a similar observation.

[^16]:    ${ }^{26}$ Because payoffs are quasilinear, considering mechanisms that do not randomize on transfers is without loss of generality.
    ${ }^{27}$ The equi-Lipschitz condition on $u$ ensures we can take the derivative inside the integral.

[^17]:    ${ }^{28}$ The designer can be viewed as an online advertising platform and the agent as an advertiser. State $\omega$ represents the click-through rate of an ad slot, and higher $\theta$ corresponds to a larger advertiser willing to pay more for exposure. The designer's payoff captures both the value created by advertising and the disutility from showing ads of large advertisers, e.g., due to user brand fatigue.

[^18]:    ${ }^{29}$ In the statement, $w_{k \theta_{i}}$ denotes the derivative of $w_{k}$ with respect to $\theta_{i}$, for $k \in\{1, \ldots, N\}$.
    ${ }^{30}$ To be sure, Theorem 3 extends to the case in which the agent's type is fully persistent and correlated with the state.

[^19]:    ${ }^{31}$ The Ionescu-Tulcea extension theorem implies this measure is always well-defined for any mechanism and any agent's

[^20]:    strategy. See Appendix C. 1 for details.
    ${ }^{32}$ Throughout this section, limits of measures should always be understood in the weak* sense.
    ${ }^{33}$ Note that in mechanism design one always focuses on mechanisms that have well-defined best responses in single-agent settings, and equilibria in multi-agent ones.

[^21]:    ${ }^{34}$ By assumption, the marginal of $\bar{v}_{\sigma}^{T}$ on $A \times \Theta \times M \times \tilde{\Omega}$ converges to $v_{\sigma}$. Moreover, we show that the marginal of $\bar{v}_{\sigma}^{T}$ on $\Delta(\tilde{\Omega})$ also converges (Lemma C.3). However, this is not enough to ensure the convergence of $\bar{v}_{\sigma}^{T}$.
    ${ }^{35}$ Whereas the above two-stage mechanism is described in terms of beliefs over $\Omega \times \mathcal{E}$, we show in the appendix how to derive from it a two-stage mechanism in terms of beliefs over $\Omega$.

[^22]:    ${ }^{36}$ This notation allows us to keep the definitions of the histories when the agent participates and does not participate symmetric, and saves us on including the agent's participation strategy in the histories.
    ${ }^{37}$ That is, starting from a dynamic mechanism $\varphi$ and a best response strategy ( $p, \sigma$ ), one can construct an alternative mechanism $\varphi^{\prime}$ such that participation and truthtelling are a best response for the agent and preserves the distribution over $(\Omega \times \Theta \times A)^{\infty}$ induced by $(p, \sigma)$ and $\varphi$.

[^23]:    ${ }^{38}$ This is easily seen in the example after Definition 7. Note the mechanism that allocates the good to the agent if and only if her type is $\theta_{2}$ satisfies monotonicity. Hence, a transfer scheme exists that implements this allocation rule with transfers.
    ${ }^{39}$ In a repeated principal-agent game with communication, Meng (2021) shows that the principal can guarantee in the patient limit his complete information payoff subject to the constraint that her actions satisfy the cyclic monotonicity condition in Rochet (1987). We view the results as complementary: We focus on implementable outcome distributions, instead of payoffs, when the agent has limit of the means preferences, which makes our notion of implementation exact.

[^24]:    ${ }^{40}$ A standard argument implies that if $\vartheta$ satisfies Equation 16, then a finite support $\beta^{\prime}$ exists such that $\vartheta$ satisfies Equation 16 with $\beta^{\prime}$ in place of $\beta$.

[^25]:    ${ }^{41}$ Let $\mu_{0}\left(\cdot \mid \theta_{i}\right) \in \Delta(\Omega)$ denote the prior of the agent with type $\theta_{i}$ and $\mu\left(\cdot \mid s_{i}\right)$ denote the updated belief of an agent with prior

[^26]:    $\mu_{0}$ upon observing signal $s_{i}^{*}$. When the signal is $s_{i}^{*}$, the agent with type $\theta_{i}$ updates her beliefs to:

    $$
    \mu_{i}\left(\cdot \mid \theta_{i}, s_{i}^{*}\right)=\frac{\mu_{0}\left(\cdot \mid \theta_{i}\right) \cdot \frac{\mu\left(\cdot \mid s_{i}^{*}\right)}{\mu_{0}(\cdot)}}{\left\|\mu_{0}\left(\cdot \mid \theta_{i}\right) \cdot \frac{\mu\left(\cdot \mid s_{i}^{*}\right)}{\mu_{0}(\cdot)}\right\|},
    $$

    where the ⋅ and / operations are meant componentwise, and $\|\cdot\|$ is the $l^{1}$-norm.
    ${ }^{42}$ Namely, Equation B. 5 implies that for all $(\theta, \omega), \vartheta(\cdot \mid \theta, \omega) \in \operatorname{clco}\{\alpha(\cdot \mid \theta, \mu): \mu \in \Delta(\Omega)\}$. Rubin and Wesler (1958) implies that our under assumptions $\operatorname{clco}\{\alpha(\cdot \mid \theta, \mu): \mu \in \Delta(\Omega)\}=\operatorname{co}\{\alpha(\cdot \mid \theta, \mu): \mu \in \Delta(\Omega)\}$, and the rest of the claim follows from Carathéodory's theorem.
    ${ }^{43}$ To extend the result to the case in which the agent's type is correlated with the state, note the following. Knowing $\mu_{0}$

[^27]:    updates to $\mu_{k}$ conditional on $s$ is enough to pin down the agent's belief $\mu(\cdot \mid \theta, s)$, with respect to which the agent's incentive compatibility and individual rationality constraints are defined (see footnote 41).

[^28]:    ${ }^{44}$ We could expand the mechanism by allowing the agent to have a message which triggers the outside option, but this is not necessary as $\alpha$ is individually rational.

[^29]:    ${ }^{45}$ Lemma C. 2 implies that $\bar{v}_{\sigma^{\prime}}^{T, 1}$ and $\bar{v}_{\sigma^{\prime}}^{T, 2}$ have the same set of subsequential limits. Indeed, let $g$

[^30]:    denote any continuous bounded function on $\Omega \times \Delta(\Omega) \times \Theta \times \Theta \times A$. Let $D_{T}(g)=\mathbb{E}_{\bar{v}_{\sigma^{\prime}}^{T, 2}}[g]-\mathbb{E}_{\bar{v}_{\sigma^{\prime}}^{T, 1}}[g]=$ $\frac{1}{T} \sum_{t=1}^{T} \mathbb{E}_{\sigma^{\prime}}\left[g\left(a_{t}, \theta_{t}, \theta_{t}^{\prime}, \omega, \mu_{t+1}\right)-g\left(a_{t}, \theta_{t}, \theta_{t}^{\prime}, \omega, \mu_{t}\right)\right]$. The argument in Lemma C. 2 implies that $D_{T}(g) \rightarrow 0$ as $T \rightarrow \infty$ (this does not rely on the existence of a limit, just the convergence of beliefs and the continuity of $g$ ). Now, let $T_{n}$ be such that $\bar{v}_{\sigma^{\prime}}^{T_{n}, 1} \xrightarrow{w^{*}} \tilde{v}$. Note that

    $$
    \int g d \bar{v}_{\sigma^{\prime}}^{T_{n}, 2}=\int g d \bar{v}_{\sigma^{\prime}}^{T_{n}, 1}+D_{T_{n}}(g) \rightarrow \int g d \tilde{v}+0,
    $$

    so a subsequential limit of $\bar{v}_{\sigma^{\prime}}^{T, 1}$ is a subsequential limit of $\bar{v}_{\sigma^{\prime}}^{T, 2}$. Switching the role of 1 and 2, we obtain the opposite set inclusion.

[^31]:    ${ }^{46}$ A set $X \subseteq \Delta(\Omega)$ is open relative to $\Delta(\Omega)$ if an open set $Y \subseteq \mathbb{R}^{|\Omega|}$ exists such that $X=Y \cap \Delta(\Omega)$. The boundary relative to $\Delta(\Omega)$ is analogously defined via open sets relative to $\Delta(\Omega)$.
    ${ }^{47}$ Lemma D. 1 provides an interval of radii $r \in\left(0, r_{0}\right)$ such that $\int_{B(\hat{\mu}, r)} U_{\text {net }}(\mu) \tau_{(p, \sigma)}(d \mu)<0$. Note that only countable many such $r$ can have $\tau_{(p, \sigma)}(\partial B(\hat{\mu}, r))>0$ (the boundaries for different radii are disjoint), so we can always pick $r$ such that $\tau_{(p, \sigma)}(\partial B(\hat{\mu}, r))=0$ and preserve the negative sign.

[^32]:    ${ }^{48}$ The property that $\tau_{(p, \sigma)}(\partial B)=0$ ensures that $\mathbb{1}\left[\mu_{t}\left(h^{\infty}\right) \in B\right]$ is eventually constant almost surely. Let $E=\left\{h^{\infty}: \mu_{\infty}\left(h^{\infty}\right) \notin\right.$ $\partial B\}$. On $E$, either $\mu_{\infty}$ is in the interior of $B$ (relative to $\Delta(\Omega)$ ) or in the interior of $B^{\complement}$ (relative to $\Delta(\Omega)$ ), that is, an $\epsilon>0$ exists such that $\left(B\left(\mu_{\infty}, \epsilon\right) \cap \Delta(\Omega)\right) \subset B$ or $B^{\complement}$. In either case, for each $h^{\infty}$, there exists $N\left(h^{\infty}\right)$ such that for all $t \geq N\left(h^{\infty}\right)$, $\mu_{t}\left(h^{\infty}\right) \in\left(B\left(\mu_{\infty}, \epsilon\right) \cap \Delta(\Omega)\right)$ and hence $\mathbb{1}\left[\mu_{t}\left(h^{\infty}\right) \in B\right]$ is eventually constant. When $\tau_{(p, \sigma)}(\partial B)=0$, we have that $E$ has probability 1 under $\mathbb{P}_{(p, \sigma)}$.
    ${ }^{49}$ Indeed, $\mathbb{E}_{(p, \sigma)}\left[\frac{1}{T} \sum_{t=1}^{T}\left(u\left(a_{t}, \theta_{t}, \omega\right)-u\left(a_{\varnothing}, \theta_{t}, \omega\right)\right) \mathbb{1}\left[\mu_{t} \in B\right]\right]=\mathbb{E}_{\bar{v}_{(p, \sigma)}^{T, 1}}\left[\left(u(a, \theta, \omega)-u\left(a_{\varnothing}, \theta, \omega\right)\right) \mathbb{1}[\mu \in B]\right]$ and our previous analysis implies it converges to $\int_{B} U_{\text {net }}(\mu) \tau_{(p, \sigma)}(d \mu)$.
    ${ }^{50}$ Choose $\bar{T}$ so that for all $T \geq \bar{T}$ :

    $$
    \begin{aligned}
    & \left|\mathbb{E}_{(p, \sigma)}\left[\frac{1}{T} \sum_{t=1}^{T}\left(u\left(a_{t}, \theta_{t}, \omega\right)-u\left(a_{\varnothing}, \theta_{t}, \omega\right)\right) \mathbb{1}\left[\mu_{t} \in B\right]\right]-\int_{B} U_{\text {net }}(\mu) \tau_{(p, \sigma)}(d \mu)\right| \\
    & +\left|\mathbb{E}\left[-\frac{1}{T} \sum_{t=1}^{L-1}\left(u\left(a_{t}, \theta_{t}, \omega\right)-u\left(a_{\varnothing}, \theta_{t}, \omega\right)\right) \mathbb{1}\left[\mu_{t} \in B\right]\right]\right| \leq \delta / 2 .
    \end{aligned}
    $$

[^33]:    ${ }^{51}$ Anticipating our construction in item 3, for each belief $\mu$ in the support of $\tau$, the allocation rule $\alpha^{\prime}(\cdot \mid \cdot, \mu)$ has no profitable undetectable deviations relative to payoff function $u^{\prime}(a, \theta)=\sum_{\omega \in \Omega} \mu(\omega) u(a, \theta, \omega)$.

[^34]:    ${ }^{52}$ Any $\|.\|_{p}$ would work, but larger $p$ results in weakly shorter adjustment phases.

[^35]:    ${ }^{53}$ In the expressions that follow, recall the full participation mechanism $\phi_{N}^{*}$ already averages over the agents' own randomization devices.

[^36]:    ${ }^{54}$ To be sure, the proof of Lemma C. 1 is written in the context of repeated mechanisms but it extends verbatim to dynamic mechanisms with simple notational adjustments.

[^37]:    ${ }^{55}$ To be sure,

    $$
    \mathbb{E}_{\sigma}\left[g\left(\mu_{t}\right)\right]=\int_{H^{\infty}} g\left(\mu_{t}\left(h^{\infty}\right)\right) \mathbb{P}_{\sigma}\left(d h^{\infty}\right)
    $$

[^38]:    ${ }^{56}$ The argument shows that the agent's equilibrium payoff is at least $U\left(\tau_{\phi}\right)$. However, it is immediate that $U\left(\tau_{\phi}\right)$ is the most the agent can make in the game as $\tau_{\phi}$ extracts all information from the mechanism.

[^39]:    ${ }^{57}$ This almost surely is under the law of $A$ under $\phi\left(\cdot \mid m, \tilde{\omega}^{\star}\right)$.

[^40]:    ${ }^{58}$ We index histories by the messages to distinguish these histories from those when the designer uses direct mechanisms.

[^41]:    ${ }^{59}$ Recall the fictitious reports encode the agent's participation decisions in the original mechanism.
    ${ }^{60}$ In fact, if we kept track of the designer's draws of fictitious messages, the new mechanism implements the same distribution over terminal histories $\mathcal{H}_{M}^{\infty}$, and a fortiori, its marginal over $\Omega \times(A \times \Theta)^{\infty}$ is the same.

[^42]:    ${ }^{61}$ Recall that to minimize notation and make history lengths symmetric across participation and nonparticipation, we record the agent's rejection of the mechanism as the empty message $\varnothing$.
    ${ }^{62}$ In fact, the construction ensures the stronger property that $\mathbb{P}_{\tilde{\varphi},(\tilde{p}, \tilde{\sigma})}^{T}=\mathbb{P}_{\varphi,\left(p^{\prime}, \sigma^{\prime}\right)}^{T}$, when in a slight abuse of notation we keep track of the fictitious messages in $\mathbb{P}_{\tilde{\varphi},(\tilde{p}, \tilde{\sigma})}^{T}$.

## Citation and provenance

**Authors:** Laura Doval; Alex Smolin

**Canonical citation:** Doval, Laura, and Alex Smolin. “Calibrated Mechanism Design.” TSE Working Paper 26-1718, 2026.

**Canonical machine-readable version:** https://alexsmolin.com/corpus/papers/calibrated-mechanism-design.md

**Source record:** https://alexsmolin.com/#research

**Attribution guidance:** Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.

Provenance metadata: https://alexsmolin.com/corpus/PROVENANCE.txt
