---
schema: "https://alexsmolin.com/corpus/schema.json"
work_id: "alex-smolin:persuasion-and-welfare"
paper_id: "alex-smolin:persuasion-and-welfare:2023-09-06"
title: "Persuasion and Welfare"
authors:
  - name: "Laura Doval"
    url: "https://www.laura-doval.com/"
  - name: "Alex Smolin"
    url: "https://alexsmolin.com/"
    orcid: "https://orcid.org/0000-0003-4740-2376"
manuscript_date: "2023-09-06"
language: "en"
version_type: "author-manuscript"
canonical_url: "https://alexsmolin.com/corpus/papers/persuasion-and-welfare.md"
source_record: "https://arxiv.org/abs/2109.03061"
doi: "https://doi.org/10.1086/729067"
citation: "Doval, Laura, and Alex Smolin. “Persuasion and Welfare.” Journal of Political Economy 132, no. 7 (2024): 2451–2487."
attribution_guidance: "Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable."
provenance_url: "https://alexsmolin.com/corpus/PROVENANCE.txt"
---

> Machine-readable author manuscript.
> Authors: Laura Doval; Alex Smolin.
> Canonical citation: Doval, Laura, and Alex Smolin. “Persuasion and Welfare.” Journal of Political Economy 132, no. 7 (2024): 2451–2487.
> Attribution and provenance: https://alexsmolin.com/corpus/PROVENANCE.txt

# Persuasion and Welfare

**Authors:** Laura Doval; Alex Smolin

**Manuscript date:** 2023-09-06

#### Abstract

Information policies such as scores, ratings, and recommendations are increasingly shaping society's choices in high-stakes domains. We provide a framework to study the welfare implications of information policies on a population of heterogeneous individuals. We define and characterize the Bayes welfare set, consisting of the population's utility profiles that are feasible under some information policy. The Pareto frontier of this set can be recovered by a series of standard Bayesian persuasion problems, in which a utilitarian planner takes the role of the information designer. We provide necessary and sufficient conditions under which an information policy exists that Pareto dominates the no-information policy. We illustrate our results with applications to data leakage, price discrimination, and credit ratings.


Keywords: Bayesian persuasion, information design, welfare economics, algorithms, information policies.

[^0]
## 1 Introduction

Information has increasingly become a tool for shaping society's choices in highstakes domains. Consider, for instance, the role of algorithms in making recommendations for bail (Angwin et al., 2016), recruiting (Raghavan et al., 2020; Li et al., 2020), health (Obermeyer et al., 2019), education (Kučak et al., 2018), and lending (Jagtiani and Lemieux, 2019), among others. The role of information as a policy instrument is not confined to the big data economy. Indeed, information policies in the form of scores and ratings have been in place long before algorithmic recommendations to determine school placement, promotions, and who receives credit. As society becomes reliant on information to guide decisions in policy-relevant domains, understanding the welfare implications of information-based policies becomes a first-order concern.

In this paper, we provide a framework to study the welfare impact of information policies in a population of heterogeneous agents. Formally, we study the following model. There is a unit mass population with types in a finite set, distributed according to a given prior distribution. We model an information policy as an information structure, which associates to each type a distribution over signals and hence, via Bayes' rule, a distribution over posterior beliefs (Kamenica and Gentzkow, 2011). Our primitive is a welfare function that represents for each (posterior) type distribution the welfare of individuals of a given type. Given the welfare function, each information structure induces a Bayes welfare profile, which describes for each type in the population their expected payoff under the information structure. We define the Bayes welfare set to be the set of all such profiles.

We characterize the Bayes welfare set and study its properties. The Bayes welfare set allows us to reduce society's choice of an information structure to the choice of a Bayes welfare profile. Instead of imposing properties that the information structure must satisfy, society's preferences over the welfare distribution in the population determine the properties of the chosen information structure. Our perspective thus complements that of the algorithmic fairness literature which remains agnostic about the population's payoffs, focusing instead on statistical properties of information structures, such as accuracy, parity, or fairness. It also complements that of the literature in Bayesian persuasion (Rayo and Segal, 2010; Kamenica and Gentzkow, 2011), which characterizes the (maximum) average welfare consistent with some information structure, but not necessarily the Bayes welfare profiles that give rise to the average welfare.

Theorem 1 characterizes the Bayes welfare set via the convex hull of a vector-valued function. In doing so, we extend the geometric characterizations of Aumann and Maschler (1995) and Kamenica and Gentzkow (2011) of the feasible set of ex ante payoffs to the characterization of the Bayes welfare set. Whereas a Bayes welfare profile depends on the distribution over posteriors conditional on each type, we show it can be alter-
natively expressed as the unconditional expectation over posteriors of a truth-adjusted payoff function, where the adjustment is proportional to the posterior likelihood ratio of each type. Evaluated at a given type, the truth-adjusted welfare function allows us to characterize the welfare individuals of a given type may obtain under some information structure. In turn, interpreting the truth-adjusted payoff function as a vectorvalued function allows us to capture the across-type restrictions imposed by Bayes' rule and precisely characterize the Bayes welfare set.

Theorem 2 characterizes the Pareto frontier of the Bayes welfare set. Points in the Pareto frontier are natural candidates for being the outcome of efficient bargaining over information structures or a social planner's choice. Theorem 2 shows the points in the Pareto frontier of the Bayes welfare set can be recovered by a series of standard Bayesian persuasion problems, in which a utilitarian planner takes the role of an information designer. We use Theorem 2 throughout the paper to characterize optimal information structures in specific applications. Leveraging Theorem 2, Corollary 3 provides a necessary and sufficient condition under which an information structure exists that Pareto dominates providing no information.

Theorem 3 enriches our characterization in the case in which the welfare function is equal to the expectation of a one-dimensional random variable, the support of which we call the reputation vector. This special case constitutes a natural benchmark and is commonly used in the literature on career concerns (Holmström, 1999), social image (Bénabou and Tirole, 2006, Tirole, 2021), and policy prediction problems (Mullainathan, 2018). Theorem 3 shows a welfare profile belongs to the Bayes welfare set if and only if it can be represented as the product between the reputation vector and a completely positive matrix that satisfies a version of Bayes plausibility. ${ }^{1}$ In addition, we show that the Bayesian persuasion problems that characterize the relative boundary of the Bayes welfare set correspond to instances of the problem in Rayo and Segal (2010). It follows that the information structures that induce welfare profiles on the relative boundary of the Bayes welfare set can be characterized using the graph-theoretic approach in Rayo and Segal (2010). We leverage their approach in Proposition 3, where we show that the information structures that maximize the welfare of a given type in the population correspond to a noisy version of the priority mechanisms studied in the matching literature (e.g., Celebi and Flynn, 2022).

Finally, we note that by interpreting our welfare function as an individual's typedependent payoff function, the Bayes welfare set is also the object of interest in more standard information design applications. For instance, the types may represent the private information of an informed principal who can commit to an information structure only after observing her type, as in Perez-Richet (2014) and Koessler and Skreta

[^1](Forthcoming). Similarly, in the study of mechanism design with limited commitment, Doval and Skreta (2022) describe the principal's mechanism as an information structure that must satisfy an informed agent's incentive constraints. Similar constraints appear in the studies of information design without commitment, as in Fréchette et al. (2022), Lipnowski and Ravid (2020), and Salamanca (2021), in the analysis of tests subject to participation constraints in Rosar (2017), and in the analysis of moral hazard in Saeedi and Shourideh (2020). Thus, the Bayes welfare set can be viewed as a unifying concept that underlies the incentive constraints the equilibrium information structure must satisfy. As we show in our first working paper version, Doval and Smolin (2021), our tools also open the door to the study of new problems in this literature.

Related Literature: Our work contributes to the literature on information design reviewed in the introduction. Starting from the work of Kamenica and Gentzkow (2011) and Rayo and Segal (2010), a series of papers investigate the limits imposed by common knowledge of Bayesian rationality (Aumann, 1987). Whereas the Bayesian persuasion literature studies the (maximum) average welfare that can be achieved under some information structure, we characterize instead the welfare profiles that are consistent with some information structure.

Whereas ours is the first characterization of the Bayes welfare set, a small literature studies certain Bayes welfare profiles within applications. Assuming the information designer is an informed principal, Perez-Richet (2014) refers to the payoff profile induced by an information structure as an interim payoff and studies the informed principal's preferred Bayes welfare profile. Recently, Galperti et al. (2023) study the Bayes welfare profile that gives rise to the sender's maximum average payoff and relate it to the Lagrange multiplier in the Bayes plausibility constraint. Restricting attention to the case in which the welfare function is linear in beliefs, Saeedi and Shourideh (2020) characterize a subset of the Bayes welfare set that satisfies certain incentive compatibility constraints, thus obtaining a different characterization. Finally, as the analysis below makes clear, a Bayes welfare profile depends on the distribution over posteriors induced by an information structure conditional on each type. Whereas Levy et al. (2021) and Arieli et al. (2022) characterize the set of conditional distributions over posteriors consistent with the prior, we follow a complementary approach that allows us to carry only the (unconditional) distribution over posteriors induced by an information structure (see Claim 1).

We also contribute to the economics literature that studies algorithmic fairness. Mullainathan (2018), Kleinberg et al. (2018), and Rambachan et al. (2020) argue for letting the social planner's objective determine the properties of algorithms. Our analysis is also related to Liang et al. (2022). Starting from a fixed joint distribution over groups, covariates, and states, Liang et al. (2022) model an algorithm as taking actions directly as a func-
tion of covariates and study the group error profiles in the fairness-accuracy frontier as the algorithm varies. Instead, we model an algorithm as an information structure that sends non-binding action recommendations to an unmodeled receiver, whose actions determine the population's welfare and characterize the set of all Bayes welfare profiles as we vary the algorithm.

By considering the welfare redistribution effects of information, our work joins the mechanism design literature that studies the role of markets in redistributing welfare (see, e.g., Dworczak et al., 2021, Akbarpour et al., Forthcoming, and Akbarpour et al., 2023). We complement this work, which typically assumes the planner can utilize transfers to achieve its objectives, by considering the role of information, which can be a powerful tool when the planner does not have access to transfers.

Finally, our work contributes indirectly to the literature on higher-order beliefs. Indeed, when the welfare function is linear in beliefs as in Section 5, the welfare profile can be seen as a profile of second-order expectations. Starting with Samet (1998), a body of work uses Markov matrices to represent such higher-order beliefs and expectations of higher-order beliefs for a given information structure (see, e.g., Cripps et al., 2008; Golub and Morris, 2017). Instead, our result in Theorem 3 identifies the set of matrices that correspond to some information structure.

In lieu of an organizational paragraph, we summarize below the notation used throughout the paper:

Notation: For ease of presentation, we sometimes find it convenient to denote a function from a set $\Theta$ to $\mathbb{R}$ as a vector in $\mathbb{R}^{N}$, where $N$ is the cardinality of $\Theta$. In this case, we reserve the italic notation $x$ for the function $x: \Theta \mapsto \mathbb{R}$ and the upright notation x for the vector in $\mathbb{R}^{N}$. Any vector $\mathrm{x} \in \mathbb{R}^{N}$ is taken to be a column vector; we denote its $i^{\text {th }}$ component by $\mathrm{x}_{i}$ or $x\left(\theta_{i}\right)$ interchangeably. If $\mathrm{x} \in \mathbb{R}^{N}$ is a column vector, $\mathrm{x}^{T}$ denotes its transpose. If $\mathrm{x}, \mathrm{y}$ are two vectors, $\mathrm{x} * \mathrm{y}$ denotes their Hadamard (elementwise) product and x/y denotes their Hadamard division. We denote by $\mathrm{e} \in \mathbb{R}^{N}$ the vector with $\mathrm{e}_{1}=\cdots=\mathrm{e}_{N}=1$. When we want to emphasize that $x$ is a random variable, we write it as $\tilde{x}$.

## 2 Model

A unit mass population has types in a finite set, $\Theta \equiv\left\{\theta_{1}, \ldots, \theta_{N}\right\}$. Letting $\Delta(\Theta)$ denote the set of probability distributions over $\Theta$, we denote by $\mu_{0} \in \Delta(\Theta)$ the frequency of types in the population. We assume that $\mu_{0}$ has full support. We denote by $\Delta(\Delta(\Theta))$ the set of distributions over posteriors, and by $\Delta_{\mu_{0}}(\Delta(\Theta))$ the set of distributions over posteriors with mean equal to the prior $\mu_{0}$.

Welfare function: An individual's welfare depends on her type $\theta$ and an (unmodeled) outside observer's belief about her type. We represent this by a welfare function $w: \Delta(\Theta) \times \Theta \mapsto \mathbb{R}$ that represents for each belief $\mu$ and each type $\theta$, the welfare of individuals of type $\theta$ under belief $\mu, w(\mu, \theta)$. We assume throughout that $w$ is bounded.

A welfare profile is a vector $\mathrm{w} \in \mathbb{R}^{N}$, where $\mathrm{w}_{i}$ describes the welfare level of individuals with type $\theta_{i}$. Any welfare profile w induces an ex ante welfare of $\mu_{0}^{T} \mathrm{w}=$ $\sum_{i=1}^{N} \mu_{0}\left(\theta_{i}\right) \mathrm{w}_{i}$. When we average a profile w using weights other than the prior $\mu_{0}$, we refer to average welfare instead. We are interested in characterizing those welfare profiles that are induced by some information structure.

Information structures: An information structure $\Pi=(\pi, S)$ consists of a countable set of labels $S$, and a mapping $\pi$, which associates to each type $\theta$ a distribution over signals $\pi(\cdot \mid \theta) \in \Delta(S)$. Given an information structure $\Pi$ and a signal realization $s \in S$, the corresponding posterior belief $\mu_{s} \in \Delta(\Theta)$ is obtained by Bayes' rule whenever possible, and is given by

$$
\mu_{s}(\theta)=\frac{\mu_{0}(\theta) \pi(s \mid \theta)}{\sum_{\theta^{\prime} \in \Theta} \mu_{0}\left(\theta^{\prime}\right) \pi\left(s \mid \theta^{\prime}\right)} .
$$

Thus, an information structure can be seen as inducing a distribution over posterior beliefs $\left\{\mu_{s}: s \in S\right\}$. In what follows, two such distributions are of interest: the distribution over posterior beliefs conditional on an individual's type-as induced by $\pi(\cdot \mid \theta)$ and the unconditional distribution over posterior beliefs-as induced by the prior $\mu_{0}$ and the signal distribution. When taking expectations using these distributions, we use the notations $\langle\Pi \mid \theta\rangle$ and $\langle\Pi\rangle$ to denote the conditional and unconditional distributions over posterior beliefs, respectively.

Bayes welfare profiles: The welfare function $w$ together with an information structure, $\Pi$, defines a welfare profile, $w_{\Pi}: \Theta \mapsto \mathbb{R}$, as

$$
w_{\Pi}(\theta) \equiv \mathbb{E}_{\langle\Pi \mid \theta\rangle}[w(\tilde{\mu}, \theta)]=\sum_{s \in S} \pi(s \mid \theta) w\left(\mu_{s}, \theta\right) .
$$

That is, for each type $\theta, w_{\Pi}(\theta)$ describes the (expected) welfare of type- $\theta$ individuals under information structure $\Pi$. Note that in computing the welfare of type- $\theta$ individuals, their type $\theta$ enters twice: directly through the welfare function, $w(\cdot, \theta)$, and indirectly through the signal distribution, $\pi(\cdot \mid \theta)$.

We now present our two main objects of study:
Definition 1 (Bayes welfare profile). A welfare profile $\mathrm{w} \in \mathbb{R}^{N}$ is a Bayes welfare profile if an information structure, $\Pi$, exists such that for all types $\theta_{i}, \mathrm{w}_{i}=w_{\Pi}\left(\theta_{i}\right)$.

Definition 2 (Bayes welfare set). The Bayes welfare set is the set of all Bayes welfare profiles; that is,

$$
\mathrm{W} \equiv\left\{\mathrm{w} \in \mathbb{R}^{N}: \exists \Pi \text { s.t. } \mathrm{w}_{i}=w_{\Pi}\left(\theta_{i}\right) \forall i \in\{1, \ldots, N\}\right\} .
$$

The Bayes welfare set W represents the utility possibility set in an economy where the allocations are given by information structures. As such, it describes the welfare effects that different information structures have for individuals with different types in applications such as grading schemes in the case of schooling (Ostrovsky and Schwarz, 2010), disclosure about job performance (Mukherjee, 2008), affirmative action in the case of college admissions or the job market, rating systems in the case of platforms (Saeedi and Shourideh, 2020), and market segmentations (Bergemann et al., 2015).

Throughout, we illustrate our results using the following examples:
Example 1 (Data leakage). Consumers concerned about how a third party may use their data wish to maximize the third party's uncertainty about their types. ${ }^{2}$ We formalize this as follows. There are two types of consumer, $\Theta=\left\{\theta_{A}, \theta_{B}\right\}$. Letting $\mu_{0}$ denote the frequency of consumers of type $\theta_{B}$ (i.e., $\mu_{0} \equiv \mu_{0}\left(\theta_{B}\right)$ ), we assume that $\mu_{0}=0.3$. When the third party believes the consumer's type is $\theta_{B}$ with probability $\mu \in[0,1]$, a consumer experiences a welfare loss of $w(\mu, \theta)=-(\mu-1 / 2)^{2}$. The consumer's welfare loss is minimal when the third party is maximally confused about the consumer's type (i.e., $\mu=\frac{1}{2}$ ). Instead, the consumer's welfare loss is maximal when the third party has precise information about the consumer's type (i.e., $\mu \in\{0,1\}$ ). ◇

Example 2 (Price discrimination). An online marketplace makes algorithmic recommendations to a seller about what price the seller should set for consumers. There are two consumer types, $\theta_{H}$ and $\theta_{L}$. Assume $\mu_{0} \equiv \mu_{0}\left(\theta_{L}\right)=0.6$. Consumers can have one of three values for the seller's good: low (1), medium (2), and high (3). A consumer's type indexes their distribution over values, with $\theta_{H}$-consumers being more likely to have high valuations, and $\theta_{L}$-consumers being more likely to have low valuations. In particular, we assume the likelihoods of the values \{1, 2, 3\} are $\left\{\frac{1}{5}, \frac{1}{5}, \frac{3}{5}\right\}$ and $\left\{\frac{3}{5}, \frac{1}{5}, \frac{1}{5}\right\}$ for $\theta_{H^{-}}$and $\theta_{L^{-}}$-consumers, respectively.

The platform can provide information only about the consumer's type and not about their value. Hence, the seller's price depends on the likelihood $\mu$ the seller attaches to the consumer's type being $\theta_{L}$. Assuming the seller breaks ties in favor of consumers,

[^2]consumers' welfare as a function of the seller's belief $\mu$ and their type $\theta$ is as follows:
$$
w\left(\mu, \theta_{H}\right)=\left\{\begin{array}{ll}
0 & \text { if } \mu<1 / 2 \\
3 / 5 & \text { if } \mu \in[1 / 2,3 / 4) \\
7 / 5 & \text { if } \mu \in[3 / 4,1]
\end{array}, w\left(\mu, \theta_{L}\right)=\left\{\begin{array}{ll}
0 & \text { if } \mu<\frac{1}{2} \\
1 / 5 & \text { if } \mu \in[1 / 2,3 / 4) \\
3 / 5 & \text { if } \mu \in[3 / 4,1]
\end{array} .\right.\right.
$$ $\square$

Example 3 (Credit ratings). A credit agency makes lending decisions based on an applicant's perceived repayment probability. An applicant of type $\theta$ repays loans with probability $\rho(\theta)$. The credit agency approves loans with probability proportional to the expected value of $\rho$. A regulator wishes to maximize the probability that applicants of a given type $\theta_{i}$ receive a loan by choosing the information on which the credit agency can condition its approval decision. $\square$

Remark 1 (Interpretation of the welfare function). The welfare function admits several interpretations. First, following Kamenica and Gentzkow (2011), the population's welfare may be determined by the actions taken by the outside observer after observing the realization of an information structure, as in Examples 2 and 3. Using the notation in that paper, denote by $v(a, \theta)$ the utility of type- $\theta$ individuals when the outside observer takes action $a \in A$. If the outside observer takes action $a(\mu)$ when her posterior belief is $\mu$, type- $\theta$ individuals obtain payoff $v(a(\mu), \theta) \equiv w(\mu, \theta) .^{3}$ Second, the welfare function may also capture that the population's welfare may be driven by image or reputation concerns, as in Bénabou and Tirole (2006) and Tirole (2021), or psychological motives, as in Lipnowski and Mathevet (2018). Indeed, Example 1 admits a psychological interpretation under which an individual derives utility from keeping the outside observer, who is attempting to guess the individual's type, in suspense (cf. Ely et al., 2015).

## 3 Characterization

Section 3 presents our characterization of the Bayes welfare set via the convex hull of the graph of a vector-valued function, in the spirit of the belief-based approach of Kamenica and Gentzkow (2011).

Truth-drifting: An apparent obstacle in following the belief approach in Kamenica and Gentzkow (2011) is that the elements of W are expressed in terms of expectations conditional on a given type $\theta \in \Theta$, rather than unconditional expectations. Indeed, as shown in Francetich and Kreps (2014), conditional expectations do not satisfy the martingale

[^3]property; rather, they drift toward the truth. More precisely, for any type $\theta$ and for any information structure $\Pi$, the expectation of the posterior probability of $\theta$ conditional on $\tilde{\theta}=\theta$ is higher than the prior probability of $\theta$. That is, ${ }^{4}$
$$
\mathbb{E}_{\langle\Pi \mid \theta\rangle}\left[\frac{\tilde{\mu}(\theta)}{\mu_{0}(\theta)}\right]=\sum_{s \in S} \pi(s \mid \theta) \frac{\mu_{s}(\theta)}{\mu_{0}(\theta)} \geq 1 .
$$
Instead of pursuing a characterization of conditional distributions of posteriors, we recover the belief approach in Kamenica and Gentzkow (2011) by studying a suitably modified welfare function.

Truth-adjusted welfare: We show any element $\mathrm{w} \in \mathrm{W}$ can be expressed as the unconditional expectation of an adjusted version of the welfare function. Indeed, define the truth-adjusted welfare function $\hat{w}: \Delta(\Theta) \times \Theta \mapsto \mathbb{R}$ to be

$$
\hat{w}(\mu, \theta) \equiv \frac{\mu(\theta)}{\mu_{0}(\theta)} w(\mu, \theta) .
$$

That is, $\hat{w}$ is the welfare function $w$ adjusted by the truth-drift $\mu(\theta) / \mu_{0}(\theta)$. For any given posterior belief $\mu$, the likelihood ratio $\mu(\theta) / \mu_{0}(\theta)$ measures the representation of type $\theta$ under $\mu$ relative to its ex ante representation under $\mu_{0}$.

The truth-adjusted welfare function combines the preferences of individuals of type $\theta$ for a particular belief- $w(\mu, \theta)$-and the resource constraint-Bayes plausibility-in our economy, where information structures take the role of allocations. Indeed, the likelihoodratio adjustment $\mu(\theta) / \mu_{0}(\theta)$ captures the constraint that comes from Bayes plausibility. Intuitively, type- $\theta$ individuals have an endowment equal to $\mu_{0}(\theta)$ that can be spread over different beliefs $\mu(\theta)$, and the Bayes plausibility constraint ensures this spread is done in a way that respects the budget. ${ }^{5}$

Example 1 (continued). We illustrate the truth-adjusted welfare function in the context of Example 1. Figure 1 depicts the welfare function $w(\mu, \theta)$ (Figure 1a) and the truth-adjusted welfare function for $\theta_{A}$ (Figure 1b) and $\theta_{B}$ (Figure 1c). Whereas in this example welfare is assumed to be type-independent, the truth-adjusted welfare function is type-dependent. This natural consequence of the likelihood-ratio adjustment reflects that individuals of different types get to benefit differently from various induced beliefs and therefore, from the same information structure. $\square$

Claim 1 justifies our interest in the truth-adjusted welfare function $\hat{w}$ :

[^4]

> [Figure omitted from this text-only corpus; refer to the source manuscript.]
Figure 1: Truth-adjusted welfare function in Example 1

Claim 1 (From conditional to unconditional expectations). For any information structure $\Pi$ and any type $\theta \in \Theta$, the following holds:

$$
w_{\Pi}(\theta)=\mathbb{E}_{\langle\Pi \mid \theta\rangle}[w(\tilde{\mu}, \theta)]=\mathbb{E}_{\langle\Pi\rangle}[\hat{w}(\tilde{\mu}, \theta)] .
$$

The proof of Claim 1 and of other results can be found in the appendix.
Claim 1 implies the expectation of $w$ under $\Pi$ conditional on type $\theta$ can be expressed as the unconditional expectation of $\hat{w}$ under $\Pi$. Because the definition of $\hat{w}$ does not depend on $\Pi$, the distribution over posteriors induced by $\Pi,\langle\Pi\rangle$, is enough to determine the unconditional expectation of $\hat{w}$ under $\Pi$. Consequently, the analysis that follows relies on the unconditional distribution over posteriors induced by an information structure $\Pi$, rather than on the family of conditional distributions over posteriors induced by $\Pi$.

Claim 1 can be obtained from the analysis in Alonso and Câmara (2016). ${ }^{6}$ Indeed, for a given type $\theta$, one can interpret our model as one of Bayesian persuasion with heterogeneous priors in which the sender assigns probability 1 to state $\theta$ and the receiver's prior belief is $\mu_{0}$. Alonso and Câmara (2016, pp. 683-684) show the sender's payoff under an information structure $\Pi$ can be equivalently obtained as the expectation over the distribution of the receiver's posterior beliefs induced by $\Pi$ of an adjusted payoff function. Specialized to the case in which the sender assigns probability 1 to state $\theta$ and the receiver's prior belief is $\mu_{0}$, the truth-adjusted welfare function, $\hat{w}(\mu, \theta)$, corresponds to the adjusted payoff function in Alonso and Câmara (2016).

Relying on Claim 1, we can immediately characterize the range of welfare values that individuals of type $\theta$ may obtain under some information structure, via the expected

[^5]value of $\hat{w}(\cdot, \theta)$ under a Bayes plausible distribution over posteriors (Aumann and Maschler, 1995; Kamenica and Gentzkow, 2011). Indeed, let $w_{*}(\theta), w^{*}(\theta)$ denote the minimum and maximum welfare individuals of type $\theta$ can obtain under some information structure. That is, $w_{*}(\theta)=\inf \{w(\theta): \mathrm{w} \in \mathrm{W}\}$ and $w^{*}(\theta)=\sup \{w(\theta): \mathrm{w} \in \mathrm{W}\}$. We have the following:

Proposition 1 (Individually feasible welfare bounds). For any type $\theta$, the following holds: ${ }^{7}$

$$
w_{*}(\theta)=\operatorname{vex} \hat{w}\left(\mu_{0}, \theta\right), w^{*}(\theta)=\operatorname{cav} \hat{w}\left(\mu_{0}, \theta\right) .
$$

Proposition 1 follows from the main result in Kamenica and Gentzkow (2011). The vertical solid lines in Figures 1b and 1c illustrate the individually feasible welfare values for $\theta_{A}$ and $\theta_{B}$ in Example 1. An implication of Proposition 1 is that any Bayes welfare profile w satisfies that for all types $\theta, w(\theta) \in\left[\operatorname{vex} \hat{w}\left(\mu_{0}, \theta\right)\right.$, cav $\left.\hat{w}\left(\mu_{0}, \theta\right)\right]$.

Whereas Proposition 1 characterizes what is individually feasible for each type in the population, it does not deliver the characterization of the Bayes welfare set. The reason is that it ignores the across-type restrictions imposed by Bayes' rule. For instance, the truth-adjusted welfare function (AW) highlights that only types on the support of belief $\mu$ get to enjoy the payoff of inducing said belief. Similarly, inspection of Figures 1b and 1c shows $\theta_{A}$ 's preferred information structure is no disclosure, whereas $\theta_{B}$ would prefer some disclosure to no disclosure. In other words, the profile $\left(w^{*}\left(\theta_{A}\right), w^{*}\left(\theta_{B}\right)\right)$ is not jointly feasible.

Instead, the characterization of the Bayes welfare set can be obtained by studying the convex hull of the graph of the vector-valued function $\hat{\mathrm{w}}, \hat{\mathrm{w}}: \Delta(\Theta) \mapsto \mathbb{R}^{N}$, where for each $i \in\{1, \ldots, N\}, \hat{\mathrm{w}}_{i}(\mu) \equiv \hat{w}\left(\mu, \theta_{i}\right)$. Indeed, we have the following:

Theorem 1 (Belief-based characterization). The Bayes welfare set W satisfies the following:

$$
\mathrm{W}=\left\{\mathrm{w} \in \mathbb{R}^{N}:\left(\mu_{0}, \mathrm{w}\right) \in \operatorname{co}(\operatorname{graph} \hat{\mathrm{w}})\right\} .
$$

Theorem 1 provides a geometric characterization of the set W : it is the section at the prior of the convex hull of the graph of the truth-adjusted welfare function $\hat{\mathrm{w}}$. Relying on the result in Kamenica and Gentzkow (2011) that any Bayes plausible distribution over posteriors is the outcome of some information structure, ${ }^{8}$ Theorem 1 characterizes a more primitive object, the set of welfare profiles that can be generated by some

[^6]information structure. Indeed, whereas the main result in Kamenica and Gentzkow (2011) would allow us to characterize the (maximal) ex ante welfare a population with welfare function $w$ can obtain, Theorem 1 characterizes the welfare profiles whose average leads to that welfare.

Calculating the Bayes welfare set: Relying on Theorem 1, Figure 2 illustrates the construction of the Bayes welfare set in Example 1. Figure 2a depicts the convex hull of the graph of $\hat{\mathrm{w}}$. Applying Theorem 1, the resulting Bayes welfare set is the section of this convex hull at $\mu_{0}=0.3$. The blue shaded area in Figure 2a depicts this section, which is represented in Figure 2b.

Figure 2b illustrates which welfare profiles are jointly feasible under some information structure in Example 1. For instance, fully revealing or concealing individuals' types is always possible, so that the full and no-disclosure profiles, $\mathrm{w}^{F D}$ and $\mathrm{w}^{N D}$, are feasible. All Bayes welfare profiles Pareto dominate the full-disclosure profile, whereas no Bayes welfare profile Pareto dominates the no-disclosure one. As discussed above, simultaneously giving all individual types their maximum welfare $w^{*}(\theta)$ is not possible, because $\theta_{B}$ 's welfare is maximized by a policy that sometimes reveals an individual is of type $\theta_{A}$.

Because the welfare function is continuous in Example 1, the Bayes welfare set is closed, but this property does not follow from the ongoing assumption that the welfare function is bounded. However, making assumptions other than that the welfare function is bounded may not be natural. To illustrate, consider the case in which the welfare function captures in reduced form that the welfare of the population is determined by the outside observer's actions after observing the realization of the information structure. As we explained in Remark 1, each selection from the outside observer's best-response correspondence induces $a$ welfare function and hence a corresponding Bayes welfare set. If one assumes-as we do in Example 2-the outside observer breaks ties in favor of the individuals, one would naturally obtain an upper-semicontinuous welfare function, but under adversarial tie-breaking, having a lower-semicontinuous welfare function would have made sense. As we illustrate in Section 4, for a fixed tie-breaking rule, the corresponding Bayes welfare set may not be closed (see, e.g., Figure 3b). When one instead considers the welfare implications of different selection rules, the analogue of the Bayes welfare set is the set of all profiles that are induced by some information structure and some selection from the best-response correspondence. As we show in Proposition A. 1 in the appendix, the characterization behind Theorem 1 delivers that this analogue of the Bayes welfare set is closed.

Cardinality: Theorem 1 has an immediate implication for the cardinality of the information structures that generate points in W :

> [Figure omitted from this text-only corpus; refer to the source manuscript.]
Figure 2: Constructing the Bayes welfare set in Example 1

Corollary 1. Let $\mathrm{w} \in \mathrm{W}$. Then, an information structure $\Pi$ with at most $2 N$ signals exists such that $\mathrm{w}_{i}=w_{\Pi}\left(\theta_{i}\right)$ for all $i \in\{1, \ldots, N\}$.

Corollary 1 stands in contrast to the result in Bayesian persuasion that finding an information structure that delivers the sender's maximal payoff and employs at most $N$ posteriors is always possible. There are two reasons behind this difference. First, Theorem 1 characterizes all welfare profiles consistent with some information structure so that without further assumptions on the truth-adjusted welfare function, we rely on Carathédory's theorem to obtain Corollary 1. Instead, Kamenica and Gentzkow (2011) characterize the sender's maximal payoff and assume the indirect utility function is upper-semicontinuous. For that reason, Kamenica and Gentzkow (2011) can rely on Fénchel-Bunt's theorem instead of Carathéodory to obtain an upper bound of $N$ instead of $N+1$. Second, we are interested not just in the payoff of one "sender," but in the payoff of $N$, one for each type.

In Example 1, the upper bound in Corollary 1 is loose. Because the truth-adjusted welfare function is continuous, we can rely on Fénchel-Bunt's theorem to reduce the upper bound in Corollary 1 by 1. Furthermore, the Bayes welfare profiles on the boundary of W are induced by information structures with at most two signals. We explain the reason for this further reduction in Section 4, where we characterize the boundary of the Bayes welfare set and, in particular, its Pareto frontier.

## 4 The Pareto frontier of the Bayes welfare set

We characterize in this section the Pareto frontier of W. Points in the Pareto frontier are natural candidates for being the outcome of efficient bargaining over information structures or a social planner's choice. Indeed, these points correspond to the solution of a utilitarian planner as we vary the weights the planner assigns to different types. Theorem 2 shows these points can be recovered as solutions to standard Bayesian persuasion problems, where an information designer takes the role of the utilitarian planner. Armed with this characterization, Corollary 3 provides a necessary and sufficient condition for the no-disclosure profile to be part of the Pareto frontier. In other words, it provides a necessary and sufficient condition for information disclosure to lead to a Pareto improvement relative to no disclosure.

We define the Pareto frontier of W to be the set of weak Pareto efficient Bayes welfare profiles. Formally,

$$
\mathrm{W}_{\mathrm{P}}=\left\{\mathrm{w} \in \mathrm{~W}:\left(\nexists \mathrm{w}^{\prime} \in \mathrm{W}\right) \mathrm{w}^{\prime}>\mathrm{w}\right\} .
$$

Because the Bayes welfare set W is convex, for any $\mathrm{w} \in \mathrm{W}_{\mathrm{P}}$, the separating hyperplane theorem implies a direction $\lambda \in \mathbb{R}_{+}^{N} \backslash\{0\}$ exists such that ${ }^{9}$

$$
\begin{aligned}
\lambda^{T} \mathrm{w}=\max \left\{\lambda^{T} \mathrm{w}^{\prime}: \mathrm{w}^{\prime} \in \mathrm{W}\right\} & =\max \left\{\lambda^{T} \mathbb{E}_{\tau}[\hat{\mathrm{w}}(\tilde{\mu})]: \tau \in \Delta_{\mu_{0}}(\Delta(\Theta))\right\} \\
& =\max \left\{\mathbb{E}_{\tau}\left[\lambda^{T} \hat{\mathrm{w}}(\tilde{\mu})\right]: \tau \in \Delta_{\mu_{0}}(\Delta(\Theta))\right\} .
\end{aligned}
$$

The first equality simply states that w is a maximizer of the support function of the Bayes welfare set in direction $\lambda$. Instead, the second equality follows from Theorem 1. Indeed, Theorem 1 implies we can exchange the maximization over welfare profiles in W for a maximization over Bayes plausible distributions over posteriors. Note we can interchangeably talk about Pareto efficient Bayes welfare profiles and Pareto efficient information structures, and we do this in what follows.

Once we note we can restrict attention to directions $\lambda \in \Delta(\Theta)$, Equation 6 has two economic interpretations. First, consider the problem of a social planner who assigns weight $\lambda(\theta)$ to type $\theta$ and wishes to maximize the weighted sum of utilities of each type. Under this interpretation, Equation 6 states that w is a solution to the social planner's problem. Second, we can interpret $\lambda^{T} \mathrm{w}$ as the expectation with respect to $\theta$ of the welfare profile w under the measure $\lambda$. In this case, Equation 6 implies w is the vector of interim payoffs of a sender with payoff function $w(\mu, \theta)$ and prior $\lambda$. For instance, when $\lambda=\mu_{0}$, so that the sender's prior coincides with that of the outside observer, the above problem coincides with that of Kamenica and Gentzkow

[^7](2011). Instead, whenever the direction $\lambda$ is any element of $\Delta(\Theta)$, the above problem coincides with that considered by Alonso and Câmara (2016). ${ }^{10}$

Moreover, Equation 6 has an important practical implication: any Pareto efficient Bayes welfare profile is induced by the solution to a supporting Bayesian persuasion problem. A supporting Bayesian persuasion problem is an instance of the model in Kamenica and Gentzkow (2011) in which the sender's indirect utility function equals

$$
\hat{v}_{\lambda}(\mu)=\lambda^{T} \hat{\mathrm{w}}(\mu)=\sum_{\theta \in \Theta} \lambda(\theta) \frac{\mu(\theta)}{\mu_{0}(\theta)} w(\mu, \theta)=\sum_{\theta \in \Theta} \mu(\theta) \frac{\lambda(\theta)}{\mu_{0}(\theta)} w(\mu, \theta) .
$$

This indirect utility is the product of the truth-adjusted welfare function and the Pareto weights, which capture the rate of substitution between the truth-adjusted welfare of different types. Theorem 2 summarizes the above discussion:

Theorem 2 (Pareto frontier). The welfare profile $\mathrm{w} \in \mathbb{R}^{N}$ is in the Pareto frontier of W if and only if a direction $\lambda \in \Delta(\Theta)$ exists such that w is the profile induced by an information structure that solves the supporting Bayesian persuasion problem with indirect utility function $\hat{v}_{\lambda}$.

Because W is convex, the characterization in Theorem 2 extends to all Bayes welfare profiles on the boundary of W, except that the supporting direction may no longer be in $\Delta(\Theta)$.

Theorem 2 characterizes the Pareto frontier of W by connecting the solution of the utilitarian planner with weights $\lambda$ to the Bayesian persuasion problem of the sender with indirect utility function $\hat{v}_{\lambda}$. Whereas Theorem 1 characterizes the Bayes welfare set via the convex hull of the graph of a vector-valued function, $\hat{\mathrm{w}}$, the results in Kamenica and Gentzkow (2011) imply the supporting Bayesian persuasion problems corresponding to Pareto efficient profiles can be solved by concavifying a real-valued function, $\hat{v}_{\lambda}$. Theorem 2 thus provides us with a tractable way of characterizing the Pareto efficient information structures and thus recover the Pareto efficient profiles (see, e.g., the analysis in Section 5). The ability to recover the Pareto frontier may be useful when maximizing non-utilitarian social welfare functions, such as those that correspond to a fairness-aware planner (e.g., Rawls' criterion or Epstein and Segal's quadratic social welfare function), or those that obtain from efficient bargaining (e.g., Nash bargaining). Even if such an objective would select a Pareto efficient profile, the supporting direction $\lambda$ may only be identifiable after characterizing the solution, so that knowledge of the whole Pareto frontier may be important.

[^8]

> [Figure omitted from this text-only corpus; refer to the source manuscript.]
Figure (a) The convex hull of the graph of $\hat{w}$

> [Figure omitted from this text-only corpus; refer to the source manuscript.]
Figure (b) The Bayes welfare set W

Figure 3: Constructing the Bayes welfare set in Example 2. The diagonal dashed line is the $45^{\circ}$ degree line. Blue dashed lines denote profiles on the boundary of the Bayes welfare set, which are not in the Bayes welfare set.

Theorem 2 has yet another practical implication: when a (Pareto efficient) w is an extreme point of W , an information structure exists that employs at most $N$ signals and generates w. We record this result below:

Corollary 2 (Extreme points of W). Let w be an extreme point of W . Then, an information structure $\Pi$ with at most $N$ signals exists such that $\mathrm{w}_{i}=w_{\Pi}\left(\theta_{i}\right)$ for all $i \in\{1, \ldots, N\}$.

In Example 1, all points in $\mathrm{W}_{\mathrm{P}}$ are extreme (Figure 2b). Corollary 2 then implies information structures that induce at most two posteriors are enough to characterize the Pareto frontier in Example 1.

We use Example 2 to illustrate properties of the Bayes welfare profiles on the Pareto frontier and why the bound in Corollary 2 applies only to extreme points on the Pareto frontier of W. In doing so, we highlight the difference between the value of the program in Equation 6 and the Bayes welfare profiles consistent with that value:

Example 2 (continued). Figure 3 illustrates the convex hull of the graph of $\hat{w}$ (Figure 3a) and the Bayes welfare set (Figure 3b) for the online marketplace example. We note the following features of the Pareto frontier of W in this example, depicted in black in Figure 3b. First, in contrast to Example 1, the full-disclosure profile, $\mathrm{w}^{F D}$, is Pareto efficient, whereas the no-disclosure profile, $\mathrm{w}^{N D}$, is not. Second, like in Example 1, there is a continuum of Bayes welfare profiles that are fair in the sense of equalizing welfare
across different consumer types. Whereas the points in the flat segment of the Pareto frontier to the right of $\mathrm{w}=(0.6,0.6)$ maximize Rawls' criterion, only $\mathrm{w}=(0.6,0.6)$ is both Pareto efficient and fair.

Third, contrary to Example 1, not every point on the boundary of W is an extreme point. Consider, for instance, the points on the boundary in the direction $\lambda=\left(\frac{2}{3}, \frac{1}{3}\right)$ in Figure 3b, which would be consistent with the platform being interested in promoting the participation of $\theta_{H}$-consumers. It is possible to show that the points in the interior of this line are generated by information structures with three signals. ${ }^{11}$ Note all the points in that line lead to the same value in the Bayesian persuasion problem with indirect utility function $\hat{v}_{\lambda}$, namely, cav $\hat{v}_{\lambda}\left(\mu_{0}\right)$. However, they correspond to different welfare profiles with different implications regarding how $\theta_{H^{-}}$and $\theta_{L}$-consumers share the payoff cav $\hat{v}_{\lambda}\left(\mu_{0}\right)$. This highlights a benefit of the perspective we develop in this paper: insofar as one cares about the cross-sectional implications of different information structures, studying only the average welfare induced by a given information structure may not be sufficient.

Because the seller breaks ties in favor of the consumer when setting prices, the welfare function in Example 2 is upper-semicontinuous, but not continuous. Consequently, the Bayes welfare set is not closed in this example: The profiles corresponding to the blue dashed lines in Figure 3b can be induced by an information structure only when the seller breaks ties against the consumer. Note, however, that upper-semicontinuity of the welfare function guarantees that the supporting Bayesian persuasion problem attains a solution in all directions $\lambda \in \Delta(\Theta)$. Whereas in general upper-semicontinuity does not guarantee the Pareto frontier is closed, it ensures that all weakly monotone social welfare functions attain a maximum in the Bayes welfare set (see Proposition A.2). ◇

We close this section by noting that when W is closed, Theorem 2 provides an alternative characterization of the Bayes welfare set. Whereas points on the Pareto frontier are natural candidates for the choice of a social planner, points outside the Pareto frontier may be relevant if, for instance, the social planner faces constraints in their choice of a Bayes welfare profile. Because such constraints may rule out Pareto efficient and even boundary profiles, knowledge of the whole Bayes welfare set may be important. For this reason, we see the characterizations in both theorems as complementary.

[^9]
### 4.1 When is (no) information disclosure efficient?

A natural question to ask is when providing some information is the efficient thing to do; in other words, when does an information structure exist that all types strictly prefer to no disclosure? When such an information structure exists, we say all types benefit from disclosure. Examples 1 and 2 offer an interesting contrast in this respect: because the no-disclosure profile $\mathrm{w}^{N D}$ is not in the Pareto frontier in Example 2, all consumer types benefit from disclosure. Instead, the no-disclosure profile $\mathrm{w}^{N D}$ is in the Pareto frontier in Example 1. Indeed, in Example 1 information disclosure necessarily hurts at least one of the types (in this case, $\theta_{A}$ ).

Corollary 3 provides a necessary and sufficient condition for all types to benefit from disclosure. To introduce this condition, recalling a definition from Kamenica and Gentzkow (2011) is useful: a sender with indirect utility function $\hat{v}$ benefits from persuasion if the concavification of $\hat{v}$ at the prior exceeds the value of $\hat{v}$ at the prior; that is, $\hat{v}\left(\mu_{0}\right)<\operatorname{cav} \hat{v}\left(\mu_{0}\right)$. We have the following:

Corollary 3 (Efficiency of disclosure). All types benefit from disclosure if and only if, for all directions $\lambda \in \Delta(\Theta)$, a sender with indirect utility $\hat{v}_{\lambda}$ benefits from persuasion.

In other words, Corollary 3 states that the no-disclosure profile $\mathrm{w}^{N D}$ is in the Pareto frontier of W if and only if a direction $\lambda^{N D} \in \Delta(\Theta)$ exists such that a sender with indirect utility $\hat{v}_{\lambda^{N D}}$ does not benefit from persuasion. Alternatively, a social planner with Pareto weights $\lambda^{N D}$ would find no disclosure to be an optimal information structure. Recall Example 1: in that case, when $\lambda=\mu_{0}$, the indirect utility function $\hat{v}_{\mu_{0}}$ equals the strictly concave function $-(\mu-1 / 2)^{2}$, so that no disclosure is optimal. Corollary A. 1 in Appendix A. 2 shows that this observation reflects a more general result: if for all $\theta \in \Theta$ the welfare function takes the form $a(\theta) w(\mu)+b(\theta)$, where $a(\theta)>0$ and $w$ is concave, $\mathrm{w}^{N D}$ is in the Pareto frontier of W.

Whereas Corollary 3 provides a condition in terms of the indirect utility function $\hat{v}_{\lambda}$, Observation 1 provides conditions on the truth-adjusted welfare function $\hat{w}$ under which all types do or do not benefit from disclosure: ${ }^{12}$

Observation 1 (Disclosure benefits). The following hold:

1. If, for all $\theta \in \Theta, \hat{w}(\cdot, \theta)$ is strictly convex in a neighborhood of the prior, all types benefit from disclosure.
2. Instead, if a type $\theta$ exists such that $\hat{w}(\cdot, \theta)$ is concave in $\mu$, no disclosure is Pareto efficient.

As in Kamenica and Gentzkow (2011), concavity and convexity properties of a pay-

[^10]off function determine whether information disclosure is beneficial. In contrast to Kamenica and Gentzkow (2011), the concavity and convexity properties of the truthadjusted welfare function are what determine whether any given type can benefit from disclosure. This result can be clearly seen in Example 1, where the truth-adjusted welfare of $\theta_{A}$ is concave around the prior and that of $\theta_{B}$ is strictly convex around the prior, even though each type's welfare function is strictly concave. ${ }^{13}$

To understand the role of the convexity (concavity) properties of the truth-adjusted welfare function in determining the benefits from disclosure, note that information disclosure affects the welfare of a given type $\theta$ through two channels: directly through its impact on the welfare function as in Kamenica and Gentzkow (2011) and indirectly through the truth-drift adjustment. The second channel is most easily seen in the case of binary types. In that case, the property in Equation TD ensures that the posterior belief drifts along a straight line toward the true type $\theta$. When the welfare function of individuals of type $\theta$ is convex and increasing in $\mu(\theta)$, both effects are positive, ensuring that individuals of type $\theta$ benefit from disclosure. Observation 2 below summarizes this discussion:

Observation 2 (Binary types). Let $N=2$. If for all $\theta \in \Theta, w(\mu, \theta)$ is strictly convex in $\mu$ and increasing in $\mu(\theta)$, both types benefit from disclosure; moreover, full disclosure is uniquely Pareto efficient. Instead, if for some $\theta \in \Theta, w(\mu, \theta)$ is concave in $\mu$ and decreasing in $\mu(\theta)$, no disclosure is Pareto efficient.

## 5 Expected Reputation

We now specialize our results to the case in which the welfare function equals the expectation of some one-dimensional variable of interest, such as an individual's productivity, quality, or trade value. This is a standard way to model reputation, image, or career concerns in economics (see, e.g., Holmström, 1999, Bénabou and Tirole, 2006). It also captures the class of prediction policy problems in Rambachan et al. (2020), in which a decision maker-our outside observer-selects a treatment (e.g., hiring, bail, loan-approval) on the basis of the prediction of an outcome of interest (e.g., productivity, recidivism, creditworthiness).

[^11]Formally, we assume a reputation vector $\rho \in \mathbb{R}^{N}$ exists such that for all $\theta_{i}, \theta_{j} \in \Theta$,

$$
w\left(\mu, \theta_{i}\right)=w\left(\mu, \theta_{j}\right)=\mathbb{E}_{\mu}[\rho(\theta)]=\sum_{k=1}^{N} \mu\left(\theta_{k}\right) \rho\left(\theta_{k}\right)=\mu^{T} \rho .
$$

We refer to $w$ as the individual's reputation. Thus, a Bayes welfare profile is a profile of expected reputations. Without loss of generality, we label $\rho$ in increasing order, that is, $\rho_{1} \leq \cdots \leq \rho_{N}$, so that types are labeled in increasing order of their values under $\rho$.

The analysis in this section allows us to focus on the redistributive role of information. Indeed, all information structures lead to the same ex ante welfare. That is, for any information structure $\Pi$ and the corresponding Bayes welfare profile w, the ex ante welfare is:

$$
\mu_{0}^{T} \mathrm{w}=\sum_{\theta \in \Theta} \mu_{0}(\theta) \mathbb{E}_{\langle\Pi \mid \theta\rangle}\left[\tilde{\mu}^{T} \rho\right]=\mathbb{E}_{\langle\Pi\rangle}\left[\tilde{\mu}^{T} \rho\right]=\mu_{0}^{T} \rho .
$$

However, as the results in this section illustrate, different information structures lead to different welfare profiles, so that the chosen information structure determines how the different types in the population share the ex ante welfare.

When $w$ is as in Equation 8, we can provide an alternative characterization of the set W . From Section 3, it follows that $\mathrm{w} \in \mathrm{W}$ if and only if we can find a Bayes plausible distribution over posteriors, $\tau \in \Delta_{\mu_{0}}(\Delta(\Theta))$, with finite support, such that

$$
\mathrm{w}=\mathbb{E}_{\tau}[\hat{\mathrm{w}}(\tilde{\mu})]=\mathbb{E}_{\tau}\left[\frac{\tilde{\mu}}{\mu_{0}}\left(\tilde{\mu}^{T} \rho\right)\right]=\mathrm{D}_{0} \mathbb{E}_{\tau}\left[\tilde{\mu} \tilde{\mu}^{T}\right] \rho,
$$

where $\mathrm{D}_{0}$ denotes a diagonal matrix with $(i, i)$-th element equal to $1 / \mu_{0}\left(\theta_{i}\right)$.
Equation 10 shows a Bayes welfare profile can be represented as the product of three terms: the reputation vector $\rho$, the prior-normalizing matrix $\mathrm{D}_{0}$, and the matrix $\mathbb{E}_{\tau}\left[\mu \mu^{T}\right]$. Furthermore, the matrix $\mathbb{E}_{\tau}\left[\mu \mu^{T}\right]$ satisfies the following two properties. First, it is a completely positive matrix (Berman, 1988): an $N \times N$ matrix C is completely positive if it can be written as $\sum_{m=1}^{M} \mathrm{x}_{\mathrm{m}} \mathrm{x}_{\mathrm{m}}^{T}$ for some finite collection of non-negative vectors $\mathrm{x}_{\mathrm{m}} \in \mathbb{R}_{+}^{N}$. ${ }^{14}$ Second, the rows of the matrix $\mathbb{E}_{\tau}\left[\mu \mu^{T}\right]$ add up to the prior: $\mathbb{E}_{\tau}\left[\mu \mu^{T}\right] \mathrm{e}=\mathbb{E}_{\tau}\left[\mu\left(\mu^{T} \mathrm{e}\right)\right]=\mathbb{E}_{\tau}[\mu]=\mu_{0}$. Theorem 3 shows these two properties are

[^12]not only necessary but also sufficient and thus fully characterize the Bayes welfare set: ${ }^{15}$

Theorem 3. Given the reputation vector $\rho, \mathrm{w} \in \mathrm{W}$ if and only if a completely positive matrix $\mathrm{C} \in \mathbb{R}^{N \times N}$ exists such that $\mathrm{Ce}=\mu_{0}$ and

$$
\mathrm{w}=\mathrm{D}_{0} \mathrm{C} \rho .
$$

Putting together the properties in Theorem 3, we obtain that any Bayes welfare profile w is the product of the reputation vector $\rho$ and a matrix P , where $\mathrm{P} \equiv \mathrm{D}_{0} \mathrm{C}$ is the transition matrix of a reversible Markov chain with invariant distribution $\mu_{0}$. That is, (i) $\mu_{0}^{T} \mathrm{P}=\mu_{0}^{T}$, (ii) $\mathrm{Pe}=\mathrm{e}$, and (iii) P satisfies the detailed balance conditions: for all $i, j \in N, \mu_{0 i} \mathrm{P}_{i j}=\mu_{0 j} \mathrm{P}_{j i}$. The first property captures the pure redistribution of welfare highlighted in Equation 9. The second property implies any Bayes welfare profile can be viewed as a garbled version of the full information profile $\rho$. The third property delineates the limits of how payoffs can be redistributed by linking how much of $\rho\left(\theta_{i}\right)$ can be attributed to $\theta_{j}$, and vice versa. Indeed, because P is the transition matrix of a reversible Markov chain, we obtain that there is mean reversion in the redistribution of payoffs across types. To see this, note that if $\mathrm{w}=\mathrm{P} \rho \in \mathrm{W}$, then also $\mathrm{Pw} \in \mathrm{W} .{ }^{16}$ Because $\mu_{0}$ is the invariant distribution of P , we have that $\mathrm{P}^{k} \mathrm{w} \rightarrow_{k \rightarrow \infty}$ $\left(\mu_{0}^{T} \mathrm{w}\right) * \mathrm{e}=\left(\mu_{0}^{T} \rho\right) * \mathrm{e}=\mathrm{w}^{N D}$, where $\mathrm{w}^{N D}$ is the no-disclosure profile.

Remark 2 (Connections to the literature). Reversible Markov chains are prominent in the study of higher-order beliefs and expectations of higher-order beliefs (Samet, 1998; Cripps et al., 2008; Golub and Morris, 2017). Indeed, note that when the welfare function is linear and type independent, a Bayes welfare profile is a vector of second-order expectations: for any type $\theta$ and any $\mathrm{w} \in \mathrm{W}, w(\theta)$ is the expectation under some information structure of the random variable $\mu^{T} \rho$ of an individual that knows $\theta$. Whereas that literature takes the information structure as given and shows (sequences of) higher-order expectations can be obtained by iteratively applying the transition matrix of a reversible Markov chain, Theorem 3 identifies which transition matrices are consistent with some information structure and shows complete positivity is the key property they must satisfy.

Theorem 3 also relates to the literature on majorization (Hardy et al., 1952): if all types are equally likely, P is doubly stochastic and $\rho$ majorizes w . However, not any profile majorized by $\rho$ is a Bayes welfare profile, because not all doubly stochastic matrices are symmetric, and hence, some do not satisfy the detailed balance conditions.

[^13]We now show how Theorem 3 delivers a more general version of the truth-drifting property discussed in Section 3. This property has been obtained in different forms in the literature that studies the feasible evolution of beliefs (e.g., Francetich and Kreps, 2014, Hart and Rinott, 2020). Truth-drifting states that whereas an information structure can occasionally "deceive" the outside observer about an individual's true type, it cannot systematically do so. This property underlies the limits of using information as a tool to distribute welfare in the population.

Formally, consider any event $X$ that is correlated with the types according to the conditional probability function $\beta \in[0,1]^{N}, \beta_{i} \equiv \operatorname{Pr}\left(X \mid \theta_{i}\right)$, so that the prior probability of the event is $\operatorname{Pr}(X)=\mu_{0}^{T} \beta .^{17}$ If all $\beta_{i} \in\{0,1\}$, the event effectively indicates a subset of types. More generally, the event may involve extraneous uncertainty, and the types may be only imperfectly informative about it. We show that if the event is true, the average posterior probability that the outside observer attaches to this event must be at least as large as the prior probability:

Claim 2 (Truth drifting). For any event $X$ and information structure $\Pi$,

$$
\mathbb{E}_{\Pi}[\operatorname{Pr}(X \mid s) \mid X] \geq \operatorname{Pr}(X) .
$$

Francetich and Kreps (2014) obtain this result, relying on the properties of KullbackLeibler divergence. ${ }^{18}$ Hart and Rinott (2020) obtain a version of Claim 2 in the special case of $X \subseteq \Theta$, relying on the monotone-likelihood ratio property. Instead, our proof of Claim 2, presented in Appendix A.3, builds on the property that the underlying matrix C is completely positive and thus necessarily positive semi-definite.

Boundary information structures: Recall that Theorem 2 characterizes the boundary of the Bayes welfare set by means of supporting Bayesian persuasion problems. In the reputation model, Equation 9 implies the Bayes welfare set lies within a hyperplane with orthogonal vector $\mu_{0}$. Thus, instead of studying the boundary of the Bayes welfare set, we focus on its relative boundary, which consists of all Bayes welfare profiles not in the relative interior of the Bayes welfare set. ${ }^{19}$ In a slight abuse of terminology, we refer to the profiles on the relative boundary of W and the information structures that induce them as boundary profiles and information structures, respectively.

[^14]As we show next, the supporting Bayesian persuasion problems in the reputation model take a well-known structure. Indeed, fix a direction $\lambda \in \mathbb{R}^{N} \backslash\{0\}$ not collinear with $\mu_{0}$ and consider the induced supporting Bayesian persuasion problem:

$$
\begin{aligned}
\max _{\tau \in \Delta_{\mu_{0}}(\Delta(\Theta))} \mathbb{E}_{\tau}\left[\lambda^{T} \hat{\mathrm{w}}(\mu)\right] & =\max _{\tau \in \Delta_{\mu_{0}}(\Delta(\Theta))} \mathbb{E}_{\tau}\left[\left(\frac{\lambda^{T}}{\mu_{0}} \mu\right)\left(\rho^{T} \mu\right)\right] \\
& =\max _{\tau \in \Delta_{\mu_{0}}(\Delta(\Theta))} \mathbb{E}_{\tau}\left[\mathbb{E}_{\mu}\left[\frac{\lambda(\theta)}{\mu_{0}(\theta)}\right] \mathbb{E}_{\mu}[\rho(\theta)]\right]
\end{aligned}
$$

where the first equality uses the form of $w$ and the definition of $\hat{w}$. Equation $\mathrm{RS}_{\lambda}$ shows that if an information structure $\Pi$ delivers a profile w on the relative boundary of W, the information structure solves an instance of the information design problem in Rayo and Segal (2010). To be precise, Rayo and Segal (2010) consider the following problem. A sender owns a prospect, and his objective is that the receiver accepts it. When the sender's type is $\theta$ and the receiver accepts the prospect, the sender and the receiver obtain a payoff $\gamma(\theta) \equiv \lambda(\theta) / \mu_{0}(\theta)$ and $\rho(\theta) \in[0,1]$, respectively. Instead, if the receiver rejects the prospect, the sender obtains a payoff of 0, whereas the receiver obtains a payoff $u$ distributed uniformly over [0,1] independently of $\theta$. The sender chooses an information structure, $\Pi$, without observing the realization of $u$. Thus, when $\Pi$ induces a belief $\mu$, the sender expects the receiver to accept the project with probability, $\rho^{T} \mu$. It follows that the last term in Equation $\mathrm{RS}_{\lambda}$ represents the sender's expected payoff when $\tau$ is the distribution over posteriors induced by information structure $\Pi$.

Proposition 2. (Boundary profiles) A welfare profile w is on the relative boundary of W if and only if a direction $\lambda \in \mathbb{R}^{N} \backslash\{0\}$ not collinear with $\mu_{0}$ exists such that w is induced by an information structure that solves the program $R S_{\lambda}$.

Proposition 2 allows us to rely on the approach of Rayo and Segal (2010) to characterize the shape of the information structures that achieve the boundary Bayes welfare profiles. This approach relies on a graphical representation of an information structure, in which the prospect values $\{(\gamma(\theta), \rho(\theta)): \theta \in \Theta\}$ are the nodes (see Remark A. 1 in the appendix). The results in Rayo and Segal (2010) have immediate implications for the information structures that induce the boundary profiles of W:

Corollary 4. In the reputation model, the following hold:

1. An information structure that induces a boundary Bayes welfare profile in the direction $\lambda$ does not pool types $\theta_{i}$ and $\theta_{j}$ whenever their ranking under the vector $\lambda / \mu_{0}$ and the reputation vector $\rho$ is the same;
2. The full- and no-disclosure profiles are on the relative boundary of W.

The first part of Corollary 4 highlights a natural feature of optimal information provision in the reputation model in terms of the alignment of preferences of the social planner, captured by $\lambda / \mu_{0}$, and of the outside observer, captured by $\rho$ : the planner should not pool any two types as long as the planner and the outside observer are in agreement about the types' relative ranking. The second part of Corollary 4 shows that the full- and no-disclosure profiles are on the boundary of the Bayes welfare set. Whereas this property holds in Examples 1 and 2, in which the welfare function is not linear, this property is not a general one, as we illustrate in Example A. 4 in the appendix.

Individual reputation bounds: We can further build on the graphical approach of Rayo and Segal (2010) to characterize the information structures that deliver maximal (or minimal) welfare to any given type. This exercise provides a rough way to bound the Bayes welfare set W and also suggests how information may be employed to boost (or dilute) the reputation of particular types in the population. In the context of our credit agency example, Example 3, the information structure that maximizes the welfare of individuals of type $\theta_{i}$ maximizes the probability that individuals of type $\theta_{i}$ obtain credit.

Formally, given a target type $\theta_{i}$, we want to solve the following problem:

$$
\max _{\mathrm{w} \in \mathrm{~W}} \mathrm{w}_{i} .
$$

(i-MAX)
Proposition 3 below shows a particular class of information structures solves the problem $i$-MAX.

Definition 3 (Noisy priority). A $\theta_{i}$-noisy-priority policy with threshold $k$ is an information structure $(\pi, S)$ such that $S=\Theta$, and the likelihood function $\pi$ satisfies:

1. If $j \neq i, \pi\left(s=\theta_{j} \mid \theta_{j}\right)=1$,
2. If $j<k, \pi\left(s=\theta_{j} \mid \theta_{i}\right)=0$, and
3. If $j \geq k, \pi\left(s=\theta_{j} \mid \theta_{i}\right)>0$.

In other words, a $\theta_{i}$-noisy-priority policy pairwise pools the target type $\theta_{i}$ with all types with indices above some threshold and separates all other types. A noisy-priority policy has an implementation akin to the priority mechanisms in the matching literature (Celebi and Flynn, 2022), and hence its name. A noisy-priority policy with threshold $k \geq i$ can be implemented by first assigning a perfectly revealing score to each type equal to their index, and then prioritizing the target type $\theta_{i}$ by increasing this type's score by a random number.

Proposition 3. The Bayes welfare profile that solves $i-M A X$ is induced by a $\theta_{i}$-noisypriority policy with threshold $k \geq i$.

The proof is in Appendix A.3. One part of Proposition 3 is straightforward: if one wishes to increase the reputation of $\theta_{i}$, then $\theta_{i}$ should be separated from all types with lower indices. What might be less obvious is that whenever $\theta_{i}$ is pooled with some other type, $\theta_{i}$ should be pooled with it pairwise. In a sense, pooling several types together redistributes the reputation from higher-quality types to lower-quality types. Pairwise pooling then allows the target type to obtain maximal reputation gains from any other type without sharing the gains with others. Finally, pairwise pooling with many types ensures no signal is overly "muddled," which in turn ensures an overall high reputation for $\theta_{i}$.

By simply reversing signs, Proposition 3 can be used to characterize the information structure that minimizes the expected reputation of individuals of type $\theta_{i}$ : this information structure should pairwise pool the target type with types whose indices are below some threshold. Such adversarial pairwise pooling inflicts maximal reputation losses and can be viewed as a noisy-degrading policy.

We conclude this section by illustrating the Bayes welfare set in the context of Example 3 and highlight the importance of population heterogeneity as captured by the number of types.

Example 3 (continued). Figure 4 illustrates the results of this section in the context of Example 3. Recall that in this case $\rho(\theta)$ denotes the probability that an individual of type $\theta$ repays the loan, and hence, $w(\mu, \theta)$ is the expected repayment probability under belief $\mu$. Like in Rayo and Segal (2010), we assume the credit agency has a uniform outside option. Assuming individuals wish to maximize the probability the lending agency approves the loan justifies that $\rho^{T} \mu$ corresponds to their welfare.

Figure 4a depicts the individually feasible welfare profiles (dashed square) and the Bayes welfare set (blue line) in the case of $N=2$. Proposition 1 implies any payoff between $\rho_{1}=0$ and $\mu_{0}^{T} \rho$ is feasible for $\theta_{1}$, whereas any payoff between $\mu_{0}^{T} \rho$ and $\rho_{2}=1$ is feasible for $\theta_{2}$. As Figure 4a illustrates, the Cartesian product $\left[\rho_{1}, \mu_{0}^{T} \rho\right] \times$ $\left[\mu_{0}^{T} \rho, \rho_{2}\right]$ is a rather lax bound in this example. In particular, the Cartesian product $\left[\rho_{1}, \mu_{0}^{T} \rho\right] \times\left[\mu_{0}^{T} \rho, \rho_{2}\right]$ ignores that all Bayes welfare profiles satisfy $\mu_{0}^{T} \mathrm{w}=\mu_{0}^{T} \rho=0.5$ (Equation 9). In the case of $N=2$, adding this restriction is enough to pin down the Bayes welfare set. The reason is that by Theorem 3, all Bayes welfare profiles can be obtained by "garbling" the full-disclosure Bayes welfare profile, $\rho$, and in the case of binary types, this garbling turns out to span a linear segment. Finally, note the structure of the Bayes welfare set implies a social planner with Pareto weights $\lambda$ finds it optimal to provide no or full information, depending on whether the planner weighs the welfare of $\theta_{1}$-individuals more than that of $\theta_{2}$-individuals (that is, $\lambda_{1} \lessgtr \lambda_{2}$ ).

Figure 4b depicts the Bayes welfare set W in the case of $N=3$. In contrast to the binary-type case, the boundary of the Bayes welfare set is non-linear and features a

> [Figure omitted from this text-only corpus; refer to the source manuscript.]
Figure 4: Expected reputation in Example 3. The blue color marks Bayes welfare set W . The dashed segments outline the Cartesian product of individual welfare sets.

continuum of extreme points. This example illustrates that the constraints imposed by Bayes plausibility are richer than the simple garbling constraint that characterizes the set when $N=2$. As we show in Appendix A.3, four classes of information structures span the boundary of the Bayes welfare set. The first two are the $\theta_{1}$-noisy-priority and $\theta_{3}$-noisy-degrading policies, which span the nonlinear segment of the boundary. The other two span the linear segments in Figure 4b and similar to the $\theta_{3}$-noisy-priority and the $\theta_{1}$-noisy-degrading policies, either maximize the loan-approval probability of $\theta_{3}$-individuals, by separating them from $\theta_{1}$ - and $\theta_{2}$-individuals, or minimize the loanapproval probability of $\theta_{1}$-individuals, by separating them from $\theta_{2}$ - and $\theta_{3}$-individuals. Unlike the $\theta_{3}$-noisy-priority and the $\theta_{1}$-noisy-degrading policies, the two classes of information structures that span the linear segments may pool together the types below $\theta_{3}$ or above $\theta_{1}$, respectively. Notably, for almost all boundary points, $\theta_{2}$-individuals are pooled with individuals of some other type. Intuitively, $\theta_{2}$-individuals exert pooling externalities on individuals of types $\theta_{1}$ and $\theta_{3}$. ${ }^{20}$ For instance, $\theta_{2}$-individuals enable boosting the loan-approval probability of $\theta_{1}$-individuals in the case of the $\theta_{1}$-noisy-priority policy, but exert a negative externality on $\theta_{3}$-individuals in the $\theta_{3}$-noisy-degrading policy. $\square$

## 6 Conclusion

[^15]We provide a framework to study the potentially disparate impact of information policies in a population of heterogeneous individuals. Because information policies increasingly shape society's choices in high-stakes domains, the Bayes welfare set describes the limits of what society can achieve under such policies and the welfare trade-offs implied by the choice between different information policies. In the spirit of mechanism design and information design, our characterization of the Bayes welfare set provides a unifying tool to evaluate the welfare implications of different information policies across a wide array of objective functions.

We see several avenues worth exploring and left for future work. First, our model assumes any information can be provided about an individual's payoff-relevant type. However, this assumption does not necessarily hold in applications of interest in which an individual's type may include protected characteristics. In the online appendix, we extend our framework to accommodate limits on how much information can be disclosed about the individuals in the population. The analysis there, however, does not consider that these limits may be designed when a potentially malicious third party selects the information structure. Liang et al. (2022) consider this case in the context of decision-making algorithms, and we expect their insights to extend to the case of recommendations algorithms like the ones we consider.

Second, because the welfare function depends only on the first-order beliefs about an individual's type, our model only accounts for strategic interactions that follow the realization of a public signal (see, e.g., Laclau and Renou, 2017). Extending the analysis to account for general strategic interactions is worth exploring. Galperti et al. (2023), which studies the Bayes welfare profile that gives rise to the sender's maximum average payoff across all Bayes correlated equilibria, is a step in this direction.

Third, motivated by recent policies, Tirole (2021) studies the use of information in the form of a social score to incentivize good behavior in the population. Whereas Tirole (2021) studies this question in the context of a parametric family of information structures, our initial explorations show the Bayes welfare set allows us to extend his results by allowing any information structure. More generally, information has been suggested as a substitute for monetary incentives, and the Bayes welfare set describes what can be achieved with information alone.

Finally, whereas we characterize individuals' welfare as a function of their type, thinking of applications in which we care instead about the welfare of groups is natural. For instance, an individual's type could encompass their gender and their ability, and the social planner is concerned with the welfare different genders may obtain. This extension can inform the study of statistical discrimination, where recent work shows Bayesian persuasion tools can shed new light to this problem (Chambers and Echenique, 2021; Escudé et al., 2022; Deb and Renou, 2022). Whereas much of the existing litera-
ture focuses on statistical properties of discrimination, our work highlights the important aspect of economic welfare, offering a complementary perspective that enriches the ongoing dialogue between statistics and economics.

## References

Akbarpour, M., E. Budish, P. Dworczak, and S. D. Kominers (2023): "An Economic Framework for Vaccine Prioritization," The Quarterly Journal of Economics, qjad022.

Akbarpour, M., P. Dworczak, and S. D. Kominers (Forthcoming): "Redistributive AIlocation Mechanisms," Journal of Political Economy.

Aliprantis, C. D. and K. C. Border (2013): Infinite Dimensional Analysis: A Hitchhiker's Guide, Springer-Verlag Berlin and Heidelberg GmbH \& Company KG.

Alonso, R. and O. Câmara (2016): "Bayesian Persuasion with Heterogeneous Priors," Journal of Economic Theory, 165, 672-706.

Angwin, J., J. Larson, S. Mattu, and L. Kirchner (2016): "Machine Bias: There's Software Used across the Country to Predict Future Criminals. And It's Biased Against Blacks." ProPublica, 23, 77-91.

Arieli, I., Y. Babichenko, and F. Sandomirskiy (2022): "Persuasion as Transportation," in Proceedings of the 23rd ACM Conference on Economics and Computation, 468.

Aumann, R. (1987): "Correlated Equilibrium as an Expression of Bayesian Rationality," Econometrica, 55, 1-18.

Aumann, R. and M. Maschler (1995): Repeated Games with Incomplete Information, MIT Press.

Bénabou, R. and J. Tirole (2006): "Incentives and Prosocial Behavior," American Economic Review, 96, 1652-1678.

Bergemann, D., B. Brooks, and S. Morris (2015): "The Limits of Price Discrimination," American Economic Review, 105, 921-57.

Berman, A. (1988): "Complete Positivity," Linear Algebra and its Applications, 107, 57 63.

Berman, A. and N. Shaked-Monderer (2003): Completely Positive Matrices, World Scientific.

Blackwell, D. (1953): "Equivalent Comparisons of Experiments," The Annals of Mathematical Statistics, 265-272.

Celebi, O. and J. Flynn (2022): "Adaptive Priority Mechanisms," Working Paper.
Chambers, C. P. and F. Echenique (2021): "A Characterisation of 'Phelpsian' Statistical Discrimination," The Economic Journal, 131, 2018-2032.

Cripps, M. W., J. C. Ely, G. J. Mailath, and L. Samuelson (2008): "Common Learning," Econometrica, 76, 909-933.

Deb, R. and L. Renou (2022): "Which Wage Distributions are Consistent with Statistical Discrimination?" Working paper.

Doval, L. and V. Skreta (2022): "Mechanism Design with Limited Commitment," Econometrica, 90, 1463-1500.

Doval, L. and A. Smolin (2021): "Information Payoffs: An Interim Perspective," arXiv preprint arXiv:2109.03061.

Dworczak, P., S. D. Kominers, and M. Akbarpour (2021): "Redistribution through Markets," Econometrica, 89, 1665-1698.

Ely, J., A. Frankel, and E. Kamenica (2015): "Suspense and Surprise," Journal of Political Economy, 123, 215-260.

Epstein, L. G. and U. Segal (1992): "Quadratic social welfare functions," Journal of Political Economy, 100, 691-712.

Escudé, M., P. Onuchic, L. Sinander, and Q. Valenzuela-Stookey (2022): "Statistical Discrimination and Statistical Informativeness," arXiv preprint arXiv:2205.07128.

Francetich, A. and D. Kreps (2014): "Bayesian Inference Does Not Lead You Astray... On Average," Economics Letters, 125, 444-446.

Fréchette, G. R., A. Lizzeri, and J. Perego (2022): "Rules and Commitment in Communication: An Experimental Analysis," Econometrica, 90, 2283-2318.

Galperti, S., A. Levkun, and J. Perego (2023): "The Value of Data Records," Review of Economic Studies, rdad044.

Gentzkow, M. and E. Kamenica (2017): "Bayesian Persuasion with Multiple Senders and Rich Signal Spaces," Games and Economic Behavior, 104, 411-429.

Golub, B. and S. Morris (2017): "Higher-Order Expectations," Available at SSRN 2979089.

Green, J. and N. Stokey (1978): "Two Representations of Information Structures and their Comparisons," IMSSS, Stanford University.

Hardy, G. H., J. E. Littlewood, G. Pólya, G. Pólya, D. Littlewood, et al. (1952): Inequalities, Cambridge University Press.

Hart, S. and Y. Rinott (2020): "Posterior Probabilities: Dominance and Optimism," Economics Letters, 194, 109352.

Hiriart-Urruty, J.-B. and C. Lemaréchal (2004): Fundamentals of Convex Analysis, Springer Science \& Business Media.

Holmström, B. (1999): "Managerial Incentive Problems - A Dynamic Perspective," Review of Economic Studies.

Jagtiani, J. and C. Lemieux (2019): "The Roles of Alternative Data and Machine Learning in Fintech Lending: Evidence from the LendingClub Consumer Platform," Financial Management, 48, 1009-1029.

Kamenica, E. and M. Gentzkow (2011): "Bayesian Persuasion," American Economic Review, 101, 2590-2615.

Kartik, N., F. X. Lee, and W. Suen (2021): "Information Validates the Prior: A Theorem on Bayesian Updating and Applications," American Economic Review: Insights, 3, 165-182.

Kleinberg, J., J. Ludwig, S. Mullainathan, and A. Rambachan (2018): "Algorithmic Fairness," in AEA Papers and Proceedings, vol. 108, 22-27.

Koessler, F. and V. Skreta (Forthcoming): "Informed Information Design," Journal of Political Economy.

Kučak, D., V. Juričić, and G. Đambić (2018): "Machine Learning in Education- a Survey of Current Research Trends," Annals of DAAAM \& Proceedings, 29.

Laclau, M. and L. Renou (2017): "Public Persuasion," Working Paper.
Levy, G., I. Moreno de Barreda, and R. Razin (2021): "Feasible Joint Distributions of Posteriors: A Graphical Approach," Working Paper.

Li, D., L. R. Raymond, and P. Bergman (2020): "Hiring as Exploration," National Bureau of Economic Research.

Liang, A., J. Lu, and X. Mu (2022): "Algorithmic Design: Fairness versus Accuracy," in Proceedings of the 23rd ACM Conference on Economics and Computation, 58-59.

Lipnowski, E. and L. Mathevet (2018): "Disclosure to a Psychological Audience," American Economic Journal: Microeconomics, 10, 67-93.

Lipnowski, E. and D. Ravid (2020): "Cheap Talk with Transparent Motives," Econometrica, 88, 1631-1660.

Mukherjee, A. (2008): "Sustaining Implicit Contracts When Agents Have Career Concerns: the Role of Information Disclosure," The RAND Journal of Economics, 39, 469-490.

Mullainathan, S. (2018): "Algorithmic Fairness and the Social Welfare Function," in Proceedings of the 2018 ACM Conference on Economics and Computation, 1-1.

Obermeyer, Z., B. Powers, C. Vogeli, and S. Mullainathan (2019): "Dissecting Racial Bias in an Algorithm Used to Manage the Health of Populations," Science, 366, 447-453.

Ostrovsky, M. and M. Schwarz (2010): "Information Disclosure and Unraveling in Matching Markets," American Economic Journal: Microeconomics, 2, 34-63.

Perez-Richet, E. (2014): "Interim Bayesian Persuasion: First Steps," American Economic Review, 104, 469-74.

Quigley, D. and A. Walther (2019): "Contradiction-Proof Information Design," Working Paper.

Raghavan, M., S. Barocas, J. Kleinberg, and K. Levy (2020): "Mitigating Bias in Algorithmic Hiring: Evaluating Claims and Practices," in Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, 469-481.

Rambachan, A., J. Kleinberg, J. Ludwig, and S. Mullainathan (2020): "An Economic Perspective on Algorithmic Fairness," in AEA Papers and Proceedings, vol. 110, 91-95.

Rayo, L. and I. Segal (2010): "Optimal Information Disclosure," Journal of Political Economy, 118, 949-987.

Rosar, F. (2017): "Test Design Under Voluntary Participation," Games and Economic Behavior, 104, 632-655.

Saeedi, M. and A. Shourideh (2020): "Optimal Rating Design," arXiv preprint arXiv:2008.09529.

Salamanca, A. (2021): "The Value of Mediated Communication," Journal of Economic Theory, 192, 105191.

Samet, D. (1998): "Iterated Expectations and Common Priors," Games and Economic Behavior, 24, 131-141.

Sayin, M. O. and T. Başar (2021): "Bayesian Persuasion with State-Dependent Quadratic Cost Measures," IEEE Transactions on Automatic Control, 67, 1241-1252.

Tirole, J. (2021): "Digital Dystopia," American Economic Review, 111, 2007-48.

## A Omitted results and proofs from the main text

## A. 1 Omitted proofs from Section 3

For completeness, we include a proof of Claim 1, which follows from the analysis in Alonso and Câmara (2016, pp. 683-684), in present notation:

Proof of Claim 1. For an information structure $\Pi$, let $\operatorname{supp}\langle\Pi\rangle$ denote the support of the distribution over posterior beliefs induced by $\Pi$ and let $\operatorname{Pr}_{\Pi}(s)$ denote the unconditional probability of signal $s$ under the prior distribution $\mu_{0}$, i.e., $\operatorname{Pr}_{\Pi}(s)=$ $\sum_{\theta \in \Theta} \mu_{0}(\theta) \pi(s \mid \theta)$. Then, for a given type $\theta$, their welfare under information structure $\Pi$ can be written as follows:

$$
\begin{aligned}
w_{\Pi}(\theta)=\mathbb{E}_{\langle\Pi \mid \theta\rangle}[w(\tilde{\mu}, \theta)] & =\sum_{\mu \in \operatorname{supp}\langle\Pi\rangle} \sum_{s \in S: \mu_{s}=\mu} \pi(s \mid \theta) w(\mu, \theta) \\
& =\sum_{\mu \in \operatorname{supp}\langle\Pi\rangle} \sum_{s \in S: \mu_{s}=\mu} \operatorname{Pr}_{\Pi}(s) \frac{1}{\mu_{0}(\theta)} \frac{\mu_{0}(\theta) \pi(s \mid \theta)}{\operatorname{Pr}_{\Pi}(s)} w(\mu, \theta) \\
& =\sum_{\mu \in \operatorname{supp}\langle\Pi\rangle} \sum_{s \in S: \mu_{s}=\mu} \operatorname{Pr}_{\Pi}(s) \frac{\mu(\theta)}{\mu_{0}(\theta)} w(\mu, \theta) \\
& =\sum_{\mu \in \operatorname{supp}\langle\Pi\rangle} \sum_{s \in S: \mu_{s}=\mu} \operatorname{Pr}_{\Pi}(s) \hat{w}(\mu, \theta)=\mathbb{E}_{\langle\Pi\rangle}[\hat{w}(\tilde{\mu}, \theta)] .
\end{aligned}
$$ $\square$

Proof of Theorem 1. By definition, the point $\left(\mu_{0}, \mathrm{w}\right) \in \operatorname{co}(\operatorname{graph} \hat{\mathrm{w}})$ if and only if a Bayes plausible distribution over posteriors $\tau$ exists such that $\mathbb{E}_{\tau}[\hat{\mathrm{w}}(\mu)]=\mathrm{w}$. At the same time, the distribution over posteriors induced by an information structure, $\langle\Pi\rangle$, is Bayes plausible, i.e., $\mathbb{E}_{\langle\Pi\rangle}[\tilde{\mu}]=\mu_{0}$. The result follows from Claim 1. $\square$

On closedness Proposition A. 1 collects the results discussed in Section 3. To introduce Proposition A.1, we first introduce the analogue of the Bayes welfare set when the population's welfare depends on the outside observer's actions and together with the information structure, we can flexibly choose a selection from the outside observer's best-response correspondence. Formally, upon observing the realization from
the information structure, the outside observer takes an action in a finite set $A$ to maximize their expected payoff. We denote the outside observer's utility by $u: A \times \Theta \mapsto \mathbb{R}$ and the welfare function by $v: A \times \Theta \mapsto \mathbb{R}$. Given an information structure, $\Pi$, let $\operatorname{supp}\langle\Pi\rangle$ denote the support of the distribution over posteriors $\langle\Pi\rangle$. For each $\mu_{s} \in \operatorname{supp}\langle\Pi\rangle$, let

$$
\alpha\left(\mu_{s}\right) \in \Delta\left(\arg \max _{a \in A} \sum_{\theta \in \Theta} \mu_{s}(\theta) u(a, \theta)\right) \equiv \Delta\left(a^{*}(\mu)\right),
$$

denote the outside observer's (possibly mixed) best response. Recall that the Theorem of the Maximum implies that the correspondence $a^{*}(\mu)$ is upper-hemicontinuous.

Denote by $F$ set of tuples $(\Pi, \alpha)$, such that selection $\alpha$ satisfies Equation A.2. Each $(\Pi, \alpha) \in F$ defines a welfare profile, $w_{\Pi, \alpha}: \Theta \mapsto \mathbb{R}^{N}$, such that

$$
w_{\Pi, \alpha}(\theta)=\sum_{s \in S} \pi(s \mid \theta) \sum_{a \in A} \alpha\left(\mu_{s}\right)(a) v(a, \theta)=\mathbb{E}_{\langle\Pi \mid \theta\rangle}\left[\sum_{a \in A} \alpha(\tilde{\mu})(a) v(a, \theta)\right] .
$$

The analogue of the Bayes welfare set, which we denote by $\mathrm{W}_{\mathrm{BP}}$, is then

$$
\mathrm{W}_{\mathrm{BP}}=\left\{\mathrm{w} \in \mathbb{R}^{N}:(\exists(\Pi, \alpha) \in F) \text { s.t. } \mathrm{w}_{i}=w_{\Pi, \alpha}\left(\theta_{i}\right) \forall i \in\{1, \ldots, N\}\right\} .
$$

We have the following result:
Proposition A.1. The following hold:

(a) If for each $\theta \in \Theta$, the welfare function $w(\cdot, \theta)$ is continuous on $\Delta(\Theta)$, the Bayes welfare set W is closed.
(b) The set $\mathrm{W}_{\mathrm{BP}}$ is closed.

Proof of Proposition A.1. The proof of part (a) is immediate and hence omitted. Consider then part (b). Define the analogue of the truth-adjusted welfare function $\hat{v}(a, \mu, \theta)$ to be

$$
\hat{v}(a, \mu, \theta)=\frac{\mu(\theta)}{\mu_{0}(\theta)} v(a, \theta) .
$$

The same arguments as in Claim 1 imply that for any pair $(\Pi, \alpha) \in F$ and all $\theta \in \Theta$,

$$
w_{\Pi, \alpha}(\theta)=\mathbb{E}_{\langle\Pi\rangle}\left[\sum_{a \in A} \alpha(\tilde{\mu})(\theta) \hat{v}(a, \tilde{\mu}, \theta)\right] .
$$

Define the correspondence $\hat{V}: \Delta(\Theta) \rightrightarrows \mathbb{R}^{N}$ as follows

$$
\hat{V}(\mu)=\operatorname{co}\left\{\hat{\mathrm{v}}(a, \mu): a \in a^{*}(\mu)\right\} .
$$

Observe that $\hat{V}$ is a non-empty valued, convex-valued, and compact-valued correspondence. Furthermore, $\hat{V}$ is upper-hemicontinuous. The first part follows immediately from noting that $\hat{V}$ is the convex hull of finitely many vectors in $\mathbb{R}^{N}-a^{*}(\mu)$ is nonempty-and $\hat{v}$ is bounded. Upper-hemicontinuity of $\hat{V}$ follows from continuity of $\hat{\mathrm{v}}$ in $\mu$ and upper-hemicontinuity of $a^{*}(\mu)$. Consequently, the graph of $\hat{V}$,

$$
\operatorname{graph} \hat{V}=\left\{(\mu, \mathrm{v}) \in \Delta(\Theta) \times \mathbb{R}^{N}: \mathrm{v} \in \hat{V}(\mu)\right\},
$$

is closed.
We now show that similar to Theorem 1, the set $\mathrm{W}_{\mathrm{BP}}$ is the section at the prior of the convex hull of the graph of the correspondence $\hat{V}$, that is

$$
\mathrm{W}_{\mathrm{BP}}=\left\{\mathrm{w} \in \mathbb{R}^{N}:\left(\mu_{0}, \mathrm{w}\right) \in \operatorname{co}(\operatorname{graph} \hat{V})\right\} .
$$

Clearly, if $\mathrm{w} \in \mathrm{W}_{\mathrm{BP}}$, then $\left(\mu_{0}, \mathrm{w}\right) \in \operatorname{co}(\operatorname{graph} \hat{V})$. To see that the opposite holds, let $\left(\mu_{0}, \mathrm{w}\right) \in \operatorname{co}($ graph $\hat{V})$. Then, a finite collection $\left(\tau_{k}, \mu_{k}, \tilde{\mathrm{v}}_{k}\right)_{k=1}^{M}$ of non-negative weights, beliefs, and correspondence values exists such that $\sum_{k=1}^{M} \tau_{k}=1, \sum_{k=1}^{M} \tau_{k} \mu_{k}=$ $\mu_{0}, \tilde{\mathrm{v}}_{k} \in \hat{V}\left(\mu_{k}\right)$ for all $k \in\{1, \ldots, M\}$, and

$$
\mathrm{w}=\sum_{k=1}^{M} \tau_{k} \tilde{\mathrm{v}}_{k} .
$$

Because for each $k \in\{1, \ldots, M\}, \tilde{\mathrm{v}}_{k} \in \hat{V}\left(\mu_{k}\right)$, the definition of $\hat{V}$ implies a finite collection of non-negative weights and actions, $\left\{\alpha_{k, l}, a_{l}\right\}_{l=1}^{L_{k}}$, exists such that $\sum_{l=1}^{L_{k}} \alpha_{l, k}=$ 1 , for all $l \in\left\{1, \ldots, L_{k}\right\}, a_{l} \in a^{*}\left(\mu_{k}\right)$ and for all $i \in\{1, \ldots, N\}$,

$$
\tilde{\mathrm{v}}_{k, i}=\sum_{l=1}^{L_{k}} \alpha_{k, l} \hat{v}\left(a_{l}, \mu_{k}, \theta_{i}\right) .
$$

Define $\Pi$ to be the information structure with signals $S=\left\{\mu_{1}, \ldots, \mu_{M}\right\}$, and signal distribution $\pi\left(\mu_{k} \mid \theta\right)=\left(\mu_{k}(\theta) / \mu_{0}(\theta)\right) \tau_{k}$. Furthermore, define $\alpha$ so that for belief $\mu_{k}$, $\alpha\left(\mu_{k}\right) \in \Delta\left(a^{*}\left(\mu_{k}\right)\right)$ coincides with $\left\{\alpha_{k, l}\right\}_{l=1}^{L_{k}}$. By construction, $(\Pi, \alpha) \in F$. Equations A. 5 and A. 6 together imply that

$$
\mathrm{w}=\sum_{k=1}^{M} \tau_{k} \sum_{l=1}^{L_{k}} \alpha_{k, l \hat{\mathrm{v}}}\left(a_{l}, \mu_{k}\right)=\mathbb{E}_{\langle\Pi\rangle}\left[\sum_{a \in A} \alpha(\tilde{\mu})(a) \hat{\mathrm{v}}(a, \mu)\right] .
$$

Finally, Equation A. 4 allows us to conclude that the set $\mathrm{W}_{\text {BP }}$ is closed: it is the section at the prior of the convex hull of the graph of the correspondence $\hat{V}$, which is closed. $\square$

## A. 2 Omitted proofs from Section 4

Proof of Theorem 2. To complete the proof of the first direction of Theorem 2, we provide the steps to show that if $\mathrm{w} \in \mathrm{W}_{\mathrm{P}}$, a direction $\lambda \in \mathbb{R}_{+}^{N} \backslash\{0\}$ exists such that

$$
\lambda^{T} \mathrm{w}=\max \left\{\lambda^{T} \mathrm{w}^{\prime}: \mathrm{w}^{\prime} \in \mathrm{W}\right\},
$$

The opposite direction in Theorem 2 immediately follows.
Fix a Pareto efficient w and let $\Gamma=\left\{\mathrm{w}^{\prime} \in \mathbb{R}^{N}: \mathrm{w}^{\prime} \geq \mathrm{w}\right\}$. Clearly, $\Gamma$ is convex and int $\Gamma$ is non-empty. Because w is Pareto efficient, int $\Gamma \cap \mathrm{W}=\emptyset$. By Minkowski's separating hyperplane theorem, a direction $\lambda \in \mathbb{R}^{N} \backslash\{0\}$ exists such that for all $\mathrm{w}^{\prime \prime} \in \mathrm{W}$ and $\mathrm{w}^{\prime} \in \Gamma$,

$$
\lambda^{T} \mathrm{w}^{\prime \prime} \leq \lambda^{T} \mathrm{w}^{\prime} .
$$

Because $\mathrm{w} \in \Gamma$, we have that $\lambda^{T} \mathrm{w}^{\prime \prime} \leq \lambda^{T} \mathrm{w}$ for all $\mathrm{w}^{\prime \prime}$ in Bayes welfare set. Thus, $\lambda^{T} \mathrm{w}=\max \left\{\lambda^{T} \mathrm{w}^{\prime \prime}: \mathrm{w}^{\prime \prime} \in \mathrm{W}\right\}$ (cf. Equation 6). Similarly, because $\mathrm{w} \in \mathrm{W}$, then we have that $\lambda^{T} \mathrm{w} \leq \lambda^{T} \mathrm{w}^{\prime}$ for all $\mathrm{w}^{\prime} \in \Gamma$.

We now show that $\lambda \geq 0$. Let $\mathrm{t}_{i}$ denote the canonical vector that has a 1 in coordinate $\theta_{i}$ and 0 otherwise. Then, $\mathrm{w}+\mathrm{t}_{i} \in \Gamma$, so that

$$
\lambda^{T} \mathrm{w} \leq \lambda^{T}\left(\mathrm{w}+\mathrm{t}_{i}\right) \Rightarrow 0 \leq \lambda\left(\theta_{i}\right) .
$$

By definition, $\lambda \neq 0$ so that without loss of generality $\lambda \in \Delta(\Theta)$. $\square$

Proposition A.2. Suppose $w(\cdot, \theta)$ is upper-semicontinuous for all $\theta \in \Theta .{ }^{21}$ Suppose $\mathrm{w}^{*}$ is a limit point of $\mathrm{W}_{\mathrm{P}}$. Then, a point $\mathrm{w}^{* *} \in \mathrm{~W}_{\mathrm{P}}$ exists such that $\mathrm{w}^{* *} \geq \mathrm{w}^{*}$. Consequently, any weakly monotone social welfare function attains a solution in W .

Proof of Proposition A.2. Let $\left(\mathrm{w}_{n}\right)_{n \in \mathbb{N}} \subset \mathrm{~W}_{\mathrm{P}}$ be such that $\mathrm{w}_{n} \rightarrow \mathrm{w}^{*}$ as $n \rightarrow \infty$. For each $n \in \mathbb{N}$ a direction $\lambda_{n} \in \Delta(\Theta)$ exists such that

$$
\left(\forall \mathrm{w}^{\prime} \in \mathrm{W}\right) \lambda_{n}^{T} \mathrm{w}_{n} \geq \lambda_{n}^{T} \mathrm{w}^{\prime} .
$$

Since $\Delta(\Theta)$ is compact, then up to a subsequence $\lambda_{n} \rightarrow \lambda_{*}$. Linearity of $\lambda_{n}^{T} \mathrm{w}^{\prime}$ implies that taking limits on both sides of Equation A. 8 we obtain

$$
\left(\forall \mathrm{w}^{\prime} \in \mathrm{W}\right) \lambda_{*}^{T} \mathrm{w}^{*} \geq \lambda_{*}^{T} \mathrm{w}^{\prime} .
$$

[^16]Thus, if $\mathrm{w}^{*} \in \mathrm{~W}$, Theorem 2 implies that $\mathrm{w}^{*} \in \mathrm{~W}_{\mathrm{P}}$ and $\lambda_{*} \in \Delta(\Theta)$ is the direction that witnesses this.

Now, for each $\mathrm{w}_{n}$, a distribution over posteriors $\tau_{n} \in \Delta_{\mu_{0}}(\Delta(\Theta))$ exists such that $\mathrm{w}_{n}=\mathbb{E}_{\tau_{n}}[\hat{\mathrm{w}}]$. Because $\Delta_{\mu_{0}}(\Delta(\Theta))$ is compact, we have that, up to a subsequence, $\tau_{n} \rightarrow \tau^{*} \in \Delta_{\mu_{0}}(\Delta(\Theta))$. Aliprantis and Border (2013, Theorem 15.5) implies $\mathbb{E}_{\tau}[\hat{\mathrm{w}}]$ is upper-semicontinuous as a function of $\tau$, thus for all $i \in\{1, \ldots, N\}$ we have that

$$
\mathbb{E}_{\tau^{*}}\left[\hat{\mathrm{w}}_{i}\right] \geq \lim _{n \rightarrow \infty} \mathbb{E}_{\tau_{n}}\left[\hat{\mathrm{w}}_{i}\right]=\lim _{n \rightarrow \infty} \mathrm{w}_{n, i}=\mathrm{w}_{i}^{*} .
$$

Let $\mathrm{w}^{* *} \equiv \mathbb{E}_{\tau^{*}}[\hat{\mathrm{w}}]$ and note that it is an element of W that dominates $\mathrm{w}^{*}$ coordinateby-coordinate. Moreover, because $\lambda_{*} \in \Delta(\Theta)$, Equation A. 10 implies that $\lambda_{*}^{T} \mathrm{w}^{* *} \geq$ $\lambda_{*}^{T} \mathrm{w}^{*}$. This, together with Equation A.9, implies $\mathrm{w}^{* *} \in \mathrm{~W}_{\mathrm{P}}$, which completes the proof. $\square$

The proof of the statements in Observation 1 follows from the following result:
Corollary A. 1 ((No) Benefit from disclosure). The following hold:

1. All types benefit from disclosure if for all $\theta \in \Theta, \hat{w}(\cdot, \theta)$ is strictly convex in a neighborhood of the prior. Furthermore, if $\hat{w}(\cdot, \theta)$ is everywhere strictly convex, then full disclosure is uniquely Pareto efficient.
2. No disclosure is Pareto efficient if either
    (a) a type $\theta \in \Theta$ exists such that $\hat{w}(\mu, \theta)$ is concave in $\mu$, or
    (b) a vector $a \in \mathbb{R}_{+}^{N}$ and a concave function $w: \Delta(\Theta) \mapsto \mathbb{R}$ exist such that for all $\theta \in \Theta$, the welfare function is given by $a(\theta) w(\mu)+b(\theta)$.

Proof of Corollary A.1. Consider first the conditions in part 1. Toward a contradiction, suppose that $\mathrm{w}^{N D}$ is in the Pareto frontier. Then, by Corollary 3 a direction $\tilde{\lambda} \in \Delta(\Theta)$ exists such that no disclosure is a solution to the supporting Bayesian persuasion problem in direction $\tilde{\lambda}$. However, under the conditions in part 1, the indirect utility function is strictly convex in a neighborhood of the prior for all directions $\lambda \in \Delta(\Theta)$, contradicting that no disclosure is a solution to the supporting Bayesian persuasion problem in some direction $\tilde{\lambda}$ in $\Delta(\Theta)$. Furthermore, when $\hat{w}(\cdot, \theta)$ is strictly convex, so is $\hat{v}_{\lambda}$ for all $\lambda \in \Delta(\Theta)$ and the value of full disclosure strictly dominates that of any other information structure.

The proof of part 2 follows from Theorem 2 by looking at the solution of the supporting Bayesian persuasion problem in certain directions. Under the conditions in part 2a, no disclosure is a solution to the supporting Bayesian persuasion problem in direction $\lambda \in \Delta(\Theta)$ such that $\lambda(\theta)=1$ and $\lambda\left(\theta^{\prime}\right)=0$ for $\theta^{\prime} \neq \theta$. Instead, under the conditions
in part 2b, no disclosure is a solution to the supporting Bayesian persuasion problem in direction $\lambda(\theta)=\mu_{0}(\theta) / a(\theta)$. $\square$

Proof of Observation 2. The result follows from Corollary 3. For the first part, the condition ensures $\mu(\theta) w(\mu, \theta)$ is strictly convex for all $\theta \in\left\{\theta_{1}, \theta_{2}\right\}$, so the indirect utility function $\hat{v}_{\lambda}$ is strictly convex for any $\lambda \in \Delta(\Theta)$ and full disclosure is the unique solution to the supporting Bayesian persuasion problem. For the second part, the condition ensures $\mu(\theta) w(\mu, \theta)$ is weakly concave for type $\theta$, and the result follows from looking at the direction $\lambda$ such that $\lambda(\theta)=1$ and $\lambda\left(\theta^{\prime}\right)=0$ for $\theta^{\prime} \neq \theta$. $\square$

## A. 3 Omitted proofs in Section 5

Proof of Theorem 3. As explained in the main text, necessity follows from noting

$$
\mathrm{w} \in \mathrm{~W} \Rightarrow \mathrm{w}=\mathrm{D}_{0} \sum_{m=1}^{M} \alpha_{m} \mu_{m} \mu_{m}^{T} \rho \equiv \mathrm{D}_{0} \mathrm{C} \rho,
$$

where $M \leq 2 N$ follows from Corollary 1. C is completely positive because it is the convex combination of rank-one non-negative matrices, $\mu_{m} \mu_{m}^{T}$. That $\mathrm{Ce}=\mu_{0}$ follows from the martingale property of beliefs.

For sufficiency, consider $\mathrm{w}=\mathrm{D}_{0} \mathrm{C} \rho$, for some completely positive matrix C, such that $\mathrm{Ce}=\mu_{0}$. Then, $\left\{\mathrm{x}_{1}, \ldots, \mathrm{x}_{\mathrm{M}}\right\} \subseteq \mathbb{R}_{+}^{N}$ exist such that

$$
\mathrm{C}=\sum_{m=1}^{M} \mathrm{x}_{\mathrm{m}} \mathrm{x}_{\mathrm{m}}^{T} .
$$

Let $\sqrt{\alpha_{m}}=\sum_{j=1}^{N} \mathrm{x}_{\mathrm{m} j}$ and note $\mathrm{x}_{\mathrm{m}} /\left(\sqrt{\alpha_{m}}\right) \equiv \mu_{m} \in \Delta(\Theta)$.

$$
\mathrm{C}=\sum_{m=1}^{M} \alpha_{m}\left(\frac{\mathrm{x}_{\mathrm{m}}}{\sqrt{\alpha_{m}}}\right)\left(\frac{\mathrm{x}_{\mathrm{m}}}{\sqrt{\alpha_{m}}}\right)^{T}=\sum_{m=1}^{M} \alpha_{m} \mu_{m} \mu_{m}^{T} .
$$

It remains to show $\sum_{m=1}^{M} \alpha_{m}=1$ and $\sum_{m=1}^{M} \alpha_{m} \mu_{m}=\mu_{0}$. Note that for all $i \in$ $\{1, \ldots, N\}$,

$$
(\mathrm{Ce})_{i}=\sum_{m=1}^{M} \alpha_{m} \mu_{m i} \sum_{j=1}^{N} \mu_{m j}=\sum_{m=1}^{M} \alpha_{m} \mu_{m i}=\mu_{0}\left(\theta_{i}\right) .
$$

Furthermore,

$$
\sum_{i=1}^{N} \mu_{0}\left(\theta_{i}\right)=1=\sum_{i=1}^{N} \sum_{m=1}^{M} \alpha_{m} \mu_{m i}=\sum_{m=1}^{M} \alpha_{m} .
$$

Thus, an information structure exists that generates the distribution over posteriors $\left\{\alpha_{m}, \mu_{m}\right\}_{m=1}^{M}$. Therefore, $\mathrm{w} \in \mathrm{W}$. $\square$

Proof of Claim 2. If $\operatorname{Pr}(X)=0$, the statement is trivial. If $\operatorname{Pr}(X)>0$, denote by $\mathrm{P}_{\mathrm{i}}$. the i-th row of the matrix $\mathrm{P} \equiv \mathrm{D}_{0} \mathrm{C}$, presented as a row-vector. By Bayes' rule, $\operatorname{Pr}(X)=\mu_{0}^{T} \beta$ and $\operatorname{Pr}\left(\theta_{i} \mid X\right)=\left(\mu_{0 i} \beta_{i}\right) /\left(\mu_{0}^{T} \beta\right)$, so

$$
\begin{gathered}
\mathbb{E}_{\Pi}\left[\operatorname{Pr}(X \mid s) \mid \theta_{i}\right]=\sum_{j=1}^{N} \mathbb{E}_{\Pi}\left[\operatorname{Pr}\left[\theta_{j} \mid s\right] \mid \theta_{i}\right] \operatorname{Pr}\left(X \mid \theta_{j}\right)=\mathrm{P}_{\mathrm{i} \cdot} \beta \\
\mathbb{E}_{\Pi}[\operatorname{Pr}(X \mid s) \mid X]=\sum_{i=1}^{N} \operatorname{Pr}\left(\theta_{i} \mid X\right) \mathbb{E}_{\Pi}\left[\operatorname{Pr}(X \mid s) \mid \theta_{i}\right]=\sum_{i=1}^{N} \frac{\mu_{0 i} \beta_{i}}{\mu_{0}^{T} \beta} \operatorname{Pi}_{\mathrm{i} \cdot \beta} .
\end{gathered}
$$

Hence, the truth-drifting condition can be restated as:

$$
\sum_{i=1}^{N} \frac{\mu_{0 i} \beta_{i}}{\mu_{0}^{T} \beta} P_{i \cdot} \beta \geq \mu_{0}^{T} \beta .
$$

Define $\hat{\mathrm{C}} \equiv \mathrm{PD}_{0}=\mathrm{D}_{0} \mathrm{CD}_{0}$. By Theorem 3, $\hat{\mathrm{C}}$ is a completely positive matrix such that $\hat{\mathrm{C}} \mu_{0}=\mathrm{e}$ and $\mu_{0}^{T} \hat{\mathrm{C}} \mu_{0}=1$. Hence, the truth-drifting condition can be restated in a matrix form as:

$$
\left(\frac{\mu_{0} * \beta}{\mu_{0}^{T} \beta}\right)^{T} \hat{\mathrm{C}}\left(\frac{\mu_{0} * \beta}{\mu_{0}^{T} \beta}\right) \geq \mu_{0}^{T} \hat{\mathrm{C}} \mu_{0} .
$$

The term $\zeta \equiv\left(\mu_{0} * \beta\right) /\left(\mu_{0}^{T} \beta\right)$ is an element of the simplex $\Delta(\Theta)$, equal to $\mu_{0}$ when $\beta=$ e. Hence, showing that $\mu_{0}$ is a minimizer of a quadratic form $\zeta^{T} \hat{\mathrm{C}} \zeta$ among all $\zeta \in \Delta(\Theta)$ is enough to prove the result. Noting that we can rely on the Lagrangian approach, at $\zeta=\mu_{0}$, the derivative of the quadratic form is collinear to e and hence, collinear to the space $\Delta(\Theta)$. Thus, first-order conditions are satisfied. At the same time, $\hat{\mathrm{C}}$ is completely positive and thus positive semi-definite. Thus, second-order conditions are satisfied. The result follows. $\square$

Example A. $4\left(\mathrm{w}^{N D}\right.$ and $\mathrm{w}^{F D}$ not on the boundary of W). Consider the case of binary types, $\Theta=\left\{\theta_{1}, \theta_{2}\right\}$. Denote by $\mu \in[0,1]$ the probability of type $\theta_{2}$ and let $\mu_{0}=1 / 2$. Consider the following welfare function:

$$
w\left(\mu, \theta_{1}\right)=\frac{\sin (2 \pi \mu)}{2(1-\mu)}, w\left(\mu, \theta_{2}\right)=\frac{\sin (4 \pi \mu)}{2 \mu},
$$

with $w\left(1, \theta_{1}\right)$ and $w\left(0, \theta_{2}\right)$ defined by continuity as equal to $-\pi$ and $2 \pi$, respectively. Given this welfare function, the truth-adjusted welfare function is

$$
\hat{w}\left(\mu, \theta_{1}\right)=\sin (2 \pi \mu), \hat{w}\left(\mu, \theta_{2}\right)=\sin (4 \pi \mu) .
$$

The corresponding indirect utility in the supporting Bayesian persuasion problem in the direction $\lambda=\left(\lambda_{1}, \lambda_{2}\right)$ is equal to

$$
\hat{v}_{\lambda}(\mu)=\lambda_{1} \sin (2 \pi \mu)+\lambda_{2} \sin (4 \pi \mu) .
$$

For any $\lambda \in \mathbb{R}^{2} \backslash\{0\}, \hat{v}_{\lambda}\left(\mu_{0}\right)=\hat{v}_{\lambda}(1 / 2)=\hat{v}_{\lambda}(0)=\hat{v}_{\lambda}(1)=0$. Hence, both full disclosure and no disclosure results in zero payoff. At the same time, for any such $\lambda$, $\hat{v}_{\lambda}(\mu)$ is a non-constant continuous function anti-symmetric around $\mu=1 / 2$. Hence, it achieves strictly positive values on [0, 1] and $\operatorname{cav} \hat{v}_{\lambda}\left(\mu_{0}\right)>0$ so that optimal disclosure outperforms both full disclosure and no disclosure. Theorem 2-extended to all boundary points-implies that $\mathrm{w}^{N D}$ and $\mathrm{w}^{F D}$ are not on the boundary of W.

The proofs of Proposition 3 and Corollary 4 rely on the graph-theoretic approach in Rayo and Segal (2010), the main properties of which we summarize in Remark A.1:

Remark A. 1 (Rayo and Segal, 2010). Rayo and Segal (2010) propose the following graphical depiction of an information structure, II. Given a direction $\lambda$, let the prospect values $\left(\frac{\lambda\left(\theta_{i}\right)}{\mu_{0}\left(\theta_{i}\right)}, \rho\left(\theta_{i}\right)\right)=\left(\gamma_{i}, \rho_{i}\right)$ for $i=1, \ldots, N$ be vertices of a graph in $\mathbb{R}^{2}$. Connect the points $\left(\gamma_{i}, \rho_{i}\right)$ and $\left(\gamma_{j}, \rho_{j}\right)$ by an edge if and only if a signal $s$ exists such that $\pi\left(s \mid \theta_{j}\right) \pi\left(s \mid \theta_{i}\right)>0$. The set of types that have positive probability under $s$ is called the pooling set of signal $s$.

Lemmas 2-5 in Rayo and Segal (2010) establish that under any optimal information structure, the following hold:

(a) the posterior expectations of the prospect values induced by any two signals are ranked (in vector order), that is, for any two signals $s, s^{\prime}$, either $\left(\mathbb{E}_{\mu_{s}}[\gamma(\tilde{\theta})], \mathbb{E}_{\mu_{s}}[\rho(\tilde{\theta})]\right) \geq$ $\left(\mathbb{E}_{\mu_{s^{\prime}}}[\gamma(\tilde{\theta})], \mathbb{E}_{\mu_{s^{\prime}}}[\rho(\tilde{\theta})]\right)$ or the opposite inequality holds.
(b) prospects appear in the support of some signal only if they lie on a straight line with non-positive slope,
(c) if the pooling segments ${ }^{22}$ of two signals do not lie on the same line, they can intersect only if they share an endpoint, and
(d) if two prospect values are ranked and appear in the support of two signals, then the posterior expectations induced by these signals are ranked in the same way.

Proof of Proposition 3. By the arguments presented in the main text, any optimal information structure solves the instance of the problem of Rayo and Segal (2010) in which the prospect values are ( $0, \rho\left(\theta_{j}\right)$ ) for $j \neq i$ and $\left(1, \rho\left(\theta_{i}\right)\right)$ for the sender and for the receiver, respectively.

[^17]Given the structure of the prospect values in our problem, the property in part (b) implies that $\theta_{i}$ is never pooled with lower-index types. Furthermore, whenever it is pooled with some type, it is pairwise pooled. The property in part (c) implies that whenever $\theta_{i}$ is pooled with some type $\theta_{j}$, then no types $\theta_{k}, \theta_{l}$ with $k<j<l$ can be pooled. Together with the property in part (d), this observation implies that whenever $\theta_{i}$ is pooled with some type $\theta_{j}$, it is also pooled with all types $\theta_{k}$ with $k>j$. Moreover, as $\theta_{i}$ is pooled with increasingly higher-index types, the corresponding posterior expectations increase in vector order, which means that higher signals induce higher reputation yet have a relatively higher proportion of $\theta_{i}$ (if all types are equally likely, then the probability of pooling $\theta_{i}$ with $\theta_{j}$ increases in $j$ ).

It is left to show that the threshold type-the lowest type with which $\theta_{i}$ is pooled-is not pooled with any type of lower index. However, because the threshold type is of higher index than $\theta_{i}$, such pooling could clearly be improved by pooling the threshold type exclusively with $\theta_{i}$. $\square$

Calculations for Example 3. Define the following parameterized family of information structures (rows correspond to types and columns to signals):

$$
\begin{aligned}
& \Pi_{1}(\alpha, \beta)=\left(\begin{array}{ccc}
\alpha & 1-\alpha & 0 \\
1-\beta & \beta & 0 \\
0 & 0 & 1
\end{array}\right), \quad \Pi_{2}(\alpha)=\left(\begin{array}{cc}
\alpha & 1-\alpha \\
1 & 0 \\
0 & 1
\end{array}\right), \\
& \Pi_{3}(\beta)=\left(\begin{array}{cc}
1 & 0 \\
0 & 1 \\
\beta & 1-\beta
\end{array}\right), \quad \Pi_{4}(\alpha, \beta)=\left(\begin{array}{ccc}
1 & 0 & 0 \\
0 & \alpha & 1-\alpha \\
0 & 1-\beta & \beta
\end{array}\right) .
\end{aligned}
$$

Note information structures $\Pi_{1}(1,1)$ and $\Pi_{4}(1,1)$ coincide and correspond to full disclosure. Likewise, information structures $\Pi_{2}(0)$ and $\Pi_{3}(1)$ both correspond to full pooling of types $\theta_{1}$ and $\theta_{3}$. Information structures $\Pi_{2}$ and $\Pi_{3}$ are the $\theta_{1}$-noisy-priority policy and the $\theta_{3}$-noisy-degrading policies, respectively. Like the $\theta_{3}$-noisy-priority policy, $\Pi_{1}$ separates $\theta_{3}$ from $\theta_{1}$ and $\theta_{2}$, but unlike the $\theta_{3}$-noisy-priority policy, it allows for $\theta_{1}$ and $\theta_{2}$ to be pooled, which does not affect $\theta_{3}$ 's expected reputation. Similarly, $\Pi_{4}$ separates $\theta_{1}$ from $\theta_{2}$ and $\theta_{3}$ like the $\theta_{1}$-noisy-degrading policy, but unlike this policy, it allows for $\theta_{2}$ and $\theta_{3}$ to be pooled.

By Proposition 2, any solution to the supporting Bayesian persuasion problem solves the instance of the problem of Rayo and Segal (2010) with prospect values $\left\{\left(\gamma_{i}, \rho_{i}\right):\right.$ $i \in\{1,2,3\}\}$ for the sender and for the receiver, respectively. For simplicity, we assume that the types are strictly ranked under $\rho$, i.e., $\rho_{1}<\rho_{2}<\rho_{3}$.

The property in part (b) in Remark A. 1 implies that if $\left(\gamma_{i}, \rho_{i}\right)<\left(\gamma_{j}, \rho_{j}\right)$, then types $\theta_{i}$ and $\theta_{j}$ are never pooled (cf. Corollary 4). We can then immediately establish the properties of the boundary information structures in the following cases:

- If $\gamma_{1}<\gamma_{2}<\gamma_{3}$, then the uniquely optimal information structure is full disclosure.
- If $\gamma_{2}<\gamma_{1}<\gamma_{3}$, then any optimal information structure separates type $\theta_{3}$ and belongs to class $\Pi_{1}(\alpha, \beta)$.
- If $\gamma_{1}<\gamma_{3}<\gamma_{2}$, then any optimal information structure separates type $\theta_{1}$ and belongs to class $\Pi_{4}(\alpha, \beta)$.
- If $\gamma_{2}<\gamma_{3}<\gamma_{1}$, then an optimal information structure never pools types $\theta_{2}$ and $\theta_{3}$ and belongs to class $\Pi_{2}(\alpha)$.
- If $\gamma_{3}<\gamma_{1}<\gamma_{2}$, then an optimal information structure never pools types $\theta_{1}$ and $\theta_{2}$ and belongs to class $\Pi_{3}(\beta)$.

In the remaining case $\gamma_{3}<\gamma_{2}<\gamma_{1}$, no two prospects are ranked. However, by the property in part (a) in Remark A.1, the induced posterior expectations are necessarily ranked. Hence, if all three prospects lie on a straight line, then no disclosure is optimal. In contrast, if the three prospects do not lie on a straight line, then an optimal information structure separates either types $\theta_{1}$ and $\theta_{2}$ or types $\theta_{2}$ and $\theta_{3}$, and thus belongs to either class $\Pi_{2}(\alpha)$ or to class $\Pi_{3}(\beta)$.

Finally, it is easy to see that by the same arguments, an optimal information structure for the cases in which $\gamma_{i}=\gamma_{j}$ for some $i$ and $j$ belongs to one of the same four classes of information structures.

Knowing the classes of boundary information structures, we can plot the Bayes welfare set in Example 3 by direct calculation. $\square$

## Online Appendix

## B Data Limits

The analysis in the paper assumes that the information structure can arbitrarily condition on an individual's payoff-relevant type. However, this assumption does not necessarily hold in many applications of interest. For instance, regulation may prevent the disclosure of protected characteristics, such as gender or race. Thus, when $\theta$ encompasses such characteristics, considering information structures that respect these restrictions is natural.

In this section, we extend our analysis by removing this assumption. Formally, we consider the following extension of the model in Section 2. Together with the individuals' types, we are given a data source that is potentially informative about these types. The data source has realizations in a finite set $D \equiv\left\{d_{1}, \ldots, d_{M}\right\}$. We describe the joint distribution over payoff-relevant types and data via the prior distribution on $\Theta, \mu_{0}$, and a system of conditional probabilities $\left\{\nu_{0}(\cdot \mid \theta): \theta \in \Theta\right\}$, describing the distribution of the data source $d$ conditional on the types $\theta$. We let $\eta_{0} \in \Delta(D)$ denote the induced marginal distribution on $D .{ }^{23}$ The model in Section 2 corresponds to the case in which $\Theta=D$ and $\nu_{0}(d \mid \theta)=\mathbb{1}[d=\theta]$.

We assume information can be provided to the outside observer only about datasource realizations, but not an individual's type. Formally, an information structure $\Pi=(\pi, S)$ consists of a countable set of labels $S$ and a mapping $\pi$, which associates to each data-source realization, $d$, a distribution over signals $\pi(\cdot \mid d) \in \Delta(S)$. Given an information structure $\Pi$ and a signal realization $s \in S$, updated beliefs about $\theta$ depend only on the updated belief about the realization of $d$. Indeed,

$$
\mu_{s}(\theta)=\sum_{d \in D} \frac{\mu_{0}(\theta) \nu_{0}(d \mid \theta)}{\eta_{0}(d)} \eta_{s}(d),
$$

where $\eta_{s}$ is the marginal on $D$ of the updated joint belief on $\Theta \times D$. It follows that we can define the welfare function as depending on beliefs about $d$ rather than about $\theta$. That is, we can define the function $w_{\dagger}: \Delta(D) \times \Theta \mapsto \mathbb{R}$ as follows:

$$
w_{\dagger}(\eta, \theta)=w(\mu(\eta), \theta),
$$

where the function $\mu(\eta)$ is determined by Equation B.1.
Given an information structure ( $\pi, S$ ), the welfare of an individual of type $\theta$ is

$$
w_{\Pi}(\theta) \equiv \mathbb{E}_{\langle\Pi \mid \theta\rangle}\left[w_{\dagger}(\tilde{\eta}, \theta)\right]=\sum_{s \in S} \sum_{d \in D} \nu_{0}(d \mid \theta) \pi(s \mid d) w_{\dagger}\left(\eta_{s}, \theta\right),
$$

[^18]and the Bayes welfare set continues to be defined as the set of Bayes welfare profiles.
We now show the analysis in the main text extends verbatim. Indeed, by the same arguments as in Section 3, the welfare of an individual of type $\theta$ under information structure $\Pi=(\pi, S)$ can be written as:
$$
w_{\Pi}(\theta)=\mathbb{E}_{\langle\Pi \mid \theta\rangle}\left[w_{\dagger}(\tilde{\eta}, \theta)\right]=\mathbb{E}_{\langle\Pi\rangle}\left[\hat{w}_{\dagger}(\tilde{\eta}, \theta)\right],
$$
where the truth-adjusted welfare function $\hat{w}_{\dagger}$ now takes the form:
$$
\hat{w}_{\dagger}(\eta, \theta)=\sum_{d \in D} \nu_{0}(d \mid \theta) \frac{\eta(d)}{\eta_{0}(d)} w_{\dagger}(\eta, \theta) .
$$
By separating the variable on which welfare is conditioned on-the payoff-relevant types, $\theta$-from the variable about which information is provided-the data source, $d-$Equation B. 4 allows us to provide further insight into the truth-adjusted welfare function in the model in Section 2. Indeed, note the likelihood correction is based on the variable $d$, highlighting that it corresponds to the variable about which information is provided. Similar to before, we can interpret the likelihood-ratio adjustment as describing that each data-source realization $d$ has a budget $\eta_{0}(d)$ to be distributed across different (data) posteriors $\eta$. Unlike the analysis before, individuals of type $\theta$ only own a fraction $\nu_{0}(d \mid \theta)$ of this ratio.

Equation B. 3 implies Theorem 1 immediately extends to this setting:
Theorem B.1. The Bayes welfare set W satisfies the following:

$$
\mathrm{W}=\left\{\mathrm{w} \in \mathbb{R}^{N}:\left(\eta_{0}, \mathrm{w}\right) \in \operatorname{co}\left(\operatorname{graph} \hat{\mathrm{w}}_{\dagger}\right)\right\} .
$$

In what follows, we explore how the Bayes welfare set changes as we change the informativeness of the data source. Intuitively, we would expect that the Bayes welfare set shrinks as data becomes less precise. Proposition B. 1 below shows that this intuition holds when the notion of less precise coincides with the notion of garbling in Blackwell (1953).

Formally, given the distribution of payoff-relevant types $\mu_{0} \in \Delta(\Theta)$, we wish to understand the effect of different data sources, as described by data-source realizations $D^{\prime}$ and conditional probability systems $\left\{\nu_{0}^{\prime}(\cdot \mid \theta) \in \Delta\left(D^{\prime}\right): \theta \in \Theta\right\}$. Following Blackwell (1953), we say $\left(D^{\prime}, \nu_{0}^{\prime}\right)$ is a garbling of ( $D, \nu_{0}$ ) if a stochastic matrix $G: D \mapsto \Delta\left(D^{\prime}\right)$ exists such that for every data-type pair $\left(d^{\prime}, \theta\right)$,

$$
\nu_{0}^{\prime}\left(d^{\prime} \mid \theta\right)=\sum_{d \in D} G\left(d^{\prime} \mid d\right) \nu_{0}(d \mid \theta)
$$

Let $\mathrm{W}\left(\mu_{0}, w, D, \nu_{0}\right)$ denote the Bayes welfare set for prior type distribution $\theta$ and welfare function $w$, as we vary the (informativeness of the) data source $\left(D, \nu_{0}\right)$. We then have the following:

Proposition B. 1 (Data Comparison). $\mathrm{W}\left(\mu_{0}, w, D^{\prime}, \nu_{0}^{\prime}\right) \subseteq \mathrm{W}\left(\mu_{0}, w, D, \nu_{0}\right)$ for all welfare functions $w$ and type distributions $\mu_{0}$ if and only if $\left(D^{\prime}, \nu_{0}^{\prime}\right)$ is a garbling of $\left(D, \nu_{0}\right)$.

Proof of Proposition B.1. One direction is straightforward: if $\left(D^{\prime}, \nu_{0}^{\prime}\right)$ is a garbling of $\left(D, \nu_{0}\right)$, any distribution of signals conditional on payoff-relevant types induced by some information structure under data source ( $D^{\prime}, \nu_{0}^{\prime}$ ) is feasible under data source $\left(D, \nu_{0}\right)$. Consequently, any welfare profile that can be induced by some information structure under ( $D^{\prime}, \nu_{0}^{\prime}$ ) can be induced under ( $D, \nu_{0}$ ).

To obtain the other direction, toward a contradiction, assume $\left(D^{\prime}, \nu_{0}^{\prime}\right)$ is not a garbling of $\left(D, \nu_{0}\right)$. Then, by Blackwell (1953), a prior $\mu_{0} \in \Delta(\Theta)$ and a payoff function $u:$ $A \times \Theta \rightarrow \mathbb{R}$ exist such that a decision maker with utility $u$ derives strictly greater value from having access to $\left(D^{\prime}, \nu_{0}^{\prime}\right)$ than to $\left(D, \nu_{0}\right)$. That is, letting $U(\mu)$ denote the decision maker's indirect utility, $\max _{a \in A} \mathbb{E}_{\mu}[u(a, \theta)]$, we have that:

$$
\sum_{\theta \in \Theta} \mu_{0}(\theta) \mathbb{E}_{\left(D^{\prime}, \nu_{0}^{\prime}\right)}[U(\mu) \mid \theta]>\sum_{\theta \in \Theta} \mu_{0}(\theta) \mathbb{E}_{\left(D, \nu_{0}\right)}[U(\mu) \mid \theta] .
$$

Consider now the Bayes welfare sets given the welfare function $w(\mu, \theta)=U(\mu)$ and prior distribution $\mu_{0}$, under data sources $\left(D, \nu_{0}\right)$ and $\left(D^{\prime}, \nu_{0}^{\prime}\right)$. We have that

$$
\begin{aligned}
\sum_{\theta \in \Theta} \mu_{0}(\theta) \mathbb{E}_{\left(D^{\prime}, \nu_{0}^{\prime}\right)}[U(\mu) \mid \theta] & =\max _{\mathrm{w} \in \mathrm{~W}\left(\cdot, D^{\prime}, \nu_{0}^{\prime}\right)} \sum_{\theta \in \Theta} \mu_{0}(\theta) \mathrm{w}(\theta) \\
& \leq \max _{\mathrm{w} \in \mathrm{~W}\left(\cdot, D, \nu_{0}\right)} \sum_{\theta \in \Theta} \mu_{0}(\theta) \mathrm{w}(\theta)=\sum_{\theta \in \Theta} \mu_{0}(\theta) \mathbb{E}_{\left(D, \nu_{0}\right)}[U(\mu) \mid \theta],
\end{aligned}
$$

where the equalities follow because the maximal ex ante payoff is obtained by having full access to available data, and the inequality follows from the assumption that $\mathrm{W}\left(\mu_{0}, w, D^{\prime}, \nu_{0}^{\prime}\right) \subseteq \mathrm{W}\left(\mu_{0}, w, D, \nu_{0}\right)$ for all $w$ and $\mu_{0}$. Comparing Equations B. 6 and B. 7 leads to the desired contradiction and the result follows. $\square$

Remark B. 1 (When to blind an algorithm). Proposition B. 1 stands in contrast with the recommendation in the algorithmic fairness literature to "blind" algorithms to sensitive inputs such as race or gender. Indeed, having taken into account the impact of the outside observer's incentives in the population's welfare, allowing the information structure to condition on the individuals' payoff-relevant types leads to the largest Bayes welfare set, thereby (weakly) increasing the value of any social welfare function that is used to choose what information structure to implement.

> [Figure omitted from this text-only corpus; refer to the source manuscript.]
Figure B.1: Noisy data in the online marketplace. Bayes welfare sets for different values of $\sigma \in\{0.55,0.7,0.85,1\}$.

There may be other reasons outside our model that could justify blinding the information structure to the individuals' types. For instance, Liang et al. (2022) show that when an agent different from the social planner selects a decision-making algorithm, the social planner may prefer to restrict the inputs into the agent's algorithm. Even though in our model algorithms are information structures that send non-binding recommendations, a similar result would hold in our setting.

We conclude this section with an example that illustrates the following two points. First, whereas Proposition B. 1 shows that less precise data sources limit the ability to generate and distribute welfare via information, the example shows that this effect is not uniform across individuals of different types. Second, the Bayes welfare set may collapse to the no-disclosure Bayes welfare profile for data sources that are strictly more informative than no information in the Blackwell order.

Example B. 2 (Example 2 continued; Noisy Data). Suppose the online marketplace only has access to a noisy estimate of the consumer's type, perhaps from past purchases or undeleted cookies. We model this as a data source that reveals a consumer's type with a fixed precision $\sigma \in[1 / 2,1]: D=\left\{d_{1}, d_{2}\right\}$ and $\nu_{0}\left(d_{i} \mid \theta_{i}\right)=\sigma$. When $\sigma=1$, the data source is perfectly informative about a consumer's type; when $\sigma=1 / 2$, the data source is pure noise. More generally, if $\sigma<\sigma^{\prime}$, the data source that corresponds to $\sigma$ is a garbling of the data source that corresponds to $\sigma^{\prime}$.

Figure B. 1 illustrates the Bayes welfare set W for different precision values. Three features are worth noting. First, in line with Proposition B.1, Bayes welfare sets resulting from data sources with lower precision are subsets of those with higher precision.

When $\sigma=1$, the Bayes welfare set naturally coincides with the one in Figure 3b in the main text. Second, at high values of $\sigma$, lower data precision has asymmetric effects across types: it decreases the maximal payoff of $\theta_{L}$-consumers without affecting their minimal payoff, yet it increases the minimal payoff of $\theta_{H}$-consumers without affecting their maximal payoff. Indeed, for sufficiently low values of $\sigma$, the unique Pareto efficient information structure is the one that maximizes the payoff of $\theta_{H}$-consumers. That is, in this example, lower data precision benefits $\theta_{H}$-consumers. Finally, whereas it is immediate that the Bayes welfare set coincides with the no-disclosure profile $\mathrm{w}^{N D}$ when $\sigma=1 / 2$, the Bayes welfare set actually collapses to this point at $\sigma=3 / 5$ : Once $\sigma<3 / 5$, generating Bayes plausible distributions over posteriors with support outside the interval $[1 / 2,3 / 4)$ is not possible, and on this interval, $w$ is constant. This feature highlights that an incrementally more informative data source may have a discontinuous impact on welfare redistribution possibilities. $\square$


[^0]:    *We thank the Editor, Emir Kamenica, and three anonymous referees for feedback that has greatly improved this paper. For valuable suggestions and comments, we would like to thank Ricardo Alonso, Odilon Câmara, Navin Kartik, Elliot Lipnowski, Antonio Penta, Jean Tirole, and Kai Hao Yang, as well as seminar participants at Toulouse School of Economics, Columbia, Bonn Winter Theory Workshop 2021, Warwick Theory Workshop 2022, Stony Brook 2022, ESSET 2022, EEA-ESEM 2022, and Virtual Seminars in Economic Theory. We thank Shunsuke Matsuno for excellent research assistance. Smolin acknowledges funding from the French National Research Agency (ANR) under the Investments for the Future (Investissements d'Avenir) program (grant ANR-17-EURE-0010).
    ${ }^{\dagger}$ Columbia Business School and CEPR. E-mail: laura.doval@columbia.edu.
    ${ }^{\ddagger}$ Toulouse School of Economics and CEPR. E-mail: alexey.v.smolin@gmail.com.

[^1]:    ${ }^{1}$ A matrix $\mathrm{C} \in \mathbb{R}^{N \times N}$ is completely positive if non-negative vectors $\mathrm{c}_{1}, \ldots, \mathrm{c}_{K} \in \mathbb{R}_{+}^{N}$ exist such that $\mathrm{C}=\sum_{i=1}^{K} \mathrm{c}_{i} \mathrm{c}_{i}^{T}$ (Berman, 1988).

[^2]:    ${ }^{2}$ Because maximizing uncertainty can sometimes be accomplished by providing some information, this notion of data privacy differs from an unambiguous preference for no information disclosure.

[^3]:    ${ }^{3}$ Note that the welfare function $w$ differs from the sender's indirect utility function in Kamenica and Gentzkow (2011), usually denoted by $\hat{v}$. The indirect utility function is the expectation under $\mu$ of the welfare function, $w(\mu, \cdot)$. That is, $\hat{v}(\mu)=\sum_{\theta \in \Theta} \mu(\theta) v(a(\mu), \theta)=\sum_{\theta \in \Theta} \mu(\theta) w(\mu, \theta)$.

[^4]:    ${ }^{4}$ Claim 2 provides a more general version of this result based on Theorem 3.
    ${ }^{5}$ Formally, for any information structure, $\Pi$, and type $\theta \in \Theta, \mathbb{E}_{\langle\Pi\rangle}\left[\tilde{\mu}(\theta) / \mu_{0}(\theta)\right]=1$.

[^5]:    ${ }^{6}$ Rosar (2017) and Quigley and Walther (2019) similarly observe that the distribution over posteriors conditional on an individual's type can be written in terms of the modified unconditional distribution.

[^6]:    ${ }^{7}$ For a real-valued function $f, \operatorname{cav} f$ denotes the smallest concave function that dominates $f$ and vex $f$ denotes the highest convex function dominated by $f$ (Hiriart-Urruty and Lemaréchal, 2004).
    ${ }^{8}$ See also Aumann and Maschler (1995) and Rayo and Segal (2010).

[^7]:    ${ }^{9}$ The maximum is attained because by definition $\mathrm{w} \in \mathrm{W}$. As we show in Appendix A.2, that $\lambda \in$ $\mathbb{R}_{+}^{N} \backslash\{0\}$ follows from $\mathrm{w} \in \mathrm{W}_{\mathrm{P}}$.

[^8]:    ${ }^{10}$ Thus, one can always interpret the heterogeneous priors model in Alonso and Câmara (2016) as a model in which the sender and the receiver share the same prior, but the sender assigns weights different than those under the prior $\mu_{0}$ to each of his possible types.

[^9]:    ${ }^{11}$ For instance, the midpoint $\mathrm{w}=(3 / 4,1 / 2)$ can only be generated by an information structure that employs at least three signals. One such information structure is given by

    $$
    \begin{array}{c|ccc}
    \theta_{H} & 1 / 4 & 3 / 8 & 3 / 8 \\
    \theta_{L} & 0 & 1 / 4 & 3 / 4
    \end{array} .
    $$

[^10]:    ${ }^{12}$ Appendix A. 2 provides weaker conditions under which all types (do not) benefit from disclosure.

[^11]:    ${ }^{13}$ For an even starker example, consider $\Theta=\left\{\theta_{1}, \theta_{2}\right\}$ and $w\left(\mu, \theta_{1}\right)=-\mu^{2}-2 \mu+3, w\left(\mu, \theta_{2}\right)=$ $-\mu^{2}+4 \mu$, where $\mu \equiv \mu\left(\theta_{2}\right)$. For each type, the welfare function is concave; however, for any prior distribution, the truth-adjusted welfare functions are globally convex. As a result, the unique Pareto efficient information structure is full disclosure.

[^12]:    ${ }^{14}$ Completely positive matrices have been studied extensively as they play an important role in optimization theory, machine learning, and other applications (Berman and Shaked-Monderer, 2003). A completely positive matrix is symmetric and positive-semidefinite, with positive elements; for $N \leq 4$, the converse is also true.

[^13]:    ${ }^{15}$ An analogous characterization appears in concurrent work by Sayin and Başar (2021), who discuss the computational advantages of working with completely positive matrices to solve information design problems.
    ${ }^{16} \mathrm{P}^{2} \mathrm{e}=\mathrm{Pe}=\mathrm{e}$ and $\mathrm{P}^{2}=\mathrm{D}_{0} \mathrm{C}^{\prime}$, where $\mathrm{C}^{\prime} \equiv \mathrm{CD}_{0} \mathrm{C}$ is completely positive because C is symmetric.

[^14]:    ${ }^{17}$ For concreteness, $X$ can be seen as a subset of $\Theta \times[0,1]$ equipped with a probability measure that agrees with $\mu_{0}$ on $\Theta$ (Green and Stokey, 1978; Gentzkow and Kamenica, 2017).
    ${ }^{18}$ In a setting in which information respects the state space's ordinal structure, Kartik et al. (2021) formalize the sense in which the drift toward the truth is stronger for more Blackwell informative information structures.
    ${ }^{19}$ Recall that the relative interior of a set $X$ is the interior of $X$ within its affine hull, which is the set of all affine combinations of elements in $X$.

[^15]:    ${ }^{20}$ We follow the terminology in Galperti et al. (2023), who highlight that individuals with certain types are valuable precisely because of the possibility of pooling them with other types.

[^16]:    ${ }^{21}$ Taking $\Theta$ to be a compact Polish space, we endow the set of Borel probability measures on $\Theta, \Delta(\Theta)$, and on $\Delta(\Theta), \Delta(\Delta(\Theta))$, with the weak* topology, so they are also compact Polish (Aliprantis and Border, 2013, Theorems 15.11 and 15.12).

[^17]:    ${ }^{22}$ By part (b), the pooling set of a signal lies on a segment.

[^18]:    ${ }^{23}$ Although we should index $\eta_{0}$ by $\left(\mu_{0},\left(\nu_{0}(\cdot \mid \theta)\right)_{\theta \in \Theta}\right)$, we omit this dependence to simplify notation.

## Citation and provenance

**Authors:** Laura Doval; Alex Smolin

**Canonical citation:** Doval, Laura, and Alex Smolin. “Persuasion and Welfare.” Journal of Political Economy 132, no. 7 (2024): 2451–2487.

**Canonical machine-readable version:** https://alexsmolin.com/corpus/papers/persuasion-and-welfare.md

**Source record:** https://arxiv.org/abs/2109.03061

**Published record:** https://doi.org/10.1086/729067

**Attribution guidance:** Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.

Provenance metadata: https://alexsmolin.com/corpus/PROVENANCE.txt
