{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0001","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"Document metadata","text":"> Machine-readable author manuscript.\n> Authors: Dirk Bergemann; Alessandro Bonatti; Alex Smolin.\n> Canonical citation: Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.\n> Attribution and provenance: https://alexsmolin.com/corpus/PROVENANCE.txt","text_sha256":"8ba5885a6ff07fa7c9efedbf72f196b06cf70729c1fde41b0684cc61339d4d4f"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0002","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"Menu Pricing of Large Language Models","text":"# Menu Pricing of Large Language Models\n\n**Authors:** Dirk Bergemann; Alessandro Bonatti; Alex Smolin\n\n**Manuscript date:** 2026-03-10\n\n#### Abstract\n\nWe develop a framework for the optimal pricing and product design of LLMs in which a provider sells menus of token budgets to users who differ in their valuations across a continuum of tasks. Under a homogeneous production technology, we show that users' high-dimensional type profiles are summarized by a scalar index, reducing the seller's problem to one-dimensional screening. The optimal mechanism takes the form of committed-spend contracts: buyers pay for a budget that they allocate across token classes priced at marginal cost. We extend the analysis to environments with multiple differentiated models and to competition between a proprietary leader and an open-source fringe, showing that competitive pressure reshapes both the intensive and extensive margins of compute provision. Each element of our theory (token-budget menus, maximum- and minimum-spend plans, multi-model versioning, and linear API pricing) has a direct counterpart in the observed pricing practices of providers such as Anthropic, OpenAI, and GitHub.\n\nKeywords: Large Language Models, Optimal Pricing, Menu Pricing, Fine-Tuning. JEL Codes: D47, D82, D83.\n\n[^0]","text_sha256":"69eacbe16bb45853512054b11921d74a5ec07628b3e4daad9b2306196c1e65b3"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0003","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"1 Introduction","text":"## 1 Introduction\n\nAccess to large language models is fast becoming a major input to economic activity. Businesses use LLMs for coding, customer support, legal research, and content generation; individual users rely on them for writing, analysis, and decision-making. The leading providers, Anthropic, OpenAI, and Google, collectively serve hundreds of millions of users and generate billions of dollars in annualized revenue. Yet the pricing of these services remains strikingly ad hoc: subscription tiers, token-based metering, credit systems, and volume commitments coexist with no clear theoretical foundation. Understanding how a profit-maximizing provider should price LLM access, and quantifying the distortions that arise when it does, is the goal of this paper. ${ }^{1}$\n\nPricing access to LLMs is fundamentally a multidimensional screening problem. An LLM user faces a continuum of tasks, each with a different value. The provider offers token budgets-bundles of input, output, and fine-tuning tokens-that the user allocates across tasks at her discretion. The provider can meter total token usage but cannot observe or contract on the user's task-by-task allocation, creating a combined adverse-selection and moral-hazard problem. User types are infinite-dimensional (a function from tasks to values), the allocation space is high-dimensional (tokens of multiple classes across many tasks), and the user's hidden action (token allocation across tasks) further compounds the difficulty. A priori, this problem seems intractable.\n\nOur central result is that it is not. Under a homogeneous production technology-meaning that the gain function is homogeneous, so the optimal composition of inference tokens is scale-invariant-each user's entire type profile collapses to a scalar index, the $a g$ gregate type. Users with the same aggregate type make the same total token demands, receive the same total surplus, and choose the same menu item, regardless of the fine structure of their task valuations. This aggregation property reduces the provider's problem to one-dimensional screening à la Mussa and Rosen (1978), despite the underlying complexity of the environment. Building on this reduction, we develop the analysis in four steps.\n\nFirst, we characterize the efficient allocation (Section 3). Efficiency requires that all tasks employ inference token classes in common proportions; only the scale of token usage varies with a task's marginal value. When the provider faces capacity constraints in each token class, the efficient allocation can be implemented through linear prices equal to inflated shadow costs (Corollary 1), a result that rationalizes the linear, pay-per-token pricing universally observed in developer-facing API markets.\n\n[^1]Second, we characterize the optimal mechanism for a single-model monopolist (Section 4). By the aggregation result, the buyer's indirect utility from a token-budget bundle takes a multiplicative form: aggregate type times aggregate quality. The optimal menu therefore excludes low types and distorts quality downward for (almost) all others, exactly as in the standard one-dimensional framework. Crucially, the optimal mechanism admits intuitive indirect implementations (Proposition 4): it can be realized as a maximum-spend mechanism (a budget of credits allocated across tokens priced at marginal cost), as a minimum-spend mechanism (volume commitments that unlock lower per-token prices), or as a menu of twopart tariffs. These implementations are not merely theoretical curiosities; they correspond precisely to the pricing structures observed at leading platforms.\n\nThird, we extend the framework to multiple differentiated models (Section 5). When models share a common returns-to-scale parameter in inference but differ in returns to finetuning, the buyer's payoff has a constant elasticity of substitution over model qualities. Cost minimization implies that each type uses a single model for all tasks (Proposition 6). This generates a natural versioning structure: higher-tier plans grant access to more capable models, not merely larger usage allocations. Anthropic's pricing illustrates the single-model benchmark of Section 4: all paid tiers access the same models, with differentiation occurring through compute budgets. OpenAI, by contrast, screens jointly on usage and model access, reserving its most compute-intensive reasoning model for the highest tier, consistent with the multi-model menu of Section 5.\n\nFourth, we study competition between a proprietary leader and an open-source fringe that sells tokens at marginal cost (Section 5). The leader designs an optimal menu subject to the buyer's ability to supplement usage from the fringe. Three regions emerge: low types purchase exclusively from the fringe; intermediate types buy from the leader at quantities that exactly deter fringe top-up; and high types are served as under pure monopoly, with standard downward distortions (Proposition 7). Competition thus reshapes both the intensive margin (how much compute is sold to each buyer) and the extensive margin (which buyers adopt the proprietary model at all).\n\nWe conclude by connecting each theoretical construct to observed pricing (Section 6). Consumer subscriptions at Anthropic and OpenAI implement the nonlinear menus of Sections 4 and 5. Model aggregators (e.g., Quora's Poe and GitHub Copilot) implement the committed-spend mechanisms of Section 4.2, differing in whether they enforce a hard budget cap (maximum spend) or allow overages at linear prices (minimum spend). API pricing is linear in tokens with no volume discounts, consistent with providers prioritizing adoption over rent extraction. Thus, pricing practices observed in the industry suggest that these mechanisms capture fundamental economic forces rather than incidental design choices.","text_sha256":"5e5cb9a206517f55525c4eb64a0f278f638a1be428f7655d63738aead4521bc5"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0004","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"1 Introduction","text":"Related Literature Our paper contributes to the rapidly growing literature on the economics of AI. Existing work has examined the impact of LLMs on price competition (Fish, Gonczarowski, and Shorrer, 2024), token auctions (Duetting, Mirrokni, Paes Leme, Xu, and Zuo, 2024), and sponsored search (Bergemann, Bojko, Dütting, Paes Leme, Xu, and Zuo, 2024), and Demirer, Fradkin, Tadelis, and Peng (2025) document pricing and market-share patterns. By contrast, our focus is on the provider's own design problem, how to price and version LLM access, which has received surprisingly little theoretical attention. For example, Mahmood (2024) studies competition among producers of generative AI who can specialize in different tasks, but restricts attention to linear token pricing. We go further and characterize the fully optimal nonlinear mechanism, leveraging the sufficient-statistic reduction generated by the homogeneity of the gain function, which renders the joint screening-and-moral-hazard problem solvable and yields sharp, implementable predictions.\n\nAt their core, LLMs are an information technology, connecting our work to the literature on selling information (Babaioff, Kleinberg, and Paes Leme, 2012; Bergemann, Bonatti, and Smolin, 2018; Yang, 2022). However, LLMs possess distinctive features: they are generalpurpose technologies deployed across many tasks, usage is metered in tokens, and precision is improved by combining inference and fine-tuning, which give rise to a structurally different design problem.\n\nMethodologically, our aggregation result relates to the literature on \"1.5-dimensional\" mechanism design (Fiat, Goldner, Karlin, and Koutsoupias, 2016; Devanur, Goldner, Saxena, Schvartzman, and Weinberg, 2020). Unlike most of that literature, however, we must address potential incentive clashes across the continuum of subproblems that arise from the buyer's hidden allocation of tokens, requiring an additional argument based on the production technology. Even in simpler settings, such as one-dimensional screening with moral hazard (Castro-Pires, Chade, and Swinkels, 2024) or multidimensional screening and bundling without moral hazard (Armstrong, 1996; Rochet and Stole, 2003; Daskalakis, Deckelbaum, and Tzamos, 2017), tractability issues arise. Our environment combines both, and our solution exhibits a constrained efficiency property as in Laffont and Tirole (1990) and Doligalski, Dworczak, Krysta, and Tokarski (2025).","text_sha256":"260a364369d207f468a8fdad5a95a33e7496d4bcd333dc41db285047357253b0"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0005","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"2 Model","text":"## 2 Model","text_sha256":"69d3a5f84483f9aafd18fefcbfe33e9bab8536442b3889a76ddff069b3a561b0"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0006","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"2.1 Task Environment","text":"### 2.1 Task Environment\n\nA buyer of LLM services faces a unit measure of tasks indexed by $i \\in[0,1]$. In our baseline setting, an LLM provider offers one type of model only. The buyer can decide on (i) how\nmany inference tokens to use for each task $i, x_{i} \\in \\mathbb{R}_{+}^{J}$ and (ii) how many fine-tuning tokens to use to improve the model's performance (e.g., precision) on all tasks, $z \\in \\mathbb{R}_{+}^{K}$. In language models, a token is a unit of text, typically a word.\n\nThe model's performance on task $i$ is given by the gain function $g\\left(x_{i}, z\\right)$. The function $g$ is common across tasks and takes the following form:\n\n$$\ng\\left(x_{i}, z\\right)=\\Psi\\left(x_{i}\\right) \\Phi(z) .\n$$\n\nThus, the gain function is multiplicatively separable in inference and fine-tuning tokens.\nDiminishing returns to inference tokens on a single task are a fundamental feature of LLM performance, consistent with empirically observed scaling laws. Thus the function $\\Psi$ is positive, increasing, strictly concave, differentiable, Inada at 0, ${ }^{2}$ and homogeneous of degree $\\sigma \\in(0,1)$, i.e., $\\Psi\\left(r x_{i}\\right)=r^{\\sigma} \\Psi\\left(x_{i}\\right)$ for all $r>0$. This function captures the returns to inference tokens. A canonical example of such a function is a CES function, $\\Psi\\left(x_{i}\\right)=\\left(\\sum_{j=1}^{J} \\alpha_{j} x_{i j}^{\\rho}\\right)^{\\sigma / \\rho}$, where $\\alpha_{j}>0, \\sum_{j} \\alpha_{j}=1, \\rho<1$. The function $\\Phi(z)$ is positive, increasing, such that $g$ is strictly concave (e.g., $\\Phi(z)^{1 /(1-\\sigma)}$ is strictly concave) and $(1-\\sigma)$-Inada at 0. ${ }^{3}$ This function captures the returns to fine-tuning.\n\nThese properties are similar in shape to the scaling laws for training LLMs. Concavity captures the fundamental assumption that time and computing resources spent on a single task exhibit diminishing marginal returns. Homogeneity is the key tractability assumption. It ensures that the optimal mix of token classes is the same for every task; only the scale of token usage varies. This property drives the aggregation result in Section 3, which reduces the seller's infinite-dimensional screening problem to a one-dimensional problem. The parameter $\\sigma$ captures the returns to scale.\n\nThe buyer's marginal value of performance on each task is captured by $w=\\left(w_{i}\\right)_{i \\in[0,1]}$,\n\n$$\nw:[0,1] \\rightarrow[0,1],\n$$\n\nwhich we refer to as the buyer's type. Using a $z$-fine-tuned model with a profile of $\\left(x_{i}\\right)_{i \\in[0,1]}$ inference tokens delivers the following total payoff for buyer type $w$ :\n\n$$\n\\int_{0}^{1} w_{i} g\\left(x_{i}, z\\right) d i .\n$$\n\nThe buyer's type is distributed according to a commonly known distribution $F_{w}$. The buyer\n\n[^2]knows his type, while the provider does not.\nThe provider bears the cost of processing tokens. We assume that the marginal processing costs are constant but can vary across different token classes. The cost of inference tokens is $c_{j}>0, j \\in[J]$, and the cost of fine-tuning tokens is $\\hat{c}_{k}>0, k \\in[K]$. We let $c \\in \\mathbb{R}_{++}^{J}$ and $\\hat{c} \\in \\mathbb{R}_{++}^{K}$ denote the corresponding cost profiles.\n\nNotation Throughout the text, we use the following notation. \" $\\triangleq$ \" indicates a definition. $[K] \\triangleq\\{1, \\ldots, K\\} . \\mathbb{R}_{+}^{K} \\triangleq\\left\\{x \\in \\mathbb{R}^{K}: x_{k} \\geq 0, k \\in[K]\\right\\}$ and $\\mathbb{R}_{++}^{K} \\triangleq\\left\\{x \\in \\mathbb{R}^{K}: x_{k}>\\right.$ $0, k \\in[K]\\}$. All stand-alone qualifiers such as \"positive,\" \"increasing,\" and \"concave\" are understood in the weak sense, that is, as \"non-negative,\" \"non-decreasing,\" and \"weakly concave,\" respectively. For constants, subscripts refer to variables or tasks; for functions, they indicate partial derivatives. All proofs are in Appendix A.","text_sha256":"73066a3f8bf9e54991f8e14e1e2262c643660c7d33df7db1afd4b8a6834ef959"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0007","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"2.2 Mechanism Design","text":"### 2.2 Mechanism Design\n\nThe provider contracts on the total number of tokens of each class used by the buyer. In other words, the provider sells budgets of various kinds of tokens. A mechanism consists of an arbitrary menu of the form\n\n$$\n\\left\\{\\left(X_{1}(w), \\ldots, X_{J}(w), Z_{1}(w), \\ldots, Z_{K}(w), t(w)\\right)\\right\\}_{w},\n$$\n\nwhere $X_{j}$ is the total number of inference tokens of class $j$. Upon purchase, the buyer can freely distribute those tokens across tasks. (We relax this restriction in Appendix B, where we allow the provider to contract on the usage of all tokens by the buyer.) $Z_{k}$ is the number of fine-tuning tokens of class $k$, which is used for fine-tuning and cannot be distributed.\n\nThe seller's problem is an optimal mechanism design problem with infinite-dimensional private information and moral hazard. However, we show that the homogeneity of the gain function renders the problem tractable.","text_sha256":"8331f3c5c69946613c07fe1f7b58f8a6f467b110a2b82dba27e3b5350296c210"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0008","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"2.3 Mapping the Setting to Practice","text":"### 2.3 Mapping the Setting to Practice\n\nBefore turning to the analysis, we describe how the primitives of our setting relate to the design and operation of contemporary large language models (LLMs).\n\nFine-tuning tokens, $\\boldsymbol{z} \\in \\mathbb{R}_{+}^{\\boldsymbol{K}}$ The case $z=0$ corresponds to a baseline model that has not been fine-tuned. ${ }^{4}$ In this case, the model can be interpreted as a foundation model. Its\n\n[^3]quality depends on its size (i.e., the number of parameters) ${ }^{5}$, its architecture, the quantity and quality of training data, and the details of the training procedure. Scaling up model size typically improves capability and accuracy, but state-of-the-art performance generally requires commensurate scaling of training data as well (Kaplan et al., 2020). In our setting, we take pretraining as given. Accordingly, we focus on inference-time compute rather than training-time compute.\n\nWhen $z>0$, this variable captures the number of tokens used to fine-tune the foundation model. Fine-tuning resembles pretraining in that it updates the model's parameters, without changing its size or architecture, but it is much smaller in scale. It is performed on a dataset directly relevant to a class of tasks (e.g., labeled X-ray scans, code examples, or conversational transcripts) and can be interpreted as injecting task-specific knowledge and behavior into the model.\n\nInference tokens, $\\boldsymbol{x}_{\\boldsymbol{i}} \\in \\mathbb{R}_{+}^{\\boldsymbol{J}}$ These are the tokens used directly to process a given task. The two standard classes of inference tokens are input and output tokens. Input tokens are those provided by the user. Increasing the number of input tokens can improve predictive quality for two reasons. First, richer prompts provide more context, enabling more tailored and appropriate responses. Second, additional input may supply the data on which the model is meant to operate. For example, retrieval-augmented generation (RAG) uses the prompt to retrieve relevant passages from an external corpus and appends them to the prompt.\n\nOutput tokens are those generated by the model. A larger number of output tokens can also improve quality. First, it allows the model to provide more detailed and nuanced answers. Second, it provides opportunities for multi-step reasoning, often described as chainof-thought (CoT) computation. This additional reasoning channel can enable the model to solve more complex tasks, at the cost of generating many intermediate tokens, often hidden from the user. This inference-time reasoning is qualitatively distinct from pretraining, remains only partially understood, and is an active frontier for improving LLM performance. ${ }^{6}$\n\nGain function, $\\boldsymbol{g}$ Tokens of different classes enter the large language model as distinct inputs. They are neither perfect substitutes nor perfect complements, but each can contribute to higher predictive quality. This motivates our gain-function formulation in (1), which is also in line with the empirically observed scaling laws for inference-time computation (e.g., Wu, Sun, Li, Welleck, and Yang, 2025).\n\n[^4]Furthermore, one can envision other contractible token categories, either by refining the distinctions above (e.g., separating reasoning tokens from response tokens) or by introducing new modalities and resources (e.g., media files, external databases, etc.). Accordingly, our model accommodates an arbitrary number of inference and fine-tuning token classes.\n\nCosts, $\\boldsymbol{c}_{\\boldsymbol{j}}, \\hat{\\boldsymbol{c}}_{\\boldsymbol{k}}$ Processing tokens is costly because it requires executing the neural network and, for fine-tuning, computing parameter updates, which in turn requires energy and specialized hardware. Because tokens within a given class are processed symmetrically from a computational perspective, we assume that the marginal cost per token within a class is constant across tasks. ${ }^{7}$ However, because different token classes correspond to different computational routines, we allow marginal costs to differ across classes.","text_sha256":"e2d5d3e879e05c3bb396202abeb48358390981632e334654232d73beef351072"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0009","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"3 Efficient Solution","text":"## 3 Efficient Solution\n\nIn this section, we begin by analyzing the socially efficient allocation of tokens to buyers and tasks. This solution provides a useful benchmark and also coincides with the optimal monopoly solution when the buyer's type is known. We then turn to the problem of a social planner who faces capacity constraints for each token class, and we use the constrainedefficient solution to characterize the buyer's utility from purchasing a fixed token budget.\n\nUnconstrained problem Given a type $w=\\left(w_{i}\\right)_{i \\in[0,1]}$, the efficient allocation solves the following problem:\n\n$$\n\\max _{\\left(x_{i}\\right)_{i \\in[0,1],}, z \\geq 0} \\int_{0}^{1} w_{i} \\Psi\\left(x_{i}\\right) \\Phi(z) d i-\\sum_{j=1}^{J} c_{j} \\int_{0}^{1} x_{i j} d i-\\sum_{k=1}^{K} \\hat{c}_{k} z_{k}\n$$\n\nConsider the allocation of inference tokens first. A key implication of the homogeneity of the production function $\\Psi$ is that inference tokens are optimally employed in the same proportions: the efficient allocation of tokens across tasks differs solely by the scale at which these inputs are employed. To see this, consider the system of $J$ first-order conditions for each $x_{i}=\\left(x_{i 1}, \\ldots, x_{i J}\\right)$ :\n\n$$\nw_{i} \\nabla \\Psi\\left(x_{i}\\right) \\Phi(z)=c .\n$$\n\nBy equation (6), for all $i, \\nabla \\Psi\\left(x_{i}\\right)$ belongs to a ray with a direction $c$. Because $\\Psi$ is homogeneous and strictly concave, the ray in the space of gradients corresponds to a ray in the\n\n[^5]space of tokens. Thus, for each $i$, any optimal $x_{i}$ can be written as:\n$$\nx_{i}=r_{i} d,\n$$\nwhere $d \\in \\mathbb{R}_{+}^{J}$ is the unique vector that solves $\\nabla \\Psi(d)=c$, i.e., the cost-minimizing input shares. Furthermore, because $\\Psi$ is homogeneous of degree $\\sigma, \\Psi_{j}$ is homogeneous of degree $\\sigma-1$, so that the optimal scale is given by\n$$\nr_{i}=w_{i}^{\\frac{1}{1-\\sigma}} \\Phi(z)^{\\frac{1}{1-\\sigma}} .\n$$\nSubstituting back in the objective function, both the total surplus and the total consumption of each token class depend on the buyer's type $w$ only through $\\int_{0}^{1} w_{i}^{1 /(1-\\sigma)} d i$. It is then convenient to define the aggregate type $\\theta(w)$ as\n$$\n\\theta(w) \\triangleq\\left(\\int_{0}^{1} w_{i}^{\\frac{1}{1-\\sigma}} d i\\right)^{1-\\sigma}\n$$\n\nOur first result establishes that the total surplus and the total amount of inference tokens depend only on the aggregate type and not on the finer details of the type profile $w$.\n\nProposition 1 (Efficient Allocation). Under the efficient allocation, all buyer types $w$ with the same aggregate type $\\theta(w)$ consume the same number of fine-tuning tokens, consume the same total number of inference tokens in each class, and obtain the same total payoff. The number of inference tokens allocated to task $i$ is proportional to $w_{i}^{\\frac{1}{1-\\sigma}}$.\n\nProposition 1 has a key implication: two buyers with very different task profiles-one who values a few tasks intensely and many tasks little, and another who values all tasks moderately-consume the same total resources and obtain the same total surplus, provided they share the same aggregate type. The seller therefore cannot distinguish between them on the basis of total token consumption, nor would it want to. This indistinguishability is not an assumption but a consequence of the technology: homogeneity of $\\Psi$ ensures that the efficient mix of token classes is task-independent, so only the scale of usage varies, and the aggregator (7) is the natural summary of how much total scale a buyer demands.\n\nWe now show that a similar logic can be used to characterize the socially efficient solution under capacity constraints in each class of tokens (training, computation, input, output).\n\nConstrained problem We now introduce capacity constraints on each token class. This detour is not merely a generalization: the special case in which all constraints bind is precisely the problem faced by a buyer who has purchased a fixed bundle of token budgets. Thus,\nthe constrained-efficient allocation provides the foundation for the seller's mechanism design problem in Section 4.\n\nFix a type $w$ and capacity constraints $X_{j}>0, Z_{k}>0$ for all $j, k$. The capacityconstrained planner's problem is\n\n$$\n\\begin{aligned}\n& \\max _{\\left(x_{i}\\right)_{i \\in[0,1], z \\geq 0}} \\int_{0}^{1} w_{i} \\Psi\\left(x_{i}\\right) \\Phi(z) d i-\\sum_{j=1}^{J} c_{j} \\int_{0}^{1} x_{i j} d i-\\sum_{k=1}^{K} \\hat{c}_{k} z_{k} \\\\\n& \\text { s.t. } \\quad \\int_{0}^{1} x_{i j} d i \\leq X_{j}, z_{k} \\leq Z_{k} \\text { for } j \\in[J], k \\in[K]\n\\end{aligned}\n$$\n\nOur next result establishes that the capacity-constrained efficient allocation coincides with the (unconstrained) efficient allocation for marginal costs inflated by the shadow costs of the capacity constraints, $\\left(c^{\\prime}, \\hat{c}^{\\prime}\\right) \\geq(c, \\hat{c})$. This observation yields an immediate implementation of the efficient solution.","text_sha256":"26a3011c3cc9d3b57ea538fb5287fdfa90789852abd65897c9abdf59161a3a5f"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0010","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"Corollary 1 (Constrained Efficient Allocation).","text":"## Corollary 1 (Constrained Efficient Allocation).\n\n1. In the capacity-constrained efficient allocation, all buyer types $w$ with the same aggregate type $\\theta(w)$ consume the same total number of fine-tuning tokens and inference tokens in each class, and they obtain the same total payoff.\n2. The capacity-constrained efficient allocation can be implemented via linear prices equal to inflated marginal costs.\n\nThus, capacity constraints do not change the qualitative properties of the solution, but lead to an inefficiency in the relative allocation of token types, and in the split between those buyer types who choose to fine-tune $(z>0)$ and those who do not $(z=0)$.\n\nThe special case where all capacity constraints bind is of particular interest, because it is instrumental in the characterization of the buyer's demand for token budgets in Section 4. When all constraints bind, the total production cost is pinned down. The planner's and the buyer's solutions then coincide, i.e., the constrained efficient allocation solves the problem of a buyer who has access to inference and fine-tuning token budgets $(X, Z)$.\n\nThe optimal allocation of a fixed budget $(X, Z)$ solves the following problem:\n\n$$\n\\max _{x_{i j} \\geq 0} \\int_{0}^{1} w_{i} \\Psi\\left(x_{i}\\right) \\Phi(Z) d i, \\quad \\text { s.t. } \\int_{0}^{1} x_{i j} d i=X_{j} \\text { for } j \\in[J] .\n$$\n\nApplying the homogeneity argument again, $\\Phi(Z)$ factors out, and the buyer optimally sets $x_{i}=r_{i} d$ with $r_{i} \\propto w_{i}^{1 /(1-\\sigma)}$. Since the budget constraint for each token class $j$ requires $d_{j} \\int_{0}^{1} r_{i} d i=X_{j}$, the common direction is $d=X / \\int_{0}^{1} r_{i} d i$, which yields a simple expression for the buyer's optimal payoff.\n\nThe following result is the key step in our analysis. It shows that the buyer's indirect utility from any token bundle is multiplicatively separable in the aggregate type and an aggregate quality index, placing the seller's problem squarely in the Mussa-Rosen framework.\n\nProposition 2 (Buyer Indirect Utility). For any inference and fine-tuning token budgets $X=\\left(X_{1}, \\ldots, X_{J}\\right) \\geq 0$ and $Z=\\left(Z_{1}, \\ldots, Z_{K}\\right) \\geq 0$, the indirect utility of buyer type $w$ is\n\n$$\nU(w, X, Z)=\\theta(w) \\Psi(X) \\Phi(Z),\n$$\n\nwhere $\\theta(w)$ is the aggregate type defined in (7).\nWe therefore obtain a tractable expression for the buyer's payoff by means of a representative task, whose value is given by the aggregate type $\\theta$, to which the buyer assigns the entire token budget. Heterogeneity across tasks washes out, and only the aggregate matters. A notable implication is that the buyer's unobserved allocation of tokens across tasks generates no additional incentive constraints beyond those of the standard screening problem. Under homogeneity, the cost-minimizing input mix is independent of scale, so the buyer's within-bundle allocation is pinned down regardless of type. This product form is the key input to the seller's problem, which we turn to next.","text_sha256":"0413eb64a512495e92a699af66e5a7020bc1b5aa6acb666f9175cebd9405932d"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0011","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"4 Menus of Token Budgets","text":"## 4 Menus of Token Budgets\n\nThe main result of this section is that the seller's optimal mechanism takes the form of a menu of committed-spend contracts: each buyer pays an upfront fee for a spending budget that she allocates freely across token classes priced at the provider's marginal cost. Higher types purchase larger budgets at higher fees, with quantity discounts that are consistent with observed industry pricing. The formal analysis proceeds in two steps: we first characterize the optimal direct mechanism (Section 4.1) and then show that it admits intuitive indirect implementations (Section 4.2).\n\nWe now characterize the provider's profit-maximizing menu of token budgets. By Proposition 2, the buyer's indirect utility from a bundle $(X, Z)$ is $\\theta(w) \\Psi(X) \\Phi(Z)$. Two features of this expression are worth emphasizing. First, private information enters only through the scalar $\\theta$ : the seller's multidimensional screening problem reduces to a one-dimensional problem. Second, the payoff is multiplicatively separable in type and allocation, placing us in the framework of Mussa and Rosen (1978).","text_sha256":"a20d3b6f34c62596475979b6557851a77cbe06d75ca144cc54bf5173898b0fe3"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0012","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"4.1 Optimal Budget Mechanism","text":"### 4.1 Optimal Budget Mechanism\n\nDefine the effective type as $\\theta$ and the effective allocation as the aggregate quality given by\n\n$$\nQ(X, Z) \\triangleq \\Psi(X) \\Phi(Z) .\n$$\n\nThe effective production costs of producing the aggregate quality $Q$ can be found via a cost-minimization problem:\n\n$$\nC(Q) \\triangleq \\min _{X, Z \\geq 0} \\sum_{j=1}^{J} c_{j} X_{j}+\\sum_{k=1}^{K} \\hat{c}_{k} Z_{k}, \\quad \\text { s.t. } \\Psi(X) \\Phi(Z)=Q .\n$$\n\nThe solution to the cost-minimization problem has the following properties.\nLemma 1 (Cost Function). The cost function $C(Q)$ is strictly increasing and strictly convex and satisfies $C_{+}^{\\prime}(0)=0 .{ }^{8}$\n\nConvexity follows from the strict concavity of the gain function $g$ : producing higher aggregate quality requires disproportionately more tokens. The property $C_{+}^{\\prime}(0)=0$ follows from the Inada conditions. Economically, $C_{+}^{\\prime}(0)=0$ means that the first unit of aggregate quality is arbitrarily cheap to produce. Combined with convexity, this implies that it is never efficient to exclude any buyer type entirely. Thus, in the monopoly problem, exclusion is driven by information rents rather than production costs.\n\nBy Lemma 1, the analysis of Mussa and Rosen (1978) applies. Denote the prior distribution of $\\theta$ by $F$, which is derived from $F_{w}$. The seller chooses an aggregate quality schedule $Q(\\theta)$ and a transfer schedule $t(\\theta)$ to solve\n\n$$\n\\max _{Q(\\cdot), t(\\cdot)} \\mathbb{E}_{F}[t(\\theta)-C(Q(\\theta))],\n$$\n\nsubject to incentive compatibility and individual rationality.\nAs usual, let\n\n$$\n\\varphi(\\theta) \\triangleq \\theta-\\frac{1-F(\\theta)}{f(\\theta)},\n$$\n\ndenote the virtual (aggregate) type. If the virtual value $\\varphi$ is not increasing, let $\\bar{\\varphi}$ denote its Myerson-ironed version. Then, all $\\theta$ with $\\bar{\\varphi}(\\theta) \\leq 0$ are excluded. For all other $\\theta$, the optimal allocation is uniquely pinned down by:\n\n$$\n\\bar{\\varphi}(\\theta)=C^{\\prime}(Q(\\theta)) .\n$$\n\n[^6]Denoting the corresponding optimal transfers by\n\n$$\nt(\\theta)=\\theta Q(\\theta)-\\int_{0}^{\\theta} Q(r) d r\n$$\n\nwe can then fully characterize the optimal mechanism.\nProposition 3 (Optimal Menu). The optimal menu of token budgets is $\\{(X(\\theta), Z(\\theta), t(\\theta))\\}_{\\theta}$, where $(X(\\theta), Z(\\theta))$ are the cost-minimizing token budgets that deliver quality $Q=Q(\\theta)$ at prices $t(\\theta)$ as defined in (10)-(13).\n\nThe optimal menu offers a continuum of plans. All types $w$ with the same aggregate type $\\theta(w)$ choose the same item. Low-type buyers are excluded. All served buyers receive distorted-downward quality (fewer tokens than the efficient allocation would prescribe) except at the very top. Buyers with higher aggregate types purchase strictly larger token budgets and pay strictly higher transfers, but enjoy quantity discounts: the average price per unit of quality falls with $\\theta$.","text_sha256":"bb136f72afb8fac5fe433a1d95dec83052d39c97be0b1811e8a3d01d9aba695e"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0013","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"4.2 Indirect Mechanisms","text":"### 4.2 Indirect Mechanisms\n\nThe optimal direct mechanism specifies token quantities for each type. In practice, providers do not announce menus of token quantities; they post prices and spending limits. We now show that three natural pricing formats implement the optimum.\n\nWe first consider committed-spend mechanisms, where the terms of the contract with a buyer are related to a maximum or minimum total expenditure.\n\nDefinition 1 (Maximum-Spend Mechanism). A maximum-spend mechanism is a collection of token prices $p$ and a menu of monetary budgets and transfers $\\left\\{\\left(B_{n}, T_{n}\\right)\\right\\}_{n}$ such that, upon selecting item $n$, the buyer pays $T_{n}$ for access to budget $B_{n}$, which he can freely spend on tokens priced at $p$.\n\nWhile the optimal budget (direct) mechanism is a cap on quantities, the maximumspend mechanism limits quality by capping expenditures at fixed prices. Those expenditures are virtual in that they do not contribute to payment beyond $T_{n}$, but they govern token allocation. As we will see in Section 6, the maximum-spend mechanism is used in practice, for example, by Quora's Poe (see Section 6), and it is closely related to the Cost-Based tariffs (where $p=c$ ) introduced in Armstrong (1996). Indeed, Proposition 8 in Appendix B derives analogous conditions to those in Armstrong (1996) under which a maximum-spend mechanism is optimal across all mechanisms, including those that specify a task-by-task token allocation.\n\nAn alternative implementation does not impose a cap on spending, but exposes the buyer to variable marginal prices, much like committed-spend mechanisms in cloud computing.\n\nDefinition 2 (Minimum-Spend Mechanism). A minimum-spend mechanism is a menu of monetary budgets and transfers $\\left\\{\\left(p_{n}, T_{n}\\right)\\right\\}_{n}$ such that, upon selecting item $n$, the buyer commits to spending at least $T_{n}$ on tokens priced at $p_{n}$.\n\nA minimum-spend mechanism allows buyers to commit to larger budgets (i.e., total expenditures) to unlock variable-price discounts. In contrast to a maximum-spend mechanism, payment comes from actual token consumption rather than from an upfront payment.\n\nDefinition 3 (Two-Part-Tariff Mechanism). A two-part-tariff mechanism is a menu of prices and transfers $\\left\\{\\left(p_{n}, T_{n}\\right)\\right\\}_{n}$ such that, upon selecting item $n$, the buyer pays $T_{n}$ and can buy any number of tokens priced at $p_{n}$.\n\nA two-part tariff is a classical mechanism that combines upfront and consumption payments but does not impose any token consumption limits.\n\nDefine a type-dependent markup $m(\\theta)$ as:\n\n$$\nm(\\theta) \\triangleq \\frac{\\theta}{\\bar{\\varphi}(\\theta)} .\n$$\n\nThe markup $m(\\theta)$ is the ratio of the true type to the virtual type; it captures the informationrent wedge that the seller imposes. Higher markups on lower types reflect greater information rents extracted from higher types in the menu.\n\nProposition 4 (Indirect Implementation). An optimal menu of token budgets can be implemented via a maximum-spend mechanism. If $m(\\theta)$ is decreasing (e.g., if $F$ has an increasing hazard rate), then an optimal menu of token budgets can be implemented via a minimum-spend mechanism and a two-part tariff mechanism.\n\nImplementability by a maximum-spend mechanism is straightforward and relies on the fact that the optimal token allocation is constrained-efficient, as in Doligalski et al. (2025): it suffices to set prices equal to marginal costs and to choose token expenditures and transfers so as to mimic those under the direct mechanism.\n\nImplementability by minimum-spend and two-part tariff mechanisms requires more care. Because $\\Psi$ is homogeneous, the cost-minimizing input mix is the same for every quality level $Q$; only the scale changes. A two-part tariff that preserves this input mix must therefore set token prices proportional to marginal costs; any other price vector would distort the buyer's allocation away from cost minimization. The proportionality factor is the markup $m(\\theta)=$\n$\\theta / \\bar{\\varphi}(\\theta)$, which inflates all marginal costs equally. The condition that $m(\\theta)$ is decreasing ensures that higher types, who select larger budgets, face lower per-token prices, so the menu is incentive-compatible. In the minimum-spend implementation, the markup equals the ratio between total revenue and total cost for type $\\theta$. Section 6 shows that each of these three implementations corresponds to an observed pricing format at leading platforms.","text_sha256":"fdadc8019529d271470c61cd7e6fc19f2b9c25fd589a41cec6048129ef6697b1"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0014","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"5 Multiple Models and Competition","text":"## 5 Multiple Models and Competition\n\nIn practice, every major LLM provider offers multiple models that differ in capability and cost (e.g., Anthropic's Haiku, Sonnet, and Opus, or OpenAI's GPT-4o-mini and o1). A buyer can assign different tasks to different models, and the provider can screen on both usage quantity and model access. This section extends our framework to this richer environment. Two new questions arise: How does a buyer optimally allocate tasks across models? And how does a profit-maximizing provider design menus that screen on both dimensions?\n\nSpecifically, we extend our setting and allow there to be $L$ models, each with a modelspecific gain function:\n\n$$\ng_{l}\\left(x_{i}, z\\right)=\\Psi_{l}\\left(x_{i}\\right) \\Phi_{l}(z),\n$$\n\nwhere $\\Psi_{l}$ and $\\Phi_{l}$ are assumed to have the same properties as in the baseline model, with $\\Psi_{l}$ being homogeneous of degree $\\sigma_{l}$. We order the models so that $\\sigma_{1} \\leq \\sigma_{2} \\leq \\cdots \\leq \\sigma_{L}$. The token costs are model-specific $\\left(c_{l}, \\hat{c}_{l}\\right)$.\n\nThe buyer can process different tasks with different models; however, he cannot use two models for the same task. ${ }^{9}$ If the buyer of type $w$ processes tasks $i \\in I_{l}$ with model $l$, his total payoff ignoring the payment is:\n\n$$\n\\sum_{l=1}^{L} \\int_{I_{l}} w_{i} g_{l}\\left(x_{l i}, z_{l}\\right) d i\n$$\n\nIn this setting, a bundle specifies token budgets for all available models $\\left(X_{l}, Z_{l}\\right)_{l=1}^{L}=$ $\\left(X_{l 1}, \\ldots, X_{l J}, Z_{l 1}, \\ldots, Z_{l K}\\right)_{l=1}^{L}$.\n\nAs a preliminary step for the subsequent analysis, we show that the buyer's value for a bundle $\\left(X_{l}, Z_{l}\\right)_{l=1}^{L}$ depends only on the aggregate quality of each model,\n\n$$\nQ_{l} \\triangleq g_{l}\\left(X_{l}, Z_{l}\\right),\n$$\n\n[^7]and admits a tractable closed-form expression. To this end, we fix $w$ and order the tasks so that $w_{i}$ is increasing. We assume that $Q_{l}>0$ for all $l$ (if $Q_{l}=0$ then the $l$-th model can be ignored). If the buyer allocates tasks in the set $I_{l}$ to model $l$, then his payoff from model $l$ is, by the arguments of Proposition 2,\n$$\nU_{l}\\left(I_{l}, Q_{l}\\right)=\\theta_{l}\\left(I_{l}\\right) Q_{l} \\text {, }\n$$\nwhere $\\theta_{l}\\left(I_{l}\\right) \\triangleq\\left(\\int_{I_{l}} w_{i}^{1 /\\left(1-\\sigma_{l}\\right)} d i\\right)^{1-\\sigma_{l}}$. The total payoff from the task-model allocation $\\left(I_{l}\\right)_{l=1}^{L}$ is\n$$\n\\sum_{l=1}^{L} U_{l}\\left(I_{l}, Q_{l}\\right)=\\sum_{l=1}^{L} \\theta_{l}\\left(I_{l}\\right) Q_{l} .\n$$\nThe optimal payoff $U^{*}\\left(Q_{1}, \\ldots, Q_{L}\\right)=\\sum_{l=1}^{L} U_{l}\\left(I_{l}^{*}, Q_{l}\\right)$ is evaluated at the optimal task-model allocation $\\left(I_{l}^{*}\\right)_{l=1}^{L}$.\n\nLemma 2 (Buyer-Optimal Payoff). If $\\sigma_{l}=\\sigma$ for all $l \\in[L]$, then\n\n$$\nU^{*}\\left(Q_{1}, \\ldots, Q_{L}\\right)=\\theta\\left(\\sum_{l=1}^{L} Q_{l}^{1 / \\sigma}\\right)^{\\sigma} .\n$$\n\nBy Lemma 2, despite the combinatorial complexity of assigning a continuum of tasks to multiple models, the buyer's payoff from a bundle of tokens across different models admits a simple product decomposition into an aggregate type and a CES-style aggregate quality over models. The elasticity of substitution across models is pinned down by the common returns-to-scale parameter $\\sigma$, and the single-model tractability of Sections 3-4 extends to the multi-model environment.","text_sha256":"9187da6f3517f2841c9a83712414f60910f92f474fec881d040b09a47087e9a6"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0015","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"5.1 Efficient Solution","text":"### 5.1 Efficient Solution\n\nWe first compute the efficient allocation in this setting. Since the buyer's value takes a simple form (20) as a function of aggregate qualities of each model, the efficient allocation delivers the efficient amount of these qualities in a cost-efficient way.\n\nAssumption 1 (Homogeneous Fine-Tuning). For each $l=1, \\ldots, L$ :\n\n1. $\\Psi_{l}$ is homogeneous of degree $\\sigma \\in(0,1)$.\n2. $\\Phi_{l}$ is homogeneous of degree $\\hat{\\sigma}_{l}<1-\\sigma$.\n\nAll subsequent analysis holds under Assumption 1, so we omit it from the statements. Assumption 1 allows us to obtain closed-form expressions because of the following result:\n\nLemma 3 (Model-Specific Cost Function). The cost function of delivering aggregate quality $Q_{l}$ through model $l$ is $C_{l}\\left(Q_{l}\\right)=\\kappa_{l} Q_{l}^{1 /\\left(\\sigma+\\hat{\\sigma}_{l}\\right)}$ for some constant $\\kappa_{l}>0$.\n\nBy Lemma 3, the minimal cost of obtaining quality $Q$ from $L$ models is\n\n$$\nC(Q)=\\min _{Q_{1}, \\ldots, Q_{L} \\geq 0} \\sum_{l=1}^{L} \\kappa_{l} Q_{l}^{1 /\\left(\\sigma+\\hat{\\sigma}_{l}\\right)}, \\quad \\text { s.t. } \\sum_{l=1}^{L} Q_{l}^{1 / \\sigma}=Q^{1 / \\sigma} .\n$$\n\nThe optimal solution utilizes a single model, resulting in\n\n$$\nC(Q)=\\min _{l \\in[L]} \\kappa_{l} Q^{1 /\\left(\\sigma+\\hat{\\sigma}_{l}\\right)} .\n$$\n\nThus, the rich functional model heterogeneity (15) reduces to heterogeneity in two parameters $\\left(\\kappa_{l}, \\hat{\\sigma}_{l}\\right)$ that roughly correspond to \"cost-effectiveness\" and \"propensity to fine-tune.\" Heterogeneity in $\\kappa_{l}$ is vertical: higher $\\kappa_{l}$ correspond to overall costlier models. Heterogeneity in $\\hat{\\sigma}_{l}$ is horizontal: models with high $\\hat{\\sigma}_{l}$ are better at generating high $Q$ whereas models with low $\\hat{\\sigma}_{l}$ are better at generating low $Q$. As a result, higher values of $Q$ are optimally produced by models with higher $\\hat{\\sigma}_{l}$. In other words, models with higher returns to fine-tuning have flatter cost curves and are cheaper for producing high qualities, introducing a natural form of horizontal differentiation among models.\n\nNotably, the cost function $C(Q)$ is not convex but piecewise convex, with kinks at the model switches, which translates into unusual properties of the efficient and profitmaximizing allocations. Specifically, the efficient quality allocation for type $\\theta$ when using model $l$ solves:\n\n$$\n\\max _{Q \\geq 0} \\theta Q-\\kappa_{l} Q^{1 /\\left(\\sigma+\\hat{\\sigma}_{l}\\right)},\n$$\n\nleading to the following characterization:\nProposition 5 (Efficient Allocation). In the efficient allocation, each type $\\theta$ processes all tasks with a single model:\n\n$$\nl^{*} \\in \\arg \\max _{l}\\left(1-\\sigma-\\hat{\\sigma}_{l}\\right)\\left(\\frac{\\left(\\sigma+\\hat{\\sigma}_{l}\\right)}{\\kappa_{l}}\\right)^{\\left(\\sigma+\\hat{\\sigma}_{l}\\right) /\\left(1-\\sigma-\\hat{\\sigma}_{l}\\right)} \\theta^{1 /\\left(1-\\sigma-\\hat{\\sigma}_{l}\\right)},\n$$\n\nwith the aggregate quality:\n\n$$\nQ^{*}(\\theta)=\\left(\\frac{\\theta\\left(\\sigma+\\hat{\\sigma}_{l^{*}}\\right)}{\\kappa_{l^{*}}}\\right)^{\\left(\\sigma+\\hat{\\sigma}_{l^{*}}\\right) /\\left(1-\\sigma-\\hat{\\sigma}_{l^{*}}\\right)}\n$$\n\nThe efficient token allocation is given by Lemma 3 evaluated at $Q^{*}$ from (23).\n\nEvery type uses a single model. Higher aggregate types $\\theta$ consume higher aggregate qualities by using models with higher fine-tuning capabilities $\\hat{\\sigma}_{l}$, i.e., $Q^{*}(\\theta)$ is increasing.\n\nBecause $C(Q)$ is not continuously differentiable, the allocation is discontinuous at the model switch. Indeed, it follows from (23) and (22) that if the efficient allocation at type $\\theta$ is indifferent between $Q_{l}$ of model $l$ and $Q_{l^{\\prime}}$ of model $l^{\\prime}$ such that $\\hat{\\sigma}_{l^{\\prime}}>\\hat{\\sigma}_{l}$, then\n\n$$\n\\frac{Q_{l^{\\prime}}}{Q_{l}}=\\frac{1-\\sigma-\\hat{\\sigma}_{l}}{1-\\sigma-\\hat{\\sigma}_{l^{\\prime}}}>1 .\n$$","text_sha256":"5132c214070fd795c7ebb5b6c63c4594c70b9899fd409131f06d2e2832958eec"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0016","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"5.2 Multi-Model Monopolist","text":"### 5.2 Multi-Model Monopolist\n\nWe now assume that all $L$ models are sold by a monopolist who offers a menu. Each item specifies a transfer and a collection of token budgets for different models $\\left(X_{l}, Z_{l}\\right)_{l=1}^{L}$.\n\nBy Lemma 2, the seller's problem is equivalent to the optimal screening problem in Mussa and Rosen (1978), in which the buyer's type is $\\theta \\in \\mathbb{R}$, the seller chooses a menu of $(Q, t)$, the buyer's payoff is $U(\\theta, Q)=\\theta Q$, and the cost of obtaining a quality $Q$ is equal to $C(Q)$ as in (21). Although $C(Q)$ is not convex, the Envelope theorem applies and this problem is equivalent to a maximization of virtual surplus:\n\n$$\n\\max _{Q(\\cdot): \\text { increasing }} \\int_{0}^{1}(\\varphi(\\theta) Q(\\theta)-C(Q(\\theta))) d F(\\theta)\n$$\n\nwhere $\\varphi(\\theta)=\\theta-(1-F(\\theta)) / f(\\theta)$.\nAssuming $\\varphi(\\theta)$ is increasing, the solution can be obtained pointwise to equal an efficient allocation with respect to the virtual type: $Q^{\\mathrm{m}}(\\theta)=Q^{*}(\\varphi(\\theta))$. By Proposition 5, if $\\varphi(\\theta)<0$, then the buyer is excluded, $Q^{\\mathrm{m}}=0$; if $\\varphi(\\theta) \\geq 0$,\n\n$$\n\\begin{aligned}\n& l^{\\mathrm{m}}(\\theta) \\in \\arg \\max _{l}\\left(1-\\sigma-\\hat{\\sigma}_{l}\\right)\\left(\\frac{\\left(\\sigma+\\hat{\\sigma}_{l}\\right)}{\\kappa_{l}}\\right)^{\\left(\\sigma+\\hat{\\sigma}_{l}\\right) /\\left(1-\\sigma-\\hat{\\sigma}_{l}\\right)} \\varphi(\\theta)^{1 /\\left(1-\\sigma-\\hat{\\sigma}_{l}\\right)}, \\\\\n& Q^{\\mathrm{m}}(\\theta)=\\left(\\frac{\\varphi(\\theta)\\left(\\sigma+\\hat{\\sigma}_{l^{\\mathrm{m}}}\\right)}{\\kappa_{l^{\\mathrm{m}}}}\\right)^{\\left(\\sigma+\\hat{\\sigma}_{l^{\\mathrm{m}}}\\right) /\\left(1-\\sigma-\\hat{\\sigma}_{l^{\\mathrm{m}}}\\right)} .\n\\end{aligned}\n$$\n\nSince $Q^{*}$ is increasing, the pointwise solution solves the problem. Optimal transfers are\n\n$$\nt^{\\mathrm{m}}(\\theta)=\\theta Q^{\\mathrm{m}}(\\theta)-\\int_{0}^{\\theta} Q^{\\mathrm{m}}(r) d r\n$$\n\nAs under the efficient allocation, because $C(Q)$ is not continuously differentiable, the quality and transfer schedules are discontinuous.\n\nProposition 6 (Multi-Model Monopoly). If $\\varphi(\\theta)$ is increasing, then the optimal token-\nbudget menu is given by $\\left\\{\\left(\\left(X_{l}(\\theta), Z_{l}(\\theta)\\right)_{l=1}^{L}, t(\\theta)\\right)\\right\\}_{\\theta}$, where $\\left(X_{l}(\\theta), Z_{l}(\\theta)\\right)_{l=1}^{L}$ are efficient token budgets that deliver quality $Q^{\\mathrm{m}}(\\theta)$ at prices $t(\\theta)$ as defined in (26) and (27).\n\nThe monopolist assigns each buyer type to a single model, with higher types using more capable models. All types $w$ with the same aggregate type $\\theta(w)$ choose the same item and use one model, $l^{\\mathrm{m}}(\\theta)$ as defined in (25), for all tasks. The quality schedule exhibits discrete jumps at model-switching thresholds: a buyer at the margin between two models jumps to a strictly higher quality when upgrading. OpenAI's tier structure, which reserves its most compute-intensive reasoning model for the highest-paying subscribers, provides a direct empirical counterpart; see the discussion in Section 6.","text_sha256":"f782d709ac6723ed96dc59ba74dbb2afc76086f612a8de3861c327ef5e8608c8"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0017","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"5.3 Leader-Fringe Competition","text":"### 5.3 Leader-Fringe Competition\n\nWe now study competition between a proprietary leader and an open-source competitive fringe. The key economic forces are: (i) the fringe provides an outside option whose value depends on the buyer's type; (ii) for intermediate types, the leader must provide enough tokens to deter fringe top-up; (iii) for high types, the leader acts as an unconstrained monopolist. The resulting allocation has three distinct regions, generating a richer pattern of distortions than either the single-model monopoly or the multi-model monopoly.\n\nSpecifically, we assume that there is a single leader, corresponding to the highest-capability model $(l=L)$ in the previous section, and a continuum of firms in a competitive fringe that sell their tokens at fixed per-token prices equal to marginal costs. We assume that the leader possesses a proprietary model characterized by the aggregate cost parameter $c_{L}>0$, returns to intensity $\\sigma_{L}=\\sigma>0$, and returns to fine-tuning $\\hat{\\sigma}_{L}>0$, such that $\\sigma+\\hat{\\sigma}_{L}<1$. The fringe possesses an open-source model characterized by the lower aggregate cost parameter $c_{F} \\in\\left[0, c_{L}\\right)$, the same returns to intensity $\\sigma_{F}=\\sigma$, and lower returns to fine-tuning $\\hat{\\sigma}_{F} \\in\\left[0, \\hat{\\sigma}_{L}\\right)$.\n\nThe buyer has a private type $w$. By Lemma 2, only the aggregate type $\\theta(w)$ matters, so we take $\\theta$ as a primitive and assume that $\\theta$ is distributed according to $F$ with strictly positive density $f$ everywhere on [0, 1]. The buyer can multi-home and can buy at most one item from the leader.\n\nWe want to solve the leader's problem, which can be viewed as a monopolistic design of an optimal menu $(q, t)$ subject to the competitive pressure from the fringe. For notational simplicity, define the quantities:\n\n$$\n\\begin{aligned}\n& q_{L} \\triangleq Q_{L}^{1 /\\left(\\sigma+\\hat{\\sigma}_{L}\\right)}=g_{L}\\left(X_{L 1}, \\ldots, X_{L J}, Z_{L 1}, \\ldots, Z_{L K}\\right)^{1 /\\left(\\sigma+\\hat{\\sigma}_{L}\\right)} \\\\\n& q_{F} \\triangleq Q_{F}^{1 /\\left(\\sigma+\\hat{\\sigma}_{F}\\right)}=g_{F}\\left(X_{F 1}, \\ldots, X_{F J}, Z_{F 1}, \\ldots, Z_{F K}\\right)^{1 /\\left(\\sigma+\\hat{\\sigma}_{F}\\right)}\n\\end{aligned}\n$$\n\nThe payoff of type $\\theta$ from having purchased quantity $q_{L}$ from the leader is\n\n$$\n\\max _{q_{F} \\geq 0} \\theta\\left(q_{L}^{\\left(\\sigma+\\hat{\\sigma}_{L}\\right) / \\sigma}+q_{F}^{\\left(\\sigma+\\hat{\\sigma}_{F}\\right) / \\sigma}\\right)^{\\sigma}-c_{F} q_{F},\n$$\n\nwith a (type-dependent) outside option corresponding to $q_{L}=0$. The first-order condition for the buyer's problem of purchasing fringe quantity is\n\n$$\n\\theta\\left(\\sigma+\\hat{\\sigma}_{F}\\right)\\left(q_{L}^{\\left(\\sigma+\\hat{\\sigma}_{L}\\right) / \\sigma}+q_{F}^{\\left(\\sigma+\\hat{\\sigma}_{F}\\right) / \\sigma}\\right)^{\\sigma-1} q_{F}^{\\hat{\\sigma}_{F} / \\sigma}=c_{F} .\n$$\n\nIn general, (29) does not admit a closed-form solution. It does so in the case of $\\hat{\\sigma}_{F}=0$. Specifically, define\n\n$$\n\\hat{q}(\\theta) \\triangleq\\left(\\frac{\\theta \\sigma}{c_{F}}\\right)^{\\sigma /\\left((1-\\sigma)\\left(\\sigma+\\hat{\\sigma}_{L}\\right)\\right)}, \\quad \\psi(\\theta) \\triangleq \\theta^{\\frac{1}{1-\\sigma}}(1-\\sigma)\\left(\\frac{\\sigma}{c_{F}}\\right)^{\\sigma /(1-\\sigma)} .\n$$\n\nIf the quantity purchased from the leader is $q_{L}<\\hat{q}(\\theta)$, the buyer will purchase fringe tokens to achieve the optimal total quantity $\\hat{q}(\\theta)$; if $q_{L}>\\hat{q}(\\theta)$, the buyer will single-home with the leader. The buyer's outside option utility corresponding to $q=0$ is $\\psi(\\theta)$. We then obtain the following characterization of the buyer's demand for the leader's quantity.\n\nLemma 4 (Buyer's Payoff). The leader's problem is equivalent to one in which the buyer has a zero outside option and the following payoff:\n\n$$\nu(\\theta, q)= \\begin{cases}c_{F} q^{\\left(\\sigma+\\hat{\\sigma}_{L}\\right) / \\sigma}, & \\text { if } q<\\hat{q}(\\theta), \\\\ \\theta q^{\\sigma+\\hat{\\sigma}_{L}}-\\psi(\\theta), & \\text { if } q \\geq \\hat{q}(\\theta) .\\end{cases}\n$$\n\nThe payoff $u(\\theta, q)$ in Lemma 4 is continuous, differentiable, and increasing in both parameters. In the region $q<\\hat{q}(\\theta)$, it is convex in $q$ and independent of $\\theta$. In the region $q>\\hat{q}(\\theta)$, it is concave in $q$ and supermodular in $q$ and $\\theta$.\n\nThe two-regime structure has a clear economic interpretation. In the first regime ( $q<$ $\\hat{q}(\\theta))$, the buyer supplements leader tokens with fringe tokens to achieve the optimal total quantity, so her marginal value of additional leader tokens is pinned down by the fringe price $c_{F}$; the leader faces a perfectly elastic residual demand. In the second regime $(q \\geq \\hat{q}(\\theta))$, the buyer single-homes with the leader, and her value is the standard concave payoff minus the outside option $\\psi(\\theta)$ she forgoes by not using the fringe. The transition between regimes is where the leader's competitive constraint binds.\n\nThe leader's problem is an instance of screening with a type-dependent outside option (Jullien, 2000); the reformulation (31) absorbs this outside option into the buyer's payoff.\n\nAn optimal menu can then be characterized by standard methods. ${ }^{10}$ Indeed, $u(\\theta, q)$ satisfies the (weak) single-crossing constraints: $u_{\\theta q}(\\theta, q)=0$ if $q<\\hat{q}(\\theta)$ and $u_{\\theta q}(\\theta, q)>0$ if $q \\geq \\hat{q}(\\theta)$. Allocation $q(\\theta)$ is implementable only if it is increasing whenever $q \\geq \\hat{q}(\\theta)$, and the profit associated with allocation $q$ is:\n\n$$\n\\Pi(q)=\\int_{0}^{1}\\left(u(\\theta, q(\\theta))-c_{L} q(\\theta)-\\frac{1-F(\\theta)}{f(\\theta)} u_{\\theta}(\\theta, q(\\theta))\\right) f(\\theta) d \\theta\n$$\n\nIn regular environments, this profit can be maximized pointwise. To this end, denote by $q^{\\mathrm{m}}(\\theta)$ the optimal monopoly quantity in the absence of the competitive fringe\n\n$$\nq^{\\mathrm{m}}(\\theta)=\\left(\\frac{\\left(\\sigma+\\hat{\\sigma}_{L}\\right) \\varphi(\\theta)}{c_{L}}\\right)^{1 /\\left(1-\\sigma-\\hat{\\sigma}_{L}\\right)} .\n$$","text_sha256":"0c1512f6a296f0299f5170badce1c78d2a52c2056e274857a4c0a0643b6014f2"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0018","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"5.3 Leader-Fringe Competition","text":"Proposition 7 (Leader-Fringe Competition). Let $F$ satisfy a monotone hazard rate and $\\hat{\\sigma}_{F}=0$. Then, there exist $\\underline{\\theta}$ and $\\bar{\\theta}, \\underline{\\theta} \\leq \\bar{\\theta}$, such that (i) for $\\theta \\leq \\underline{\\theta}, q(\\theta)=0$, (ii) for $\\theta \\in(\\underline{\\theta}, \\bar{\\theta})$, $q(\\theta)=\\hat{q}(\\theta)$ given by (30), and (iii) for $\\theta \\geq \\bar{\\theta}, q(\\theta)=q^{\\mathrm{m}}(\\theta)$ given by (33).\n\nUnder the optimal mechanism, low types purchase exclusively from the fringe. All higher types purchase exclusively from the leader. Out of those, the midline types are on the margin of whether to buy from the fringe, whereas the highline types strictly prefer to purchase from the leader. Depending on parameters, the midline region may not exist, $\\underline{\\theta}=\\bar{\\theta}$, or the highline region may not exist, $\\bar{\\theta}>1$.\n\nThe three-region structure reflects three distinct economic regimes. In the fringe-only region $(\\theta \\leq \\underline{\\theta})$, the buyer's willingness to pay for the leader's superior fine-tuning capability is too low to justify adoption. In the deterrence region $(\\underline{\\theta}<\\theta<\\bar{\\theta})$, the leader offers exactly enough tokens to make the buyer indifferent between single-homing with the leader and supplementing with the fringe; this pins the leader's quantity to $\\hat{q}(\\theta)$, which rises with type. In the monopoly region $(\\theta \\geq \\bar{\\theta})$, the buyer's value is high enough that fringe competition no longer constrains the leader, who reverts to the standard downward-distorted monopoly quantity $q^{\\mathrm{m}}(\\theta)$. The novel region is the deterrence region, which has no analog in either the single-model monopoly or the multi-model monopoly.\n\n[^8]Comparison with Efficient Allocation It is instructive to compare the allocation under leader-fringe competition with the efficient benchmark. By Proposition 5, defining\n\n$$\nu_{l}^{*}(\\theta)=\\left(1-\\sigma_{l}-\\hat{\\sigma}_{l}\\right) \\theta^{1 /\\left(1-\\sigma_{l}-\\hat{\\sigma}_{l}\\right)}\\left(\\frac{\\sigma_{l}+\\hat{\\sigma}_{l}}{c_{l}}\\right)^{\\left(\\sigma_{l}+\\hat{\\sigma}_{l}\\right) /\\left(1-\\sigma_{l}-\\hat{\\sigma}_{l}\\right)}\n$$\n\ntype $\\theta$ uses the fringe model if and only if $u_{F}^{*}(\\theta)>u_{L}^{*}(\\theta)$. The indifference type at which $u_{F}^{*}(\\theta)=u_{L}^{*}(\\theta)$ is (recall that $\\sigma_{L}=\\sigma_{F}=\\sigma$ and $\\hat{\\sigma}_{F}=0$ )\n\n$$\n\\hat{\\theta}=\\left(\\frac{1-\\sigma}{1-\\sigma-\\hat{\\sigma}_{L}}\\right)^{\\left(1-\\sigma-\\hat{\\sigma}_{L}\\right)(1-\\sigma) / \\hat{\\sigma}_{L}}\\left(\\frac{\\sigma}{c_{F}}\\right)^{\\sigma\\left(1-\\sigma-\\hat{\\sigma}_{L}\\right) / \\hat{\\sigma}_{L}}\\left(\\frac{c_{L}}{\\sigma+\\hat{\\sigma}_{L}}\\right)^{\\left(\\sigma+\\hat{\\sigma}_{L}\\right)(1-\\sigma) / \\hat{\\sigma}_{L}} .\n$$\n\nIf $\\theta<\\hat{\\theta}$, then the type uses the fringe model: $q_{F}(\\theta)=\\left(\\theta \\sigma / c_{F}\\right)^{1 /(1-\\sigma)}$ and $q_{L}(\\theta)=0$. If $\\theta \\geq \\hat{\\theta}$, then the type uses the leader model: $q_{F}(\\theta)=0$ and $q_{L}(\\theta)=\\left(\\theta\\left(\\sigma+\\hat{\\sigma}_{L}\\right) / c_{L}\\right)^{1 /\\left(1-\\sigma-\\hat{\\sigma}_{L}\\right)}$.\n\nComparing this with the fringe-competition allocation characterized in Proposition 7, we see two main differences: First, we see that the \"highline\" leader quantity $q^{\\mathrm{m}}(\\theta)$ is distorted downward from efficiency because $\\varphi(\\theta)$ is smaller than $\\theta$. This is natural, since in that range the leader is unconstrained by the fringe and acts as a monopolist. Second, there is an additional extensive margin distortion, because $\\hat{\\theta}<\\underline{\\theta} .{ }^{11}$\n\nComparison with Multi-Model Monopoly Another natural benchmark is the case of a monopolist possessing the leader and fringe models. By Proposition 6, the monopolist effectively faces a Mussa-Rosen problem with the buyer's payoff:\n\n$$\nu\\left(\\theta, q_{L}, q_{F}\\right)=\\theta\\left(q_{F}+q_{L}^{\\left(\\sigma+\\hat{\\sigma}_{L}\\right) / \\sigma}\\right)^{\\sigma} .\n$$\n\nand will be supplying an efficient allocation with respect to the virtual type. If $\\theta<\\theta^{\\mathrm{m}}$, where $\\varphi\\left(\\theta^{\\mathrm{m}}\\right) \\triangleq \\hat{\\theta}$ and $\\hat{\\theta}$ is defined at (34), then the type uses the fringe model and $q_{F}(\\theta)=$ $\\left(\\varphi(\\theta) \\sigma / c_{F}\\right)^{1 /(1-\\sigma)}$ and $q_{L}(\\theta)=0$. If $\\theta \\geq \\theta^{\\mathrm{m}}$, then the type uses the leader model with $q_{F}(\\theta)=0$ and $q_{L}(\\theta)=\\left(\\varphi(\\theta)\\left(\\sigma+\\hat{\\sigma}_{L}\\right) / c_{L}\\right)^{1 /\\left(1-\\sigma-\\hat{\\sigma}_{L}\\right)}$.\n\nRelative to the efficiency benchmark, the allocation is distorted downward within each model, except at the very top. In addition, fewer types consume the leader model.\n\nRelative to the leader-fringe environment of Proposition 7, the integrated multi-model monopolist internalizes the fringe technology and therefore does not face a competitive outside option. This has two implications. First, the fringe allocation is no longer pinned down by marginal-cost pricing: types that use the fringe model receive $q_{F}(\\theta)=\\left(\\varphi(\\theta) \\sigma / c_{F}\\right)^{1 /(1-\\sigma)}$\n\n[^9](and are excluded whenever $\\varphi(\\theta)<0$ ), whereas under leader-fringe competition the corresponding types purchase the efficient quantity $\\left(\\theta \\sigma / c_{F}\\right)^{1 /(1-\\sigma)}$ from the fringe. Second, the \"midline\" region $\\theta \\in(\\underline{\\theta}, \\bar{\\theta})$ in which the leader supplies $q=\\hat{q}(\\theta)$ to deter top-up purchases from the fringe disappears. Because the monopolist controls both models, she uses the fringe model directly as the low-quality product in the Mussa-Rosen menu and assigns each type to a single model with a single cutoff $\\theta^{\\mathrm{m}}$. Finally, once the fringe constraint is slack, the intensive-margin provision coincides and equals the monopoly quantity $q^{\\mathrm{m}}(\\theta)$; differences between the two environments are therefore concentrated among low and intermediate types. ${ }^{12}$\n\nFigure 1 illustrates the optimal allocation and distortions in a uniform example. All calculations for this example are in Appendix A.","text_sha256":"e66cce1dcfb17aaa9f4dc06187ff79429f91557440dd8281c9e46c8097318c74"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0019","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"5.3 Leader-Fringe Competition","text":"The top panel plots the leader quantity $q(\\theta)$, and the bottom panel plots the fringe quantity $q_{F}(\\theta)$. Four cutoffs, ordered $\\hat{\\theta}<\\underline{\\theta}<\\theta^{\\mathrm{m}}<\\bar{\\theta}$, partition the type space. The efficient allocation (solid) features a single switch at $\\hat{\\theta}$ : types below $\\hat{\\theta}$ use only the fringe, types above use only the leader. The leader-fringe allocation (dotted) shifts this switch rightward to $\\underline{\\theta}$, reflecting the extensive-margin distortion. Between $\\underline{\\theta}$ and $\\bar{\\theta}$, the leader supplies the deterrence quantity $\\hat{q}(\\theta)$, which lies strictly below the efficient leader quantity. Above $\\bar{\\theta}$, the fringe constraint is slack and the leader reverts to the monopoly quantity $q^{\\mathrm{m}}(\\theta)$, which coincides with the integrated monopolist's allocation (dashed). In the bottom panel, the efficient and leader-fringe curves for $q_{F}$ coincide over their respective ranges, since the fringe prices at marginal cost in both cases; the difference is that fringe usage persists up to $\\underline{\\theta}$ rather than $\\hat{\\theta}$. The integrated monopoly distorts fringe provision downward and switches to the leader at the later cutoff $\\theta^{\\mathrm{m}}$.","text_sha256":"56ef4b156a982661f8f0ae6d10ca28a58c97a7405669bdb515afb1ffd080925b"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0020","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"6 LLM Pricing in Practice","text":"## 6 LLM Pricing in Practice\n\nA striking feature of current LLM pricing is that different market segments implement each of the optimal mechanisms in our theoretical framework: consumer subscriptions implement the nonlinear menus of Sections 4 and 5; model aggregators implement the committed-spend mechanisms of Section 4.2; and developer APIs implement the constrained-efficient linear pricing of Corollary 1.\n\nOur main findings are as follows. The two largest proprietary providers, Anthropic and OpenAI, offer remarkably similar price points (\\$20 and \\$200) but implement fundamentally different screening mechanisms. Anthropic screens through quantity alone, holding model ac-\n\n[^10]\n\n> [Figure omitted from this text-only corpus; refer to the source manuscript.]\nFigure 1: Leader-Fringe allocations across different regimes. Example with $\\theta \\sim U[0,1]$, $\\sigma=1 / 2, \\hat{\\sigma}_{L}=1 / 4, c_{F}=1 / 10, c_{L}=1 / 8$.\n\ncess constant across paid tiers, matching the token-budget mechanism of Proposition 3. OpenAI screens through both quantity and model access, reserving its most compute-intensive reasoning model for the highest tier, matching the multi-model menu of Proposition 6. This contrast illustrates the distinction between the single-model screening of Section 4 and the multi-model versioning of Section 5. Model aggregators, which resell access to upstream providers through a single interface, implement the committed-spend mechanisms of Section 4.2: Quora's Poe enforces a hard budget cap (maximum spend), while GitHub Copilot allows overage purchasing at linear prices (minimum spend). Finally, API pricing across all major providers is linear in tokens with no volume discounts, resembling the constrainedefficient benchmark of Corollary 1 rather than any profit-maximizing mechanism, consistent with providers prioritizing market share over rent extraction in the developer segment.","text_sha256":"20a653b55af1958e7c731cf12f583f4d28728d208a7204c8c5eae161cba8c948"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0021","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"6.1 Anthropic","text":"### 6.1 Anthropic\n\nAnthropic's consumer pricing for Claude best illustrates the token-budget mechanism of Section 4. Anthropic offers the same model family across all paid tiers, with differentiation occurring through usage allocations. As in the optimal menu characterized in Proposition 3, users with higher aggregate types $\\theta$ select items with larger token budgets while consuming the same underlying technology.\n\nFigure 2 summarizes the structure of paid tiers as of January 2026. All paid tiers, namely Pro (\\$20/month), Max 5x (\\$100/month), and Max 20x (\\$200/month), grant access to Haiku 4.5, Sonnet 4.5, and Opus 4.5. The key distinction is the quantity of compute: Anthropic measures usage limits in computational intensity. For example, Opus queries consume resources approximately $5 \\times$ faster than Sonnet. This aligns with Proposition 3, where the optimal mechanism defines users' budgets in total tokens consumed across classes. Although Anthropic offers multiple models, all paid tiers grant access to the same set, so the relevant screening dimension is quantity rather than model access. The different depletion rates across models (e.g., Opus consuming resources $5 \\times$ faster than Sonnet) function as heterogeneous per-unit costs within a single compute-budget mechanism, analogous to the token-class costs $c_{j}$ in Proposition 3.\n\n> [Figure omitted from this text-only corpus; refer to the source manuscript.]\nFigure 2: Anthropic subscription tiers (January 2026). Model access is constant across paid tiers; differentiation occurs through usage allocations.","text_sha256":"dcad30128d41736d275ff7d7913a75887ea1b98a92f1a15206142e1f3e728bff"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0022","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"6.2 OpenAI","text":"### 6.2 OpenAI\n\nUnlike Anthropic's tier structure, which differentiates mostly through constraints on usage, OpenAI's ChatGPT pricing differentiates through both quantity and exclusive model access: higher tiers grant access to more capable models and relax usage constraints. This joint screening aligns with the multi-model monopolist of Section 5.2.\n\nFigure 3 summarizes OpenAI's tier structure as of January 2026. The Free tier provides limited GPT-4o access with automatic downgrade to GPT-4o-mini when capacity is constrained. Plus (\\$20/month) expands to approximately 80 GPT-4o messages per 3-hour window and adds reasoning models (o3, o4-mini) with a separate weekly limit of roughly 100 queries. Pro (\\$200/month) removes most constraints and provides exclusive access to o1pro, OpenAI's most compute-intensive reasoning model. Thus, both Anthropic and OpenAI converge on similar price points (\\$20 standard, \\$200 premium), but implement very different screening strategies.\n\n> [Figure omitted from this text-only corpus; refer to the source manuscript.]\nFigure 3: OpenAI ChatGPT subscription tiers (January 2026). Higher tiers grant access to more capable models and larger usage allocations.\n\nIn Proposition 6, we showed that when models differ in cost curvature, the optimal mechanism assigns higher-capability models to higher buyer types. OpenAI's decision to reserve its o1-pro model for the highest tier is consistent with this prediction. In consumer subscriptions, the relevant cost differences across models may reflect inference-time compute intensity rather than user fine-tuning, but the qualitative implication of a monotone assignment with discrete jumps at upgrade thresholds is the same.","text_sha256":"63f1e6c4dbcff2c7d2bccf06a7eed7edf5523cd344e87f52f1a93e449983b885"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0023","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"6.3 Model Aggregators: GitHub and Quora","text":"### 6.3 Model Aggregators: GitHub and Quora\n\nA number of AI platforms, including GitHub Copilot and Quora's Poe, do not develop their own models. Instead, they aggregate models produced by others (OpenAI, Anthropic, Google, Mistral, and others) and sell access to users through a single interface. A common contractual arrangement is one whereby subscribers pay a monthly fee, receive a budget of platform-specific credits, and allocate those credits across the available models. Selecting a more capable model typically depletes the budget faster. We now describe Quora's and GitHub's platforms in detail and then relate their pricing structures to the indirect implementation of the optimal mechanism discussed in Section 4.2.\n\nQuora-Poe Poe offers subscribers access to over 100 AI models through a unified chat interface. A subscriber pays a fixed monthly fee and receives an allocation of \"compute points,\" which serve as the platform's internal currency. Each model is priced in points per message at a rate that reflects its inference cost. ${ }^{13}$\n\nTable 1 summarizes Poe's tier structure. The entry-level Basic plan provides 300,000 points per month for \\$5; the top-tier Premium plan provides 12,500,000 points for \\$250. Above the Basic tier, the effective price per million points is constant at \\$20, so that higher tiers simply offer a proportionally larger budget at a fixed unit rate. Table 2 reports point costs for selected models, illustrating the wide dispersion across model classes.\n\n| Tier | Monthly Price | Points/Month | Eff. \\$/M Points |\n| :--- | :--- | :--- | :--- |\n| Basic | \\$5 | 300,000 | \\$16.67 |\n| Standard | \\$20 | 1,000,000 | \\$20.00 |\n| Pro | \\$50 | 2,500,000 | \\$20.00 |\n| Advanced | \\$100 | 5,000,000 | \\$20.00 |\n| Premium | \\$250 | 12,500,000 | \\$20.00 |\n\nTable 1: Poe subscription tiers (2025).\n\nThe connection to the maximum-spend framework of Section 4.2 is immediate. The monthly subscription fee corresponds to the transfer $T_{n}$; the point allocation corresponds to the budget $B_{n}$; and the per-model point costs correspond to marginal-cost pricing of different token classes. The subscriber chooses how to allocate the budget taking the marginal prices into account. Finally, once the budget is exhausted, access to all models is suspended until the next billing cycle, and points do not roll over. Thus, Poe's menu is best captured by a maximum-spend mechanism.\n\n[^11]| Class | Model | Points/Message |\n| :--- | :--- | :--- |\n| Budget | GPT-4o-mini | 9 |\n|  | Gemini-2.0-Flash | 9 |\n| Mid-tier | GPT-4o | 224 |\n|  | Claude-3.5-Sonnet | 276 |\n| Frontier | Claude-3-Opus | 1,697 |\n|  | Claude-Opus-4 | 4,105 |\n\nTable 2: Poe per-model point costs (2025).\n\nGitHub Copilot GitHub Copilot provides AI-assisted coding within a developer's integrated environment. Like Poe, Copilot offers multiple pricing tiers, each bundling a monthly fee with a budget of \"premium requests\" that the developer allocates across model calls. Table 3 summarizes the plan structure.\n\n| Plan | Monthly Fee | Premium Requests | Notes |\n| :--- | :--- | :--- | :--- |\n| Free | \\$0 | 50 | Limited quota, basic completions |\n| Pro | \\$10 | 300 | Unlimited completions, chat, IDE support |\n| Pro+ | \\$39 | 1,500 | Larger quota, advanced models |\n| Business | \\$19/user | 300/user | Org controls, policy management |\n| Enterprise | \\$39/user | 1,000/user | Enterprise features |\n\nTable 3: GitHub Copilot pricing tiers and premium request budgets (2026).\n\nThe mechanism by which Copilot meters usage differs slightly from Poe's point system but serves the same function. Rather than assigning each model a point cost per message, Copilot assigns each model a multiplier that determines how many premium requests a single interaction consumes. A set of baseline models, including GPT-5 mini, GPT-4.1, and GPT-4o, carry a multiplier of 0 on paid plans: interactions with these models are included at no additional budget cost. Efficient reasoning models such as Claude Haiku 4.5 and o4-mini carry multipliers of 0.25-0.33. Standard models such as Claude Sonnet 4.x and GPT-5.x carry a multiplier of 1. At the frontier, Claude Opus 4.5 carries a multiplier of 3, and GPT-4.5 carries a multiplier of 50.\n\nThis mechanism matches the framework of Section 4.2 quite closely. The plan fee corresponds to $T_{n}$; the monthly premium-request allocation corresponds to $B_{n}$; and the model multipliers play the role of heterogeneous marginal costs. A distinctive feature of Copilot is that it extends this budget mechanism with an overage option. When a developer (i.e., a buyer) exhausts her monthly allocation, Copilot allows continued use at a fixed charge\nof \\$0.04 per premium request. The minimum-spend framework of Section 4.2 matches this feature rather well: each user commits to a certain amount of monthly spending, there are no refunds, but the developer can increase her consumption ex post at linear prices. ${ }^{14}$\n\nThus, both Poe and Copilot implement the core logic of the committed-spend mechanisms: a non-bankable budget, priced up front via a subscription fee, that the user freely allocates across inputs with heterogeneous per-unit costs. The key distinction lies in how each platform handles the overages. Poe enforces a hard cap: once points are exhausted, access is suspended until the next billing cycle. This corresponds exactly to the maximumspend mechanism. Copilot, by contrast, allows the cap to be relaxed at a known marginal price (which varies across models through the multiplier system), effectively implementing a minimum-spend mechanism.","text_sha256":"675bb18864ba64032bdc3bf6d5cf0b575dfcddf5fcca80b0b61ecdf3658a09b7"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0024","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"6.4 API Pricing","text":"### 6.4 API Pricing\n\nAlongside consumer subscriptions, every major LLM provider operates a developer-facing API through which applications can programmatically submit prompts and receive completions. A developer who calls the API pays per token, with separate rates for input tokens (the prompt) and output tokens (the model's response). There is no subscription fee, no bundled allocation, and no minimum commitment: the developer pays only for what she uses. Table 4 reports prices for selected models across three providers as of January 2025.\n\n| Provider | Model | Input (\\$/M) | Output (\\$/M) |\n| :--- | :--- | :--- | :--- |\n| OpenAI | GPT-4o-mini | 0.15 | 0.60 |\n| OpenAI | GPT-4o | 2.50 | 10.00 |\n| OpenAI | o1 (reasoning) | 15.00 | 60.00 |\n| Anthropic | Claude 3.5 Haiku | 0.80 | 4.00 |\n| Anthropic | Claude 3.5 Sonnet | 3.00 | 15.00 |\n| Anthropic | Claude 3 Opus | 15.00 | 75.00 |\n| Google | Gemini 1.5 Flash | 0.075 | 0.30 |\n| Google | Gemini 1.5 Pro | 1.25 | 5.00 |\n\nTable 4: API pricing across major providers (January 2025). Prices per million tokens.\n\nThe pricing structure is notably uniform. Output tokens are priced at $3-5 \\times$ input tokens across all providers, reflecting that output generation is sequential and autoregressive (each token requires a forward pass conditioned on all preceding tokens), whereas input encod-\n\n[^12]ing is parallelizable and therefore substantially cheaper per token. Pricing is strictly linear: there are no quantity discounts, volume commitments, or tiered rates in standard API access. Although outside our model, all providers offer batch discounts (typically 50\\%) for asynchronous processing and prompt-caching discounts (up to 90\\%) for repeated inputs.\n\nThis linear, pay-per-token structure stands in sharp contrast to the nonlinear menus that characterize consumer subscriptions. It corresponds instead to the constrained-efficient allocation of Corollary 1: when a provider faces capacity constraints but wishes to maximize total surplus rather than profit (for instance, to lock in consumers in anticipation of future monetization), the constrained-efficient allocation can be implemented via linear prices equal to marginal costs inflated by shadow costs. The absence of nonlinear screening in API pricing is consistent with providers currently prioritizing adoption over rent extraction.\n\nThe uniformity of the pricing format across providers is itself evidence of competitive pressure: if any single provider were to introduce nonlinear screening in the API segment, it would risk losing developers to competitors who maintain simpler, linear pricing.\n\nEvidence from pricing dynamics is consistent with this interpretation. First, API prices have declined rapidly: GPT-4-class capability fell from \\$30/\\$60 per million tokens at launch (March 2023) to \\$2.50/\\$10 by August 2024, a 90\\% reduction in 16 months. Second, providers appear to cut prices in response to competitive launches rather than cost improvements; OpenAI's August 2024 reductions followed Claude 3.5 Sonnet and Llama 3.1 releases. Third, gross margins of 50-75\\% imply prices remain above marginal cost but are being compressed by competition. The entry of DeepSeek in January 2025 with reasoning-model pricing at \\$0.55/\\$2.19, versus OpenAI's \\$15/\\$60 for comparable capability, suggested that further compression is viable. See Demirer et al. (2025) for a comprehensive treatment of market share dynamics in the API segment.","text_sha256":"a5ff46c255113052345add83735e8d2016cdd8de30204852f77fac54cf9f1c0e"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0025","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"7 Conclusion","text":"## 7 Conclusion\n\nA priori, the problem of pricing LLM access appears intractable: users have high-dimensional private information, the allocation space is also high-dimensional, and the user's hidden allocation of tokens across tasks introduces moral hazard. In this paper, we have shown that homogeneity of the gain function generates a sufficient-statistic reduction that makes the problem solvable. Users' high-dimensional type profiles collapse to a scalar aggregate type; the seller's problem reduces to one-dimensional screening; and the optimal mechanism admits simple implementations (committed-spend contracts and two-part tariffs) that correspond precisely to the pricing structures observed at leading providers.\n\nWhile we have phrased the analysis in terms of large language models, the framework applies more broadly to any setting in which a provider sells access to a general-purpose technology with multiple input classes and heterogeneous users. Cloud computing, where providers offer a multiplicity of services with both horizontal and vertical differentiation, is a natural example. In both settings, the core pricing problem is how to allocate and monetize costly computing resources, particularly when operating under capacity constraints.\n\nSeveral extensions would enrich the analysis. First, we assumed that all tasks are homogeneous in their use of input, output, and fine-tuning tokens, and differ only vertically in how valuable a given task is to the buyer. Allowing the gain function parameters to vary across tasks would break the aggregation result. Understanding how the optimal mechanism changes when the scalar sufficient statistic no longer exists is a challenging but important open question, especially in the context of competing specialized models.\n\nSecond, we assumed that all buyer types have fine-tuning data readily available. In practice, buyers face constraints on data availability or may be unwilling to share data with the provider because of privacy, compliance, or intellectual-property concerns. A data-poor buyer may then substitute toward inference-time tokens, while a data-rich buyer substitutes toward fine-tuning tokens. When data availability is the user's private information, the provider's problem involves screening on two dimensions (aggregate type and data availability), which may require new techniques.\n\nThird, LLM platforms exhibit network effects and data externalities that our static framework abstracts from. A provider that attracts more users may improve its model through fine-tuning, reinforcement learning from human feedback, or simply from the revenue that funds further training. These dynamic complementarities between adoption and model quality could substantially reshape optimal pricing, potentially justifying the aggressive belowaverage-cost pricing observed in the API segment not just as a means to capture market share but as investment in model improvement.\n\nFourth, our analysis treats each provider as either a monopolist or a leader facing a competitive fringe. A full oligopoly analysis with multiple differentiated proprietary models would capture the strategic interactions among Anthropic, OpenAI, Google, and others that increasingly shape the market. The tractability of our framework through the reduction to one-dimensional types suggests that such an extension may be feasible and would yield further predictions about equilibrium pricing and product differentiation.\n\nThe LLM industry is at an inflection point between growth-oriented pricing and profitmaximizing pricing. Our framework provides a theoretical foundation for understanding the mechanisms through which this transition will unfold, and for evaluating whether the resulting market structure efficiently allocates this critical economic input.","text_sha256":"7884e609966af13d3b7dde61d1368b1c3c80a55d1c5cc8ae93e08cbb0768109c"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0026","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"A Proofs and Derivations","text":"## A Proofs and Derivations\n\nProof of Proposition 1 For any $z$ and $i$, the optimal allocation of inference tokens on task $i$ solves\n\n$$\n\\max _{x_{i j} \\geq 0} w_{i} \\Psi\\left(x_{i}\\right) \\Phi(z)-\\sum_{j=1}^{J} c_{j} x_{i j} .\n$$\n\nIf $w_{i}=0$ or $\\Phi(z)=0$, then $x_{i}=0$. Otherwise, by the Inada assumption, the problem admits an interior solution, which satisfies the system of $J$ first-order conditions:\n\n$$\nw_{i} \\nabla \\Psi\\left(x_{i}\\right) \\Phi(z)=c .\n$$\n\nBy (35), for all $i, \\nabla \\Psi\\left(x_{i}\\right)$ belongs to a ray with a direction $c$. Because $\\Psi$ is homogeneous and strictly concave, the ray in the space of gradients corresponds to a ray in the space of tokens. ${ }^{15}$ Thus, for each $i$, any optimal $x_{i}$ can be written as:\n\n$$\nx_{i}=d y_{i},\n$$\n\nwhere $d \\in \\mathbb{R}_{+}^{J}$ is the unique vector that solves $\\nabla \\Psi(d)=c$.\nTherefore, by the homogeneity of $\\Psi$, we have $\\nabla \\Psi\\left(x_{i}\\right)=y_{i}^{\\sigma-1} \\nabla \\Psi(d)$, and (35) can be rewritten as:\n\n$$\nw_{i} y_{i}^{\\sigma-1} \\nabla \\Psi(d) \\Phi(z)=\\nabla \\Psi(d),\n$$\n\nwhich pins down the optimal scale as:\n\n$$\ny_{i}=w_{i}^{\\frac{1}{1-\\sigma}} \\Phi(z)^{\\frac{1}{1-\\sigma}} .\n$$\n\nThe resulting total surplus ignoring fine-tuning costs is:\n\n$$\n\\begin{aligned}\n& \\int_{0}^{1} w_{i} \\Psi\\left(x_{i}\\right) \\Phi(z) d i-\\sum_{j=1}^{J} c_{j} \\int_{0}^{1} x_{i j} d i=\\int_{0}^{1} w_{i} y_{i}^{\\sigma} \\Psi(d) \\Phi(z)-y_{i}(d \\cdot c) d i \\\\\n& =\\int_{0}^{1} w_{i}^{\\frac{1}{1-\\sigma}} \\Phi(z)^{\\frac{1}{1-\\sigma}}(\\Psi(d)-d \\cdot c) d i=\\theta(w)^{\\frac{1}{1-\\sigma}} \\Phi(z)^{\\frac{1}{1-\\sigma}}(1-\\sigma) \\Psi(d)\n\\end{aligned}\n$$\n\n[^13]where $\\theta(w)$ is the aggregate type defined as\n$$\n\\theta(w) \\triangleq\\left(\\int_{0}^{1} w_{i}^{\\frac{1}{1-\\sigma}} d i\\right)^{1-\\sigma}\n$$\nand the last equality in (36) follows from the definition of $d$ and Euler's theorem for homogeneous functions: $d \\cdot \\nabla \\Psi(d)=\\sigma \\Psi(d)$.\n\nSimilarly, the total amount of inference tokens of class $j$ is:\n\n$$\n\\int_{0}^{1} x_{i j} d i=\\int_{0}^{1} d_{j} y_{i} d i=\\int_{0}^{1} d_{j} w_{i}^{\\frac{1}{1-\\sigma}} \\Phi(z)^{\\frac{1}{1-\\sigma}} d i=\\theta(w)^{\\frac{1}{1-\\sigma}} d_{j} \\Phi(z)^{\\frac{1}{1-\\sigma}}\n$$\n\nThis completes the proof. $\\square$\n\nProof of Corollary 1 The Lagrangian approach applies. Thus, associating Lagrange multipliers $\\lambda_{j} \\geq 0$ and $\\hat{\\lambda}_{k} \\geq 0$ with the capacity constraints for the different token classes, the optimal solution solves\n\n$$\n\\max _{\\left(x_{i}\\right)_{i \\in[0,1], z \\geq 0}} \\int_{0}^{1} w_{i} g\\left(x_{i}, z\\right) d i-\\sum_{j=1}^{J}\\left(c_{j}+\\lambda_{j}\\right) \\int_{0}^{1} x_{i j} d i-\\sum_{k=1}^{K}\\left(\\hat{c}_{k}+\\hat{\\lambda}_{k}\\right) z_{k} .\n$$\n\nAs such, the solution is an efficient allocation given the adjusted costs $c_{j}^{\\prime} \\triangleq c_{j}+\\lambda_{j}$ and $\\hat{c}_{k}^{\\prime} \\triangleq \\hat{c}_{k}+\\hat{\\lambda}_{k}$. The result follows. $\\square$\n\nProof of Proposition 2 Consider the buyer-optimal inference token allocation across tasks for any given token budgets $(X, Z)$ such that $\\Phi(Z)>0$ :\n\n$$\n\\max _{x_{i j} \\geq 0} \\int_{0}^{1} w_{i} \\Psi\\left(x_{i}\\right) \\Phi(Z) d i, \\quad \\text { s.t. } \\int_{0}^{1} x_{i j} d i=X_{j} \\text { for } j \\in[J] .\n$$\n\nBecause $\\Phi(Z)$ is a strictly positive constant, it can be ignored. Assume that for each $j$, $X_{j}>0$; otherwise, $x_{i j}=0$ for all $i$. Lagrangian approach applies. We associate Lagrange multipliers $\\lambda=\\left(\\lambda_{1}, \\ldots, \\lambda_{J}\\right)$ with the budget constraints of (38); the resulting first-order conditions are:\n\n$$\nw_{i} \\nabla \\Psi\\left(x_{i}\\right)=\\lambda .\n$$\n\nBy the same logic as in the efficiency analysis, these conditions imply that the relative ratios of input tokens are the same across all tasks:\n\n$$\nx_{i}=y_{i} d,\n$$\n\nfor some $d \\in \\mathbb{R}_{+}^{J}$. The budget constraints imply that $\\left(\\int_{0}^{1} y_{i} d i\\right) d=X$ and hence $d$ is proportional to $X$ and can be normalized to satisfy $\\sum_{j=1}^{J} d_{j}=1$, leading to\n\n$$\nd=\\frac{1}{\\sum_{j=1}^{J} X_{j}} X\n$$\n\nThe optimal scales $y_{i}$ solve:\n\n$$\n\\max _{y_{i} \\geq 0} \\int_{0}^{1} w_{i} y_{i}^{\\sigma} \\Psi(d) d i, \\quad \\text { s.t. } \\quad \\int_{0}^{1} y_{i} d i=\\sum_{j=1}^{J} X_{j} .\n$$\n\nAt the optimum, the marginal gains $w_{i} \\sigma y_{i}^{\\sigma-1}$ are equalized across tasks, so:\n\n$$\ny_{i}=w_{i}^{\\frac{1}{1-\\sigma}} \\frac{\\sum_{j=1}^{J} X_{j}}{\\int_{0}^{1} w_{r}^{\\frac{1}{1-\\sigma}} d r} .\n$$\n\nThe optimal buyer's payoff is:\n\n$$\n\\begin{aligned}\nU & =\\int_{0}^{1} w_{i}\\left(w_{i}^{\\frac{1}{1-\\sigma}} \\frac{\\sum_{j=1}^{J} X_{j}}{\\int_{0}^{1} w_{r}^{\\frac{1}{1-\\sigma}} d r}\\right)^{\\sigma} \\Psi(d) \\Phi(Z) d i \\\\\n& =\\left(\\int_{0}^{1} w_{i}^{\\frac{1}{1-\\sigma}} d i\\right)^{1-\\sigma}\\left(\\sum_{j=1}^{J} X_{j}\\right)^{\\sigma} \\Psi(d) \\Phi(Z)=\\theta(w) \\Psi(X) \\Phi(Z)\n\\end{aligned}\n$$\n\nwhere in the last equality we used the homogeneity of $\\Psi$ and (39). $\\square$\n\nProof of Lemma 1 The fact that $C(Q)$ is strictly increasing and strictly convex is immediate because $g(X, Z)=\\Psi(X) \\Phi(Z)$ is strictly increasing and strictly concave in its arguments whenever $g(X, Z)>0$.\n\nTo show $C_{+}^{\\prime}(0)=0$, observe that since $C(Q)$ is convex, $C_{+}^{\\prime}(0)=\\inf _{Q>0} C(Q) / Q$. Pick any $X_{0}$ such that $\\Psi\\left(X_{0}\\right)>0$ and $Z_{0}$ as in Footnote 3. For $r>0$, consider $\\left(X_{r}, Z_{r}\\right)=r\\left(X_{0}, Z_{0}\\right)$ and set $Q_{r}=g\\left(X_{r}, Z_{r}\\right)$. We have,\n\n$$\n\\frac{C\\left(Q_{r}\\right)}{Q_{r}} \\leq \\frac{c \\cdot X_{r}+\\hat{c} \\cdot Z_{r}}{Q_{r}}=\\frac{r\\left(c \\cdot X_{0}+\\hat{c} \\cdot Z_{0}\\right)}{r^{\\sigma} \\Psi\\left(X_{0}\\right) \\Phi\\left(r Z_{0}\\right)}=\\frac{c \\cdot X_{0}+\\hat{c} \\cdot Z_{0}}{\\Psi\\left(X_{0}\\right)} \\frac{1}{\\Phi\\left(r Z_{0}\\right) / r^{1-\\sigma}} .\n$$\n\nBy the choice of $Z_{0}, \\Phi\\left(r Z_{0}\\right) / r^{1-\\sigma} \\rightarrow+\\infty$ as $r \\downarrow 0$, hence the right-hand side goes to 0 . Since $Q_{r} \\downarrow 0$, it follows that $\\inf _{Q>0} C(Q) / Q=0$. $\\square$\n\nProof of Proposition 3 The result is an application of Mussa and Rosen (1978) to the buyer's payoff $\\theta Q$ from Proposition 2 and the cost function $C(Q)$ from Lemma 1. The optimal quality schedule (12) and transfers (13) follow from standard arguments. $\\square$","text_sha256":"d98dbc4f79243325fbbd9dd1bdf9c57cadd3307341b90036316d571993339eeb"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0027","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"A Proofs and Derivations","text":"Proof of Proposition 4 Fix an optimal direct mechanism $\\left(X^{*}(\\theta), Z^{*}(\\theta), T^{*}(\\theta)\\right)$ with the associated aggregate quality $Q^{*}(\\theta)=\\Psi\\left(X^{*}(\\theta)\\right) \\Phi\\left(Z^{*}(\\theta)\\right)$.\n\nMaximum-spend mechanism. Consider a maximum-spend mechanism with $p=$ $(c, \\hat{c}), T(\\theta)=T^{*}(\\theta)$, and $B(\\theta)=C\\left(Q^{*}(\\theta)\\right)$. Under this mechanism, since tokens are priced at marginal cost, each type $w$ purchasing budget $B\\left(\\theta^{\\prime}\\right)$ would optimally allocate tokens in a constrained-efficient way, thus deriving payoff $\\theta(w) Q$ where $C(Q)=B\\left(\\theta^{\\prime}\\right)$, i.e., $Q=Q^{*}\\left(\\theta^{\\prime}\\right)$. Hence, the choice across items in this mechanism is equivalent to the choice in the direct mechanism, and this mechanism implements the allocation of the direct mechanism.\n\nTwo-part-tariff mechanism. When type $w$ faces any two-part tariff, his optimal token allocation problem is equivalent to the efficient allocation problem (5) with prices taking the roles of costs. Therefore, all types $w$ with the same $\\theta(w)$ have the same payoff from any possible two-part tariff and can be treated as a single type.\n\nFaced with a given item, the problem of type $\\theta$ can be written as value-maximizationpayment-minimization:\n\n$$\n\\max _{Q \\geq 0} \\theta Q-P(Q),\n$$\n\nwhere\n\n$$\nP(Q) \\triangleq \\min _{X_{1}, \\ldots, X_{J}, Z_{1}, \\ldots, Z_{K} \\geq 0} \\sum_{j=1}^{J} p_{j} X_{j}+\\sum_{k=1}^{K} \\hat{p}_{k} Z_{k}, \\quad \\text { s.t. } \\Psi(X) \\Phi(Z)=Q .\n$$\n\nConsider the following two-part-tariff mechanism:\n\n$$\n\\begin{aligned}\n& p_{j}(\\theta)=m(\\theta) c_{j}, \\hat{p}_{k}(\\theta)=m(\\theta) \\hat{c}_{k}, j \\in[J], k \\in[K], \\\\\n& p_{0}(\\theta)=t(\\theta)-m(\\theta) C(Q(\\theta)),\n\\end{aligned}\n$$\n\nwhere $C(Q)$ is as defined in (10), $Q(\\theta)$ is as defined in (12), and $t(\\theta)$ is as defined in (13). Under this mechanism, $\\sum_{j=1}^{J} p_{j} X_{j}+\\sum_{k=1}^{K} \\hat{p}_{k} Z_{k}=m(\\theta)\\left(\\sum_{j=1}^{J} c_{j} X_{j}+\\sum_{k=1}^{K} \\hat{c}_{k} Z_{k}\\right)$. Thus, the buyer-optimal allocation is efficient, and $P(Q)=m(\\theta) C(Q)$.\n\nThe rest of the argument is standard (e.g., Tirole (1988, pp. 154-157)). A buyer-optimal $Q(\\theta)$ is determined by the first-order condition:\n\n$$\n\\theta=P^{\\prime}(Q(\\theta))=m(\\theta) C^{\\prime}(Q(\\theta)),\n$$\n\nand thus the buyer-optimal level of quality satisfies the optimality condition:\n\n$$\n\\bar{\\varphi}(\\theta)=C^{\\prime}(Q(\\theta)) .\n$$\n\nIf $m(\\theta)$ is decreasing, then the implied payment schedule in units of $C(Q)$ is concave. Thus, a menu of two-part tariffs with markups $m(\\theta)$, constant across inputs, implements the desired quality schedule $Q(\\theta)$, with each type $\\theta$ consuming the optimal amount of tokens and paying the optimal total transfer.\n\nMinimum-spend mechanism. Consider a minimum-spend mechanism that consists of $p(\\theta)=r(\\theta)(c, \\hat{c})$ and $B(\\theta)=T^{*}(\\theta)$ for all types such that $Q^{*}(\\theta)>0$, where $r(\\theta)=$ $\\frac{T^{*}(\\theta)}{C\\left(Q^{*}(\\theta)\\right)}$. By construction, since prices are proportional to costs, type $w$, when consuming aggregate quality $Q$, would optimally derive value $\\theta(w) Q$. Thus, the buyer's problem of optimal spending can be formulated in terms of $Q$. The restriction of minimum spend when reporting $\\theta^{\\prime}$ can be written as $r\\left(\\theta^{\\prime}\\right) C(Q) \\geq T^{*}\\left(\\theta^{\\prime}\\right)$, which simplifies to $Q \\geq Q^{*}\\left(\\theta^{\\prime}\\right)$. Thus, the optimal payoff of type $\\theta$ when reporting type $\\theta^{\\prime}$ is\n\n$$\nU\\left(\\theta, \\theta^{\\prime}\\right)=\\max _{Q \\geq Q^{*}\\left(\\theta^{\\prime}\\right)} \\theta Q-r\\left(\\theta^{\\prime}\\right) C(Q) .\n$$\n\nThis mechanism implements the allocation of the direct mechanism if each type prefers to report truthfully and choose $Q=Q^{*}(\\theta)$. Since $B(\\theta)$ is increasing, for this to happen, $r(\\theta)$ must be decreasing.\n\nIn fact, $r(\\theta)$ being decreasing is not only necessary but also sufficient for implementability. To see this, note that $Q^{*}(\\theta)$ is continuously increasing and $r(\\theta) \\geq 1$. Thus, the optimal deviation is always weakly below the maximum quality offered in the direct menu:\n\n$$\n\\arg \\max _{Q \\geq Q^{*}\\left(\\theta^{\\prime}\\right)} \\theta Q-r\\left(\\theta^{\\prime}\\right) C(Q) \\leq \\arg \\max _{Q \\geq Q^{*}\\left(\\theta^{\\prime}\\right)} \\bar{\\theta} Q-C(Q)=Q^{*}(\\bar{\\theta}),\n$$\n\nwhere $\\bar{\\theta}$ is the maximum point in $\\operatorname{supp} F$. Moreover, if type $\\theta$ deviates to type $\\theta^{\\prime}$ and consumes quality $Q \\leq Q^{*}(\\bar{\\theta})$, then there exists a type $\\theta^{\\prime \\prime} \\geq \\theta^{\\prime}$ such that $Q=Q^{*}\\left(\\theta^{\\prime \\prime}\\right)$. Since $r$ is decreasing, it is more profitable for type $\\theta$ to deviate to type $\\theta^{\\prime \\prime}$ and consume $Q^{*}\\left(\\theta^{\\prime \\prime}\\right)$. But such a deviation is equivalent to the deviation to $\\theta^{\\prime \\prime}$ in the direct menu and is therefore not profitable.\n\nIt remains to show that decreasing $m$ implies decreasing $r$. Assume that $m$ is decreasing. Since $m$ and $r$ are continuous, it suffices to show that their derivatives are negative at all\ndifferentiable points. At those points, $m^{\\prime}(\\theta) \\leq 0$, which implies $\\bar{\\varphi}(\\theta)-\\bar{\\varphi}^{\\prime}(\\theta) \\theta \\leq 0$, and\n\n$$\nr^{\\prime}(\\theta)=\\frac{T^{* \\prime}(\\theta) C\\left(Q^{*}(\\theta)\\right)-T^{*}(\\theta) C^{\\prime}\\left(Q^{*}(\\theta)\\right) Q^{* \\prime}(\\theta)}{C\\left(Q^{*}(\\theta)\\right)^{2}}=\\frac{Q^{\\prime}(\\theta)}{C\\left(Q^{*}(\\theta)\\right)^{2}}\\left(\\theta C\\left(Q^{*}(\\theta)\\right)-T^{*}(\\theta) \\bar{\\varphi}(\\theta)\\right),\n$$\n\nwhere we used the fact that $T^{* \\prime}(\\theta)=\\theta Q^{* \\prime}(\\theta)$ and $C^{\\prime}\\left(Q^{*}(\\theta)\\right)=\\bar{\\varphi}(\\theta)$. Since $Q$ is increasing, it suffices to show that $\\frac{\\theta}{\\bar{\\varphi}(\\theta)} \\leq \\frac{T^{*}(\\theta)}{C\\left(Q^{*}(\\theta)\\right)}$, i.e., the markup in the two-part tariff implementation is lower than the markup in the minimum-spend implementation. To show this, observe that at the maximal excluded type $\\theta_{0}, \\theta C\\left(Q^{*}(\\theta)\\right)-T^{*}(\\theta) \\bar{\\varphi}(\\theta)=0$, and for all types $\\theta>\\theta_{0}$,","text_sha256":"510ed134ddfb2ddca87893a87daad2b9d2f05dece9bf3cf4b7df933f15e1691f"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0028","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"A Proofs and Derivations","text":"$$\n\\left(\\frac{\\theta C\\left(Q^{*}(\\theta)\\right)-T^{*}(\\theta) \\bar{\\varphi}(\\theta)}{\\theta}\\right)^{\\prime}=\\frac{T(\\theta)}{\\theta^{2}}\\left(\\bar{\\varphi}(\\theta)-\\bar{\\varphi}^{\\prime}(\\theta) \\theta\\right) \\leq 0 .\n$$\n\nTherefore, $\\theta C\\left(Q^{*}(\\theta)\\right)-T^{*}(\\theta) \\bar{\\varphi}(\\theta) \\leq 0$, and $r$ is decreasing. $\\square$\n\nProof of Lemma 2 For general, not necessarily identical, $\\sigma_{l}$, the marginal gain from assigning an infinitesimal task with value $w_{i}$ to model $l$ is, by (18):\n\n$$\n\\delta_{l}\\left(w_{i}\\right)=\\frac{1-\\sigma_{l}}{\\left(\\int_{I_{l}} w_{r}^{1 /\\left(1-\\sigma_{l}\\right)} d r\\right)^{\\sigma_{l}}} Q_{l} w_{i}^{1 /\\left(1-\\sigma_{l}\\right)}\n$$\n\nUnder an optimal task split, if a task with value $w_{i}$ is assigned to a model $l$, then it must be that for all $m, \\delta_{l}\\left(w_{i}\\right) \\geq \\delta_{m}\\left(w_{i}\\right)$. Because $\\delta_{m}\\left(w_{i}\\right) / \\delta_{l}\\left(w_{i}\\right)$ is increasing in $w_{i}$ for all $m>l$, it follows that an optimal task-model allocation exists that forms a monotone partition, with models of higher $l$ (and thus higher $\\sigma_{l}$ ) being allocated the more valuable tasks.\n\nDenoting by $i_{l}$ the delimiting points, with $i_{0}=0$ and $i_{L}=1$, the optimal partition is determined by equalizing the marginal gains, $\\delta_{l}\\left(w_{i_{l}}\\right)=\\delta_{l+1}\\left(w_{i_{l}}\\right)$ for $l \\in[L-1]$ :\n\n$$\n\\frac{\\left(1-\\sigma_{l}\\right) Q_{l} w_{i_{l}}^{1 /\\left(1-\\sigma_{l}\\right)}}{\\left(1-\\sigma_{l+1}\\right) Q_{l+1} w_{i_{l}}^{1 /\\left(1-\\sigma_{l+1}\\right)}}=\\frac{\\left(\\int_{I_{l}} w_{i}^{1 /\\left(1-\\sigma_{l}\\right)} d i\\right)^{\\sigma_{l}}}{\\left(\\int_{I_{l+1}} w_{i}^{1 /\\left(1-\\sigma_{l+1}\\right)} d i\\right)^{\\sigma_{l+1}}} .\n$$\n\nWhen $\\sigma_{l} \\equiv \\sigma$, the optimality condition (42) simplifies to the following expression: ${ }^{16}$\n\n$$\n\\frac{Q_{l}}{Q_{l+1}}=\\frac{\\left(\\int_{I_{l}^{*}} w_{i}^{1 /(1-\\sigma)} d i\\right)^{\\sigma}}{\\left(\\int_{I_{l+1}^{*}} w_{i}^{1 /(1-\\sigma)} d i\\right)^{\\sigma}}\n$$\n\n[^14]Thus,\n\n$$\n\\int_{I_{l}^{*}} w_{i}^{1 /(1-\\sigma)} d i=\\frac{Q_{l}^{1 / \\sigma}}{\\sum_{m} Q_{m}^{1 / \\sigma}} \\int_{0}^{1} w_{i}^{1 /(1-\\sigma)} d i\n$$\n\nor, equivalently,\n\n$$\n\\theta_{l}\\left(I_{l}^{*}\\right)=\\left(\\frac{Q_{l}^{1 / \\sigma}}{\\sum_{m} Q_{m}^{1 / \\sigma}}\\right)^{1-\\sigma} \\theta .\n$$\n\nThe collection $\\left(\\theta_{l}\\left(I_{l}^{*}\\right)\\right)_{l=1}^{L}$ determines an optimal (monotone) partition $\\left(I_{l}^{*}\\right)_{l=1}^{L} .{ }^{17}$ All models with $Q_{l}>0$ are employed at some tasks. The resulting optimal payoff is then (20). $\\square$\n\nProof of Lemma 3 Drop the $l$ index. First, consider the case $\\hat{\\sigma}>0$. Define the unit-cost indices\n\n$$\ne \\triangleq \\min _{x \\geq 0}\\{c \\cdot x: \\Psi(x) \\geq 1\\}, \\quad \\hat{e} \\triangleq \\min _{z \\geq 0}\\{\\hat{c} \\cdot z: \\Phi(z) \\geq 1\\} .\n$$\n\nFor strictly concave, increasing, homogeneous $\\Psi, \\Phi$ these are finite and attained, and the respective minimizers $d, \\hat{d}$ are unique.\n\nBy homogeneity and the definition of $e$, the minimal cost to reach $\\Psi(x)=\\Psi_{0}$ is\n\n$$\n\\min _{x \\geq 0}\\left\\{c \\cdot x: \\Psi(x) \\geq \\Psi_{0}\\right\\}=\\Psi_{0}^{1 / \\sigma} e,\n$$\n\nattained at $x=\\Psi_{0}^{1 / \\sigma} d$. Similarly, $\\min _{z \\geq 0}\\left\\{\\hat{c} \\cdot z: \\Phi(z) \\geq \\Phi_{0}\\right\\}=\\Phi_{0}^{1 / \\hat{\\sigma}} \\hat{e}$, attained at $z=\\Phi_{0}^{1 / \\hat{\\sigma}} \\hat{d}$.\nThe cost minimization reduces to\n\n$$\n\\min _{\\Psi_{0}, \\Phi_{0} \\geq 0} \\Psi_{0}^{1 / \\sigma} e+\\Phi_{0}^{1 / \\hat{\\sigma}} \\hat{e} \\quad \\text { s.t. } \\quad \\Psi_{0} \\Phi_{0}=Q\n$$\n\nStraightforward calculation gives:\n\n$$\n\\Psi_{0}=Q^{\\frac{\\sigma}{\\sigma+\\tilde{\\sigma}}}\\left(\\frac{\\hat{e} \\sigma}{e \\hat{\\sigma}}\\right)^{\\frac{\\sigma \\hat{\\sigma}}{\\sigma+\\tilde{\\sigma}}}, \\quad \\Phi_{0}=Q^{\\frac{\\hat{\\sigma}}{\\sigma+\\hat{\\sigma}}}\\left(\\frac{\\hat{e} \\sigma}{e \\hat{\\sigma}}\\right)^{-\\frac{\\sigma \\hat{\\sigma}}{\\sigma+\\hat{\\sigma}}},\n$$\n\nwhich corresponds to the optimal token allocation:\n\n$$\nx^{*}(Q)=Q^{\\frac{1}{\\sigma+\\hat{\\sigma}}}\\left(\\frac{\\hat{e} \\sigma}{e \\hat{\\sigma}}\\right)^{\\frac{\\hat{\\sigma}}{\\sigma+\\hat{\\sigma}}} d, \\quad z^{*}(Q)=Q^{\\frac{1}{\\sigma+\\hat{\\sigma}}}\\left(\\frac{\\hat{e} \\sigma}{e \\hat{\\sigma}}\\right)^{-\\frac{\\sigma}{\\sigma+\\hat{\\sigma}}} \\hat{d}\n$$\n\nThe resulting cost function is:\n\n$$\nC(Q)=e \\Psi_{0}^{1 / \\sigma}+\\hat{e} \\Phi_{0}^{1 / \\hat{\\sigma}}=e^{\\frac{\\sigma}{\\sigma+\\hat{\\sigma}}} \\hat{e}^{\\frac{\\hat{\\sigma}}{\\sigma+\\hat{\\sigma}}}\\left[\\left(\\frac{\\sigma}{\\hat{\\sigma}}\\right)^{\\frac{\\hat{\\sigma}}{\\sigma+\\hat{\\sigma}}}+\\left(\\frac{\\hat{\\sigma}}{\\sigma}\\right)^{\\frac{\\sigma}{\\sigma+\\hat{\\sigma}}}\\right] Q^{\\frac{1}{\\sigma+\\hat{\\sigma}}}=C_{0} Q^{\\frac{1}{\\sigma+\\hat{\\sigma}}} .\n$$\n\n[^15]The case $\\hat{\\sigma}=0$ is analogous and simpler because in that case $\\Phi(z) \\equiv \\Phi_{0}>0$ and fine-tuning can be ignored. The resulting cost function is\n\n$$\nC(Q)=e\\left(\\frac{Q}{\\Phi_{0}}\\right)^{\\frac{1}{\\sigma}}=C_{0} Q^{\\frac{1}{\\sigma}} .\n$$\n\nThis completes the proof. $\\square$\n\nProof of Proposition 5 By Lemma 3, each model's surplus $\\theta Q-\\kappa_{l} Q^{1 /\\left(\\sigma+\\hat{\\sigma}_{l}\\right)}$ is concave in $Q$ and can be maximized in closed form. The efficient allocation selects the model that yields the highest surplus, yielding (22)-(23). Single-model usage follows because the cost function (21) is concave in $Q^{1 / \\sigma}$. $\\square$\n\nProof of Proposition 6 The result follows by pointwise maximization of virtual surplus with $\\varphi(\\theta)$ in place of $\\theta$ in Proposition 5. Monotonicity of $\\varphi$ ensures that the pointwise solution $Q^{\\mathrm{m}}(\\theta)=Q^{*}(\\varphi(\\theta))$ is increasing and therefore implementable. $\\square$\n\nProof of Lemma 4 When $\\hat{\\sigma}_{F}=0$, (29) simplifies into:\n\n$$\nq_{L}^{\\left(\\sigma+\\hat{\\sigma}_{L}\\right) / \\sigma}+q_{F}=\\left(\\frac{\\theta \\sigma}{c_{F}}\\right)^{1 /(1-\\sigma)}\n$$\n\nThe outside option corresponding to $q_{L}=0$ is\n\n$$\n\\psi(\\theta)=\\theta^{\\frac{1}{1-\\sigma}}(1-\\sigma)\\left(\\frac{\\sigma}{c_{F}}\\right)^{\\sigma /(1-\\sigma)}\n$$\n\nThe formulation (31) follows. $\\square$\n\nProof of Proposition 7 Denote the expression in the brackets of (32) by $\\Pi(\\theta, q)$ and observe that","text_sha256":"0c76f4ff2de2aff96ad45f26aa70b1d2b9eab8f60b890d7394dafbb51bab9342"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0029","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"A Proofs and Derivations","text":"$$\n\\Pi(\\theta, q(\\theta))= \\begin{cases}c_{F} q(\\theta)^{\\left(\\sigma+\\hat{\\sigma}_{L}\\right) / \\sigma}-c_{L} q(\\theta), & \\text { if } q(\\theta)<\\hat{q}(\\theta), \\\\ \\varphi(\\theta) q(\\theta)^{\\sigma+\\hat{\\sigma}_{L}}-c_{L} q(\\theta)-\\psi(\\theta)+\\frac{1-F(\\theta)}{f(\\theta)} \\psi^{\\prime}(\\theta), & \\text { if } q(\\theta)>\\hat{q}(\\theta) .\\end{cases}\n$$\n\nConsider the pointwise maximization of $\\Pi(\\theta, q(\\theta))$. In the region $q(\\theta) \\in[0, \\hat{q}(\\theta)], \\Pi(\\theta, q(\\theta))$ is convex in $q(\\theta)$ and therefore attains its maximum at a corner, $q(\\theta)=0$ or $q(\\theta)=\\hat{q}(\\theta)$.\n\nDirect calculation shows that $\\Pi(\\theta, 0)>\\Pi(\\theta, \\hat{q}(\\theta))$ if and only if $\\theta<\\theta_{1}$, where\n\n$$\n\\theta_{1} \\triangleq \\frac{c_{F}}{\\sigma}\\left(\\frac{c_{L}}{c_{F}}\\right)^{(1-\\sigma)\\left(\\sigma+\\hat{\\sigma}_{L}\\right) / \\hat{\\sigma}_{L}}\n$$\n\nIn the region $q(\\theta) \\geq \\hat{q}(\\theta)$, if $\\varphi(\\theta) \\leq 0$, then optimally $q(\\theta)=\\hat{q}(\\theta)$. If $\\varphi(\\theta)>0$, then $\\Pi(\\theta, q(\\theta))$ is concave in $q(\\theta)$ and therefore the maximum is either at the corner, $q(\\theta)=\\hat{q}(\\theta)$, or in the interior, in which case it equals the optimal monopoly quantity\n\n$$\nq^{\\mathrm{m}}(\\theta)=\\left(\\frac{\\left(\\sigma+\\hat{\\sigma}_{L}\\right) \\varphi(\\theta)}{c_{L}}\\right)^{1 /\\left(1-\\sigma-\\hat{\\sigma}_{L}\\right)}\n$$\n\nThus, the solution over the region $q(\\theta) \\geq \\hat{q}(\\theta)$ is interior if and only if $\\Pi_{q}(\\theta, \\hat{q}(\\theta))>0$, or\n\n$$\n\\varphi(\\theta)>\\frac{c_{L}}{\\sigma+\\hat{\\sigma}_{L}}\\left(\\frac{\\theta \\sigma}{c_{F}}\\right)^{\\frac{\\sigma\\left(1-\\sigma-\\sigma_{L}\\right)}{(1-\\sigma)\\left(\\sigma+\\tilde{\\sigma}_{L}\\right)}} .\n$$\n\nCondition (48) can hold on disjoint intervals of $\\theta$ even if $\\varphi(\\theta)$ is increasing, because the right-hand side is increasing in $\\theta$. However, the power of $\\theta$ on the right-hand side is strictly less than 1 . Thus, if $F$ satisfies monotone hazard rate, then $\\varphi^{\\prime}(\\theta) \\geq 1$, and condition (48) holds, if at all, for all $\\theta>\\theta_{2}$ where $\\theta_{2}$ solves:\n\n$$\n\\varphi\\left(\\theta_{2}\\right)=\\frac{c_{L}}{\\sigma+\\hat{\\sigma}_{L}}\\left(\\frac{\\theta_{2} \\sigma}{c_{F}}\\right)^{\\frac{\\sigma\\left(1-\\sigma-\\sigma_{L}\\right)}{(1-\\sigma)\\left(\\sigma+\\sigma_{L}\\right)}}\n$$\n\nIf (49) doesn't admit a solution, then the solution is never interior and thus for all $\\theta<\\theta_{1}$, $q(\\theta)=0$ and for all $\\theta>\\theta_{1}, q(\\theta)=\\hat{q}(\\theta)$. If (49) admits a solution at $\\theta_{2}>\\theta_{1}$, then: for all $\\theta<\\theta_{1}, q(\\theta)=0$; for all $\\theta \\in\\left(\\theta_{1}, \\theta_{2}\\right), q(\\theta)=\\hat{q}(\\theta)$; for all $\\theta \\geq \\theta_{2}, q(\\theta)=q^{\\mathrm{m}}(\\theta)$. Finally, if (49) admits a solution at $\\theta_{2}<\\theta_{1}$, then $\\theta_{3} \\in\\left[\\theta_{2}, \\theta_{1}\\right]$ exists such that for all $\\theta<\\theta_{3}, q(\\theta)=0$ and for all $\\theta \\geq \\theta_{3}, q(\\theta)=q^{\\mathrm{m}}(\\theta)$. (The threshold $\\theta_{3}>\\theta_{2}$ is determined by the condition $\\Pi\\left(\\theta_{3}, q^{\\mathrm{m}}\\left(\\theta_{3}\\right)\\right)=0$.)\n\nSince all these (relaxed) solutions are implementable, the result follows. $\\square$","text_sha256":"60cfcd69a532adccf99b0c1ed661b4577e7a8abbc15fa1f6509bb14be5684e3f"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0030","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"Leader-Fringe Competition with Uniform Types","text":"## Leader-Fringe Competition with Uniform Types\n\nCompetitive Mechanism- With $\\hat{\\sigma}_{F}=0$ the buyer's first-order condition implies\n\n$$\nq_{L}^{\\left(\\sigma+\\hat{\\sigma}_{L}\\right) / \\sigma}+q_{F}=\\left(\\frac{\\theta \\sigma}{c_{F}}\\right)^{1 /(1-\\sigma)}\n$$\n\nOn the fringe-indifference boundary $\\left(q_{F}=0\\right)$ the leader's quantity solves\n\n$$\n\\hat{q}(\\theta)=\\left(\\frac{\\theta \\sigma}{c_{F}}\\right)^{\\frac{\\sigma}{(1-\\sigma)\\left(\\sigma+\\hat{\\sigma}_{L}\\right)}} .\n$$\n\nWhen the buyer single-homes on the leader and the fringe constraint is slack,\n\n$$\nq^{\\mathrm{m}}(\\theta)=\\left(\\frac{\\left(\\sigma+\\hat{\\sigma}_{L}\\right) \\varphi(\\theta)}{c_{L}}\\right)^{\\frac{1}{1-\\sigma-\\hat{\\sigma}_{L}}}=\\left(\\frac{\\left(\\sigma+\\hat{\\sigma}_{L}\\right)(2 \\theta-1)}{c_{L}}\\right)^{\\frac{1}{1-\\sigma-\\hat{\\sigma}_{L}}} .\n$$\n\nWe want the optimal menu to be characterized by $0<\\theta_{1}<\\theta_{2}<1$ such that\n\n$$\nq^{L F}(\\theta)=\\left\\{\\begin{array}{ll}\n0, & \\theta \\leq \\theta_{1}, \\\\\n\\hat{q}(\\theta), & \\theta_{1}<\\theta<\\theta_{2}, \\\\\nq^{\\mathrm{m}}(\\theta), & \\theta \\geq \\theta_{2},\n\\end{array} \\quad q_{F}^{L F}(\\theta)= \\begin{cases}\\left(\\frac{\\theta \\sigma}{c_{F}}\\right)^{\\frac{1}{1-\\sigma}}, & \\theta \\leq \\theta_{1}, \\\\\n0, & \\theta>\\theta_{1} .\\end{cases}\\right.\n$$\n\nThe entry cutoff is\n\n$$\n\\theta_{1}=\\frac{c_{F}}{\\sigma}\\left(\\frac{c_{L}}{c_{F}}\\right)^{\\frac{(1-\\sigma)\\left(\\sigma+\\hat{\\sigma}_{L}\\right)}{\\hat{\\sigma}_{L}}} .\n$$\n\nThe boundary-interior switch $\\theta_{2}$ is the unique solution to\n\n$$\n\\varphi(\\theta)=\\frac{c_{L}}{\\sigma+\\hat{\\sigma}_{L}}\\left(\\frac{\\theta \\sigma}{c_{F}}\\right)^{\\frac{\\sigma\\left(1-\\sigma-\\hat{\\sigma}_{L}\\right)}{(1-\\sigma)\\left(\\sigma+\\hat{\\sigma}_{L}\\right)}},\n$$\n\nwhich exists in (1/2, 1) if and only if $\\frac{c_{L}}{\\sigma+\\hat{\\sigma}_{L}}\\left(\\frac{\\sigma}{c_{F}}\\right)^{\\frac{\\sigma\\left(1-\\sigma-\\hat{\\sigma}_{L}\\right)}{(1-\\sigma)\\left(\\sigma+\\hat{\\sigma}_{L}\\right)}}<1$. Furthermore, $\\theta_{2}>\\theta_{1}$ if and only if $\\theta_{1}<\\left(\\sigma+\\hat{\\sigma}_{L}\\right) /\\left(\\sigma+2 \\hat{\\sigma}_{L}\\right)$. Therefore, for $0<\\theta_{1}<\\theta_{2}<1$ to take place, the necessary and sufficient conditions are:\n\n$$\n\\frac{c_{L}}{\\sigma+\\hat{\\sigma}_{L}}\\left(\\frac{\\sigma}{c_{F}}\\right)^{\\frac{\\sigma\\left(1-\\sigma-\\hat{\\sigma}_{L}\\right)}{(1-\\sigma)\\left(\\sigma+\\hat{\\sigma}_{L}\\right)}}<1, \\quad \\frac{c_{F}}{\\sigma}\\left(\\frac{c_{L}}{c_{F}}\\right)^{\\frac{(1-\\sigma)\\left(\\sigma+\\hat{\\sigma}_{L}\\right)}{\\hat{\\sigma}_{L}}}<\\frac{\\sigma+\\hat{\\sigma}_{L}}{\\sigma+2 \\hat{\\sigma}_{L}} .\n$$\n\nEfficient Allocation - Efficiency features single-homing with a model-switch cutoff $\\hat{\\theta}$ :\n\n$$\n\\left(q_{L}^{*}(\\theta), q_{F}^{*}(\\theta)\\right)= \\begin{cases}\\left(0,\\left(\\frac{\\theta \\sigma}{c_{F}}\\right)^{\\frac{1}{1-\\sigma}}\\right), & \\theta<\\hat{\\theta} \\\\ \\left(\\left(\\frac{\\theta\\left(\\sigma+\\hat{\\sigma}_{L}\\right)}{c_{L}}\\right)^{\\frac{1}{1-\\sigma-\\hat{\\sigma}_{L}}}, 0\\right), & \\theta \\geq \\hat{\\theta}\\end{cases}\n$$\n\nwith\n\n$$\n\\hat{\\theta}=\\left(\\frac{1-\\sigma}{1-\\sigma-\\hat{\\sigma}_{L}}\\right)^{\\frac{\\left(1-\\sigma-\\hat{\\sigma}_{L}\\right)(1-\\sigma)}{\\hat{\\sigma}_{L}}}\\left(\\frac{\\sigma}{c_{F}}\\right)^{\\frac{\\sigma\\left(1-\\sigma-\\hat{\\sigma}_{L}\\right)}{\\hat{\\sigma}_{L}}}\\left(\\frac{c_{L}}{\\sigma+\\hat{\\sigma}_{L}}\\right)^{\\frac{\\left(\\sigma+\\hat{\\sigma}_{L}\\right)(1-\\sigma)}{\\hat{\\sigma}_{L}}} .\n$$\n\nMulti-Model Monopoly- Virtual type is $\\varphi(\\theta)=2 \\theta-1$. Types $\\theta<1 / 2$ are excluded. Among\nserved types the monopolist assigns a single model, switching at $\\theta^{\\mathrm{m}}=(1+\\hat{\\theta}) / 2$. Optimal quantities are\n\n$$\n\\left(q_{L}^{\\mathrm{m}}(\\theta), q_{F}^{\\mathrm{m}}(\\theta)\\right)= \\begin{cases}(0,0), & \\theta<\\frac{1}{2}, \\\\ \\left(0,\\left(\\frac{\\sigma \\varphi(\\theta)}{c_{F}}\\right)^{\\frac{1}{1-\\sigma}}\\right), & \\frac{1}{2} \\leq \\theta<\\theta^{\\mathrm{m}}, \\\\ \\left(\\left(\\frac{\\left(\\sigma+\\hat{\\sigma}_{L}\\right) \\varphi(\\theta)}{c_{L}}\\right)^{\\frac{1}{1-\\sigma-\\hat{\\sigma}_{L}}}, 0\\right), & \\theta \\geq \\theta^{\\mathrm{m}} .\\end{cases}\n$$","text_sha256":"921a7fd405703fec8b30ea39c65e675752befb24d58c3944286fa480396722ec"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0031","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"B Contractible Tasks","text":"## B Contractible Tasks\n\nThroughout the manuscript, we focus on the more realistic case in which the LLM provider cannot contract on the allocation of tokens across the buyer's tasks. However, the fullcontracting setting provides a natural benchmark and we study it in this section. Thus, the seller designs a direct menu that specifies a token allocation across different tasks:\n\n$$\n\\left\\{\\left(\\left(x_{i 1}(w), \\ldots, x_{i J}(w)\\right)_{i \\in[0,1]}, z_{1}(w), \\ldots, z_{K}(w), t(w)\\right)\\right\\}_{w} .\n$$\n\nThe allocation being contractible means the buyer has no freedom to reallocate tokens across tasks. This naturally increases the scope for screening. This also makes the problem intractable in full generality, because the reduction to a one-dimensional aggregate type and quality as in Proposition 2 does not apply. Nevertheless, in this section we obtain complete characterizations in two special settings: the case of separable type distributions (Section B.1) and the case of two types (Section B.2). In the former case, the optimal solution can be obtained via a menu of token budgets. In the latter case, the optimal solution can be obtained by identifying the structure of binding incentive constraints.","text_sha256":"4d81ea807c765293d55e1ec7d4d1d0f0a1d404831f0fde8d418e9c0c0ddb768b"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0032","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"B. 1 Separable Type Distribution and Cost-Based Pricing","text":"## B. 1 Separable Type Distribution and Cost-Based Pricing\n\nIn this section, we provide a sufficient condition on type distribution under which contracting on tasks is not profitable. In fact, under that condition an even less contractually restrictive class of mechanisms, cost-based tariffs, is optimal:\n\nDefinition 4 (Cost-Based Tariff). A cost-based tariff is a menu of monetary budgets and transfers $\\left\\{\\left(B_{n}, T_{n}\\right)\\right\\}_{n}$ such that, upon purchasing item $n$, the buyer pays $T_{n}$ for access to budget $B_{n}$, which he can freely spend on tokens priced at their marginal costs.\n\nWe build on the analysis of Armstrong (1996) and introduce the conditions that are jointly sufficient for the optimality of cost-based tariffs: demand separability and type separability. We show that the former is always satisfied in our setting, whereas the latter is equivalent to the aggregate type not being informative about the relative task weights.\n\nTo this end, consider a problem of optimal token allocation by type $w$ given a monetary budget $B$ and token prices equal to marginal costs:\n\n$$\nV(w, B)=\\max _{\\left\\{\\left(x_{i}\\right)_{i \\in[0,1], z \\geq 0\\}}\\right.} \\int_{0}^{1} w_{i} g\\left(x_{i}, z\\right) d i, \\quad \\text { s.t. } \\int_{0}^{1} \\sum_{j=1}^{J} c_{j} x_{i j} d i+\\sum_{k=1}^{K} \\hat{c}_{k} z_{k}=B,\n$$\n\nDefinition 5 (Demand Separability). Demand is separable if there exist functions $V_{1}(w)$ and $V_{2}(B)$ such that\n\n$$\nV(w, B)=V_{1}(w) V_{2}(B) .\n$$\n\nIf demand is separable, then faced with a cost-based tariff, all types $w$ with the same $V_{1}(w)$ purchase the same item.\n\nDemand separability always holds in our setting. Indeed, by Proposition 2, if type $w$ purchases a total amount of inference tokens $X=\\left(X_{1}, \\ldots, X_{J}\\right)$ and fine-tuning tokens $Z=$ $\\left(Z_{1}, \\ldots, Z_{K}\\right)$, then his optimal payoff is $\\theta(w) g(X, Z)$. Therefore, (51) holds with $V_{1}(w)=$ $\\theta(w)$ and\n\n$$\nV_{2}(B)=\\max _{X \\geq 0, Z \\geq 0} g(X, Z), \\quad \\text { s.t. } \\sum_{j=1}^{J} c_{j} X_{j}+\\sum_{k=1}^{K} \\hat{c}_{k} Z_{k}=B .\n$$\n\nDefinition 6 (Type Separability). Type distribution is separable if $f_{1}, f_{2}$ exist such that $f(w)=f_{1}(\\theta(w)) f_{2}(w)$ for all $w$ and $f_{2}$ is homogeneous of degree zero.\n\nA separable type distribution means that knowing $\\theta(w)$ provides no information about which ray from the origin $w$ lies on, that is, about the relative weights the buyer assigns to different tasks. Equivalently, denoting by $\\|\\cdot\\|_{p}$ a standard $L_{p}$ norm, $\\theta(w)=\\|w\\|_{1 /(1-\\sigma)}$ and the type separability is equivalent to $\\|w\\|_{1 /(1-\\sigma)}$ and $w /\\|w\\|_{1 /(1-\\sigma)}$ to be independent. Thus, any type distribution generated by a draw of a \"total size\" $\\|w\\|_{1 /(1-\\sigma)}$ according to an arbitrary distribution together with an independent draw of \"relative weights\" $w /\\|w\\|_{1 /(1-\\sigma)}$ according to any other distribution, is separable.\n\nProposition 8 (Cost-Based Optimality). If the type distribution is separable, then a cost-based tariff is optimal.\n\nProof. Since in our setting demand is always separable, the result follows from the arguments presented by Armstrong (1996). Specifically, we can relax the IC constraints between types located on different rays from the origin. Then, applying the Envelope Theorem ray by ray, we can find a solution to the relaxed problem as follows:\n\n$$\n\\left(x^{*}, z^{*}\\right) \\in \\arg \\max _{\\left\\{\\left(x_{i}\\right)_{i \\in[0,1],} z \\geq 0\\right\\}}\\left(1-\\frac{\\int_{1}^{\\infty} r^{n-1} f(r w) d r}{f(w)}\\right) u(w, x, z)-c(x, z) .\n$$\n\nGiven the demand separability, this condition can be rewritten as:\n\n$$\n\\left(x^{*}, z^{*}\\right) \\in \\arg \\max _{\\left\\{\\left(x_{i}\\right)_{i \\in[0,1]}, z \\geq 0\\right\\}} u(w, x, z), \\quad \\text { s.t. } c(x, z) \\leq B^{*}(w),\n$$\n\nand\n\n$$\nB^{*}(w) \\in \\arg \\max _{B \\geq 0}\\left(1-\\frac{\\int_{1}^{\\infty} r^{n-1} f(r w) d r}{f(w)}\\right) \\theta(w) V_{2}(B)-B .\n$$\n\nAt the same time, the optimal cost-based tariff implements an allocation\n\n$$\n\\left(x^{*}, z^{*}\\right) \\in \\arg \\max _{\\left\\{\\left(x_{i}\\right)_{i \\in[0,1]}, z \\geq 0\\right\\}} u(w, x, z), \\quad \\text { s.t. } c(x, z) \\leq B^{*}(w),\n$$\n\nwhere\n\n$$\nB^{*}(w) \\in \\arg \\max _{B}\\left(\\theta(w)-\\frac{1-F(\\theta(w))}{f(\\theta(w))}\\right) V_{2}(B)-B .\n$$\n\nIf the type distribution is separable, then\n\n$$\n\\left(1-\\frac{\\int_{1}^{\\infty} r^{n-1} f(r w) d r}{f(w)}\\right) \\theta(w)=\\theta(w)-\\frac{1-F(\\theta(w))}{f(\\theta(w))},\n$$\n\nand the cost-based tariff implements the solution to the relaxed problem. Therefore, it is (indirectly) optimal. $\\square$","text_sha256":"ba7d99e745658521e37ae7b96c507249dec22738cd072132058705e025a28a8b"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0033","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"B. 2 Binary Types","text":"## B. 2 Binary Types\n\nAlternatively, let there be only two types, $w_{1}$ and $w_{2}$, which occur with prior strictly positive probabilities $f_{1}$ and $f_{2}$, respectively. In this case, the optimal mechanism depends on what happens if the seller attempts to extract the first-best level of surplus, that is, if she offers a menu containing the efficient amounts of tokens for each type with prices equal to their respective added values. If this menu is incentive compatible, then it is clearly optimal, and we call the type with the higher payment \"high\" and associate the label $H$ with it. If this menu is not incentive compatible, then we call \"high\" the type whose incentive constraint is violated. We call the other type \"low\" and associate the label $L$ with it.\n\nTo design an optimal menu, it is important to determine which type out of $w_{1}$ and $w_{2}$ is high and which one is low. If $w_{2}$ dominates $w_{1}$ for every task, then it is clearly high; alternatively, the types are in some sense horizontally differentiated and the distinction is less clear. However, it turns out that the ranking can be derived from the aggregate types, and the optimal menu admits a simple characterization.\n\nProposition 9 (Binary Types). The high type is the one with the higher aggregate type, i.e., $\\theta\\left(w_{H}\\right) \\geq \\theta\\left(w_{L}\\right)$. In the optimal menu, the token allocation of $w_{H}$ is always efficient. If\n\n$$\n\\int_{0}^{1}\\left(w_{H i}-w_{L i}\\right) w_{L i}^{\\frac{\\sigma}{1-\\sigma}} d i \\leq 0\n$$\n\nthen the token allocation of $w_{L}$ is also efficient and the seller extracts full surplus. Otherwise, the token allocation of $w_{L}$ is efficient with respect to a virtual type $w_{L}-\\left(w_{H}-w_{L}\\right) f_{H} / f_{L}$.\n\nProof. With a small abuse of notation relative to previous sections, denote by $q=\\left(q_{i}\\right)_{i \\in[0,1]}$ the profile of qualities delivered to each task, $q_{i} \\triangleq g\\left(x_{i}, z\\right)$, and denote by $C(q)$ the minimal total cost of generating a given profile $q$ :\n\n$$\nC(q) \\triangleq \\min _{x_{i}, z \\geq 0} \\int_{0}^{1} \\sum_{j=1}^{J} c_{j} x_{i j} d i+\\sum_{k=1}^{K} \\hat{c}_{k} z_{k}, \\quad \\text { s.t. } g\\left(x_{i}, z\\right)=q_{i}, \\forall i \\in[0,1] .\n$$\n\nBecause the set of feasible profiles $q$ is convex, it follows from the analysis of Haghpanah and Siegel (2024) on general screening problems with two buyer types that in our setting either (i) the seller extracts full surplus; or (ii) the incentive constraint of type $w_{H}$ and the individual rationality constraint of type $w_{L}$ bind.\n\nIt follows that if the seller cannot extract full surplus, then the seller's problem can be written as\n\n$$\n\\begin{aligned}\n\\max _{q_{L}, q_{H}, t_{L}, t_{H}} & f_{L}\\left(t_{L}-C\\left(q_{L}\\right)\\right)+f_{H}\\left(t_{H}-C\\left(q_{H}\\right)\\right) \\\\\n\\text { s.t. } & \\int_{0}^{1} w_{H i} q_{H i} d i-t_{H}=\\int_{0}^{1} w_{H i} q_{L i} d i-t_{L}, \\int_{0}^{1} w_{L i} q_{L i} d i-t_{L}=0\n\\end{aligned}\n$$\n\nSolving for transfers from the constraints, the problem can be restated as:\n\n$$\n\\max _{q_{L}, q_{H}} f_{L}\\left(\\int_{0}^{1}\\left(w_{L i}-\\frac{f_{H}}{f_{L}}\\left(w_{H i}-w_{L i}\\right)\\right) q_{L i} d i-C\\left(q_{L}\\right)\\right)+f_{H}\\left(\\int_{0}^{1} w_{H i} q_{H i} d i-C\\left(q_{H}\\right)\\right),\n$$\n\nwhich is solved by the allocation efficient relative to the virtual types.\nIt is left to determine which type is high and provide conditions for full surplus extraction. Consider a mechanism that attempts full surplus extraction. By the efficiency analysis behind Proposition 1, under this mechanism type $w$, when reporting type $\\tilde{w}$, obtains added value\n\n$$\nu(w, \\tilde{w})=\\int_{0}^{1} w_{i} \\Psi\\left(x_{i}(\\tilde{w})\\right) \\Phi(z(\\tilde{w})) d i=\\int_{0}^{1} w_{i} \\tilde{w}_{i}^{\\frac{\\sigma}{1-\\sigma}} \\Psi(d) \\Phi(z(\\tilde{w}))^{\\frac{1}{1-\\sigma}} d i\n$$\n\nBecause under truth-telling each type obtains zero rents, the corresponding payment is\n$t(w)=u(w, w)$, and the incentive constraint $u(w, w)-t(w) \\geq u(w, \\tilde{w})-t(\\tilde{w})$ is violated if and only if:\n\n$$\n\\Psi(d) \\Phi(z(\\tilde{w}))^{\\frac{1}{1-\\sigma}} \\int_{0}^{1}\\left(w_{i}-\\tilde{w}_{i}\\right) \\tilde{w}_{i}^{\\frac{\\sigma}{1-\\sigma}} d i>0\n$$\n\nBy Jensen's inequality and the concavity of the logarithm:\n\n$$\nw_{i} \\tilde{w}_{i}^{\\frac{\\sigma}{1-\\sigma}}=w_{i}^{\\frac{1-\\sigma}{1-\\sigma}} \\tilde{w}_{i}^{\\frac{\\sigma}{1-\\sigma}} \\leq(1-\\sigma) w_{i}^{\\frac{1}{1-\\sigma}}+\\sigma \\tilde{w}_{i}^{\\frac{1}{1-\\sigma}} .\n$$\n\nTherefore, inequality (53) implies\n\n$$\n\\int_{0}^{1}\\left(w_{i}^{\\frac{1}{1-\\sigma}}-\\tilde{w}_{i}^{\\frac{1}{1-\\sigma}}\\right) d i>0\n$$\n\nand the incentive violation is possible only if $\\theta(w)>\\theta(\\tilde{w})$, i.e., from high to low aggregate type. Therefore, if full surplus extraction is not incentive compatible, then the high type is the one with the higher aggregate type. If full surplus extraction is incentive compatible, then the high type is the one with the higher aggregate type directly by the analysis behind Proposition 1. The result follows. $\\square$\n\nNote that even though the virtual types in Proposition 9 follow the standard formula of the single-dimensional case, each type there is infinite-dimensional.\n\nOne might wonder whether the correspondence between the incentive order and aggregate types extends beyond the binary-type case. This is not the case. In Section B.3, we study a special case of our model in which the buyer's type can be effectively parameterized by two variables: the number of ex ante homogeneous tasks to which the buyer attaches positive value, and the value he attaches to each of those tasks. We show that the incentive constraints that bind in the optimal mechanism do not admit a one-dimensional structure. Instead, they form an infinite collection of one-dimensional segments. In that setting the optimal mechanisms also admit a natural implementation via a menu of two-part tariffs. However, in contrast to the indirect implementations described in Section 4.2, each item in the menu must be accompanied by task caps, that is, restrictions on the number of tasks the buyer can process.","text_sha256":"c4338dbc2ecb0a9a62acdc5671431a0c8a184ccaa8fd8e13671c348da6556ee9"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0034","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"B. 3 Value-Scale Heterogeneity","text":"## B. 3 Value-Scale Heterogeneity\n\nIn this section, we characterize an optimal menu with contractible token allocations across tasks in the case of value-scale heterogeneity. In this case, each multidimensional type $w$ is\ncharacterized (with a small abuse of notation) by two parameters $(w, s)$ such that\n\n$$\nw_{i}=\\left\\{\\begin{array}{l}\nw, \\text { if } i \\leq s \\\\\n0, \\text { if } i>s\n\\end{array}\\right.\n$$\n\nin which $w$ and $s$ are independently distributed according to CDFs $F_{w}$ and $F_{s}$, with $F_{w}$ featuring increasing virtual values. We will argue that in this case, under additional assumptions, the binding incentive constraints are those within each scale, and the seller is able to generate the same profit as if the scale were observable.\n\nFor this section, we drop the product structure and the homogeneity requirements of (1) and instead require that $g\\left(x_{i}, z\\right)$ is (i) positive and continuous on $\\mathbb{R}_{+}^{J+K}$, (ii) strictly monotone, twice continuously differentiable, strictly concave with negative definite Hessian, and with all cross-partial derivatives strictly positive on $\\mathbb{R}_{++}^{J+K}$, (iii) $g(0)=0$, and (iv) Inada at zero, i.e., $\\lim _{y_{m} \\downarrow 0} g_{y_{m}}\\left(y_{m}, y_{-m}\\right)=+\\infty$ for all $y_{-m} \\in \\mathbb{R}_{+}^{J+K-1}$ such that $g\\left(0, y_{-m}\\right)>0 .{ }^{18}$\n\nConsider the problem in which the scale is commonly known to be $s>0$. The buyer's payoff from any given item on the menu (50) is\n\n$$\nw \\int_{i=0}^{s} g\\left(x_{i}, z\\right) d i-t\n$$\n\nFor any reported $w$ the seller should optimize token allocation to deliver a promised level of (total) quality $q$,\n\n$$\nw q-t,\n$$\n\nwith the minimal cost function $C(q, s)$ of delivering a given quality being:\n\n$$\nC(q, s)=\\min _{\\left(x_{i}\\right)_{i \\in[0,1], z \\geq 0}} \\int_{i=0}^{s} \\sum_{j=1}^{J} c_{j} x_{i j} d i+\\sum_{k=1}^{K} \\hat{c}_{k} z_{k}, \\quad \\text { s.t. } \\int_{i=0}^{s} g\\left(x_{i}, z\\right) d i=q \\text {. }\n$$\n\nSince $g$ is strictly concave, the solution to this problem is achieved by allocating the inference tokens uniformly across the $s$ tasks. The problem can be equivalently stated as:\n\n$$\nC(q, s)=\\min _{x, z \\geq 0} s \\sum_{j=1}^{J} c_{j} x_{j}+\\sum_{k=1}^{K} \\hat{c}_{k} z_{k}, \\quad \\text { s.t. } s g(x, z)=q .\n$$\n\nThe resulting cost function $C(q, s)$ satisfies, over the domain of admissible $(q, s)$, the following properties:\n\n[^16]","text_sha256":"7e05b8802c4689932b1e1bfc8dbd3171e5c67dad498f954ac06debe74dcd684a"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0035","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"Lemma 5 (Cost Function).","text":"## Lemma 5 (Cost Function).\n\n1. $C(q, s)$ is strictly increasing and strictly convex in $q$ with $C_{q}(0, s)=0$.\n2. $C(q, s)$ is strictly decreasing in $s$ for $q>0$.\n3. $C(q, s)$ is submodular, i.e., $C_{q}(q, s)$ is decreasing in $s$ for all $q$.\n\nProof. The first two statements follow directly from our assumptions on $g$. To establish the third property, posit the Lagrangian for the cost minimization problem:\n\n$$\nL=s \\sum_{j=1}^{J} c_{j} x_{j}+\\sum_{k=1}^{K} \\hat{c}_{k} z_{k}+\\lambda(q-s g(x, z)) .\n$$\n\nBy the Envelope Theorem, $C_{q}(q, s)=\\lambda(q, s)$. Thus, it suffices to show that $\\lambda(q, s)$ is decreasing in $s$ for a fixed $q$.\n\nTo this end, for any $q>0$, by the Inada condition, the solution must be interior, $(x, z) \\gg$ 0 . The constraint binds, so $\\lambda>0$. Denoting $(x, z)$ by $y$, the first-order conditions are:\n\n$$\ng_{x_{j}}(y)=c_{j} / \\lambda, \\quad g_{z_{k}}(y)=\\hat{c}_{k} /(\\lambda s) .\n$$\n\nDefine the Hessian function $H(y) \\triangleq \\nabla^{2} g(y)$. Denote by $A, B \\in \\mathbb{R}^{J+K}$ vectors such that $A_{r}=c_{r} / \\lambda^{2}$ if $r \\leq J$ and $=\\hat{c}_{r-J} /\\left(\\lambda^{2} s\\right)$ if $r>J, B_{r}=0$ if $r \\leq J$ and $=\\hat{c}_{r-J} /\\left(\\lambda s^{2}\\right)$ if $r>J$. Then, differentiating (57) with respect to $s$, we obtain:\n\n$$\nH \\frac{d y}{d s}=-A \\frac{d \\lambda}{d s}-B .\n$$\n\nAt the same time, differentiating the constraint $s g(y(s))=q$ with respect to $s$, we obtain:\n\n$$\ng(y)+s \\nabla g(y)^{\\top} \\frac{d y}{d s}=0 .\n$$\n\nSolving for $d y / d s$ in (58) and substituting it into (59), we obtain, omitting the dependence on $y$ :\n\n$$\n\\frac{d \\lambda}{d s} s \\nabla g^{\\top}\\left(-H^{-1} A\\right)=-g-s \\nabla g^{\\top}\\left(-H^{-1} B\\right) .\n$$\n\nNow, observe that for all $y \\in \\mathbb{R}_{++}^{J+K}$, by assumption on $g, H(y)$ is symmetric negative definite; moreover, by the positivity of cross-derivative, all off-diagonal entries of $H(y)$ are strictly positive. Then, $-H(y)$ is a Stieltjes matrix and, consequently, all elements of $H(y)^{-1}$ are negative. It immediately follows from (60) that whenever $q>0$ and $s>0, d \\lambda / d s<0$. □\n\nAs such, the seller's problem for any given $s$ is analogous to Mussa and Rosen (1978). Since $w$ and $s$ are independently distributed, we can drop the dependence on $s$ and define the virtual value as\n\n$$\n\\varphi(w) \\triangleq w-\\frac{1-F_{w}(w)}{f_{w}(w)} .\n$$\n\nSince $\\varphi(w)$ is increasing, all $w$ with $\\varphi(w) \\leq 0$ are excluded and all other $w$ receive the quality level $q(w, s)$ that solves:\n\n$$\n\\varphi(w)=C_{q}(q(w, s), s) .\n$$\n\nThe corresponding optimal transfers are\n\n$$\nt(w, s)=w q(w, s)-\\int_{0}^{w} q(r, s) d r\n$$\n\nLemma 5 shows that it is cheaper to generate an extra unit of (total) quality when you have more tasks. This property is intuitive given that the returns on each task are diminishing. Thus, the optimal quality, and hence the buyer's rent, increase in scale for any given value $w$.\n\nAssumption 2 (Bounded Rent Increase). For all $w, s$, the function $q(w, s)$ defined in (62) satisfies $\\int_{0}^{w} s q_{s}(r, s) d r \\leq w q(w, s)$.\n\nAssumption 2 requires that the buyer's rent does not grow too quickly and, specifically, that the marginal increase of buyer rent from having an additional task is smaller than the average equilibrium value generated by LLM across existing tasks.\n\nProposition 10 (Optimal Menu of Token Allocations). Under Assumption 2, an optimal menu is\n\n$$\n\\left(\\left(x_{i}(w, s)\\right)_{i \\in[0,1]}, z(w, s), t(w, s)\\right)_{(w, s)},\n$$\n\nwhere for each $(w, s),\\left(\\left(x_{i}(w, s)\\right)_{i \\in[0,1]}, z(w, s)\\right)$ are cost-minimizing tokens from (55) that deliver quality $q(w, s)$ as defined in (62), and $t(w, s)$ is as defined in (63).\n\nProof. If each type reports truthfully, then the menu attains the profits of the observablescale benchmark and is thus optimal.\n\nIf type $(w, s)$ deviates to $(w, \\tilde{s})$ with $\\tilde{s} \\leq s$, then, under the proposed menu, he obtains exactly the same payoff as type $(w, \\tilde{s})$, because he processes the same number of tasks with the same willingness to pay for quality. By Lemma $5, C_{q}(q, s)$ is decreasing in $s$ for all $q$, and thus $q(w, s)$ is increasing in $s$ for all $w$. Therefore, the rents accrued by type $(w, s)$ under truth-telling,\n\n$$\nU(w, s)=\\int_{0}^{w} q(r, s) d r\n$$\n\nare increasing in $s$ for all $w$. Therefore, $(w, s)$ does not want to deviate to $(w, \\tilde{s})$ with $\\tilde{s} \\leq s$. Furthermore, by incentive compatibility within a given $\\tilde{s},(w, \\tilde{s})$ does not want to deviate to $(\\tilde{w}, \\tilde{s})$. Therefore, $(w, s)$ does not want to deviate to any $(\\tilde{w}, \\tilde{s})$ with $\\tilde{s} \\leq s$.\n\nIf type $(w, s)$ deviates to $(\\tilde{w}, \\tilde{s})$ with $\\tilde{s}>s$, then he obtains gross payoff $w q(\\tilde{w}, \\tilde{s}) s / \\tilde{s}$ and pays the transfer $t(\\tilde{w}, \\tilde{s})$. Therefore, the optimal double deviation strategy for a misreporting type solves\n\n$$\n\\max _{\\tilde{w} \\geq 0}\\left[w q(\\tilde{w}, \\tilde{s}) \\frac{s}{\\tilde{s}}-\\tilde{w} q(\\tilde{w}, \\tilde{s})+\\int_{0}^{\\tilde{w}} q(r, \\tilde{s}) d r\\right] .\n$$\n\nBecause the mechanism incentivizes truthful reporting by any $\\tilde{s}$-truthtelling type, including the type ( $w s / \\tilde{s}, \\tilde{s}$ ), it follows that\n\n$$\n\\tilde{w}^{*}=\\frac{w s}{\\tilde{s}}, \\quad U(w ; s, \\tilde{s})=\\int_{0}^{\\frac{w s}{\\tilde{s}}} q(r, \\tilde{s}) d r\n$$\n\nThe condition that discourages local deviations to $\\tilde{s}>s$ is:\n\n$$\ns \\int_{0}^{w} q_{s}(r, s) d r \\leq w q(w, s)\n$$\n\nfor all $(w, s)$, which is precisely Assumption 2. Furthermore, observe that for $\\tilde{s}>s$ :\n\n$$\n\\begin{aligned}\nU_{\\tilde{s}}(w ; s, \\tilde{s}) & =\\int_{0}^{\\frac{w s}{\\tilde{s}}} q_{s}(r, \\tilde{s}) d r-\\frac{w s}{\\tilde{s}^{2}} q\\left(\\frac{w s}{\\tilde{s}}, \\tilde{s}\\right) \\\\\n& =\\frac{1}{\\tilde{s}}\\left(\\tilde{s} \\int_{0}^{\\frac{w s}{\\tilde{s}}} q_{s}(r, \\tilde{s}) d r-\\frac{w s}{\\tilde{s}} q\\left(\\frac{w s}{\\tilde{s}}, \\tilde{s}\\right)\\right) \\leq 0\n\\end{aligned}\n$$\n\nwhere the inequality holds by Assumption 2 with $(w, s)$ replaced by $(s w / \\tilde{s}, \\tilde{s})$. Therefore, the upward global deviations in $\\tilde{s}$ are suboptimal and the result follows. $\\square$","text_sha256":"f544c919c9630a8f9268538385fda241ab5788e0fda6f48f29ed66c3a3d05461"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0036","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"References","text":"## References\n\nArmstrong, M. (1996): \"Multiproduct Nonlinear Pricing,\" Econometrica, 64, 51-76.\nBabaioff, M., R. Kleinberg, and R. Paes Leme (2012): \"Optimal Mechanisms for Selling Information,\" in Proceedings of the 13th ACM Conference on Electronic Commerce, 92-109.\n\nBergemann, D., M. Bojko, P. Dütting, R. Paes Leme, H. Xu, and S. Zuo (2024): \"Data-Driven Mechanism Design: Jointly Eliciting Preferences and Information,\" arXiv preprint arXiv:2412.16132.\n\nBergemann, D., A. Bonatti, and A. Smolin (2018): \"The Design and Price of Information,\" American Economic Review, 108, 1-48.\n\nCalzolari, G. and V. Denicolò (2013): \"Competition with Exclusive Contracts and Market-Share Discounts,\" American Economic Review, 103, 2384-2411.\n\n- (2015): \"Exclusive Contracts and Market Dominance,\" American Economic Review, 105, 3321-3351.\n\nCastro-Pires, H., H. Chade, and J. Swinkels (2024): \"Disentangling Moral Hazard and Adverse Selection,\" American Economic Review, 114, 1-37.\n\nDaskalakis, C., A. Deckelbaum, and C. Tzamos (2017): \"Strong Duality for a Multiple-Good Monopolist,\" Econometrica, 85, 735-767.\n\nDemirer, M., A. Fradkin, N. Tadelis, and S. Peng (2025): \"The Emerging Market for Intelligence: Pricing, Supply, and Demand for LLMs,\" Tech. rep., National Bureau of Economic Research.\n\nDevanur, N. R., K. Goldner, R. R. Saxena, A. Schvartzman, and S. M. Weinberg (2020): \"Optimal Mechanism Design for Single-Minded Agents,\" in Proceedings of the 21st ACM Conference on Economics and Computation, 193-256.\n\nDoligalski, P., P. Dworczak, J. Krysta, and F. Tokarski (2025): \"Incentive separability,\" Journal of Political Economy Microeconomics, 3, 539-567.\n\nDuetting, P., V. Mirrokni, R. Paes Leme, H. Xu, and S. Zuo (2024): \"Mechanism design for large language models,\" in Proceedings of the ACM on Web Conference 2024, 144-155.\n\nFiat, A., K. Goldner, A. R. Karlin, and E. Koutsoupias (2016): \"The Fedex Problem,\" in Proceedings of the 2016 ACM Conference on Economics and Computation, 21-22.\n\nFish, S., Y. A. Gonczarowski, and R. I. Shorrer (2024): \"Algorithmic Collusion by Large Language Models,\" arXiv preprint arXiv:2404.00806.\n\nHaghpanah, N. and R. Siegel (2024): \"Screening Two Types,\" Tech. rep., Penn State.\nJullien, B. (2000): \"Participation Constraints in Adverse Selection Models,\" Journal of Economic Theory, 93, 1-47.\n\nKaplan, J., S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei (2020): \"Scaling Laws for Neural Language Models,\" arXiv preprint arXiv:2001.08361.\n\nLaffont, J.-J. and J. Tirole (1990): \"The regulation of multiproduct firms: Part I: Theory,\" Journal of Public Economics, 43, 1-36.\n\nMahmood, R. (2024): \"Pricing and Competition for Generative AI,\" arXiv preprint arXiv:2411.02661.\n\nMussa, M. and S. Rosen (1978): \"Monopoly and Product Quality,\" Journal of Economic Theory, 18, 301-317.\n\nRochet, J.-C. and L. A. Stole (2003): \"The Economics of Multidimensional Screening,\" Econometric Society Monographs, 35, 150-197.\n\nTirole, J. (1988): The Theory of Industrial Organization, Cambridge: MIT Press.\nWu, Y., Z. Sun, S. Li, S. Welleck, and Y. Yang (2025): \"Inference Scaling Laws: An Empirical analysis of Compute-Optimal Inference for LLM Problem-Solving,\" in The Thirteenth International Conference on Learning Representations.\n\nYang, K. H. (2022): \"Selling Consumer Data for Profit: Optimal Market-Segmentation Design and Its Consequences,\" American Economic Review, 112, 1364-1393.\n\n[^0]:    *Bergemann: Department of Economics, Yale University, dirk.bergemann@yale.edu. Bonatti: MIT Sloan, bonatti@mit.edu. Smolin: Toulouse School of Economics, alexey.v.smolin@gmail.com. We thank Mark Armstrong, Mert Demirer, Scott Kominers, Antonio Russo, Ron Siegel, and Frank Yang for valuable comments and discussions, as well as audiences at the Triangle Conference, USC, Caltech, Virginia, the Chicago workshop, the Sciences Po workshop, the NBER Market Design workshop, VSET, Brown, LSE, the TSE Digital Economics Conference, the IESE Economics of AI Conference, and Cambridge. An extended abstract of an earlier version of this paper appeared as \"The Economics of Large Language Models: Token Allocation, Fine-Tuning, and Optimal Pricing\" in the proceedings of EC'25. Dirk Bergemann gratefully acknowledges financial support from NSF SES 2049754 and ONR MURI. Alessandro Bonatti gratefully acknowledges financial support from NSF SES 2519401. Alex Smolin gratefully acknowledges funding from the French National Research Agency (ANR) under the Investments for the Future program (grant ANR-17-EURE-0010) and through the AI Interdisciplinary Institute ANITI (grant ANR-23-IACL-0002).\n\n[^1]:    ${ }^{1}$ For the practical challenges associated with LLM pricing, see https://www.wsj.com/articles/no-one-knows-how-to-price-ai-tools-f346ea8a.\n\n[^2]:    ${ }^{2}$ That is, for all $j$ and $x_{i,-j} \\in \\mathbb{R}_{+}^{J-1}$ such that $\\Psi\\left(0, x_{i,-j}\\right)>0, \\lim _{x_{i j} \\downarrow 0} \\Psi_{j}\\left(x_{i j}, x_{i,-j}\\right)=+\\infty$.\n    ${ }^{3}$ That is, there exists $z_{0} \\geq 0, z_{0} \\neq 0$, such that $\\lim _{r \\downarrow 0} \\Phi\\left(r z_{0}\\right) / r^{1-\\sigma}=+\\infty$. Sufficient conditions for this property are (i) $\\Phi(0)>0$ or (ii) $\\Phi$ is homogeneous of degree $\\hat{\\sigma} \\in(0,1)$ with $\\sigma+\\hat{\\sigma}<1$.\n\n[^3]:    ${ }^{4}$ For tractability, we will often study the case $g(x, 0)=0$. This can be viewed as a normalization that doesn't affect the economics of the problem.\n\n[^4]:    ${ }^{5}$ The parameter counts of leading models are estimated to be in the several-billion range; see https: //codingscape.com/blog/most-powerful-llms-large-language-models.\n    ${ }^{6}$ Recent progress has leveraged this channel; see, for instance, https://arcprize.org/blog/oai-o3-pub-breakthrough.\n\n[^5]:    ${ }^{7}$ Note that fine-tuning typically does not change the model's architecture or size and therefore should not directly affect marginal inference costs.\n\n[^6]:    ${ }^{8}$ In Section 5, we show that if $\\Phi(Z)$ is also a homogeneous function, then $C(Q)$ is simply a power function.\n\n[^7]:    ${ }^{9}$ Equivalently, we could assume that the buyer can combine tokens from two models on any task, with a performance given by the sum of individual gain functions.\n\n[^8]:    ${ }^{10}$ See Calzolari and Denicolò (2013) for a related analysis of nonlinear competition between a dominant firm and a competitive fringe under private information about demand.","text_sha256":"a87faabcf5086a74fa54b3e4708ee585016ce6f691a68e365d1085e40f378013"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0037","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"References","text":"[^9]:    ${ }^{11}$ Indeed, $\\underline{\\hat{\\theta}}=\\left[\\left(\\frac{1-\\sigma}{1-\\sigma-\\hat{\\sigma}_{L}}\\right)^{1-\\sigma-\\hat{\\sigma}_{L}}\\left(\\frac{\\sigma}{\\sigma+\\hat{\\sigma}_{L}}\\right)^{\\sigma+\\hat{\\sigma}_{L}}\\right]^{(1-\\sigma) / \\hat{\\sigma}_{L}}$, and thus, by the strict concavity of the logarithm and Jensen's inequality, $\\ln \\left(\\frac{\\hat{\\theta}}{\\underline{\\theta}}\\right)=\\frac{1-\\sigma}{\\hat{\\sigma}_{L}}\\left[\\left(1-\\sigma-\\hat{\\sigma}_{L}\\right) \\ln \\left(\\frac{1-\\sigma}{1-\\sigma-\\hat{\\sigma}_{L}}\\right)+\\left(\\sigma+\\hat{\\sigma}_{L}\\right) \\ln \\left(\\frac{\\sigma}{\\sigma+\\hat{\\sigma}_{L}}\\right)\\right]<0$.\n\n[^10]:    ${ }^{12}$ Calzolari and Denicolò (2015) show that a dominant firm with a competitive advantage can profitably impose exclusive dealing on privately informed buyers; the mechanism here is analogous, though exclusivity arises from the optimal nonlinear tariff rather than from contractual restrictions.\n\n[^11]:    ${ }^{13}$ For instance, a message to GPT-4o-mini costs 9 points, while a message to Claude Opus 4 costs 4,105 points, a ratio of roughly 450×.\n\n[^12]:    ${ }^{14}$ As in the theory, the overage charge interacts with the multiplier system: a query to a 1 × model costs \\$0.04 in the overage region, while a query to a $3 \\times$ model (Claude Opus 4.5) costs \\$0.12.\n\n[^13]:    ${ }^{15}$ By homogeneity, for all $x \\in \\mathbb{R}_{+}^{J}$ and $r>0, \\nabla \\Psi(r x)=r^{\\sigma-1} \\nabla \\Psi(x)$, so any token ray corresponds to a gradient ray. By strict concavity, if $x_{1} \\neq x_{2}$, then $\\nabla \\Psi\\left(x_{1}\\right) \\neq \\nabla \\Psi\\left(x_{2}\\right)$, so distinct token rays map to distinct gradient rays.\n\n[^14]:    ${ }^{16}$ In the general case of heterogeneous $\\sigma_{l}$, the optimal task split and the resulting buyer's payoff do not admit tractable closed-form solutions. In particular, the fine details of the profile $w$ could matter.\n\n[^15]:    ${ }^{17}$ Non-monotone partitions may also be optimal, but $\\theta_{l}\\left(I_{l}^{*}\\right)$ are uniquely determined.\n\n[^16]:    ${ }^{18}$ The Inada condition is imposed only for simplicity of arguments in Lemma 5.","text_sha256":"dc7812fef8b565ca82166cb8178dd3cb00bb4286f5019d549e1c67f475fa5f84"}
{"schema_version":"1.0","chunk_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10:0038","work_id":"alex-smolin:menu-pricing-of-large-language-models","paper_id":"alex-smolin:menu-pricing-of-large-language-models:2026-03-10","title":"Menu Pricing of Large Language Models","authors":[{"name":"Dirk Bergemann","url":"https://campuspress.yale.edu/dirkbergemann/"},{"name":"Alessandro Bonatti","url":"https://mitmgmtfaculty.mit.edu/abonatti/"},{"name":"Alex Smolin","url":"https://alexsmolin.com/","orcid":"https://orcid.org/0000-0003-4740-2376"}],"manuscript_date":"2026-03-10","language":"en","version_type":"working-paper","canonical_url":"https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md","source_record":"https://arxiv.org/abs/2502.07736","citation":"Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.","attribution_guidance":"Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.","provenance_url":"https://alexsmolin.com/corpus/PROVENANCE.txt","section":"Citation and provenance","text":"## Citation and provenance\n\n**Authors:** Dirk Bergemann; Alessandro Bonatti; Alex Smolin\n\n**Canonical citation:** Bergemann, Dirk, Alessandro Bonatti, and Alex Smolin. “Menu Pricing of Large Language Models.” Working paper, 2026.\n\n**Canonical machine-readable version:** https://alexsmolin.com/corpus/papers/menu-pricing-of-large-language-models.md\n\n**Source record:** https://arxiv.org/abs/2502.07736\n\n**Attribution guidance:** Preserve the supplied title, complete author list, citation, and canonical URL when technically practicable.\n\nProvenance metadata: https://alexsmolin.com/corpus/PROVENANCE.txt","text_sha256":"7588c67b11cf7d011ee922af374a51de915b585a0b55d962b7b8d12d0308069b"}
