Adaptive Sequential Experiments with Unknown Information Arrival Processes

Gur, Yonatan; Momeni, Ahmadreza

Adaptive Sequential Experiments with Unknown Information Arrival Processes

Yonatan Gur and Ahmadreza Momeni
Additional contact information
Yonatan Gur: Stanford U
Ahmadreza Momeni: Stanford U

Research Papers from Stanford University, Graduate School of Business

Abstract: Sequential experiments are often designed to strike a balance between maximizing immediate payoffs based on available information, and acquiring new information that is essential for maximizing future payoffs. This trade-off is captured by the multi-armed bandit (MAB) framework that has been studied and applied, typically when at each time epoch feedback is received only on the action that was selected at that epoch. However, in many practical settings, including product recommendations, dynamic pricing, retail management, and health care, additional information may become available between decision epochs. We introduce a generalized MAB formulation in which auxiliary information may appear arbitrarily over time. By obtaining matching lower and upper bounds, we characterize the minimax complexity of this family of problems as a function of the information arrival process, and study how salient characteristics of this process impact policy design and achievable performance. In terms of achieving optimal performance, we establish that: (i) upper confidence bound and posterior sampling policies possess natural robustness with respect to the information arrival process without any adjustments, which uncovers a novel property of these popular families of policies and further lends credence to their appeal; and (ii) policies with exogenous exploration rate do not possess such robustness. For such policies, we devise a novel virtual time indices method for dynamically controlling the effective exploration rate. We apply our method for designing Epsilon_{t}-greedy-type policies that, without any prior knowledge on the information arrival process, attain the best performance (in terms of regret rate) that is achievable when the information arrival process is a priori known. We use data from a large media site to analyze the value that may be captured in practice by leveraging auxiliary information for designing content recommendations.

Date: 2020-04
References: Add references at CitEc
Citations:

Downloads: (external link)
https://www.gsb.stanford.edu/gsb-cmis/gsb-cmis-download-auth/463336
Our link check indicates that this URL is bad, the error code is: 404 Not Found

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:ecl:stabus:3693

Access Statistics for this paper

More papers in Research Papers from Stanford University, Graduate School of Business Contact information at EDIRC.
Bibliographic data for series maintained by ().