Name: Bandits with Unobserved Confounders: A Causal Approach
Start: 2015-12-07T19:00:00-0500
End: 2015-12-07T23:59:00-0500

Back To Schedule

Bandits with Unobserved Confounders: A Causal Approach

The Multi-Armed Bandit problem constitutes an archetypal setting for sequential decision-making, permeating multiple domains including engineering, business, and medicine. One of the hallmarks of a bandit setting is the agent's capacity to explore its environment through active intervention, which contrasts with the ability to collect passive data by estimating associational relationships between actions and payouts. The existence of unobserved confounders, namely unmeasured variables affecting both the action and the outcome variables, implies that these two data-collection modes will in general not coincide. In this paper, we show that formalizing this distinction has conceptual and algorithmic implications to the bandit setting. The current generation of bandit algorithms implicitly try to maximize rewards based on estimation of the experimental distribution, which we show is not always the best strategy to pursue. Indeed, to achieve low regret in certain realistic classes of bandit problems (namely, in the face of unobserved confounders), both experimental and observational quantities are required by the rational agent. After this realization, we propose an optimization metric (employing both experimental and observational distributions) that bandit agents should pursue, and illustrate its benefits over traditional algorithms.

Speakers

Andrew Forney

Monday December 7, 2015 19:00 - 23:59 EST
210 C #60

Posters

NIPS 2015

Andrew Forney

Attendees (0)

NIPS 2015

Sign up or log in to save this to your schedule, view media, leave feedback and see who's attending!

Andrew Forney

Attendees (0)