Off-policy Evaluation with General Logging Policies:Implementation at Mercari(22.10月)NARITA Yusuke. RIETI

Off-policy evaluation (OPE) attempts to predict the performance of counterfactual policies using log data from a different policy. We extend its applicability by developing an OPE method for a class of both full support and deficient support logging policies in contextual-bandit settings. This class includes deterministic bandit (such as Upper Confidence Bound) as well as deterministic decision-making based on supervised and unsupervised learning. We prove that our method's prediction converges in probability to the true performance of a counterfactual policy as the sample size increases. We validate our method with experiments on partly and entirely deterministic logging policies. Finally, we apply it to evaluate coupon targeting policies by a major online platform and show how to improve the existing policy.

公共政策からビジネスまで、アルゴリズムを利用した意思決定が広がっている。その際に重要なのが、過去に使用された方策(意思決定アルゴリズム)が蓄積したデータを用いて、過去に使われたことのない新しい方策の性能を予測することだ。「方策外評価」などと呼ばれるこの予測によって、データに基づいて意思決定・資源配分アルゴリズム・メカニズムを設計していくことが可能になる。本論文では、従来の手法では分析することの難しかった、より広いクラスの方策が生成したデータに適用可能な方策外評価手法を開発する。そして、提案手法をフリマアプリ・メルカリにおけるクーポン割当方策の評価に適用し、既存の方策を改善する方法を示す。

22e097.pdf

1.92MB

'정책칼럼' 카테고리의 다른 글

The scarring effects of deep contractions(22-10-3)/David Aikman.BIS (0)	2022.10.11
데이터 변조방지에 기초한 디지털 화폐 지갑 연구동향 - 익명성과 투명성의 양립을 위해 (In Japanese)/오오츠카 아키라.BOJ (0)	2022.10.11
The Impact of COVID-19 on Global Inequality and Poverty(22-10-5)/Daniel Gerszon Mahler.World Bank (0)	2022.10.11
Tax Incentives and the Global MinimumCorporate Tax-Reconsidering Tax Incentives after the GloBE Rules(22-10-6)/OECD (0)	2022.10.11
Progress Report on the Administration and Tax Certainty Aspects of Amount A of Pillar One - Two-Pillar Solution to the Tax Challenges of the Digitalisation of the Economy(22.10月)/OECD (0)	2022.10.11

yknet

Off-policy Evaluation with General Logging Policies:Implementation at Mercari(22.10月)NARITA Yusuke. RIETI

'정책칼럼' 카테고리의 다른 글

티스토리툴바

Off-policy Evaluation with General Logging Policies:Implementation at Mercari(22.10月)NARITA Yusuke. RIETI

'정책칼럼' 카테고리의 다른 글

'정책칼럼' Related Articles

티스토리툴바