Recursive partitioning for heterogeneous causal effects

doi:10.1073/PNAS.1510489113

Open AccessJournal ArticleDOI

Recursive partitioning for heterogeneous causal effects

Susan Athey, +1 more

- 05 Jul 2016 -

Proceedings of the National Academy of S...

- Vol. 113, Iss: 27, pp 7353-7360

Chats0

TLDR

This paper provides a data-driven approach to partition the data into subpopulations that differ in the magnitude of their treatment effects, and proposes an “honest” approach to estimation, whereby one sample is used to construct the partition and another to estimate treatment effects for each subpopulation.

Abstract:

In this paper we propose methods for estimating heterogeneity in causal effects in experimental and observational studies and for conducting hypothesis tests about the magnitude of differences in treatment effects across subsets of the population. We provide a data-driven approach to partition the data into subpopulations that differ in the magnitude of their treatment effects. The approach enables the construction of valid confidence intervals for treatment effects, even with many covariates relative to the sample size, and without “sparsity” assumptions. We propose an “honest” approach to estimation, whereby one sample is used to construct the partition and another to estimate treatment effects for each subpopulation. Our approach builds on regression tree methods, modified to optimize for goodness of fit in treatment effects and to account for honest estimation. Our model selection criterion anticipates that bias will be eliminated by honest estimation and also accounts for the effect of making additional splits on the variance of treatment effect estimates within each subpopulation. We address the challenge that the “ground truth” for a causal effect is not observed for any individual unit, so that standard approaches to cross-validation must be modified. Through a simulation study, we show that for our preferred method honest estimation results in nominal coverage for 90% confidence intervals, whereas coverage ranges between 74% and 84% for nonhonest approaches. Honest estimation requires estimating the model with a smaller sample size; the cost in terms of mean squared error of treatment effects for our preferred method ranges between 7–22%.

Recursive partitioning for heterogeneous causal effects

Citations

"Improving" prediction of human behavior using behavior modification

Using Machine Learning to Identify Heterogeneous Impacts of Agri-Environment Schemes in the EU: A Case Study

Can Consumer-Posted Photos Serve as a Leading Indicator of Restaurant Survival? Evidence from Yelp

Assessment of Heterogeneous Treatment Effect Estimation Accuracy via Matching

How Magic a Bullet Is Machine Learning for Credit Analysis? An Exploration with FinTech Lending Data

References

Random Forests

Regression Shrinkage and Selection via the Lasso

The Nature of Statistical Learning Theory

Statistical learning theory

The central role of the propensity score in observational studies for causal effects

Related Papers (5)

Estimation and Inference of Heterogeneous Treatment Effects using Random Forests

Estimating causal effects of treatments in randomized and nonrandomized studies.

The central role of the propensity score in observational studies for causal effects

Random Forests

Causality: models, reasoning, and inference