Uplift modeling and causal inference with machine learning algorithms


This project is stable and being incubated for long-term support. It may contain new experimental code, for which APIs are subject to change.

Causal ML: A Python Package for Uplift Modeling and Causal Inference with ML

Causal ML is a Python package that provides a suite of uplift modeling and causal inference methods using machine learning algorithms based on recent research. It provides a standard interface that allows user to estimate the Conditional Average Treatment Effect (CATE) or Individual Treatment Effect (ITE) from experimental or observational data. Essentially, it estimates the causal impact of intervention T on outcome Y for users with observed features X, without strong assumptions on the model form. Typical use cases include

  • Campaign targeting optimization: An important lever to increase ROI in an advertising campaign is to target the ad to the set of customers who will have a favorable response in a given KPI such as engagement or sales. CATE identifies these customers by estimating the effect of the KPI from ad exposure at the individual level from A/B experiment or historical observational data.

  • Personalized engagement: A company has multiple options to interact with its customers such as different product choices in up-sell or messaging channels for communications. One can use CATE to estimate the heterogeneous treatment effect for each customer and treatment option combination for an optimal personalized recommendation system.

The package currently supports the following methods

  • Tree-based algorithms
    • Uplift tree/random forests on KL divergence, Euclidean Distance, and Chi-Square
    • Uplift tree/random forests on Contextual Treatment Selection
  • Meta-learner algorithms
    • S-learner
    • T-learner
    • X-learner
    • R-learner
  • Instrumental variables algorithms
    • 2-Stage Least Squares (2SLS)



Install dependencies:

$ pip install -r requirements.txt

Install from pip:

$ pip install causalml

Install from source:

$ git clone https://github.com/uber/causalml.git
$ cd causalml
$ python setup.py build_ext --inplace
$ python setup.py install

Quick Start

Average Treatment Effect Estimation with S, T, X, and R Learners

from causalml.inference.meta import LRSRegressor
from causalml.inference.meta import XGBTRegressor, MLPTRegressor
from causalml.inference.meta import BaseXRegressor
from causalml.inference.meta import BaseRRegressor
from xgboost import XGBRegressor
from causalml.dataset import synthetic_data

y, X, treatment, _, _, e = synthetic_data(mode=1, n=1000, p=5, sigma=1.0)

lr = LRSRegressor()
te, lb, ub = lr.estimate_ate(X, treatment, y)
print('Average Treatment Effect (Linear Regression): {:.2f} ({:.2f}, {:.2f})'.format(te[0], lb[0], ub[0]))

xg = XGBTRegressor(random_state=42)
te, lb, ub = xg.estimate_ate(X, treatment, y)
print('Average Treatment Effect (XGBoost): {:.2f} ({:.2f}, {:.2f})'.format(te[0], lb[0], ub[0]))

nn = MLPTRegressor(hidden_layer_sizes=(10, 10),
te, lb, ub = nn.estimate_ate(X, treatment, y)
print('Average Treatment Effect (Neural Network (MLP)): {:.2f} ({:.2f}, {:.2f})'.format(te[0], lb[0], ub[0]))

xl = BaseXRegressor(learner=XGBRegressor(random_state=42))
te, lb, ub = xl.estimate_ate(X, e, treatment, y)
print('Average Treatment Effect (BaseXRegressor using XGBoost): {:.2f} ({:.2f}, {:.2f})'.format(te[0], lb[0], ub[0]))

rl = BaseRRegressor(learner=XGBRegressor(random_state=42))
te, lb, ub =  rl.estimate_ate(X=X, p=e, treatment=treatment, y=y)
print('Average Treatment Effect (BaseRRegressor using XGBoost): {:.2f} ({:.2f}, {:.2f})'.format(te[0], lb[0], ub[0]))

See the Meta-learner example notebook for details.

Interpretable Causal ML

Causal ML provides methods to interpret the treatment effect models trained as follows:

Meta Learner Feature Importances

from causalml.inference.meta import BaseSRegressor, BaseTRegressor, BaseXRegressor, BaseRRegressor
from causalml.dataset.regression import synthetic_data

# Load synthetic data
y, X, treatment, tau, b, e = synthetic_data(mode=1, n=10000, p=25, sigma=0.5)
w_multi = np.array(['treatment_A' if x==1 else 'control' for x in treatment]) # customize treatment/control names

slearner = BaseSRegressor(LGBMRegressor(), control_name='control')
slearner.estimate_ate(X, w_multi, y)
slearner_tau = slearner.fit_predict(X, w_multi, y)

model_tau_feature = RandomForestRegressor()  # specify model for model_tau_feature

slearner.get_importance(X=X, tau=slearner_tau, model_tau_feature=model_tau_feature,
                        normalize=True, method='auto', features=feature_names)

# Using the feature_importances_ method in the base learner (LGBMRegressor() in this example)
slearner.plot_importance(X=X, tau=slearner_tau, normalize=True, method='auto')

# Using eli5's PermutationImportance
slearner.plot_importance(X=X, tau=slearner_tau, normalize=True, method='permutation')

# Using SHAP
shap_slearner = slearner.get_shap_values(X=X, tau=slearner_tau)

# Plot shap values without specifying shap_dict
slearner.plot_shap_values(X=X, tau=slearner_tau)

# Plot shap values WITH specifying shap_dict
slearner.plot_shap_values(X=X, shap_dict=shap_slearner)

# interaction_idx set to 'auto' (searches for feature with greatest approximate interaction)

See the feature interpretations example notebook for details.

Uplift Tree Visualization

from IPython.display import Image
from causalml.inference.tree import UpliftTreeClassifier, UpliftRandomForestClassifier
from causalml.inference.tree import uplift_tree_string, uplift_tree_plot

uplift_model = UpliftTreeClassifier(max_depth=5, min_samples_leaf=200, min_samples_treatment=50,
                                    n_reg=100, evaluationFunction='KL', control_name='control')


graph = uplift_tree_plot(uplift_model.fitted_uplift_tree, features)

See the Uplift Tree visualization example notebook for details.


We welcome community contributors to the project. Before you start, please read our code of conduct and check out contributing guidelines first.



We document versions and changes in our changelog.


This project is licensed under the Apache 2.0 License - see the LICENSE file for details.




To cite CausalML in publications, you can refer to the following sources:

Whitepaper: CausalML: Python Package for Causal Machine Learning


@misc{chen2020causalml, title={CausalML: Python Package for Causal Machine Learning}, author={Huigang Chen and Totte Harinen and Jeong-Yoon Lee and Mike Yung and Zhenyu Zhao}, year={2020}, eprint={2002.11631}, archivePrefix={arXiv}, primaryClass={cs.CY} }


  • Nicholas J Radcliffe and Patrick D Surry. Real-world uplift modelling with significance based uplift trees. White Paper TR-2011-1, Stochastic Solutions, 2011.
  • Yan Zhao, Xiao Fang, and David Simchi-Levi. Uplift modeling with multiple treatments and general response types. Proceedings of the 2017 SIAM International Conference on Data Mining, SIAM, 2017.
  • Sören R. Künzel, Jasjeet S. Sekhon, Peter J. Bickel, and Bin Yu. Metalearners for estimating heterogeneous treatment effects using machine learning. Proceedings of the National Academy of Sciences, 2019.
  • Xinkun Nie and Stefan Wager. Quasi-Oracle Estimation of Heterogeneous Treatment Effects. Atlantic Causal Inference Conference, 2018.

Related projects

  • uplift: uplift models in R
  • grf: generalized random forests that include heterogeneous treatment effect estimation in R
  • rlearner: A R package that implements R-Learner
  • DoWhy: Causal inference in Python based on Judea Pearl's do-calculus
  • EconML: A Python package that implements heterogeneous treatment effect estimators from econometrics and machine learning methods
  • v0.13.0(Sep 2, 2022)

    • CausalML surpassed 1MM downloads on PyPI and 3,200 stars on GitHub. Thanks for choosing CausalML and supporting us on GitHub.
    • We have 7 new contributors @saiwing-yeung, @lixuan12315, @aldenrogers, @vincewu51, @AlkanSte, @enzoliao, and @alexander-pv. Thanks for your contributions!
    • @alexander-pv revamped CausalTreeRegressor and added CausalRandomForestRegressor with more seamless integration with scikit-learn's Cython tree module. He also added integration with shap for causal tree/ random forest interpretation. Please check out the example notebook.
    • We dropped the support for Python 3.6 and removed its test workflow.

    What's Changed

    • Fix typo (% -> $) by @saiwing-yeung in https://github.com/uber/causalml/pull/488
    • Add function for calculating PNS bounds by @t-tte in https://github.com/uber/causalml/pull/482
    • Fix hard coding bug by @t-tte in https://github.com/uber/causalml/pull/492
    • Update README of conda install and instruction of maintain in conda-forge by @ppstacy in https://github.com/uber/causalml/pull/485
    • Update examples.rst by @lixuan12315 in https://github.com/uber/causalml/pull/496
    • Fix incorrect effect_learner_objective in XGBRRegressor by @jeongyoonlee in https://github.com/uber/causalml/pull/504
    • Fix Filter F doesn't work with latest statsmodels' F test f-value format by @paullo0106 in https://github.com/uber/causalml/pull/505
    • Exclude tests in setup.py by @aldenrogers in https://github.com/uber/causalml/pull/508
    • Enabling higher orders feature importance for F filter and LR filter by @zhenyuz0500 in https://github.com/uber/causalml/pull/509
    • Ate pretrain 0506 by @vincewu51 in https://github.com/uber/causalml/pull/511
    • Update methodology.rst by @AlkanSte in https://github.com/uber/causalml/pull/518
    • Fix the bug of incorrect result in qini for multiple models by @enzoliao in https://github.com/uber/causalml/pull/520
    • Test get_qini() by @enzoliao in https://github.com/uber/causalml/pull/523
    • Fixed typo in uplift_trees_with_synthetic_data.ipynb by @jroessler in https://github.com/uber/causalml/pull/531
    • Remove Python 3.6 test from workflows by @jeongyoonlee in https://github.com/uber/causalml/pull/535
    • Causal trees update by @alexander-pv in https://github.com/uber/causalml/pull/522
    • Causal trees interpretation example by @alexander-pv in https://github.com/uber/causalml/pull/536

    New Contributors

    • @saiwing-yeung made their first contribution in https://github.com/uber/causalml/pull/488
    • @lixuan12315 made their first contribution in https://github.com/uber/causalml/pull/496
    • @aldenrogers made their first contribution in https://github.com/uber/causalml/pull/508
    • @vincewu51 made their first contribution in https://github.com/uber/causalml/pull/511
    • @AlkanSte made their first contribution in https://github.com/uber/causalml/pull/518
    • @enzoliao made their first contribution in https://github.com/uber/causalml/pull/520
    • @alexander-pv made their first contribution in https://github.com/uber/causalml/pull/522

    Full Changelog: https://github.com/uber/causalml/compare/v0.12.3...v0.13.0

    Source code(tar.gz)
    Source code(zip)
  • v0.12.3(Mar 14, 2022)

    This patch is to release a version without the constraint of Shap which can be used for conda-forge.

    What's Changed

    • Modify the requirement version of Shap by @ppstacy in https://github.com/uber/causalml/pull/483

    Full Changelog: https://github.com/uber/causalml/compare/v0.12.2...v0.12.3

    Source code(tar.gz)
    Source code(zip)
  • v0.12.2(Feb 18, 2022)

    This patch includes three updates by our latest contributors, @tonkolviktor and @heiderich. We also start using black, a Python formatter. Please check out the updated contribution guideline to learn how to use it.

    What's Changed

    • Opens up scipy dependency version range towards newer releases (#441) by @tonkolviktor in https://github.com/uber/causalml/pull/473
    • Merely define preferred backend for joblib instead of hard-coding it by @heiderich in https://github.com/uber/causalml/pull/476
    • Allow parallel prediction and make joblib's backend configurable for UpliftRandomForestClassifier by @heiderich in https://github.com/uber/causalml/pull/477
    • Reformat code using black by @jeongyoonlee in https://github.com/uber/causalml/pull/474

    New Contributors

    • @tonkolviktor made their first contribution in https://github.com/uber/causalml/pull/473
    • @heiderich made their first contribution in https://github.com/uber/causalml/pull/476

    Full Changelog: https://github.com/uber/causalml/compare/v0.12.1...v0.12.2

    Source code(tar.gz)
    Source code(zip)
  • v0.12.1(Feb 5, 2022)

    This patch includes two bug fixes for UpliftRandomForestClassifier as follows:

    • #462 by @paullo0106: Use the correct treatment_idx for fillTree() when applying validation data set
    • #468 by @jeongyoonlee: Switch the joblib backend for UpliftRandomForestClassifier to threading to avoid memory copy across trees
    Source code(tar.gz)
    Source code(zip)
  • v0.12.0(Jan 14, 2022)

    0.12.0 (Jan 2022)

    • CausalML surpassed 637K downloads on PyPI and 2,500 stars on Github!
    • We have 4 new community contributors, Luis (@lgmoneda ), Ravi (@raviksharma), Louis (@LouisHernandez17) and JackRab (@JackRab). Thanks for the contribution!
    • We refactored and speeded up UpliftTreeClassifier/UpliftRandomForestClassifier by 5x with Cython (#422 #440 by @jeongyoonlee)
    • We revamped our API documentation, it now includes the latest methodology, references, installation, notebook examples, and graphs! (#413 by @huigangchen @t-tte @zhenyuz0500 @jeongyoonlee @paullo0106)
    • Our team gave talks at 2021 Conference on Digital Experimentation @ MIT ([email protected]), Causal Data Science Meeting 2021, and KDD 2021 Tutorials on CausalML introduction and applications. Please take a look if you missed them! Full list of publications and talks can be found here.


    • Update documentation on Instrument Variable methods @huigangchen (#447)
    • Add benchmark simulation studies example notebook by @t-tte (#443)
    • Add sample_weight support for R-learner by @paullo0106 (#425)
    • Fix incorrect binning of numeric features in UpliftTreeClassifier by @jeongyoonlee (#420)
    • Update papers, talks, and publication info to README and refs.bib by @zhenyuz0500 (#410 #414 #433)
    • Add instruction for contributing.md doc by @jeongyoonlee (#408)
    • Fix incorrect feature importance calculation logic by @paullo0106 (#406)
    • Add parallel jobs support for NearestNeighbors search with n_jobs parameter by @paullo0106 (#389)
    • Fix bug in simulate_randomized_trial by @jroessler (#385)
    • Add GA pytest workflow by @ppstacy (#380)
    Source code(tar.gz)
    Source code(zip)
  • v0.11(Jul 29, 2021)

    0.11.0 (2021-07-28)

    (sorry for the spam, attempting to correctly update to the right files)

    • CausalML surpassed 2K stars!
    • We have 3 new community contributors, Jannik (@jroessler), Mohamed (@ibraaaa), and Leo (@lleiou). Thanks for the contribution!

    Major Updates

    • Make tensorflow dependency optional and add python 3.9 support by @jeongyoonlee (#343)
    • Add delta-delta-p (ddp) tree inference approach by @jroessler (#327)
    • Add conda env files for Python 3.6, 3.7, and 3.8 by @jeongyoonlee (#324)

    Minor Updates

    • Fix inconsistent feature importance calculation in uplift tree by @paullo0106 (#372)
    • Fix filter method failure with NaNs in the data issue by @manojbalaji1 (#367)
    • Add automatic package publish by @jeongyoonlee (#354)
    • Fix typo in unit_selection optimization by @jeongyoonlee (#347)
    • Fix docs build failure by @jeongyoonlee (#335)
    • Convert pandas inputs to numpy in S/T/R Learners by @jeongyoonlee (#333)
    • Require scikit-learn as a dependency of setup.py by @ibraaaa (#325)
    • Fix AttributeError when passing in Outcome and Effect learner to R-Learner by @paullo0106 (#320)
    • Fix error when there is no positive class for KL Divergence filter by @lleiou (#311)
    • Add versions to cython and numpy in setup.py for requirements.txt accordingly by @maccam912 (#306)
    Source code(tar.gz)
    Source code(zip)
  • v0.10.0(Feb 19, 2021)

    0.10.0 (2021-02-19)

    • CausalML surpassed 235,000 downloads!
    • We have 5 new community contributors, Suraj (@surajiyer), Harsh (@HarshCasper), Manoj (@manojbalaji1), Matthew (@maccam912) and Václav (@vaclavbelak). Thanks for the contribution!

    Major Updates

    • Add Policy learner, DR learner, DRIV learner by @huigangchen (#292)
    • Add wrapper for CEVAE, a deep latent-variable and variational autoencoder based model by @ppstacy (#276)

    Minor Updates

    • Add propensity_learner to R-learner by @jeongyoonlee (#297)
    • Add BaseLearner class for other meta-learners to inherit from without duplicated code by @jeongyoonlee (#295)
    • Fix installation issue for Shap>=0.38.1 by @paullo0106 (#287)
    • Fix import error for sklearn>= 0.24 by @jeongyoonlee (#283)
    • Fix KeyError issue in Filter method for certain dataset by @surajiyer (#281)
    • Fix inconsistent cumlift score calculation of multiple models by @vaclavbelak (#273)
    • Fix duplicate values handling in feature selection method by @manojbalaji1 (#271)
    • Fix the color spectrum of SHAP summary plot for feature interpretations of meta-learners by @paullo0106 (#269)
    • Add IIA and value optimization related documentation by @t-tte (#264)
    • Fix StratifiedKFold arguments for propensity score estimation by @paullo0106 (#262)
    • Refactor the code with string format argument and is to compare object types, and change methods not using bound instance to static methods by @harshcasper (#256, #260)
    Source code(tar.gz)
    Source code(zip)
  • v0.9.0(Oct 23, 2020)

    0.9.0 (2020-10-23)

    • CausalML won the 1st prize at the poster session in UberML'20
    • DoWhy integrated CausalML starting v0.4 (release note)
    • CausalML team welcomes new project leadership, Mert Bay
    • We have 4 new community contributors, Mario Wijaya (@mwijaya3), Harry Zhao (@deeplaunch), Christophe (@ccrndn) and Georg Walther (@waltherg). Thanks for the contribution!

    Major Updates

    • Add feature importance and its visualization to UpliftDecisionTrees and UpliftRF by @yungmsh (#220)
    • Add feature selection example with Filter methods by @paullo0106 (#223)

    Minor Updates

    • Implement propensity model abstraction for common interface by @waltherg (#223)
    • Fix bug in BaseSClassifier and BaseXClassifier by @yungmsh and @ppstacy (#217, #218)
    • Fix parentNodeSummary for UpliftDecisionTrees by @paullo0106 (#238)
    • Add pd.Series for propensity score condition check by @paullo0106 (#242)
    • Fix the uplift random forest prediction output by @ppstacy (#236)
    • Add functions and methods to init for optimization module by @mwijaya3 (#228)
    • Install GitHub Stale App to close inactive issues automatically @jeongyoonlee (#237)
    • Update documentation by @deeplaunch, @ccrndn, @ppstacy(#214, #231, #232)
    Source code(tar.gz)
    Source code(zip)
  • v0.8.0(Oct 21, 2020)

    0.8.0 (2020-07-17)

    CausalML surpassed 100,000 downloads! Thanks for the support.

    Major Updates

    • Add value optimization to optimize by @t-tte (#183)
    • Add counterfactual unit selection to optimize by @t-tte (#184)
    • Add sensitivity analysis to metrics by @ppstacy (#199, #212)
    • Add the iv estimator submodule and add 2SLS model to it by @huigangchen (#201)

    Minor Updates

    • Add GradientBoostedPropensityModel by @yungmsh (#193)
    • Add covariate balance visualization by @yluogit (#200)
    • Fix bug in the X learner propensity model by @ppstacy (#209)
    • Update package dependencies by @jeongyoonlee (#195, #197)
    • Update documentation by @jeongyoonlee, @ppstacy and @yluogit (#181, #202, #205)
    Source code(tar.gz)
    Source code(zip)
