\(\newcommand{\ci}{\perp\!\!\!\perp}\) \(\newcommand{\cA}{\mathcal{A}}\) \(\newcommand{\cB}{\mathcal{B}}\) \(\newcommand{\cC}{\mathcal{C}}\) \(\newcommand{\cD}{\mathcal{D}}\) \(\newcommand{\cE}{\mathcal{E}}\) \(\newcommand{\cF}{\mathcal{F}}\) \(\newcommand{\cG}{\mathcal{G}}\) \(\newcommand{\cH}{\mathcal{H}}\) \(\newcommand{\cI}{\mathcal{I}}\) \(\newcommand{\cJ}{\mathcal{J}}\) \(\newcommand{\cK}{\mathcal{K}}\) \(\newcommand{\cL}{\mathcal{L}}\) \(\newcommand{\cM}{\mathcal{M}}\) \(\newcommand{\cN}{\mathcal{N}}\) \(\newcommand{\cO}{\mathcal{O}}\) \(\newcommand{\cP}{\mathcal{P}}\) \(\newcommand{\cQ}{\mathcal{Q}}\) \(\newcommand{\cR}{\mathcal{R}}\) \(\newcommand{\cS}{\mathcal{S}}\) \(\newcommand{\cT}{\mathcal{T}}\) \(\newcommand{\cU}{\mathcal{U}}\) \(\newcommand{\cV}{\mathcal{V}}\) \(\newcommand{\cW}{\mathcal{W}}\) \(\newcommand{\cX}{\mathcal{X}}\) \(\newcommand{\cY}{\mathcal{Y}}\) \(\newcommand{\cZ}{\mathcal{Z}}\) \(\newcommand{\bA}{\mathbf{A}}\) \(\newcommand{\bB}{\mathbf{B}}\) \(\newcommand{\bC}{\mathbf{C}}\) \(\newcommand{\bD}{\mathbf{D}}\) \(\newcommand{\bE}{\mathbf{E}}\) \(\newcommand{\bF}{\mathbf{F}}\) \(\newcommand{\bG}{\mathbf{G}}\) \(\newcommand{\bH}{\mathbf{H}}\) \(\newcommand{\bI}{\mathbf{I}}\) \(\newcommand{\bJ}{\mathbf{J}}\) \(\newcommand{\bK}{\mathbf{K}}\) \(\newcommand{\bL}{\mathbf{L}}\) \(\newcommand{\bM}{\mathbf{M}}\) \(\newcommand{\bN}{\mathbf{N}}\) \(\newcommand{\bO}{\mathbf{O}}\) \(\newcommand{\bP}{\mathbf{P}}\) \(\newcommand{\bQ}{\mathbf{Q}}\) \(\newcommand{\bR}{\mathbf{R}}\) \(\newcommand{\bS}{\mathbf{S}}\) \(\newcommand{\bT}{\mathbf{T}}\) \(\newcommand{\bU}{\mathbf{U}}\) \(\newcommand{\bV}{\mathbf{V}}\) \(\newcommand{\bW}{\mathbf{W}}\) \(\newcommand{\bX}{\mathbf{X}}\) \(\newcommand{\bY}{\mathbf{Y}}\) \(\newcommand{\bZ}{\mathbf{Z}}\) \(\newcommand{\ba}{\mathbf{a}}\) \(\newcommand{\bb}{\mathbf{b}}\) \(\newcommand{\bc}{\mathbf{c}}\) \(\newcommand{\bd}{\mathbf{d}}\) \(\newcommand{\be}{\mathbf{e}}\) \(\newcommand{\bg}{\mathbf{g}}\) \(\newcommand{\bh}{\mathbf{h}}\) \(\newcommand{\bi}{\mathbf{i}}\) \(\newcommand{\bj}{\mathbf{j}}\) \(\newcommand{\bk}{\mathbf{k}}\) \(\newcommand{\bl}{\mathbf{l}}\) \(\newcommand{\bm}{\mathbf{m}}\) \(\newcommand{\bn}{\mathbf{n}}\) \(\newcommand{\bo}{\mathbf{o}}\) \(\newcommand{\bp}{\mathbf{p}}\) \(\newcommand{\bq}{\mathbf{q}}\) \(\newcommand{\br}{\mathbf{r}}\) \(\newcommand{\bs}{\mathbf{s}}\) \(\newcommand{\bt}{\mathbf{t}}\) \(\newcommand{\bu}{\mathbf{u}}\) \(\newcommand{\bv}{\mathbf{v}}\) \(\newcommand{\bw}{\mathbf{w}}\) \(\newcommand{\bx}{\mathbf{x}}\) \(\newcommand{\by}{\mathbf{y}}\) \(\newcommand{\bz}{\mathbf{z}}\) \(\newcommand{\RR}{\mathbb{R}}\) \(\newcommand{\NN}{\mathbb{N}}\) \(\newcommand{\balpha}{\boldsymbol{\alpha}}\) \(\newcommand{\bbeta}{\boldsymbol{\beta}}\) \(\newcommand{\btheta}{\boldsymbol{\theta}}\) \(\newcommand{\hpi}{\widehat{\pi}}\) \(\newcommand{\bpi}{\boldsymbol{\pi}}\) \(\newcommand{\hbpi}{\widehat{\boldsymbol{\pi}}}\) \(\newcommand{\bxi}{\boldsymbol{\xi}}\) \(\newcommand{\bmu}{\boldsymbol{\mu}}\) \(\newcommand{\bepsilon}{\boldsymbol{\epsilon}}\) \(\newcommand{\bzero}{\mathbf{0}}\) \(\newcommand{\T}{\text{T}}\) \(\newcommand{\Trace}{\text{Trace}}\) \(\newcommand{\Cov}{\text{Cov}}\) \(\newcommand{\Var}{\text{Var}}\) \(\newcommand{\E}{\mathbb{E}}\) \(\newcommand{\Pr}{\text{Pr}}\) \(\newcommand{\pr}{\text{pr}}\) \(\newcommand{\pdf}{\text{pdf}}\) \(\newcommand{\P}{\text{P}}\) \(\newcommand{\p}{\text{p}}\) \(\newcommand{\One}{\mathbf{1}}\) \(\newcommand{\argmin}{\operatorname*{arg\,min}}\) \(\newcommand{\argmax}{\operatorname*{arg\,max}}\) \(\newcommand{\dtheta}{\frac{\partial}{\partial\theta} }\) \(\newcommand{\ptheta}{\nabla_\theta}\) \(\newcommand{\alert}[1]{\color{darkorange}{#1}}\) \(\newcommand{\alertr}[1]{\color{red}{#1}}\) \(\newcommand{\alertb}[1]{\color{blue}{#1}}\)

1 Overview

We have now seen several examples of estimating the optimal treatment regime. In this section, our goal is to explore several issues and topics around this problem. Hopefully they can provide some guidelines for you to think about when you are working on your own project.

2 Outcome Weighted Learning

Outcome Weighted Learning (Zhao et al. 2012) is probably one of the most popular direct learning approaches. Formally, let’s be precise about our target of estimation. Following (Qian and Murphy 2011), let us define a quantity called the value function of a decision rule \(\pi(x)\). In fact, this is also a concept in reinforcement learning, and it is more common to use \(R\) (reward) to denote the outcome rather than \(Y\). Hence we will follow the notation in the literature. Let \(R(a)\) denote the potential reward under treatment \(a\). Under consistency, conditional exchangeability \(R(a) \perp A \mid X\), and positivity, the value function can be identified as follows:

\[ \begin{aligned} {\cV}(\pi) =& \E^\pi(R)\\ =& \E\left[R\{\pi(X)\}\right]\\ =& \E_X\big[ \E_R\{R \mid X, A = \pi(X)\} \big]. \end{aligned} \]

This means that we first (inner expectation \(E_R\)) let everyone take the treatment label suggested by the decision rule \(\pi\), i.e., \(A = \pi(X)\), and obtain the expected reward, and then average over the entire population (\(E_X\)). Hence, our goal is to maximize \({\cV}(\pi)\) by searching for the best \(\pi(X)\):

\[ \pi_\text{opt} = \underset{\pi}{\arg\max} \,\, {\cV}(\pi) \]

Note that this is the same as our previous definition of the optimal treatment regime: choosing the treatment with the larger conditional expected reward at each \(X=x\) also maximizes the population average reward. However, a tricky problem here is that we do not observe the distribution under which everyone follows the suggested \(\pi(X)\) treatment. Instead, our observational study is collected under its own mechanism, or distribution, denoted as \((R, X, A) \sim \mu\). This is called a behavior policy. On the other hand, our target is to estimate the value function under a restricted policy that gives \((R, X, A = \pi(X)) \sim \mu^\pi\). Let \(g(a \mid x) = \Pr(A=a \mid X=x)\) denote the behavior propensity. A tool we could utilize is the Radon-Nikodym theorem1, which suggests that

\[ \begin{aligned} \cV(\pi) =& \int R d\mu^\pi \\ =& \int R \frac{d\mu^\pi}{d\mu}d\mu \\ =& \int \frac{R \cdot \mathbf{1}\{A = \pi(X)\}}{g(A \mid X)} d\mu. \end{aligned} \]

Now, since \(d\mu\) is observed, we can use the sample (empirical) estimation of this value function.

\[ \widehat{\pi}_\text{opt} = \underset{\pi}{\arg\max} \,\, \frac{1}{n} \sum_{i=1}^n \frac{R_i \cdot \mathbf{1}\{A_i = \pi(X_i)\} }{\widehat g(A_i \mid X_i)}. \]

We can simply change the equal sign to unequal sign and this would become a minimization problem:

\[ \widehat{\pi}_\text{opt} = \underset{\pi}{\arg\min} \frac{1}{n} \sum_{i=1}^n \frac{R_i \cdot \mathbf{1}\{A_i \neq \pi(X_i)\} }{\widehat g(A_i \mid X_i)}. \]

The interesting view of this formulation is that, this becomes a classification problem, where

  • \(\pi(X_i)\) is the decision rule
  • \(A_i\) is the observed class label
  • \(\frac{R_i}{\widehat g(A_i \mid X_i)}\) is the subject-specific weight2.

Hence, our goal is to fit a weighted classification model. We want to encourage the model to do a better job for those with high rewards (large weights). For the binary-treatment implementation below, we encode the treatment as \(A \in \{-1,+1\}\). The following code visually demonstrates how this estimator works.

  set.seed(1)
  n = 800
  p = 2
  x1 <- runif(n, 0, 1)
  x2 <- runif(n, 0, 1)
  a = rbinom(n, 1, 0.5)*2-1
  side = sign(x2 - sin(2*pi*x1)/3-0.5)
  
  R <- rnorm(n, mean = ifelse(side==a, 1.5, 0.5), sd = 0.5)

  # the observed data, with class labels
  par(mar=rep(1,4))
  plot(x1, x2, pch = 19, xaxt = "n", yaxt = "n", xlab = "", ylab = "", 
       col = ifelse(a==1, "deepskyblue", "darkorange"))
  legend("topright", c("Treatment = 1", "Treatment = -1"), pch = 19, 
         col = c("deepskyblue", "darkorange"), cex = 1.3)

Now let us add the reward as weights, indicated by the point size. It is clear that on the lower side, treatment \(-1\) is more dominant, while on the upper side, treatment \(1\) is more dominant. The true decision boundary is also given below.

  par(mfrow = c(1, 2))
  par(mar=rep(1,4))
  
  plot(x1, x2, pch = 19, xaxt = "n", yaxt = "n", xlab = "", ylab = "", 
       cex = (R+2)/3, col = ifelse(a==1, "deepskyblue", "darkorange"))
  legend("topright", c("Treatment = 1", "Treatment = -1"), pch = 19, 
         col = c("deepskyblue", "darkorange"), cex = 1.3)
  
  plot(x1, x2, pch = 19, xaxt = "n", yaxt = "n", xlab = "", ylab = "", 
       cex = (R+2)/3, col = ifelse(a==1, "deepskyblue", "darkorange"))
  legend("topright", c("Treatment = 1", "Treatment = -1"), pch = 19, 
         col = c("deepskyblue", "darkorange"), cex = 1.3)
  x1_grid = sort(x1)
  lines(x1_grid, sin(2*pi*x1_grid)/3+0.5, lwd = 3)

2.1 Choosing the Functional Class

The outcome weighted learning framework permits a wide range of functional classes for estimating \(\pi\). For example, we can use a simple linear model, a function of reproducing kernel Hilbert space, a decision tree, a random forest, or a neural network. But also be aware that the functional class cannot be too flexible, otherwise it would exactly fit the observed data with \(\widehat\pi(X_i) = A_i\). The following code demonstrates the idea using a simple linear model. This can fit reasonably well, but its straight decision boundary cannot represent the curved boundary in this example.

  mydata = data.frame("x1" = x1, "x2" = x2, "R" = R, "a" = a)

  # estimate Pr(A = 1 | X), although we know this is a randomized trial
  pscore = glm(as.factor(a) ~ x1 + x2, data = mydata, family = "binomial")$fitted.values

  # convert to the probability of each subject's observed treatment
  pscore.observed = ifelse(a == 1, pscore, 1 - pscore)

  # add a constant to the reward to make the classification weights positive
  W = (mydata$R + 2)/pscore.observed
  
  # fit a weighted classification model to estimate the OTR
  fit = glm(as.factor(a) ~ x1 + x2, data = mydata, family = "binomial", weights = W)

  # the decision rule
  table(fit$fitted.values > 0.5, side)
##        side
##          -1   1
##   FALSE 370  60
##   TRUE   48 322
  mean((fit$fitted.values > 0.5) == (side == 1))
## [1] 0.865

Since the linear model does not perform well, we now use a more flexible functional class based on the Reproducing Kernel Hilbert Space (RKHS). Let \(f \in \mathcal{H}\) be a real-valued function, and define the decision rule as \(\pi(x) = \mathrm{sign}\{f(x)\}\). A common approach is to use the weighted support vector machine (SVM) framework, which replaces the nonconvex misclassification loss \(1\{A_i \neq \pi(X_i)\}\) with the convex hinge loss (see this note). The optimization problem becomes

\[ \widehat{f} = \arg\min_{f \in \mathcal{H}} \frac{1}{n} \sum_{i=1}^n \underbrace{\frac{R_i^{+}}{\widehat g(A_i \mid X_i)}}_{\text{subject weight}} \, \big[1 - A_i\, f(X_i)\big]_+ + \lambda \| f \|_{\mathcal{H}}^2, \qquad \widehat{\pi}(x) = \mathrm{sign}\{\widehat{f}(x)\}, \]

where \(A_i \in \{-1, +1\}\), \([u]_+ = \max(u, 0)\), and \(R_i^{+} = R_i + c\) for any constant \(c\) chosen so that \(R_i^{+} \ge 0\). The term \(\| f \|_{\mathcal{H}}^2\) controls the smoothness or complexity of the function.
By the Representer Theorem, the solution takes the form

\[ \widehat{f}(x) = \sum_{i=1}^n \alpha_i K(x, X_i), \]

where \(K(\cdot, \cdot)\) is a kernel function and \(\bK\) is the corresponding kernel matrix with entries \(K(X_i, X_j)\). The weighted SVM can be efficiently implemented using the WeightSVM package by specifying the subject-specific weights. The following code demonstrates this idea using a radial basis function (RBF) kernel.

  library(WeightSVM)
  owl.fit = wsvm(a ~ x1 + x2, data = mydata, 
                 kernel = "radial", gamma = 0.15,
                 weight = W)
  
  table(owl.fit$fitted > 0, side)
##        side
##          -1   1
##   FALSE 368  46
##   TRUE   50 336
  mean((owl.fit$fitted > 0) == (side == 1))
## [1] 0.88
  
  par(mfrow = c(1, 1))
  par(mar=rep(1,4))
  plot(x1, x2, pch = 19, xaxt = "n", yaxt = "n", xlab = "", ylab = "", 
       col = ifelse(owl.fit$fitted > 0, "deepskyblue", "darkorange"))
  legend("topright", c("Treatment = 1", "Treatment = -1"), pch = 19, 
         col = c("deepskyblue", "darkorange"), cex = 1.3)

2.2 Example: lalonde Data

Let’s use the lalonde data again. Recall that the outcome is re78, the treatment treat is whether the subject received a job training program. For this example, we will use the DTRlearn2 package.

  library(Matching)
## Loading required package: MASS
## ## 
## ##  Matching (Version 4.10-15, Build Date: 2024-10-14)
## ##  See https://www.jsekhon.com for additional documentation.
## ##  Please cite software as:
## ##   Jasjeet S. Sekhon. 2011. ``Multivariate and Propensity Score Matching
## ##   Software with Automated Balance Optimization: The Matching package for R.''
## ##   Journal of Statistical Software, 42(7): 1-52. 
## ##
  data(lalonde)
  head(lalonde)
  
  library(DTRlearn2)
## Loading required package: kernlab
## Loading required package: Matrix
## Loading required package: foreach
## Loading required package: glmnet
## Loaded glmnet 5.0
  X = data.frame("age" = scale(lalonde$age),
                 "educ" = scale(lalonde$educ),
                 "black" = lalonde$black,
                 "hisp" = lalonde$hisp,
                 "married" = lalonde$married,
                 "nodegr" = lalonde$nodegr,
                 "re74" = lalonde$re74/max(lalonde$re74),
                 "re75" = lalonde$re75/max(lalonde$re75))
  A = 2*(as.numeric(lalonde$treat) - 0.5)
  R = lalonde$re78
  
  # fit the OWL model
  set.seed(1)
  OWL.fit = owl(H = data.matrix(X), AA = A, RR = R,
            n = nrow(X),
            K = 1, # for our problem, it is only one stage of treatment so we use K = 1
            kernel = "rbf", # radial basis function kernel
            sigma = 0.5) # bandwidth for each covariate
  
  # the estimated value function
  OWL.fit$valuefun
## [1] 7031.391

The package reports an estimated value function for the learned rule. Because the same data are used to learn and evaluate the rule here, this is an in-sample value. It should not by itself be used to conclude that individualization improves outcomes relative to a one-size-fits-all rule. Such a comparison requires held-out or cross-fitted policy evaluation, together with an assessment of uncertainty. The doubly robust idea from the ATE lecture can also be used to evaluate a rule on held-out data.

3 Continuous Treatment with Stochastic Policy View

In the previous sections, we have mainly focused on binary treatment. However, in many cases, the treatment can be continuous. The most common example is the dosage of a drug. In this case, we need to modify the methods we have learned to accommodate continuous treatment. There are several simple approaches. For example, we can discretize the treatment into several levels and then apply the methods we have learned. However, this would greatly reduce the information we have for each treatment level. We can also use a linear regression and add interaction terms, possibly even second-order terms of \(A\), between the treatment and the covariates. When we need to choose the best treatment, we can either solve for the optimal treatment directly or use a grid search. However, a linear model can be biased if its functional form is misspecified. Hence some dedicated methods are needed. A natural idea is to utilize kernel smoothing to modify the outcome weighted learning approach. This leads to the method proposed in Chen et al. (2016).

Let’s recall that the outcome weighted learning approach for binary/multi-categorical treatment is to solve the following optimization problem:

\[ \widehat{\pi}_\text{opt} = \underset{\pi}{\arg\max} \frac{1}{n} \sum_{i=1}^n \frac{R_i \cdot \mathbf{1}\{A_i = \pi(X_i)\} }{\widehat g(A_i \mid X_i)} \]

When the treatment is continuous, we need to modify the indicator function since the probability that \(A_i = \pi(X_i)\) is zero. It might be helpful to revisit our definition of the value function and the Radon-Nikodym theorem, which converted the observed behavior-policy distribution to the target-policy distribution. A deterministic target policy with \(A=\pi(X)\) is degenerate relative to a continuous behavior density. A natural alternative is to let the target treatment follow a distribution centered at \(\pi(X)\). This is connected with the reinforcement learning literature, where we use a stochastic target policy3. For example, let the stochastic target policy be Uniform on \([\pi(X) - h, \pi(X) + h]\), where \(h\) is a bandwidth. Its conditional density is

\[ q_{\pi,h}(a \mid x) = \frac{1}{2h} \cdot \mathbf{1}\left\{|a-\pi(x)| \le h\right\}. \]

Then if we revisit the definition of the value function, we would have

\[ \begin{aligned} \cV_h(\pi) =& \int R d\mu^{\pi,h} \\ =& \int R \frac{d\mu^{\pi,h}}{d\mu}d\mu \\ =& \int \frac{R \cdot q_{\pi,h}(A \mid X)}{g(A \mid X)} d\mu \\ =& \int \frac{R \cdot \mathbf{1}\{|A-\pi(X)| \le h\}}{2h \cdot g(A \mid X)} d\mu. \\ \end{aligned} \]

This leads to the sample version of the estimator, since we observe samples from \(\mu\):

\[ \widehat{\pi}_{h,\text{opt}} = \underset{\pi}{\arg\max} \frac{1}{n} \sum_{i=1}^n \frac{R_i \cdot \mathbf{1}\left\{|A_i-\pi(X_i)| \le h\right\} }{2h\,\widehat g(A_i \mid X_i)}. \]

This permits a variety of choices for the kernel density. Instead of using a Uniform kernel, we can in general write

\[ \widehat{\pi}_{h,\text{opt}} = \underset{\pi}{\arg\max} \frac{1}{n} \sum_{i=1}^n \frac{R_i \cdot K_h\{A_i-\pi(X_i)\} }{\widehat g(A_i \mid X_i)}. \]

where \(K_h(u) = K(u/h)/h\) is a kernel density. If \(K\) is nonnegative and integrates to one, this corresponds to the stochastic target-policy density \(q_{\pi,h}(a \mid x) = K_h\{a-\pi(x)\}\). For the Uniform-window objective, Chen et al. (2016) replace the discontinuous disagreement indicator with a truncated absolute surrogate. This surrogate can be written as a difference of convex functions and optimized with the DC algorithm (Tao and An 1998)4.

4 Variance Reduction of OWL

We now return to binary treatment rules. The variance-reduction idea below applies regardless of the class of decision rules used to fit OWL.

One potential issue with OWL is that it may be sensitive to the scale or location of the outcome weight \(R\). Subjects with much larger weights can dominate the learning process. To mitigate this issue, we notice that, at the population level with the correct behavior propensity, subtracting any function of \(X\) from \(R\) does not change the maximizer of the problem:

\[ \begin{aligned} & \quad \,\, \E \left[ \frac{ \left\{ R - m(X) \right\} \mathbf{1}\{A = \pi(X)\} }{g(A \mid X)} \right] \\ &= \E \left[ \frac{ R \mathbf{1}\{A = \pi(X)\} }{g(A \mid X)} \right] - \E \left[ \frac{ m(X) \mathbf{1}\{A = \pi(X)\} }{g(A \mid X)} \right] \\ &= \cV(\pi) - \E_X\left[ m(X) \E_A \left\{ \frac{\mathbf{1}\{A = \pi(X)\}}{g(A \mid X)} \Biggm| X \right\} \right] \\ &= \cV(\pi) - \E_X\left[ m(X) \sum_{a \in \mathcal A} \frac{\mathbf{1}\{\pi(X)=a\}g(a \mid X)}{g(a \mid X)} \right] \\ &= \cV(\pi) - \E_X\left[ m(X) \sum_{a \in \mathcal A}\mathbf{1}\{\pi(X)=a\} \right] \\ &= \cV(\pi) - \E_X\left[ m(X) \right]. \end{aligned} \]

Hence, this quantity does not depend on the choice of the decision rule \(\pi\). A well-chosen \(m(X)\) can reduce the variability of the weights, although an arbitrary or poorly estimated \(m(X)\) need not reduce variance. A convenient choice used in Zhou et al. (2017) is

\[ m(X) = \E\left[ \frac{R}{2g(A \mid X)} \Biggm| X \right] = \frac{\E[R \mid X, A = 1] + \E[R \mid X, A = 0]}{2}. \]

A second useful choice, used in Zhu et al. (2017), is:

\[ m(X) = \E[ R \mid X] = \E[ R \mid X, A = 1] g(1 \mid X) + \E[ R \mid X, A = 0] g(0 \mid X). \]

Both choices are intended to reduce variability. The realized reduction depends on how well \(m(X)\) is estimated.


Chen, Guanhua, Donglin Zeng, and Michael R Kosorok. 2016. “Personalized Dose Finding Using Outcome Weighted Learning.” Journal of the American Statistical Association 111 (516): 1509–21.
Qian, Min, and Susan A Murphy. 2011. “Performance Guarantees for Individualized Treatment Rules.” Annals of Statistics 39 (2): 1180.
Tao, Pham Dinh, and Le Thi Hoai An. 1998. “A DC Optimization Algorithm for Solving the Trust-Region Subproblem.” SIAM Journal on Optimization 8 (2): 476–505.
Zhao, Yingqi, Donglin Zeng, A John Rush, and Michael R Kosorok. 2012. “Estimating Individualized Treatment Rules Using Outcome Weighted Learning.” Journal of the American Statistical Association 107 (499): 1106–18.
Zhou, Xin, Nicole Mayer-Hamblett, Umer Khan, and Michael R Kosorok. 2017. “Residual Weighted Learning for Estimating Individualized Treatment Rules.” Journal of the American Statistical Association 112 (517): 169–87.
Zhu, Ruoqing, Ying-Qi Zhao, Guanhua Chen, Shuangge Ma, and Hongyu Zhao. 2017. “Greedy Outcome Weighted Tree Learning of Optimal Personalized Treatment Rules.” Biometrics 73 (2): 391–400.

  1. Let \((\Omega,\mathcal{F})\) be a measurable space and let \(\nu\) and \(\mu\) be \(\sigma\)-finite measures with \(\nu \ll \mu\) (i.e., whenever \(\mu(E)=0\) then \(\nu(E)=0\)). The RN theorem guarantees the existence of a measurable function \(\frac{d\nu}{d\mu}:\Omega\to[0,\infty)\) (the RN derivative, or “density” of \(\nu\) w.r.t. \(\mu\)) such that, for every integrable \(f\), \[ \int f\,d\nu \;=\; \int f\,\frac{d\nu}{d\mu}\,d\mu. \] In our notation, \(\mu\) is the observed/behavior data-generating distribution of \((R,X,A)\) and \(\mu^\pi\) is the (counterfactual) distribution induced by the target policy \(\pi\). Under positivity (\(\mu^\pi \ll \mu\)), the value can be rewritten as \[ \mathcal V(\pi) \;=\; \int R\, d\mu^\pi \;=\; \int R\,\frac{d\mu^\pi}{d\mu}\, d\mu. \] When actions are discrete and we factor measures by the behavior propensity \(g(a \mid x)\), the RN derivative reduces to the importance weight: \[ \frac{d\mu^\pi}{d\mu}(R,X,A) \;=\; \frac{\Pr_\pi(A\mid X)}{g(A\mid X)} \quad\Longrightarrow\quad \mathcal V(\pi) \;=\; \mathbb E\!\left[\,R\,\frac{\Pr_\pi(A\mid X)}{g(A\mid X)}\,\right]. \]↩︎

  2. In some cases, the observed outcome can be negative. At the population level, the optimal treatment rule is invariant to a location shift of \(R\). With the correct behavior propensity, we can verify that \(\E\left[ \frac{\alpha \cdot \mathbf{1}\{A \neq \pi(X)\}}{g(A \mid X)}\right] = \alpha\) for any constant \(\alpha\). Hence, a common implementation adds a constant to all \(R_i\) to make the classification weights positive. In a finite sample, especially with estimated propensities and regularization, the fitted rule can still be affected by this shift.↩︎

  3. Note that in Chen et al. (2016), the authors started from a different perspective by replacing exact agreement with a small tolerance while still viewing the problem as a deterministic policy with \(A = \pi(X)\). The resulting computational framework also aligns with this stochastic-policy view.↩︎

  4. The Difference of Convex (DC) algorithm is a general optimization framework for problems whose objective can be expressed as the difference between two convex functions, i.e., \(f(\theta) = g(\theta) - h(\theta)\) where both \(g\) and \(h\) are convex. The key idea is to linearize the concave part \(-h(\theta)\) at the current iterate \(\theta^{(t)}\) and solve the resulting convex subproblem iteratively: \[ \theta^{(t+1)} = \arg\min_{\theta} \; g(\theta) - \langle s^{(t)}, \theta \rangle, \qquad s^{(t)}\in\partial h(\theta^{(t)}). \] This yields a sequence of convex optimizations that decreases the original nonconvex objective and seeks a local minimum. In the formulation of Chen et al. (2016), the Uniform-window disagreement indicator is replaced by the truncated absolute surrogate loss \[ L(u) = \min(|u|, 1), \qquad u = \frac{a - \pi(x)}{h}. \] This loss is zero at exact agreement, increases linearly within the bandwidth, and remains at one for larger deviations. It upper-bounds the outside-window disagreement indicator \(\mathbf{1}\{|u|>1\}\). Using the identity \[ \min(|u|,1)=|u|-(|u|-1)_+, \] the loss can be expressed as the difference of two convex functions. Hence, the empirical risk in continuous-treatment OWL can be optimized using the DC algorithm.↩︎