23  Lazy evaluation

You write filter(penguins, species == "Adelie") and it works. But species isn’t a variable in your environment; it’s a column buried inside a data frame, invisible to ordinary evaluation rules. So how does R find it?

The answer starts with what R does with every argument you pass to a function, and a two-line experiment shows it. The same habit is behind default arguments that refer to other arguments, functions that read your code before running it, and the tidyverse convention of writing column names as if they were variables. Full metaprogramming lives in Chapter 26; here, the goal is understanding the machinery well enough to use it, and well enough to know when it’s working against you.

23.1 Promises

Give a function an argument that announces when it is evaluated:

f <- function(x) {
  cat("inside f\n")
  x
}

f({ cat("evaluating the argument\n"); 5 })
#> inside f
#> evaluating the argument
#> [1] 5

The function body started running before its argument had been evaluated. When you call f(x + 1), R does not compute x + 1 and pass the result. It bundles the expression x + 1 together with the environment where that expression should be evaluated (your calling environment), and that bundle, a promise, travels into f unevaluated.

The first time f actually touches the argument, R opens the promise: it takes the stored expression, evaluates it in the stored environment, and caches the result. If the argument is never used, the promise is never evaluated, and side effects inside it never happen. If the argument is used twice, the second access returns the cached value. One evaluation, at most. This sounds like a harmless optimization, but it changes what it means for a function to “receive” an argument.

Promises are almost invisible from R code. Base R has no is.promise(), str() on an argument forces the promise it was asked to inspect, and the only way to build one by hand is delayedAssign(). The invisibility is deliberate: inspecting a promise’s value means evaluating it, which changes program behavior, so the abstraction holds only as long as you cannot look behind it.

Early programming languages split over when arguments should be evaluated. Fortran and C computed each argument before handing it to the function, which is call-by-value. Algol 60 tried the opposite, passing the raw expression and re-evaluating it every time the function touched it, call-by-name. If the expression had side effects, they fired on every access; if the expression was expensive, the cost multiplied by the number of times the function used it. The compromise came in 1976: pass the expression unevaluated, but cache the result after the first evaluation. In the terminology of lambda calculus this is call-by-need, and the technique became known as lazy evaluation. Haskell made laziness the default for everything when it launched in 1990. R inherited from S a narrower design in which function arguments are lazy and everything else is eager, which gives you expression capture without pervasive laziness.

The Church-Rosser theorem guarantees that if a reduction terminates, every evaluation order reaches the same normal form; the orders differ only in whether they terminate at all. The Y combinator in Section 22.8 is a case where call-by-need terminates and call-by-value does not.

What does laziness buy you in practice? Default arguments that would be impossible under eager evaluation:

f <- function(x, y = x * 2) y
f(5)
#> [1] 10

The default for y is a promise containing the expression x * 2. When R needs y’s value, it evaluates x * 2 inside the function body, where x is already bound to 5, yielding 10. Supply y explicitly and the default promise is never touched:

f(5, 99)
#> [1] 99

Defaults can depend on computations that happen inside the function itself, which gets stranger:

g <- function(x, n = length(x)) n
g(c(1, 2, 3))
#> [1] 3

The default n = length(x) is an expression, not a precomputed value; R evaluates it when n is first accessed, by which time x is already bound. If R evaluated defaults at the moment of the call, before the function body had a chance to run, none of this would work. But what happens when laziness interacts with side effects?

Exercises

  1. Predict the output of this code, then run it:

    h <- function(x) {
      cat("first use\n")
      x
      cat("second use\n")
      x
    }
    h({ cat("evaluating argument\n"); 5 })

    How many times does “evaluating argument” print? Why?

  2. Write a function f(x, y = x + 1) and call f(10). What is y? Now call f(10, 50). What changed?

  3. What happens if you call a function that never uses its argument?

    ignore <- function(x) 42
    ignore(stop("this should error"))

    Does it error? Why or why not?

23.2 Consequences of laziness

A function that ignores its argument never forces the promise:

quiet <- function(x) "I ignore my argument"
quiet(print("you will never see this"))
#> [1] "I ignore my argument"

No output from print(). The promise sat there, inert, and was eventually garbage-collected without ever running. That is evaluate-only-when-needed taken literally, and it means any side effect buried in an argument expression (printing, writing a file, raising an error) happens only if the function actually touches the argument.

The same window before evaluation is what lets a function tell whether an argument was supplied at all:

report <- function(x) {
  if (missing(x)) "not supplied" else "supplied"
}
report()
#> [1] "not supplied"
report(42)
#> [1] "supplied"

If every argument were evaluated before the function saw it, report() would have failed before its body ran. missing() reads the promise’s state without forcing it, and it stays reliable until the function assigns to the argument, after which it reports FALSE.

The default n = length(x) you saw earlier is the same mechanism from the other side: a default is a stored expression, so it can depend on other arguments, on computations earlier in the function body, or on variables in the enclosing environment.

The lazy evaluation trap from Section 20.2 is also a promise. A factory that never touches its argument hands the promise, still unevaluated, to the function it returns:

make <- function(i) function() i

funs <- list()
for (k in 1:3) funs[[k]] <- make(k)
funs[[1]]()
#> [1] 3

make(k) received k as a promise, nothing in make used it, and by the time funs[[1]]() finally forces it the loop is over and k is 3. Forcing the promise inside the factory pins the value:

make <- function(i) {
  force(i)
  function() i
}

for (k in 1:3) funs[[k]] <- make(k)
funs[[1]]()
#> [1] 1

What is force(x) actually doing? Nothing clever. Its definition is just x, which accesses the argument and thereby triggers evaluation and caching. The name exists purely to communicate intent: “evaluate this now, don’t wait.” Which raises a broader question: if laziness can cause subtle bugs in closures, why does R use it at all?

Exercises

  1. Predict: does quiet(log(-1)) produce a warning? Why or why not? (Use the quiet function defined above.)

  2. Write a function greet(name = "stranger") that returns paste("Hello,", name). Call it with and without an argument. Explain which default mechanism makes this work.

23.3 What is non-standard evaluation?

Because arguments travel as unevaluated promises, a function can look at the expression before R reduces it to a value. Ordinarily it does not. When you write mean(c(1, 2, 3)), R computes the vector [1] 1 2 3 first, then passes it to mean, and mean never sees the expression c(1, 2, 3) at all.

Now try subset(df, x > 3). If R evaluated x > 3 in your environment before passing it to subset, the call would fail: there is no x in your workspace. Instead, subset reaches into the promise, extracts the expression before evaluation, and decides for itself where that expression should run:

df <- data.frame(x = 1:5, y = c(10, 20, 30, 40, 50))
subset(df, x > 3)
#>   x  y
#> 4 4 40
#> 5 5 50

x here is a column of df, not a variable in your global environment. subset never evaluates x > 3 in your environment. It captures the expression with substitute(), then evaluates it inside df with eval():

# Simplified subset internals
subset.data.frame <- function(x, subset, ...) {
  expr <- substitute(subset)
  row_mask <- eval(expr, x, parent.frame())
  x[row_mask & !is.na(row_mask), , drop = FALSE]
}

substitute() grabs the expression out of the promise. eval() evaluates it in an environment where the columns of x are visible as variables. That’s why x > 3 finds the column instead of failing with “object ‘x’ not found.”

Intercepting the expression and choosing where to evaluate it is non-standard evaluation, NSE for short, in its simplest form. Standard evaluation, the mean(c(1, 2, 3)) path, is what every other function gets. Side by side:

# Standard evaluation: explicit, verbose
df[df$x > 3, ]

# Non-standard evaluation: concise, readable
subset(df, x > 3)

NSE is the reason the tidyverse feels like a domain-specific language for data analysis. filter(penguins, species == "Adelie") reads almost like English because you write column names directly, without $ or quotes. The cost is that the rules become less transparent: the meaning of species depends on which function you’re inside, not just on what’s in your environment. That tension between conciseness and predictability runs through everything that follows.

Exercises

  1. Run subset(df, x > 3) with the data frame above, then try df[df$x > 3, ]. Verify they produce the same result. Which is easier to read?

  2. What happens if you define x <- 100 in your global environment and then run subset(df, x > 3) again? Does it use the column or the variable? Why?

23.4 Data masking

The call this chapter opened with does the same thing as subset():

library(dplyr)
filter(penguins, species == "Adelie")

R looks for species first in penguins, then in your calling environment. The data masks the environment: if the data frame has a column called species, that column wins over any variable named species in your workspace. Data masking is the tidyverse’s name for NSE with that lookup order, and aes(x = bill_length_mm) in ggplot2 is the same principle: the expression is captured and evaluated against the data at plot time, with columns treated as ordinary variables.

The benefit for interactive analysis is enormous. You type column names hundreds of times in a session; not having to write penguins$species or penguins[["species"]] each time keeps your code compact and readable, closer to how you think about the data than to how the computer stores it.

The cost appears the moment you try to program with data-masked functions. Suppose you want to write a function that filters by a column whose name is stored in a variable:

col <- "species"
filter(penguins, col == "Adelie")

filter looks for a column named col in penguins, doesn’t find one, then finds col in your environment (the string "species"), and compares that string to "Adelie". No error, and zero rows.

To see the mechanism concretely, here is a simplified filter() you could write yourself in about ten lines:

my_filter <- function(.data, expr) {
  e <- substitute(expr)                    # capture the caller's expression
  env <- list2env(.data, parent = parent.frame())  # build an environment from the data frame
  mask <- eval(e, envir = env)             # evaluate the expression in that environment
  .data[mask & !is.na(mask), , drop = FALSE]
}

df <- data.frame(x = 1:5, y = c(10, 20, 30, 40, 50))
my_filter(df, x > 3)
#>   x  y
#> 4 4 40
#> 5 5 50

substitute(expr) captures the caller’s expression (x > 3) without evaluating it, exactly as a promise holds an unevaluated expression (Section 23.1). list2env(.data, parent = parent.frame()) creates a new environment whose bindings are the columns of the data frame, with the caller’s environment as the parent so that variables not in the data frame are still found. eval(e, envir = env) forces the expression in that constructed environment, where x resolves to the column rather than to anything in the global environment. dplyr::filter() does the same, with more machinery for tidy evaluation, error handling, and grouped data frames layered on top.

Data masking is optimized for the common case (interactive analysis) at the expense of the less common case (writing reusable functions). That’s the problem tidy evaluation exists to solve.

Exercises

  1. Define x <- 1000 in your global environment. Create a data frame df <- data.frame(x = 1:5). What does dplyr::filter(df, x > 3) return? Does it use the column or the variable?

  2. Explain in one sentence why aes(x = bill_length_mm) doesn’t need quotes around bill_length_mm.

23.5 Tidy evaluation basics

Data masking works because substitute() captures the promise from the immediate caller. But when you wrap a data-masked function inside your own function, substitute() captures your wrapper’s argument name, not the expression the original caller wrote. Tidy evaluation exists to thread promises through that extra layer.

The embrace operator { } (called “curly-curly”) solves the most common programming-with-NSE problem: passing a column name through your function to a dplyr verb without it getting lost along the way.

library(dplyr)

my_summary <- function(data, var) {
  data |>
    summarise(mean = mean({{ var }}, na.rm = TRUE))
}

{ var } says: “take whatever the caller passed as var, capture it as a data-masked expression, and inject it here.” The caller writes column names unquoted, exactly as they would with dplyr directly:

my_summary(penguins, body_mass_g)

Without { }, you’d have to teach your users a different syntax for your wrapper than they use for dplyr itself, which defeats the purpose of wrapping. With it, the wrapper is transparent.

For tidy selection (used in select(), across(), and similar), { } works the same way:

my_select <- function(data, cols) {
  data |> select({{ cols }})
}
my_select(penguins, starts_with("bill"))

The caller passes tidy-select expressions exactly as they would to select() directly. Your function just forwards them.

When { } isn’t enough, there are more tools. .data[["col_name"]] lets you use string column names inside data-masked functions. !! (bang-bang) and enquo() give finer control over quoting and unquoting. These are part of the full tidy evaluation framework in the rlang package, and they matter when you’re building complex programmatic interfaces. But they belong in Chapter 26, not here.

TipOpinion

For most dplyr wrapper functions, { } is sufficient. The full tidy evaluation system (quosures, enquo(), !!, !!!) exists for cases where you need to build expressions programmatically, splice multiple variables, or mix quoted and unquoted inputs.

Exercises

  1. Write a function my_count(data, group_var) that uses { } to count rows per group. Test it with any data frame.

  2. Write a function my_arrange(data, sort_var) that arranges a data frame by a column passed by the caller. Use { }.

  3. What happens if you forget { } and write summarise(data, mean = mean(var, na.rm = TRUE)) inside a wrapper function? What error do you get?

23.6 The trade-offs of NSE

Promises enable NSE, NSE enables data masking, and { } makes data masking programmable. Each link adds a layer of indirection, and each layer has a price.

filter(df, x > 5) is concise and readable, but if a variable x exists in your environment and a column x exists in your data frame, the column wins without a message. When that’s wrong, nothing tells you.

NSE is optimized for the console, where you type select(df, name, age) and it works. Writing a function that wraps select requires { }, an extra concept that sits between you and the dplyr you already know. The interactive user pays nothing; the programmer pays, and the price scales with the complexity of the interface you’re building.

“Object ‘species’ not found” could be a typo, missing data, or a masking scope problem. The error message is the same in all three cases.

Base R has used NSE since S: subset(), with(), transform(), and the formula interface (lm(y ~ x, data = df)) all capture their arguments. The tidyverse applies the idea more systematically and more pervasively.

TipOpinion

NSE is the right trade-off for data analysis. You type column names hundreds of times a day, and quoting each one would be painful. The cost falls on the programmer writing reusable functions, not on the analyst exploring data interactively, and that’s a reasonable place for the cost to land since most R users spend more time analyzing than abstracting. Accept the trade-off, learn { } for when you cross the line into programming, and spend your energy on the next question: when you call print(x), how does R decide which print to use?