Debugging and Troubleshooting

Bugs are normal, not a sign you’re bad at this. What separates a smooth debugging session from a frustrating one is a calm, systematic process. Panic-editing random lines is the slow path; the steps below are the fast one (Figure 1).

Debugging is a calm, ordered process, not panic-editing. Read the whole error message, reduce the problem to a minimal reproducible example, isolate the failure by printing intermediate values, and check the types and shapes of your data. If it still fails, that same reprex is exactly what you need to search effectively or ask for help — then you loop back with what you learned.
Figure 1. Debugging is a calm, ordered process, not panic-editing. Read the whole error message, reduce the problem to a minimal reproducible example, isolate the failure by printing intermediate values, and check the types and shapes of your data. If it still fails, that same reprex is exactly what you need to search effectively or ask for help — then you loop back with what you learned.

Read the error message#

It sounds obvious, but read the whole message, slowly. It usually names the problem, the function, and the line. Most “mysterious” errors are explained in the text people skip.

Error in log(x) : non-numeric argument to mathematical function

That is telling you x isn’t numeric. Believe it, and go check the type of x.

Reproduce it minimally (a reprex)#

Strip the problem down to the smallest snippet that still fails. Doing so often reveals the cause on its own, and it’s exactly what you’ll need if you ask for help. This minimal, self-contained example is a reproducible example, or reprex.

R
# A good reprex: self-contained, runs on its own, shows the error
x <- c("1", "2", "3")   # oops, character not numeric
mean(x)
#> Error in mean.default(x): argument is not numeric or logical

R’s reprex package renders a snippet with its output so you can paste a faithful copy anywhere:

R
reprex::reprex({
  x <- c("1", "2", "3")
  mean(x)
})

Isolate with print and interactive tools#

Find where reality diverges from your expectation by looking at intermediate values.

R
# @show-style tracing in R
f <- function(v) {
  cat("class(v):", class(v), "\n")   # inspect before it blows up
  mean(v)
}
Python
def f(v):
    breakpoint()          # drops into pdb here; type `v`, `n`, `c`
    return sum(v) / len(v)
Julia
f(v) = (@show typeof(v); sum(v) / length(v))

Check assumptions: types and shapes#

A huge share of bugs are a value that isn’t what you assumed — a string where you expected a number, a length-0 vector, a matrix transposed, an NA sneaking through.

R
str(x)          # structure, type, first values
length(x); dim(m); anyNA(x)
Python
print(type(x), np.shape(x), np.isnan(x).any())
Julia
@show typeof(x) size(m) any(isnan, x)

Rubber-duck it#

Explain the code out loud, line by line, to a rubber duck (or a patient colleague). Forcing yourself to state what each line should do surfaces the line that doesn’t.

Search effectively#

When and how to ask for help#

Ask once you’ve read the error, made a reprex, and checked types — that work often solves it, and if not, it makes your question answerable. A good help request includes:

  1. What you’re trying to do (the goal, not just the broken line).
  2. A minimal reproducible example anyone can run.
  3. The full error message.
  4. What you expected vs. what happened, and what you already tried.
  5. Your environment (sessionInfo() / versions).

This is the spirit of Jenny Bryan’s What They Forgot to Teach You About R: a self-contained reprex plus a clean environment respects the helper’s time and, more often than not, hands you the answer while you build it.

A worked example#

R
# BAD: fails, and it's not obvious why at a glance
rates <- read.csv("rates.csv")$rate   # comes in as "0.3","0.5",...
adjusted <- rates * 100
#> Error: non-numeric argument to binary operator

Diagnose it:

R
str(rates)
#>  chr [1:200] "0.3" "0.5" ...     # <- the culprit: character, not numeric

Fix it:

R
rates <- as.numeric(read.csv("rates.csv")$rate)
adjusted <- rates * 100