Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions src/_quarto.yml
Original file line number Diff line number Diff line change
Expand Up @@ -199,6 +199,7 @@ website:
- reference-manual/diagnostics.qmd
- section: "Usage"
contents:
- reference-manual/glossary.qmd
- reference-manual/reproducibility.qmd
- reference-manual/licenses.qmd

Expand Down
1 change: 1 addition & 0 deletions src/reference-manual/_quarto.yml
Original file line number Diff line number Diff line change
Expand Up @@ -64,6 +64,7 @@ book:

- part: "Usage"
chapters:
- glossary.qmd

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the glossary should be last in the list here.

- reproducibility.qmd
- licenses.qmd

Expand Down
110 changes: 110 additions & 0 deletions src/reference-manual/glossary.qmd
Original file line number Diff line number Diff line change
@@ -0,0 +1,110 @@
---
pagetitle: Glossary
---

This glossary offers simplified explanations of common technical terms found
in Stan's documentation and interfaces. The glossary will aid users that are
familiar with basic concepts in Bayesian inference and computation, but
unfamiliar with Stan's vocabulary.

The explanations included here are not intended to be exhaustive or detailed.
For more information, consult our chapter [*How to Diagnose and Resolve

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please keep to third person and remove the "our".

Convergence
Problems*](https://mc-stan.org/learn-stan/diagnostics-warnings.html) or our
section [*Algorithms*](https://mc-stan.org/docs/reference-manual/mcmc.html)
in the Reference Manual.

**Bulk ESS** A version of the effective sample size that measures the
reliability of the center, or "bulk", of the posterior distribution, i.e., the
region summarized by a mean or a median. A high bulk ESS indicates that

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Which is it, mean or median? I've never quite understood this concept myself, so I'm not sure. But with many distributions medians and means can be pretty far apart with enough skew.

estimates of the distribution's central tendency are reliable. See also:
*Effective sample size*, *Tail ESS*.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It would be better if these were links.


**Chain** An ordered sequence of random draws produced by a sampler to

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this should say that it's a Markov chain, but then that's not true during warmup.

describe a posterior distribution (see a formal definition
[here](https://mc-stan.org/docs/reference-manual/analysis.html#markov-chains)).
Fitting a model typically involves several chains running in parallel, which
lets us check that they agree with one another. See also: *Iteration*, *Monte
Carlo standard error (MCSE)*, *R-hat*.

**Divergence** A warning that the sampler ran into trouble and may have

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would say that the divergence is referring to not preserving the Hamiltonian in the Hamiltonian simulation, as I'm not sure people understand what is supposed to be diverging here.

skipped part of the posterior distribution. Skipping this region biases the
posterior sample and hinders inferences. So, divergences signal that the draws
may be unreliable---even if the R-hat and effective sample size look fine. The
Reference Manual offers a more formal explanation
[here](https://mc-stan.org/docs/reference-manual/mcmc.html#divergent-transitions).
And Stan's [*Learning Resources on Diagnostics and
warnings*](https://mc-stan.org/learn-stan/diagnostics-warnings.html#divergent-transitions-after-warmup)
explicate some causes and solutions. See also: *Maximum tree depth*,
*Reparametrize*.

**Effective sample size (ESS)** A measure of how much information a set of
draws actually provides. Draws within a chain tend to be related to one
another, making them less informative than the same number of independent
draws would be. For a given number of draws in a model, the ESS reports the

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd lead with the "For a given number" sentence as that is the real explanation.

equivalent number of independent draws. The [chapter on Posterior
Analysis](https://mc-stan.org/docs/reference-manual/analysis.html#effective-sample-size.section)
of the Reference Manual offers a formal definition. And our chapter on
*How to Diagnose and Resolve Convergence Problems*
[suggests](https://mc-stan.org/learn-stan/diagnostics-warnings.html#bulk-and-tail-ess)
some useful thresholds. See also: *Bulk ESS*, *Tail ESS*.

**Iteration** A single step in a sampler's computational process. Each step
produces one draw in a chain. The total number of iterations includes warmup
iterations, which are discarded; and post-warmup iterations, which are kept

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The warmup iterations can be saved, but they shouldn't be used for inference (i.e., calculating means or quantiles).

for analysis. See also: *Chain*, *Warmup*.

**Maximum treedepth** A limit on how long Stan’s NUTS sampler may search for a

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd link this to say that the maximum number of leapfrog steps is 2^(max treedepth), and each leapfrog step requires a log density and gradient evaluation.

usually if you hit the limit, the problem is that you need to parameterize. Very rarely is increasing the max depth to 11 or 12 or 13 the right answer.

new draw during each iteration. If the sampler repeatedly reaches this limit,
it may be stopping its search too early and exploring the posterior
inefficiently. See more details
[here](https://mc-stan.org/learn-stan/diagnostics-warnings.html#maximum-treedepth)
and suggestions for dealing with this problem in point 7
[here](https://mc-stan.org/learn-stan/diagnostics-warnings.html#diagnosing-and-resolving-problems).
See also: *Divergence*, *Reparametrize*.

**Monte Carlo standard error (MCSE)** A measure of the uncertainty of the
Monte Carlo samples used to get the posterior mean of a parameter. Stan relies

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wouldn't say there's uncertainty in the samples. The uncertainty we care about is in the estimate and what would happen under new samples. As such, MCSE is an estimate of the standard deviation of the posterior mean estimator under replicated data.

on Monte Carlo random sampling, so different runs of the same simulation can
yield different results. The MCSE quantifies how much an estimate (such as a
posterior mean) is expected to fluctuate due to this randomness. See a formal
definition
[here](https://mc-stan.org/docs/reference-manual/analysis.html#estimation-of-mcmc-standard-error).
For final results, the MCSE should be small enough not to affect the digits
intended for reporting. So, if we plan to report two decimal places, the MCSE
should be below $0.01$. See also: *Effective sample size*, *Iteration*.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The 0.01 value only holds if the estimate is 1. So this needs to be removed. MCSE is relative the estimate. For example, 0.0012 is reported to two digits of accuracy and requires MCSE be below 0.0001. But this isn't how people report. People regularly write things like 0.32 +/- 0.05 where 0.05 is the standard error.


**R-hat** A diagnostic that evaluates whether multiple chains have converged

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

diagnostic -> statistic of multiple chains

It doesn't evaluate, it tests. And it's not a foolproof test in that there are false positives and false negatives.

The two types of variation do not become equal. Rhat is essential sqrt(1 + between/within). The idea is that the between variation goes to 0 and the within stays constant at its nominal value based on the model, so that Rhat converges to 1.

Charles Margossian's nested R-hat paper is super useful for understanding what R-hat is doing.

to the same posterior distribution. R-hat compares the variation of draws
within each chain to the variation across different chains. When chains
converge, these two types of variation become roughly equal, so R-hat gets
close to 1. When chains diverge, the variation across different chains is
larger than the variation within chains, so R-hat is above 1. See a formal

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's happening is that the within chain variation scales the between chain variation to make R-hat unit free.

definition in [this paper](https://arxiv.org/abs/1903.08008) and some useful
thresholds [here](https://mc-stan.org/learn-stan/diagnostics-warnings.html#r-hat).
See also: *Chain*.

**Reparametrize** To rewrite a model in a different but mathematically

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

They're not mathematically equivalent in that they define two different densities (otherwise there wouldn't be much point). The idea is that the posteriors have enough information to derive each others random variables of interest.

Stan can always represent a model in infinitely many different ways.

Centering a predictor does not leave the model the same---it shifts things from coefficients into the intercept in a regression, for example. The point is that the centered posterior draws can be transformed to non-centered posterior draws.

equivalent form. Stan can often represent a given model in multiple ways, some
of which are easier to sample than others. For example, centering a continuous
predictor by subtracting its mean leaves the fitted model unchanged but
reduces the posterior dependency between the intercept and the slope.
Reparametrizing a model can help avoid divergences and increase effective
sample sizes. See
[here](https://mc-stan.org/docs/stan-users-guide/reparameterization.html)
for examples in Stan, and see point 15
[here](https://mc-stan.org/learn-stan/diagnostics-warnings.html#diagnosing-and-resolving-problems)
for more resources. See also: *Divergence*, *Effective sample size*,
*Maximum tree depth*.

**Tail ESS** A version of the effective sample size that measures the
reliability of the tails of the posterior distribution, i.e., the regions
summarized by the 5% or 95% posterior quantiles. A high tail ESS indicates

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this really the 5% and 95% quantiles? I don't know.

that these quantiles are reliable. See also: *Bulk ESS*,
*Effective sample size*.

**Warmup** An initial set of iterations that are discarded rather than used

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're describing what people call "burn in" despite Gelman's objections that it's a bad analogy. In Stan, warmup critically does adaptation. As such, it doesn't form a Markov chain. We just assume that the adaptation will also take care of burn in.

There are three warmup phases. Phase I attempts to burn in with a unit mass matrix. Phase II attempts to estimate the mass matrix. Phase III does a final tuning of step size.

for analysis. Early draws in a chain are usually too inaccurate to describe
the posterior distribution reliably, so the sampler uses this warmup period to
settle into the region of plausible parameter values before collecting the
draws that are kept. See also: *Iteration*.