-
-
Notifications
You must be signed in to change notification settings - Fork 133
Added draft of glossary to reference manual. #979
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
base: master
Are you sure you want to change the base?
Changes from all commits
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -64,6 +64,7 @@ book: | |
|
|
||
| - part: "Usage" | ||
| chapters: | ||
| - glossary.qmd | ||
| - reproducibility.qmd | ||
| - licenses.qmd | ||
|
|
||
|
|
||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,110 @@ | ||
| --- | ||
| pagetitle: Glossary | ||
| --- | ||
|
|
||
| This glossary offers simplified explanations of common technical terms found | ||
| in Stan's documentation and interfaces. The glossary will aid users that are | ||
| familiar with basic concepts in Bayesian inference and computation, but | ||
| unfamiliar with Stan's vocabulary. | ||
|
|
||
| The explanations included here are not intended to be exhaustive or detailed. | ||
| For more information, consult our chapter [*How to Diagnose and Resolve | ||
|
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Please keep to third person and remove the "our". |
||
| Convergence | ||
| Problems*](https://mc-stan.org/learn-stan/diagnostics-warnings.html) or our | ||
| section [*Algorithms*](https://mc-stan.org/docs/reference-manual/mcmc.html) | ||
| in the Reference Manual. | ||
|
|
||
| **Bulk ESS** A version of the effective sample size that measures the | ||
| reliability of the center, or "bulk", of the posterior distribution, i.e., the | ||
| region summarized by a mean or a median. A high bulk ESS indicates that | ||
|
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Which is it, mean or median? I've never quite understood this concept myself, so I'm not sure. But with many distributions medians and means can be pretty far apart with enough skew. |
||
| estimates of the distribution's central tendency are reliable. See also: | ||
| *Effective sample size*, *Tail ESS*. | ||
|
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. It would be better if these were links. |
||
|
|
||
| **Chain** An ordered sequence of random draws produced by a sampler to | ||
|
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. I think this should say that it's a Markov chain, but then that's not true during warmup. |
||
| describe a posterior distribution (see a formal definition | ||
| [here](https://mc-stan.org/docs/reference-manual/analysis.html#markov-chains)). | ||
| Fitting a model typically involves several chains running in parallel, which | ||
| lets us check that they agree with one another. See also: *Iteration*, *Monte | ||
| Carlo standard error (MCSE)*, *R-hat*. | ||
|
|
||
| **Divergence** A warning that the sampler ran into trouble and may have | ||
|
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. I would say that the divergence is referring to not preserving the Hamiltonian in the Hamiltonian simulation, as I'm not sure people understand what is supposed to be diverging here. |
||
| skipped part of the posterior distribution. Skipping this region biases the | ||
| posterior sample and hinders inferences. So, divergences signal that the draws | ||
| may be unreliable---even if the R-hat and effective sample size look fine. The | ||
| Reference Manual offers a more formal explanation | ||
| [here](https://mc-stan.org/docs/reference-manual/mcmc.html#divergent-transitions). | ||
| And Stan's [*Learning Resources on Diagnostics and | ||
| warnings*](https://mc-stan.org/learn-stan/diagnostics-warnings.html#divergent-transitions-after-warmup) | ||
| explicate some causes and solutions. See also: *Maximum tree depth*, | ||
| *Reparametrize*. | ||
|
|
||
| **Effective sample size (ESS)** A measure of how much information a set of | ||
| draws actually provides. Draws within a chain tend to be related to one | ||
| another, making them less informative than the same number of independent | ||
| draws would be. For a given number of draws in a model, the ESS reports the | ||
|
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. I'd lead with the "For a given number" sentence as that is the real explanation. |
||
| equivalent number of independent draws. The [chapter on Posterior | ||
| Analysis](https://mc-stan.org/docs/reference-manual/analysis.html#effective-sample-size.section) | ||
| of the Reference Manual offers a formal definition. And our chapter on | ||
| *How to Diagnose and Resolve Convergence Problems* | ||
| [suggests](https://mc-stan.org/learn-stan/diagnostics-warnings.html#bulk-and-tail-ess) | ||
| some useful thresholds. See also: *Bulk ESS*, *Tail ESS*. | ||
|
|
||
| **Iteration** A single step in a sampler's computational process. Each step | ||
| produces one draw in a chain. The total number of iterations includes warmup | ||
| iterations, which are discarded; and post-warmup iterations, which are kept | ||
|
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. The warmup iterations can be saved, but they shouldn't be used for inference (i.e., calculating means or quantiles). |
||
| for analysis. See also: *Chain*, *Warmup*. | ||
|
|
||
| **Maximum treedepth** A limit on how long Stan’s NUTS sampler may search for a | ||
|
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. I'd link this to say that the maximum number of leapfrog steps is 2^(max treedepth), and each leapfrog step requires a log density and gradient evaluation. usually if you hit the limit, the problem is that you need to parameterize. Very rarely is increasing the max depth to 11 or 12 or 13 the right answer. |
||
| new draw during each iteration. If the sampler repeatedly reaches this limit, | ||
| it may be stopping its search too early and exploring the posterior | ||
| inefficiently. See more details | ||
| [here](https://mc-stan.org/learn-stan/diagnostics-warnings.html#maximum-treedepth) | ||
| and suggestions for dealing with this problem in point 7 | ||
| [here](https://mc-stan.org/learn-stan/diagnostics-warnings.html#diagnosing-and-resolving-problems). | ||
| See also: *Divergence*, *Reparametrize*. | ||
|
|
||
| **Monte Carlo standard error (MCSE)** A measure of the uncertainty of the | ||
| Monte Carlo samples used to get the posterior mean of a parameter. Stan relies | ||
|
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. I wouldn't say there's uncertainty in the samples. The uncertainty we care about is in the estimate and what would happen under new samples. As such, MCSE is an estimate of the standard deviation of the posterior mean estimator under replicated data. |
||
| on Monte Carlo random sampling, so different runs of the same simulation can | ||
| yield different results. The MCSE quantifies how much an estimate (such as a | ||
| posterior mean) is expected to fluctuate due to this randomness. See a formal | ||
| definition | ||
| [here](https://mc-stan.org/docs/reference-manual/analysis.html#estimation-of-mcmc-standard-error). | ||
| For final results, the MCSE should be small enough not to affect the digits | ||
| intended for reporting. So, if we plan to report two decimal places, the MCSE | ||
| should be below $0.01$. See also: *Effective sample size*, *Iteration*. | ||
|
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. The 0.01 value only holds if the estimate is 1. So this needs to be removed. MCSE is relative the estimate. For example, 0.0012 is reported to two digits of accuracy and requires MCSE be below 0.0001. But this isn't how people report. People regularly write things like 0.32 +/- 0.05 where 0.05 is the standard error. |
||
|
|
||
| **R-hat** A diagnostic that evaluates whether multiple chains have converged | ||
|
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. diagnostic -> statistic of multiple chains It doesn't evaluate, it tests. And it's not a foolproof test in that there are false positives and false negatives. The two types of variation do not become equal. Rhat is essential sqrt(1 + between/within). The idea is that the between variation goes to 0 and the within stays constant at its nominal value based on the model, so that Rhat converges to 1. Charles Margossian's nested R-hat paper is super useful for understanding what R-hat is doing. |
||
| to the same posterior distribution. R-hat compares the variation of draws | ||
| within each chain to the variation across different chains. When chains | ||
| converge, these two types of variation become roughly equal, so R-hat gets | ||
| close to 1. When chains diverge, the variation across different chains is | ||
| larger than the variation within chains, so R-hat is above 1. See a formal | ||
|
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. What's happening is that the within chain variation scales the between chain variation to make R-hat unit free. |
||
| definition in [this paper](https://arxiv.org/abs/1903.08008) and some useful | ||
| thresholds [here](https://mc-stan.org/learn-stan/diagnostics-warnings.html#r-hat). | ||
| See also: *Chain*. | ||
|
|
||
| **Reparametrize** To rewrite a model in a different but mathematically | ||
|
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. They're not mathematically equivalent in that they define two different densities (otherwise there wouldn't be much point). The idea is that the posteriors have enough information to derive each others random variables of interest. Stan can always represent a model in infinitely many different ways. Centering a predictor does not leave the model the same---it shifts things from coefficients into the intercept in a regression, for example. The point is that the centered posterior draws can be transformed to non-centered posterior draws. |
||
| equivalent form. Stan can often represent a given model in multiple ways, some | ||
| of which are easier to sample than others. For example, centering a continuous | ||
| predictor by subtracting its mean leaves the fitted model unchanged but | ||
| reduces the posterior dependency between the intercept and the slope. | ||
| Reparametrizing a model can help avoid divergences and increase effective | ||
| sample sizes. See | ||
| [here](https://mc-stan.org/docs/stan-users-guide/reparameterization.html) | ||
| for examples in Stan, and see point 15 | ||
| [here](https://mc-stan.org/learn-stan/diagnostics-warnings.html#diagnosing-and-resolving-problems) | ||
| for more resources. See also: *Divergence*, *Effective sample size*, | ||
| *Maximum tree depth*. | ||
|
|
||
| **Tail ESS** A version of the effective sample size that measures the | ||
| reliability of the tails of the posterior distribution, i.e., the regions | ||
| summarized by the 5% or 95% posterior quantiles. A high tail ESS indicates | ||
|
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Is this really the 5% and 95% quantiles? I don't know. |
||
| that these quantiles are reliable. See also: *Bulk ESS*, | ||
| *Effective sample size*. | ||
|
|
||
| **Warmup** An initial set of iterations that are discarded rather than used | ||
|
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. You're describing what people call "burn in" despite Gelman's objections that it's a bad analogy. In Stan, warmup critically does adaptation. As such, it doesn't form a Markov chain. We just assume that the adaptation will also take care of burn in. There are three warmup phases. Phase I attempts to burn in with a unit mass matrix. Phase II attempts to estimate the mass matrix. Phase III does a final tuning of step size. |
||
| for analysis. Early draws in a chain are usually too inaccurate to describe | ||
| the posterior distribution reliably, so the sampler uses this warmup period to | ||
| settle into the region of plausible parameter values before collecting the | ||
| draws that are kept. See also: *Iteration*. | ||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
I think the glossary should be last in the list here.