Bollinger Bands

How to use rolling-average bands to set bounds on a KPI

Featured image

The short version

Bollinger Bands draw a line through the average of the last 20 values, then a band two standard deviations above and below it. Because the width comes from the data's own recent spread, the band widens when a number gets jumpy and narrows when it calms down. That self-adjusting width is why they are tempting for KPI bounds: you don't have to hand-pick a threshold for every metric.

Used as-is on business data, though, they break in predictable ways. Weekly cycles inflate the width. One bad day widens the band for three weeks. A real, lasting shift gets absorbed as "normal" within a window. The "2 standard deviations = 95%" idea is wrong even by Bollinger's own account.

Recommendation: keep the idea (bounds that come from the metric's own history) but build them the way monitoring teams do. Compare each day with the same weekday, use a median and a robust spread instead of mean and standard deviation, keep flagged days out of the baseline, and add a "run of 8" rule to catch lasting shifts. Section 6 has a concrete design.

Why this matters to you

This came from a card on my board: "Look up Bollinger Bands (rolling-average bands) for setting KPI parameter bounds." The product I'm working on is a retail-analytics system. It watches operational KPIs, things like shipping cost per order, the share of orders split across locations, and store-level margin, and makes recommendations when something moves. Parameter design, meaning how we decide what counts as "normal" for each KPI, sits on the critical path.

So the question here isn't "can Bollinger Bands make money trading stocks?" It's "is a rolling mean ± k standard deviations a good way to set the normal range for a business metric, and if not as-is, what should we build?" All numbers in the worked examples are simulated; none come from a real business.

The basics Remember

Simple moving average (SMA)
The plain average of the last n values, recomputed every period as the window slides forward.
Standard deviation (SD, σ)
A measure of how far values typically sit from their average; bigger means more spread out.
Middle band
The 20-period SMA — the centre line of a Bollinger Band.
Upper and lower bands
The middle band plus and minus k standard deviations of the same 20 values (k = 2 by default).
%b ("percent b")
Where the latest value sits inside the band: 0 at the lower band, 0.5 in the middle, 1 at the upper band, above 1 or below 0 when it's outside.
BandWidth
The band's width divided by the middle band — a single number for how volatile the series currently is.
The Squeeze
A period when BandWidth is unusually low (traders often use "lowest in about six months"), taken as a sign a big move may follow.
Envelope
The older idea Bollinger improved on: lines a fixed percentage (say 3%) above and below a moving average.
Control chart
A chart from factory quality control with a centre line and limits (usually ±3σ) used to spot when a process has changed.
False alarm / missed detection
Flagging a normal value as unusual, versus failing to flag a real change. Every threshold trades one for the other.

Where they came from

John Bollinger developed the bands around 1980 while trading options, and made them public in the early 1980s while he was chief market analyst at the Financial News Network [1]. The name was an accident: an anchor asked him on air what he called them.

Traders already drew bands around a moving average, but at a fixed distance, such as 5% above and below. Bollinger's complaint was that markets have calm spells and wild spells, and a fixed-width band is too wide in the first and too narrow in the second. His fix was to let the band's width come from the series' own recent standard deviation [1]. He wrote the definitive book on them, Bollinger on Bollinger Bands, in 2001, and published 22 rules for using them [2].

The formula, named

For the latest period \(t\), with window length \(n\) (default 20) and width multiplier \(k\) (default 2):

\[ \text{Middle}_t = \bar{x}_t = \frac{1}{n}\sum_{i=0}^{n-1} x_{t-i} \qquad \sigma_t = \sqrt{\frac{1}{n}\sum_{i=0}^{n-1}\left(x_{t-i}-\bar{x}_t\right)^2} \] \[ \text{Upper}_t = \bar{x}_t + k\,\sigma_t \qquad \text{Lower}_t = \bar{x}_t - k\,\sigma_t \]

In words: average the last 20 values; measure how spread out they are; put the band two spreads above and below the average. Three details matter later:

  • Bollinger divides by \(n\), not \(n-1\) — the "population" standard deviation [1]. Many tools (pandas, for one) default to \(n-1\), which makes 20-period bands about 2.6% wider than Bollinger's [9].
  • \(k\) should change with \(n\). His Rule 11: use 1.9 for a 10-period window and 2.1 for a 50-period one [2].
  • The derived numbers: \(\%b = (x_t - \text{Lower}_t)/(\text{Upper}_t - \text{Lower}_t)\) and \(\text{BandWidth} = (\text{Upper}_t-\text{Lower}_t)/\text{Middle}_t\). With the default settings, BandWidth is exactly four times the coefficient of variation (SD divided by mean) [2][3].

Check yourself

What are the three lines of a Bollinger Band, and what are the default settings?

A middle band (the 20-period simple moving average) and upper and lower bands at the middle band ± 2 standard deviations of the same 20 values.

What problem were they designed to fix?

Fixed-percentage envelopes don't adapt to changing volatility. Bollinger Bands take their width from the series' recent standard deviation, so they widen in volatile periods and narrow in quiet ones.

A value has %b = 1.3. Where is it?

Above the upper band, by 30% of the band's width.

How it works Understand

An analogy first

Think of a road where the lane markings are repainted every night, based on how much drivers swerved over the past 20 days. On a calm stretch, the lanes get narrow, so a small drift looks unusual. On a stretch where everyone swerves (potholes, wind), the lanes widen, so the same drift looks normal. The lanes don't tell you where the road should go. They only describe where traffic has recently been.

That last point is the key to everything that follows. Bollinger Bands are a description of the recent past, not a model of what ought to happen. Bollinger says so himself in Rule 1: the bands define "high" and "low" relative to recent values, nothing more [2].

Now without the analogy

Each period, three things happen:

  1. The window slides forward one step: the newest value enters, the oldest leaves.
  2. The middle band moves toward wherever the window's values now sit.
  3. The width is recomputed from how scattered those values are around their new average.

Because the width is measured from the same window as the centre, anything that makes the window's values scatter widens the band. That includes real noise, but also a trend (values climbing steadily look "spread out" around their average), a weekly pattern (Mondays low, Saturdays high), and a single extreme value. The band can't tell these apart; all of them are just "spread."

Why the band widens when a series gets volatile

Standard deviation squares each value's distance from the average before averaging. A few large distances therefore dominate. When the series starts swinging harder, those large distances enter the window, the SD jumps, and the band opens. When the swings leave the window, 20 periods later, the band closes again — often abruptly, which is why Bollinger allows an exponentially weighted version that fades old values gradually instead of dropping them [2].

Common misunderstandings

BeliefReality
"The bands contain 95% of values, so a break is a rare 2-sigma event."95% only holds for independent, bell-shaped data. Bollinger's own Rule 14 says to "make no statistical assumptions" and puts containment near 90% for prices [2]; StockCharts measures 88–89% [3]. For an arbitrary distribution, the only guarantee (Chebyshev's inequality) is 75% [19].
"Touching the upper band is a sell signal."Rule 6: "Tags of the bands are just that, tags not signals." Rule 7: in a trend, values can "walk" along the band for a long time [2].
"A Squeeze tells you which way it will break."It only says volatility is unusually low. The first break is often a fake-out [3].
"You can change the window and keep k = 2."Shorter windows should use smaller k, longer ones larger k (Rule 11) [2].
"It's a complete method."Bollinger presents it as one input to combine with unrelated indicators — "two momentum indicators aren't better than one" [2].

Check yourself

In your own words: why does a steady upward trend widen the bands even if day-to-day noise hasn't changed?

The SD measures scatter around the window's average. In a trend, early values sit well below that average and late values well above it, so the window looks spread out. The band reads trend as noise.

Why does a single extreme day affect the band for exactly one window length?

It stays inside the window — pulling the average and inflating the SD — until it's 20 periods old and drops out. Then the band snaps back.

Using it Apply

A worked example

Take a made-up KPI: average shipping cost per online order, in US dollars, for ten days. To keep the arithmetic readable, use a 10-day window. (Strictly, Rule 11 says k = 1.9 for 10 periods; I'll use 2 so the numbers match the general formula.)

Day12345678910
Cost14.213.814.514.013.614.914.313.914.414.1

Step 1 — the middle band. The ten values sum to 141.7, so the average is \(141.7 / 10 = 14.17\).

Step 2 — the squared distances. Subtract 14.17 from each value and square it:

0.0009, 0.1369, 0.1089, 0.0289, 0.3249, 0.5329, 0.0169, 0.0729, 0.0529, 0.0049. These sum to 1.281.

Step 3 — the standard deviation. Divide by \(n = 10\) and take the square root: \(\sqrt{1.281/10} = \sqrt{0.1281} \approx 0.358\).

Step 4 — the bands. Upper \(= 14.17 + 2(0.358) = 14.886\). Lower \(= 14.17 - 2(0.358) = 13.454\). So for day 11, "normal" is roughly $13.45 to $14.89.

Step 5 — judge day 11. Day 11 comes in at $16.20. Its %b is \((16.20 - 13.454) / (14.886 - 13.454) = 2.746 / 1.432 \approx 1.92\) — far above the band. That deserves a look.

Step 6 — notice something. Day 6 ($14.90) sits just above the upper band of 14.886. But that band was computed from a window that includes day 6 itself. That's the trap in naive implementations: a value helps set the limit it's being judged against. The fix is to judge day \(t\) against a band built from days \(t-n\) to \(t-1\) only.

Rule of thumb: always build today's band from yesterday and earlier. Otherwise a spike inflates its own band and partly hides itself — and in a backtest, you've let the model see the answer.

How to do it yourself

  1. Pick the KPI and the period (daily is typical). Make sure missing days are marked as missing, not zero.
  2. Choose a window. For daily business data, use a multiple of 7 (28 days is a good start) so every weekday is equally represented.
  3. For each day, compute the mean and SD of the previous \(n\) days — not including today.
  4. Set the band at mean ± k·SD. Start with k = 3 if you'll monitor many metrics, 2 if only a few (section 4 explains why).
  5. Flag days outside the band. Look at a few months of history before switching alerts on: how many flags, and were they real?

In code

A minimal pandas version. Note shift(1) (use only earlier days) and ddof=0 (Bollinger's population SD):

import pandas as pd

s = df.set_index("date")["ship_cost_per_order"].asfreq("D")   # gaps become NaN

window = s.shift(1).rolling(28, min_periods=21)
mid    = window.mean()
sd     = window.std(ddof=0)
k      = 3
upper, lower = mid + k * sd, mid - k * sd

pct_b = (s - lower) / (upper - lower)
flags = s[(s > upper) | (s < lower)]

In a spreadsheet, with values in column B starting at row 2, the upper band for row 30 is =AVERAGE(B2:B29) + 3*STDEV.P(B2:B29). STDEV.P is the population version.

Check yourself

Given the last five values 55, 45, 51, 49, 53, compute 5-period bands with k = 2 (population SD). Is a next value of 56 outside?

Sum = 253, so the mean is 50.6. Squared distances: 19.36, 31.36, 0.16, 2.56, 5.76; sum 59.2; divide by 5 = 11.84; SD ≈ 3.44. Bands: 50.6 ± 6.88 → 43.72 to 57.48. So 56 is inside.

Why must the rolling window be shifted by one day?

So the day being judged isn't part of its own baseline. Otherwise a spike raises both the mean and the SD, shrinking its own apparent size.

Taking it apart Analyze

What each dial controls

SettingTurn it upTurn it down
Window length \(n\)Steadier band; slower to accept a genuine new normal; slower to follow seasonal drift.Adapts quickly, but chases every wobble. After a real step change the band widens and swallows the shift within days.
Multiplier \(k\)Fewer false alarms, more missed changes.More sensitive, many more false alarms.
Centre (mean vs median)The mean is dragged by extreme values; the median isn't.
Spread (SD vs MAD)One outlier can make the SD as large as you like; the median absolute deviation (MAD) ignores up to half the window being bad [15].

How much k matters, in numbers

If the data were independent and bell-shaped (they won't be, but it sets the scale), the chance that a normal day falls outside the band is:

  • k = 2: about 4.6%, or 1 day in 22.
  • k = 3: about 0.27%, or 1 day in 370 [5].

Now multiply by the number of KPIs. Watch 200 metrics at k = 2 and you'd expect about 9 false alarms every single day. At k = 3, about one every two days [20]. Real business data has heavier tails and day-to-day carry-over, which push these numbers up, not down [17][18]. This is the single biggest reason trading defaults (k = 2) are wrong for a KPI monitoring product.

The assumptions, and what happens when they're false

1. Values are independent. Business metrics aren't: a high-cost day tends to follow a high-cost day. With that kind of carry-over, a short window sees less spread than the series really has, so the band hugs the local level and breaks whenever the series drifts [17]. (The experts disagree on how much this matters; see section 5.)

2. The level is stable within the window. A trend or step change inflates the SD and moves the centre. The result: the series "walks" along the band edge rather than breaking it, and a lasting shift becomes the new normal within about one window [2].

3. No repeating pattern. This is the one that bites daily retail data hardest. Look at a simulated KPI with a weekly cycle — weekends cost more — plus a one-day spike on day 38 and a real, lasting step up from day 60:

12 14 16 18 day 0 day 14 day 28 day 42 day 56 day 70 real step up 20-day Bollinger Bands (k = 2) on a KPI with a weekly cycle
Simulated data. Red circles are flags. The band is about $3 wide because the weekday/weekend swing counts as "noise." It catches the day-38 spike but also flags an ordinary Sunday (day 27), and after the real step it flags only two Saturdays — by accident, because Saturdays are the highest days.

The 20-day band spends its width on the weekly cycle. It's too wide to see modest problems on weekdays, yet still trips on high weekends. The step change on day 60 is never clearly flagged; it just drifts into the baseline.

4. The baseline is clean. The day-38 spike sits in the window for the next 20 days, widening the band and making it blind to a second problem in that stretch. A week-long outage would become next month's "normal" [6].

5. The spread is the same at every volume. For a ratio KPI — split-order rate, conversion, cost per shipment — a store with 40 orders a day is naturally far noisier than one with 4,000. A single SD band is too tight for the small store and too loose for the big one. The quality-control answer is a band that depends on that day's denominator (a "p-chart"): \(\bar{p} \pm 3\sqrt{\bar{p}(1-\bar{p})/n_t}\) [16].

The nearest alternatives

MethodHow the band is builtGood atWeak at
Bollinger BandsRolling mean ± k·SD of last n valuesSimple; adapts to changing volatilityWeekly cycles, outliers, lasting shifts
Fixed-percentage envelopeRolling mean ± fixed %Predictable; easy to explainIgnores how noisy each metric is
Shewhart / XmR chartFixed limits from a baseline period; spread from average day-to-day change (mean ± 2.66 × average moving range) [7]Big, sudden changes; trends don't inflate the limitsSmall lasting drifts (a 1σ shift takes ~44 points to catch) [5]
EWMA chartExponentially weighted average; tight limits on itSmall gradual drifts [4]Slower to react to a single spike
Same-weekday median ± MAD (Hampel-style)Median of the last 6–8 same weekdays ± k × 1.4826 × MADWeekly cycles; outliers can't widen the bandNeeds 6–8 weeks of history; MAD can hit zero on flat data
Forecast interval (Holt-Winters, Prophet)Model trend + season + holidays; band = prediction intervalComplex seasonality, holidays, retail calendars [14]More to build, tune and explain

Here is the same simulated series with a same-weekday baseline: each day is compared with the median of the previous six same weekdays, with a robust spread.

12 14 16 18 day 0 day 14 day 28 day 42 day 56 day 70 real step up Same-weekday baseline: median of last 6 same weekdays, ±3.5 robust SD
Simulated data. The band now follows the weekly shape and is much narrower on weekdays. The spike is caught and the ordinary Sunday is not flagged. The step is still hard to see day by day (only one Monday, day 70, trips the band) — but a "run of 8 days above the centre line" rule fires on day 67, a week after the change.

Check yourself

Your KPI has a strong weekly cycle. What goes wrong with 20-day bands, and what would you do instead?

The weekday-to-weekend swing inflates the SD, so the band is too wide on typical days while high weekends can still trip it. Compare each day with the same weekday over the last 6–8 weeks, or remove the weekly pattern first and put the band on what's left.

A cost KPI jumps 6% and stays there. A week later the 28-day band has stopped flagging it. Diagnose.

Rolling bands forget: the higher values have entered the window, raised the centre and widened the SD. Single-point bands are poor at persistent shifts. Add a persistence rule (e.g. 8 in a row above the centre) or an EWMA/CUSUM check, and consider freezing the baseline when a shift is flagged.

Judging it Evaluate

The evidence from trading

Bollinger Bands were built for trading, so that's where they've been tested most. The results are not encouraging for the standard rules:

  • Lento, Gradojevic and Wright (2007) tested US and Canadian indices and currencies from 1995 to 2004. After trading costs, band rules "are consistently unable to earn profits in excess of the buy-and-hold" strategy [10].
  • Leung and Chong (2003) compared Bollinger Bands with plain fixed-percentage envelopes. The adaptive width — the whole point of the method — did not improve results [11].
  • Fang, Jacobsen and Qin (2017) found the bands would have been very profitable before they were published, and lost most of their predictive power once they became popular [12].
  • A survey of about 95 studies of technical trading rules found early profits shrink sharply once researchers account for having tried many rules and reported the best [13].

Most of that doesn't transfer to KPI monitoring. A store's shipping cost doesn't react to other people watching the band. But one lesson does: the adaptive width, on its own, has not been shown to beat simpler baselines.

The statistical criticisms

  • "k" is not a probability. Non-normal data can raise the false-alarm rate of limits several times over [18]. Treat k as a sensitivity dial you tune against history, not as "95% confidence."
  • The SD breaks easily. One outlier inflates it; several outliers hide each other. Median and MAD resist up to half the window being bad [15].
  • The width is itself noisy. An SD estimated from 20 points has a relative error of roughly 16% even on clean data, so the effective false-alarm rate wanders from day to day.
  • Rolling bands forget. They're built to accept the new normal, which is exactly wrong when the new normal is a problem you want fixed.

Where experts disagree

The quality-control world has a real split. Douglas Montgomery's textbook tradition, following Alwan and Roberts (1988), says that when each value is correlated with the one before it, you should model that correlation and chart what's left over [17]. Donald Wheeler disagrees: correlations up to about 0.6 barely change the limits, and past that the pattern is usually obvious by eye. His line: whenever you're told you can't chart correlated data, "you are listening to an SPC novice" [8].

Wheeler also argues that Shewhart's 3σ limits were never meant as probability limits. They were chosen because they work on real data, whatever its shape [18]. For a KPI product, the practical upshot is the same from both sides: judge bounds by whether they give useful alerts on your history, not by a textbook percentage.

Monitoring practitioners land in the middle. Datadog offers three algorithms precisely because none wins everywhere: a fast rolling one, a seasonal forecasting one, and a decomposition one that holds steady through long anomalies [6]. Grafana's open framework uses mean ± 2 SD but adds a minimum band width so flat periods don't produce a zero-width band [21]. Prometheus's early guidance alerts only when a value is more than 2 SD above the mean and at least 20% above it, so tiny-but-"significant" wobbles don't page anyone [22]. LinkedIn's ThirdEye compares with the same weekday in prior weeks and treats holidays separately [23].

Verdict for KPI bounds

Use the idea, not the defaults. Bounds that come from each metric's own history are right for a product that watches many KPIs — nobody wants to hand-set thousands of thresholds. But plain 20-day, 2-SD Bollinger Bands will produce too many false alarms, stumble on weekly cycles, and let lasting problems fade into the baseline. Build the same-weekday, robust version, with a persistence rule and a minimum practical size, and tune k against history.

Check yourself

A teammate proposes 3σ bands instead of 2σ to cut false alarms. Argue for and against.

For: with many KPIs, 2σ produces a flood of false alarms (≈9/day for 200 metrics); 3σ cuts that roughly 17-fold, and people stop ignoring alerts. Against: small real changes will be missed for longer; a 1σ shift may take weeks to trip a 3σ band. Middle ground: 3σ for single-day flags, plus a persistence rule (e.g. 8 days above the centre) to catch small lasting shifts.

The academic trading studies found Bollinger Bands don't beat buy-and-hold. Does that mean they're useless for KPI bounds?

Not directly — trading returns depend on markets pricing in public rules, which KPIs don't do. The transferable lesson is narrower: the adaptive width alone hasn't beaten simpler baselines, so it has to earn its place against, say, a same-weekday median.

Making something with it Create

A proposed design for KPI bounds

Here's how I'd specify "normal range" for each KPI in a retail-analytics product. It keeps what's good about Bollinger Bands — self-adjusting, per-metric bounds — and fixes the failure modes from section 4.

  1. Classify each KPI first. Is it a total (orders, revenue), an average (cost per order), or a ratio with a denominator (split-order rate, margin %)? Ratios get denominator-aware bounds (step 6). Everything else follows steps 2–5.
  2. Baseline = same weekday. Centre line = median of the last 8 same weekdays, excluding flagged days and known holidays or promotions. With under 4 same-weekday values, don't alert; show "learning."
  3. Spread = robust. Spread = 1.4826 × MAD of those same values, with a floor (say 3% of the centre) so it never collapses to zero on flat data.
  4. Three kinds of flag:
    • Spike: today is more than 3.5 robust spreads from the centre.
    • Shift: 8 days in a row on the same side of the centre (Nelson-style), which catches small lasting changes that no single day shows [24].
    • Practical size: in either case, only flag if the move is also bigger than a minimum business-relevant size (e.g. ±5% or a set dollar amount), so statistically real but trivial moves don't create work.
  5. Protect the baseline. When a day is flagged, drop it from future baselines until someone marks it "expected" (e.g. a planned price change). That stops a problem from becoming normal.
  6. Ratios. Pool the baseline as total numerator ÷ total denominator over the window (not an average of daily ratios), then use \(\bar{p} \pm 3\sqrt{\bar{p}(1-\bar{p})/n_t}\). Skip days where \(n_t \bar{p} < 5\). If everything looks out of control at large volumes, widen by the observed spread (a Laney p′ adjustment) [16].
  7. Retail calendar. For year-over-year comparisons and holiday baselines, align on fiscal weeks in the retail 4-5-4 calendar, not calendar dates, so each comparison has the same weekday mix [25].
  8. Tune on history. Replay the last 6–12 months. Count flags per KPI per month and check a sample by hand. Set k so the total alert volume across all KPIs is something a person will actually read — for example, no more than a handful a day.

Written as a rule the engineering team could implement:

for each KPI, each store, each day t:
    hist   = values on the same weekday, weeks t-1 .. t-8, excluding flagged/holiday days
    if len(hist) < 4: status = "learning"; continue
    centre = median(hist)
    spread = max(1.4826 * MAD(hist), 0.03 * centre)
    z      = (x_t - centre) / spread
    spike  = abs(z) > 3.5 and abs(x_t - centre) > min_practical_change
    shift  = last 8 days all on the same side of their centres
    if spike or shift: flag(t); exclude t from future baselines

Extensions worth trying

  • Use %b as a feature, not just an alarm. A robust version of %b (the z above) puts every KPI on the same scale. That lets a recommendation engine rank "which store's shipping cost is most unusual today" across metrics with very different units.
  • Graduate to forecast intervals where it pays. For the handful of KPIs with strong holiday and promotional effects, a Holt-Winters or Prophet forecast with a 95–99% interval (Prophet's default is only 80%) will beat any rolling band [14]. Keep the simple method for the long tail of metrics.
  • Compare across stores on the same day. A funnel plot — each store's rate against limits that depend on its volume — finds the outlier store rather than the outlier day [16].

Open questions

  • Are bounds for alerting people, or constraints inside the recommendation logic? The second needs stability more than sensitivity.
  • How much history is there per store and KPI? Same-weekday baselines want 6–8 weeks; seasonal models want at least a year.
  • Who marks a shift as "expected," and how does that feed back into the baseline?

Check yourself

Sketch the rule for a split-order-rate KPI at 50 stores: inputs, parameters, what triggers a flag.

Inputs: daily split orders and total orders per store. Baseline: pooled rate over the last 8 same weekdays, excluding flagged days. Limits: p̄ ± 3√(p̄(1−p̄)/n_t) using that day's order count; skip days with n_t·p̄ < 5. Flags: outside the limits and at least 2 percentage points from p̄; or 8 days in a row above p̄. Across stores, a same-day funnel plot highlights which store is off.

Further reading

Sources

  1. John Bollinger, "Bollinger Bands" (bollingerbands.com)
  2. John Bollinger, "Bollinger Band Rules" (bollingerbands.com)
  3. StockCharts ChartSchool, "Bollinger Bands"; also %b and BandWidth pages
  4. NIST/SEMATECH e-Handbook, "EWMA Control Charts"
  5. NIST/SEMATECH e-Handbook, "CUSUM average run length"; "What are variables control charts?"
  6. Datadog, "Anomaly Monitor" documentation
  7. Stacey Barr, "How to build an XmR chart for your KPI"
  8. Donald J. Wheeler, "Autocorrelated Data," Quality Digest (2017)
  9. pandas documentation, Rolling.std
  10. Lento, Gradojevic & Wright, "Investment information content in Bollinger Bands?", Applied Financial Economics Letters (2007)
  11. Leung & Chong, "An empirical comparison of moving average envelopes and Bollinger Bands," Applied Economics Letters (2003)
  12. Fang, Jacobsen & Qin, "Popularity versus Profitability: Evidence from Bollinger Bands," Journal of Portfolio Management (2017)
  13. Park & Irwin, "What do we know about the profitability of technical analysis?", Journal of Economic Surveys (2007)
  14. Prophet documentation, "Uncertainty intervals"
  15. Rousseeuw & Hubert, "Anomaly detection by robust statistics," WIREs (2018); SAS blog, "The Hampel filter"
  16. p-chart (Wikipedia); Minitab, "Laney p′ charts"; funnel plots in health surveillance
  17. Alwan & Roberts, "Time-Series Modeling for Statistical Process Control," JBES (1988); Winkel & Zhang, NIST (2004)
  18. Donald J. Wheeler, "Probability Limits," Quality Digest (2015)
  19. Chebyshev's inequality (Wikipedia)
  20. Columbia Public Health, "False Discovery Rate"
  21. Grafana Labs, "How to use Prometheus to efficiently detect anomalies at scale"
  22. Prometheus blog, "Practical Anomaly Detection" (2015)
  23. LinkedIn Engineering, "Smart alerts in ThirdEye" (2019)
  24. Western Electric rules (Wikipedia); Wheeler, "Using Extra Detection Rules"
  25. National Retail Federation, "4-5-4 Calendar"