In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

When a constant arrival rate misses the pattern

A Poisson process makes a strong claim: counts in equal-length windows share a distribution and disjoint-window counts are independent. How much does a constant average hide in an actual event sequence?

This case examines 700 cataloged earthquakes of magnitude at least 5 during March 2011, from the USGS us catalog, globally. It is a historical model-checking example. The count and gap plots use observed catalog records, not simulated data. They do not provide earthquake forecasts.

Observed daily earthquake counts and fitted expected counts05010015020025030051015202530March 2011: UTC daycatalog events (magnitude ≥5)

Blue bars: observed catalog counts. Gold line: the selected model's fitted mean. The line is an expectation, not a confidence band.

QuantityValue
Events / observed days700 / 31
Overall daily mean22.5806
Sample variance of daily counts2570.5849
Variance / mean113.8402
One-rate residual deviance1403.452
Two-rate residual deviance1203.561
Days 1–10 fitted daily rate6.8000
Days 11–31 fitted daily rate30.0952

Both models are fitted to these same days. Two fitted rates add flexibility; a smaller in-sample deviance is not evidence of better forecasts or independent arrivals. Moving the split after seeing the data is exploratory. The calculation does not identify a causal effect.

The count distribution loses the time order

Observed daily-count frequencies compared with fitted Poisson probabilities00.050.10.15050100150200250events in a dayfraction / model probability

All 31 days and their largest count are included. Gold joins discrete probabilities as a visual guide. With two rates, it is their duration-weighted Poisson mixture. Neither this histogram nor matching its variance can establish independent increments.

Interarrival gaps give another view

Empirical interarrival CDF and exponential CDF at the fitted overall rate00.20.40.60.81020406080100120140160180gap between catalog origins (minutes)cumulative fraction

Blue: empirical CDF of all 699 consecutive-origin gaps. Gold: Exponential(rate 0.015681 per minute), using the one-rate fit. 71 gaps extend beyond this display; the empirical denominator stays 699. This is a descriptive comparison, not a calibrated goodness-of-fit test. First and last observation-boundary gaps are excluded from the completed interarrival sample.

Inspect all daily counts and fitted means
UTC dateObservedFitted mean
2011-03-01322.5806
2011-03-02222.5806
2011-03-03622.5806
2011-03-04422.5806
2011-03-05422.5806
2011-03-06722.5806
2011-03-07422.5806
2011-03-08322.5806
2011-03-092822.5806
2011-03-10722.5806
2011-03-1128122.5806
2011-03-128622.5806
2011-03-133922.5806
2011-03-143122.5806
2011-03-152122.5806
2011-03-161822.5806
2011-03-172222.5806
2011-03-181522.5806
2011-03-191122.5806
2011-03-202122.5806
2011-03-21622.5806
2011-03-222622.5806
2011-03-23722.5806
2011-03-24722.5806
2011-03-25822.5806
2011-03-26722.5806
2011-03-27622.5806
2011-03-28422.5806
2011-03-29522.5806
2011-03-30522.5806
2011-03-31622.5806
Pinned USGS catalog observations, not simulated counts. Credit: U.S. Geological Survey. The models are teaching comparisons fitted to this historical window.

Define exactly what was observed

The pinned query selects catalog=us, eventtype=earthquake, and minmagnitude=5, from 1 March 2011 at 00:00 UTC through 1 April at 00:00 UTC. The service includes both endpoints; preparation explicitly keeps start ≤ origin time < end. This snapshot has no record at the excluded endpoint, so all 700 input records remain.

The timestamp is the estimated earthquake origin time, not when a server received a report. Records are sorted by that field. Each UTC calendar day has 24 hours here; all 31 days are represented, including a zero-count day if one existed. The sample is the specified catalog extract, not a guaranteed census of every physical earthquake.

Event IDs are unique, and time, magnitude, event type, and source network are present for every retained row. Other columns have missing measurements: all 700 rows lack dmin, horizontal error, and magnitude error; 14 lack RMS, 223 lack depth error, and 297 lack magnitude-station counts. Those fields are not required for the displayed count/gap calculations and are not filled with invented measurements.

The magnitude field combines documented magnitude types (mb, ms, mwb, mwc, mwr, mww), with magnitude sources us and gcmt. The threshold follows the catalog’s reported value; it does not reconcile scales or prove uniform detection completeness. Catalog revisions can also change historical records, which is why this lesson pins its retrieved bytes.

Fit one daily rate, then inspect what it misses

The constant-rate estimate is 700/31 = 22.5806 events per day. Under the fitted Poisson count model, daily mean and variance would both equal that rate. The observed sample variance, using denominator 30, is 2570.5849: a variance-to-mean ratio of 113.8402.

That ratio is a descriptive mismatch, not a proof of one specific mechanism. The time plot shows information the ratio discards: 281 selected records have origin times on 11 March, followed by 86 on 12 March. A constant horizontal mean does not represent this burst and decline.

The window contains the 11 March 2011 Tohoku earthquake. That historical context motivates inspecting change around the date. This global extract does not itself classify every selected event as a foreshock, mainshock, or aftershock, and the date split does not estimate a causal effect.

Make a prediction

If you shuffled the order of the 31 daily counts, would their histogram and sample variance change?

Explore the answer

No. Both would stay the same, while the time pattern could change completely. Matching a marginal count distribution is not enough to establish stationary or independent increments.

A more flexible comparison still has limitations

The two-rate model fits one rate to 1–10 March and another to 11–31 March. The estimates are 6.8 and 30.0952 per day. It represents a change in average frequency, but still assumes constant rates within each segment and independent Poisson daily counts.

The residual Poisson deviance sums 2·[c·log(c/m)−c+m] over days, using c·log(c/m)=0 when c=0. Here c is the observed count and m its fitted mean. It falls from 1403.452 for one rate to 1203.561 for the default two-rate fit. The second model has an additional rate parameter; improved in-sample fit is expected from its flexibility.

Move the split to explore that dependence. The two rates and deviance update, but no held-out forecast score or calibrated significance level is produced. A split chosen after seeing the counts also uses information from the same data being assessed. The large burst and subsequent decline remain more detailed than a two-level mean can describe.

The count histogram overlays either one Poisson distribution or the duration-weighted mixture of the two fitted Poisson distributions. Total variance explains why such a mixture can have variance above its mean. It does not explain away serial dependence or establish that the mixture is adequate.

Check gaps without changing their denominator

There are 699 gaps between consecutive retained origin times and no tied origin times in this snapshot. The blue curve is their empirical CDF. The gold curve uses an exponential rate of (700/31)/1440 per minute, the same constant-rate estimate as the count model.

Changing the display limit hides long gaps outside the plot; it does not renormalize the remaining gaps. The first boundary-to-origin interval is 53.7723 minutes, and the last origin-to-boundary interval is 264.9352 minutes. Neither is included as a complete consecutive-event gap. A finite-window sample of completed gaps also has boundary-selection effects. This comparison is therefore descriptive, with an estimated rate, not a ready-made exponential goodness-of-fit test.

Make a prediction

Would fitting a negative-binomial histogram by matching the observed mean and variance prove that earthquake arrivals are independent?

Explore the answer

No. It could improve the marginal count fit while leaving the time dependence unexplained. An adequate process model has to account for the order and timing as well as count frequencies. Fitting and checking are separate tasks.

Reproduce from the pinned input

Download the original CSV snapshot, source and retrieval metadata, and preparation script. Place the files together, then run:

uv run earthquake-arrivals.py --input earthquakes-2011-03.csv --output earthquake-arrivals.json

The script pins Python 3.13 and Polars 1.34.0, verifies the source checksum, validates required fields and unique IDs, builds the complete daily calendar, and computes the gaps and reported statistics. It stops if the source bytes differ; it never downloads changing data during a page visit.

Source SHA-256: def12dc84a8ce6cbaee42d9cc31017c171a7f2a72a43f54efa11ddfb37ce4b62.

Retrieved 12 September 2026. Credit: U.S. Geological Survey, National Earthquake Information Center, with contributing magnitude sources retained in the CSV. The USGS public-domain notice covers USGS-produced data; this lesson republishes numerical catalog records and its own analysis, with no third-party images or narrative copied. See the catalog query documentation and ComCat field definitions for the source conventions.

Return to learning paths and foundations, or revisit Poisson counts and rate uncertainty with the difference between fitting and checking in mind.

Reset all settings