The question, in words
Drop random points in a square, count how many land inside a quarter circle, and you get an estimate of pi. With more points the estimate gets better. Does its error shrink as one over the square root of the number of points, so that four times the points buys half the error? That is the lab's first question.
What we expected, and why
Each point is an independent yes or no, so the estimate is a scaled average of coin flips, and its spread should follow the square-root law. On a log-log plot the error should fall on a line of slope -0.5, the value the protocol tests. The protocol fixed, before the run, the rules for keeping or dropping that idea, a kill rule for an error that does not shrink at all, and four predictions: the slope close to its theoretical value, the error at each size close to the theoretical value, a small error at the largest size, and a mean estimate close to pi there.
What we did
We ran the frozen estimator at five sample sizes, each four times the last, with one hundred seeds per size, and measured the root mean square error against pi at each size. The one thing varied was the number of points. There was one run, and nothing was voided or re-run.
What happened
The RMS error fell with a fitted log-log slope of -0.493 (R-2), against -0.5 in theory.
The first decision rule fired, so the hypothesis is kept. The other two did not fire, the kill rule included. Two predictions hit: the slope, and the small error at the largest size. Two missed. The error scaled by the square root of the sample size left the 10% band around its theoretical value that the protocol set, at one size. And at the largest size the mean estimate was further from pi than we had predicted.
An earlier result, R-1, since superseded, fitted the slope to the mean absolute error where the protocol names the RMS error.
What we learned
The square-root law holds for this estimator across the range we ran, and the slope is a robust way to see it: it pools all five sizes. The two misses come from asking too much of one size at a time. With one hundred seeds, an RMS error is itself only known to within several percent, and a mean estimate to within about a thousandth, so the band and the distance from pi we predicted were tighter than the run could promise.
What this does not show
- Anything about another estimator, another random generator, or estimates in more than two dimensions.
- Any sample size outside 64 to 16,384 points.
- Whether the estimator is biased: no result measures the bias, and one mean estimate missing its prediction is not evidence of one.
- Running time, which was not measured.