n = z² · p(1-p) / E². Runs convert observations into a prompt count. Engines that re-retrieve on every call (Perplexity, Google AI Overviews) need the top of the 8-12 run range.
The Wilson score interval is correct at the low rates most brands actually have, where the normal approximation breaks badly.
Your prompt count appears here.
Enter appearances and total responses to bound the rate.
Sample size is the part everyone skips
Running each prompt once and reporting the result as a rate is not a measurement, it is a coin flip recorded as a fact. Language models are probabilistic: the same prompt, sent twice, returns a different brand set. This planner does the two calculations that make a Share of Model number trustworthy, how many observations you need for a target precision, and how wide the confidence interval is on a rate you already measured.
How to use it
- Enter your expected inclusion rate, the margin of error you want to hold, your confidence level and how many times you run each prompt.
- Read the observations needed and the prompt count, then round to a 250 to 500 prompt portfolio.
- For a rate you already measured, enter appearances and total responses to get the Wilson interval.
- Report the number with the interval attached, never a bare point estimate.
How many runs per prompt do I need?
Eight to twelve per prompt per engine, minimum. At one run your estimate is either 0% or 100%; at three runs it can still be off by 30 points. The estimate only settles into an actionable range around eight to twelve runs, and high live-retrieval engines like Perplexity and Google AI Overviews need the top of that range.
How many prompts make a decision-grade portfolio?
To hold a two-point margin at a 27% rate you need roughly 1,900 scored observations per brand per engine, about 190 prompts at ten runs. Decision-grade programmes land at a 250 to 500 prompt portfolio. Precision flattens hard after about 2,500 observations, so more engines usually beats more prompts on a tight budget.
Why the Wilson interval instead of the normal approximation?
Because the normal (Wald) approximation breaks badly at the low inclusion rates most brands actually have, producing intervals that can dip below zero or overstate precision. The Wilson score interval stays correct at small n and extreme p, which is exactly the regime AI-visibility measurement lives in.
Does my data leave the browser?
No. Every calculation runs entirely in your browser. Nothing you enter is uploaded, stored or sent to any server.