On this page
Lesson details
- Estimated study time
- 2 hr
Learning objectives (14)
Juniper's controller asks, “Did customers pay faster in Year 2?” The question sounds numerical. It is not ready for arithmetic.
What counts as a customer? Does “pay” mean first receipt, bank settlement, or final application? Are open invoices included? Does one row represent an invoice, invoice line, payment, or customer? A precise answer to the wrong version of the question is still wrong.
Build the seven-part contract
Use the Juniper source audit as you work through these fields:
| Part | Juniper declaration | Why it changes the result |
|---|---|---|
| Unit | One qualifying invoice | Prevents payment or line rows from receiving unintended weight |
| Variable | Whole days from issue through final payment application | Fixes the start, endpoint, unit, and precision |
| Population | Ordinary external trade-credit invoices issued and fully settled within each year | Excludes open invoices and therefore narrows the claim |
| Frame | Deduplicated list of every qualifying invoice after cutoff | Makes coverage and duplicate tests possible |
| Sample | 16 IDs selected without replacement within each year | Separates observed rows from the target population |
| Source and cutoff | AR subledger plus cash-application log, extracted February 15, 2027 | Controls versions, dates, and late changes |
| Policies | Reject missing dates; retain and investigate valid unusual values | Prevents silent case deletion |
Predict the error before reading further: if the source join produces one row per payment, which invoices receive extra weight? Any invoice paid in several installments. That changes the observational unit even if the final column is labeled invoice ID.
Population, frame, and sample are three objects
The target population answers what the question is about. The frame is what the selection mechanism can reach. The sample is what was selected and observed.
A random sample protects selection only within its frame. If late-posted invoices are missing, cash sales remain, or one invoice appears twice, the sample can be random and still point at the wrong target. Reconcile frame counts and amounts to the source, inspect exclusions, test unique keys, and freeze the query and extraction time.
Depth checkpoint: repeated customers
Suppose one customer generates 30% of the invoices. Invoice-level simple random selection still gives each invoice equal probability, but invoices from that customer may share terms, disputes, or payment processes. A defensible graduate response would say that equal selection probability does not establish independent outcomes: repeated-customer concentration could make 16 invoices carry less independent information than 16 unrelated invoices. Request the customer key and invoices-per-customer distribution before choosing a variance method. Computing a cluster-adjusted variance is deferred to a later module.
Classify the variables
Juniper's dictionary contains:
- invoice ID: nominal identifier;
- customer segment: nominal category;
- review priority: ordinal category; and
- days to settlement: ratio-scale quantitative response.
Digits do not make an identifier quantitative. An ordinal code supports order, not equal numerical gaps. Complete the scale check before selecting a mean or standard deviation.
Separate sampling and nonsampling error
Sampling variability arises because only some frame units are selected. Nonsampling error includes undercoverage, duplicates, wrong dates, missing values, classification, cutoff, join, and processing failures. Increasing the sample size addresses only part of that list.
Exit check
Write two sentences. The first defines the population mean Juniper might seek. The second explains why the supplied target, settled-within-year invoices, does not answer a question about ending open receivables. If either sentence omits the unit, variable, or period, repair it.