Worked example · EX:statistics-estimation-and-uncertainty/audit-juniper-source-contract

Audit Juniper's unit, population, frame, and sample

Turn the invoice packet into a reproducible source contract before computing a statistical conclusion.

Updated Aug 7, 2026 Review due Nov 7, 2026
On this page
  1. Start from the requested claim
  2. Audit the row and dictionary
  3. Reconcile the frame
  4. Preserve the stop conditions
Worked-example setupScope and assumptions
  • Every entity, record, date, selection event, and value is fictional.
  • The packet stipulates independent simple random selection without replacement after construction of each year-specific frame.
  • The source audit evaluates the stated contract; it does not prove that an equivalent real extract would satisfy it.
Period
Fictional calendar years 2025 and 2026; extract dated February 15, 2027
Units
Invoice records, whole calendar days, and unitless selection probabilities
Rounding
Counts exact; selection probabilities shown to four decimals; statistics retain full precision

Start from the requested claim

“Did customers pay faster?” contains at least four undefined terms. Replace it with a target:

Estimate and describe invoice-weighted calendar days from issue through final payment application among ordinary external trade-credit invoices both issued and fully settled within each named year.

That sentence excludes open invoices. Record the limitation now rather than after the result is known.

Audit the row and dictionary

One row represents one qualifying invoice. Verify a unique invoice key before joining payments; otherwise installments could duplicate the unit.

Field Role Scale Permitted first summaries
invoice_id Opaque identifier Nominal uniqueness and count
customer_segment Category Nominal counts and proportions
review_priority Ordered label Ordinal ordered counts; positional summaries with care
days_to_settlement Quantitative response Ratio distribution, mean, median, spread

An average invoice ID or average segment code would be numerical output without measurement meaning.

Reconcile the frame

For each year, retain the query, source-system version, extraction timestamp, eligibility predicates, excluded-record counts, record count, and control amount. Test:

  • one row per eligible invoice;
  • issue and final-application dates present and in the named year;
  • external trade-credit classification;
  • no cash sale, credit memo, intercompany, duplicate, or unmatched payment;
  • frame count of 480 for Year 1 and 520 for Year 2; and
  • sample IDs all belong to the corresponding frame exactly once.

The selection probabilities are (16/480=0.0333) and (16/520\approx0.0308). Equal probability within each year does not imply equal probability across the combined years, and the analysis does not pool them.

Preserve the stop conditions

Proceed with sample description because the learner packet stipulates the controls. In real work, stop or narrow the claim if frame totals do not reconcile, selected IDs cannot be reproduced, key dates are missing, or the row grain changes after a join.

The typed result confirms 16 finite observations per year and retains the 78-day value. That is computation evidence, not source-authenticity evidence.

Verified calculation · descriptive estimation analysis

The curriculum loader recomputed this example before it entered the site build. Expand any structured input to inspect the stated facts.

comparisons
1 field
Inspect data
{
  "year_2_vs_year_1": {
    "direction": "right-minus-left",
    "left": "year_1",
    "right": "year_2"
  }
}
series
2 fields
Inspect data
{
  "year_1": {
    "dispersion_convention": "sample-n-minus-one",
    "inference_basis": "simple-random-sample",
    "mean_confidence_interval": {
      "confidence_level": 0.95,
      "critical_value": 2.131449545559323,
      "degrees_of_freedom": 15,
      "interval_scope": "two-sided-one-sample-mean",
      "reference_distribution": "student-t-supplied"
    },
    "missing_value_policy": "reject",
    "observations": [
      22,
      25,
      26,
      27,
      28,
      29,
      30,
      31,
      32,
      33,
      34,
      35,
      36,
      38,
      40,
      44
    ],
    "outlier_policy": "retain-and-flag",
    "population_role": "sample",
    "quartile_method": "median-of-halves-exclusive"
  },
  "year_2": {
    "dispersion_convention": "sample-n-minus-one",
    "inference_basis": "simple-random-sample",
    "mean_confidence_interval": {
      "confidence_level": 0.95,
      "critical_value": 2.131449545559323,
      "degrees_of_freedom": 15,
      "interval_scope": "two-sided-one-sample-mean",
      "reference_distribution": "student-t-supplied"
    },
    "missing_value_policy": "reject",
    "observations": [
      18,
      20,
      21,
      22,
      23,
      24,
      25,
      26,
      27,
      28,
      29,
      30,
      31,
      33,
      35,
      78
    ],
    "outlier_policy": "retain-and-flag",
    "population_role": "sample",
    "quartile_method": "median-of-halves-exclusive"
  }
}

Recomputed result

Values recomputed by the curriculum loader
MeasureValue
1 count16
1 mean31.875
2 count16
2 mean29.375
2 potential outlier count1