How to run a check 01 Pick a panel from the tab strip or open its address; the four panels share one verdict language. 02 Enter the claim. Equivalence takes expressions A and B; Answer Check takes a problem, your answer and the variable, plus an optional dimensional check. 03 Paste evidence, not prose. Data Check takes one numeric column separated by commas, spaces or newlines. Number Audit takes N, M, SD and decimals, then t, df and reported p; two groups' N/M/SD recompute the independent t. 04 Run the check with the panel button; nothing fires on stale input. 05 Read the card verdict, trust badge, method, evidence and warnings, then add it to the report if useful.
Worked readouts Every figure below comes from the panels' demo inputs; press a check button to reproduce them.
Panel
Input
Readout
Equivalence
x^2 - 1 ≟ (x-1)(x+1)
✓ Equivalent, Exact — A−B simplifies to 0 structurally; 32/32 sample points agree
Answer Check
x^2 - 5x + 6 = 0, answer x=2, x=3
✓ Correct, Verified — residual 0 for both roots; engine reference x = 2, x = 3
Data Check
SPC demo, 11 values ending 15
✗ Fail (4 triggers): CL 10.4727 , UCL 12.3338 , LCL 8.61165 , σ̂ 0.620359 ; point 11 at z = 7.30 ; two 9-point runs below the centerline
Number Audit
N 28 , M 4.63 , SD 1.21 , 2 decimals
GRIM impossible — nearest 4.64 / 4.61 ; SD not assessable; t(27) = 2.31, p = 0.03 reads 0.0288 two-tailed
All six two-group fields (n₁ = 28, M₁ = 4.63, SD₁ = 1.21; n₂ = 26, M₂ = 3.94, SD₂ = 1.12) give t(52) = 2.1698 , p = 0.0346 .
Three verdicts, four trust levels The verdict separates "no" from "not proven". Fail needs an explicit witness — a counterexample point, a residual beyond tolerance, a flagged control point. Pass needs the strongest available evidence. Conditional covers the middle: domains that differ on one side, an identity no proof reaches, a missing root, a sample too small to decide. The trust badge describes the method: Exact is symbolic or exact-rational work, Verified is a numeric result carrying a residual or witness, Approximate is unquantified floating point, Heuristic is sampling that cannot prove anything. Numeric agreement alone never earns Equivalent: x + 1e12 and x + 1e12 + 1 sample 32/32 yet read Unconfirmed.
Caps and scope
Input and size caps. I-MR charts need at least three observations; Anderson–Darling normality needs at least eight and rejects zero variance; reference-interval checks need an upper above the lower. GRIM loses all power once N ≥ 10^decimals (28 is fine at 2 decimals), and simplified GRIMMER skips above N = 200. Answer Check accepts numeric or constant roots, and confirms "no real solution" only when the engine solves the equation. Near 10¹², a genuine +1 gap sits inside sampling resolution and stays Unconfirmed. What it does not do: it does not show step-by-step derivations, fit models, or prove that reported data are genuine — GRIM/GRIMMER and the Nelson rules test self-consistency, not authenticity. For step-by-step solving use the scientific calculator ; for tests and summaries from raw data use Statistics ; for study design and control-chart setup use Applied Statistics .
The residual is not the error The signature trap is a tiny residual on a huge scale. For 100000x = 1e12 answered x = 1e7+1, the raw residual is 100000 ; against the local scale of 10¹² it normalizes to 1.0e-7 — inside the 1e-6 residual tolerance, so a residual-only check would accept the answer. The second arm converts the residual into a variable-space error, the Newton step |f(x̂)/f′(x̂)| = 1 , far beyond the 0.01 cap; the card reads Incorrect and reports the missing root 10000000 . Rounding is treated differently: 3x = 1 answered x = 0.333333333 passes with a root error of 3.333e-10 , while x = 0.33 fails with residual 0.01 and root error 0.003333 . A small |f(x̂)| is not a small error in x.
Where it helps Marking or checking a solution (x-6)/(x-6) simplifies to 1 yet is only conditionally equivalent: at x = 6 only the right side is defined. Answer Check gives per-root evidence — x=2, x=3 reads residual 0 while x=2 alone is missing root 3 — and d/dx(x³ + sin x) = 3x² + cos x is Exact.
Screening a data column The SPC demo's spike is caught three ways: UCL 12.3338 , point 11 at z = 7.30, two 9-point runs below the centerline. A clean 12-value sample passes normality (p = 0.995 , small-sample warning); append 25 and it fails (A²* = 6.93 , p = 5.82e-17 ) at point 20, z = 4.25.
Auditing a paper's numbers The Number Audit defaults already fail GRIM — N = 28 cannot round to M = 4.63 — which blocks the SD check too. t(27) = 2.31 with p = 0.03 is consistent two-tailed (0.0288), but reading it as one-tailed p = 0.045 is off by 0.0306, beyond the 0.005 tolerance.
Privacy All four panels compute inside your browser; nothing you type or paste is uploaded.
References
Brown N.J.L. & Heathers J.A.J., The GRIM test , Social Psychological and Personality Science 8(4):363–369 (2017), doi.org (访问日期:2026-10-01)— granularity limits on reported means.
Anaya J., The GRIMMER test , PeerJ Preprints 4:e2400v1 (2016), doi.org (访问日期:2026-10-01)— SD feasibility.
Nelson L.S., The Shewhart Control Chart — Tests for Special Causes , Journal of Quality Technology 16(4):237–239 (1984), doi.org (访问日期:2026-10-01)— the eight detection rules.
NIST/SEMATECH, e-Handbook of Statistical Methods , §6.3.1 What are Control Charts?, itl.nist.gov (访问日期:2026-10-01)— 3σ limits.
NIST/SEMATECH, e-Handbook of Statistical Methods , §1.3.5.14 Anderson-Darling Test, itl.nist.gov (访问日期:2026-10-01)— normality testing.
NIST/SEMATECH, e-Handbook of Statistical Methods , §1.3.6.6.4 t Distribution, itl.nist.gov (访问日期:2026-10-01)— tail probabilities.
Calculators in this hub
Hand-picked tools, one click away. The mini versions compute live and carry your values into the full calculator.
Sources & review
Reviewed by CalcX Editorial Team
Updated 2026-10-01