Pseudoreplication
Cells from one donor are not independent samples. Testing 40,000 cells as 40,000 observations when they came from four donors inflates the p-value massively.
Point a general-purpose AI at your single-cell analysis and it cries wolf on 74% of clean results. Redline re-runs the load-bearing statistics on your own data and cries wolf on none of them, then marks the real false discoveries on the figures you already made.
Live demo. It runs on a locked fixture with zero API keys. Skip to the workbench →
Claim, struck through: IL2RA knockdown significantly increased FOXP3 expression (p < 0.001, n = 51,842).
Corrected: IL2RA knockdown did not significantly change FOXP3 expression at the donor level (Welch’s t, p = 0.21, n = 4 donors).
Your Reviewer 2, on the bench.Four founding pillars and four rigor checks, every one on the same module interface. Each finding names the failure mode, cites the method paper that fixes it, and shows the corrected result beside the claim.
Cells from one donor are not independent samples. Testing 40,000 cells as 40,000 observations when they came from four donors inflates the p-value massively.
Clusters defined on the data and then tested for their own marker genes on that same data manufacture false positives. It is the default in standard pipelines.
The biological story often rides on an arbitrary clustering resolution the scientist never justified. Move the knob and the cell state can vanish.
The comparison of interest can be inseparable from a technical variable, for instance when treated and control samples ran on different days.
Calling genes significant on raw p-values across thousands of tests inflates false discoveries.
A batch or covariate that is separable from the effect but omitted from the model can carry a spurious result.
A cluster count chosen without a stability criterion is a story built on a default nobody justified.
A test whose assumptions the data violate reports a number that does not mean what it claims.
services/rigor/bench.Everything Redline asserts is shown, reproducible, and cited. The corrected code is downloadable and runs, and the preview is its output.
When a design is unsalvageable, Redline says so plainly and shows no corrected result anywhere. The contract refuses to carry a fix that cannot exist.
A clean analysis is a real answer. A passed check renders as Verified in green, stated with the same confidence Redline gives a flag.
A foundation step resolves the design, an agent proposes the claims, and the registered checks run on the roles you confirmed. A ComputeTarget seam decides where the statistics actually run, behind one return contract, so the interface never changes.
Redline reads your obs columns and proposes the design: which is the biological replicate, which is the comparison, which are technical nuisances. Nothing runs until you confirm it.
An agent inspects your stored results and proposes each auditable claim, already routed to the checks that can test it. You confirm, edit, or remove the list.
Each confirmed claim runs its checks. Every finding is numbers, a named failure mode, a citation, and a conclusion rewritten in defensible language.
A plain-English report with a citation behind every call, plus a downloadable bundle of runnable Python that reproduces the honest re-analysis.
Everything is open and configurable through env vars, with no hidden paths. The core drops into a browser, an agent, Claude Science, or a downloadable script that runs on your laptop.
A plots-first workbench that renders your figures and marks each finding on them. One panel per check, every knob exposed, the corrected result shown beside the claim.
Every check is an independent MCP tool behind one return contract. The same rigor drops into any agent or your own pipeline without touching the driver.
The same engine packaged as a Claude Skill, so it loads natively into Claude Science and runs on a scientist’s own data the day the hackathon ends.
For every flagged check you download runnable Python that reproduces the honest re-analysis: a README, a consolidated notebook, one script per finding. What you saw is the output of that code.
@redline/contracts holds the Zod shapes every surface speaks. The fixture, the Python engine, the reasoning layer, and the UI all agree on one contract, so a finding means the same thing everywhere it lands.Redline is dataset-agnostic, but it is validated against the Marson and Pritchard genome-scale CD4+ T-cell Perturb-seq data, Gladstone’s flagship single-cell resource, with raw counts so the re-runs are real.
One hard rule. The authors did their analysis rigorously, and there is no error in their published work to catch. Redline audits a naive foil instead, the standard cluster-then-annotate-then-DE workflow a less-experienced scientist would run on the same data. Pointed at a clean analysis, Redline reports clean.
Drop in the data you analyzed and the analysis you ran. The demo runs on a locked fixture with zero cloud credentials, then point it at the Python engine to run the real statistics on your own data.