SAME-DAY SHIPPING ON ORDERS BY 5PM ET. USA SHIPPED. FREE FEDEX SHIPPING ON ORDERS $250+. PREMIUM RESEARCH PRODUCTS • FOR RESEARCH USE ONLY.

FOR RESEARCH USE ONLY — NOT FOR HUMAN OR VETERINARY USE

Reading Preclinical Literature Critically

A preclinical paper reports what happened in one specified system, under one set of conditions, on one occasion. This article covers the reading practices that keep a finding attached to that system: the evidence hierarchy and why results do not transfer up it, measured endpoints versus outcomes, design and reporting standards, statistical interpretation, publication bias, the reproducibility record, journal provenance, and the limitations section.

Educational reference only. Nothing in this article describes, recommends, or makes any claim about American Alpha Labs products, and nothing here constitutes medical, dosing, administration, or preparation guidance.

The Ladder of Evidence, and Why Results Do Not Climb It

Preclinical evidence is conventionally described as a ladder: computational prediction, then biochemical and cell-based assays, then whole-animal experiments, then human studies. The metaphor is accurate about sequence and misleading about inference. Each rung is not a lower-resolution version of the rung above it. It is a different physical system, answering a different question, and a result obtained on one rung is a statement about that system only. Nothing in the structure of the ladder licenses carrying a conclusion upward; the ladder is a filter, not a conveyor.

Consider what each level actually constrains. In-silico work — computational modelling of several kinds — produces ranked hypotheses conditional on a force field, a protonation state, a chosen protein conformation or ensemble, and a scoring function whose known failure modes are documented in the computational chemistry literature. A high score is a reason to run an assay, not an experimental finding. In-vitro work introduces real molecules but removes almost everything else that an intact organism would impose on them, including the concentration ceiling that whole-organism physiology sets. Cell-culture concentrations are typically chosen to place a response inside the assay’s dynamic range, not to match anything an intact organism would present, and a compound active only at such concentrations has demonstrated chemistry rather than biology. Animal work reintroduces intact physiology at the cost of species difference, and adds a second layer of abstraction, because an animal model is usually an acute induction standing in for a chronic process.

The attrition record quantifies how much filtering remains after all of this. Analysing more than 400,000 clinical trial entries covering over 21,000 compounds from 2000 to 2015, Wong, Siah and Lo estimated that, in that dataset, a compound entering Phase I had a probability of roughly 14% of eventually reaching approval.1 Every one of those compounds had already cleared a preclinical package that its sponsors considered persuasive. The base rate for a preclinical positive result predicting a human outcome is therefore low, and it is low even for the subset that survived enough internal scrutiny to justify an investigational filing.

A useful illustration of how contestable translational claims are: in 2013, Seok and colleagues reported in PNAS that genomic responses in a set of mouse models poorly mimicked the human data they were compared against.2 In 2015, Takao and Miyakawa reanalysed substantially the same data in the same journal and reached the opposite headline conclusion, the difference resting largely on which genes were selected for comparison and how correlation was computed.3 The episode does not settle the question.It demonstrates that a stated resemblance between a model and its target is an analytical choice as much as an empirical finding, and that a reader who encounters such a claim should look for the specific evidence behind it.

Measured Endpoints Are Not Outcomes

An endpoint is what an instrument recorded. An outcome is what someone actually wanted to know about. The two are joined by an inferential chain, and a paper’s credibility depends on how many links in that chain were tested rather than assumed. A band on a Western blot is a statement about immunoreactive material of a given apparent mass in a lysate. A fluorescence ratio is a statement about a reporter. A dissociation constant is a statement about an interaction under a defined buffer, temperature and ionic strength. A volume measured at day 21, a behavioural score, a concentration at a single timepoint — each is a number produced by a specified procedure, and each is separated by several unmeasured steps from the biological process a reader may have in mind.

The clinical literature contains the canonical demonstration of what happens when a measured endpoint is treated as an outcome. Fleming and DeMets set out the general problem of surrogate endpoints in 1996;4 the Cardiac Arrhythmia Suppression Trial had already supplied the empirical case. Agents that moved the surrogate endpoint in the anticipated direction did not produce the anticipated clinical outcome.5 The measured quantity behaved as predicted. The outcome did not follow. Nothing about the mechanism was obviously wrong in advance; the surrogate simply was not the outcome, and only an experiment powered on the outcome could reveal it.

Two further mismatches are worth checking explicitly in any preclinical report. The first is concentration: an effect observed in culture at concentrations far above anything a whole animal would ever experience is an observation about the assay system, and the paper should say whether that was measured at all. The second is time: a short culture experiment or a two-week animal study is a short window, and a marker that moves within it may or may not correspond to anything that persists.

The practical habit is small and reliable. Before repeating a claim, write the sentence that the figure legend alone can defend — including the system, the conditions, the measured quantity with units, the timepoint, and the number of independent units. Then compare that sentence with the one in the discussion or the abstract. Where the second has more scope than the first, the extra scope was contributed by the authors’ interpretation, not by the data, and it should be carried forward, if at all, clearly labelled as such.

Design: Sample Size, Randomisation, Blinding, and Reporting Standards

Preclinical group sizes are usually small, and small studies fail in two directions at once. They miss real effects, which is well understood, and — less intuitively — the effects they do detect as statistically significant are systematically overestimated. Button and colleagues laid out the arithmetic in 2013: when power is low, the subset of results that clear a significance threshold is enriched for exaggerated estimates, so a low-powered literature produces both false negatives and inflated true positives.6 The consequence for a reader is that a striking effect size in a study with six animals per group should be treated as an upper bound, not a point estimate. Note also whether the sample size was justified in advance; a stated a-priori power calculation, with the assumed effect size and variance, is far more informative than the retrospective assertion that the study was adequately powered because the result was significant.

Randomisation and blinding address a different failure. Allocation to groups, order of processing, cage position on the rack, and time of day are all capable of generating differences on their own. Blinded assessment matters most where an endpoint involves judgement — microscopy scoring, behavioural rating, manual selection of regions of interest for densitometry. An overview of systematic reviews of animal studies found that studies with inadequate or unreported randomisation tended to report larger effect estimates than those with adequate randomisation, though the size and consistency of that difference varied across the reviews examined.7 The relevant question when reading is not whether the authors are trustworthy but whether the design made the result independent of their expectations.

Two reporting instruments exist for this. The ARRIVE guidelines, first published in 20108 and revised as ARRIVE 2.0 in 2020,9 specify what an animal study should report, with an “Essential 10” covering study design, sample size, inclusion and exclusion criteria, randomisation, blinding, outcome measures, statistical methods, experimental animals and procedures, and results. SYRCLE’s risk-of-bias tool, adapted from the Cochrane instrument for animal studies, provides a structured way to appraise a paper once read.10 Both are reporting and appraisal frameworks, not conduct standards: a paper that is silent on randomisation may still have randomised, but the silence leaves a reader unable to distinguish that case from the alternative, and unresolved ambiguity is the reader’s problem to price in.

Finally, check the unit of analysis. If the statistics were computed over wells, technical replicates, or individual cells drawn from a single culture or a single animal, the p-value describes that culture or that animal, however large the apparent sample. This is pseudoreplication, it is common, and it is usually detectable by comparing the n stated in the figure legend with the number of independent biological units described in the methods.

P-Values, Effect Sizes, and Intervals

A p-value is the probability, computed under a specified statistical model in which the null hypothesis holds, of obtaining data at least as extreme as those observed. That definition already excludes most of the uses it is put to. The American Statistical Association’s 2016 statement on p-values,11 and the longer catalogue of misinterpretations compiled by Greenland and colleagues the same year,12 set out the negatives explicitly: a p-value is not the probability that the null hypothesis is true, not the probability the result arose by chance, not a measure of the size or importance of an effect, and not — when above a conventional threshold — evidence that no effect exists. A study reporting p = 0.049 and one reporting p = 0.051 have produced nearly identical evidence, whatever the surrounding prose asserts.

What a reader needs instead is a magnitude and a range of values compatible with the data. An effect size with a confidence interval answers the question the p-value cannot: how large, and how precisely determined. In small preclinical studies the interval is frequently wide enough to include both a trivial and a substantial effect, and a paper that reports only asterisks has withheld the information required to notice that. If a claimed finding cannot be restated as “a difference of approximately X units, with an interval running from A to B, in n independent units,” the paper has not supplied enough for the claim to be evaluated, and that absence is itself informative.

Multiplicity is the third consideration. Count the comparisons the figures imply — several timepoints, several markers, several sample types, several groups — and ask how many tests were performed against how many were reported. Preclinical work is rarely preregistered, so the analysis plan is usually reconstructed by the reader from the methods section. Where the analysis appears to have been shaped after seeing the data, where subgroups appear without prior justification, or where a primary endpoint is not identified anywhere, the nominal p-values understate the true false-positive rate by an amount nobody can quantify.

The Literature You Cannot See: Publication Bias and the File Drawer

Rosenthal named the file-drawer problem in 1979: the published literature is a non-random sample of the experiments performed, because null and disconfirming results are less likely to be written up, submitted, and accepted.13 Every subsequent estimate of an effect drawn from published sources therefore inherits a censoring process that no amount of careful reading of an individual paper can detect.

The scale in preclinical work has been measured directly. Sena and colleagues examined 16 systematic reviews from a single field of animal research, covering 525 unique publications. Only ten of those publications — about 2% — reported no significant result on the primary endpoint. Statistical adjustment for the resulting asymmetry suggested that publication bias accounted for roughly a third of the effect size reported in that literature.14 A body of work in which 98% of papers report a positive finding is not describing a field in which almost everything works; it is describing a field in which almost everything that does not work goes unpublished.

Selective outcome reporting operates within papers as well as between them. Where several endpoints were measured and only some are reported, or where an endpoint declared in a methods section does not reappear in the results, the paper has performed the same censoring internally. Hypothesising after results are known produces prose in which an exploratory finding is presented as a confirmed prediction, and it leaves no trace a reader can identify with certainty.

The defensive reading practices are straightforward. Treat a single positive paper as one draw from a distribution you cannot see. Prefer systematic reviews that assess small-study effects — funnel plot asymmetry, trim-and-fill, Egger regression — over narrative reviews that count positive citations. Search deliberately for null results in the same model, including preprints and registered reports, and note when a search returns none at all in an area where many groups are working.

The Reproducibility Record and What It Implies for a Reader

Two industrial audits set the terms of the current discussion. Begley and Ellis reported in Nature in 2012 that an Amgen research group had attempted to confirm the findings of 53 papers it regarded as landmark work, and that the scientific findings were confirmed in six cases.15 Prinz, Schlange and Asadullah had reported in 2011 that across 67 in-house Bayer replication projects, published data were completely in line with internal replication attempts in roughly a fifth to a quarter of cases.16 Both figures are widely quoted and both deserve a caveat that is itself instructive: neither publication listed the papers examined or the replication protocols used, so neither can be independently checked. The most influential evidence about reproducibility is, in these two instances, not itself reproducible, which is a reasonable thing to hold in mind when reading them.

Survey data point the same direction with a different method. Baker’s 2016 Nature survey of 1,576 researchers found that a large majority reported having failed to reproduce another scientist’s experiment, and a majority reported having failed to reproduce one of their own.17

The most methodologically transparent evidence comes from a large systematic replication project whose overall findings Errington and colleagues published in eLife in 2021.18 The project set out to repeat 193 experiments from 53 high-impact papers and completed 50 experiments from 23 papers. The obstacles are the part a reader should absorb. None of the 193 experiments was described in sufficient detail in the original paper to allow a repetition protocol to be designed from that paper alone. The descriptive and inferential data needed to compute effect sizes and conduct power analyses were publicly accessible for four of the 193, and despite contacting the original authors the team was unable to obtain those data for 68% of the experiments. The dominant problem was not fabrication but underspecification: papers that do not contain enough information to be repeated cannot be confirmed or refuted, only believed or not.

The reading consequence follows directly. An unreplicated result is a lead. Independent replication — a different laboratory, ideally a different model system, with the original authors not involved in the analysis — is what converts a lead into something worth building on. When a finding has been cited hundreds of times, check whether the citations are replications or merely repetitions of the original claim, because a citation count measures circulation, not confirmation.

Venue, Limitations, and the One Habit Worth Keeping

Where a paper appeared is weak evidence about its content but strong evidence about the scrutiny it did or did not receive. A 2019 consensus definition, published in Nature by Grudniewicz and colleagues, characterises predatory journals as entities that prioritise self-interest at the expense of scholarship, marked by false or misleading information, deviation from best editorial and publication practices, lack of transparency, and aggressive indiscriminate solicitation.19 The cross-sectional comparison by Shamseer and colleagues in BMC Medicine identified concrete distinguishing features: a very broad or incoherent stated scope, promises of unusually rapid peer review, article-processing charges that are low or disclosed only after acceptance, non-professional contact details, spelling and grammatical errors on the journal site, editorial boards whose members cannot be verified or did not consent, and claimed indexing or impact metrics that do not check out.20

The verification steps take a few minutes. Confirm whether the journal is listed in the Directory of Open Access Journals and whether the publisher belongs to COPE or follows ICMJE recommendations. Distinguish genuine MEDLINE indexing from mere deposit in PubMed Central and from a claim of being “indexed in Google Scholar,” which is close to meaningless. Check any asserted impact factor against Journal Citation Reports rather than against the journal’s own page. Look up two or three named editorial board members and see whether they list the role themselves. A legitimate venue, however, is necessary rather than sufficient: retractions and corrections occur in the highest-profile journals, so check Retraction Watch, the Crossmark record, and PubPeer for the specific article before relying on it.

The limitations section is the most under-read part of most papers, and it repays being read early — before the discussion, sometimes before the results. Separate ritual limitations, which restate that further work is warranted, from load-bearing ones, which name a specific thing the study could not determine: no concentration measurements were taken, only one cell line was used, only one sex or one strain was studied, the observation period ended before the relevant timescale, the comparison group differed in a way that was not controlled. Then read for what the limitations section omits, particularly anything a careful reader noticed in the methods that the authors did not acknowledge. Finally, set the limitations paragraph beside the last sentence of the abstract. The gap between them is a measure of the paper’s own overreach, and it is usually the same gap that appears, widened, in any secondary account of the work.

The single habit that does the most work is naming the system before repeating the finding. Reconstruct the claim in full — this quantity, measured this way, in this system, under these conditions, over this interval, in this many independent units, with this uncertainty — and see whether it still says what you thought it said. If the reconstructed sentence becomes unwieldy, that is the diagnostic result, not an inconvenience: the shorter version was carrying scope the experiment never established. Results do not travel between systems on their own. The system travels with the result, or the result should not be repeated.

References

  1. Wong CH, Siah KW, Lo AW. Estimation of clinical trial success rates and related parameters. Biostatistics. 2019;20(2):273–286. doi:10.1093/biostatistics/kxx069; PMID 29394327.
  2. Seok J, Warren HS, Cuenca AG, et al. Genomic responses in mouse models poorly mimic human inflammatory diseases. Proc Natl Acad Sci USA. 2013;110(9):3507–3512. doi:10.1073/pnas.1222878110.
  3. Takao K, Miyakawa T. Genomic responses in mouse models greatly mimic human inflammatory diseases. Proc Natl Acad Sci USA. 2015;112(4):1167–1172. doi:10.1073/pnas.1401965111.
  4. Fleming TR, DeMets DL. Surrogate end points in clinical trials: are we being misled? Ann Intern Med. 1996;125(7):605–613. doi:10.7326/0003-4819-125-7-199610010-00011.
  5. Echt DS, Liebson PR, Mitchell LB, et al. Mortality and morbidity in patients receiving encainide, flecainide, or placebo: the Cardiac Arrhythmia Suppression Trial. N Engl J Med. 1991;324(12):781–788. doi:10.1056/NEJM199103213241201.
  6. Button KS, Ioannidis JPA, Mokrysz C, Nosek BA, Flint J, Robinson ESJ, Munafò MR. Power failure: why small sample size undermines the reliability of neuroscience. Nat Rev Neurosci. 2013;14:365–376. doi:10.1038/nrn3475; PMID 23571845.
  7. Hirst JA, Howick J, Aronson JK, Roberts N, Perera R, Koshiaris C, Heneghan C. The need for randomization in animal trials: an overview of systematic reviews. PLoS ONE. 2014;9(6):e98856. doi:10.1371/journal.pone.0098856; PMID 24906117.
  8. Kilkenny C, Browne WJ, Cuthill IC, Emerson M, Altman DG. Improving bioscience research reporting: the ARRIVE guidelines for reporting animal research. PLoS Biol. 2010;8(6):e1000412. doi:10.1371/journal.pbio.1000412.
  9. Percie du Sert N, Hurst V, Ahluwalia A, et al. The ARRIVE guidelines 2.0: updated guidelines for reporting animal research. PLoS Biol. 2020;18(7):e3000410. doi:10.1371/journal.pbio.3000410.
  10. Hooijmans CR, Rovers MM, de Vries RBM, Leenaars M, Ritskes-Hoitinga M, Langendam MW. SYRCLE’s risk of bias tool for animal studies. BMC Med Res Methodol. 2014;14:43. doi:10.1186/1471-2288-14-43.
  11. Wasserstein RL, Lazar NA. The ASA statement on p-values: context, process, and purpose. The American Statistician. 2016;70(2):129–133. doi:10.1080/00031305.2016.1154108.
  12. Greenland S, Senn SJ, Rothman KJ, Carlin JB, Poole C, Goodman SN, Altman DG. Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations. Eur J Epidemiol. 2016;31:337–350. doi:10.1007/s10654-016-0149-3.
  13. Rosenthal R. The file drawer problem and tolerance for null results. Psychological Bulletin. 1979;86(3):638–641. doi:10.1037/0033-2909.86.3.638.
  14. Sena ES, van der Worp HB, Bath PMW, Howells DW, Macleod MR. Publication bias in reports of animal stroke studies leads to major overstatement of efficacy. PLoS Biol. 2010;8(3):e1000344. doi:10.1371/journal.pbio.1000344; PMID 20361022.
  15. Begley CG, Ellis LM. Drug development: raise standards for preclinical cancer research. Nature. 2012;483:531–533. doi:10.1038/483531a; PMID 22460880.
  16. Prinz F, Schlange T, Asadullah K. Believe it or not: how much can we rely on published data on potential drug targets? Nat Rev Drug Discov. 2011;10:712. doi:10.1038/nrd3439-c1; PMID 21892149.
  17. Baker M. 1,500 scientists lift the lid on reproducibility. Nature. 2016;533:452–454. doi:10.1038/533452a.
  18. Errington TM, Mathur M, Soderberg CK, Denis A, Perfito N, Iorns E, Nosek BA. Investigating the replicability of preclinical cancer biology. eLife. 2021;10:e71601. doi:10.7554/eLife.71601.
  19. Grudniewicz A, Moher D, Cobey KD, et al. Predatory journals: no definition, no defence. Nature. 2019;576:210–212. doi:10.1038/d41586-019-03759-y.
  20. Shamseer L, Moher D, Maduekwe O, Turner L, Barbour V, Burch R, Clark J, Galipeau J, Roberts J, Shea BJ. Potential predatory and legitimate biomedical journals: can you tell the difference? A cross-sectional comparison. BMC Med. 2017;15:28. doi:10.1186/s12916-017-0785-9.

All products supplied by American Alpha Labs are for laboratory research use only. Not for human or veterinary use. Not for diagnostic or therapeutic use.