← All posts

2026-09-07

Why bioinformatics pipelines fail quietly

Four incompatible genomes all called hg19, strandedness flags that silently halve counts, TPM matrices fed to count-model tests: the documented silent errors behind BioNodulo's contract work.

The dangerous failures in bioinformatics are not the crashes. The tools install, the run completes, every process exits zero, and the output looks plausible. The result is wrong anyway, because two steps disagreed about something neither of them reported: which reference genome, which strand convention, which normalization state.

Four genomes, one name

GATK's own documentation records that at least four mutually incompatible reference builds are all colloquially called hg19: UCSC hg19, b37 (the GATK bundle's Homo_sapiens_assembly19), GRCh37, and humanG1Kv37 differ in contig names, decoys, and unplaced sequences. Files built on different variants are coordinate-compatible but dictionary-incompatible, so the error is invisible until two files happen to reach the same tool, and annotation files on the wrong build never trigger any error at all.

This is not a theoretical risk. Li and colleagues re-called 1,572 exomes on both GRCh37 and GRCh38 and found that 1.5 percent of SNVs and 2.0 percent of indels were discordant per build, with 206 genes enriched for discordance, including eight Mendelian disease genes. A family with a diagnostic variant in one of those genes can get a different answer depending on which build the pipeline happened to mix in. Nothing crashes.

The flag that halves your reads

RNA-seq quantification tools encode library strandedness in three mutually incompatible CLI conventions: htseq-count uses yes, no, reverse; featureCounts uses 0, 1, 2; RSEM uses a forward probability. The HTSeq documentation itself warns that the wrong setting does not error, it just drops or contaminates roughly half the reads. A dedicated analysis in Briefings in Functional Genomics worked through the consequences: plausible differential expression results computed from counts that are wrong by a factor near two, with no signal anywhere in the logs.

Counts, TPM, and honest tests

Normalization state is the third quiet killer. Count-model differential expression tools such as DESeq2 and edgeR require raw counts; TPM and FPKM inputs produce plausible gene lists with wrong statistics. Wagner and colleagues proved the mathematical reason in 2012 (within-sample length normalization is inconsistent across samples), and a 2020 paper in RNA documented how widespread the unintentional misuse still is. The spreadsheet world has its own version: gene symbols silently converted to dates in published supplementary files, audited across thousands of papers.

What exists today, and what does not

Galaxy checks file formats with sniffers and auto-inserts format converters, which catches the syntactic 80 percent. nf-core pipelines validate their input sheets against JSON schemas, and some, like the RNA-seq pipeline, fail when detected strandedness contradicts the metadata. Those are real gains, but they are per-pipeline patches. No workflow system today carries the semantic state of the data (assembly, coordinate space, strandedness, normalization) as a first-class type that is checked on every edge between every pair of tools.

What we are building

BioNodulo is adding a semantic contract layer to its node graph. Every tool node will declare what it assumes about its inputs and guarantees about its outputs, in terms the graph can check: reference assembly, sort order, strandedness, normalization state. Every edge verifies that upstream guarantees satisfy downstream assumptions before anything runs. Where a legal conversion exists, such as sorting an unsorted BAM, the graph inserts the converter automatically; where none exists, it rejects the edge and explains which node guaranteed what and which node needed what. Every run will record what was checked in a standard provenance crate.

The same discipline applies to the agents. FlowBench, a 2026 benchmark of agentic bioinformatics, found that frontier models plan pipelines reasonably but attempt unsafe fault repairs in up to half of unrecoverable failure cases: fixes that produce a clean exit while leaving the data invalid. An agent composing workflows inside a contract-checked graph inherits the type system, so its proposals get the same per-edge verification a human's connections get.

This work is in active development, and we will report the measured results, including the failures, as it lands. If quiet failure modes have bitten your lab, we would like to hear about them.

References