Data vs hypothesis driven research: the false dichotomy
This post is adapted from a talk I gave at the Plantae Presents webinar More Money, More Data: Data-Driven versus Hypothesis-Driven Research (recording here). The webinar was moderated by Plantae Fellows Krishna Alamuru, Dennis Baffour-Awuah, Aditi Bhat, and Kestrel Maio. They moderated the discussion and invited me and Ivan Baxter of the Donald Danforth Plant Science Center to explore the evolving landscape of plant biology in the era of big data.
My contribution tackled the presumed tension between traditional hypothesis-driven science and big, data-driven approaches. The slides from my presentation are available on Zenodo.
What follows is a blog version of that talk.
Every few years, science seems to find itself in a new “versus” debate. In plant biology right now, it’s data-driven versus hypothesis-driven research. The rise of high-throughput technologies and massive datasets has put this tension front and center.
I recently spoke at the Plantae Presents webinar More Money, More Data: Data-Driven versus Hypothesis-Driven Research (link), and my main message was simple:
👉 It’s a false dichotomy.
These aren’t opposing modes of science. In fact, data-driven work is best seen as a type of exploratory research. And once your explorations yield a tangible finding? You have no excuse but to switch into the classic hypothesis–experiment cycles. That’s where the real science happens.
The real issue here is… but wait, we’ll come back to that later.
What business are we in, anyway?
I’d like to zoom out before tackling this debate. As scientists, we are in the business of knowledge. That’s our currency.
- If you’re a chef, you’re in the business of food.
- If you’re a banker, you’re in the business of money.
- If you’re a scientist, you’re in the business of producing and communicating knowledge.
It’s worth reminding ourselves: the goal isn’t simply to publish papers. The goal is to produce knowledge that yields predictable outcomes — science that stands the test of time.
That’s why my obsession isn’t flashy stats, wishy-washy studies, or ‘one-time-wonder’ papers. It’s reproducibility. Nothing beats replication, especially orthogonal replication where a finding holds up across independent methods or systems.
I don’t build my program around one exciting, fragile experiment. I build it around a series of experiments that converge on a robust, reproducible conclusion.
How to science (yes, there’s a mnemonic)
You can’t do science without a BBQ: a Big Biological Question.
The art is in balancing that big question against the systems and methods you actually have. Some of our most ambitious questions are still out of reach — not because they aren’t worth asking, but because the technology or biomaterials aren’t there yet. Good science lives in that tension.
In my lab, we use a couple of simple frameworks to keep us grounded:
- GOHREP: Goal — Hypothesis — Rationale — Experimental Plan. Every project needs this structure.
- PLESI cycle: Planning — Execution — Scoring — Interpretation. Each step is independent, and all are equally important.
And research isn’t one heroic experiment. It’s a Sisyphean cycle: pushing that boulder up the hill over and over until, hopefully, it becomes a mountain of solid evidence.
Failed experiments, negative results, and a better mindset
Students often ask: What’s a failed experiment?
Here’s my answer:
- A failed experiment isn’t one that disproves your hypothesis. That’s valuable!
- A failed experiment is one from which you can draw no conclusion.
Negative data isn’t the enemy. In fact, using unexpected results as a springboard into new directions can make your lab more productive. There’s even research on this — labs that embrace “serendipitous pivots” outperform those that grind away trying to force a pre-set narrative.
The real issue behind “data-driven research”
Here’s where I poke the bear. The problem isn’t data-driven science per se. It’s when it becomes a shortcut to “easy” papers.
I’ve seen it:
- A lab lands a big grant.
- Generates tons of -omics data.
- Publishes a descriptive paper with little meaningful insight.
- High-impact journal obliges (you paid the APCs, after all).
I’m not impressed by datasets dressed up as discoveries. Data is just the start. Its value is unlocked when you extract impactful findings and follow up with proper hypothesis testing.
That said, raw datasets are only powerful if you share them openly. Preprints, Zenodo communities, mini-papers (aka micro-papers) — these formats credit your work, trigger collaborations, and don’t prevent you from publishing afterwards in journals.
Our #OpenWheatBlast project is a great example: a dozen of small datasets first published in a Zenodo community and then turned into a global community paper, with everyone who authored the Zenodo mini-papers credited.
To sum up
- You’re in the business of generating and disseminating knowledge. Don’t lose sight of that.
- Data vs. hypothesis-driven research? It’s a false dichotomy. Exploratory work quickly needs to become hypothesis testing.
- Reproducibility beats flashy results every time.
- Share your data early — preprints, mini/micro-papers, Zenodo. Open science starts there.
- And please, don’t deliver “Deliveroo science”: half-baked experiments rushed to please a PI or a journal. Follow the data. Do the work. Build the knowledge.
Acknowledgements
I thank the organizers for inviting me to the event and for prompting me to reflect on the topic. This article was written with assistance from ChatGPT.
This article is available on a CC-BY license via Zenodo.
Cite as: Kamoun, S. (2025). Data vs hypothesis driven research: the false dichotomy. Zenodo. https://doi.org/10.5281/zenodo.16733366
