Every reference dataset has a creation myth — the clean, compressed version that appears in the methods section of a paper. Ours is not that story. Between January 2019 and March 2025, the team at Stem Lab Zone collected, processed, and annotated 1,400 biological samples to build what is now the lab's primary in-house reference cohort for stem cell characterization work. The process involved four recruitment cycles, two protocol revisions, one IRB amendment that cost us eight weeks of momentum, and a demographic composition that still keeps the team up at night. This article is an attempt to document what actually happened, so that anyone who uses the dataset — or who is building something similar — understands exactly what they are working with.
Why build a reference dataset at all ¶
By late 2018, the team was running characterization assays against publicly available reference panels that had been assembled in the early 2010s. The problem was not that those panels were wrong — they were carefully made. The problem was that they reflected donor populations, cell-processing technologies, and passage conventions that did not match what the lab was seeing in its own pipelines. Batch effects between external reference and internal experimental data were eating roughly 12 to 18 percent of statistical variance in routine differential expression analyses, based on the team's internal benchmarking runs from Q3 2018. Building a matched in-house reference was, in the end, a practical engineering decision more than a scientific ambition.
Recruitment: four cycles and what each one taught us ¶
The team ran four distinct recruitment cycles. Cycle 1 (2019) drew from existing research-participant registries at two affiliated university hospitals and yielded 218 donors in nine months — faster than expected, but skewed heavily toward adults aged 35 to 55 who had previously participated in cardiovascular studies. Cycle 2 (2020 to 2021) was the hardest: pandemic restrictions halved in-person screening capacity, and the team pivoted to a postal consent model that ultimately produced 310 samples but introduced a new confound around cold-chain handling time that required a separate QC flag column in the master manifest. Cycle 3 (2022) brought in a community partnership with a network of regional clinics, adding 490 samples and meaningfully widening geographic spread within the country. Cycle 4 (2023 to 2025) focused deliberately on under-represented age brackets — specifically donors under 25 and over 70 — and closed with 382 samples, bringing the total to 1,400.
Inclusion and exclusion criteria: where the arguments happened ¶
The inclusion criteria sound simple on paper: age 18 to 80, no active oncological diagnosis at time of donation, no immunosuppressant therapy within 90 days, and signed informed consent under the lab's IRB protocol (amended in 2021 to extend the biobank retention clause from 10 to 20 years). The exclusion decisions were where the team spent most of its disagreement. Naomi Reyes, who led sample QC through Cycles 2 and 3, pushed hard to exclude donors with any documented autoimmune condition, citing contamination of baseline pluripotency markers. The counter-argument, led by senior analyst David Okonkwo, was that excluding autoimmune donors would make the reference useless for a large fraction of the lab's translational work. The compromise: autoimmune donors are included but carry a categorical flag, AUTOIMMUNE_PRESENT, and the team publishes stratified summary statistics for flagged and unflagged subsets separately.
Processing protocol: the two versions and why version 1 still haunts samples 1 through 311 ¶
Samples collected in Cycle 1 and the first half of Cycle 2 were processed under what the team now calls Protocol v1 — a fibroblast-derived iPSC reprogramming workflow using an episomal plasmid system with a 21-day culture timeline before characterization. In early 2021, the lab adopted Protocol v2, which shortened the timeline to 16 days by modifying the seeding density and switching to a chemically defined medium that had become commercially available. The change improved consistency of OCT4 and SOX2 expression scores across samples, but it introduced a systematic offset relative to v1 samples that the team has not yet fully resolved with a single correction factor. Currently, the dataset includes a PROTOCOL_VERSION column, and any cross-version analysis in the lab's publications uses a mixed-effects model with protocol version as a random effect. This is disclosed in every paper that draws on the cohort.
The demographic composition, honestly described ¶
The team does not want to bury this. As of the final Cycle 4 close, the cohort is 61 percent female, 39 percent male, with no non-binary or intersex demographic capture because the consent form used a binary sex-at-birth field until the 2023 revision — a gap that is now corrected for future recruitment but cannot be retroactively fixed for the 1,018 samples collected before that point. Ethnicity data relies on self-report using a national classification schema, and the largest single group represents 54 percent of the cohort. Donors from two of the country's northern regions are structurally under-represented because Cycle 3's clinic network did not extend past the central corridor. The team has been explicit about this in the two papers published from the dataset so far, and the data dictionary distributed with any external data-sharing request includes a dedicated limitations annex.
What the dataset is good for, and what it is not ¶
The 1,400-sample cohort is well-powered for analyses of baseline pluripotency marker variance across age groups, protocol-matched batch correction benchmarking, and longitudinal consistency testing within the lab's own workflows. It is not an appropriate reference for population-level epidemiological claims, for research questions requiring representation of rare genetic variants, or for any work that depends on precise ethnic stratification beyond the broad categories available. The team has fielded three external collaboration requests since 2023; two were accepted with data-sharing agreements that include the limitations annex as a required attachment, and one was declined because the proposed analysis required ethnic subgroup sizes the cohort cannot support.
Where the dataset goes next ¶
The team is in early planning for a Cycle 5 expansion targeting 300 additional samples, with recruitment focused on closing the northern-region gap and adding a longitudinal arm — meaning a subset of existing donors will be re-recruited for a second donation three years after their first, allowing the lab to study intra-individual variability over time. There is also an ongoing internal debate about whether to seek a formal data-sharing partnership with one of the national biobank networks, which would mean standardizing parts of the manifest to a common schema and potentially opening the dataset to external researchers under a managed access model. No decision has been made, and the team is documenting the tradeoffs in a working paper that Naomi Reyes is leading.
The 1,400-sample figure looks tidy from the outside. From the inside, it is a living record of protocol arguments, pandemic pivots, consent form revisions, and demographic gaps the team is still working to close. The lab's position is that a dataset's limitations are as scientifically meaningful as its findings, and that publishing both honestly is the only version of this work worth doing.