Data Governance in Drug Discovery: Designing Audit-Ready Storage for Regulated R&D
Key takeaways
- Metadata, lineage and retention determine whether discovery data can be reconstructed, reused and defended later.
- A tiered storage model lets exploratory research move quickly without losing the path to audit-ready evidence.
- Immutable raw zones, versioned analytic context, individual access and tested archive/export controls are the minimum for datasets likely to support assay transfer, AI model retraining, biomarker work or external diligence.
- Backup, archive and export are different controls and should be designed, tested and owned separately.
- Draft EU/PIC/S updates show regulators are moving toward more explicit requirements for data integrity, audit trails, system security, supplier oversight and AI/ML governance.
In regulated drug discovery, data governance is not an administrative layer added after the science. It is part of experimental design. The way teams store raw files, metadata, analysis code, model inputs and review-ready outputs determines whether results can be reconstructed, reused for AI, transferred to development, or defended during diligence and inspection. The practical question is not whether every exploratory dataset should be treated like a validated GMP record. It is how to design storage zones that let research move quickly while preserving the evidence that may later matter.
FAIR principles frame stewardship as a condition for discovery and reuse. NIH repository guidance treats metadata, provenance, retention policy, security and long-term sustainability as core selection criteria, and both FDA and MHRA emphasize that metadata is part of the record required to reconstruct regulated activity [1-6].
What regulated data governance must prove
Regulators ask whether electronic records are trustworthy, reliable, retrievable, attributable, reviewable and fit for their intended use. In the U.S., 21 CFR Part 11 defines the baseline for electronic records and signatures. Its closed-system controls include validation, accurate and complete copies, protection of records through the retention period, authorized access, secure time-stamped audit trails, authority checks and controls over system documentation [7]. FDA’s Part 11 scope-and-application guidance places those controls in a documented, risk-based context [8].
Across Europe and international inspectorates, the same pattern appears. Annex 11 requires lifecycle risk management, supplier oversight, data storage controls, backups, audit trails, security, business continuity and archiving that preserves accessibility, readability and integrity [9]. PIC/S and MHRA go further by positioning data governance as part of the pharmaceutical quality system, tied to data ownership, accountability, process design, monitoring and controls across the full data lifecycle [4,10].
For discovery organizations, that matters because a dataset may begin as exploratory biology, then become evidence for assay transfer, partner due diligence, model qualification, a biomarker package or support for later GLP, GMP or GCP activities. OECD guidance for GLP computerized systems reflects this reality by requiring identification of study-relevant electronic records, defined backup and archiving requirements, protection against alteration or loss, continued readability and access, and traceability to validated system configuration across the lifecycle [11].
Flexibility and control can coexist
Discovery teams need fast ingestion, heterogeneous formats and room to test new assays and models. Auditors, quality units and later-stage development teams need stable definitions, clear ownership, locked releases, reviewable audit trails and retrievable records.
Trying to introduce one storage mode that satisfies both goals usually combines the worst of both solutions: either rigid infrastructure that researchers work around or permissive storage that becomes expensive to remediate later. This tension is why regulators repeatedly emphasize justified, documented risk assessment and proportional controls [8-11].
PIC/S is especially useful here because it distinguishes data criticality from data risk and states that not all data or processing steps are equally important to product quality and patient safety. Control intensity should therefore follow the importance and vulnerability of the data, not the popularity of a tool or the preferences of a single team [10].
A sensible storage strategy is a tiered governance model: keep exploratory work in a research-friendly zone, promote reusable project outputs into a managed zone and move evidence-bearing datasets into an audit-ready zone with stronger controls [3,4,8-11].
Dimension | Research-friendly storage | Audit-ready storage |
|---|---|---|
Typical write behavior | Frequent ad hoc ingestion, transformation, and re-analysis | Controlled ingestion, defined release points, restricted modification |
Metadata posture | Useful but often uneven unless automated | Required, structured, preserved with the record, sufficient for reconstruction |
Lineage | Often partial unless workflow tooling is disciplined | Explicit chain of custody from raw data to approved output |
Versioning | Dataset snapshots may exist, but context can drift | Raw, curated, released data, code, configuration, and reference assets are versioned together |
Access model | Broad team access is common | Least privilege, individual accounts, admin segregation, and access review |
Auditability | Logging may be inconsistent across tools | Secure, time-stamped audit trails and reviewable change history |
Retention and retrieval | Active-project focused | Retention, archive, readability, restorability, and export are designed and tested |
Tooling bias | Flexible analytics stack, notebooks, lake, or object storage | Validated or qualification-ready components, controlled interfaces, and documented operating model |
Main failure mode | Results cannot be reconstructed later | Friction and over-control, if applied too early or too broadly |
Table 1. Infrastructure layers for trustworthy foundation models in pharma R&D.
Governance should start before data becomes evidence
Late remediation is usually expensive or impossible. If a raw data file arrives without operator identity, sample context, software version, parameter settings or a link to the protocol in force at the time, any subsequent reconstruction may be incomplete or non-defensible.
FDA treats data integrity as a lifecycle problem from creation through archive and disposition. MHRA and PIC/S apply governance across generation, processing, retention, retrieval and destruction. OECD adds that study-relevant electronic records, backup, archiving and configuration traceability should be defined across the computerized system lifecycle [3,4,10,11].
The same is true for identity and change control. Shared logins, generic admin access, untracked spreadsheet edits and post hoc data movement create ambiguity that is difficult to unwind later. FDA recommends restricting changes to authorized personnel and documenting access privileges. MHRA states that shared or generic logins should not be used for systems that generate, amend or store GxP data and that system administrator rights should be tightly restricted and separated from direct data access [3,4].
A practical maturity model is therefore more useful than a binary “regulated versus unregulated” mindset. The storage lifecycle should make clear how data moves from speed-oriented research work into stronger evidence-bearing controls.
Figure 1. Five-state storage lifecycle for regulated discovery data.
State | Purpose | Control emphasis | Promotion trigger |
|---|---|---|---|
Landing/raw | Capture original files, instrument outputs and source metadata. | Immutable capture, source IDs, checksum/hash where appropriate, operator and timestamp. | Data may influence a project decision or be reused outside the originating experiment. |
Curated/standardized | Normalize formats, resolve identifiers and add controlled metadata. | Schema/version control, curation rules, controlled vocabulary, quality checks. | Data becomes reusable across teams, models, assays or diligence packages. |
Analysis workspace | Run notebooks, pipelines, statistical analyses and model experiments. | Versioned code, containers, parameters, reference assets and feature-set releases. | An analysis output is selected for review, assay transfer, model qualification or external sharing. |
Controlled release | Freeze evidence-bearing outputs for review, transfer, partner diligence or regulatory support. | Release approvals, least-privilege access, audit trail review, controlled change management. | The record must remain reproducible, defensible and retrievable after active project work. |
Archive | Preserve records and context through the retention period. | Retention schedule, readability, restorability, exportability, vendor-exit path and periodic restore testing. | The project closes, moves to later-stage development or requires long-term evidentiary support. |
Table 2. Five-stage data storage lifecycle in regulated drug discovery, outlining the purpose, governance controls, and promotion criteria for progressing data from initial capture to long-term archival.
A practical blueprint for storage that is fit for purpose
- Classify data by evidentiary potential before you choose controls. At project kickoff, decide which data are exploratory, which inform project decisions and which may later support transfer, filing or external diligence. Controls should follow data criticality and data risk, rather than be dictated by organizational habit. That is the clearest route to staying flexible without becoming careless [8-10].
- Store by data state, not by department folder. In practice, most organizations benefit from at least five states: landing or raw, curated or standardized, analysis workspace, controlled release and archive. That model maps well to the lifecycle language in FDA, MHRA, PIC/S, Annex 11 and OECD guidance because each state can carry different rules for mutability, review, access and retention [3,4,9-11].
- Capture metadata automatically whenever possible. If metadata entry depends on memory, it will drift. Instrument ID, sample or batch ID, operator ID, acquisition time, protocol version, parameter settings, material status, software version and reference dataset version should be captured close to the source. FDA defines metadata as contextual information required to understand data and says it must be preserved in a secure and traceable manner throughout retention; MHRA likewise treats metadata as integral to the original record [3,4].
- Version the full analytic context, not only the file. In drug discovery, reproducibility depends on the combination of data, code, notebooks, model weights, containers, configuration, reference genomes, ontologies and processing parameters. Provenance literature shows that reproducibility improves when results remain linked to both computational and non-computational steps across the research lifecycle [5,6].
- Make identity and authorization individual and reviewable. Named accounts, least-privilege roles, admin segregation and documented access changes should be standard for any dataset that matters. FDA and MHRA both warn against weak access models; MHRA is explicit that shared or generic logins should not be used for critical GxP data paths, and FDA requires that only authorized individuals can change records [3,4,7,9].
- Treat backup, archive and export as different controls. Backup is for recovery. Archive is for retention and later verification. Export is for review, collaboration, migration or submission. They are not interchangeable. MHRA states that validated, periodically tested backup does not replace long-term archive; FDA requires backup data to be exact, complete and secure from alteration or loss; and Annex 11 and OECD both require restorability, readability and access across the retention period [3,4,9,11].
- Preserve dynamic records when their behavior matters. If a record can be reprocessed, queried or reinterpreted only in its original dynamic form, a PDF or image export may not be enough. FDA and MHRA both say that true copies must preserve meaning and, when needed, the dynamic functionality of the original record [3,4].
- Design for vendor exit before you go live. Cloud or SaaS does not transfer accountability. MHRA, PIC/S and Annex 11 all point to supplier oversight, risk-based review and contractual clarity. Before adopting a service, consider document ownership, geographic location, export format, access to metadata and audit trails, retention obligations and how data will remain readable if the system is replaced [4,9,10].
Related Ardigen reading: Why Data Quality Matters in AI-powered Drug Discovery and Lab-in-the-Loop: reclaiming the 50% of scientific time lost to data.
Regulatory watch: draft Annex 11 and Annex 22
Regulatory expectations are also becoming more explicit. As of June 2026, the European Commission lists the current EU GMP Annex 11 on Computerised Systems as the January 2011 revision [9]. In July 2025, the European Commission and PIC/S opened stakeholder consultation on draft revisions to Chapter 4 and Annex 11 and a new Annex 22 on Artificial Intelligence [14,15].
The practical signal for discovery and development organizations is clear: data integrity, audit trails, supplier oversight, system security, lifecycle validation and AI/ML governance are moving toward more prescriptive expectations. The exact scope of final obligations will depend on the adopted text, so teams should track the consultation outcomes rather than treating draft clauses as binding.
Why this matters in drug discovery
Drug discovery already carries enough uncertainty that cannot be reduced. A 2015 estimate put the annual cost of preclinical irreproducibility at roughly US$28 billion in the United States alone – a figure still cited as the standard reference a decade later, since no updated calculation has replaced it. A 2024 survey of biomedical researchers found that 72% still perceive an ongoing reproducibility crisis, with 27% calling it significant [12,13,16].
Good storage design will not rescue a weak research hypothesis, but it allows a team to move fast without losing control of the evidence that may later matter most. It protects the integrity of assay transfer, model retraining, biomarker work, partner diligence and eventual regulatory support [3-6,9-11].
In drug discovery, storage is part of experimental design. The right goal is to ensure that any dataset with scientific or regulatory consequences can be promoted to stronger control without a messy archaeological dig through shared drives, local notebooks and undocumented transformations [3,4,8-11].
Want to know which datasets can safely remain exploratory and which need stronger controls? Book a 30-minute data governance readiness review with Ardigen’s Data Management team. We will map your current storage landscape across raw, curated, analysis, release and archive zones, then identify the highest-risk gaps in metadata, lineage, access and retention. Where relevant, the review can connect those gaps to Ardigen capabilities in data engineering, automated pipelines, cloud-native infrastructure, access portals, compliance/governance and MLOps/ML engineering.
Frequently Asked Questions
What is data governance in drug discovery?
It is the set of policies, technical controls and operating practices that preserve the meaning, ownership, lineage, access rights and retention status of discovery data across its lifecycle. In practice, it determines whether data can be reconstructed, reused for AI, transferred to development or defended during diligence and inspection.
What makes drug discovery storage audit-ready?
Audit-ready storage preserves raw data, metadata, access history, versioned analytic context, review status, retention rules and exportability in a way that can be reconstructed later. It does not mean every exploratory file is over-controlled; it means evidence-bearing data has a defined path into stronger controls.
How do 21 CFR Part 11 and Annex 11 affect discovery data?
They become relevant when electronic records support regulated decisions, submissions, GxP activities or later evidence packages. Discovery teams should use risk-based controls so that high-value datasets can be promoted without retroactive reconstruction.
Why does metadata matter for AI-ready drug discovery?
AI models learn from both data and context. Missing sample, assay, protocol, software, parameter and reference-dataset metadata can make model training, retraining and interpretation unreliable or impossible to defend.
What is lineage or provenance in discovery data?
Lineage shows how data moved and changed from raw source to curated dataset, analysis output, reviewed result and archive. Provenance adds the scientific and computational context needed to understand who did what, when, with which protocol, code, parameters, reference assets and system configuration.
What is the difference between backup and archive?
Backup supports operational recovery after failure. Archive supports long-term retention, readability, retrieval and verification. Export supports review, collaboration, migration or submission. They should not be treated as interchangeable.
When should exploratory data move into stronger controls?
When it begins to influence project decisions, assay transfer, partner diligence, biomarker work, model qualification, GLP/GMP/GCP planning or regulatory support. The trigger should be evidentiary potential, not department ownership.
How should cloud or SaaS vendor exit be handled?
Before go-live, define ownership, location, export formats, metadata and audit-trail access, retention obligations, migration path, and how records remain readable if the vendor or system changes.
Author: Martyna Piotrowska
Technical editing: Ardigen expert: Dawid Rymarczyk, PhD
References
- Wilkinson MD, Dumontier M, Aalbersberg IJJ, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:160018. https://doi.org/10.1038/sdata.2016.18
- National Institutes of Health. Selecting a Data Repository [Internet]. Bethesda (MD): NIH; 2025. Available from: https://grants.nih.gov/policy-and-compliance/policy-topics/sharing-policies/dms/selecting-a-data-repository
- U.S. Food and Drug Administration. Data Integrity and Compliance With Drug CGMP: Questions and Answers. Guidance for Industry [Internet]. Silver Spring (MD): FDA; 2018. Available from: https://www.fda.gov/regulatory-information/search-fda-guidance-documents/data-integrity-and-compliance-drug-cgmp-questions-and-answers
- Medicines and Healthcare products Regulatory Agency. GxP Data Integrity Guidance and Definitions [Internet]. London: MHRA; 2018. Available from: https://assets.publishing.service.gov.uk/media/5aa2b9ede5274a3e39 1e37f3/MHRA_GxP_data_integrity_guide_March_edited_Final.pdf
- Gierend K, Krueger F, Genehr S, et al. Provenance Information for Biomedical Data and Workflows: Scoping Review. J Med Internet Res. 2024;26:e51297. https://doi.org/10.2196/51297
- Samuel S, Koenig-Ries B. End-to-End provenance representation for the understandability and reproducibility of scientific experiments using a semantic approach. J Biomed Semantics. 2022;13:1. https://doi.org/10.1186/s13326-021-00253-1
- Electronic Code of Federal Regulations. 21 CFR Part 11–Electronic Records; Electronic Signatures [Internet]. Current version. Available from: https://www.ecfr.gov/current/title-21/chapter-I/subchapter-A/part-11
- U.S. Food and Drug Administration. Part 11, Electronic Records; Electronic Signatures–Scope and Application. Guidance for Industry [Internet]. Silver Spring (MD): FDA; 2003. Available from: https://www.fda.gov/media/75414/download
- European Commission. EudraLex Volume 4, Annex 11: Computerised Systems [Internet]. Brussels: European Commission; 2011. Available from: https://health.ec.europa.eu/system/files/2016-11/annex11_01-2011_en_0.pdf
- Pharmaceutical Inspection Co-operation Scheme. Guidance on Data Integrity. PI 041-1 [Internet]. Geneva: PIC/S; 2021. Available from: https://picscheme.org/docview/4234
- Organisation for Economic Co-operation and Development. The Application of the Principles of GLP to Computerised Systems [Internet]. Paris: OECD; 1995. Available from: https://www.oecd.org/content/dam/oecd/en/publications/reports/1995/10/the-application-of-the-principles-of-glp-to-computerised-systems_g1gh32fb/9789264078710-en.pdf
- Freedman LP, Cockburn IM, Simcoe TS. The Economics of Reproducibility in Preclinical Research. PLoS Biol. 2015;13(6):e1002165. https://doi.org/10.1371/journal.pbio.1002165
- Prinz F, Schlange T, Asadullah K. Believe it or not: how much can we rely on published data on potential drug targets? Nat Rev Drug Discov. 2011;10:712. https://doi.org/10.1038/nrd3439-c1
- European Commission. Stakeholders’ Consultation on EudraLex Volume 4 – Good Manufacturing Practice Guidelines: Chapter 4, Annex 11 and New Annex 22 [Internet]. Brussels: European Commission; 2025. Available from: https://health.ec.europa.eu/consultations/stakeholders-consultation-eudralex-volume-4-good-manufacturing-practice-guidelines-chapter-4-annex_en
- Pharmaceutical Inspection Co-operation Scheme. Joint stakeholders consultation on the revision of Chapter 4 on Documentation, Annex 11 on Computerised Systems and on the new Annex 22 on Artificial Intelligence of the PIC/S and EU GMP Guides [Internet]. Geneva: PIC/S; 2025. Available from: https://picscheme.org/en/news/joint-stakeholders-consultation-on-the-revision-of-chapter-4
- Cobey KD, Ebrahimzadeh S, Page MJ, Thibault RT, Nguyen P-Y, et al. Biomedical researchers’ perspectives on the reproducibility of research. PLOS Biology. 2024;22(11):e3002870. https://doi.org/10.1371/journal.pbio.3002870