Data Governance in Drug Discovery: Designing Audit-Ready Storage for Regulated R&D

Scientist working with laboratory and data analysis tools to support data governance in regulated drug discovery.

Data Governance in Drug Discovery: Designing Audit-Ready Storage for Regulated R&D

Key takeaways

  • Metadata, lineage and retention determine whether discovery data can be reconstructed, reused and defended later.
  • A tiered storage model lets exploratory research move quickly without losing the path to audit-ready evidence.
  • Immutable raw zones, versioned analytic context, individual access and tested archive/export controls are the minimum for datasets likely to support assay transfer, AI model retraining, biomarker work or external diligence.
  • Backup, archive and export are different controls and should be designed, tested and owned separately.
  • Draft EU/PIC/S updates show regulators are moving toward more explicit requirements for data integrity, audit trails, system security, supplier oversight and AI/ML governance.

In regulated drug discovery, data governance is not an administrative layer added after the science. It is part of experimental design. The way teams store raw files, metadata, analysis code, model inputs and review-ready outputs determines whether results can be reconstructed, reused for AI, transferred to development, or defended during diligence and inspection. The practical question is not whether every exploratory dataset should be treated like a validated GMP record. It is how to design storage zones that let research move quickly while preserving the evidence that may later matter.

FAIR principles frame stewardship as a condition for discovery and reuse. NIH repository guidance treats metadata, provenance, retention policy, security and long-term sustainability as core selection criteria, and both FDA and MHRA emphasize that metadata is part of the record required to reconstruct regulated activity [1-6].

What regulated data governance must prove

Regulators ask whether electronic records are trustworthy, reliable, retrievable, attributable, reviewable and fit for their intended use. In the U.S., 21 CFR Part 11 defines the baseline for electronic records and signatures. Its closed-system controls include validation, accurate and complete copies, protection of records through the retention period, authorized access, secure time-stamped audit trails, authority checks and controls over system documentation [7]. FDA’s Part 11 scope-and-application guidance places those controls in a documented, risk-based context [8].

Across Europe and international inspectorates, the same pattern appears. Annex 11 requires lifecycle risk management, supplier oversight, data storage controls, backups, audit trails, security, business continuity and archiving that preserves accessibility, readability and integrity [9]. PIC/S and MHRA go further by positioning data governance as part of the pharmaceutical quality system, tied to data ownership, accountability, process design, monitoring and controls across the full data lifecycle [4,10].

For discovery organizations, that matters because a dataset may begin as exploratory biology, then become evidence for assay transfer, partner due diligence, model qualification, a biomarker package or support for later GLP, GMP or GCP activities. OECD guidance for GLP computerized systems reflects this reality by requiring identification of study-relevant electronic records, defined backup and archiving requirements, protection against alteration or loss, continued readability and access, and traceability to validated system configuration across the lifecycle [11].

Flexibility and control can coexist

Discovery teams need fast ingestion, heterogeneous formats and room to test new assays and models. Auditors, quality units and later-stage development teams need stable definitions, clear ownership, locked releases, reviewable audit trails and retrievable records.

Trying to introduce one storage mode that satisfies both goals usually combines the worst of both solutions: either rigid infrastructure that researchers work around or permissive storage that becomes expensive to remediate later. This tension is why regulators repeatedly emphasize justified, documented risk assessment and proportional controls [8-11].

PIC/S is especially useful here because it distinguishes data criticality from data risk and states that not all data or processing steps are equally important to product quality and patient safety. Control intensity should therefore follow the importance and vulnerability of the data, not the popularity of a tool or the preferences of a single team [10].

A sensible storage strategy is a tiered governance model: keep exploratory work in a research-friendly zone, promote reusable project outputs into a managed zone and move evidence-bearing datasets into an audit-ready zone with stronger controls [3,4,8-11].

Dimension

Research-friendly storage

Audit-ready storage

Typical write behavior

Frequent ad hoc ingestion, transformation, and re-analysis

Controlled ingestion, defined release points, restricted modification

Metadata posture

Useful but often uneven unless automated

Required, structured, preserved with the record, sufficient for reconstruction

Lineage

Often partial unless workflow tooling is disciplined

Explicit chain of custody from raw data to approved output

Versioning

Dataset snapshots may exist, but context can drift

Raw, curated, released data, code, configuration, and reference assets are versioned together

Access model

Broad team access is common

Least privilege, individual accounts, admin segregation, and access review

Auditability

Logging may be inconsistent across tools

Secure, time-stamped audit trails and reviewable change history

Retention and retrieval

Active-project focused

Retention, archive, readability, restorability, and export are designed and tested

Tooling bias

Flexible analytics stack, notebooks, lake, or object storage

Validated or qualification-ready components, controlled interfaces, and documented operating model

Main failure mode

Results cannot be reconstructed later

Friction and over-control, if applied too early or too broadly

Table 1. Infrastructure layers for trustworthy foundation models in pharma R&D.

Governance should start before data becomes evidence

Late remediation is usually expensive or impossible. If a raw data file arrives without operator identity, sample context, software version, parameter settings or a link to the protocol in force at the time, any subsequent reconstruction may be incomplete or non-defensible.

FDA treats data integrity as a lifecycle problem from creation through archive and disposition. MHRA and PIC/S apply governance across generation, processing, retention, retrieval and destruction. OECD adds that study-relevant electronic records, backup, archiving and configuration traceability should be defined across the computerized system lifecycle [3,4,10,11].

The same is true for identity and change control. Shared logins, generic admin access, untracked spreadsheet edits and post hoc data movement create ambiguity that is difficult to unwind later. FDA recommends restricting changes to authorized personnel and documenting access privileges. MHRA states that shared or generic logins should not be used for systems that generate, amend or store GxP data and that system administrator rights should be tightly restricted and separated from direct data access [3,4].

A practical maturity model is therefore more useful than a binary “regulated versus unregulated” mindset. The storage lifecycle should make clear how data moves from speed-oriented research work into stronger evidence-bearing controls.

Five-step data workflow from Landing/Raw to Archive: Landing/Raw → Curated/Standardized → Analysis Workspace → Controlled Release → Archive

Figure 1. Five-state storage lifecycle for regulated discovery data.

State

Purpose

Control emphasis

Promotion trigger

Landing/raw

Capture original files, instrument outputs and source metadata.

Immutable capture, source IDs, checksum/hash where appropriate, operator and timestamp.

Data may influence a project decision or be reused outside the originating experiment.

Curated/standardized

Normalize formats, resolve identifiers and add controlled metadata.

Schema/version control, curation rules, controlled vocabulary, quality checks.

Data becomes reusable across teams, models, assays or diligence packages.

Analysis workspace

Run notebooks, pipelines, statistical analyses and model experiments.

Versioned code, containers, parameters, reference assets and feature-set releases.

An analysis output is selected for review, assay transfer, model qualification or external sharing.

Controlled release

Freeze evidence-bearing outputs for review, transfer, partner diligence or regulatory support.

Release approvals, least-privilege access, audit trail review, controlled change management.

The record must remain reproducible, defensible and retrievable after active project work.

Archive

Preserve records and context through the retention period.

Retention schedule, readability, restorability, exportability, vendor-exit path and periodic restore testing.

The project closes, moves to later-stage development or requires long-term evidentiary support.

Table 2. Five-stage data storage lifecycle in regulated drug discovery, outlining the purpose, governance controls, and promotion criteria for progressing data from initial capture to long-term archival.

A practical blueprint for storage that is fit for purpose

  1. Classify data by evidentiary potential before you choose controls. At project kickoff, decide which data are exploratory, which inform project decisions and which may later support transfer, filing or external diligence. Controls should follow data criticality and data risk, rather than be dictated by organizational habit. That is the clearest route to staying flexible without becoming careless [8-10].
  2. Store by data state, not by department folder. In practice, most organizations benefit from at least five states: landing or raw, curated or standardized, analysis workspace, controlled release and archive. That model maps well to the lifecycle language in FDA, MHRA, PIC/S, Annex 11 and OECD guidance because each state can carry different rules for mutability, review, access and retention [3,4,9-11].
  3. Capture metadata automatically whenever possible. If metadata entry depends on memory, it will drift. Instrument ID, sample or batch ID, operator ID, acquisition time, protocol version, parameter settings, material status, software version and reference dataset version should be captured close to the source. FDA defines metadata as contextual information required to understand data and says it must be preserved in a secure and traceable manner throughout retention; MHRA likewise treats metadata as integral to the original record [3,4].
  4. Version the full analytic context, not only the file. In drug discovery, reproducibility depends on the combination of data, code, notebooks, model weights, containers, configuration, reference genomes, ontologies and processing parameters. Provenance literature shows that reproducibility improves when results remain linked to both computational and non-computational steps across the research lifecycle [5,6].
  5. Make identity and authorization individual and reviewable. Named accounts, least-privilege roles, admin segregation and documented access changes should be standard for any dataset that matters. FDA and MHRA both warn against weak access models; MHRA is explicit that shared or generic logins should not be used for critical GxP data paths, and FDA requires that only authorized individuals can change records [3,4,7,9].
  6. Treat backup, archive and export as different controls. Backup is for recovery. Archive is for retention and later verification. Export is for review, collaboration, migration or submission. They are not interchangeable. MHRA states that validated, periodically tested backup does not replace long-term archive; FDA requires backup data to be exact, complete and secure from alteration or loss; and Annex 11 and OECD both require restorability, readability and access across the retention period [3,4,9,11].
  7. Preserve dynamic records when their behavior matters. If a record can be reprocessed, queried or reinterpreted only in its original dynamic form, a PDF or image export may not be enough. FDA and MHRA both say that true copies must preserve meaning and, when needed, the dynamic functionality of the original record [3,4].
  8. Design for vendor exit before you go live. Cloud or SaaS does not transfer accountability. MHRA, PIC/S and Annex 11 all point to supplier oversight, risk-based review and contractual clarity. Before adopting a service, consider document ownership, geographic location, export format, access to metadata and audit trails, retention obligations and how data will remain readable if the system is replaced [4,9,10].

Related Ardigen reading: Why Data Quality Matters in AI-powered Drug Discovery and Lab-in-the-Loop: reclaiming the 50% of scientific time lost to data.

Regulatory watch: draft Annex 11 and Annex 22

Regulatory expectations are also becoming more explicit. As of June 2026, the European Commission lists the current EU GMP Annex 11 on Computerised Systems as the January 2011 revision [9]. In July 2025, the European Commission and PIC/S opened stakeholder consultation on draft revisions to Chapter 4 and Annex 11 and a new Annex 22 on Artificial Intelligence [14,15].

The practical signal for discovery and development organizations is clear: data integrity, audit trails, supplier oversight, system security, lifecycle validation and AI/ML governance are moving toward more prescriptive expectations. The exact scope of final obligations will depend on the adopted text, so teams should track the consultation outcomes rather than treating draft clauses as binding.

Why this matters in drug discovery

Drug discovery already carries enough uncertainty that cannot be reduced. A 2015 estimate put the annual cost of preclinical irreproducibility at roughly US$28 billion in the United States alone – a figure still cited as the standard reference a decade later, since no updated calculation has replaced it. A 2024 survey of biomedical researchers found that 72% still perceive an ongoing reproducibility crisis, with 27% calling it significant [12,13,16]. 

Good storage design will not rescue a weak research hypothesis, but it allows a team to move fast without losing control of the evidence that may later matter most. It protects the integrity of assay transfer, model retraining, biomarker work, partner diligence and eventual regulatory support [3-6,9-11].

In drug discovery, storage is part of experimental design. The right goal is to ensure that any dataset with scientific or regulatory consequences can be promoted to stronger control without a messy archaeological dig through shared drives, local notebooks and undocumented transformations [3,4,8-11].

Want to know which datasets can safely remain exploratory and which need stronger controls? Book a 30-minute data governance readiness review with Ardigen’s Data Management team. We will map your current storage landscape across raw, curated, analysis, release and archive zones, then identify the highest-risk gaps in metadata, lineage, access and retention. Where relevant, the review can connect those gaps to Ardigen capabilities in data engineering, automated pipelines, cloud-native infrastructure, access portals, compliance/governance and MLOps/ML engineering.

Frequently Asked Questions

It is the set of policies, technical controls and operating practices that preserve the meaning, ownership, lineage, access rights and retention status of discovery data across its lifecycle. In practice, it determines whether data can be reconstructed, reused for AI, transferred to development or defended during diligence and inspection.

Audit-ready storage preserves raw data, metadata, access history, versioned analytic context, review status, retention rules and exportability in a way that can be reconstructed later. It does not mean every exploratory file is over-controlled; it means evidence-bearing data has a defined path into stronger controls.

They become relevant when electronic records support regulated decisions, submissions, GxP activities or later evidence packages. Discovery teams should use risk-based controls so that high-value datasets can be promoted without retroactive reconstruction.

AI models learn from both data and context. Missing sample, assay, protocol, software, parameter and reference-dataset metadata can make model training, retraining and interpretation unreliable or impossible to defend.

Lineage shows how data moved and changed from raw source to curated dataset, analysis output, reviewed result and archive. Provenance adds the scientific and computational context needed to understand who did what, when, with which protocol, code, parameters, reference assets and system configuration.

Backup supports operational recovery after failure. Archive supports long-term retention, readability, retrieval and verification. Export supports review, collaboration, migration or submission. They should not be treated as interchangeable.

When it begins to influence project decisions, assay transfer, partner diligence, biomarker work, model qualification, GLP/GMP/GCP planning or regulatory support. The trigger should be evidentiary potential, not department ownership.

Before go-live, define ownership, location, export formats, metadata and audit-trail access, retention obligations, migration path, and how records remain readable if the vendor or system changes.

Author: Martyna Piotrowska

Technical editing:  Ardigen expert: Dawid Rymarczyk, PhD

References

  1. Wilkinson MD, Dumontier M, Aalbersberg IJJ, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:160018. https://doi.org/10.1038/sdata.2016.18
  2. National Institutes of Health. Selecting a Data Repository [Internet]. Bethesda (MD): NIH; 2025. Available from: https://grants.nih.gov/policy-and-compliance/policy-topics/sharing-policies/dms/selecting-a-data-repository
  3. U.S. Food and Drug Administration. Data Integrity and Compliance With Drug CGMP: Questions and Answers. Guidance for Industry [Internet]. Silver Spring (MD): FDA; 2018. Available from: https://www.fda.gov/regulatory-information/search-fda-guidance-documents/data-integrity-and-compliance-drug-cgmp-questions-and-answers
  4. Medicines and Healthcare products Regulatory Agency. GxP Data Integrity Guidance and Definitions [Internet]. London: MHRA; 2018. Available from: https://assets.publishing.service.gov.uk/media/5aa2b9ede5274a3e39   1e37f3/MHRA_GxP_data_integrity_guide_March_edited_Final.pdf  
  5. Gierend K, Krueger F, Genehr S, et al. Provenance Information for Biomedical Data and Workflows: Scoping Review. J Med Internet Res. 2024;26:e51297. https://doi.org/10.2196/51297
  6. Samuel S, Koenig-Ries B. End-to-End provenance representation for the understandability and reproducibility of scientific experiments using a semantic approach. J Biomed Semantics. 2022;13:1. https://doi.org/10.1186/s13326-021-00253-1
  7. Electronic Code of Federal Regulations. 21 CFR Part 11–Electronic Records; Electronic Signatures [Internet]. Current version. Available from: https://www.ecfr.gov/current/title-21/chapter-I/subchapter-A/part-11
  8. U.S. Food and Drug Administration. Part 11, Electronic Records; Electronic Signatures–Scope and Application. Guidance for Industry [Internet]. Silver Spring (MD): FDA; 2003. Available from: https://www.fda.gov/media/75414/download
  9. European Commission. EudraLex Volume 4, Annex 11: Computerised Systems [Internet]. Brussels: European Commission; 2011. Available from: https://health.ec.europa.eu/system/files/2016-11/annex11_01-2011_en_0.pdf
  10. Pharmaceutical Inspection Co-operation Scheme. Guidance on Data Integrity. PI 041-1 [Internet]. Geneva: PIC/S; 2021. Available from: https://picscheme.org/docview/4234
  11. Organisation for Economic Co-operation and Development. The Application of the Principles of GLP to Computerised Systems [Internet]. Paris: OECD; 1995. Available from: https://www.oecd.org/content/dam/oecd/en/publications/reports/1995/10/the-application-of-the-principles-of-glp-to-computerised-systems_g1gh32fb/9789264078710-en.pdf
  12. Freedman LP, Cockburn IM, Simcoe TS. The Economics of Reproducibility in Preclinical Research. PLoS Biol. 2015;13(6):e1002165. https://doi.org/10.1371/journal.pbio.1002165
  13. Prinz F, Schlange T, Asadullah K. Believe it or not: how much can we rely on published data on potential drug targets? Nat Rev Drug Discov. 2011;10:712. https://doi.org/10.1038/nrd3439-c1
  14. European Commission. Stakeholders’ Consultation on EudraLex Volume 4 – Good Manufacturing Practice Guidelines: Chapter 4, Annex 11 and New Annex 22 [Internet]. Brussels: European Commission; 2025. Available from: https://health.ec.europa.eu/consultations/stakeholders-consultation-eudralex-volume-4-good-manufacturing-practice-guidelines-chapter-4-annex_en
  15. Pharmaceutical Inspection Co-operation Scheme. Joint stakeholders consultation on the revision of Chapter 4 on Documentation, Annex 11 on Computerised Systems and on the new Annex 22 on Artificial Intelligence of the PIC/S and EU GMP Guides [Internet]. Geneva: PIC/S; 2025. Available from: https://picscheme.org/en/news/joint-stakeholders-consultation-on-the-revision-of-chapter-4
  16. Cobey KD, Ebrahimzadeh S, Page MJ, Thibault RT, Nguyen P-Y, et al. Biomedical researchers’ perspectives on the reproducibility of research. PLOS Biology. 2024;22(11):e3002870. https://doi.org/10.1371/journal.pbio.3002870

You might be also interested in:

Ardigen and VERAXA Biotech announce AI-enabled drug discovery collaboration
Ardigen and VERAXA Biotech announce AI-enabled drug discovery collaboration
Scientist reviewing biological data for multimodal foundation models in drug discovery
Foundation Models in Drug Discovery: What Pharma Needs to Scale Multimodal AI
Poster: Bridging the Phenotype-Proteome Gap: A Multi-Modal AI Framework for analysis of Cell Painting images
Biomedical knowledge graph connecting drug discovery data, disease biology, targets and indications to support target prioritization and indication expansion
Knowledge Graphs in Drug Discovery: Bringing Context to Internal Data Assets

Contact

Ready to transform drug discovery?

Discover how one of the top AI CROs in the world, can be your trusted partner in revolutionizing drug discovery through AI.

Contact us today to learn more about our tailored solutions for empowering your drug development journey.

Send us a message and we will contact you back within 48 hours.

Newsletter

Become an insider

Be the first to know about Ardigen’s latest news and get access to our publications, webinars and more!