Knowledge Graphs in Drug Discovery for Connected Biomedical Data
Key takeaways
- Proprietary internal data creates strategic advantage only when it is connected to broader biomedical context (e.g. disease, target, pathway, tissue and evidence context).
- A biomedical knowledge graph for drug discovery turns fragmented biological information into a structured, queryable context layer, crucially in both AI and ML – ready formats.
- Efficient Docking method of proprietary data into a graph requires identifier harmonization, ontology alignment, provenance, metadata and confidence scoring and is critical to realize the full potential of internal data assets.
- Knowledge Graph Semantic representation enables graph methods to support target prioritization, disease clustering and indication expansion.
- Ardigen can help teams work from ARDDKG or from an existing client graph by adding docking strategy, graph ML, evidence reporting and governed AI-agent interfaces.
The context problem: why unique data is not enough
Pharma and biotech companies sit on valuable proprietary experimental, omics and clinical data. Yet those assets often struggle to reach their full strategic potential because they are analyzed inside a narrow platform or project context. A compound, biologic, assay readout or clinical observation may be promising, but the next portfolio question is broader: where else can this asset go, and what evidence supports that move?
A biomedical knowledge graph for drug discovery helps answer that question by connecting internal data to a wider map of human biology. Instead of treating an asset, target or readout as a standalone object, a knowledge graph places it in relation to diseases, genes, proteins, pathways, tissues, phenotypes, drugs and evidence sources.
This does not replace scientific judgment. It gives R&D, data science and portfolio teams a structured way to ask better questions: what is already known, what is adjacent, what is missing and where does the proprietary signal create a differentiated opportunity?
What kind of context do internal data assets need?
Most organizations already know that data needs biomedical context. The hard part is choosing the right context and making it operational. Commercial datasets can be valuable, but access, licensing and reuse rules may limit how they can be combined with internal assets. Public resources can be extremely useful, but they vary in curation, provenance, coverage and harmonization.
Even when data access is solved, scale creates a second problem. When teams face millions or billions of biological data points, they need a way to separate the signal that is relevant to a pipeline decision from context that is merely available.
The practical goal is not to collect every possible data point. It is to create a context layer that is broad enough to capture disease biology, yet structured enough to support focused questions from scientists, data teams and decision makers.
Knowledge graphs: a practical context layer for drug discovery
Knowledge graphs are a pragmatic middle layer between isolated internal datasets and overwhelming unstructured biomedical information. A knowledge graph connects entities such as diseases, genes, proteins, tissues, pathways, drugs and phenotypes through typed relationships. Depending on the graph model, those relationships may be represented as triples or as typed edges in a property graph.
For example, a graph may encode relationships such as:
- (Drug, has_indication, Disease)
- (Disease, associated_with, Gene)
- (Gene, encodes, Protein)
- (Gene, expressed_in, Tissue)
- (Protein, participates_in, Pathway)
The important point is not the syntax. The important point is that biological facts, assumptions and evidence become explicit and navigable. When the graph is built with a defined schema, ontology mapping, provenance and quality control, it becomes a reusable context layer for discovery questions.
Public biomedical graphs such as PrimeKG, OREGANO and OptimusKG show how multimodal biomedical knowledge can be represented as connected entities and relationships. Enterprise use cases add another layer: the ability to connect a company’s proprietary evidence to that broader context under governance.
Figure 1. Knowledge graph harmonization in PrimeKG.
Reproduced from Chandak et al., Nature Scientific Data (2023) [9].
What does it mean to dock proprietary data into a graph?
The real value begins when internal assets are docked into this broader biomedical graph. Docking is the process of mapping proprietary data – such as compounds, biologics, assays, targets, omics signatures, phenotypic readouts, clinical observations or model outputs – to standard entities and relationships in the graph.
This requires more than data ingestion. Effective docking includes identifier harmonization, ontology alignment, metadata capture, confidence scoring, source provenance and evidence weighting. Without those steps, a graph can become a visually appealing but poorly governed data swamp.
Done properly, docking turns a narrow internal signal into a contextualized knowledge asset. A platform-specific readout can be interpreted against disease mechanisms, tissue expression, target safety, druggability, pathway adjacency, competitive activity and known therapeutic evidence.
1. Internal assets | 2. Docking layer | 3. Biomedical KG | 4. Discovery workflows |
|---|---|---|---|
Molecules, biologics, assays, omics, phenotypic readouts, clinical observations, model outputs | Identifier mapping, ontology alignment, provenance, metadata, confidence and evidence weighting | Diseases, genes, proteins, tissues, pathways, drugs, phenotypes and typed relationships | Target prioritization, disease clustering, indication expansion, evidence-backed exploration |
Table 1.Docking proprietary data into a biomedical knowledge graph.
How graph methods support discovery workflows
Once internal data is docked into a wider biomedical graph, several computational workflows become possible. Each should be framed as decision support and hypothesis generation, not automated truth.
Target prioritization through network centrality. Personalized PageRank and related centrality methods can rank targets by their proximity to a disease module, internal asset signature or mechanism of interest. The key comparison is global relevance versus local relevance: a target may be broadly important in biology, but especially relevant to the proprietary biology represented by a specific asset or platform.
Asset grouping and disease clustering through community detection. Community detection can reveal subgraphs that connect diseases, targets, phenotypes or assets through shared molecular mechanisms. This helps teams move beyond surface-level clinical similarity and find biologically coherent groups that may support portfolio clustering or indication expansion.
Indication expansion as link prediction. Knowledge graph embeddings and link-prediction models can ask which plausible relationships are missing from the current graph. For example, a model may rank candidate disease indications for a docked asset by learning patterns across known asset-target, target-disease, pathway and tissue relationships. These outputs are ranked hypotheses, not validated indications.
LLM-ready exploration through governed graph interfaces. Because graph relationships are semantically typed, they can be exposed through controlled query layers, retrieval workflows or AI-agent interfaces. Model Context Protocol servers and similar access patterns can help internal agents query graph-backed context, but they should be implemented with source attribution, access control, audit trails and guardrails against unsupported biological claims.
Related reading: Target Identification: From Poor Data to Quality Predictions | AI Lab Loop for knowledge graph model training
Build your own, or build on an enterprise foundation?
Teams can build a biomedical knowledge graph internally. Public starting points include PrimeKG, OptimusKG, OREGANO and resources listed through KG-Hub. Data teams can use schemas and frameworks such as the Biolink Model or BioCypher, graph databases such as Neo4j or Amazon Neptune, and graph ML libraries such as PyTorch Geometric, DGL, PyKEEN, graph-tool, igraph and NetworKit.
That route is feasible, but it is not a weekend exercise. The hard work is usually not drawing nodes and edges. It is maintaining entity resolution, schema governance, provenance, evidence scoring, release management, quality assurance, model validation and user-facing workflows as the underlying science evolves.
For many organizations, the strategic decision is therefore not build versus buy in the abstract. It is where the company wants to invest scarce scientific data engineering time: in constructing a context layer from scratch, or in using an enterprise-ready foundation and focusing effort on proprietary data, validation and decision workflows.
Related reading: Why Data Quality Matters in AI-powered Drug Discovery | Scaling AI in Life Sciences: Why Data Infrastructure Determines Success | Leveraging Public Datasets for AI-Driven Discovery
How Ardigen can help
Ardigen can help teams accelerate this path through the Ardigen Drug Discovery Knowledge Graph (ARDDKG), Data Universe, AI Lab Loop experience and graph machine learning expertise. ARDDKG provides a biomedical context layer that can be adapted to client-specific discovery questions, proprietary evidence and infrastructure requirements.
The practical value is not only the graph itself. Ardigen can support the docking workflow around it: mapping internal assets to standard entities, designing evidence-aware relationships, building graph ML and link-prediction workflows, evaluating outputs, and exposing the graph through controlled interfaces for scientists and AI agents.
If a company already has an internal knowledge graph or a preferred third-party provider, Ardigen can build on that foundation instead of replacing it. The engagement can focus on graph ML, validation, evidence reporting, data-docking strategy and governed LLM or agent access layers.
Commercial, deployment and IP terms should be defined in the project agreement, but the operating principle is simple: proprietary data, generated hypotheses and client-specific model improvements should be governed in a way that protects the client’s competitive advantage.
Conclusion: turn siloed data into contextualized hypotheses
Internal data is a strategic asset, but it becomes more useful when it is connected to the biological, clinical and evidential context around it. Knowledge graphs give pharma and biotech teams a way to make those connections explicit, computable and reusable across discovery workflows.
The question is not whether your organization has data. It is whether your data is connected well enough to answer the next portfolio question: where else can this asset go, and what evidence supports that move?
Next step
Book a 30-minute Knowledge Graph Fit Assessment with Ardigen. We will map one priority asset, indication or target area against the graph-docking workflow and identify the highest-value context layers, validation questions and next actions.
Frequently Asked Questions
What is a biomedical knowledge graph in drug discovery?
It is a structured representation of biomedical entities, such as genes, proteins, diseases, drugs, tissues and pathways, connected by typed relationships. In drug discovery, it helps teams explore how assets and targets relate to broader disease biology and existing evidence.
How does a knowledge graph help with indication expansion?
A knowledge graph can connect an internal asset to disease mechanisms, target biology, tissue expression, pathway proximity and existing drug evidence. Graph ML and link prediction can then rank plausible disease-asset relationships for scientific review and experimental validation.
What does docking proprietary data into a knowledge graph mean?
Docking means mapping internal data to standard graph entities and relationships with identifiers, ontologies, metadata, provenance and confidence scores. It is a controlled integration process, not simply uploading files into a graph database.
How is a knowledge graph different from a normal database?
A conventional database usually stores records in tables or documents. A knowledge graph emphasizes relationships, semantics and evidence paths, making it easier to traverse connected biological context across diseases, genes, proteins, tissues, pathways and assets.
Can graph machine learning prioritize drug targets?
Yes. Methods such as centrality analysis, community detection, embeddings and link prediction can support target prioritization. These methods produce ranked hypotheses that still require domain review, validation design and experimental follow-up.
Are knowledge graph predictions automatically reliable?
No. Prediction quality depends on graph schema, source quality, entity resolution, provenance, training labels, validation design and leakage controls. A knowledge graph improves context, but it does not remove the need for expert review and validation.
Can an LLM talk to a biomedical knowledge graph?
Yes, but the safer approach is to connect the LLM through governed query, retrieval or agent layers. The interface should include access control, audit trails, source attribution and guardrails against unsupported biological claims.
How can Ardigen support a company that already has a knowledge graph?
Ardigen can build on an existing graph by adding docking workflows, graph ML models, evidence reporting, evaluation pipelines and controlled AI-agent interfaces. This avoids rebuilding infrastructure that already works while increasing discovery value from the graph.
Technical editing: Ardigen expert: Sergiusz Wesolowski, PhD
References
[1] Hogan, A., Blomqvist, G., Cochez, M., d’Amato, C., de Melo, G., Gutierrez, C., … & Zimmermann, A. (2021). Knowledge graphs. ACM Computing Surveys, 54(4), 1-37.
[2] Bonner, S., Barrett, I. P., Ye, M., Swiers, R., Engkvist, O., Hoover, B., … & Bender, A. (2021). Evaluating the use of knowledge graphs for drug repurposing. Computational and Structural Biotechnology Journal, 19, 6532-6541.
[3] Baghel, T., Rawat, P., & Kuhlmann, L. (2021). A survey on graph kernels and neural networks for drug discovery and development. IEEE Access, 9, 151589-151609.
[4] Ioannidis, V. N., Song, X., Tong, H., & Maciejewski, R. (2020). Towards deep learning on graphs with heterogeneous attributes. In 2019 IEEE International Conference on Data Mining (ICDM) (pp. 349-358). IEEE.
[5] Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J., Appleton, G., Axton, M., Baak, A., … & Bourne, P. E. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3(1), 1-9.
[6] Kitano, H. (2002). Computational systems biology. Nature, 420(6912), 206-210.
[7] Angles, R., & Gutierrez, C. (2008). Survey of graph database models. ACM Computing Surveys, 40(1), 1-39.
[8] Wang, Q., Mao, Z., Wang, B., & Guo, L. (2017). Knowledge graphs completion via complex tensor factorization. Journal of Machine Learning Research, 18(73), 1-49.
[9] Chandak, P., Huang, K., & Zitnik, M. (2023). Building a knowledge graph to enable precision medicine. Nature Scientific Reports, 13, 1243.
[10] Rojas-Carulla, M., Asano, Y. M., Lamb, R., Costabello, L., & Minervini, P. (2019). Oracle-guided curriculum learning for visual question answering. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 5843-5851). IEEE.
[11] Zitnik, M., Agrawal, M., & Leskovec, J. (2018). Modeling polypharmacy side effects with graph convolutional networks. Bioinformatics, 34(13), i457-i466.
[12] Caruana, R. (2005). Intelligent tutoring systems. In Handbook of Educational Data Mining (pp. 3-22). CRC Press.
[13] Morrison, A. L., & Miranda, A. T. (2018). Algorithms and data structures for sparse Boolean matrix multiplication. In 2018 Proceedings of the Twenty-First International Workshop on Software and Compilers for Embedded Systems (pp. 55-62). ACM.
[14] Blondel, V. D., Guillaume, J. L., Lambiotte, R., & Lefebvre, E. (2008). Fast unfolding of communities in large networks. Journal of Statistical Mechanics: Theory and Experiment, 2008(10), P10008.
[15] Fortunato, S., & Hric, D. (2016). Community detection in networks: A user guide. Physics Reports, 659, 1-44.
[16] Schaub, M. T., Delvenne, J. C., Rosvall, M., & Lambiotte, R. (2019). The many facets of community detection in complex networks. Applied Network Science, 4(1), 1-26.
[17] Nickel, M., Rosasco, L., & Poggio, T. (2016). Holographic embeddings of knowledge graphs. In AAAI (pp. 1955-1961).
[18] Bordes, A., Usunier, N., Garcia-Durán, A., Weston, J., & Yakhnenko, O. (2013). Translating embeddings for modeling multi-relational data. In Advances in Neural Information Processing Systems (pp. 2787-2795).
[19] Alshahrani, M., Khan, M. A., Maddouri, O., Kinjo, A. R., & Queralt-Rosinach, N. (2017). Neuro-symbolic representation learning on biological knowledge graphs. Bioinformatics, 33(17), 2723-2730.
[20] Sequeda, J. F., & Lassila, O. (2021). Designing and building knowledge graphs. Morgan & Claypool Publishers.
[21] Shadbolt, N., Berners-Lee, T., & Hall, W. (2006). The semantic web revisited. IEEE Intelligent Systems, 21(3), 96-101.
[22] Putman, T. E., Vempati, U. D., Greene, A., Tseng, T., Hamer, M. K., Pilarczyk, M., … & Pillich, R. T. (2021). Harmonizome connection specificity and network centrality in the genome-scale integrated protein interaction and pathway databases. Database, 2021, baab016.
[23] Moxon, S. A. R., Susaki, T. W., & Harris, B. D. (2021). Biolink Model: A universal ontology for data integration and a flexible input schema for all knowledge graph ingestion efforts. Journal of Chemical Information and Modeling, 61(1), 2-14.
[24] Meert, K., Zoch, A., & Müller, B. (2022). BioCypher: A framework for knowledge graph construction from heterogeneous biomedical data. bioRxiv, 2022-11.
[25] Neo4j Inc. (2023). Neo4j Graph Database. Retrieved from https://neo4j.com/
[26] Amazon Web Services. (2023). Amazon Neptune – Fully managed graph database service. Retrieved from https://aws.amazon.com/neptune/
[27] Fey, M., & Lenssen, J. E. (2019). Fast graph representation learning with PyTorch Geometric. arXiv preprint arXiv:1903.02428.
[28] Wang, M., Yu, L., Gan, D., Yu, Y., Jiang, S., Liu, Q., … & Zhou, J. (2019). DGL: A graph learning library for OpenMolecules. In Advances in Neural Information Processing Systems (pp. 5505-5516).
[29] Ali, M., Berrendorf, M., Hoover, B., Natarajan, S., Cariaso, M., Expósito, F., … & Smaili, F. Z. (2021). PyKEEN 1.0: A Python library for reproducible knowledge graph embeddings. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management (pp. 4808-4817).
[30] Peixoto, T. P. (2014). The graph-tool Python library. figshare.
[31] Csárdi, G., & Nepusz, T. (2006). The igraph software package for complex network research. InterJournal Complex Systems, 1695(5), 1-9.
[32] Schlag, R., Gottesburen, B., Schlag, S., Sanders, P., & Walshaw, C. (2023). High-quality graph partitioning for large-scale computing. In 2021 IEEE International Parallel and Distributed Processing Symposium (IPDPS) (pp. 923-933). IEEE.