# The CFDE Workbench

> **NIH NIH OT2** · ICAHN SCHOOL OF MEDICINE AT MOUNT SINAI · 2024 · $1,750,000

## Abstract

Abstract
The NIH Common Fund (CF) programs have produced transformative datasets, databases,
methods, bioinformatics tools and workflows that are significantly advancing biomedical research
in the United States and worldwide. Currently, CF programs are mostly isolated. However,
integrating data from across CF programs has the potential for synergistic discoveries. In addition,
since CF programs have a time limit of 10 years, sustainability of the widely used CF digital
resources after the programs expire is critical. To address these challenges, the NIH established
the Common Fund Data Ecosystem (CFDE) program which has been recently approved to
continue to its second new phase. For the second phase of the CFDE, this project will establish
the Data Resource Center (DRC) and the Knowledge Center (KC). Our efforts will culminate in
producing The CFDE Workbench which will be composed of three main products: the CFDE
information portal, the CFDE data resource portal, and the CFDE knowledge portal. These three
web portals will be full-stack web-based applications with a backend database and will be
integrated into one public site.
The CFDE information portal will be the entry point to the other two portals. It will contain
information about the CFDE in a dedicated About page, information about each participating and
non-participating CF program, information about each data coordination center (DCC), a link to a
catalog of CF datasets, and a link to a catalog of CF tools and workflows, news, events, funding
opportunities, standards and protocols, educational programs and opportunities, social media
feeds, and publications.
The CFDE data resource portal will contain metadata, data, workflows, and tools which are the
products of the CF programs, and their data coordination centers (DDCs). We will adopt the C2M2
data model for storing information about metadata describing DCC datasets. We will also archive
relatively small omics datasets that do not have a home in widely established repositories and do
not require PHI protection. In addition, we will expand the cataloging to CF tools, APIs, and
workflows. Importantly, we will develop a search engine that will index and present results from
all these assembled digital assets. In addition, continuing the work established in the CFDE pilot
phase, users of the data portal will be able to fetch identified datasets through links provided by
the DCCs via the DRS protocol. This will include links to raw and processed data.
The CFDE knowledge portal will provide access to CF programs processed data in various
formats including: 1) knowledge graph assertions; 2) gene, drug, metabolite, and other set
libraries; 3) data matrices ready for machine learning and other AI applications; 4) signatures; and
5) bipartite graphs. In addition, the extract, transform, and load (ETL) scripts to process the data
into these formats will be provided. Since such processed data is relatively small, we will archive
and serve this proc...

## Key facts

- **NIH application ID:** 11080094
- **Project number:** 3OT2OD036435-01S1
- **Recipient organization:** ICAHN SCHOOL OF MEDICINE AT MOUNT SINAI
- **Principal Investigator:** Avi Ma'ayan
- **Activity code:** OT2 (R01, R21, SBIR, etc.)
- **Funding institute:** NIH
- **Fiscal year:** 2024
- **Award amount:** $1,750,000
- **Award type:** 3
- **Project period:** 2023-09-18 → 2028-09-17

## Primary source

NIH RePORTER: https://reporter.nih.gov/project-details/11080094

## Citation

> US National Institutes of Health, RePORTER application 11080094, The CFDE Workbench (3OT2OD036435-01S1). Retrieved via AI Analytics 2026-05-23 from https://api.ai-analytics.org/grant/nih/11080094. Licensed CC0.

---

*[NIH grants dataset](/datasets/nih-grants) · CC0 1.0*
