Integrative Proteome-Wide Association Study and Large Language Model-Based Prioritization of Causal Protein Targets of Type 2 Diabetes
Trial Snapshot
- Phase
- Not Applicable
- Status
- Not yet recruiting
- Enrollment
- 31,994
- Primary Endpoint
- Performance of Population-Specific Protein Prediction Models
Study Overview
Brief Summary
Type 2 diabetes is a common condition in which the body has difficulty controlling blood sugar. This study will use existing genetic, protein, and health data from prior research studies to identify blood proteins that may play a role in type 2 diabetes. The study will not recruit participants, provide treatment, or collect new samples. Researchers will use computer-based analyses to identify and prioritize protein targets for future laboratory and clinical research. The goal is to support the development of better approaches for understanding, preventing, and treating type 2 diabetes.
Detailed Description
Type 2 diabetes is a major cause of illness and health disparities. Genetic association studies can identify regions of the genome associated with disease risk, but they do not always identify the proteins or biological mechanisms that contribute to disease development. This project will conduct a retrospective secondary analysis of existing, controlled-access genetic, proteomic, and phenotype data, including data accessed through the UK Biobank and other previously collected datasets. No new participants will be recruited, enrolled, contacted, treated, or followed as part of this study.
The study will use proteome-wide association methods to evaluate whether genetically predicted circulating protein levels are associated with type 2 diabetes risk. Analyses will consider evidence across African American and European American datasets when available, with attention to population-specific and shared signals. Statistical genetic evidence will be integrated with relevant biological and clinical information to prioritize protein targets that may have a causal role in type 2 diabetes.
Large language model-based methods will be used as a structured evidence-synthesis tool to organize and summarize publicly available information relevant to prioritized proteins, including biological function, disease relevance, and potential therapeutic tractability. All computational results will be reviewed by the research team. The project will generate reproducible analytic workflows, a ranked list of candidate protein targets, and hypotheses for future experimental validation. Findings are intended for research use and will not be used to make clinical decisions for individual patients.
Study Design
- Study Type
- Observational
- Observational Model
- Other
- Time Perspective
- Retrospective
Eligibility Criteria
- Ages
- 40 Years to 84 Years (Adult, Older Adult)
- Sex
- All
- Accepts Healthy Volunteers
- Yes
Inclusion Criteria
- •No new participants will be recruited for this study. The study will conduct a retrospective secondary analysis of existing, controlled-access datasets. Eligible records are from adult participants aged 40 to 84 years at enrollment in the source studies who have available genetic data and plasma proteomic data for population-specific protein prediction modeling. Participants with and without type 2 diabetes may be included, depending on the source dataset and analytic objective. Type 2 diabetes genome-wide association summary statistics will also be used; no individual-level participant contact or enrollment will occur.
Exclusion Criteria
- Not provided
Arms & Interventions
African American Participants
Retrospective analysis of existing genetic and OLINK proteomic data to develop and validate population-specific protein prediction models and evaluate genetically predicted proteins associated with type 2 diabetes risk.
European American Participants
Retrospective analysis of existing genetic and OLINK proteomic data to develop and validate population-specific protein prediction models and evaluate genetically predicted proteins associated with type 2 diabetes risk.
Outcomes
Primary Outcomes
Performance of Population-Specific Protein Prediction Models
Time Frame: Up to 12 months
Cross-validated and external validation R-squared values for cis-SNP-based prediction models of 2,943 plasma proteins. Models with reproducible performance (R-squared greater than 0.01) will be retained
Secondary Outcomes
- Genetically Predicted Protein Associations With Type 2 Diabetes Risk(Up to 12 months)
- Prioritized Protein Targets With Citation-Grounded Functional Evidence(Up to 12 months)
