Skip to main content
Clinical Trials/NCT07832656
NCT07832656Not yet recruitingNot Applicable

Integrative Proteome-Wide Association Study and Large Language Model-Based Prioritization of Causal Protein Targets of Type 2 Diabetes

Louisiana State University Health Sciences Center in New Orleans0 sites31,994 target enrollmentStarted: September 15, 2026Last updated:
Conditions

Trial Snapshot

Phase
Not Applicable
Status
Not yet recruiting
Enrollment
31,994
Primary Endpoint
Performance of Population-Specific Protein Prediction Models

Study Overview

Brief Summary

Type 2 diabetes is a common condition in which the body has difficulty controlling blood sugar. This study will use existing genetic, protein, and health data from prior research studies to identify blood proteins that may play a role in type 2 diabetes. The study will not recruit participants, provide treatment, or collect new samples. Researchers will use computer-based analyses to identify and prioritize protein targets for future laboratory and clinical research. The goal is to support the development of better approaches for understanding, preventing, and treating type 2 diabetes.

Detailed Description

Type 2 diabetes is a major cause of illness and health disparities. Genetic association studies can identify regions of the genome associated with disease risk, but they do not always identify the proteins or biological mechanisms that contribute to disease development. This project will conduct a retrospective secondary analysis of existing, controlled-access genetic, proteomic, and phenotype data, including data accessed through the UK Biobank and other previously collected datasets. No new participants will be recruited, enrolled, contacted, treated, or followed as part of this study.

The study will use proteome-wide association methods to evaluate whether genetically predicted circulating protein levels are associated with type 2 diabetes risk. Analyses will consider evidence across African American and European American datasets when available, with attention to population-specific and shared signals. Statistical genetic evidence will be integrated with relevant biological and clinical information to prioritize protein targets that may have a causal role in type 2 diabetes.

Large language model-based methods will be used as a structured evidence-synthesis tool to organize and summarize publicly available information relevant to prioritized proteins, including biological function, disease relevance, and potential therapeutic tractability. All computational results will be reviewed by the research team. The project will generate reproducible analytic workflows, a ranked list of candidate protein targets, and hypotheses for future experimental validation. Findings are intended for research use and will not be used to make clinical decisions for individual patients.

Study Design

Study Type
Observational
Observational Model
Other
Time Perspective
Retrospective

Eligibility Criteria

Ages
40 Years to 84 Years (Adult, Older Adult)
Sex
All
Accepts Healthy Volunteers
Yes

Inclusion Criteria

  • No new participants will be recruited for this study. The study will conduct a retrospective secondary analysis of existing, controlled-access datasets. Eligible records are from adult participants aged 40 to 84 years at enrollment in the source studies who have available genetic data and plasma proteomic data for population-specific protein prediction modeling. Participants with and without type 2 diabetes may be included, depending on the source dataset and analytic objective. Type 2 diabetes genome-wide association summary statistics will also be used; no individual-level participant contact or enrollment will occur.

Exclusion Criteria

  • Not provided

Arms & Interventions

African American Participants

Retrospective analysis of existing genetic and OLINK proteomic data to develop and validate population-specific protein prediction models and evaluate genetically predicted proteins associated with type 2 diabetes risk.

European American Participants

Retrospective analysis of existing genetic and OLINK proteomic data to develop and validate population-specific protein prediction models and evaluate genetically predicted proteins associated with type 2 diabetes risk.

Outcomes

Primary Outcomes

Performance of Population-Specific Protein Prediction Models

Time Frame: Up to 12 months

Cross-validated and external validation R-squared values for cis-SNP-based prediction models of 2,943 plasma proteins. Models with reproducible performance (R-squared greater than 0.01) will be retained

Secondary Outcomes

  • Genetically Predicted Protein Associations With Type 2 Diabetes Risk(Up to 12 months)
  • Prioritized Protein Targets With Citation-Grounded Functional Evidence(Up to 12 months)

Investigators

Sponsor Class
Other
Responsible Party
Sponsor

Similar Trials