跳至主要内容
临床试验/NCT04574882
NCT04574882已完成不适用

Using Digital Data to Predict Cardiovascular Health and Health Care Utilization

University of Pennsylvania1 个研究点 分布在 1 个国家目标入组 781 人开始时间: 2020年9月25日最近更新:
适应症

试验速览

阶段
不适用
状态
已完成
入组人数
781
试验地点
1
主要终点
Latent Dirichlet Allocation (LDA) Topics - Topics / Themes Discussed Between Patients With and Without Heart Disease

研究概览

简要总结

This project seeks to identify and characterize features derived from digital data (e.g. social media, online search, mobile media) which are associated with coronary heart disease (CHD) and related risk factors, and develop models that use digital data and conventional predictive models to predict CHD risk and health care utilization.

详细描述

Cardiovascular disease is the leading cause of death in the US. While secondary prevention approaches have improved longevity of patients, risk factors and adverse health behaviors (e.g., physical inactivity, smoking) are highly prevalent, and in most contemporary series, less than 1% of adults meet all factors of ideal CV health. The logistics and practicalities of meeting the goal of ideal CV health have not been clearly elucidated. Practice guidelines recommend using the Framingham risk score (FRS) or other risk prediction tools to classify patients' risk of CV disease. These models however are imprecise and there is increasing focus on identifying markers that provide better measures of risk. As digital platforms are increasingly used to document lifestyle and health behaviors, data from digital sources may provide a window into manifestations of novel risk factors and potentially a better characterization of existing risk factors. While it seems like a cliche to mention the profound impact of digital data on everyday lives, there is indeed great substance in the opportunities these new media provide for understanding behavioral, social, and environmental determinants of health. This project seeks to identify and characterize features derived from digital data (e.g. social media, online search, mobile media) which are associated with coronary heart disease (CHD) and related risk factors, and develop models that use digital data and conventional predictive models to predict CHD risk and health care utilization.

研究设计

研究类型
Observational
观察模型
Case Control
时间视角
Cross Sectional

入排标准

年龄范围
30 Years 至 74 Years(Adult, Older Adult)
性别
All
接受健康志愿者
是

入选标准

  • •30 - 74 years of age
  • •Willing to sign informed consent
  • •Primarily English speaking (for language analysis)
  • •Has an account on any of the following digital data platforms (Facebook, Instagram, Twitter Reddit, Google (gmail), or smartphone or wearable device such as Apple Health, Fitbit, Samsung Health, MapMyFitness or Garmin) and willing to share data
  • •If has social media account, Instagram or Facebook, willing to share historical and prospective data (60 days) If has Google (gmail) account, willing to download and share google takeout zip file
  • •If has smartphone or wearable device, willing to share step data
  • •Willing to share access to medical health records
  • •Willing to share healthcare insurance information

排除标准

  • •Patient does not meet age inclusion criteria above
  • •Does not use and post on digital data sources we are studying or unwilling to donate data
  • •Patient is in severe distress, e.g. respiratory, physical, or emotional distress
  • •Patient is intoxicated, unconscious, or unable to appropriately respond to questions

结局指标

主要结局

Latent Dirichlet Allocation (LDA) Topics - Topics / Themes Discussed Between Patients With and Without Heart Disease

时间窗: Through study completion, an average of 3 years

The primary outcome is topics and features (derived using the LDA method for clustering language data). For each participant, we included all available Facebook wall posts from the start of their account history through data collection, regardless of whether they occurred before or after a CHD diagnosis. We examined associations between linguistic features (unigrams, LIWC categories, LDA topics) and cardiovascular case status (CHD presence vs absence) using Pearson correlation and logistic regression. Latent LDA, a systematic method to identify text-based themes, was applied to generate 200 clusters of co-occurring words ("topics"). For each feature type (unigram, LIWC category, LDA topic), we fit separate logistic regression models and calculated Pearson correlation coefficients to assess predictive value for case status. Each language-derived feature was encoded as a normalized frequency count per user to enable consistent comparison across participants.

次要结局

未报告次要终点

研究者

申办方类型
Other
责任方
Sponsor

研究点 (1)

Loading locations...

相似试验