Skip to main content
Clinical Trials/NCT07590154
NCT07590154CompletedNot Applicable

Unsupervised Deep Representation Learning for Clinical Stratification in Substance Use Disorders

Lauro Gutiérrez Castro1 site in 1 country155 target enrollmentStarted: March 25, 2024Last updated:
Conditions
Interventions

Trial Snapshot

Phase
Not Applicable
Status
Completed
Sponsor
Enrollment
155
Locations
1
Primary Endpoint
Latent dimension scores

Study Overview

Brief Summary

Substance use disorders (SUDs) show considerable clinical heterogeneity that limits the usefulness of traditional categorical diagnoses. This observational, cross-sectional study aims to apply an unsupervised deep learning method - an autoencoder - to learn continuous latent representations from standardised psychometric data and to explore whether those representations can help stratify clinical subpopulations. The investigators will recruit 155 adults undergoing residential treatment for SUD. Participants will complete six validated instruments assessing impulsivity (BIS-11), anger regulation (STAXI-2), behavioural activation/avoidance (BADS), borderline symptomatology (BSL-23), generalised anxiety (GAD-7), and environmental reward (EROS). Demographic and clinical variables (age, sex, primary substance, years of use, prior treatments) will also be recorded.

After data cleaning and standardisation (z-scores), a symmetric autoencoder with a 12-dimensional bottleneck (architecture 21-32-24-12-24-32-21) will be trained using mean squared error loss. Regularisation includes L2 weight decay and dropout. The model will be trained 30 times with different random seeds to assess stability; the five best models (by validation pseudo-R²) will be combined into a weighted ensemble. Five-fold cross-validation will evaluate generalisation. For comparison, principal component analysis (PCA) will be applied to the same data. Gaussian mixture models (GMM) will be fitted on the latent space to explore potential clinical subgroups.

The primary outcome is the stability of the latent representation (coefficient of variation of validation MSE across runs). Secondary outcomes include reconstruction performance (pseudo-R²) of the ensemble, comparison with PCA, and the interpretability of latent dimensions via correlations with original variables. GMM results will be described using BIC, silhouette width, bootstrap stability, and clinical characterisation of clusters.

This study does not involve any intervention. Results will be hypothesis-generating and require external validation. No automated clinical decisions will be made.

Detailed Description

Substance use disorders (SUDs) are characterised by substantial heterogeneity in clinical presentation, behavioural patterns, emotional regulation difficulties, impulsivity, and treatment response. Individuals with the same categorical diagnosis may differ considerably in symptom severity, comorbid psychopathology, and psychosocial functioning. This variability limits the explanatory value of traditional diagnostic classifications and supports the development of dimensional and data-driven approaches for patient characterisation.

Recent advances in machine learning provide methods capable of identifying latent structures within complex clinical datasets. Autoencoders, a form of unsupervised deep learning, can learn compact nonlinear representations of multidimensional data while preserving relevant information from the original variables. Compared with traditional linear dimensionality reduction methods such as principal component analysis (PCA), autoencoders may better capture complex interactions among psychological and behavioural variables. When combined with probabilistic clustering approaches such as Gaussian mixture models (GMM), these latent representations may facilitate the identification of clinically meaningful patient subgroups.

The purpose of this observational study is to apply an autoencoder model to psychometric and clinical data obtained from adults receiving residential treatment for substance use disorders. The study aims to explore latent dimensions underlying symptom and behavioural variability and to evaluate whether these dimensions support stable subgroup identification.

Primary Objective:

To learn a 12-dimensional latent representation from standardised psychometric and clinical variables using an autoencoder model and evaluate the stability of this representation across repeated training procedures.

Study Design

Study Type
Observational
Observational Model
Case Only
Time Perspective
Cross Sectional

Eligibility Criteria

Ages
18 Years to 60 Years (Adult)
Sex
All
Accepts Healthy Volunteers
No

Inclusion Criteria

  • DSM-5 diagnosis of Substance Use Disorder (SUD), confirmed by a psychiatrist or clinical psychologist.
  • Age ≥ 18 years.
  • Currently admitted to a residential addiction treatment center at the time of assessment.
  • Ability to complete the psychometric questionnaires independently.
  • Written informed consent.

Exclusion Criteria

  • Active psychotic disorder (e.g., schizophrenia, delusional disorder) not stabilized pharmacologically.
  • Severe cognitive impairment (dementia, severe brain injury) that prevents understanding the questionnaire items.
  • Language barriers or illiteracy that prevent self-administration of the scales.
  • Scheduled discharge from the center within 7 days of the assessment date.

Arms & Interventions

Total sample (residential treatment)

Adult patients (N=155) with DSM-5 TR substance use disorder receiving residential treatment. All participants completed six psychometric scales (BIS-11, STAXI-2, BADS, BSL-23, GAD-7, EROS) and provided demographic/clinical data in a single cross-sectional session. No intervention was administered.

Intervention: No intervention (observational only) (Other)

Outcomes

Primary Outcomes

Latent dimension scores

Time Frame: Baseline (single assessment, cross-sectional)

Twelve continuous latent dimensions derived from the bottleneck layer of a symmetric autoencoder trained on 21 standardized clinical variables. Each dimension represents a compressed, nonlinear combination of the original psychometric indicators (impulsivity, emotion regulation, behavioral activation, borderline symptoms, anxiety, and environmental reward). The dimensions are extracted for each participant after averaging the predictions of an ensemble of the five best autoencoder runs. Unit of Measure: Standardized z-score (mean = 0, SD = 1 in the training sample)

Secondary Outcomes

  • Gaussian mixture model cluster membership(Baseline)
  • Autoencoder reconstruction pseudo-R²(Baseline (computed on the validation split and on the full sample after training))
  • Autoencoder reconstruction mean squared error(Baseline)
  • Coefficient of variation of reconstruction MSE(Baseline (after all runs are completed))
  • Cross-validated reconstruction R²(Baseline)
  • Explained variance by 12 principal components(Baseline)

Investigators

Sponsor
Lauro Gutiérrez Castro
Sponsor Class
Other
Responsible Party
Sponsor Investigator
Principal Investigator

Lauro Gutiérrez Castro

Principal Investigator

Under The Tree Therapeutic Community

Study Sites (1)

Loading locations...

Similar Trials

Cross-sectional Functional Stratification... | Clinical Trial