30-Day Hospital Readmission Risk Prediction
End-to-end machine learning pipeline for identifying hospital encounters involving patients with diabetes at elevated risk of readmission within 30 days.
A recall-first operating point, evaluated once.
799 of 1,122 actual 30-day readmissions identified on the locked test set.
The selected operating point intentionally prioritized recall and therefore produced a substantial false-positive burden.
Prioritizing limited follow-up resources.
Hospitals have limited resources for post-discharge follow-up. This retrospective project explores whether machine learning can help prioritize hospital encounters involving patients with diabetes that appear at elevated risk of 30-day readmission.
Diabetes 130-US Hospitals
More than 100,000 encounters across 130 U.S. hospitals from 1999–2008.
Encounters involving diabetes
Results should not be generalized to all patients or hospital populations.
From raw encounters to an untouched test.
- 01
Raw data
100K+ encounters
- 02
Eligibility + target
Define 30-day readmission label
- 03
Patient-level split
Zero patient overlap
- 04
Feature engineering
ICD-9 grouping · administrative mappings · semantic missing handling
- 05
Model development
Dummy · Logistic Regression · Random Forest · XGBoost · CatBoost
- 06
Grouped CV + tuning
5-fold patient-aware CV · PR-AUC optimization
- 07
Threshold + selection
Validation-based operating point · XGBoost selected
- 08
Locked test evaluation
Final untouched test assessment
- 09
Explainability + errors
TreeSHAP · TP / TN / FP / FN analysis
Comparison before selection.
DummyClassifier, Logistic Regression, Random Forest, XGBoost, and CatBoost were evaluated. PR-AUC was the primary selection metric because only approximately 11% of encounters were positive.
| Model | PR-AUC | Relative bar |
|---|---|---|
| CatBoost | 0.222551 | |
| Random Forest | 0.221958 | |
| XGBoost | 0.221426 | |
| Logistic Regression | 0.212923 |
CatBoost achieved slightly higher validation PR-AUC, but XGBoost was selected because it provided comparable predictive performance, slightly better precision and fewer false-positive alerts around the selected ~70% recall operating point, with substantially lower computational cost.
The false positives stay visible.
The operating threshold was locked before test evaluation and intentionally prioritized recall over precision.
- Test encounters
- 10,004
- Actual positives
- 1,122
- PR-AUC
- 0.232455
- ROC-AUC
- 0.680048
- Precision
- 16.35%
- Recall
- 71.21%
- F1
- 0.265979
- Flagged
- 48.84%
Confusion matrix
Predicted outcome →
↑ Actual outcome
What shaped model behavior.
TreeSHAP analysis highlighted the following grouped features. The portfolio includes the final ranked findings here; the source image will be added when a verified export is available.
- 01Discharge disposition
- 02Prior inpatient utilization
- 03Primary diagnosis group
- 04Medical specialty
- 05Payer information
- 06Secondary diagnosis group
- 07Number of diagnoses
- 08Diabetes medication status
- 09Insulin status
- 10Age
Where the classifier struggled.
Missed readmissions generally had weaker historical utilization signals than correctly identified readmissions, including fewer prior inpatient and emergency encounters, and were more frequently discharged home.
False-positive encounters often showed stronger risk-like patterns, including greater prior utilization, longer stays, more medications, more diagnoses, and more laboratory procedures.
A flag prompts review—not a decision.
- Patient discharge
- Risk score
- Elevated-risk flag
- Care-team review
- Possible follow-up
Clinical decisions remain with healthcare professionals.
Evidence boundaries matter.
Historical data
1999–2008 encounters may not reflect modern clinical practice.
False-positive burden
Higher recall required flagging many encounters that were not readmitted.
No prospective clinical validation
The model was evaluated retrospectively and has not been validated for clinical deployment.
Population scope
Results apply to this diabetic hospital-encounter dataset and should not be generalized to all patients.
Further work should investigate subgroup disparities. Results are associative, not causal, and deployment would require prospective validation plus a clinically validated threshold.
Want the technical details?
The repository contains the preprocessing pipeline, tests, tuning, SHAP analysis, and error-analysis artifacts. A verified public repository link will appear here once configured.
Back to selected work