Everything K-culture — comebacks to K-beauty, straight to your inboxGet it in your inbox

METAL MEDIA

CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation

arXiv:2608.129442026-08-12

One shared AI model trained jointly on heart electricity, pulse, and heart sound outperforms models trained on each signal alone

ECG, PPG, and PCG each observe the same heartbeat but at different moments: electrical activation, pulse arrival at the periphery, and heart sound generation. CardioState-JEPA runs all three through one shared Transformer encoder, training it to predict masked hidden cardiac states rather than raw waveforms, and uses a learned delay aligner to match events across signals that are shifted in time. Across 25 downstream tasks, this single encoder beat the best single-signal self-supervised baseline by 8.2 AUROC points on PPG classification, 18.8 points on PCG murmur detection, and 15.5 points on ECG classification.

METAL MEDIA explanatory visual

CardioState-JEPA's two-stage training pipeline

Evidence statusMeasured results reported

  1. Stage I: unimodal pretrainingECG, PPG, and PCG are each trained separately on abundant single-signal data, masking parts of the signal and predicting the hidden cardiac state at those positions with a shared Transformer encoder
  2. Stage II: delay-aware cross-modal trainingUsing a smaller set of simultaneously recorded pairs (ECG-PPG, ECG-PCG, and trimodal recordings), a learned delay aligner estimates the time offset and one modality predicts another's hidden state at the matching physiological moment
  3. Shared representation spaceAfter training, representations from all three signal types mix together regardless of sensor type (modality silhouette drops from 0.121 to -0.006), while class-relevant structure like arrhythmia grouping improves
  4. Frozen encoder evaluationThe trained encoder is frozen (no further updates) and only a small linear classifier is attached per task, then tested on 25 ECG, PPG, and PCG downstream tasks
  5. Performance comparisonThe shared encoder outperformed the strongest single-signal self-supervised baseline by 8.2 AUROC points on PPG classification, 18.8 points on PCG murmur detection, and 15.5 points on ECG classification
An explanatory diagram made by METAL MEDIA, not a figure supplied by the paper's authors.

What they did

  1. ECG measures the heart's electrical activity, PPG measures blood pulse waveforms (often from a fingertip), and PCG captures the sound of the heartbeat through a microphone; the paper's premise is that these are the same heartbeat observed at different physiological moments.
  2. In Stage I, the model first learns each signal's internal structure separately from large amounts of unimodal (single-signal) data using masked latent prediction.
  3. In Stage II, using a much smaller set of simultaneously recorded pairs, a learned delay aligner estimates the time offset between signals and the model learns to predict one modality's hidden state from another at the matching physiological moment.
  4. Rather than reconstructing raw waveforms, the training target is a hidden ('latent') cardiac state at masked positions, which pushes the model toward shared physiology instead of sensor-specific noise or shape.
  5. The trained encoder is frozen and only a small linear classifier is added on top to evaluate it across 25 real downstream tasks (arrhythmia, murmur detection, activity recognition, blood pressure and respiration estimation, etc.).
Figure 1: Motivation of CardioState-JEPA. ECG, PPG, and PCG observe the same cardiac cycle at different physiological times, motivating a shared cardiac representation.
Figure 1: Motivation of CardioState-JEPA. ECG, PPG, and PCG observe the same cardiac cycle at different physiological times, motivating a shared cardiac representation.
Table 1: PPG Baseline Linear Probe Results (100% training data). ↑: macro-AUROC ×100 (higher is better); ↓: MAE (lower is better). Bold: best; underline: second best. Average (Classification) is the unweighted mean macro-AUROC across the 6 classification tasks. Average (Regression) is the unweighted mean MAE across the 11 regression tasks.
TaskPaPaGei-SPaPaGei-PAnyPPGPulsePPGChronosBoltCardioState-JEPA
PPG Classification Tasks
WESAD ↑65.365.770.166.864.074.5
DaLiA Activity ↑68.471.179.077.377.290.3
MIMIC AF ↑48.481.193.332.446.197.7
PPG Arrhythmia ↑86.388.795.892.690.496.8
BIDMC RR ↑51.338.253.243.538.078.4
UQVital RR ↑40.641.841.829.530.544.6
Average (Classification) ↑60.164.472.257.057.780.4
PPG Regression Tasks
DaLiA HR ↓11.811.55.88.68.23.8
EarSet HR ↓10.79.59.88.98.38.2
WildPPG HR ↓10.09.95.47.77.55.4
Sensors DBP ↓7.27.06.97.46.96.8
Sensors SBP ↓20.019.416.718.618.016.8
UCI DBP ↓7.87.96.76.97.25.8
UCI SBP ↓18.518.115.216.217.410.3
BCG DBP ↓10.112.94.511.210.82.8
BCG SBP ↓13.710.213.512.711.84.4
UQVital SpO2 ↓0.760.850.780.840.830.82
WildPPG RMSSD ↓36.637.634.936.735.934.8
Average (Regression) ↓13.413.210.912.312.19.1
Figure 2: Overview of CardioState-JEPA. Stage I learns unimodal cardiac codes with masked latent prediction from abundant ECG, PPG, and PCG recordings. Stage II uses paired recordings for delay-aware cross-modal prediction, where a learned aligner estimates the physiological offset so that one modality predicts another at the corresponding cardiac time.
Figure 2: Overview of CardioState-JEPA. Stage I learns unimodal cardiac codes with masked latent prediction from abundant ECG, PPG, and PCG recordings. Stage II uses paired recordings for delay-aware cross-modal prediction, where a learned aligner estimates the physiological offset so that one modality predicts another at the corresponding cardiac time.
Table 2: ECG linear probing results under 1%, 10%, and 100% label fractions across six datasets. Avg is the mean macro-AUROC across all 18 settings (six datasets × three label fractions). Label-supervised and ECG-text models are shown as reference models but are excluded from the SSL bold/underline comparison.
PTB-XL SuperPTB-XL SubPTB-XL FormPTB-XL RhythmCPSCCSN
Method1%10%100%1%10%100%1%10%100%1%10%100%1%10%100%1%10%100%Avg
Self-supervised ECG-only models
SimCLR63.469.873.560.868.373.455.057.062.551.469.477.759.868.576.559.067.373.265.9
BYOL71.773.876.557.267.471.648.761.670.842.074.477.260.974.478.854.271.974.767.1
BarlowTwins72.976.078.462.670.874.352.160.466.150.173.577.655.172.878.460.771.677.468.4
MoCo-v373.276.778.355.969.276.750.363.771.351.471.774.362.176.775.354.674.377.768.5
SimSiam73.272.775.662.569.376.455.262.971.349.369.575.958.472.975.358.368.677.468.0
TS-TCC70.775.978.953.567.077.948.061.871.243.369.578.257.173.678.755.368.576.867.0
CLOCS68.973.476.357.972.676.252.058.072.747.271.976.359.677.877.554.471.976.167.8
ASTCL72.577.381.061.968.876.544.160.967.052.472.076.157.977.079.556.470.975.868.2
CRT69.778.277.262.070.878.746.459.568.747.473.574.458.076.482.056.273.778.868.4
ST-MEM61.166.971.454.157.963.655.760.066.151.165.474.956.763.370.459.866.971.463.2
Label-supervised ECG model
ECGFounder83.187.289.673.676.581.761.169.785.187.091.694.687.693.696.371.981.791.983.5
ECG-Text foundation models
MERL82.486.388.764.980.684.758.372.479.753.382.988.370.385.390.666.682.788.078.1
HeartLang78.984.486.769.875.183.956.466.179.572.677.391.171.184.891.368.476.189.878.0
AnyChat82.184.486.872.173.282.258.469.777.173.184.794.583.788.992.671.279.589.980.2
ESI72.380.883.761.669.075.656.959.372.856.671.481.768.278.983.860.165.179.471.0
D-BETA83.288.490.177.782.985.270.178.984.086.692.896.785.591.494.980.087.490.785.9
ECG-FM81.387.089.471.776.983.064.676.386.177.091.096.384.792.695.776.087.795.084.0
CardioState-JEPA (Ours)81.386.489.171.079.186.460.570.083.987.892.496.487.693.495.875.183.193.884.1
Figure 3: T-SNE of cardiac codes from co-recorded ECG, PPG, and PCG samples after Stage I (left) and Stage II (right).
Figure 3: T-SNE of cardiac codes from co-recorded ECG, PPG, and PCG samples after Stage I (left) and Stage II (right).
Table 3: PCG linear probe results on the murmur detection and abnormal heart sound detection tasks (macro-AUROC ×100, ↑; mean±std over 3 seeds).
MethodCirCor ↑CinC2016 ↑
StethoLM64.9 ± 1.245.7 ± 0.4
CLAP75.3 ± 1.844.3 ± 0.7
AudioMAE79.1 ± 2.762.4 ± 0.7
CardioState-JEPA (Ours)97.9 ± 0.166.8 ± 0.4
Figure 8: Learned delay alignment on paired recordings. The top panel shows ECG to PPG alignment on VitalDB and the bottom panel shows ECG to PCG alignment on EPHNOGRAM. For each ECG R-peak tR, the predicted event time tR+τ^ (red dashed) is overlaid on the target signal alongside the actual event (green), namely the systolic upstroke for PPG and the first heart sound S1 for PCG. The learned delay tracks the physiological offset between modalities.
Figure 8: Learned delay alignment on paired recordings. The top panel shows ECG to PPG alignment on VitalDB and the bottom panel shows ECG to PCG alignment on EPHNOGRAM. For each ECG R-peak tR, the predicted event time tR+τ^ (red dashed) is overlaid on the target signal alongside the actual event (green), namely the systolic upstroke for PPG and the first heart sound S1 for PCG. The learned delay tracks the physiological offset between modalities.
Table 4: Ablation studies on pretraining modalities, self-supervised objectives, and auxiliary losses.
Model / VariantECG Avg ↑PPG Cls Avg ↑PPG Reg Avg ↓PCG Avg ↑
Pretraining Modality Combination
ECG only89.7 ± 0.5
PPG only78.5 ± 2.610.7 ± 0.6
PCG only60.2 ± 0.9
ECG + PPG91.4 ± 0.480.6 ± 2.810.1 ± 0.8
ECG + PCG88.4 ± 0.679.7 ± 0.7
CardioState-JEPA (Ours)90.9 ± 0.380.4 ± 2.59.1 ± 0.682.3 ± 0.2
Self-Supervised Objective
SimCLR89.5 ± 0.570.1 ± 2.911.1 ± 0.680.0 ± 0.7
BYOL89.3 ± 0.466.7 ± 3.811.0 ± 0.580.3 ± 0.4
BarlowTwins88.6 ± 0.464.1 ± 4.411.3 ± 0.578.8 ± 1.0
MAE84.4 ± 0.562.0 ± 3.811.8 ± 0.578.1 ± 0.3
CardioState-JEPA (Ours)90.9 ± 0.380.4 ± 2.59.1 ± 0.682.3 ± 0.2
Auxiliary Loss Terms
w/o Cross-modal prediction83.8 ± 0.563.9 ± 3.510.7 ± 0.578.4 ± 0.4
w/o State alignment84.4 ± 0.566.3 ± 3.511.2 ± 0.579.3 ± 0.5
w/o Delay modeling85.3 ± 0.668.9 ± 3.711.1 ± 0.681.3 ± 0.5
w/o Phase supervision88.7 ± 0.571.0 ± 3.110.4 ± 0.678.3 ± 0.5
CardioState-JEPA (Ours)90.9 ± 0.380.4 ± 2.59.1 ± 0.682.3 ± 0.2
Figure 9: t-SNE of PPG features on the binary atrial fibrillation task (MIMICPerform-AF) from the Stage I (left) and Stage II (right) checkpoints, colored by class. The class silhouette rises from 0.03 to 0.12.
Figure 9: t-SNE of PPG features on the binary atrial fibrillation task (MIMICPerform-AF) from the Stage I (left) and Stage II (right) checkpoints, colored by class. The class silhouette rises from 0.03 to 0.12.
Table 5: Training configurations for pretraining and downstream linear probing. Regression tasks are shown with a dash in the class column and are scored with mean absolute error, while classification tasks are scored with macro-AUROC.
Dataset# ClassesTrainValidTestOpt.EpochsBSLR
ECG Pretraining
MIMIC-IV-ECG [13]710,56078,951AdamW1961.2e-4
PPG Pretraining
PPG-EXT [29]4,611,607512,401AdamW1961.2e-4
PCG Pretraining
BMD-HS [1]3,436382AdamW1961.2e-4
ECG Downstream
PTB-XL Super [40]517,0842,1462,158AdamW100161e-3
PTB-XL Sub2317,0842,1462,158AdamW100161e-3
PTB-XL Form197,197901880AdamW100161e-3
PTB-XL Rhythm1216,8322,1002,098AdamW100161e-3
CPSC 2018 [27]94,9505511,376AdamW100161e-3
CSN [48]3816,5461,8604,620AdamW100161e-3
PPG Downstream
WESAD22,104297597AdamW100161e-3
DaLiA829,2944,6025,320AdamW100161e-3
MIMIC AF23,239240717AdamW100161e-3
PPG Arrhythmia236,8204,7645,243AdamW100161e-3
BIDMC9,4121,6521,398AdamW100161e-3
UQVital22,1582,5298,158AdamW100161e-3
EarSet1,36844364AdamW100161e-3
WildPPG179,49215,00045,000AdamW100161e-3
Sensors BP1,631180250AdamW100161e-3
UCI BP89,05411,28611,411AdamW100161e-3
BCG BP5218664AdamW100161e-3
PCG Downstream
CirCor Murmur [32]31,745608654AdamW100161e-3
CinC [26]22,006627607AdamW100161e-3
Figure 10: t-SNE of PPG features on the six-class arrhythmia task from the Stage I (left) and Stage II (right) checkpoints, colored by class. The class silhouette rises from 0.01 to 0.09.
Figure 10: t-SNE of PPG features on the six-class arrhythmia task from the Stage I (left) and Stage II (right) checkpoints, colored by class. The class silhouette rises from 0.01 to 0.09.
Table 6: Input length ablation on ECG tasks (macro-AUROC ×100, ↑, 10% training data). Bold marks the best value in each column. Mean and Std are the mean and standard deviation across the three input lengths.
PTB-XL Super ↑PTB-XL Sub ↑PTB-XL Form ↑PTB-XL Rhythm ↑
Method2.5s5s10sMeanStd2.5s5s10sMeanStd2.5s5s10sMeanStd2.5s5s10sMean
ST-MEM65.369.569.268.01.958.960.760.059.90.752.458.153.554.72.562.658.762.261.2
ECG-FM80.581.986.583.02.673.372.275.173.51.262.267.876.068.75.772.377.082.077.1
ECGFounder83.386.587.085.61.674.876.076.275.70.662.668.269.366.72.973.581.790.281.8
CardioState-JEPA85.285.686.485.70.578.478.779.178.70.365.668.370.068.01.887.087.392.488.9
Table 7: Sensitivity to the cross-modal loss weight λcross. The CardioState-JEPA row uses the default weighting. Values other than MAE are macro-AUROC ×100, where higher is better, and lower is better for MAE.
VariantECG Avg ↑PPG Cls ↑PPG Reg ↓PCG Avg ↑
CardioState-JEPA (λcross=1.0)90.980.49.182.3
λcross=0.590.978.510.3479.4
λcross=2.090.677.310.5480.9
Table 8: Sensitivity to the delay supervision loss weight λdelay. The CardioState-JEPA row uses the default weighting. Values other than MAE are macro-AUROC ×100, where higher is better, and lower is better for MAE.
VariantECG Avg ↑PPG Cls ↑PPG Reg ↓PCG Avg ↑
CardioState-JEPA (λdelay=1.0)90.980.49.182.3
λdelay=0.590.878.79.6879.3
λdelay=2.090.878.010.2082.5
Table 9: Sensitivity to the state loss weight λstate. The CardioState-JEPA row uses the default weighting. Values other than MAE are macro-AUROC ×100, where higher is better, and lower is better for MAE.
VariantECG Avg ↑PPG Cls ↑PPG Reg ↓PCG Avg ↑
CardioState-JEPA (λstate=0.05)90.980.49.182.3
λstate=0.2590.774.510.2082.7
λstate=1.090.275.89.6582.7

Findings

  • On PPG tasks, average classification AUROC improved from 72.2 to 80.4 versus the strongest baseline, and average regression error (MAE) dropped from 10.9 to 9.1.
  • On ECG, average AUROC across 18 settings (six datasets times three label fractions) reached 84.1, a 15.5-point improvement over the strongest self-supervised baseline, MoCo-v3, at 68.5.
  • On PCG, the model achieved the best linear-probe results on both murmur detection and abnormal heart-sound detection, improving murmur detection by 18.8 AUROC points over the strongest self-supervised baseline.
  • CardioState-JEPA matched or exceeded some ECG models trained with privileged clinical text or supervised labels on several benchmarks, despite using no such extra supervision.
  • t-SNE visualizations showed modality-specific clusters after Stage I that became thoroughly mixed after Stage II (modality silhouette falling from 0.121 to -0.006), while downstream class separability (e.g., for arrhythmia classification) increased in the same transition.

Where it can be used

  • Using PPG signals alone from wearable devices (e.g., smartwatches) to estimate arrhythmia, blood pressure, or respiratory rate by leveraging knowledge transferred from ECG and PCG pretraining.
  • Building automated murmur-screening systems from stethoscope (PCG) recordings that draw on a shared encoder pretrained with ECG and PPG to potentially improve performance with less PCG-specific data.
  • Consolidating cardiac data collected from different sensor types in clinical settings into a single shared representation model instead of maintaining separate models per sensor.
  • Informing pretraining strategy design for clinical settings with limited labeled data, since the approach was tested under low label fractions (1% and 10%) as well as full labels.

Limits and open work

  • Synchronized multi-sensor (paired) recordings are far scarcer than unimodal data, so the delay aligner is trained on relatively few clean, detectable heartbeats.
  • PCG pretraining data is the smallest of the three modalities, which limits how far the acoustic representation could be pushed on its own without cross-modal transfer.
  • For ECG specifically, the trimodal model is within noise of the strongest bimodal variant, so the benefit of adding a third modality was clearest for PPG and PCG rather than ECG.
  • The delay aligner relies on being able to detect a reliable reference event in the source signal; very noisy recordings fall back to unsupervised alignment without physiological anchoring.
  • All reported results use a frozen encoder with linear probing only; full fine-tuning of the encoder and testing larger encoder sizes are left for future work.

Why it matters

Cardiac AI models have historically been built separately for each sensor type, but this work shows that ECG, PPG, and PCG can act as mutual teachers for one another during training. If a single shared representation can serve wearables, clinical ECG monitors, and stethoscope recordings alike, it could reduce the need to build and maintain separate models per device and help in settings where labeled data for any one modality is scarce.

Terms in this paper

  • JEPA (Joint-Embedding Predictive Architecture) · A training method that predicts a hidden ('latent') representation at masked positions instead of reconstructing the raw input
  • AUROC · A score measuring how well a classifier distinguishes positive from negative cases; closer to 1 is better
  • linear probing · An evaluation method where the pretrained encoder is kept frozen and only a small linear classifier on top is trained
  • delay aligner · A learned module that estimates the time offset between two signals of the same heartbeat, since electrical, sound, and pulse signals arrive at different moments
  • momentum encoder · A slowly-updated copy of the main encoder used to produce stable training targets, formed by exponentially averaging the main encoder's weights

Original abstract (English)

Electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG) provide complementary views of the same cardiac cycle, yet existing cardiac foundation models are trained for a single sensing modality, leaving the shared physiology across sensors unexploited. We introduce CardioState-JEPA, a cardiac foundation model to learn a single shared representation jointly across ECG, PPG, and PCG, built on a physiology-aware joint-embedding predictive architecture. The model maps heterogeneous waveforms into a common token space, processes them with a single shared Transformer encoder, and learns by predicting masked latent cardiac states, placing the pretraining target on shared physiology rather than sensor-specific waveform appearance. To handle the temporal offsets between electrical, mechanical, and hemodynamic events, cross-modal prediction uses a learned delay aligner that matches signals at the corresponding cardiac time. Because synchronized multi-sensor recordings are scarce, CardioState-JEPA first learns within-modality structure from abundant unimodal data and then uses paired data to align modalities in latent cardiac time. Evaluated as a frozen encoder across 25 downstream tasks spanning ECG, PPG, and PCG, our encoder improves average PPG classification by 8.2 AUROC points, PCG murmur detection by 18.8 AUROC points, and ECG classification by 15.5 AUROC points over the best self-supervised signal baseline and matches or exceeds cardiac models trained with privileged clinical text or supervised labels on several ECG benchmarks. These results establish that heterogeneous cardiac signals can mutually supervise a single foundation model of cardiac physiology.

Authors · Hamza Shafiq

Read on arXiv

Latest papers

All papers →

Latest from METAL MEDIA

Figures: Hamza Shafiq et al., arXiv:2608.12944, CC BY 4.0