Google Research Launches SensorFM Trained on 1 Trillion Minutes

Google Research's SensorFM, trained on over 1 trillion minutes of wearable data, sets new benchmarks across 35 health-related tasks.

By Central
SensorFM uses a ViT-1D encoder and Adaptive and Inherited Masking to handle missing wearable data.
Highlights
  • SensorFM is pre-trained on over 1 trillion minutes of sensor data from 5 million participants across 100 countries.
  • The model's largest variant, SensorFM-B, outperforms smaller versions on 33 of 35 health tasks.
  • SensorFM's Adaptive and Inherited Masking treats missing data as signal, not noise, improving accuracy.

Google Research has introduced SensorFM, a foundation model for wearable health data pre-trained on more than 1 trillion minutes of sensor readings from 5 million people. The model represents a departure from the conventional approach of building separate predictive models for each health outcome, instead learning a general representation of physiological signals that can be adapted to dozens of tasks with minimal supervision.

What Is SensorFM?

SensorFM is a large sensor foundation model for wearable time-series representation learning. It ingests 34 one-minute aggregate features drawn from five sensors: photoplethysmography (PPG), accelerometer, electrodermal activity (EDA), skin temperature, and altimeter. Those features are organized into seven categories over a 24-hour context window.

The backbone is a ViT-1D encoder trained with a masked-autoencoder objective and a patch size of [20, 1]. Pretraining used 5 million consented participants sampled between September 2024 and September 2025, spanning more than 100 countries, all 50 U.S. states, and over 20 Fitbit and Pixel Watch models. The corpus totals over two billion hours, or more than one trillion minutes.

Four model variants exist, each paired with a proportional data volume during training:

  • XXS (138,740 parameters) — 5K subjects, 2×10⁶ sensor-hours
  • XS (933,204 parameters) — 50K subjects, 2×10⁷ sensor-hours
  • S (7,290,068 parameters) — 500K subjects, 2×10⁸ sensor-hours
  • B (110,763,412 parameters) — 5M subjects, 2×10⁹ sensor-hours

Evaluation uses separate data covering 13,985 subjects across three prospective IRB-approved studies in metabolic, cardiac, respiratory health, sleep, and mental health. The 35 tasks span cardiovascular, metabolic, mental health, sleep, demographics, and lifestyle categories.

Scaling Delivers Measurable Gains

The research team swept four model sizes against four data volumes to test whether scale buys anything measurable. SensorFM-B on the 5M corpus cuts reconstruction validation loss by 31% versus SensorFM-XXS. Generative loss drops 28% on average. Downstream, it gains ΔAUC = 0.09 on classification and Δr = 0.21 on regression. Across variants, B wins 33 of 35 tasks, and XXS ranks last on 33 of 35.

The failure case is equally instructive. SensorFM-B trained on only 5K subjects posts a 1.082 validation loss — worse than every smaller variant at the same data volume. Pretraining was stopped early because the model overfit. This confirms that large capacity without proportional data actively hurts performance.

When data volume scales proportionally to model capacity, the results are unambiguous. Mean ROC AUC moves from 0.664 (XXS) to 0.752 (B), and mean Pearson r moves from 0.386 to 0.612. The trend has not saturated.

AIM: Handling Missing Data as Signal, Not Noise

Real wearable streams fragment during charging, off-wrist periods, and power-saving modes. Conventional methods either impute gaps, injecting bias, or drop windows, discarding data. SensorFM instead uses Adaptive and Inherited Masking (AIM). The applied mask is the union of the inherited missingness mask and the artificial mask. Loss is computed only on artificially masked patches that had ground truth. Two-stage token masking, using token dropout and attention masking, keeps this efficient.

Because the decoder learns to reconstruct ablated observations, imputation and forecasting come for free. Against the best baseline, SensorFM improves random imputation by 74.8% and sensor signal imputation by 83.7%. Key reconstruction MSE results on the held-out test set show SensorFM-B achieving 0.215 on random imputation (80% masked) versus 0.854 for linear interpolation, and 0.468 on temporal interpolation of 60 minutes versus 0.777 for linear interpolation.

Adapting Embeddings for Downstream Tasks

Turning SensorFM representations into predictions follows a straightforward protocol. The encoder stays frozen. Embeddings are aggregated per person using the mean and standard deviation across days, reduced to 50 principal components, and a linear head trains under five-fold person-independent cross-validation.

This linear probe beats a supervised feature-engineered baseline on 34 of 35 tasks. Selected results include age prediction with r = 0.920 versus 0.662 for feature engineering, mental health medication classification with ROC AUC = 0.819 versus 0.773, PHQ-8 with r = 0.450 versus 0.354, and insulin resistance with ROC AUC = 0.761 versus 0.710.

A notable exception is Framingham 30-year risk, where demographic-only models win by construction because the score is calculated from demographic features. The research team reports SensorFM best on 31 of 35 tasks, not all of them. Demographics still help SensorFM on 22 of 30 tasks, though the lift shrinks with scale. In very-low-label regimes, demographic priors alone remain strong.

Agentic Automation of Model Tuning

Even a linear probe needs per-task tuning. To automate that, the research team ran a “classroom” of five LLM student agents spanning Gemini 2.5 Flash through Gemini 3.1 Pro Preview. Agents generated, executed, scored, and refined Python heads over 20 cycles using unreduced embeddings, running a total of 30,516 experiments. Agentaaaa-found heads beat the linear probe on 16 of 20 classification tasks measured by F1 and raised Pearson correlation on 12 of 15 regression tasks. The winning solutions were conservative—almost all reduced the embedding space to 50–100 dimensions, and linear models outnumbered non-linear ones.

Grounding a Personal Health Agent

In the final experiment, Gemini 3 Flash generated health summaries for 31 real participant profiles. Every condition received demographics and feature-engineered daily metrics. Conditions then added SensorFM predictions, ground-truth targets, or nothing. Four board-certified physicians, blinded to condition, produced 1,860 ratings across five rubric dimensions. Adding SensorFM predictions beat the baseline overall (W = 10110, p < 0.001) and on each dimension. Its predictions were statistically indistinguishable from ground truth (p = 0.396).

Practical Use Cases

  • Screening and risk stratification: A frozen encoder plus one linear head flags candidates for confirmatory lab work. The paper scopes this to screening, not diagnosis.
  • Repairing daily summaries: With 60 contiguous minutes ablated, SensorFM retains 99.7% step-count and 99.9% deep-sleep accuracy.
  • Label-scarce studies: Probe frozen embeddings instead of training end-to-end. Compare against a demographics-only baseline first.
  • Grounded coaching: The agent prompt forbids emitting raw regression values or boolean flags. Predictions are interpreted qualitatively instead.

What This Means for Practitioners

SensorFM demonstrates that a single foundation model pretrained at sufficient scale on diverse wearable sensor data can outperform task-specific feature engineering across a broad range of health endpoints. For researchers and developers working with wearable data, the immediate takeaway is that frozen embeddings from SensorFM can serve as a drop-in replacement for hand-engineered features, reducing the need for per-task supervision while improving predictive performance. The model’s ability to handle missing data natively through AIM also eliminates a perennial preprocessing headache. The full paper and model details are available on arXiv, and practitioners should evaluate whether SensorFM’s pretrained representations are available for their own use cases, keeping in mind that the demographic-only baseline remains a strong competitor in low-label regimes and for scores directly derived from demographic variables.

Share This Article