Tutorial 01
Convert raw accelerometry to NPY
Turn one raw wrist recording into quality-controlled daily arrays ready for Sensori.
What you need
Sensori requires Python 3.13. Clone the repository, create an isolated environment and install the package:
git clone https://github.com/OxWearables/Sensori.git
cd Sensori
conda create -n sensori python=3.13 pip
conda activate sensori
pip install -e .
The script accepts .cwa, .gt3x, .bin and .csv files, with optional
.gz compression. A CSV must contain time,x,y,z: timestamps must be unique
and increasing, and XYZ acceleration must be finite and expressed in g.
- Note: CWA, GT3X and BIN files need Java 8 or newer, while CSV files do not need Java.
Convert one recording
From the repository root, run the command that matches your input. For a device file:
python scripts/get_npy.py \
--file raw/participant_001.cwa.gz \
--output processed
The CSV sampling rate is inferred from its timestamps and must be greater than
10 Hz. If you know the nominal rate, you can supply it explicitly with, for
example, --input-sample-rate 100.
The input filename becomes the output-directory name by default:
processed/
└── participant_001/
├── day_0.npy
├── day_1.npy
├── info.json
└── wear_duration.csv
If you have no recording to hand, tutorials/nhanes_preprocessing.ipynb runs
this same parser on public data: it downloads NHANES participant archives from
the CDC, combines their hourly files into one CSV and writes the same
day_*.npy layout, one call per participant.
What the script does
The preprocessing path is fixed for compatibility with the released model:
gravity calibration → 5 Hz low-pass filter → 10 Hz resampling
→ non-wear detection → calendar-day segmentation → quality control → NPY
A day is retained only when it contains a complete, finite 24-hour signal and at least 22 hours of wear. The recording must also pass reading, calibration and filtering checks, contain fewer than 10 interruptions, have mean ENMO no greater than 200 mg, and contain at least one eligible day. Non-wear is used for quality control; the retained signal is not replaced with zeros or missing values.
Eligible days are written chronologically as float32 XYZ acceleration in g.
Every array has shape (2880, 300, 3): 2,880 consecutive 30-second windows,
300 samples per window at 10 Hz, and three axes.
Check the result
Check one output before starting inference:
import numpy as np
day = np.load("processed/participant_001/day_0.npy", allow_pickle=False)
assert day.shape == (2880, 300, 3)
assert day.dtype == np.float32
assert np.isfinite(day).all()
Things to watch out for:
- The script processes one recording per command. Run it again for each input.
- Partial calendar days are skipped.
day_0is not a date; usewear_duration.csvto map output files to dates and inspect exclusion reasons. - Exit code
3means preprocessing completed but strict quality control found no eligible day. Inspectinfo.jsonfor the failed criteria. - A matching completed run is skipped. Use
--overwriteonly after inspecting the target; it replaces the parser-generated metadata andday_*.npyfiles. - Do not change the scientific constants when preparing data for the released model.