File size: 17,286 Bytes
2dd8e79 74b6bc5 2dd8e79 0782212 2dd8e79 74b6bc5 2dd8e79 59d9957 2dd8e79 0782212 2dd8e79 0782212 2dd8e79 0782212 2dd8e79 0782212 2dd8e79 0782212 2dd8e79 0782212 2dd8e79 0782212 2dd8e79 0782212 2dd8e79 0782212 2dd8e79 7265b4e f73be5f b304465 f73be5f 2dd8e79 0782212 2dd8e79 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 | ---
license: cc-by-nc-4.0
language:
- en
pipeline_tag: feature-extraction
tags:
- embeddings
- imu
- inertial
- motion
- accelerometer
- sensor
- robotics
- human-activity-recognition
- qwen3-vl
base_model: EximiusLabs/fusion-embedding-2-2b-preview
---
# fusion-embedding-2-tremor
<p align="center">
<img src="assets/tremor-banner.png" alt="Tremor β the motion sense for Fusion Embedding 2 (2B-Preview), Eximius Labs" width="100%">
</p>
<div align="center">
[](https://github.com/Eximius-Labs/fusion-embedding)
[](https://github.com/Eximius-Labs/fusion-embedding)
[](#license)
[](#)
[](https://github.com/Eximius-Labs/fusion-embedding)
</div>
**Tremor** is the inertial (IMU) sensor pack for Eximius Labs' fusion-embedding stack.
It maps a short window of body-worn motion (a 3-axis accelerometer stream) into the
[Qwen3-VL-Embedding-2B](https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B) text
embedding space, so a robot's or wearable's motion history becomes searchable in plain
language: "find when it was walking", "find when it picked something up".
Tremor is part of the **fusion-embedding family** and plugs into the same shared space as
[fusion-embedding-2](https://huggingface.co/EximiusLabs/fusion-embedding-2-2b-preview). It
targets the frozen Qwen3-VL-Embedding-2B text embeddings that the family is built on, and
because fusion-embedding-2's text, image, and video paths are byte-identical to that base,
Tremor's motion embeddings are **directly comparable to that model's audio, text, image, and
video** in one 2048-d space. Motion becomes a first-class, language-addressable modality
alongside the rest.
[GitHub](https://github.com/Eximius-Labs/fusion-embedding) | [fusion-embedding-2](https://huggingface.co/EximiusLabs/fusion-embedding-2-2b-preview) | Technical report: [arXiv:2607.18666](https://arxiv.org/abs/2607.18666)
**The family.** Each sense is a separately loadable pack over the same frozen base: [Tactus](https://huggingface.co/EximiusLabs/fusion-embedding-2-tactus) reads touch from a 32x32 pressure glove, [Tactus Mat](https://huggingface.co/EximiusLabs/fusion-embedding-2-tactus-mat) reads a 64x32 body pressure mat, [Ember](https://huggingface.co/EximiusLabs/fusion-embedding-2-ember) reads heat, and [Tremor](https://huggingface.co/EximiusLabs/fusion-embedding-2-tremor) reads motion, with a [Unitree-G1 head](https://huggingface.co/EximiusLabs/fusion-embedding-2-tremor-g1). Because the base is never modified, adding a sense costs a small trained head and an afternoon of compute rather than a new foundation model.
## Model Overview
<p align="center">
<img src="assets/tremor_model_overview.png" alt="Tremor architecture: a frozen UniMTS encoder feeds a trained projector into the fusion-embedding shared space, where motion becomes searchable in natural language alongside every other modality" width="880px">
</p>
Tremor is a **frozen [UniMTS](https://github.com/xiyuanzh/UniMTS) motion encoder plus a
small trained projector**. A single accelerometer window is placed at one joint of
UniMTS's orientation-invariant skeleton-graph encoder (ST-GCN), which produces a 512-d
motion feature; a projector of roughly 2.6M parameters maps that into the frozen Qwen
base's 2048-d text space. Nothing is trained except the projector. Motion and text
embeddings are L2-normalized and compared by cosine similarity.
The design choice that matters is the encoder: rather than train an IMU encoder from
scratch, Tremor reuses a frozen, orientation-invariant pretrained one and learns only
the projection into the language space. The pretraining prior, not our training, is
what carries cross-subject and cross-mounting generalization.
| Feature | Value |
| --- | --- |
| Parameters | ~2.06B frozen Qwen base + frozen UniMTS encoder; **2.6M trained** |
| Modality | inertial motion (3-axis accelerometer) |
| Supported tasks | `zero-shot activity retrieval`, `zero-shot activity classification` |
| Input | one accelerometer window `[3, T]` in m/s^2; pass the sensor's sample rate to resample to a 200-sample, 20 Hz (10 s) window |
| Sensor placement | single IMU, mapped to one skeleton joint (robots/wearables have a known, fixed mount) |
| Embedding dimension | 2048 |
| Pooling strategy | last-token pooling (text side) |
| Base model | Qwen/Qwen3-VL-Embedding-2B (frozen) |
| Motion encoder | UniMTS ST-GCN, accelerometer-only (frozen) |
| Trained components | projector 2.6M `LayerNorm(512) -> Linear(512,1024) -> GELU -> Linear(1024,2048)` |
| Distribution | ~10 MB trained projector; the frozen encoder and base download from their own repos |
## See it in action
**A full activity sequence, read live.** A **Unitree G1** performs each activity in simulation β sitting, walking, running, climbing β while a window slides across a real, held-out body-worn accelerometer stream ([RealWorld](https://sensor.informatik.uni-mannheim.de/), out of training) and Tremor's scores update in real time, tracking the motion as it changes. The robot shows each activity full-body; the signal and the predictions are real, frame-synced data.
<p align="center">
<img src="assets/tremor_robot_stream.gif" alt="A Unitree G1 performs sitting, walking, running and climbing full-body while Tremor's live scores track the activity on a real held-out accelerometer stream" width="860px">
</p>
**On the robot's own IMU and camera.** On a real **Unitree G1 humanoid** ([Humanoid-Everyday](https://usc-psi-lab.github.io/humanoid-everyday/)), Tremor reads the robot's own IMU as language, live and frame-synced to its **egocentric camera** β zero simulation. As it switches from adjusting a phone stand to walking to a door, its IMU signature changes from flat to dynamic and Tremor's read tracks the activity in real time.
<p align="center">
<img src="assets/tremor_robot_demo.gif" alt="Live demo on a real Unitree G1: the robot's own egocentric camera with its frame-synced IMU (zero simulation); Tremor's read tracks the activity, from adjusting a phone stand to walking to a door" width="840px">
</p>
**Any body-worn sensor.** No robot is required β the same works on a wristband or wearable. Here the raw accelerometer stream slides by directly and Tremor's scores track sitting β walking β running β climbing on the held-out RealWorld set.
<p align="center">
<img src="assets/tremor_demo.gif" alt="Live classification on a held-out human-activity stream: a window slides across the accelerometer signal and Tremor's activity scores update in real time" width="720px">
</p>
**Text β motion retrieval.** Query an activity in plain language and Tremor returns the nearest motion windows. Below, on the RealWorld set held out of training entirely (zero-shot), each thumbnail is a retrieved window's 3-axis accelerometer trace; correct retrievals (the window's true activity matches the query) are outlined in green.
<p align="center">
<img src="assets/tremor_retrieval_gallery.png" alt="Text to motion retrieval on held-out RealWorld: each query returns its top-5 nearest accelerometer windows, correct retrievals outlined green" width="840px">
</p>
## Training and Evaluation
Tremor is trained with InfoNCE against the frozen base's native chat-template text
embeddings, on a **diverse 16-source pool of real inertial data**: 13 standard
human-activity-recognition datasets, [Ego4D](https://ego4d-data.org/) (egocentric
human), Humanoid-Everyday (Unitree robot), and
[Capture24](https://ora.ox.ac.uk/objects/uuid:99d7c092-d865-4a19-b096-cc16440cd001)
(151 subjects of free-living wrist accelerometry). Only the projector learns.

The single most important lesson of the project is the **evaluation metric**. Instance
retrieval (finding the exact paired window) is misleading for this task and makes a
working activity recognizer look dead. The right metric is **k-way activity
classification on held-out subjects and held-out datasets**, and all numbers below use
it (chance = 1/k).
### Diversity is what makes the base general
Trained on one source (Ego4D alone), an IMU encoder generalizes poorly to sensor
streams that look nothing like it. Tremor's headline result is **zero-shot transfer to
datasets held out of training entirely**. On five held-out datasets, a single-source
baseline (the identical architecture trained on Ego4D only) sits at or below chance,
while Tremor generalizes.
<p align="center">
<img src="assets/tremor_v01_results.png" alt="Trained on 16 IMU sources, Tremor generalizes zero-shot to datasets held out of training entirely, far exceeding a single-source baseline that sits at or below chance" width="820px">
</p>
| Held-out dataset (5-way, zero-shot) | Single-source baseline | **Tremor v0.1** |
| --- | --- | --- |
| RealWorld | 0.16 | **0.65** |
| USC-HAD | 0.13 | **0.58** |
| DSADS | 0.29 | **0.51** |
| UTD-MHAD | 0.21 | **0.27** |
| **held-out mean** | **0.20** (chance) | **0.502** |
These are the single-sensor numbers the shipped `inference.py` produces: one accelerometer
window in m/s^2, resampled from the sensor's rate to 20 Hz and placed at one fixed skeleton
joint. That default joint is selected on in-domain held-out data disjoint from the four
zero-shot sets above, so these numbers are not tuned on the benchmark. Feeding the model the
sensor's true sample rate and unit convention is what recovers this accuracy; a rate-agnostic
window scores well below it. (A research configuration that uses each dataset's own
multi-sensor body placement reaches 0.545 mean, but that requires per-dataset sensor-mount
knowledge and is not available to a general single-IMU input.)
On the two in-domain holds (held-out subjects/episodes of sources seen in training),
Tremor reaches 0.40 5-way on both Ego4D and the Humanoid-Everyday robot set. The
held-out datasets above are excluded from training entirely, so those numbers are
genuine cross-dataset generalization.
## Usage
<details>
<summary>Requirements</summary>
- `torch` (CUDA recommended), `numpy`, `scipy`, `transformers>=4.46`, `huggingface_hub`
- The frozen UniMTS encoder code and weights (Apache-2.0 code):
```bash
git clone https://github.com/xiyuanzh/UniMTS # provides model.py::ST_GCN_18
```
Point `UNIMTS_REPO` (or `unimts_repo=`) at the clone. The UniMTS weights download
automatically from the hub.
- The frozen Qwen3-VL-Embedding-2B base downloads from its original repository.
</details>
<details open>
<summary>via <code>inference.py</code> (this repository)</summary>
```python
import numpy as np
from inference import TremorEmbedder
tr = TremorEmbedder.from_pretrained(
"EximiusLabs/fusion-embedding-2-tremor",
revision="v0.1-preview",
unimts_repo="UniMTS", # path to your git clone of xiyuanzh/UniMTS
)
# accel: a 3-axis accelerometer window in m/s^2, shape [3, T] (any length)
accel = np.load("window.npy") # e.g. [3, 300]
# rank candidate activities; pass your sensor's true sample rate (Hz) so the window is
# resampled to the model's 20 Hz. Pass unit="g" if your sensor reports g rather than m/s^2.
for text, score in tr.rank(accel, ["walking", "sitting", "running", "climbing stairs"],
sample_rate_hz=50):
print(f"{score:+.3f} {text}")
# or embed motion and text into the shared space directly
m = tr.embed_motion(accel, sample_rate_hz=50) # 2048-d, L2-normalized
t = tr.embed_text(["walking", "sitting"]) # 2048-d each
```
Give the model the sensor's true sample rate and an accelerometer in m/s^2 (gravity ~9.8);
these are what the held-out numbers above are measured with. Embedding quality is also
sensitive to the base's chat-template formatting; `inference.py` applies it for you, so match
candidate activities against motion through this API rather than embedding text yourself.
</details>
## Deploy on RunPod
[](https://www.runpod.io/console/hub/Eximius-Labs/tremor)
One-click deploy the endpoint from the
[RunPod Hub](https://www.runpod.io/console/hub/Eximius-Labs/tremor) (serverless,
scales to zero when idle). Once it is running, call it:
```bash
curl -s https://api.runpod.ai/v2/<ENDPOINT_ID>/runsync \
-H "Authorization: Bearer $RUNPOD_API_KEY" \
-H "Content-Type: application/json" \
-d '{"input": {"motion": [[...], [...], [...]]}}'
```
`motion` is a 3-axis accelerometer window shaped `[3, T]`. Use `text` instead to
embed an activity query. Returns 2048-d vectors, so motion and text land in one
space. The Unitree G1 head deploys separately at
[the G1 hub page](https://www.runpod.io/console/hub/Eximius-Labs/tremor-g1).
## Related models
Tremor comes in two tiers, both built on
[fusion-embedding-2](https://huggingface.co/EximiusLabs/fusion-embedding-2-2b-preview):
| Model | Best for | Accuracy |
| --- | --- | --- |
| **This model β general base** | any sensor / platform, unseen datasets (zero-shot) | **0.502** held-out 5-way (unseen datasets) |
| **[fusion-embedding-2-tremor-g1](https://huggingface.co/EximiusLabs/fusion-embedding-2-tremor-g1)** | deployment on the **Unitree humanoid** (G1/H1) | **0.740** in-domain 5-way |
Start with this **general base** β it is the tier that generalizes to sensors it has never seen. Move
to the per-fleet **Tremor-G1** head only when you deploy on that platform and want maximum in-domain
accuracy. Specializing a head raises in-domain accuracy but does not improve cross-dataset transfer;
the two do different jobs. Deployment guidance (sensor contract, streaming, on-device inference,
per-fleet fine-tuning) is in the Tremor-G1 card.
The same shared space also has a touch sense:
[fusion-embedding-2-tactus](https://huggingface.co/EximiusLabs/fusion-embedding-2-tactus) embeds
32x32 pressure/taxel arrays (FSR gloves, e-skins, robot hands) and matches to exceeds the STAG
(Nature 2019) supervised baseline while remaining open-vocabulary.
## Engram
This pack is one of the modalities [Engram](https://github.com/Eximius-Labs/engram) searches. Engram is
the open cross-modal memory layer for physical AI: it indexes a robot's video, audio, and motion into
one embedding space and answers questions about it in plain language, including temporal reasoning that
retrieval alone cannot.
```bash
pip install engram-robomem
```
Repo: https://github.com/Eximius-Labs/engram · PyPI: https://pypi.org/project/engram-robomem · Playground: https://www.eximiuslabs.com/playground
## License
The trained projector weights in this repository are released under
**[CC-BY-NC-4.0](https://creativecommons.org/licenses/by-nc/4.0/)** (non-commercial).
This posture is deliberate and reflects the model's lineage. Tremor's frozen encoder is
[UniMTS](https://github.com/xiyuanzh/UniMTS), which was pretrained on synthetic IMU
derived from the [AMASS](https://amass.is.tue.mpg.de/) / HumanML3D motion corpora. AMASS
is released under a non-commercial research license, so we ship Tremor non-commercial to
respect that lineage. A commercially-clean encoder (re-pretrained on permissively
licensed motion) is future work; a commercial license may follow.
## Limitations
- **Accelerometer only.** The released UniMTS encoder is accelerometer-only; gyroscope
is not used. Rotational-rate information is not currently modeled.
- **Cross-dataset ceiling.** Zero-shot cross-dataset accuracy plateaus around 0.50
(5-way). This is a research preview, not a production activity classifier. The right
use is language-addressable motion search, not high-stakes classification.
- **Single-sensor, single-joint.** One IMU mapped to one skeleton joint, matching a
fixed-mount deployment (a robot torso, a wristband). Multi-sensor fusion is not
modeled.
- **English text only**, and the base's chat-template formatting must be used.
- **Specialization vs. generalization.** Fine-tuning on a specific platform's data
raises accuracy on that platform substantially but does not improve cross-dataset
transfer. Tremor is the general base; per-fleet fine-tuning is a separate recipe.
## Citation
If you use Tremor, please cite this repository and acknowledge the frozen components it
builds on:
```bibtex
@misc{tremor2026,
title = {Tremor: an inertial-motion sensor pack for the fusion-embedding space},
author = {Tonmoy, Abdul Basit},
year = {2026},
note = {Eximius Labs. Model weights CC-BY-NC-4.0.},
url = {https://huggingface.co/EximiusLabs/fusion-embedding-2-tremor}
}
```
Tremor's frozen motion encoder is **UniMTS** (Zhang et al., *UniMTS: Unified Pre-training
for Motion Time Series*, NeurIPS 2024), pretrained on synthetic IMU derived from
**AMASS** (Mahmood et al., ICCV 2019). The text space is **Qwen3-VL-Embedding-2B**.
Please cite those works as well.
|