Upload README.md with huggingface_hub
Browse files
README.md
ADDED
|
@@ -0,0 +1,77 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: mit
|
| 3 |
+
tags:
|
| 4 |
+
- onnx
|
| 5 |
+
- skinning
|
| 6 |
+
- skin-weights
|
| 7 |
+
- rigging
|
| 8 |
+
- 3d
|
| 9 |
+
- qtmesheditor
|
| 10 |
+
library_name: onnx
|
| 11 |
+
pipeline_tag: other
|
| 12 |
+
---
|
| 13 |
+
|
| 14 |
+
# SkinTokens / TokenRig — ONNX export
|
| 15 |
+
|
| 16 |
+
**ONNX re-export of [VAST-AI SkinTokens / TokenRig](https://github.com/VAST-AI-Research)**
|
| 17 |
+
(MIT code + MIT weights, Qwen3-0.6B backbone, trained on Articulation-XL2.0 /
|
| 18 |
+
CC-BY-4.0) — autoregressive ML **skin-weight prediction**: given a mesh and a
|
| 19 |
+
skeleton, it predicts per-vertex bone weights. All credit for the original
|
| 20 |
+
weights goes to VAST-AI-Research.
|
| 21 |
+
|
| 22 |
+
Exported for **[QtMeshEditor](https://github.com/fernandotonon/QtMeshEditor)**
|
| 23 |
+
(issue #819), where it is the **default skinner** (`qtmesh skin`, the GUI
|
| 24 |
+
"Compute Skin Weights" dialog, the rig→skin chain, and the
|
| 25 |
+
`compute_skin_weights` MCP tool), running locally via ONNX Runtime with a
|
| 26 |
+
geodesic-voxel fallback.
|
| 27 |
+
|
| 28 |
+
> The files QtMeshEditor downloads at runtime live in the shared
|
| 29 |
+
> [`fernandotonon/QtMeshEditor-models`](https://huggingface.co/fernandotonon/QtMeshEditor-models)
|
| 30 |
+
> repo under `skintokens/`. This repo is the standalone model card + mirror
|
| 31 |
+
> for people who want the converted weights themselves.
|
| 32 |
+
|
| 33 |
+
## Files (five graphs + manifest)
|
| 34 |
+
|
| 35 |
+
| file | role |
|
| 36 |
+
|---|---|
|
| 37 |
+
| `mesh_cond.onnx` | Michelangelo point-cloud encoder → LLM mesh-conditioning prefix |
|
| 38 |
+
| `vae_cond.onnx` | skin-CVAE conditioning encoder over the sampled points |
|
| 39 |
+
| `embed.onnx` | token id → LLM embedding |
|
| 40 |
+
| `decoder.onnx` + `decoder.onnx.data` | Qwen3-0.6B causal-LM **KV-cache step** (external weights — ONNX Runtime 1.20.1 segfaults parsing a >1.6 GB single-file proto) |
|
| 41 |
+
| `skin_decode.onnx` | FSQ skin tokens → per-joint, per-sampled-point weights (FSQ folded in) |
|
| 42 |
+
| `skintokens.json` | manifest: every config value the host needs (below) |
|
| 43 |
+
|
| 44 |
+
## Inference contract
|
| 45 |
+
|
| 46 |
+
1. Surface-sample `num_points` (8192) points + normals; normalise mesh+joints
|
| 47 |
+
per upstream `AugmentAffine` (joints **included** in the AABB, exact
|
| 48 |
+
`[-1,1]` fit).
|
| 49 |
+
2. Tokenize the skeleton **teacher-forced** (DFS order; per bone
|
| 50 |
+
`[branch?, parent-joint xyz, joint xyz]` discretised to 256 bins over
|
| 51 |
+
`[-1,1]`; `bos=257`, cls `"articulation"=266`, stream ends with the
|
| 52 |
+
switch token `eos=258`). Multi-root topologies must be re-parented to the
|
| 53 |
+
first root + DFS-reordered first.
|
| 54 |
+
3. Prefix = mesh-cond embeddings + skeleton token embeddings → autoregressive
|
| 55 |
+
greedy decode of `J × tokens_per_skin (4)` skin tokens, constrained to the
|
| 56 |
+
FSQ range `[267, 33035)`; global EOS `33035`; full vocab `33036`.
|
| 57 |
+
4. Per joint: `skin_decode` on its 4 FSQ ids (− 267 offset) → weights over the
|
| 58 |
+
8192 sampled points; transfer to full-res vertices by 8-NN inverse-distance.
|
| 59 |
+
|
| 60 |
+
LLM dims (manifest): hidden 896, 28 layers, 8 KV heads, head_dim 128;
|
| 61 |
+
`tokens_skin_cond=384`, CVAE latent 512, FSQ codebook 32768.
|
| 62 |
+
|
| 63 |
+
Raw predictions are deliberately diffuse — the upstream demo voxel-masks them
|
| 64 |
+
by default. QtMeshEditor applies a geodesic-localisation pass (filter to
|
| 65 |
+
geodesically-local bone sets + renormalise); consumers are advised to do the
|
| 66 |
+
same (bleed 0.74 → 0.05 in our measurements).
|
| 67 |
+
|
| 68 |
+
## Reproducing
|
| 69 |
+
|
| 70 |
+
`scripts/export-skintokens-onnx.py` in the QtMeshEditor repo (one-time,
|
| 71 |
+
offline; bf16→fp32, forced eager attention, decomposed RMSNorm for opset 18,
|
| 72 |
+
trace-friendly FPS). Parity vs PyTorch ≈ 1e-5 on every graph.
|
| 73 |
+
|
| 74 |
+
## License
|
| 75 |
+
|
| 76 |
+
MIT (same as the upstream code and weights). Training data: Articulation-XL2.0
|
| 77 |
+
(CC-BY-4.0) — credit VAST-AI-Research.
|