Humaneness Voice Small β 557 Rank-1 specialist LoRAs
This repository contains 557 separate, completed Rank-1 LoRA adapters for the same Humaneness Voice Small S3 base model: 500 voice identities, 40 emotion subsets, and 17 selected VoiceNet attribute subsets. It also contains a packed index, training/validation statistics, independent analysis results, a loader, a CUDA inference example, and training/provenance code. This is a research release; adapter training completion is not evidence that every emotion or speaking style is reliably controllable.
The base is the original-learning-rate S3 checkpoint at step 32,337. Its semantic model uses Qwen3-0.6B pretrained language-model initialization (not random initialization), with a separate approximately 112.76M-parameter hierarchical Talker and 12 MOSS Audio Tokenizer v2 codebooks per 80 ms frame. The adapters modify only the semantic transformer's 28 layers of q/k/v/o self-attention projections. The Talker, other base parameters, and codec remain frozen. The base S3 weight SHA-256 is 60079723bf797a81107def1e6fc479e5353b3cc51f3780f98ea65a9f5c109f97.
These adapters are distinct from the six Humaneness Voice Small DPO LoRAs, which use ranks 16/64/128 and also target Talker modules. They are also distinct from historical SFT-3 adapters trained on a much larger MOSS base.
What's here
| Family | Count | Typical purpose | Rank / alpha | Trainable values per adapter |
|---|---|---|---|---|
profile |
500 | Reproduce one selected voice profile without a reference recording | 1 / 1 | 286,720 |
emotion |
40 | Adapt toward one high-score emotion subset | 1 / 1 | 286,720 |
vn |
17 | Adapt toward one selected high/low VoiceNet attribute tail | 1 / 1 | 286,720 |
Every adapter was trained for two passes over its selected data, peak learning rate 1e-4, 5% warmup, cosine decay to 10% of peak, and no reference audio in the specialist training prompts. The source text prompts use GENERAL: / SCRIPT: with preserved inline delivery cues, pauses, vocal bursts, and segment durations. The 500 profile adapters use 1,000 selected training rows each; the selection is not the same stringent speaker-similarity threshold for every profile. Read training/TRAINING_LOG.md before drawing identity-fidelity conclusions.
Two profile entries (anime_000, mediathek_0047) reuse the separately evaluated Rank-1 q/k/v/o pilot checkpoints. They are included in the 500 count, not added twice. The extra Rank-1 q/v and Rank-2 pilot variants are not part of this 557-adapter release.
The 17 VoiceNet subset names are in index.csv. They are not all 57 VoiceNet dimensions, and not all are delivery styles: EXPL_high concerns content explicitness, and VFLX_high concerns tempo shift. An adapter name identifies its training subset, not a guarantee of clean one-axis control.
Packed format and individual loading
To avoid hundreds of loose files, each family is an uncompressed TAR container of individual safetensors members:
| File | Members | Contents |
|---|---|---|
adapters/profile.tar |
500 | Voice identity specialists |
adapters/emotion.tar |
40 | Emotion specialists |
adapters/vn.tar |
17 | VoiceNet attribute specialists |
index.json gives every adapter's ID, archive/member name, rank, alpha, tensor count contract, training rows, best validation step and loss, source checkpoint hash, and exported member hash. manifest.json records archive and statistics hashes. Values are stored in FP32 with exact tensor parity to the training checkpoint's adapter portion; no optimizer states or source audio are in the archives. The TAR member is read in memory and never extracted to thousands of files.
Example IDs: profile/anime_000, emotion/Affection, vn/S_RANT_high.
The exact base model supplies the architecture, tokenizer assets, full S3 weight, MOSS v2 codec setup, and baseline inference code. On a CUDA machine:
hf download laion/Humaneness-Voice-Small --local-dir ./Humaneness-Voice-Small
hf download laion/Humaneness-Voice-Small-Rank1-LoRAs --local-dir ./Humaneness-Voice-Small-Rank1-LoRAs
cd Humaneness-Voice-Small-Rank1-LoRAs/code
python infer_rank1.py \
--base-root ../../Humaneness-Voice-Small \
--adapter-root .. \
--adapter-id emotion/Affection \
--prompt $'CAPTION: warmly affectionate, natural speech\nTRANSCRIPT: "I am so glad you are here."' \
--text 'I am so glad you are here.' --frames 55 --language en \
--seed 777 --output affectionate.wav
code/load_rank1.py exposes attach_rank1(model, repository, adapter_id, scale=1.0) for a caller that already built and loaded the S3 base. Attach to a fresh model; do not stack multiple adapters unless you have independently evaluated that mixture. scale is an experimental dose, not a documented quality improvement. code/infer_rank1.py is the full single-clip example. Its prompt, reference-audio, frame-budget and codec behavior follow the base model's inference code.
Evidence and limitations
| Question | Current evidence | What it does not establish |
|---|---|---|
| Did all adapters finish training? | 557/557 have completed runs, best-checkpoint validation records, and finite losses in stats/specialist_status.json. |
Training loss is not an audio-quality or controllability score. Losses across different data subsets should not be ranked as if they were the same task. |
| Can a tiny adapter hold a voice? | In a matched two-voice pilot, Rank-1 q/k/v/o reached Orange speaker cosine 0.921 for anime_000 and 0.968 for mediathek_0047, versus 0.088 and 0.018 for unadapted S3. See stats/identity_rank1_rank2_pilot.json. |
This is not a 500-voice, human-preference or consent evaluation. Orange similarity is a proxy, not proof of identity. |
| Do emotion and VoiceNet LoRAs improve their intended attribute? | Their training and validation logs are released. A matched-seed, multiple-dose, audio-level study is planned/running separately. | No broad efficacy claim yet. A lower NLL on selected high-score data does not show that the wanted attribute is promptably realized without side effects. |
| Can the bank be compressed by one shared PCA/autoencoder? | On held-out adapter identities, the best tested shared AE (64 latent dimensions) had relative update RMSE 0.9839; global PCA/full training rank stayed near 0.96β0.99. See stats/compression/. |
1.0 means roughly omitting the target update entirely. High training explained variance does not mean a new adapter is reconstructed. A text-to-LoRA generator has not been trained. |
| Can unimportant modules be omitted? | A 557-adapter functional study found that absolute-Taylor selection at 84 of 112 modules adds mean 0.0066 NLL/frame; a 20-voice audio canary averaged Orange 0.7146 vs 0.7516 for full adapters. See stats/pruning/. |
This is an inference mask, not a released sparse weight archive or a universal no-loss guarantee. Some voices regress; one clip per voice is insufficient for a blanket recommendation. |
| Can other voice LoRAs be blended to approximate a new voice? | Orange-guided sparse search was measured over 500 targets in embedding space, with only a six-target generated-audio canary. See stats/composition/. |
Embedding fit is not equivalent to audible voice quality or identity. More WER/listening checks are required. |
All raw per-adapter training records and validation events are in stats/specialist_status.json. All 23,468 logged train/validation events across the 557 completed runs are additionally packed in stats/training_events.parquet, including the original event JSON and normalized columns. Additional per-adapter module norms, concentration, importance and held-out NLL records are in stats/pruning/*.parquet; the public can inspect exact adapter-level results rather than rely only on the summary table. The composition folder includes all 500-target coefficient fits, both balanced PCA reports and the generated-audio score records. The compression folder includes PCA and all seven autoencoder result JSON files. The stats/context_dpo_study.json file is a separate DPO comparison included only to contextualize why the specialist family should not be conflated with preference-tuned LoRAs.
No original training audio is republished here. The public Humaneness Voice Small model card links the underlying ladder, sidecars, prompt guide and evaluation spaces. The DPO adapter repository provides separate preference-tuning experiments and comparison audio. A matched 57-adapter emotion/style evaluation will be added to this card once it is measured and validated; do not infer its outcome from the existing training losses.
Reproducibility and intended use
training/ contains the specialist training, source-selection and run log. The local production jobs used packed Parquet inputs and atomic adapter-plus-optimizer checkpoints; the public package contains only best adapter weights and metadata. The index preserves source-manifest/checkpoint hashes and exact best-step provenance. Decode with the MOSS Audio Tokenizer v2 setup from the base repository. Prompting follows the base model's GENERAL:/SCRIPT:, CAPTION:/TRANSCRIPT:, and transcript-only modes; a Rank-1 specialist does not replace correct text conditioning.
This is research software for voice-acting and controlled speech synthesis. Use voice identity adapters only with appropriate rights and consent; an embedding score does not verify speaker consent. Speech may reproduce dataset artifacts, accents and sensitive traits. Use human review before deployment. This repository is released under CC BY 4.0; see the base model and source dataset cards for attribution and upstream conditions.
Model tree for laion/Humaneness-Voice-Small-Rank1-LoRAs
Base model
laion/Humaneness-Voice-Small