Kokoro-82M β Core ML, segmented, Neural Engine
Kokoro-82M converted to Core ML as
five chained .mlpackages, for on-device synthesis on iOS. Built for
OpenReader and published so
that the exact weights the app asserts against are pinned somewhere stable.
config.json
albert.mlmodelc 512-token ALBERT
prosody.mlmodelc durations, F0 and noise curves
text_encoder.mlmodelc
source.mlmodelc harmonic excitation (float32)
decoder.mlmodelc ISTFT-Net vocoder (float16)
voices/*.npy 28 English voice packs, 510 style rows each
*.mlpackage the uncompiled sources the above were built from
What an app should fetch is the .mlmodelc. Compiling an .mlpackage on the
device costs minutes and produces exactly what xcrun coremlcompiler compile
produces on a build machine in seconds, so the compiled form is published and
the packages are kept only as what a future re-export starts from.
Inputs and outputs follow the segmentation in kokoro-swift (Apache 2.0), from which the conversion descends, with four deliberate differences.
ALBERT is exported at 512 tokens
Upstream's converter takes one --max-tokens defaulting to 512, and the
published ALBERT reads as an artefact of a different run than its siblings: it
declares [1, 64] where prosody, text_encoder and decoder declare
[1, 512]. All the shapes are fixed, so the chain was capped by the smallest β
62 phonemes, about ten words, per forward pass. Kokoro predicts prosody over
the whole sequence, so a cut every ten words restarts the intonation
mid-sentence. This is ALBERT re-exported at 512 from the same
kokoro-v1_0.pth weights; on text short enough for both, the two produce
bitwise-identical audio.
The recurrent layers are real lstm ops
Upstream's exporter runs every bidirectional LSTM through a helper that loops
over the time axis in Python, twice, freezing the state on padded steps with a
torch.where per step. Traced, that is what it becomes: straight-line MIL, 512
steps by two directions by five gate operations, per layer. Prosody has five
such layers, and its model.mil came to 16.9 MB against the decoder's 0.37 MB
despite holding a third as many weights.
Core ML never has to see that. It has an lstm op that stays one node whatever
the sequence length, and the only thing standing between the exporter and it is
padding β a bidirectional nn.LSTM over a padded buffer is correct forwards
and wrong backwards, because the reverse pass starts at the end of the buffer
and accumulates state across the padding before it reaches real data.
So the reverse direction is run as a forward one. Gather the sequence through
an index map that reverses the first n positions and leaves the padding
where it is, run a unidirectional LSTM over that, and gather back through the
same map β the map is an involution, so the same indices undo it. Position t
then carries the state of a reverse pass that started at n-1, which is what
the loop computed, and the padding is masked to zero as before.
That is exact rather than approximate. Against the previous export, on real
inputs: pred_dur and f0_pred are bitwise identical and n_pred differs in
2 of 1024 elements by one float16 ulp.
What it buys is the first load. Core ML specialises the MIL program for the
hardware on a first MLModel(contentsOf:) and caches the result; the cost
tracks the size of the program, not the weights. In an iOS 27 simulator, cold:
| model | before | after |
|---|---|---|
albert |
106 ms | 100 ms |
prosody |
309 609 ms | 221 ms |
text_encoder |
13 851 ms | 70 ms |
source |
24 ms | 24 ms |
decoder |
212 ms | 354 ms |
| 323.8 s | 0.77 s |
albert and source are unchanged by this and measured the same, which is the
control. Steady-state inference improved too β prosody 241 ms to 163 ms, text
encoder 47 ms to 34 ms β and the specialisation cache these five build fell
from 1.6 GB to 200 MB. A simulator has no Neural Engine, so all of these are
CPU program builds and a device may differ.
The sine generator uses floor-modulo
SineGen._f02sine computes (f0 / sampling_rate) % 1. PyTorch's % floors;
Core ML's mod truncates. They agree on positive inputs and disagree on
negative ones, and F0 is negative on unvoiced frames β so a frame that should
contribute 0.99996 of a cycle contributed 0. The gap is small per frame and
the phase is a running sum, so it accumulated into roughly twelve cycles of
drift across a sentence and the harmonics decorrelated from the fundamental.
Written as x - floor(x) instead, the converted model matches PyTorch to
1.00000000 correlation.
The iSTFT is a real inverse FFT
CustomSTFT.inverse overlap-adds a transposed convolution against a basis that
weights all eleven bins alike and never divides by the window envelope. A real
one-sided spectrum needs bins 1..N/2-1 doubled while DC and Nyquist are
counted once β and since the decoder's magnitude is exp(...), which is always
positive, the undoubled DC bin came out twice as hot as everything else and
landed as a constant offset on the waveform. About 18% of the output's power
sat below 80 Hz.
Doubling the interior bins and dividing by the precomputed hannΒ² envelope
gives 1.000000 correlation against torch.istft.
Why the source module is its own package
The decoder is float16: that is what makes its 53 M parameters eligible for the Neural Engine, and it halves the download. But the phase accumulator inside the sine generator counts cycles into the tens of thousands, and float16 has an 11-bit mantissa β past 2048 the spacing between representable values exceeds the fraction of a cycle a sample advances by, alternate periods collapse into each other, and every voice drops an octave (median F0 81 Hz against a reference 188 Hz).
Splitting the excitation out puts that one accumulator in float32, where it
needs about 15 bits, and leaves the weights β which are perfectly happy at
float16 β in the package that holds essentially all of them.
source.mlpackage carries a single 9β1 linear layer and comes to 33 KB.
source takes f0_curve [1, 1024] and returns har [1, 22, 61441], which is
the fifth input to decoder alongside asr, f0_curve, n and
acoustic_style.
Licence
Apache 2.0, as Kokoro-82M and kokoro-swift are.
- Downloads last month
- 102