Kokoro-82M β€” Core ML, segmented, Neural Engine

Kokoro-82M converted to Core ML as five chained .mlpackages, for on-device synthesis on iOS. Built for OpenReader and published so that the exact weights the app asserts against are pinned somewhere stable.

config.json
albert.mlmodelc          512-token ALBERT
prosody.mlmodelc         durations, F0 and noise curves
text_encoder.mlmodelc
source.mlmodelc          harmonic excitation (float32)
decoder.mlmodelc         ISTFT-Net vocoder (float16)
voices/*.npy             28 English voice packs, 510 style rows each
*.mlpackage              the uncompiled sources the above were built from

What an app should fetch is the .mlmodelc. Compiling an .mlpackage on the device costs minutes and produces exactly what xcrun coremlcompiler compile produces on a build machine in seconds, so the compiled form is published and the packages are kept only as what a future re-export starts from.

Inputs and outputs follow the segmentation in kokoro-swift (Apache 2.0), from which the conversion descends, with four deliberate differences.

ALBERT is exported at 512 tokens

Upstream's converter takes one --max-tokens defaulting to 512, and the published ALBERT reads as an artefact of a different run than its siblings: it declares [1, 64] where prosody, text_encoder and decoder declare [1, 512]. All the shapes are fixed, so the chain was capped by the smallest β€” 62 phonemes, about ten words, per forward pass. Kokoro predicts prosody over the whole sequence, so a cut every ten words restarts the intonation mid-sentence. This is ALBERT re-exported at 512 from the same kokoro-v1_0.pth weights; on text short enough for both, the two produce bitwise-identical audio.

The recurrent layers are real lstm ops

Upstream's exporter runs every bidirectional LSTM through a helper that loops over the time axis in Python, twice, freezing the state on padded steps with a torch.where per step. Traced, that is what it becomes: straight-line MIL, 512 steps by two directions by five gate operations, per layer. Prosody has five such layers, and its model.mil came to 16.9 MB against the decoder's 0.37 MB despite holding a third as many weights.

Core ML never has to see that. It has an lstm op that stays one node whatever the sequence length, and the only thing standing between the exporter and it is padding β€” a bidirectional nn.LSTM over a padded buffer is correct forwards and wrong backwards, because the reverse pass starts at the end of the buffer and accumulates state across the padding before it reaches real data.

So the reverse direction is run as a forward one. Gather the sequence through an index map that reverses the first n positions and leaves the padding where it is, run a unidirectional LSTM over that, and gather back through the same map β€” the map is an involution, so the same indices undo it. Position t then carries the state of a reverse pass that started at n-1, which is what the loop computed, and the padding is masked to zero as before.

That is exact rather than approximate. Against the previous export, on real inputs: pred_dur and f0_pred are bitwise identical and n_pred differs in 2 of 1024 elements by one float16 ulp.

What it buys is the first load. Core ML specialises the MIL program for the hardware on a first MLModel(contentsOf:) and caches the result; the cost tracks the size of the program, not the weights. In an iOS 27 simulator, cold:

model before after
albert 106 ms 100 ms
prosody 309 609 ms 221 ms
text_encoder 13 851 ms 70 ms
source 24 ms 24 ms
decoder 212 ms 354 ms
323.8 s 0.77 s

albert and source are unchanged by this and measured the same, which is the control. Steady-state inference improved too β€” prosody 241 ms to 163 ms, text encoder 47 ms to 34 ms β€” and the specialisation cache these five build fell from 1.6 GB to 200 MB. A simulator has no Neural Engine, so all of these are CPU program builds and a device may differ.

The sine generator uses floor-modulo

SineGen._f02sine computes (f0 / sampling_rate) % 1. PyTorch's % floors; Core ML's mod truncates. They agree on positive inputs and disagree on negative ones, and F0 is negative on unvoiced frames β€” so a frame that should contribute 0.99996 of a cycle contributed 0. The gap is small per frame and the phase is a running sum, so it accumulated into roughly twelve cycles of drift across a sentence and the harmonics decorrelated from the fundamental.

Written as x - floor(x) instead, the converted model matches PyTorch to 1.00000000 correlation.

The iSTFT is a real inverse FFT

CustomSTFT.inverse overlap-adds a transposed convolution against a basis that weights all eleven bins alike and never divides by the window envelope. A real one-sided spectrum needs bins 1..N/2-1 doubled while DC and Nyquist are counted once β€” and since the decoder's magnitude is exp(...), which is always positive, the undoubled DC bin came out twice as hot as everything else and landed as a constant offset on the waveform. About 18% of the output's power sat below 80 Hz.

Doubling the interior bins and dividing by the precomputed hannΒ² envelope gives 1.000000 correlation against torch.istft.

Why the source module is its own package

The decoder is float16: that is what makes its 53 M parameters eligible for the Neural Engine, and it halves the download. But the phase accumulator inside the sine generator counts cycles into the tens of thousands, and float16 has an 11-bit mantissa β€” past 2048 the spacing between representable values exceeds the fraction of a cycle a sample advances by, alternate periods collapse into each other, and every voice drops an octave (median F0 81 Hz against a reference 188 Hz).

Splitting the excitation out puts that one accumulator in float32, where it needs about 15 bits, and leaves the weights β€” which are perfectly happy at float16 β€” in the package that holds essentially all of them. source.mlpackage carries a single 9β†’1 linear layer and comes to 33 KB.

source takes f0_curve [1, 1024] and returns har [1, 22, 61441], which is the fifth input to decoder alongside asr, f0_curve, n and acoustic_style.

Licence

Apache 2.0, as Kokoro-82M and kokoro-swift are.

Downloads last month
102
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for richardr1126/Kokoro-82M-CoreML-ANE

Quantized
(83)
this model