Gist

Generate topics and tags for posts and articles.

Multilingual on-device content topic tagging across a 36-topic taxonomy.

Drop in a title, a post or a longer description and get its topics back, from a fixed list of 36, across 101 languages. A compact two-stream classifier (static embedding + hashed n-grams), with no transformer at inference. The deployable model is 74MB (int8 vocab-pruned multilingual embedding + a small fp16 head) and runs fully on device with zero per-call cost. Multi-label by design: most items carry two or three topics, and per-item scores can be aggregated across a collection (for example into channel- or feed-level topics).

"How to film a two-person podcast with two iPhones"technology, creator-economy · "Cómo invertir en fondos indexados"finance · "Tips for adopting a rescue dog"pets-animals · "投资指数基金入门"finance

Try it

Platforms iOS, macOS, tvOS, visionOS, Android, Linux, Windows, Browser, Node
Languages 101
Weights v2.2.0

Install

Swift (requirements)

.package(url: "https://github.com/Desert-Ant-Labs/desert-ant-core.git", from: "3.1.0")

Then add the Gist product to your target.

Kotlin (requirements)

implementation("ai.desertant:gist:3.1.0")

JavaScript (requirements)

npm i @desert-ant-labs/gist @litertjs/core   # browser
npm i @desert-ant-labs/gist                  # Node, prebuilt native core

Use

Drop-in SDKs run the model on device; each pins this repo's revision.

import { Gist } from "@desert-ant-labs/gist";           // browser (wasm + LiteRT.js)
// import { Gist } from "@desert-ant-labs/gist/native"; // Node (native)
const gist = await Gist.load();
await gist.classify("How to start a podcast with just your iPhone");
// [{ slug: "technology", name: "Technology & Software", score: 0.91 },
//  { slug: "creator-economy", name: "Creator Economy & Marketing", score: 0.44 }]
import Gist
let gist = Gist()
let topics = try await gist.classify("How to start a podcast with just your iPhone")

Files

File Format Size Contents
gist_embedding.i8 + .json int8 static embedding 64MB 101-language static embedding, the semantic feature extractor
gist.mlmodelc Core ML 6MB The classifier head: fused features → 36 topic probabilities
gist.tflite LiteRT 13MB The same head, float32
gist_tokenizer.bin Unigram 4MB The multilingual tokenizer
gist_config.json JSON tiny Slugs, feature dims, threshold
taxonomy.json JSON 8KB The 36 topics (slug, name, description, IAB + Apple category)

Inputs and outputs

  • Input: a plain text string (title, or title + description). Best on short text like posts, titles, and descriptions.
  • Output: a probability over the 36 topics (features [1, 8448]topic_probs [1, 36]); take the top-k above the threshold in gist_config.json. Optimized for multi-label use, an item's 2, 3 topics, optionally aggregated across a collection.

Topics and standard taxonomy

The 36 topics map to two industry-standard taxonomies so gist output can be rolled up or joined into existing systems: IAB Content Taxonomy 2.2 (with each node's stable integer ID) and Apple Podcasts categories. The full, machine-readable crosswalk ships in this repo as taxonomy_crosswalk.json (e.g. law → IAB 383 News & Politics › Law, crafts-hobbies → IAB 248 Arts and Crafts, finance → IAB 391 Personal Finance).

Five topics have no dedicated IAB 2.2 node and are flagged as gist extensions (society-culture, creator-economy, outdoors-nature map to a nearest parent; history and self-improvement have no IAB node); film-tv is a roll-up of IAB Movies + Television.

Languages

Topic tagging covers 101 languages. A diverse 15-language spot check (across Latin, Cyrillic, Arabic, CJK, Devanagari, Hebrew, Thai, and Greek scripts) gives 88% top-3, with CJK, Arabic, and Cyrillic scripts matching or beating the Latin ones.

Model variants

Two builds of the same 36-topic model live in this repo:

Variant Location Size Coverage
Multilingual (default) repo root 74MB 101 languages
English-only en/ 15MB English / Latin script only

The English build is the same model with a smaller embedding and tokenizer, so it is topic-identical to the multilingual model on English input. The English build does not cover non-Latin scripts (CJK, Arabic, Cyrillic, …); use it only when the input is reliably English/Latin. The Swift SDK selects it with Gist(variant: .english). The JS and Kotlin SDKs currently load the multilingual build only: variant selection has to cross the shared native ABI, which has no slot for it yet.

Evaluation

Recall on a held-out set of 572 human-labeled real posts (36 topics), zero-shot for the LLMs and zero-shot classifiers. Embedding classifiers get a light logistic head; recall@3 is the product metric (downstream aggregation consumes the top few topics).

Model Type Size recall@1 recall@3
Qwen2.5-7B (cloud) LLM zero-shot server 79% n/a
multilingual-e5-small + head transformer embed 110MB 74% 92%
bge-small-en + head transformer embed 130MB 71% 92%
gist on-device 74MB 71% 91%
all-MiniLM-L6-v2 + head transformer embed 90MB 68% 90%
mDeBERTa-v3-mnli-xnli zero-shot NLI 560MB 50% 73%
GLiClass-base zero-shot 400MB 44% 65%

gist is tied on recall@3 with the best small models, at a fraction of the size and one on-device pass, and it beats every zero-shot classifier decisively (they never learned the taxonomy or the distribution). Only a 7B cloud LLM clearly leads on recall@1.

License

Desert Ant Labs Source-Available License. Free for most apps, and a commercial license is required at scale. Full terms are at the link. Licensing: licensing@desertant.com.

See THIRD_PARTY_NOTICES.md.

Citation

@software{gist_2026,
  title  = {Gist: Multilingual on-device content topic tagging across a 36-topic taxonomy},
  author = {Desert Ant Labs},
  year   = {2026},
  url    = {https://huggingface.co/desert-ant-labs/gist},
}

© 2026 Desert Ant Labs · https://desertant.com

Downloads last month
1,339
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using desert-ant-labs/gist 1