Instructions to use hmarchant/gliner2.5-decide-person-resolution with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- GLiNER2
How to use hmarchant/gliner2.5-decide-person-resolution with GLiNER2:
from gliner2 import AutoExtractor extractor = AutoExtractor.from_pretrained("hmarchant/gliner2.5-decide-person-resolution") # Extract entities text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday." result = extractor.extract_entities(text, ["company", "person", "product", "location"]) print(result) - Notebooks
- Google Colab
- Kaggle
GLiNER2.5-Decide ยท Person Resolution
A fine-tune of fastino/GLiNER2.5-Decide
that makes the closed-set decisions behind resolving person-name mentions
in noisy text, especially automatic speech recognition (ASR) transcripts of
news and current-affairs broadcasts. Given a name mention in context and a set
of candidate people, it decides who the mention refers to. It can also answer
that the mention is a person not among the candidates, or not a person at all.
It picks the best display name for a group of co-referent mentions, and tells
genuine name variants apart from misspellings and transcription errors.
It uses the same architecture and API as the base model: one encoder forward pass, no generated tokens, and the label set is supplied at call time.
| Base model | fastino/GLiNER2.5-Decide (span architecture, DeBERTa-v3-large encoder) |
| Parameters | 486M (incl. embeddings) |
| Language | English |
| Max input | 512 tokens (text + prompt + labels) |
| Revisions | main: fp32 weights (1.9 GB) ยท bf16: bfloat16 weights (0.97 GB) |
What it decides
The model was trained on three decision types. Each one is a single-choice classification over labels you pass at inference time, and each label comes with a natural-language description.
1. Mention โ person (occurrence). Does the mention refer to one of the
listed candidates?
| label | meaning |
|---|---|
g0, g1, โฆ |
a known person from your entity store, described by name and known aliases, e.g. "Fatih Birol (known names: Birol, Fati Biro)" |
p0, p1, โฆ |
a person already mentioned earlier in the same document, described as "a person referred to in this media as 'Dave'" |
new |
a person who is not among the candidates |
not_a_person |
the span is not a person's name (organisations, places, programmes, generic roles, ASR junk) |
Titles that identify a specific person resolve to that person when context
makes it clear, e.g. "the President" when a specific president is being
discussed. Generic roles and metonyms ("the White House") are
not_a_person.
2. Canonical name (canonical). Given the surface forms used for one
person in a document (s0, s1, โฆ), pick the most complete, correctly spelt
full name. For example, it picks Fatih Birol over Fati Biro, Birol and
Fatih Biro.
3. Name variant check (known_name). Is a surface form a legitimate name
this person is known by (legit), or a misspelling or transcription error
(error)? For example, Chng for Chung is an error.
Usage
pip install gliner2
from gliner2 import AutoExtractor
model = AutoExtractor.from_pretrained("hmarchant/gliner2.5-decide-person-resolution")
# bfloat16 weights (half the download). gliner2 2.0.0 does not forward
# `revision=` to the weight download, so fetch the branch first:
# from huggingface_hub import snapshot_download
# model = AutoExtractor.from_pretrained(snapshot_download(
# "hmarchant/gliner2.5-decide-person-resolution", revision="bf16"))
instructions = (
"Which person does the surface refer to in this transcript excerpt? "
"Choose the matching person from the criteria, 'new' for a different person "
"not listed, or 'not_a_person' if the surface is not a person's name."
)
state = (
"surface: Birol\n"
"context_before: Oil prices jumped again this morning, and the energy "
"agency chief \n"
"context_after: said stock releases could be coordinated within days."
)
result = model.classify_text(
"Surface: Birol", # short focus text
{
"occurrence": {
"labels": {
"g0": "Fatih Birol (known names: Birol, Fati Biro)",
"g1": "Fatima Bibi",
"new": "a different person not listed above",
"not_a_person": "not a person's name",
},
"prompt": f"{instructions} | {state}", # instructions | full state
"multi_label": True,
"cls_threshold": 0.0, # return a score for every label
}
},
include_confidence=True,
)
print(result)
Input convention used in training. Follow it at inference for best results.
- text: a short focus string. Use
Surface: <mention>foroccurrence,Surface: <mention> | Canonical name: <name>forknown_name, and a one-line situation sentence forcanonical. - prompt:
<instructions> | <state>, where the state iskey: valuelines. Foroccurrencethese aresurface,context_beforeandcontext_after(roughly 300 characters each). Forcanonicalthey aresituation,surfaces(ass0: โฆ; s1: โฆ) andtask. Forknown_namethey arecanonical_name,surface,situationandobserved_surfaces. - labels: a dict from label key to description, as in the tables above.
Set
multi_label: Trueandcls_threshold: 0.0, then take the highest-scoring label. The per-label sigmoid score of the chosen label is a usable confidence (see calibration below).
Recommended instructions for the other two decision types:
canonical: "These surfaces all refer to the same person in one media. Choose the canonical name: the single most complete, correctly spelt full name among the options. Prefer a full name (given name plus surname) over a surname or given name alone; never choose a misspelt or truncated form."known_name: "Is '<surface>' a legitimate name for <canonical>, or a spelling/transcription error?" with labelslegit: "'<surface>' is a correctly spelt name or accepted short form this person is known by"anderror: "'<surface>' is a misspelling, transcription error, or not how '<canonical>' is actually written".
Training data
About 29k decisions built from English-language radio and TV news and
current-affairs broadcast transcripts (ASR output, 2026). Name mentions were
found with a name-span detector. Candidate lists were simulated the way an
entity-linking system would present them: previously seen people with alias
lists, same-document people, new and not_a_person. Hard negatives
(organisations, programmes, outlets, role phrases) were mined from the same
text.
Splits by broadcast. No broadcast contributes to more than one split. Cross-split duplicates were removed.
split occurrence known_name canonical total train 19,733 3,267 473 23,473 dev 1,994 811 50 2,855 test 1,804 746 52 2,602 LLM-curated labels. Every decision was labelled twice by an LLM (Grok 4.7, low reasoning effort), using two independently framed prompts with shuffled option order. A label was kept only when both passes agreed with at least medium confidence. Disagreements went to a third adjudication pass and were kept only for dev/test. A 500-label sample audit by the same LLM judged 98.6% of labels correct (Wilson 95% CI 97.1โ99.3%). This measures annotator consistency, not human-verified ground truth.
Augmentation. Some
known_nameexamples use synthetic ASR-style corruptions of real names. In train, these are capped at 1.5ร the number oflegitexamples.
Training procedure
Full fine-tune with the gliner2 trainer on a single RTX 3060 (12 GB),
taking 3.4 h.
| setting | value |
|---|---|
| epochs | 4 (5,868 optimiser steps) |
| batch | 8 ร 2 gradient accumulation (effective 16) |
| learning rate | encoder 1e-5, task heads 1e-4, cosine schedule, 5% warmup |
| precision | bf16 mixed precision |
| max length | 448 |
| selection | best dev accuracy (epoch 4) |
Dev accuracy by epoch: 0.941 โ 0.944 โ 0.948 โ 0.949.
Evaluation
Held-out test split (2,602 decisions, unseen broadcasts)
Accuracy is the argmax label against the curated label. Baselines were run on the same inputs.
| model | overall | occurrence (n=1,804) | known_name (n=746) | canonical (n=52) |
|---|---|---|---|---|
| this model | 0.940 | 0.932 | 0.971 | 0.808 |
fastino/GLiNER2.5-Decide (base) |
0.646 | 0.576 | 0.834 | 0.385 |
Laya typed-decisions (laya 0.3.11) |
0.373 | 0.454 | 0.154 | 0.692 |
Jev (jev-1.13.0, TypeSafe AI hosted API) |
0.908 | 0.885 | 0.967 | 0.885 |
not_a_person detection on occurrence decisions:
| model | precision | recall |
|---|---|---|
| this model | 0.967 | 0.955 |
| base | 0.000 | 0.000 |
Laya typed-decisions |
0.596 | 0.478 |
Jev (jev-1.13.0) |
0.952 | 0.935 |
By gold label (occurrence): new 733/768, graph candidate g0 556/591,
not_a_person 322/337, other candidates 70/108.
Paired with Jev on the same 2,602 decisions, this model is right where Jev
is wrong on 160, and Jev is right where this model is wrong on 76 (exact
McNemar p โ 5e-8). Most of the gap is in occurrence decisions: Jev assigns
more gold new mentions to an existing candidate (101/768, vs 29/768 for
this model). The known_name (p = 0.74) and canonical (p = 0.34) differences
are not significant.
Calibration: keeping only decisions whose top-label score is at least 0.9 retains 2,552 of 2,602 (98.1%) at 0.947 accuracy, against 0.940 overall. Jev keeps 56.6% of decisions at that threshold, at 0.986 accuracy. Jev's scores separate easy cases from hard ones more sharply; this model's scores are high more uniformly.
Latency: this model takes about 52 ms per decision (median) running locally on an RTX 3060, and Jev about 267 ms (median) as a hosted API call.
These test labels come from the same LLM curation process as the training labels, so they partly measure agreement with that labeller's conventions.
Separately labelled set
This set was labelled by a different LLM annotation process, with per-case rationales and accepted-alternative sets for ambiguous cases. It was not human-verified. It has 80 real broadcast mentions plus a hand-built stress-test suite: misspellings, partial names, shared surnames, non-person spans, canonical selection, and variant judgements.
| model | real occurrence (n=80) | synthetic occurrence (n=25) | canonical (n=8) | known_name (n=20) |
|---|---|---|---|---|
| this model (bf16) | 0.925 | 0.80 | 1.00 | 0.65 |
| base | 0.85 | 0.20 | 0.75 | 0.55 |
Laya typed-decisions |
0.375 | 0.68 | 0.375 | 0.60 |
Jev (jev-1.13.0) |
0.95 | 0.88 | 0.875 | 0.90 |
On this set, Jev is clearly ahead on the variant check (0.90 vs 0.65). It is slightly ahead on real mentions (0.95 vs 0.925), but that difference is not significant: on the discordant cases Jev wins 5 and this model wins 3 (p = 0.73).
Precision (fp32 vs bf16)
Both revisions were evaluated on the same RTX 3060, one decision per forward pass.
main (fp32) |
bf16 |
|
|---|---|---|
| test overall (n=2,602) | 0.9404 | 0.9400 |
| test occurrence / known_name / canonical | 0.9318 / 0.9705 / 0.8077 | 0.9313 / 0.9705 / 0.8077 |
test not_a_person precision / recall |
0.967 / 0.955 | 0.964 / 0.955 |
| real occurrence (n=80) | 0.925 | 0.925 |
| stress-test suite (n=53) | 0.774 | 0.774 |
| median latency per decision | 52 ms | 54 ms |
The two revisions choose the same label on 2,601 of 2,602 test decisions
and on all 80 real ones. The single flip is a near-tie (new 0.52 against
not_a_person 0.47 in fp32) that bf16 gets wrong. The top-label score moves
by 0.0002 on average (at most 0.058).
gliner2 2.0.0 loads the bf16 weights into fp32 parameters, so the bf16
revision halves download and disk size but not memory use or compute time.
Use main by default. Use bf16 when download size or storage matters;
the accuracy cost is negligible.
Limitations
- Domain. English news and current-affairs speech transcripts. Other domains, languages and text styles are untested.
- Short and partial names. The variant check is the weakest skill. On the
stress-test suite it over-predicts
errorfor legitimate short forms and nicknames (0.65 accuracy). - Earlier-mention vs new person. Some confusion remains between "a person
mentioned earlier in this document" (
p*) andnew. - Canonical selection is trained and tested on few examples (473 train, 52 test), so treat the 0.81 test accuracy as noisy.
- Labels come from an LLM and were not human-reviewed.
- The model only chooses among the options you give it. Resolution quality depends on your candidate retrieval.
- Downloads last month
- 42