GLiNER2.5-Decide ยท Person Resolution

A fine-tune of fastino/GLiNER2.5-Decide that makes the closed-set decisions behind resolving person-name mentions in noisy text, especially automatic speech recognition (ASR) transcripts of news and current-affairs broadcasts. Given a name mention in context and a set of candidate people, it decides who the mention refers to. It can also answer that the mention is a person not among the candidates, or not a person at all. It picks the best display name for a group of co-referent mentions, and tells genuine name variants apart from misspellings and transcription errors.

It uses the same architecture and API as the base model: one encoder forward pass, no generated tokens, and the label set is supplied at call time.

Base model fastino/GLiNER2.5-Decide (span architecture, DeBERTa-v3-large encoder)
Parameters 486M (incl. embeddings)
Language English
Max input 512 tokens (text + prompt + labels)
Revisions main: fp32 weights (1.9 GB) ยท bf16: bfloat16 weights (0.97 GB)

What it decides

The model was trained on three decision types. Each one is a single-choice classification over labels you pass at inference time, and each label comes with a natural-language description.

1. Mention โ†’ person (occurrence). Does the mention refer to one of the listed candidates?

label meaning
g0, g1, โ€ฆ a known person from your entity store, described by name and known aliases, e.g. "Fatih Birol (known names: Birol, Fati Biro)"
p0, p1, โ€ฆ a person already mentioned earlier in the same document, described as "a person referred to in this media as 'Dave'"
new a person who is not among the candidates
not_a_person the span is not a person's name (organisations, places, programmes, generic roles, ASR junk)

Titles that identify a specific person resolve to that person when context makes it clear, e.g. "the President" when a specific president is being discussed. Generic roles and metonyms ("the White House") are not_a_person.

2. Canonical name (canonical). Given the surface forms used for one person in a document (s0, s1, โ€ฆ), pick the most complete, correctly spelt full name. For example, it picks Fatih Birol over Fati Biro, Birol and Fatih Biro.

3. Name variant check (known_name). Is a surface form a legitimate name this person is known by (legit), or a misspelling or transcription error (error)? For example, Chng for Chung is an error.

Usage

pip install gliner2
from gliner2 import AutoExtractor

model = AutoExtractor.from_pretrained("hmarchant/gliner2.5-decide-person-resolution")
# bfloat16 weights (half the download). gliner2 2.0.0 does not forward
# `revision=` to the weight download, so fetch the branch first:
# from huggingface_hub import snapshot_download
# model = AutoExtractor.from_pretrained(snapshot_download(
#     "hmarchant/gliner2.5-decide-person-resolution", revision="bf16"))

instructions = (
    "Which person does the surface refer to in this transcript excerpt? "
    "Choose the matching person from the criteria, 'new' for a different person "
    "not listed, or 'not_a_person' if the surface is not a person's name."
)
state = (
    "surface: Birol\n"
    "context_before: Oil prices jumped again this morning, and the energy "
    "agency chief \n"
    "context_after:  said stock releases could be coordinated within days."
)
result = model.classify_text(
    "Surface: Birol",                         # short focus text
    {
        "occurrence": {
            "labels": {
                "g0": "Fatih Birol (known names: Birol, Fati Biro)",
                "g1": "Fatima Bibi",
                "new": "a different person not listed above",
                "not_a_person": "not a person's name",
            },
            "prompt": f"{instructions} | {state}",   # instructions | full state
            "multi_label": True,
            "cls_threshold": 0.0,               # return a score for every label
        }
    },
    include_confidence=True,
)
print(result)

Input convention used in training. Follow it at inference for best results.

  • text: a short focus string. Use Surface: <mention> for occurrence, Surface: <mention> | Canonical name: <name> for known_name, and a one-line situation sentence for canonical.
  • prompt: <instructions> | <state>, where the state is key: value lines. For occurrence these are surface, context_before and context_after (roughly 300 characters each). For canonical they are situation, surfaces (as s0: โ€ฆ; s1: โ€ฆ) and task. For known_name they are canonical_name, surface, situation and observed_surfaces.
  • labels: a dict from label key to description, as in the tables above. Set multi_label: True and cls_threshold: 0.0, then take the highest-scoring label. The per-label sigmoid score of the chosen label is a usable confidence (see calibration below).

Recommended instructions for the other two decision types:

  • canonical: "These surfaces all refer to the same person in one media. Choose the canonical name: the single most complete, correctly spelt full name among the options. Prefer a full name (given name plus surname) over a surname or given name alone; never choose a misspelt or truncated form."
  • known_name: "Is '<surface>' a legitimate name for <canonical>, or a spelling/transcription error?" with labels legit: "'<surface>' is a correctly spelt name or accepted short form this person is known by" and error: "'<surface>' is a misspelling, transcription error, or not how '<canonical>' is actually written".

Training data

About 29k decisions built from English-language radio and TV news and current-affairs broadcast transcripts (ASR output, 2026). Name mentions were found with a name-span detector. Candidate lists were simulated the way an entity-linking system would present them: previously seen people with alias lists, same-document people, new and not_a_person. Hard negatives (organisations, programmes, outlets, role phrases) were mined from the same text.

  • Splits by broadcast. No broadcast contributes to more than one split. Cross-split duplicates were removed.

    split occurrence known_name canonical total
    train 19,733 3,267 473 23,473
    dev 1,994 811 50 2,855
    test 1,804 746 52 2,602
  • LLM-curated labels. Every decision was labelled twice by an LLM (Grok 4.7, low reasoning effort), using two independently framed prompts with shuffled option order. A label was kept only when both passes agreed with at least medium confidence. Disagreements went to a third adjudication pass and were kept only for dev/test. A 500-label sample audit by the same LLM judged 98.6% of labels correct (Wilson 95% CI 97.1โ€“99.3%). This measures annotator consistency, not human-verified ground truth.

  • Augmentation. Some known_name examples use synthetic ASR-style corruptions of real names. In train, these are capped at 1.5ร— the number of legit examples.

Training procedure

Full fine-tune with the gliner2 trainer on a single RTX 3060 (12 GB), taking 3.4 h.

setting value
epochs 4 (5,868 optimiser steps)
batch 8 ร— 2 gradient accumulation (effective 16)
learning rate encoder 1e-5, task heads 1e-4, cosine schedule, 5% warmup
precision bf16 mixed precision
max length 448
selection best dev accuracy (epoch 4)

Dev accuracy by epoch: 0.941 โ†’ 0.944 โ†’ 0.948 โ†’ 0.949.

Evaluation

Held-out test split (2,602 decisions, unseen broadcasts)

Accuracy is the argmax label against the curated label. Baselines were run on the same inputs.

model overall occurrence (n=1,804) known_name (n=746) canonical (n=52)
this model 0.940 0.932 0.971 0.808
fastino/GLiNER2.5-Decide (base) 0.646 0.576 0.834 0.385
Laya typed-decisions (laya 0.3.11) 0.373 0.454 0.154 0.692
Jev (jev-1.13.0, TypeSafe AI hosted API) 0.908 0.885 0.967 0.885

not_a_person detection on occurrence decisions:

model precision recall
this model 0.967 0.955
base 0.000 0.000
Laya typed-decisions 0.596 0.478
Jev (jev-1.13.0) 0.952 0.935

By gold label (occurrence): new 733/768, graph candidate g0 556/591, not_a_person 322/337, other candidates 70/108.

Paired with Jev on the same 2,602 decisions, this model is right where Jev is wrong on 160, and Jev is right where this model is wrong on 76 (exact McNemar p โ‰ˆ 5e-8). Most of the gap is in occurrence decisions: Jev assigns more gold new mentions to an existing candidate (101/768, vs 29/768 for this model). The known_name (p = 0.74) and canonical (p = 0.34) differences are not significant.

Calibration: keeping only decisions whose top-label score is at least 0.9 retains 2,552 of 2,602 (98.1%) at 0.947 accuracy, against 0.940 overall. Jev keeps 56.6% of decisions at that threshold, at 0.986 accuracy. Jev's scores separate easy cases from hard ones more sharply; this model's scores are high more uniformly.

Latency: this model takes about 52 ms per decision (median) running locally on an RTX 3060, and Jev about 267 ms (median) as a hosted API call.

These test labels come from the same LLM curation process as the training labels, so they partly measure agreement with that labeller's conventions.

Separately labelled set

This set was labelled by a different LLM annotation process, with per-case rationales and accepted-alternative sets for ambiguous cases. It was not human-verified. It has 80 real broadcast mentions plus a hand-built stress-test suite: misspellings, partial names, shared surnames, non-person spans, canonical selection, and variant judgements.

model real occurrence (n=80) synthetic occurrence (n=25) canonical (n=8) known_name (n=20)
this model (bf16) 0.925 0.80 1.00 0.65
base 0.85 0.20 0.75 0.55
Laya typed-decisions 0.375 0.68 0.375 0.60
Jev (jev-1.13.0) 0.95 0.88 0.875 0.90

On this set, Jev is clearly ahead on the variant check (0.90 vs 0.65). It is slightly ahead on real mentions (0.95 vs 0.925), but that difference is not significant: on the discordant cases Jev wins 5 and this model wins 3 (p = 0.73).

Precision (fp32 vs bf16)

Both revisions were evaluated on the same RTX 3060, one decision per forward pass.

main (fp32) bf16
test overall (n=2,602) 0.9404 0.9400
test occurrence / known_name / canonical 0.9318 / 0.9705 / 0.8077 0.9313 / 0.9705 / 0.8077
test not_a_person precision / recall 0.967 / 0.955 0.964 / 0.955
real occurrence (n=80) 0.925 0.925
stress-test suite (n=53) 0.774 0.774
median latency per decision 52 ms 54 ms

The two revisions choose the same label on 2,601 of 2,602 test decisions and on all 80 real ones. The single flip is a near-tie (new 0.52 against not_a_person 0.47 in fp32) that bf16 gets wrong. The top-label score moves by 0.0002 on average (at most 0.058).

gliner2 2.0.0 loads the bf16 weights into fp32 parameters, so the bf16 revision halves download and disk size but not memory use or compute time. Use main by default. Use bf16 when download size or storage matters; the accuracy cost is negligible.

Limitations

  • Domain. English news and current-affairs speech transcripts. Other domains, languages and text styles are untested.
  • Short and partial names. The variant check is the weakest skill. On the stress-test suite it over-predicts error for legitimate short forms and nicknames (0.65 accuracy).
  • Earlier-mention vs new person. Some confusion remains between "a person mentioned earlier in this document" (p*) and new.
  • Canonical selection is trained and tested on few examples (473 train, 52 test), so treat the 0.81 test accuracy as noisy.
  • Labels come from an LLM and were not human-reviewed.
  • The model only chooses among the options you give it. Resolution quality depends on your candidate retrieval.
Downloads last month
42
Safetensors
Model size
0.5B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for hmarchant/gliner2.5-decide-person-resolution

Finetuned
(9)
this model