Instructions to use chrullis/relweave-4b-zeroshot with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use chrullis/relweave-4b-zeroshot with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/qwen3-4b-unsloth-bnb-4bit") model = PeftModel.from_pretrained(base_model, "chrullis/relweave-4b-zeroshot") - Notebooks
- Google Colab
- Kaggle
relweave-4b-zeroshot (experimental)
Relation extraction for relation types you define yourself, without training. You give a passage, its entities and, for each relation type, a one-sentence definition; the model answers, for every entity pair and type, how likely it is that the text states that relation. It is a LoRA adapter on Qwen3-4B (4-bit), trained to answer such yes/no questions on many relation types, and tested on types it never saw.
It is the experimental zero-shot companion of relweave-4b-base, which extracts a fixed business schema of 21 relation types at about 0.78 F1.
The easiest way to use it is the relweave library: your schema in, a graph out, with the entities found by relweave-4b-base:
from relweave import Extractor
from relweave.schema import Relation, Schema
from relweave.schema.business import Org, Person
class DonatedTo(Relation[Person | Org, Org]):
"""The source has given money, goods or other gifts to the target organisation."""
ex = Extractor(schema=Schema(name="charity", entities=[Person, Org], relations=[DonatedTo]), zeroshot=True)
graph = ex.run(open("report.txt").read())
or relweave run report.txt --schema charity.py:CHARITY --zeroshot --out graph.json. The scripts
below use the adapter directly, with your own entities.
How it works
For each ordered entity pair (source, target) and each relation type whose endpoint types fit, the model reads:
<passage>
Entities:
E1 Org: Harbour Light Foundation | the foundation
E2 Person: Ingrid Valtersen
...
For each question below, answer yes if the text states or clearly implies that the relation holds from the source to the target, otherwise no.
Q: Source: Ingrid Valtersen | Target: Harbour Light Foundation | Relation: The source has given money, goods or other gifts to the target organisation. A:
and the probability of " yes" (against " no") at the end is the score. A relation is predicted when the score is above
a threshold (see "Choosing the threshold"). Many questions about one passage can share one forward pass
(evaluate.py does this).
Usage
example.py in this repository is a minimal, runnable example (needs a CUDA GPU, transformers, peft,
bitsandbytes):
from example import load, score
tok, model = load()
text = "Retired surgeon Ingrid Valtersen gave the Harbour Light Foundation 400,000 kroner in March. ..."
entities = {"Harbour Light Foundation": ("Org", ["Harbour Light Foundation"]),
"Ingrid Valtersen": ("Person", ["Ingrid Valtersen"])}
donated = "The source has given money, goods or other gifts to the target organisation."
print(score(tok, model, text, entities, [("Ingrid Valtersen", "Harbour Light Foundation", donated)]))
Its built-in demo prints:
0.994 Ingrid Valtersen -> Harbour Light Foundation: has given money, goods or other gifts ... (true)
0.001 Fenwick & Rowe -> Harbour Light Foundation: has given money, goods or other gifts ... (only promised)
0.223 Fenwick & Rowe -> Harbour Light Foundation: has promised a future gift ... (true, weak)
0.835 Tomas Ekdahl -> Harbour Light Foundation: works for ... without pay (true)
0.006 Ola Brandt -> Harbour Light Foundation: works for ... without pay (paid director)
Write definitions as a sentence about "the source" and "the target", say what the relation is and, for near types, what it is not ("a future gift that has not been given yet"). Ask only types whose endpoint types fit the pair.
Testing it
evaluate.py reproduces the numbers below and scores your own labelled passages:
python evaluate.py # the published benchmark, splits "short" and "dense"
python evaluate.py --data my_passages.jsonl --types my_types.json
Your passages are JSON lines {"id", "text", "entities": [{"name", "type", "mentions"}], "relations": [{"source", "target", "type"}]}; the types file follows types.json
(per set: allowed source and target entity types, and NAME: definition). It reports F1 twice: tuned (threshold
chosen on half of your passages, F1 on the other half) and untuned (a fixed threshold, default 0.5). Label 10-20
passages of your own text to see which one applies to you and to pick your threshold.
Results
Strict F1 over (source, target, type) triples; entities are given.
| benchmark | types unseen in training | tuned threshold | untuned (0.5) |
|---|---|---|---|
| relweave-zeroshot-eval short: 100 fictional passages, 4-8 entities, 12 types | 12 | 0.82 | 0.77 |
| relweave-zeroshot-eval dense: 50 passages, 18-24 entities, incl. 10 list-heavy | 12 | 0.74 | 0.57 (0.73 at 0.93) |
| Re-DocRED dev, real Wikipedia, 15 held-out Wikidata types (threshold on other documents) | 15 | 0.50 global, 0.56 per type | - |
| GLiREL (a dedicated zero-shot relation model) on the same Re-DocRED task | - | 0.18 | - |
- Near types are told apart well: on the synthetic sets the highest-scoring type of a pair's sibling set is the gold one 96-98% of the time, and reversed directions almost never pass.
- On Re-DocRED a blind audit found that about a third of the model's false positives are stated in the text but missing from the gold, and only 57% of the gold is stated in the text (the rest is world knowledge); the official 0.56 understates how well it reads.
- Lists: 0.74 F1 on the list-heavy passages, but only about two thirds of list members are found.
Choosing the threshold
The best raw threshold depends on the text: about 0.7 on short passages with few entities, about 0.9 on dense ones and
on Re-DocRED, because every additional entity pair is another chance for a false yes. relweave calibrates for this by
default: each score is adjusted for the number n of entity pairs its type is asked on in the chunk,
logit(p) - 0.5 ln(n / 5), and compared with one threshold (0.6); counting per type keeps a type independent of how many
other types the schema has. Fitted on the synthetic benchmark this gave 0.84 and 0.73 on short and dense passages
(best fixed threshold per set 0.83 and 0.74) and 0.47 on held-out Re-DocRED (its best fixed threshold 0.51, a raw 0.8
0.47): it removes most of the density effect, not differences between kinds of text, and a single document can still
sit near the threshold. For your own text, calibrate on a few labelled
passages (relweave.zeroshot.calibrate, or evaluate.py --data for the raw scores).
Limitations
- Experimental. Relation types only: entities come from relweave-4b-base (Person, Org, Place, Object, Event, Coordinate) or from you.
- English only. The cost grows with entity pairs x types: a passage with 20 entities and 12 types is about 1,800
questions, about 35 seconds on an 8 GB RTX 3070 Ti with packed questions (
evaluate.py).
Training
A LoRA (rank 16) on the 4-bit Qwen3-4B, trained 1 epoch with cross-entropy on the " yes" / " no" answer of packed questions: gold triples plus negatives (the same pair with another fitting type, the reversed direction, random fitting pairs). Data: 600 Re-DocRED training documents with their relation types except the 15 held-out types above (director, composer, author, performer, record label, located in or next to body of water, mouth of the watercourse, league or competition, platform, developer, publisher, present in work, cast member, religion or worldview, official language) and their near neighbours (member of, member of political party, member of sports team, father, mother, spouse, child, sibling, owned by); 600 business passages from relweave-business-data with all business types except MEMBER_OF, FAMILY_OF and OWNS_STAKE_IN; and Re-DocRED passages labelled with open relation types by a large language model (types close to held-out ones removed). 1,693 passages, about 80,000 questions. The 12 types of the synthetic benchmark occur in none of this data.
Licences and attribution
The adapter weights are released under the Apache License 2.0, like the base model Qwen3-4B. Training text comes from English Wikipedia (CC BY-SA 4.0; attribution: Wikipedia contributors, https://en.wikipedia.org) through Re-DocRED (MIT) and relweave-business-data (CC BY-SA 4.0); redistributing that text or labels derived from it falls under CC BY-SA 4.0.
Files
| file | contents |
|---|---|
adapter_model.safetensors, adapter_config.json |
the LoRA adapter (PEFT) for unsloth/qwen3-4b-unsloth-bnb-4bit |
tokenizer.json, tokenizer_config.json, chat_template.jinja |
the tokenizer and chat template the prompt is built with |
example.py |
minimal scoring example |
evaluate.py |
benchmark and your-own-data evaluation |
- Downloads last month
- 21