File size: 2,390 Bytes
190f3ed
 
 
 
 
 
 
 
 
 
 
 
 
802a28a
190f3ed
 
 
 
 
 
 
802a28a
 
190f3ed
 
 
 
 
 
802a28a
 
190f3ed
802a28a
190f3ed
 
 
 
 
 
802a28a
 
190f3ed
 
 
802a28a
 
190f3ed
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
---
license: other
license_name: baa-proprietary
library_name: sentence-transformers
tags:
- retrieval
- embeddings
- reranker
- cross-encoder
- rag
- sentence-similarity
pipeline_tag: sentence-similarity
---

# baa.ai · Merino-Pro-4bit

**The premium unified retrieval model — bi-encoder embedding *and* cross-encoder reranking in one package, over a single shared word-embedding table.** A 1024-dimensional multilingual model, by BAA AI (Black Sheep AI). ~872M params, 4-bit quantized (~0.5 GB).

## Get the optimal model for *your* data

Merino-Pro-4bit is baa.ai's flagship default. But the best embedder + reranker is **corpus-specific** — the ideal choice depends on your documents and your notion of relevance. **baa.ai offers exclusive tooling that identifies the optimal embedding and reranking models for your specific data**, so you ship the smallest models that maximize document recovery on your corpus. For a tailored recommendation, **reach out to baa.ai**.

## What it is

A two-role retrieval model over a **shared input word-embedding matrix** (~256M params, stored once). The bi-encoder embedder and cross-encoder reranker are both built on the `xlm-roberta-large` backbone and **co-trained** to share a single word-embedding table at **no quality loss**, while each role keeps its native layers and head. 4-bit weight quantization is lossless on this stack (the embedder is the limiter; rerank tolerates lower bits).

- **Embed role:** bi-encoder, 1024-d, L2-normalized. Prepend `"query: "` to queries.
- **Rerank role:** cross-encoder, single relevance logit per (query, document) pair.
- **Router:** call `.embed(...)` or `.rerank(...)`.

## Usage

```python
from modeling_baa import BaaEmbeddingReranker   # included in this repo

m = BaaEmbeddingReranker("baa-ai/Merino-Pro-4bit")
qv = m.embed(["my query"], is_query=True)[0]     # 1024-d normalized
dv = m.embed(["doc a", "doc b"])
ranked = m.rerank("my query", ["doc a", "doc b"], top_k=10)   # [(doc, score), ...]
```

## License & attribution

- **BAA Contributions** (shared-embedding architecture, router/loader code, packaging, weights, docs) are **proprietary to BAA AI (Black Sheep AI)** — see `LICENSE`.
- Incorporates the `xlm-roberta-large` backbone under the **MIT License** — see `LICENSE-xlm-roberta-large.txt`.

© 2026 BAA AI (Black Sheep AI) — baa.ai. Provided "as is" without warranty.