Sentence Similarity
sentence-transformers
Safetensors
Korean
qwen3_vl
multimodal
embedding
visual-document-retrieval
korean
matryoshka
Instructions to use whybe-choi/kovre with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use whybe-choi/kovre with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("whybe-choi/kovre") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
Update README.md
Browse files
README.md
CHANGED
|
@@ -32,7 +32,7 @@ KoVRE is a 2B-parameter single-vector embedding model designed for **Korean visu
|
|
| 32 |
KoVRE is initialized from [`Qwen/Qwen3-VL-Embedding-2B`](https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B) and trained in two stages. The first stage performs contrastive learning on 708,729 Korean and English query-page pairs with positive-aware hard-negative mining. The second stage distills relevance scores from `Qwen3-VL-Reranker-8B` using Korean data. Matryoshka Representation Learning enables flexible embedding dimensions from 128 to 2,048.
|
| 33 |
|
| 34 |
- **Code:** https://github.com/whybe-choi/kovre
|
| 35 |
-
- **Paper:**
|
| 36 |
|
| 37 |
## Highlights
|
| 38 |
|
|
@@ -327,9 +327,10 @@ KoVRE is released under the Apache 2.0 License.
|
|
| 327 |
If you find KoVRE useful, please cite:
|
| 328 |
|
| 329 |
```bibtex
|
| 330 |
-
@
|
| 331 |
-
title
|
| 332 |
-
author
|
| 333 |
-
|
|
|
|
| 334 |
}
|
| 335 |
```
|
|
|
|
| 32 |
KoVRE is initialized from [`Qwen/Qwen3-VL-Embedding-2B`](https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B) and trained in two stages. The first stage performs contrastive learning on 708,729 Korean and English query-page pairs with positive-aware hard-negative mining. The second stage distills relevance scores from `Qwen3-VL-Reranker-8B` using Korean data. Matryoshka Representation Learning enables flexible embedding dimensions from 128 to 2,048.
|
| 33 |
|
| 34 |
- **Code:** https://github.com/whybe-choi/kovre
|
| 35 |
+
- **Paper:** https://arxiv.org/abs/2608.01389
|
| 36 |
|
| 37 |
## Highlights
|
| 38 |
|
|
|
|
| 327 |
If you find KoVRE useful, please cite:
|
| 328 |
|
| 329 |
```bibtex
|
| 330 |
+
@article{choi2026kovre,
|
| 331 |
+
title={KoVRE: Training an Efficient Embedding Model for Korean Visual Document Retrieval},
|
| 332 |
+
author={Choi, Yongbin and Shim, Gyuho and Jang, Youngjoon},
|
| 333 |
+
journal={arXiv preprint arXiv:2608.01389},
|
| 334 |
+
year={2026}
|
| 335 |
}
|
| 336 |
```
|