shikharshahi commited on
Commit
17eb2cc
·
verified ·
1 Parent(s): aebe9af

Update Sifter reranker model card

Browse files
Files changed (1) hide show
  1. README.md +191 -39
README.md CHANGED
@@ -2,62 +2,214 @@
2
  library_name: transformers
3
  license: apache-2.0
4
  base_model: distilbert-base-uncased
 
5
  tags:
6
- - generated_from_trainer
 
 
 
 
 
 
 
 
 
 
7
  model-index:
8
- - name: sifter-redrob-reranker
9
- results: []
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
10
  ---
11
 
12
- <!-- This model card has been generated automatically according to the information the Trainer had access to. You
13
- should probably proofread and complete it, then remove this comment. -->
14
 
15
- # sifter-redrob-reranker
16
 
17
- This model is a fine-tuned version of [distilbert-base-uncased](https://huggingface.co/distilbert-base-uncased) on an unknown dataset.
18
- It achieves the following results on the evaluation set:
19
- - Loss: 0.0443
20
- - Rmse: 0.2104
21
- - Mae: 0.1884
22
- - Spearman: 0.7526
23
 
24
- ## Model description
 
25
 
26
- More information needed
27
 
28
- ## Intended uses & limitations
29
 
30
- More information needed
 
 
 
31
 
32
- ## Training and evaluation data
33
 
34
- More information needed
 
 
 
35
 
36
- ## Training procedure
37
 
38
- ### Training hyperparameters
 
 
39
 
40
- The following hyperparameters were used during training:
41
- - learning_rate: 2e-05
42
- - train_batch_size: 8
43
- - eval_batch_size: 8
44
- - seed: 42
45
- - optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
46
- - lr_scheduler_type: linear
47
- - num_epochs: 3
48
 
49
- ### Training results
50
 
51
- | Training Loss | Epoch | Step | Validation Loss | Rmse | Mae | Spearman |
52
- |:-------------:|:-----:|:----:|:---------------:|:------:|:------:|:--------:|
53
- | 0.1236 | 1.0 | 21 | 0.0443 | 0.2104 | 0.1884 | 0.7526 |
54
- | 0.0473 | 2.0 | 42 | 0.0424 | 0.2060 | 0.1675 | 0.7192 |
55
- | 0.0476 | 3.0 | 63 | 0.0356 | 0.1886 | 0.1553 | 0.7192 |
56
 
 
 
 
 
 
 
 
 
 
 
 
57
 
58
- ### Framework versions
59
 
60
- - Transformers 5.9.0
61
- - Pytorch 2.11.0+cu128
62
- - Datasets 4.0.0
63
- - Tokenizers 0.22.2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2
  library_name: transformers
3
  license: apache-2.0
4
  base_model: distilbert-base-uncased
5
+ pipeline_tag: text-classification
6
  tags:
7
+ - sifter
8
+ - redrob
9
+ - reranker
10
+ - reward-model
11
+ - recruitment
12
+ - explainable-ai
13
+ - human-feedback
14
+ metrics:
15
+ - spearmanr
16
+ - rmse
17
+ - mae
18
  model-index:
19
+ - name: sifter-redrob-reranker
20
+ results:
21
+ - task:
22
+ type: text-classification
23
+ name: Job-candidate fit regression
24
+ dataset:
25
+ name: Redrob Challenge human-reviewed validation split
26
+ type: custom-redrob-sifter-human-feedback-data
27
+ split: validation
28
+ metrics:
29
+ - type: spearmanr
30
+ value: 0.7526
31
+ name: Spearman rank correlation
32
+ - type: rmse
33
+ value: 0.2104
34
+ name: RMSE
35
+ - type: mae
36
+ value: 0.1884
37
+ name: MAE
38
+ widget:
39
+ - text: |
40
+ Job description:
41
+ Senior AI Engineer for production retrieval, embeddings, vector search, hybrid retrieval, LLM reranking, ranking evaluation, Python, model serving, monitoring, and ownership.
42
+
43
+ Candidate profile:
44
+ Senior AI Engineer with 7.8 years experience. Skills: retrieval, ranking, evaluation, embeddings, vector search, Python, production ML systems.
45
  ---
46
 
47
+ # Sifter Redrob Reranker
 
48
 
49
+ This is the first trained reranker for **Sifter**, an AI hiring-ranking system built for the Redrob challenge.
50
 
51
+ The model reads a job description and one candidate profile together, then predicts a `0-1` fit score. In Sifter, it is used as a learned second opinion on the finalist pool after the full 100,000-candidate explainable ranker has already run.
 
 
 
 
 
52
 
53
+ Project repo: [Sifter_Redrob_Hackathon](https://github.com/shikhar1809/Sifter_Redrob_Hackathon)
54
+ Live app: [https://sifter1011.web.app](https://sifter1011.web.app)
55
 
56
+ ## What This Model Does
57
 
58
+ Sifter already has a deterministic evidence ranker that can process the full Redrob candidate pool locally. This model adds a trainable layer on top:
59
 
60
+ 1. Sifter ranks the full candidate pool using explainable evidence.
61
+ 2. The backend sends only the finalist pool to this Hugging Face model.
62
+ 3. The model returns a learned fit score.
63
+ 4. Sifter blends the scores and keeps the explanation/bias guardrails visible.
64
 
65
+ Current blend in the Sifter backend:
66
 
67
+ ```text
68
+ 70% explainable Sifter evidence score
69
+ 30% learned reranker score
70
+ ```
71
 
72
+ Default rerank scope:
73
 
74
+ ```text
75
+ top 25 finalist candidates
76
+ ```
77
 
78
+ ## Training Data
 
 
 
 
 
 
 
79
 
80
+ This revised public model was trained on Redrob-derived Sifter preference data with human-reviewed recruiter-style labels, not on a generic public ranking benchmark.
81
 
82
+ Training run:
 
 
 
 
83
 
84
+ | Item | Value |
85
+ | --- | --- |
86
+ | Source | Redrob candidate profiles + human-reviewed Sifter candidate review set |
87
+ | Total examples | 180 job-candidate examples |
88
+ | Train split | 166 examples |
89
+ | Validation split | 14 examples |
90
+ | Job description | Redrob Senior AI Engineer style role brief |
91
+ | Label type | Continuous fit score from `0.0` to `1.0` |
92
+ | Label source | Human-reviewed labels from the 180-candidate review set |
93
+ | Human label mix | 46 `strong_fit`, 58 `maybe`, 76 `not_fit` |
94
+ | Human independent holdout | Small reviewed validation split; no separate multi-recruiter panel yet |
95
 
96
+ Each training example is shaped like this:
97
 
98
+ ```text
99
+ Job description + candidate profile -> fit score
100
+ ```
101
+
102
+ The candidate profile text includes title, summary/headline, years of experience, location, career history, skills, certifications, assessments, and Redrob behavioral/logistics signals.
103
+
104
+ ## Label Scale
105
+
106
+ The revised run uses human-reviewed labels so the model learns from actual recruiter-style judgment instead of only bootstrapped scores.
107
+
108
+ | Label area | Meaning |
109
+ | --- | --- |
110
+ | `0.90 - 1.00` | Strong shortlist / interview-style fit |
111
+ | `0.55 - 0.72` | Review or maybe-fit candidates |
112
+ | `0.08 - 0.15` | Weak fit, rejected, or unranked lower-priority candidates |
113
+
114
+ Recruiter labels are supported by the training script and override weak labels when present:
115
+
116
+ | Recruiter label | Score |
117
+ | --- | --- |
118
+ | `hire` | `1.00` |
119
+ | `strong_fit` | `0.95` |
120
+ | `interview` | `0.90` |
121
+ | `review` | `0.62` |
122
+ | `maybe` | `0.55` |
123
+ | `not_fit` | `0.08` |
124
+ | `reject` | `0.00` |
125
+
126
+ Important: these labels are stronger than weak supervision, but they are still a compact review set. The next stronger version should add more reviewers and a separate held-out recruiter panel.
127
+
128
+ ## Metrics
129
+
130
+ Validation results from the human-reviewed revised run:
131
+
132
+ | Metric | Value |
133
+ | --- | ---: |
134
+ | Validation loss | `0.0443` |
135
+ | RMSE | `0.2104` |
136
+ | MAE | `0.1884` |
137
+ | Spearman rank correlation | `0.7526` |
138
+
139
+ What Spearman means in plain language: when the human-reviewed labels say candidate A should usually rank above candidate B, the model's scores mostly move in the same direction. `0.7526` is a strong sign that the learned reranker is now aligned with the reviewed candidate judgments.
140
+
141
+ ## Training Procedure
142
+
143
+ Base model:
144
+
145
+ ```text
146
+ distilbert-base-uncased
147
+ ```
148
+
149
+ Fine-tuning method:
150
+
151
+ ```text
152
+ Supervised reward-model regression fine-tuning
153
+ ```
154
+
155
+ Training setup:
156
+
157
+ | Hyperparameter | Value |
158
+ | --- | --- |
159
+ | Epochs | `3.0` |
160
+ | Training steps | Colab GPU run on 166 reviewed training rows |
161
+ | Batch size | `8` |
162
+ | Learning rate | `2e-5` |
163
+ | Max sequence length | `256` |
164
+ | Optimizer | AdamW |
165
+ | Precision | FP32 |
166
+
167
+ The model head is a single regression output (`num_labels=1`) trained with mean squared error loss.
168
+
169
+ ## Why This Is Still Human-In-The-Loop
170
+
171
+ This model is not treated as an automatic hiring decision system. The reviewed-label run improves the learned ranking signal, but Sifter still keeps human-facing checks:
172
+
173
+ - every rank still shows evidence and concern text,
174
+ - the bias guardrail stays visible,
175
+ - reviewer-agent questions challenge the result,
176
+ - recruiters can add more labels for future retraining.
177
+
178
+ ## How It Is Integrated Into Sifter
179
+
180
+ The model is wired into the Sifter backend:
181
+
182
+ | Code path | Purpose |
183
+ | --- | --- |
184
+ | `apps/api/src/learned-rerank.ts` | Calls this Hugging Face model, parses the returned score, blends it into finalist ranking, and falls back safely |
185
+ | `apps/api/src/config.ts` | Reads `HF_TOKEN`, `SIFTER_RERANKER_MODEL`, rerank weight, and finalist limit |
186
+ | `apps/api/src/server.ts` | Exposes learned reranking through the Redrob API flow |
187
+ | `apps/web/src/App.tsx` | Shows learned-reranker status in the UI |
188
+
189
+ The model is not allowed to become an unchecked black box. The deterministic Sifter reason, score breakdown, bias guardrail, and reviewer-agent questions remain visible after reranking.
190
+
191
+ ## Limitations
192
+
193
+ - The model is trained for the Redrob/Sifter Senior AI Engineer ranking setup, not general hiring across every role.
194
+ - The revised run uses 180 human-reviewed examples, so it is stronger than weak supervision but still small.
195
+ - The validation metric is measured on a 14-row reviewed validation split, not a large independent recruiter panel.
196
+ - The model can learn patterns present in the review labels, so Sifter keeps deterministic explanations and bias guardrails in the final product.
197
+ - The Redrob dataset does not include protected demographic labels, so this model card does not claim protected-class fairness parity.
198
+
199
+ ## Responsible Use
200
+
201
+ Use this model as a recruiter-assist reranker, not as an automatic hiring decision system. It should support human review by providing an additional fit signal while Sifter continues to show evidence, concerns, and bias checks.
202
+
203
+ Recommended use:
204
+
205
+ - rerank finalist pools,
206
+ - compare candidate-job fit,
207
+ - support interview shortlist review,
208
+ - collect recruiter labels for a better second version.
209
+
210
+ Not recommended:
211
+
212
+ - automatic rejection without human review,
213
+ - ranking based on identity or protected traits,
214
+ - claiming fairness parity without a protected-label audit,
215
+ - using the score without reading the explanation and evidence.