LoGoPPI Bernett
LoGoPPI is a protein–protein interaction prediction model built on ESM-2. It combines a Global sequence representation with a residue-level Maxsim score. The two calibrated scores are combined with fixed weights of 0.5 and 0.5.
Model overview
This checkpoint was trained independently on the Bernett human PPI benchmark. It does not load or continue training from the LoGoPPI Cross-species checkpoint.
Training and evaluation data
The data originate from the Bernett human PPI benchmark, Figshare version 3, which is based on high-confidence human interactions from HIPPIE v2.3 and uses leakage-controlled dataset construction.
The data/bernett/ directory contains proteins.fasta and query,text,label CSV files for training, checkpoint selection, calibration, and testing.
| Split | Pairs |
|---|---|
| Training | 163,192 |
| Validation selection | 29,630 |
| Validation calibration | 29,630 |
| Test | 52,048 |
Results
| Metric | Value |
|---|---|
| AUPR | 0.684319 |
| AUROC | 0.701470 |
| NLL | 0.632073 |
| F1 | 0.673916 |
Usage
Install LoGoPPI v2.0.0 and download the model files:
git clone https://github.com/netbiolab/LoGoPPI.git
cd LoGoPPI
conda env create -f environment.yml
conda activate logoppi
python - <<'PY'
from huggingface_hub import snapshot_download
snapshot_download(
"netbiolab/LoGoPPI-Bernett",
revision="v2.0.0",
local_dir="models/bernett",
ignore_patterns=["data/*"],
)
PY
Run inference with a FASTA file and a query,text pair CSV:
python inference.py \
--model_dir models/bernett \
--fasta_path proteins.fasta \
--pair_csv pairs.csv \
--output_path predictions.csv \
--gpus 0
Training, testing, sequence-only input, and embedding reuse are documented in the GitHub repository.
License and citation
LoGoPPI is released under the Apache License 2.0. Please cite the accompanying LoGoPPI study and the Bernett benchmark when using this model or the distributed data.
- Downloads last month
- 32