YOLOE-26n-Seg
This version of YOLOE-26n-Seg has been converted to run on the Axera NPU using U16 mixed-precision quantization.
Converted with Pulsar2 version: 7.0
Model Splitting
Original model repository:
git clone https://huggingface.co/openvision/yoloe26-n-seg
| Model | Description | Input | Output |
|---|---|---|---|
models/0_yoloe-26n-seg.pt |
Original PyTorch model | RGB image | Detection boxes, classes, and instance masks |
models/4_m1_cut_yoloe-26n-seg_640_pulsar2.onnx |
ONNX graph to be compiled by Pulsar2 | images, FP32 NCHW [1,3,640,640] |
bbox, class_scores, mask_coeff, proto |
models/4_m1_cut_yoloe-26n-seg_640_post.onnx |
Host post-processing graph | bbox, class_scores, mask_coeff |
output0 |
models/yoloe_26n_seg_m1.axmodel |
AX650 deployment model | RGB NHWC U8 [1,640,640,3] |
bbox, class_scores, mask_coeff, proto |
The interfaces of the split models are as follows:
Iutput of AXModel:
images [1, 3, 640, 640] FP32
Outputs of AXModel:
bbox [1, 4, 8400] FP32
class_scores [1, 80, 8400] FP32
mask_coeff [1, 32, 8400] FP32
proto [1, 32, 160, 160] FP32
Output of the Host post-processing graph:
output0 [1, 300, 38] FP32
The 38 attributes of each detection in output0 are arranged as follows:
[0:4] xyxy detection box
[4] score
[5] class_id
[6:38] 32-dimensional mask coefficient
The proto tensor and the mask coefficients in output0 are used together to decode instance masks. The Pulsar2 input graph, Host post-processing graph, and axmodel must come from the same model version and must not be mixed with other versions.
Scripts
| Script | Runtime | Description |
|---|---|---|
src/tool/export_onnx.py |
Development machine | Sets the COCO 80-class text features for YOLOE and exports a fixed 640×640 FP32 ONNX model from the PT model. |
src/tool/modify_onnx.py |
Development machine | Replaces Mod operators in the ONNX model and optionally downgrades the opset to 18. |
src/tool/e2e_cut.py |
Development machine | Extracts the ONNX graph submitted to Pulsar2 and the Host post-processing graph from the full ONNX model. |
src/infer/e2e_infer_onnx.py |
Development machine | Runs the PT or FP32 ONNX model with Ultralytics or ONNX Runtime, supporting single-image inference, dataset inference, and performance measurements. |
src/infer/e2e_infer_axmodel.py |
AX650 | Runs the axmodel with axengine and executes the Host post-processing graph with ONNX Runtime. |
src/tool/evaluate_npz.py |
Development machine | Compares single-image ONNX and axmodel NPZ outputs using cosine similarity, MSE, and maximum absolute error. |
src/tool/evaluate_yoloe.py |
Development machine | Converts inference results to COCO format, evaluates bbox/segm mAP, and compares two metrics.json files. |
The compilation configuration is pulsar2_u16_ax650.json, and the calibration dataset is datasets/coco2017val_64_calibration_dataset.tar.
The datasets/ directory contains only example images and calibration data. For official accuracy validation, download the complete COCO 2017 validation dataset.
Results
YOLOE-26n-Seg axmodel Evaluation Results
bbox
========== COCO bbox evaluation ==========
Average Precision (AP) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.264
Average Precision (AP) @[ IoU=0.50 | area= all | maxDets=100 ] = 0.394
Average Precision (AP) @[ IoU=0.75 | area= all | maxDets=100 ] = 0.280
Average Precision (AP) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.118
Average Precision (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.296
Average Precision (AP) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = 0.391
Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets= 1 ] = 0.285
Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets= 10 ] = 0.465
Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.490
Average Recall (AR) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.276
Average Recall (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.535
Average Recall (AR) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = 0.682
segm
========== COCO segm evaluation ==========
Average Precision (AP) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.222
Average Precision (AP) @[ IoU=0.50 | area= all | maxDets=100 ] = 0.372
Average Precision (AP) @[ IoU=0.75 | area= all | maxDets=100 ] = 0.229
Average Precision (AP) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.076
Average Precision (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.239
Average Precision (AP) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = 0.354
Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets= 1 ] = 0.244
Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets= 10 ] = 0.383
Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.401
Average Recall (AR) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.206
Average Recall (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.450
Average Recall (AR) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = 0.575
Images: 5000
Predictions: 790687
bbox AP@[0.5:0.95]: 0.2644
segm AP@[0.5:0.95]: 0.2219
YOLOE-26n-Seg onnx Evaluation Results
========== COCO bbox evaluation ==========
Average Precision (AP) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.280
Average Precision (AP) @[ IoU=0.50 | area= all | maxDets=100 ] = 0.415
Average Precision (AP) @[ IoU=0.75 | area= all | maxDets=100 ] = 0.296
Average Precision (AP) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.123
Average Precision (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.306
Average Precision (AP) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = 0.412
Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets= 1 ] = 0.293
Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets= 10 ] = 0.480
Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.506
Average Recall (AR) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.303
Average Recall (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.548
Average Recall (AR) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = 0.702
========== COCO segm evaluation ==========
Average Precision (AP) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.232
Average Precision (AP) @[ IoU=0.50 | area= all | maxDets=100 ] = 0.391
Average Precision (AP) @[ IoU=0.75 | area= all | maxDets=100 ] = 0.239
Average Precision (AP) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.083
Average Precision (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.250
Average Precision (AP) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = 0.367
Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets= 1 ] = 0.250
Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets= 10 ] = 0.393
Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.412
Average Recall (AR) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.221
Average Recall (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.464
Average Recall (AR) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = 0.585
Images: 5000
Predictions: 767463
bbox AP@[0.5:0.95]: 0.2796
segm AP@[0.5:0.95]: 0.2321
YOLOE-26n-Seg axmodel Inference Performance
| Model | CMM (MB) | NPU Latency (ms, Avg) | FPS | Test Conditions |
|---|---|---|---|---|
yoloe_26n_seg_m1.axmodel |
48.61 | 32.099 | 31.15 | AX650, input 640×640, batch=1, warmup=10, repeat=100 |
cmm size: 48,611,218 bytes, approximately 48.61 MB in decimal units or 46.36 MiB;- Average NPU inference latency: 32.099 ms;
- FPS calculated as
1000 / avg_ms:1000 / 32.099 = 31.15 FPS; - The
ax_run_modelresult includes onlyaxmodelinference. It excludes image reading, preprocessing, Host post-processing, NMS, mask decoding, encoding, visualization, and file I/O.
FP32 and axmodel Accuracy Comparison
Evaluation dataset:
Image count:
metric | FP32 | axmodel | drop
------------------------|------------|------------|-----
bbox AP@[0.5:0.95] | 0.279609 | 0.264382 | 0.015228
bbox AP@0.5 | 0.414795 | 0.394218 | 0.020577
bbox AP@0.75 | 0.295700 | 0.279766 | 0.015934
segm AP@[0.5:0.95] | 0.232135 | 0.221909 | 0.010226
segm AP@0.5 | 0.390976 | 0.372190 | 0.018786
segm AP@0.75 | 0.238784 | 0.228958 | 0.009826
Basic Workflow
# 1. Clone the model repository
git clone https://huggingface.co/AXERA-TECH/yoloe-26n-seg
cd yoloe-26n-seg
# 2. Prepare the full FP32 ONNX model
# Update the model paths in src/tool/export_onnx.py for the current directory layout.
python src/tool/export_onnx.py
# 3. Optionally replace Mod and downgrade the opset to 18
python src/tool/modify_onnx.py \
--input <full_FP32_ONNX> \
--output <modified_ONNX> \
--replace-mod \
--downgrade-opset 18 \
--check
# 4. Split the ONNX model
python src/tool/e2e_cut.py \
--input <modified_ONNX> \
--pulsar2-output models/4_m1_cut_yoloe-26n-seg_640_pulsar2.onnx \
--post-output models/4_m1_cut_yoloe-26n-seg_640_post.onnx \
--check
# 5. Compile the front-end ONNX graph with Pulsar2
# Update input, calibration_dataset, and output_dir in the configuration first.
pulsar2 build --config pulsar2_u16_ax650.json
# 6. Run single-image inference on the AX650
# Place the script, axmodel, post-processing ONNX, class JSON, and test image together.
python3 e2e_infer_axmodel.py \
--axmodel yoloe_26n_seg_m1.axmodel \
--post-model 4_m1_cut_yoloe-26n-seg_640_post.onnx \
--class-names yoloe-26n-seg_class_names.json \
--image test.jpg \
--output results/test.jpg
# 7. Run inference on the COCO validation dataset on the AX650
python3 e2e_infer_axmodel.py \
--axmodel yoloe_26n_seg_m1.axmodel \
--post-model 4_m1_cut_yoloe-26n-seg_640_post.onnx \
--class-names yoloe-26n-seg_class_names.json \
--image val2017 \
--output results/ax \
--eval-score-threshold 0.001 \
--full-output
# 8. Evaluate COCO bbox/segm mAP on the development machine
python src/tool/evaluate_yoloe.py \
--input results/ax/predictions.json \
--ann datasets/annotations_val2017/annotations/instances_val2017.json \
--class-names models/yoloe-26n-seg_class_names.json \
--image-dir <full_val2017_image_directory> \
--result-dir results/ax
# 9. Compare FP32 and axmodel metrics.json files
python src/tool/evaluate_yoloe.py \
--compare-metrics \
results/onnx/metrics.json \
results/ax/metrics.json
Dataset inference defaults to quick performance mode and runs image reading, preprocessing, model inference, and the Host post-processing ONNX graph. To generate predictions.json for COCO evaluation, specify --full-output.