XYZAILab commited on
Commit
6d96321
·
verified ·
1 Parent(s): d83c323

update model card

Browse files
README.md CHANGED
@@ -1 +1,141 @@
 
 
 
 
 
 
 
 
 
1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qwen/Qwen3.5-397B-A17B
3
+ pipeline_tag: text-generation
4
+ library_name: transformers
5
+ tags:
6
+ - safetensors
7
+ - qwen3.6
8
+ - agentic-search
9
+ ---
10
 
11
+ <div align="center">
12
+ <a href="https://xyz-lab.ai/">
13
+ <img src="./assets/xyz-ai-lab-slogan.svg" width="520" alt="XYZ AI Lab — We Build The Minds That Build" />
14
+ </a>
15
+ </div>
16
+
17
+ <h1 align="center">XYZ-Aquila-pro</h1>
18
+
19
+ <p align="center"><strong>An open-weight thinking model for Deep Search.</strong></p>
20
+
21
+ <p align="center">
22
+ <a href="https://xyz-lab.ai/"><img alt="Homepage" src="https://img.shields.io/badge/Homepage-XYZ%20AI%20Lab-111827?style=flat-square&logo=googlechrome&logoColor=white"></a>
23
+ <a href="https://xyz-lab.ai/demo/"><img alt="Demo AI4AI" src="https://img.shields.io/badge/Demo-AI4AI-0f766e?style=flat-square&logo=googlechrome&logoColor=white"></a>
24
+ <a href="https://xyz-lab.ai/try-it-out/"><img alt="Demo Search Agent" src="https://img.shields.io/badge/Demo-Search%20Agent-2563eb?style=flat-square&logo=googlechrome&logoColor=white"></a>
25
+ </p>
26
+
27
+ <p align="center">
28
+ <a href="https://github.com/XYZ-AI-Lab"><img alt="GitHub" src="https://img.shields.io/badge/GitHub-XYZ%20AI%20Lab-111827?style=flat-square&logo=github&logoColor=white"></a>
29
+ <a href="https://huggingface.co/datasets/XYZAILab/XYZ-Aquila-SFT"><img alt="Data" src="https://img.shields.io/badge/Data-XYZ--Aquila--SFT-7c3aed?style=flat-square&logo=huggingface&logoColor=white"></a>
30
+ <a href="https://xyz-lab.ai/blogs/ai4ai-at-scale/assets/bounded-exploration-ai4ai-system-optimization.pdf"><img alt="Technical Report" src="https://img.shields.io/badge/Technical%20Report-PDF-b45309?style=flat-square&logo=readthedocs&logoColor=white"></a>
31
+ </p>
32
+
33
+ ## Introduction
34
+
35
+ **XYZ-Aquila** is a family of open-weight Deep Search agents developed by [XYZ AI Lab](https://xyz-lab.ai/). XYZ-Aquila-pro is post-trained from [Qwen3.5-397B-A17B](https://huggingface.co/Qwen/Qwen3.5-397B-A17B) through a bounded-exploration **AI4AI** pipeline: humans define the target capability, development evidence, constraints, risk boundaries, and acceptance policy, while AI agents diagnose failures and propose scoped interventions across data, post-training, runtime, context management, tools, evaluation, and infrastructure.
36
+
37
+ The released checkpoint is a **thinking model** with Qwen-compatible reasoning and tool-call formats. It is optimized for agentic search, including long-horizon planning, English and Chinese web browsing, multi-source evidence aggregation, source verification, and recovery from failed environment interactions. The open-source [AxisAgentic harness](https://github.com/XYZ-AI-Lab/AxisAgentic) provides the concrete `search` / `scrape` / `python` tool implementations, fixed tool contract, replayable context management, and benchmark evaluation workflow; these capabilities are supplied by the surrounding harness rather than by the checkpoint alone.
38
+
39
+ ## Benchmark Results
40
+
41
+ The external benchmark suite was held out from routine AI4AI optimization. Following the technical report, evaluation uses a ReAct-style search harness with web search, webpage extraction, stateful Python execution, and a maximum 256K context. XYZ-Aquila-pro obtains the highest reported score in every column of the evaluated sub-400B open-weight comparison.
42
+
43
+ <div align="center">
44
+ <img src="./assets/benchmark_results_candidate_2.svg" width="100%" alt="XYZ-Aquila benchmark results across six agentic search benchmarks" />
45
+ </div>
46
+
47
+ The figure provides a visual overview across six agentic search benchmarks. The tables below transpose the comparison: each row is a benchmark and each column is a model within the group.
48
+
49
+ ### Small-scale open-weight (<40B)
50
+
51
+ Open-weight systems with fewer than 40B parameters, including XYZ-Aquila-mini.
52
+
53
+ | Benchmark | XYZ-Aquila-mini | Agents-A1 | Nex-N2-mini | apodex-mini | MiroThinker 1.7 mini |
54
+ |:--|--:|--:|--:|--:|--:|
55
+ | BrowseComp | **78.8** | 75.5 | 74.1 | 71.5 | 67.9 |
56
+ | BrowseComp-ZH | **82.9** | -- | 79.6† | 80.6 | -- |
57
+ | DeepSearchQA | **89.5** | -- | 87.2† | 82.2 | -- |
58
+ | GAIA | **97.1** | 96.0 | -- | -- | 80.3 |
59
+ | LiveBrowseComp | **48.7** | 29.6† | 41.4† | 32.8† | 34.9† |
60
+ | HLE | **51.1** | 47.6 | 37.1† | 46.8 | 36.4 |
61
+ | WideSearch | **80.8** | -- | 62.0 | -- | 73.3† |
62
+
63
+ ### Large-scale open-weight (<400B)
64
+
65
+ Open-weight systems with fewer than 400B parameters, including XYZ-Aquila-pro.
66
+
67
+ | Benchmark | XYZ-Aquila-pro | Nex-N2-Pro | MiroThinker 1.7 | apodex-1.0 |
68
+ |:--|--:|--:|--:|--:|
69
+ | BrowseComp | **84.8** | 83.7† | 74.0 | 75.5 |
70
+ | BrowseComp-ZH | **85.1** | 79.6† | 75.3 | 82.6 |
71
+ | DeepSearchQA | **92.5** | 92.3† | -- | 84.6 |
72
+ | LiveBrowseComp | **53.7** | 50.4† | 34.1† | -- |
73
+ | HLE | **53.3** | 50.0† | 42.9 | 49.0 |
74
+ | WideSearch | **81.2** | 75.6 | -- | -- |
75
+
76
+ ### XYZ-Aquila-pro vs. Larger-Scale and Closed-Source Models
77
+
78
+ XYZ-Aquila-pro compared with larger-scale open-weight and closed-source models.
79
+
80
+ | Benchmark | XYZ-Aquila-pro | apodex-h1 | DeepSeek-V4-<br>Pro-Max | Kimi-K2.6 | Claude<br>Opus 4.7 | GPT-5.5<br>xhigh |
81
+ |:--|--:|--:|--:|--:|--:|--:|
82
+ | BrowseComp | 84.8 | **90.3** | 83.4 | 83.2 | 79.3 | 84.4 |
83
+ | BrowseComp-ZH | **85.1** | 84.1 | -- | -- | -- | -- |
84
+ | DeepSearchQA | 92.5 | **94.4** | -- | 92.5 | 89.1 | -- |
85
+ | LiveBrowseComp | **53.7** | -- | 38.3 | 31.7 | -- | -- |
86
+ | HLE | 53.3 | **60.8** | -- | 55.5 | 54.7 | 52.2 |
87
+ | WideSearch | **81.2** | -- | -- | 80.8 | -- | -- |
88
+
89
+ All values are percentages. HLE denotes Humanity's Last Exam. DeepSearchQA uses F1, WideSearch uses Item F1 Max@4, and the remaining benchmarks use accuracy. Rows with no reported result across an entire group are omitted; `--` indicates an unreported result within an otherwise populated row. `†` marks results reproduced under the common evaluation setup; other baseline values come from public reports or benchmark submissions. See the [technical report](https://xyz-lab.ai/blogs/ai4ai-at-scale/assets/bounded-exploration-ai4ai-system-optimization.pdf) for full provenance and analysis.
90
+
91
+ ## Quickstart
92
+
93
+ ### SGLang Deployment
94
+
95
+ Use a recent SGLang release (`sglang>=0.5.10`). The example below launches an OpenAI-compatible endpoint with Qwen reasoning and tool-call parsers. It uses tensor parallelism across eight GPUs; reduce the context length if the deployment does not have enough memory.
96
+
97
+ ```bash
98
+ uv pip install "sglang[all]>=0.5.10"
99
+
100
+ MODEL_PATH=XYZAILab/XYZ-Aquila-pro
101
+ SERVED_MODEL=XYZ-Aquila-pro
102
+
103
+ python -m sglang.launch_server \
104
+ --model-path "${MODEL_PATH}" \
105
+ --served-model-name "${SERVED_MODEL}" \
106
+ --port 8000 \
107
+ --tp-size 8 \
108
+ --mem-fraction-static 0.8 \
109
+ --context-length 262144 \
110
+ --reasoning-parser qwen3 \
111
+ --tool-call-parser qwen3_coder
112
+ ```
113
+
114
+ For `XYZ-Aquila-mini`, replace `MODEL_PATH` and `SERVED_MODEL` with the mini repository and use the tensor-parallel configuration appropriate for your hardware.
115
+
116
+ ### Recommended Sampling Config
117
+
118
+ These are recommended starting values for thinking-mode agentic search. Tune them for the target task and harness.
119
+
120
+ ```yaml
121
+ temperature: 1.0
122
+ top_p: 0.95
123
+ repetition_penalty: 1.05
124
+ chat_template_kwargs:
125
+ enable_thinking: true
126
+ preserve_thinking: true
127
+ ```
128
+
129
+ Keep the input and generated response within the 262,144-token context window.
130
+
131
+ ## Citation
132
+
133
+ ```bibtex
134
+ @techreport{xyz_aquila_2026,
135
+ title = {AI4AI at Scale: A Full-Pipeline System for Enhancing LLM Agentic Capabilities},
136
+ author = {{XYZ Agentic Team}},
137
+ institution = {XYZ AI Lab},
138
+ year = {2026},
139
+ url = {https://xyz-lab.ai/blogs/ai4ai-at-scale/assets/bounded-exploration-ai4ai-system-optimization.pdf}
140
+ }
141
+ ```
assets/benchmark_results_candidate_2.svg ADDED
assets/xyz-ai-lab-slogan.svg ADDED