|
Download docs/architecture.md from ashutosh111/negotiation-arena-master: direct link, hf CLI and curl.
- Browser
- Download file 2.92 kB
-
https://huggingface.co/spaces/ashutosh111/negotiation-arena-master/resolve/main/docs/architecture.md
- Command line
-
hf download hf://spaces/ashutosh111/negotiation-arena-master/docs/architecture.md
-
curl -L -o architecture.md https://huggingface.co/spaces/ashutosh111/negotiation-arena-master/resolve/main/docs/architecture.md
2.92 kB
Architecture Overview
Placeholder. Fill in once V3 implementation lands.
V2 (current, deployed)
┌─────────────────────────────────────┐
│ inference.py │
│ Llama-3.3-70B Vendor agent │
└─────────────────┬───────────────────┘
│ POST /reset, /step
▼
┌─────────────────────────────────────┐
│ FastAPI app (server/app.py) │
│ /health /reset /step /state │
└─────────────────┬───────────────────┘
▼
┌──────────────────────────────────────────┐
│ NegotiationArenaEnvironment │
└─────────────────┬────────────────────────┘
│
┌───────────┴────────────┐
▼ ▼
┌──────────────────┐ ┌────────────────────┐
│ ScriptedClient │ │ LLMClient │
│ (training) │ │ (eval only) │
└──────────────────┘ └────────────────────┘
│
▼
server/arena.py — turn manager
server/graders.py — composite reward
server/utility.py — per-role scoring
server/private_briefs — hidden priorities
server/contract_fixtures — 3 tasks
V3 (planned additions)
server/tom.py— theory-of-mind module (vendor predicts opponent's next action; correctness shapes reward)server/tribunal.py— 3-judge ensemble grader; reward = median, anti-collusion penalty on disagreementserver/curriculum.py— auto-difficulty scheduler driven by recent win rateaudit/— adversarial robustness suite (5 exploit agents probing the trained policy)evaluation/— generalization tests + 5 diagnostic plotsweb/landing/+web/replay_ui/— public landing page and JSON-driven episode replayer
To be illustrated with a proper diagram (assets/architecture_diagram.png) once
V3 is implemented.