Instructions to use johannhartmann/decider-2b-a11y-crawler-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use johannhartmann/decider-2b-a11y-crawler-lora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Mapika/decider-2b") model = PeftModel.from_pretrained(base_model, "johannhartmann/decider-2b-a11y-crawler-lora") - Notebooks
- Google Colab
- Kaggle
Decider-2B A11y Crawler (LoRA) — v2
A small "System 1" decision model for web crawling: given the objective, the page's accessibility tree and the previous actions, it picks the next browser action from an enumerated list (click / type / scroll / go_back / stop), decision by decision, instead of asking a large LLM at every step. Typing is a second decision that picks the text from spans of the objective. One forward pass per decision, no text generation.
LoRA adapter (r=16, all linear layers) on Mapika/decider-2b
(v10, commit d61c1c1). Built for the talk "Lokale Modelle selbst trainieren" (Mayflower, 2026).
The previous version is still available as revision v1.
What changed from v1
- All usable training decisions instead of a 2,400-decision sample: 26,969 action decisions + 2,613 typing decisions from 175 training sites (311 M tokens), one epoch; checkpoint chosen on a validation split of 160 training-site tasks.
- Objective filter: every NNetNav-live task was judged once by Claude Haiku 4.5 (objective coherent and consistent with the
recorded trajectory?). 672 of 5,017 tasks were dropped; only training data is filtered. Results, prompt, response cache and
cost ledger are public in
johannhartmann/nnetnav-live-objective-filter($6.91). - Faster forward pass, same decisions:
inference.pyno longer passes an attention mask (a single right-padded sequence in a causal model; the answer slot cannot see padding). This lets PyTorch SDPA use its flash kernel for head_dim 256: measured 1.42x faster inference and 1.80x faster training on an RTX A6000; slot logits differ only by bf16 noise. Thecausal-conv1dkernel adds another 4-6 %. - Decision procedure, prompt, candidate labels and base model are unchanged; the v1 adapter works with the new
inference.py.
Usage
# pip install "decider-ai==1.5.0" "transformers==5.17.0" "peft==0.21.0" flash-linear-attention torch
from huggingface_hub import hf_hub_download
import importlib.util
spec = importlib.util.spec_from_file_location("inference", hf_hub_download("johannhartmann/decider-2b-a11y-crawler-lora", "inference.py"))
inference = importlib.util.module_from_spec(spec); spec.loader.exec_module(inference)
dec = inference.ActionDecider.from_pretrained("johannhartmann/decider-2b-a11y-crawler-lora") # v2
# dec = inference.ActionDecider.from_pretrained("johannhartmann/decider-2b-a11y-crawler-lora", revision="v1")
probs = dec.action_probs(state_text, candidate_actions) # state_text in BrowserGym/NNetNav format, see inference.py
The input must match the training format exactly: BrowserGym's flattened accessibility tree (tab indentation,
[bid] role 'name', clickable/visible flags), then URL:, OBJECTIVE: and PREVIOUS ACTIONS: in NNetNav's
style. Candidate labels are listed in decision_contract.json. The complete crawler (BrowserGym tree extraction in
Playwright, candidate generation, episode loop, data extraction) is in the training notebook under training/.
Training
- Data:
stanfordnlp/nnetnav-live(Apache-2.0), commit7f69dca, converted to BrowserGym candidate lists (41,275 of 48,333 steps usable). Split by website: allrecipes.com, amazon.com and cambridge.org and every task touching them are held out. Deduplicated by state; no held-out tree occurs in training. - After the token budget (pages up to 24,576 tokens, untruncated; 104 longer decisions skipped) and the objective filter (5,344 decisions removed): 26,969 action decisions (click 14,198, stop 5,245, type 4,645, scroll 1,628, go_back 1,253) and 2,613 typing-value decisions.
- LoRA r=16, alpha 32, all linear layers; cross-entropy over the option slots; lr 1e-4 cosine, 5 % warmup, effective batch 8, 3,698 optimizer steps (1 epoch), bf16, gradient checkpointing, peak 11.4 GiB. Validation every 500 steps; step 3,500 selected (validation 54.3 %). 29.0 h on an RTX A6000 shared with other jobs.
continued/: the same adapter after a correction round on 10 separate live tasks (23 reviewed states, 4 corrections from recorded reference paths, 3x NNetNav replay, 3 epochs).
Evaluation
Held-out step accuracy against the demonstrated action (600 decisions on the three unseen sites, identical set for all rows; typing value n=400). All three conditions were measured with the same code in the same run:
| action | n | base | v1 | v2 (this) |
|---|---|---|---|---|
| overall | 600 | 8.5 % | 38.8 % | 42.5 % |
| click | 222 | 18.0 % | 17.1 % | 19.4 % |
| type | 176 | 4.5 % | 51.1 % | 61.9 % |
| stop | 120 | 0.8 % | 78.3 % | 71.7 % |
| scroll | 44 | 4.5 % | 6.8 % | 18.2 % |
| go_back | 38 | 0.0 % | 21.1 % | 23.7 % |
| typed text | 400 | 22.8 % | 73.3 % | 75.8 % |
v2 vs. v1: 51 decisions only v2 gets right, 29 only v1 (sign test p ≈ 0.014). (v1's card reported 39.2 % with the masked attention path; with the mask-free path it measures 38.8 %.) Secondary views: tasks kept by the objective filter (n=506) 44.3 % vs. 40.7 %; "equivalent target" (a click on an element nested in / around the demonstrated one, or the same input field with a different Enter choice) 43.2 % vs. 39.3 %. Validation (160 training-site tasks, 300 decisions): 54.3 % vs. 48.7 %.
Live crawl on 16 checkable tasks on public scraping sandboxes (books.toscrape.com, quotes.toscrape.com, webscraper.io,
scrapethissite.com), max 12 steps, data extracted by the base model after stop:
| base | v1 | v2 | v2 continued/ |
|
|---|---|---|---|---|
| target page reached | 6 | 13 | 13 | 14 |
| correct data extracted | 6 | 12 | 12 | 13 |
Reproduced after upload: loading this repository (commit ebd312b) with inference.py gives 42.5 % held-out and 75.75 % typed text, identical to the training notebook on all 600 decisions.
Limitations
- Clicks remain weak on unseen sites (19.4 %): of 222 demonstrated clicks, 65 are predicted as typing and 48 as stop. The held-out sites are search-heavy; the model prefers the search box.
- The labels are noisy: in NNetNav-live only 35 % of labelled actions equal the action that was executed next, and 3,330 stop labels sit in the middle of continuing trajectories. Accuracy against these labels understates useful behaviour and caps it.
- Measured alternatives that did not help at the 2,400-decision scale: the objective filter alone (36.5 % vs. 38.8 % for v1), a hierarchical decision (action type first, then element; 30.3 %), per-type logit offsets (+0.7 points).
- The objective filter is a single LLM judgement per task without human review; it is fairly strict (it also drops tasks whose trajectory never completes the objective).
- Typing only covers values that appear in the objective or earlier input (57.6 % of NNetNav typing steps).
- The live tasks are short (1-4 steps) sandbox tasks, not a sample of real customer sites. Pages with more than 255 candidate actions fall back to visible elements only. Respect the terms of the sites you crawl.
License and attribution
Adapter: Apache-2.0. Base model Mapika/decider-2b: Apache-2.0 (based on
Qwen/Qwen3.5-2B-Base, Apache-2.0). Training data: NNetNav-live
(Stanford NLP), Apache-2.0. Objective filter: johannhartmann/nnetnav-live-objective-filter
(Apache-2.0, verdicts generated with Claude Haiku 4.5). Tree extraction follows BrowserGym
v0.13.3 (Apache-2.0).
- Downloads last month
- 31