PEFT
Safetensors
English
lora
web-agents
browser-automation
accessibility-tree
crawling
decider

Decider-2B A11y Crawler (LoRA) — v2

A small "System 1" decision model for web crawling: given the objective, the page's accessibility tree and the previous actions, it picks the next browser action from an enumerated list (click / type / scroll / go_back / stop), decision by decision, instead of asking a large LLM at every step. Typing is a second decision that picks the text from spans of the objective. One forward pass per decision, no text generation.

LoRA adapter (r=16, all linear layers) on Mapika/decider-2b (v10, commit d61c1c1). Built for the talk "Lokale Modelle selbst trainieren" (Mayflower, 2026). The previous version is still available as revision v1.

What changed from v1

  • All usable training decisions instead of a 2,400-decision sample: 26,969 action decisions + 2,613 typing decisions from 175 training sites (311 M tokens), one epoch; checkpoint chosen on a validation split of 160 training-site tasks.
  • Objective filter: every NNetNav-live task was judged once by Claude Haiku 4.5 (objective coherent and consistent with the recorded trajectory?). 672 of 5,017 tasks were dropped; only training data is filtered. Results, prompt, response cache and cost ledger are public in johannhartmann/nnetnav-live-objective-filter ($6.91).
  • Faster forward pass, same decisions: inference.py no longer passes an attention mask (a single right-padded sequence in a causal model; the answer slot cannot see padding). This lets PyTorch SDPA use its flash kernel for head_dim 256: measured 1.42x faster inference and 1.80x faster training on an RTX A6000; slot logits differ only by bf16 noise. The causal-conv1d kernel adds another 4-6 %.
  • Decision procedure, prompt, candidate labels and base model are unchanged; the v1 adapter works with the new inference.py.

Usage

# pip install "decider-ai==1.5.0" "transformers==5.17.0" "peft==0.21.0" flash-linear-attention torch
from huggingface_hub import hf_hub_download
import importlib.util
spec = importlib.util.spec_from_file_location("inference", hf_hub_download("johannhartmann/decider-2b-a11y-crawler-lora", "inference.py"))
inference = importlib.util.module_from_spec(spec); spec.loader.exec_module(inference)

dec = inference.ActionDecider.from_pretrained("johannhartmann/decider-2b-a11y-crawler-lora")          # v2
# dec = inference.ActionDecider.from_pretrained("johannhartmann/decider-2b-a11y-crawler-lora", revision="v1")
probs = dec.action_probs(state_text, candidate_actions)   # state_text in BrowserGym/NNetNav format, see inference.py

The input must match the training format exactly: BrowserGym's flattened accessibility tree (tab indentation, [bid] role 'name', clickable/visible flags), then URL:, OBJECTIVE: and PREVIOUS ACTIONS: in NNetNav's style. Candidate labels are listed in decision_contract.json. The complete crawler (BrowserGym tree extraction in Playwright, candidate generation, episode loop, data extraction) is in the training notebook under training/.

Training

  • Data: stanfordnlp/nnetnav-live (Apache-2.0), commit 7f69dca, converted to BrowserGym candidate lists (41,275 of 48,333 steps usable). Split by website: allrecipes.com, amazon.com and cambridge.org and every task touching them are held out. Deduplicated by state; no held-out tree occurs in training.
  • After the token budget (pages up to 24,576 tokens, untruncated; 104 longer decisions skipped) and the objective filter (5,344 decisions removed): 26,969 action decisions (click 14,198, stop 5,245, type 4,645, scroll 1,628, go_back 1,253) and 2,613 typing-value decisions.
  • LoRA r=16, alpha 32, all linear layers; cross-entropy over the option slots; lr 1e-4 cosine, 5 % warmup, effective batch 8, 3,698 optimizer steps (1 epoch), bf16, gradient checkpointing, peak 11.4 GiB. Validation every 500 steps; step 3,500 selected (validation 54.3 %). 29.0 h on an RTX A6000 shared with other jobs.
  • continued/: the same adapter after a correction round on 10 separate live tasks (23 reviewed states, 4 corrections from recorded reference paths, 3x NNetNav replay, 3 epochs).

Evaluation

Held-out step accuracy against the demonstrated action (600 decisions on the three unseen sites, identical set for all rows; typing value n=400). All three conditions were measured with the same code in the same run:

action n base v1 v2 (this)
overall 600 8.5 % 38.8 % 42.5 %
click 222 18.0 % 17.1 % 19.4 %
type 176 4.5 % 51.1 % 61.9 %
stop 120 0.8 % 78.3 % 71.7 %
scroll 44 4.5 % 6.8 % 18.2 %
go_back 38 0.0 % 21.1 % 23.7 %
typed text 400 22.8 % 73.3 % 75.8 %

v2 vs. v1: 51 decisions only v2 gets right, 29 only v1 (sign test p ≈ 0.014). (v1's card reported 39.2 % with the masked attention path; with the mask-free path it measures 38.8 %.) Secondary views: tasks kept by the objective filter (n=506) 44.3 % vs. 40.7 %; "equivalent target" (a click on an element nested in / around the demonstrated one, or the same input field with a different Enter choice) 43.2 % vs. 39.3 %. Validation (160 training-site tasks, 300 decisions): 54.3 % vs. 48.7 %.

Live crawl on 16 checkable tasks on public scraping sandboxes (books.toscrape.com, quotes.toscrape.com, webscraper.io, scrapethissite.com), max 12 steps, data extracted by the base model after stop:

base v1 v2 v2 continued/
target page reached 6 13 13 14
correct data extracted 6 12 12 13

Reproduced after upload: loading this repository (commit ebd312b) with inference.py gives 42.5 % held-out and 75.75 % typed text, identical to the training notebook on all 600 decisions.

Limitations

  • Clicks remain weak on unseen sites (19.4 %): of 222 demonstrated clicks, 65 are predicted as typing and 48 as stop. The held-out sites are search-heavy; the model prefers the search box.
  • The labels are noisy: in NNetNav-live only 35 % of labelled actions equal the action that was executed next, and 3,330 stop labels sit in the middle of continuing trajectories. Accuracy against these labels understates useful behaviour and caps it.
  • Measured alternatives that did not help at the 2,400-decision scale: the objective filter alone (36.5 % vs. 38.8 % for v1), a hierarchical decision (action type first, then element; 30.3 %), per-type logit offsets (+0.7 points).
  • The objective filter is a single LLM judgement per task without human review; it is fairly strict (it also drops tasks whose trajectory never completes the objective).
  • Typing only covers values that appear in the objective or earlier input (57.6 % of NNetNav typing steps).
  • The live tasks are short (1-4 steps) sandbox tasks, not a sample of real customer sites. Pages with more than 255 candidate actions fall back to visible elements only. Respect the terms of the sites you crawl.

License and attribution

Adapter: Apache-2.0. Base model Mapika/decider-2b: Apache-2.0 (based on Qwen/Qwen3.5-2B-Base, Apache-2.0). Training data: NNetNav-live (Stanford NLP), Apache-2.0. Objective filter: johannhartmann/nnetnav-live-objective-filter (Apache-2.0, verdicts generated with Claude Haiku 4.5). Tree extraction follows BrowserGym v0.13.3 (Apache-2.0).

Downloads last month
31
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for johannhartmann/decider-2b-a11y-crawler-lora

Adapter
(2)
this model

Datasets used to train johannhartmann/decider-2b-a11y-crawler-lora