Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures Paper • 2607.28802 • Published 16 days ago • 10
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Paper • 2601.11868 • Published Jan 17 • 38
MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources Paper • 2509.25531 • Published Sep 29, 2025 • 11
garak: A Framework for Security Probing Large Language Models Paper • 2406.11036 • Published Jun 16, 2024 • 2
Semantic Consistency for Assuring Reliability of Large Language Models Paper • 2308.09138 • Published Aug 17, 2023 • 2
Representation noising effectively prevents harmful fine-tuning on LLMs Paper • 2405.14577 • Published May 23, 2024 • 1
Intrinsic Sliced Wasserstein Distances for Comparing Collections of Probability Distributions on Manifolds and Graphs Paper • 2010.15285 • Published Oct 28, 2020 • 1