AI & ML interests
None defined yet.
Recent Activity
View all activity
Papers
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
datasets 312
benchflow/frontierphysics-pr857-evidence
Updated • 54
benchflow/frontierphysics-pr850-evidence
Updated • 42
benchflow/frontierphysics-pr822-evidence
Updated • 54
benchflow/frontierphysics-pr823-evidence
Updated • 76
benchflow/frontierphysics-pr821-evidence
Updated • 44
benchflow/frontierphysics-pr831-evidence
Updated • 45
benchflow/frontierphysics-pr851-evidence
Updated • 47
benchflow/frontierphysics-pr852-evidence
Updated • 29
benchflow/frontierphysics-pr846-evidence
Updated • 46
benchflow/frontierphysics-pr843-evidence
Updated • 49