Agent Safety
-
AG→
ConstraintBench
by oriolmirolf
It evaluates LLM-based agents across 50 PDDL planning tasks using the VAL 4.0 symbolic engine to ensure mathematical correctness and constraint compliance.
-
AG→
Agentsz
by Juanalbertw
We implemented a minimal prompt-ablation version of the Pi-Bench purple server, keeping the reference A2A/LiteLLM scaffold intact while adding env-var-gated prompt suffixes. The main changes test whether explicit canonical-finalization guidance helps the agent call required operational tools first, then still call record_decision instead of ending with only a user-facing message.
-
AG→
caum-agentbeats-purple
by caum-systems
A2A Purple Agent wrapped with CAUM structural observation. Includes benchmark-only control mode to study whether structural loop/stall signals improve agent behavior without exposing private task content.
-
AG→
STRIDE Pi-Bench Agent
by chaeritas
STRIDE XAI-optimized Purple Agent for Pi-Bench policy compliance. By Chaestro Inc.