Announcing Artificial Analysis Intelligence Index v4.3, upgrading Terminal-Bench to v4 and adding AutomationBench-AA, an agentic workflow automation benchmark with a private test set. This is a continuation of our rollout of Intelligence Index v5 Changelog (Index v4.2 → Index v4.3): → Terminal-Bench: v2.1 → v4, completing our upgrade to the latest version of Terminal-Bench → Replacing 𝜏³-Banking with AutomationBench-AA, our implementation of Zapier’s business workflow automation benchmark We are continuing to prioritize keeping Intelligence Index as useful as possible by bringing forward a subset of the changes we had planned for Index v5. Each change in v4.2 and v4.3 stands on its own merits and brings the index closer to real-world problem solving, adds more private test sets to prevent gaming, and reduces saturation. Intelligence Index v4.3 raises the difficulty of agentic coding tasks and broadens the types of agentic workflows tested. Because we use a held-out test set for AutomationBench in collaboration with Zapier , the weight assigned to evaluations with private tasks or answers increases from 40% to 45%. Category weights are unchanged from v4.2: Agents 30%, Coding 20%, General 30%, Scientific Reasoning 20%. Detailed changes: ➤ Upgraded Terminal-Bench v2.1 to v4.0: 66 multi-step tasks testing agents on
Read More












