Skip to main content

longhorizon-harness

amap-ml

530PythonframeworkApache-2.0
View on GitHub

LongHorizon-Harness is a standardized benchmarking framework designed to evaluate and compare the performance of AI agents on long-horizon tasks. It provides a structured environment for testing agent capabilities in complex, multi-step scenarios. The project is aimed at researchers and developers working on agentic AI systems who need reliable metrics for evaluation.

Install

pip install -e .