longhorizon-harness
amap-ml
530PythonframeworkApache-2.0
View on GitHub LongHorizon-Harness is a standardized benchmarking framework designed to evaluate and compare the performance of AI agents on long-horizon tasks. It provides a structured environment for testing agent capabilities in complex, multi-step scenarios. The project is aimed at researchers and developers working on agentic AI systems who need reliable metrics for evaluation.
Install
pip install -e .