AMB / FAMA / WhisperBench
Existing evidence supports two evaluation targets: AMB is an open and replaceable memory backend agent memory evaluation framework; Memora is a long-term memory benchmark, FAMA is the metric that penalizes reliance on outdated or invalid memories within it.
Who it is for: Developers and evaluators who need to compare different memory backends using a unified process, researchers and practitioners who need to check original materials, queries, and standard answers, researchers studying long-term memory, evolving memory, and personalized agents
Core capabilities
- Open Evaluation Assets: AMB provides public datasets, prompts, scoring logic, and results, and allows different memory backends to be connected to the same evaluation tool.
- End-to-End Evaluation Process: AMB sequentially performs data ingestion, memory retrieval, answer generation, and answer evaluation, and separately records ingestion time and retrieval time.
- Command Line Execution and Result Browsing: AMB provides commands to list datasets, memory providers, and modes, run evaluations, and browse JSON results.
- Dataset Content Browsing: AMB's dataset browser displays original sessions and documents, queries, and standard answers.
- Dual-Mode Multi-Dimensional Evaluation: AMB supports single-query and agentic modes, and compares accuracy, latency, and token cost.
Pricing
Pricing is unknown.
Updated 2026-08-24