AGENTHUBINTEL TERMINAL

Langfuse

Langfuse is an open platform for tracking, evaluating, and improving AI agents, also described by GitHub README as an open-source LLM engineering platform that helps teams develop, monitor, evaluate, and debug AI applications.

Who it is for: Engineering teams developing and maintaining LLM applications or AI agents, teams needing to view production traces, evaluate model outputs, and manage prompts, teams requiring OpenTelemetry, integrations, or self-hosted deployment paths

Core capabilities

  • LLM Application Observability: Records LLM calls, tool calls, and retrieval steps, and supports filtering by user, session, cost, latency, or custom metadata.
  • Model Output Evaluation: Supports LLM-as-a-judge, heuristic functions, or manual reviews, and can run evaluators during production data or experimental periods.
  • Prompt Management: Supports separating prompts from code and provides one-click deployment and rollback.
  • Playground and Experimentation: Playground supports testing prompts with real production inputs and side-by-side model comparisons; the experiment module supports defining test cases, running experiments, and comparing results.
  • Manual Annotation and Datasets: Supports collaborative manual trace review and creation of golden datasets.

Pricing

The free tier offers 50,000 observations per month without a credit card; full paid plans, billing units, enterprise pricing, support, and compliance add-ons are unknown.

Updated 2026-08-18

Visit Langfuse